Query 040351
Match_columns 232
No_of_seqs 111 out of 276
Neff 5.5
Searched_HMMs 46136
Date Fri Mar 29 07:36:31 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/040351.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/040351hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 smart00452 STI Soybean trypsin 100.0 1.9E-57 4.1E-62 384.1 20.5 169 30-222 1-172 (172)
2 cd00178 STI Soybean trypsin in 100.0 1.4E-56 3E-61 378.8 19.6 167 29-220 1-172 (172)
3 PF00197 Kunitz_legume: Trypsi 100.0 2.7E-56 5.8E-61 377.9 18.2 169 29-220 1-176 (176)
4 PF07951 Toxin_R_bind_C: Clost 76.4 7.5 0.00016 34.4 6.0 120 27-155 7-146 (214)
5 KOG3858 Ephrin, ligand for Eph 53.5 13 0.00029 33.4 3.0 55 34-90 117-177 (233)
6 PF02402 Lysis_col: Lysis prot 45.4 18 0.00039 24.5 2.0 19 24-42 18-36 (46)
7 PF07172 GRP: Glycine rich pro 39.6 18 0.0004 27.9 1.5 8 2-9 4-11 (95)
8 PF14009 DUF4228: Domain of un 37.1 21 0.00046 28.8 1.6 21 33-53 64-84 (181)
9 COG5341 Uncharacterized protei 36.0 2.5E+02 0.0054 23.1 8.1 18 93-110 103-120 (132)
10 PF05474 Semenogelin: Semenoge 32.3 23 0.00049 35.3 1.2 15 1-15 1-15 (582)
11 PF08194 DIM: DIM protein; In 32.1 62 0.0013 20.9 2.8 16 1-17 1-16 (36)
12 PF05550 Peptidase_C53: Pestiv 30.3 35 0.00076 28.8 1.8 25 17-41 11-35 (168)
13 PF00879 Defensin_propep: Defe 29.9 54 0.0012 22.9 2.4 18 1-19 1-18 (52)
14 PF09466 Yqai: Hypothetical pr 29.7 41 0.00088 24.9 1.9 21 28-48 24-44 (71)
15 PRK11354 kil FtsZ inhibitor pr 27.9 41 0.00089 25.0 1.7 47 35-87 19-69 (73)
16 COG4851 CamS Protein involved 27.2 63 0.0014 30.6 3.1 59 1-61 1-71 (382)
17 PRK10220 hypothetical protein; 26.5 79 0.0017 25.4 3.2 27 28-54 42-68 (111)
18 TIGR00686 phnA alkylphosphonat 25.7 85 0.0018 25.1 3.2 27 28-54 41-67 (109)
19 TIGR02052 MerP mercuric transp 24.9 47 0.001 22.8 1.5 14 1-17 1-14 (92)
20 COG5510 Predicted small secret 23.0 64 0.0014 21.8 1.7 17 1-17 2-18 (44)
21 COG2824 PhnA Uncharacterized Z 22.8 1E+02 0.0022 24.7 3.2 27 28-54 43-69 (112)
22 PRK10159 outer membrane phosph 22.6 51 0.0011 30.7 1.7 36 1-42 1-36 (351)
23 PF00812 Ephrin: Ephrin; Inte 22.4 55 0.0012 27.2 1.6 21 34-54 101-121 (145)
24 PLN03207 stomagen; Provisional 21.9 43 0.00094 26.5 0.9 8 8-15 14-21 (113)
25 PF03831 PhnA: PhnA protein; 20.9 52 0.0011 23.3 1.0 21 30-50 2-22 (56)
No 1
>smart00452 STI Soybean trypsin inhibitor (Kunitz) family of protease inhibitors.
Probab=100.00 E-value=1.9e-57 Score=384.13 Aligned_cols=169 Identities=30% Similarity=0.556 Sum_probs=152.4
Q ss_pred eecCCCCcccCCCCEEEEecccCCCCCCceEEecCCCCCCCCceEEcCCCCCCCceEEEecCC-CCcceecccceEEEEe
Q 040351 30 ILDVYGNQVDSSHRYYLVSALWGVKTGGGISADKGKNGQCPTDVIQLSPKDKRGKNLVLLPND-NSIIVRESTNIKLKFS 108 (232)
Q Consensus 30 VlDtdG~~L~~G~~YYIlPa~~g~G~gGGl~l~~t~n~~CPl~VvQ~~~~~~~GlPV~Fsp~~-~~~vI~e~t~lnI~F~ 108 (232)
|+|+|||||++|++|||+|++|+.| |||++++++|++||++|+|++++..+|+||+|+|++ ++.+|+|+|+|||+|.
T Consensus 1 VlDt~G~~l~~G~~YyI~p~~~g~G--GGl~l~~~~n~~CPl~VvQ~~~~~~~GlPV~Fs~~~~~~~ii~e~t~lnI~F~ 78 (172)
T smart00452 1 VLDTDGNPLRNGGTYYILPAIRGHG--GGLTLAATGNEICPLTVVQSPNEVDNGLPVKFSPPNPSDFIIRESTDLNIEFD 78 (172)
T ss_pred CCCCCCCCCcCCCcEEEEEccccCC--CCEEEccCCCCCCCCeeEECCCCCCCceeEEEeecCCCCCEEecCceEEEEeC
Confidence 7999999999999999999999976 899999999999999999999999999999999976 7889999999999999
Q ss_pred cCCCCCCcCCCCcEEEecCCCcCcceEEEeCC-CCCCcCCeEEEEcccccchhccccCCCCCeEEEeccCCCCCCCCccc
Q 040351 109 RVSSLQQCNKDSLWKVDNDNASLGKQFITIGE-GKTCQNLFKLEKVSASIFDMKIALDIPCLYKIVHCSTPVNGSCDTLC 187 (232)
Q Consensus 109 ~~~~~~~C~~st~W~V~~~d~~~g~~~V~tGg-~~~~~~~FkIeK~~~~~~~~~~~~~~~~~YKLvfCp~~~~~~c~~~C 187 (232)
.. +.|++||+|+|++ ++..++|||+||| .+...|||||||++.. .+.|||+|||..|+ ...|
T Consensus 79 ~~---~~C~~st~W~V~~-~~~~~~~~V~~gg~~~~~~~~FkIek~~~~----------~~~YKLv~Cp~~~~---~~~C 141 (172)
T smart00452 79 AP---PLCAQSTVWTVDE-DSAPEGLAVKTGGYPGVRDSWFKIEKYSGE----------SNGYKLVYCPNGSD---DDKC 141 (172)
T ss_pred CC---CCCCCCCEEEEec-CCccccEEEEeCCcCCCCCCeEEEEECCCC----------CCCEEEEEcCCCCC---CCcc
Confidence 87 7899999999996 5677899999999 5556799999999842 25799999998764 5789
Q ss_pred ceeeEEee-CCeeEEEEEcCCCCCCCCeeEEEEEcC
Q 040351 188 KDVGVSNV-DGVQRLVVVDDNDQPNLPLPVVLFPAD 222 (232)
Q Consensus 188 ~dVGi~~d-~g~rrL~l~~d~~~~~~pf~V~F~ka~ 222 (232)
+|||+++| +|+||||| +++ +||.|+|+|++
T Consensus 142 ~~vGi~~d~~g~rrL~l-s~~----~p~~v~F~k~~ 172 (172)
T smart00452 142 GDVGIFIDPEGGRRLVL-SNE----NPLVVVFKKAE 172 (172)
T ss_pred CccCeEECCCCcEEEEE-cCC----CCeEEEEEECC
Confidence 99999987 89999999 764 39999999985
No 2
>cd00178 STI Soybean trypsin inhibitor (Kunitz) family of protease inhibitors. Inhibit proteases by binding with high affinity to their active sites. Trefoil fold, common to interleukins and fibroblast growth factors.
Probab=100.00 E-value=1.4e-56 Score=378.78 Aligned_cols=167 Identities=37% Similarity=0.661 Sum_probs=150.3
Q ss_pred ceecCCCCcccCCCCEEEEecccCCCCCCceEEecCCCCCCCCceEEcCCCCCCCceEEEecCC-CCcceecccceEEEE
Q 040351 29 PILDVYGNQVDSSHRYYLVSALWGVKTGGGISADKGKNGQCPTDVIQLSPKDKRGKNLVLLPND-NSIIVRESTNIKLKF 107 (232)
Q Consensus 29 ~VlDtdG~~L~~G~~YYIlPa~~g~G~gGGl~l~~t~n~~CPl~VvQ~~~~~~~GlPV~Fsp~~-~~~vI~e~t~lnI~F 107 (232)
||||+|||||++|.+|||+|++||.| |||++++++|++||++|+|++++..+|+||+|+|++ ++.+|||+++|||+|
T Consensus 1 ~VlD~~G~~l~~g~~YyI~p~~~g~G--GGl~l~~~~~~~CPl~VvQ~~~~~~~GlPv~Fs~~~~~~~~I~e~t~lnI~F 78 (172)
T cd00178 1 PVLDTDGNPLRNGGRYYILPAIRGGG--GGLTLAATGNETCPLTVVQSPSELDRGLPVKFSPPNPKSDVIRESTDLNIEF 78 (172)
T ss_pred CcCcCCCCCCcCCCeEEEEEceeCCC--CcEEEcCCCCCCCCCeeEECCCCCCCCeeEEEEeCCCCCCEEECCCcEEEEe
Confidence 69999999999999999999999976 999999999999999999999999999999999987 899999999999999
Q ss_pred ecCCCCCCc-CCCCcEEEecCCCcCcceEEEeCCCC--CCcCCeEEEEcccccchhccccCCCCCeEEEeccCCCCCCCC
Q 040351 108 SRVSSLQQC-NKDSLWKVDNDNASLGKQFITIGEGK--TCQNLFKLEKVSASIFDMKIALDIPCLYKIVHCSTPVNGSCD 184 (232)
Q Consensus 108 ~~~~~~~~C-~~st~W~V~~~d~~~g~~~V~tGg~~--~~~~~FkIeK~~~~~~~~~~~~~~~~~YKLvfCp~~~~~~c~ 184 (232)
... +.| ++|++|+|+++++ .++|||+|||.. +.++||||||++.. .+.|||+|||+.| +
T Consensus 79 ~~~---~~c~~~st~W~V~~~~~-~~~~~V~~Gg~~~~~~~~~FkIek~~~~----------~~~YKL~~Cp~~~----~ 140 (172)
T cd00178 79 DAP---TWCCGSSTVWKVDRDST-PEGLFVTTGGVKGNTLNSWFKIEKVSEG----------LNAYKLVFCPSSC----D 140 (172)
T ss_pred CCC---CcCCCCCCEEEEeccCC-ccCeEEEeCCcCCCcccceEEEEECCCC----------CCcEEEEEcCCCC----C
Confidence 987 566 9999999997665 789999999943 37999999999852 2679999999864 5
Q ss_pred cccceeeEEee-CCeeEEEEEcCCCCCCCCeeEEEEE
Q 040351 185 TLCKDVGVSNV-DGVQRLVVVDDNDQPNLPLPVVLFP 220 (232)
Q Consensus 185 ~~C~dVGi~~d-~g~rrL~l~~d~~~~~~pf~V~F~k 220 (232)
..|+|||+++| +|.||||| +++ +||.|+|+|
T Consensus 141 ~~C~~VGi~~d~~g~rrL~l-~~~----~p~~V~F~k 172 (172)
T cd00178 141 SKCGDVGIFIDPEGVRRLVL-SDD----NPLVVVFKK 172 (172)
T ss_pred CceeecccEECCCCcEEEEE-cCC----CCeEEEEeC
Confidence 68999999987 79999999 764 399999997
No 3
>PF00197 Kunitz_legume: Trypsin and protease inhibitor; InterPro: IPR002160 Peptide proteinase inhibitors can be found as single domain proteins or as single or multiple domains within proteins; these are referred to as either simple or compound inhibitors, respectively. In many cases they are synthesised as part of a larger precursor protein, either as a prepropeptide or as an N-terminal domain associated with an inactive peptidase or zymogen. This domain prevents access of the substrate to the active site. Removal of the N-terminal inhibitor domain either by interaction with a second peptidase or by autocatalytic cleavage activates the zymogen. Other inhibitors interact direct with proteinases using a simple noncovalent lock and key mechanism; while yet others use a conformational change-based trapping mechanism that depends on their structural and thermodynamic properties. The Kunitz-type soybean trypsin inhibitor (STI) family consists mainly of proteinase inhibitors from Leguminosae seeds []. They belong to MEROPS inhibitor family I3, clan IC. They exhibit proteinase inhibitory activity against serine proteinases; trypsin (MEROPS peptidase family S1, IPR001254 from INTERPRO) and subtilisin (MEROPS peptidase family S8, IPR000209 from INTERPRO), thiol proteinases (MEROPS peptidase family C1, IPR000668 from INTERPRO) and aspartic proteinases (MEROPS peptidase family A1, IPR001461 from INTERPRO) []. Inhibitors from cereals are active against subtilisin and endogenous alpha-amylases, while some also inhibit tissue plasminogen activator. The inhibitors are usually specific for either trypsin or chymotrypsin, and some are effective against both. They are thought to protect the seeds against consumption by animal predators, while at the same time existing as seed storage proteins themselves - all the actively inhibitory members contain 2 disulphide bridges. The existence of a member with no inhibitory activity, winged bean albumin 1, suggests that the inhibitors may have evolved from seed storage proteins. Proteins from the Kunitz family contain from 170 to 200 amino acid residues and one or two intra-chain disulphide bonds. The best conserved region is found in their N-terminal section. The crystal structures of soybean trypsin inhibitor (STI), trypsin inhibitor DE-3 from the Kaffir tree Erythrina caffra (ETI) [] and the bifunctional proteinase K/alpha-amylase inhibitor from wheat (PK13) have been solved, showing them to share the same 12-stranded beta-sheet structure as those of interleukin-1 and heparin-binding growth factors []. The beta-sheets are arranged in 3 similar lobes around a central axis, 6 strands forming an anti-parallel beta-barrel. Despite the structural similarity, STI shows no interleukin-1 bioactivity, presumably as a result of their primary sequence disparities. The active inhibitory site containing the scissile bond is located in the loop between beta-strands 4 and 5 in STI and ETI. The STIs belong to a superfamily that also contains the interleukin-1 proteins, heparin binding growth factors (HBGF) and histactophilin, all of which have very similar structures, but share no sequence similarity with the STI family.; GO: 0004866 endopeptidase inhibitor activity; PDB: 3TC2_B 3S8J_A 3S8K_A 1TIE_A 2GZB_A 3E8L_C 2IWT_B 3BX1_C 1AVA_D 3IIR_A ....
Probab=100.00 E-value=2.7e-56 Score=377.88 Aligned_cols=169 Identities=34% Similarity=0.648 Sum_probs=148.0
Q ss_pred ceecCCCCcccCCCCEEEEecccCCCCCCceEEecCCCCCCCCceEEcCCCCCCCceEEEecC--C-CCcceecccceEE
Q 040351 29 PILDVYGNQVDSSHRYYLVSALWGVKTGGGISADKGKNGQCPTDVIQLSPKDKRGKNLVLLPN--D-NSIIVRESTNIKL 105 (232)
Q Consensus 29 ~VlDtdG~~L~~G~~YYIlPa~~g~G~gGGl~l~~t~n~~CPl~VvQ~~~~~~~GlPV~Fsp~--~-~~~vI~e~t~lnI 105 (232)
||+|+|||||++|++|||+|++|+.| |||++++++|++||++|+|++++..+|+||+|+|+ . .+++|||+++|||
T Consensus 1 pVlD~~G~~l~~g~~YyI~p~~~~~G--GGl~l~~~~n~~CPl~Vvq~~~~~~~GlPv~Fs~~~~~~~~~~ir~st~l~I 78 (176)
T PF00197_consen 1 PVLDTDGNPLRNGGEYYILPAIRGAG--GGLTLAKTGNETCPLDVVQSPSELSRGLPVKFSPPYRNSFDTVIRESTDLNI 78 (176)
T ss_dssp B-BETTSCB-BTTSEEEEEESSTGCS--EEEEEECCTTSSSSEEEEEESSTTS-BSEEEEEESSSSSSTBCTBTTSEEEE
T ss_pred CcCCCCCCCCcCCCCEEEEeCccCCC--CeeEecCCCCCCCChheEEccCCCCCceeEEEEeCCcccCCCeeEcceEEEE
Confidence 79999999999999999999999997 88999999999999999999999999999999993 3 7889999999999
Q ss_pred EEecCCCCCCcCCCCcEEEecCCCcCcceEEEeCCC---CCCcCCeEEEEcccccchhccccCCCCCeEEEeccCCCCCC
Q 040351 106 KFSRVSSLQQCNKDSLWKVDNDNASLGKQFITIGEG---KTCQNLFKLEKVSASIFDMKIALDIPCLYKIVHCSTPVNGS 182 (232)
Q Consensus 106 ~F~~~~~~~~C~~st~W~V~~~d~~~g~~~V~tGg~---~~~~~~FkIeK~~~~~~~~~~~~~~~~~YKLvfCp~~~~~~ 182 (232)
+|... +.|..+++|+|+++++++++ ||+|||. ++..+||||||++.. ..+.|||+|||..|+
T Consensus 79 ~F~~~---~~c~~~~~W~V~~~~~~~~~-~V~~gg~~~~~~~~~~FkIek~~~~---------~~~~YKLvfCp~~~~-- 143 (176)
T PF00197_consen 79 EFSSP---TSCACSTVWKVVKDDPETGQ-FVKTGGVKGPETVDSWFKIEKYEDG---------FNNAYKLVFCPSVCC-- 143 (176)
T ss_dssp EESSE---CTTSSSSBEEEEEETTTTEE-EEEEESSSSSGCGCCEEEEEEESSS---------STTEEEEEEESSSSS--
T ss_pred EEccC---CCCCccCEEEEeecCcccce-EEEeCCcccCCccCcEEEEEEeCCC---------CCCcEEEEECCCccc--
Confidence 99987 78999999999987777666 8999993 368999999999851 136899999998753
Q ss_pred CCcccceeeEEee-CCeeEEEEEcCCCCCCCCeeEEEEE
Q 040351 183 CDTLCKDVGVSNV-DGVQRLVVVDDNDQPNLPLPVVLFP 220 (232)
Q Consensus 183 c~~~C~dVGi~~d-~g~rrL~l~~d~~~~~~pf~V~F~k 220 (232)
...|+||||++| +|+||||| +++ +||.|+|||
T Consensus 144 -~~~C~dvGi~~d~~g~rrL~l-~~~----~p~~V~F~K 176 (176)
T PF00197_consen 144 -DSLCGDVGIYFDDNGNRRLAL-SDD----NPFVVVFQK 176 (176)
T ss_dssp -TSSEEEEEEEEETTSEEEEEE-ESS----SB-EEEEEE
T ss_pred -cCccceeeEEEcCCCeEEEEE-CCC----CcEEEEEEC
Confidence 789999999987 79999999 764 399999998
No 4
>PF07951 Toxin_R_bind_C: Clostridium neurotoxin, C-terminal receptor binding; InterPro: IPR013104 The Clostridium neurotoxin family is composed of tetanus neurotoxins and seven serotypes of botulinum neurotoxin. The structure of the botulinum neurotoxin reveals a four domain protein. The N-terminal catalytic domain (IPR000395 from INTERPRO), the central translocation domains and two receptor-binding domains []. This domain is the C-terminal receptor-binding domain, which adopts a modified beta-trefoil fold with a six stranded beta-barrel and a beta-hairpin triplet capping the domain []. The first step in the intoxication process is a binding event between this domain and the pre-synaptic nerve ending []. ; PDB: 3AZW_A 3N7L_A 3AZV_A 3N7M_A 3RSJ_B 3FUQ_A 3RMX_D 3OBT_A 3RMY_B 3OGG_A ....
Probab=76.45 E-value=7.5 Score=34.44 Aligned_cols=120 Identities=13% Similarity=0.208 Sum_probs=70.5
Q ss_pred CCceecCCCCcccCCCCEEEEecccCCC----CCCce-EEecCCCCCCCCceEEcCCCCCCCceEEEecCC----CCcce
Q 040351 27 SEPILDVYGNQVDSSHRYYLVSALWGVK----TGGGI-SADKGKNGQCPTDVIQLSPKDKRGKNLVLLPND----NSIIV 97 (232)
Q Consensus 27 ~~~VlDtdG~~L~~G~~YYIlPa~~g~G----~gGGl-~l~~t~n~~CPl~VvQ~~~~~~~GlPV~Fsp~~----~~~vI 97 (232)
+..+.|==||||+-..+||++++..-.- ...++ ....++.. =-+.+.-..+++-.|++|++.... .+..|
T Consensus 7 ~niLKDfWGN~L~YdkeYYl~N~~~~n~yi~~~~~~~~~~n~~r~~-~~~ni~~n~r~LY~G~k~iIkr~~~~~~~Dn~V 85 (214)
T PF07951_consen 7 TNILKDFWGNYLRYDKEYYLLNVLYPNKYIKRKSDSILSINNQRGT-GVFNIYLNYRDLYTGIKFIIKRYADNSNNDNRV 85 (214)
T ss_dssp TTB-BBTTSSB-BTTSEEEEEESSSTTEEEEEETTSEEEEEEEEEE-EEEEEESEETSSSSS-EEEEEESSTSSSTSSB-
T ss_pred ccHHHHhcCCccccCceeEEEecCCcccceeecccceeeecccccc-cceeeeeeehhhccCceEEEEEccCCCCCccee
Confidence 4688899999999999999999865321 00222 22221111 011233334457799999998742 78899
Q ss_pred ecccceEEEEecCCCCCCcCCCCcEEEec---C--CCcCcceEEEeCC---CC-CC--cCCeEEEEccc
Q 040351 98 RESTNIKLKFSRVSSLQQCNKDSLWKVDN---D--NASLGKQFITIGE---GK-TC--QNLFKLEKVSA 155 (232)
Q Consensus 98 ~e~t~lnI~F~~~~~~~~C~~st~W~V~~---~--d~~~g~~~V~tGg---~~-~~--~~~FkIeK~~~ 155 (232)
|.+..+.|.|... ...|.|-- + +.+..+-.+.+.+ .. .. -..|+|++..+
T Consensus 86 r~~D~iy~n~~~~--------n~ey~l~~~~~Y~~~~~~~~kli~l~~l~~~~~~~~~~~vmqik~~~~ 146 (214)
T PF07951_consen 86 RNGDYIYFNVVIN--------NKEYRLYADTMYKNSKNQSEKLIYLLRLSDSNDNINQYIVMQIKNYNS 146 (214)
T ss_dssp BTTEEEEEEEEET--------TEEEEEEEETEECTTSSSSEEEEEEEEEECSCTTTCEECEEEEEEEEE
T ss_pred ecCCEEEEEEEeC--------CceEEEEeeeecccccccchheeeEEecccCCCCcCceEEEEEEeccc
Confidence 9999999999765 56799821 1 1222234454444 21 11 24799999864
No 5
>KOG3858 consensus Ephrin, ligand for Eph receptor tyrosine kinase [Signal transduction mechanisms]
Probab=53.54 E-value=13 Score=33.37 Aligned_cols=55 Identities=24% Similarity=0.266 Sum_probs=35.9
Q ss_pred CCCcccCCCCEEEEecccCCC-----CCCceEEecCCCCCCCCceEEcCCC-CCCCceEEEec
Q 040351 34 YGNQVDSSHRYYLVSALWGVK-----TGGGISADKGKNGQCPTDVIQLSPK-DKRGKNLVLLP 90 (232)
Q Consensus 34 dG~~L~~G~~YYIlPa~~g~G-----~gGGl~l~~t~n~~CPl~VvQ~~~~-~~~GlPV~Fsp 90 (232)
.|.+-++|.+||+++...|.- +-||+- .+.+.+|-..|.|++.. ...=.++.+.|
T Consensus 117 ~G~EF~pG~~YY~IStStg~~~g~~~~~ggvc--~~~~mk~~~~V~~~~~~~~~~~~~~~~~p 177 (233)
T KOG3858|consen 117 LGFEFQPGHTYYYISTSTGDAEGLCNLRGGVC--VTRNMKLLMKVGQSPRSGVTPEKPVSEEP 177 (233)
T ss_pred CCccccCCCeEEEEeCCCccccccchhhCCEe--ccCCceEEEEecccCCCCccccccccccc
Confidence 699999999999999876642 123444 34466787888888653 11223555544
No 6
>PF02402 Lysis_col: Lysis protein; InterPro: IPR003059 The DNA sequence of the entire colicin E2 operon has been determined []. The operon comprises the colicin activity gene (ceaB), the colicin immunity gene (ceiB) and the lysis gene (celB), which is essential for colicin release from producing cells []. A putative LexA binding site is located upstream from ceaB, and a rho-independent terminator structure is located downstream from celB []. Comparison of the amino acid sequences of colicin E2 and cloacin DF13 reveal extensive similarity. These colicins have different modes of action and recognise different cell surface receptors; the two major regions of heterology at the C terminus, and in the C-terminal end of the central region are thought to correspond to the catalytic and receptor-recognition domains, respectively []. Sequence similarities between colicins E2, A and E1 [] are less striking. The colicin E2 (pyocin) immunity protein does not share similarity with either the colicin E3 or cloacin DF13 [] immunity proteins. By contrast, the lysis proteins of the ColE2, ColE1 and CloDF13 plasmids are almost identical except in the N-terminal regions, which themselves are similar to lipoprotein signal peptides []. Processing of the ColE2 prolysis protein to the mature form is prevented by globomycin, a specific inhibitor of the lipoprotein signal peptidase []. The mature ColE2 lysis protein is located in the cell envelope [].; GO: 0009405 pathogenesis, 0019835 cytolysis, 0019867 outer membrane
Probab=45.35 E-value=18 Score=24.48 Aligned_cols=19 Identities=37% Similarity=0.419 Sum_probs=15.2
Q ss_pred CCCCCceecCCCCcccCCC
Q 040351 24 ASESEPILDVYGNQVDSSH 42 (232)
Q Consensus 24 ~a~~~~VlDtdG~~L~~G~ 42 (232)
+.+.+.|.|+.|--+.+..
T Consensus 18 aCQaN~iRDvqGGtVaPSS 36 (46)
T PF02402_consen 18 ACQANYIRDVQGGTVAPSS 36 (46)
T ss_pred HhhhcceecCCCceECCCc
Confidence 5678899999998887653
No 7
>PF07172 GRP: Glycine rich protein family; InterPro: IPR010800 This family consists of glycine rich proteins. Some of them may be involved in resistance to environmental stress [].
Probab=39.62 E-value=18 Score=27.93 Aligned_cols=8 Identities=38% Similarity=0.351 Sum_probs=3.5
Q ss_pred cchhHHHH
Q 040351 2 KTSLVTTL 9 (232)
Q Consensus 2 K~~~~~~l 9 (232)
|+.+||.|
T Consensus 4 K~~llL~l 11 (95)
T PF07172_consen 4 KAFLLLGL 11 (95)
T ss_pred hHHHHHHH
Confidence 34444444
No 8
>PF14009 DUF4228: Domain of unknown function (DUF4228)
Probab=37.12 E-value=21 Score=28.79 Aligned_cols=21 Identities=10% Similarity=0.143 Sum_probs=17.2
Q ss_pred CCCCcccCCCCEEEEecccCC
Q 040351 33 VYGNQVDSSHRYYLVSALWGV 53 (232)
Q Consensus 33 tdG~~L~~G~~YYIlPa~~g~ 53 (232)
.-.++|++|.-||++|..+-.
T Consensus 64 ~~d~~L~~G~~Y~llP~~~~~ 84 (181)
T PF14009_consen 64 PPDEELQPGQIYFLLPMSRLQ 84 (181)
T ss_pred CccCeecCCCEEEEEEccccC
Confidence 456789999999999987643
No 9
>COG5341 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=36.03 E-value=2.5e+02 Score=23.15 Aligned_cols=18 Identities=11% Similarity=0.126 Sum_probs=11.4
Q ss_pred CCcceecccceEEEEecC
Q 040351 93 NSIIVRESTNIKLKFSRV 110 (232)
Q Consensus 93 ~~~vI~e~t~lnI~F~~~ 110 (232)
.+++|-.-..|-|++...
T Consensus 103 GetIVclPh~lvIev~~~ 120 (132)
T COG5341 103 GETIVCLPHKLVIEVKSK 120 (132)
T ss_pred CCEEEEcCCeEEEEEEcc
Confidence 455666666677777654
No 10
>PF05474 Semenogelin: Semenogelin; InterPro: IPR008836 This family consists of several mammalian semenogelin (I and II) proteins. Freshly ejaculated Homo sapiens semen has the appearance of a loose gel in which the predominant structural protein components are the seminal vesicle secreted semenogelins (Sg) [].; GO: 0005198 structural molecule activity, 0019953 sexual reproduction, 0005576 extracellular region, 0030141 stored secretory granule
Probab=32.34 E-value=23 Score=35.25 Aligned_cols=15 Identities=33% Similarity=0.545 Sum_probs=12.1
Q ss_pred CcchhHHHHHHHHHH
Q 040351 1 MKTSLVTTLSFLILA 15 (232)
Q Consensus 1 MK~~~~~~lsfll~a 15 (232)
||++++|.||+||++
T Consensus 1 MK~~I~F~lSLLLiL 15 (582)
T PF05474_consen 1 MKSIIFFVLSLLLIL 15 (582)
T ss_pred CCceeehHHHHHHHH
Confidence 999988888876654
No 11
>PF08194 DIM: DIM protein; InterPro: IPR013172 Drosophila immune-induced molecules (DIMs) are short proteins induced during the immune response of Drosophila []. This entry includes DIMs 1 to 4 and DIM23.
Probab=32.13 E-value=62 Score=20.94 Aligned_cols=16 Identities=38% Similarity=0.586 Sum_probs=9.1
Q ss_pred CcchhHHHHHHHHHHHh
Q 040351 1 MKTSLVTTLSFLILALT 17 (232)
Q Consensus 1 MK~~~~~~lsfll~a~~ 17 (232)
||...+ ++.++++|++
T Consensus 1 Mk~l~~-a~~l~lLal~ 16 (36)
T PF08194_consen 1 MKCLSL-AFALLLLALA 16 (36)
T ss_pred CceeHH-HHHHHHHHHH
Confidence 887765 2335555554
No 12
>PF05550 Peptidase_C53: Pestivirus Npro endopeptidase C53; InterPro: IPR008751 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This group of cysteine peptidases belong to MEROPS peptidase family C53 (clan C-). The active site residues occur in the order E, H, C in the sequence which is unlike that in any other family. They are unique to pestiviruses. The N-terminal cysteine peptidase (Npro) encoded by the bovine viral diarrhoea virus genome is responsible for the self-cleavage that releases the N terminus of the core protein. This unique protease is dispensable for viral replication, and its coding region can be replaced by a ubiquitin gene directly fused in frame to the core [, , , ].; GO: 0016032 viral reproduction, 0019082 viral protein processing
Probab=30.27 E-value=35 Score=28.84 Aligned_cols=25 Identities=32% Similarity=0.320 Sum_probs=18.4
Q ss_pred hcCCCCCCCCCCceecCCCCcccCC
Q 040351 17 TTKPQLGASESEPILDVYGNQVDSS 41 (232)
Q Consensus 17 ~t~~~l~~a~~~~VlDtdG~~L~~G 41 (232)
-|+..-|+...|||+|..|+||.-.
T Consensus 11 kt~kq~P~Gv~EPVyd~~g~plfGe 35 (168)
T PF05550_consen 11 KTYKQKPAGVEEPVYDSAGNPLFGE 35 (168)
T ss_pred hhcCCCCcccccccccCCCCCccCC
Confidence 3444445566799999999999864
No 13
>PF00879 Defensin_propep: Defensin propeptide The pattern for this Prosite entry doesn't match the propeptide.; InterPro: IPR002366 Defensins are 2-6 kDa, cationic, microbicidal peptides active against many Gram-negative and Gram-positive bacteria, fungi, and enveloped viruses [], containing three pairs of intramolecular disulphide bonds []. On the basis of their size and pattern of disulphide bonding, mammalian defensins are classified into alpha, beta and theta categories. Alpha-defensins, which have been identified in humans, monkeys and several rodent species, are particularly abundant in neutrophils, certain macrophage populations and Paneth cells of the small intestine. Every mammalian species explored thus far has beta-defensins. In cows, as many as 13 beta-defensins exist in neutrophils. However, in other species, beta-defensins are more often produced by epithelial cells lining various organs (e.g. the epidermis, bronchial tree and genitourinary tract). Theta-defensins are cyclic and have so far only been identified in primate phagocytes. Defensins are produced constitutively and/or in response to microbial products or proinflammatory cytokines. Some defensins are also called corticostatins (CS) because they inhibit corticotropin-stimulated corticosteroid production. The mechanism(s) by which microorganisms are killed and/or inactivated by defensins is not understood completely. However, it is generally believed that killing is a consequence of disruption of the microbial membrane. The polar topology of defensins, with spatially separated charged and hydrophobic regions, allows them to insert themselves into the phospholipid membranes so that their hydrophobic regions are buried within the lipid membrane interior and their charged (mostly cationic) regions interact with anionic phospholipid head groups and water. Subsequently, some defensins can aggregate to form `channel-like' pores; others might bind to and cover the microbial membrane in a `carpet-like' manner. The net outcome is the disruption of membrane integrity and function, which ultimately leads to the lysis of microorganisms. Some defensins are synthesized as propeptides which may be relevant to this process - in neutrophils only the mature peptides have been identified but in Paneth cells, the propeptide is stored in vesicles [] and appears to be cleaved by trypsin on activation. ; GO: 0006952 defense response
Probab=29.91 E-value=54 Score=22.87 Aligned_cols=18 Identities=33% Similarity=0.508 Sum_probs=12.5
Q ss_pred CcchhHHHHHHHHHHHhcC
Q 040351 1 MKTSLVTTLSFLILALTTK 19 (232)
Q Consensus 1 MK~~~~~~lsfll~a~~t~ 19 (232)
||+..||+- .||+||-+.
T Consensus 1 MRTL~LLaA-lLLlAlqaQ 18 (52)
T PF00879_consen 1 MRTLALLAA-LLLLALQAQ 18 (52)
T ss_pred CcHHHHHHH-HHHHHHHHh
Confidence 888766544 678887654
No 14
>PF09466 Yqai: Hypothetical protein Yqai; InterPro: IPR018474 The hypothetical protein YqaI is expressed in bacteria, particularly Bacillus subtilis. It forms a homo-dimer, with each monomer containing an alpha helix and four beta strands.; PDB: 2DSM_B.
Probab=29.69 E-value=41 Score=24.94 Aligned_cols=21 Identities=29% Similarity=0.744 Sum_probs=12.2
Q ss_pred CceecCCCCcccCCCCEEEEe
Q 040351 28 EPILDVYGNQVDSSHRYYLVS 48 (232)
Q Consensus 28 ~~VlDtdG~~L~~G~~YYIlP 48 (232)
-++.|.-|+++.+|.+|+|.|
T Consensus 24 ~~i~D~yG~EI~~~D~y~i~~ 44 (71)
T PF09466_consen 24 HPIEDFYGDEIFPGDDYFISP 44 (71)
T ss_dssp -B---TTSS-B-TTS-EEE-E
T ss_pred cceeeeeccccccCCeEEEeC
Confidence 378899999999999999966
No 15
>PRK11354 kil FtsZ inhibitor protein; Reviewed
Probab=27.89 E-value=41 Score=24.95 Aligned_cols=47 Identities=21% Similarity=0.144 Sum_probs=32.7
Q ss_pred CCcccCCCCEEEEecccCCCCCCceEEecCCCCC----CCCceEEcCCCCCCCceEE
Q 040351 35 GNQVDSSHRYYLVSALWGVKTGGGISADKGKNGQ----CPTDVIQLSPKDKRGKNLV 87 (232)
Q Consensus 35 G~~L~~G~~YYIlPa~~g~G~gGGl~l~~t~n~~----CPl~VvQ~~~~~~~GlPV~ 87 (232)
|.-+.-.+++|+.+|..-.. |.|+|......+ |-..|..+ .+|.|++
T Consensus 19 G~~v~~~grty~ASAN~~~r--~~LYl~~~~e~~~i~d~~IeVyL~----~~G~Plt 69 (73)
T PRK11354 19 GDYVLHEGRTYIASANNIKK--RKLYIRTLTTKTCITDCMIKVFLG----RDGLPVK 69 (73)
T ss_pred ceEEEEcCcEEEEEechhhC--ceEEEEeeeEEEEEeeeEEEEEEc----CCCCccc
Confidence 66677778999999864454 799987643333 66666655 4788885
No 16
>COG4851 CamS Protein involved in sex pheromone biosynthesis [General function prediction only]
Probab=27.25 E-value=63 Score=30.57 Aligned_cols=59 Identities=19% Similarity=0.276 Sum_probs=32.6
Q ss_pred CcchhHHHHHHHHHHHhcCCCCCCCCCCceecCCCCccc----------CCCCEE--EEecccCCCCCCceEE
Q 040351 1 MKTSLVTTLSFLILALTTKPQLGASESEPILDVYGNQVD----------SSHRYY--LVSALWGVKTGGGISA 61 (232)
Q Consensus 1 MK~~~~~~lsfll~a~~t~~~l~~a~~~~VlDtdG~~L~----------~G~~YY--IlPa~~g~G~gGGl~l 61 (232)
||.++.+++-=+.|.|+........+.+.|.+.+|++-. ....|| ++|-..|.. -||..
T Consensus 1 Mkktl~i~~ta~vliLs~C~~~~dd~edkv~qk~~~sseq~k~Iv~k~nis~n~YktvLpyk~~ka--RGl~~ 71 (382)
T COG4851 1 MKKTLGIAATASVLILSGCFPFVDDTEDKVVQKEGKSSEQEKGIVPKANISENYYKTVLPYKAGKA--RGLGS 71 (382)
T ss_pred CchhhhHHHHHHHHHHhhccCccCCccchhhhccCCchhhcccccccccccccchheeeeeccccc--cchhh
Confidence 888765433111122222221223456788888888766 246788 888766654 34443
No 17
>PRK10220 hypothetical protein; Provisional
Probab=26.51 E-value=79 Score=25.35 Aligned_cols=27 Identities=19% Similarity=0.099 Sum_probs=20.8
Q ss_pred CceecCCCCcccCCCCEEEEecccCCC
Q 040351 28 EPILDVYGNQVDSSHRYYLVSALWGVK 54 (232)
Q Consensus 28 ~~VlDtdG~~L~~G~~YYIlPa~~g~G 54 (232)
..|.|.+|++|..|.+--++=.+.=.|
T Consensus 42 ~~vkDsnG~~L~dGDsV~viKDLkVKG 68 (111)
T PRK10220 42 LIVKDANGNLLADGDSVTIVKDLKVKG 68 (111)
T ss_pred ceEEcCCCCCccCCCEEEEEeeccccc
Confidence 468999999999998887776544343
No 18
>TIGR00686 phnA alkylphosphonate utilization operon protein PhnA. The protein family includes an uncharacterized member designated phnA in Escherichia coli, part of a large operon associated with alkylphosphonate uptake and carbon-phosphorus bond cleavage. This protein is not related to the characterized phosphonoacetate hydrolase designated PhnA by Kulakova, et al. (2001, 1997).
Probab=25.68 E-value=85 Score=25.11 Aligned_cols=27 Identities=19% Similarity=0.128 Sum_probs=20.8
Q ss_pred CceecCCCCcccCCCCEEEEecccCCC
Q 040351 28 EPILDVYGNQVDSSHRYYLVSALWGVK 54 (232)
Q Consensus 28 ~~VlDtdG~~L~~G~~YYIlPa~~g~G 54 (232)
..|.|.+|++|..|.+--|+=.+.=.|
T Consensus 41 ~~~kDsnG~~L~dGDsV~liKDLkVKG 67 (109)
T TIGR00686 41 LIVKDCNGNLLANGDSVILIKDLKVKG 67 (109)
T ss_pred ceEEcCCCCCccCCCEEEEEeeccccC
Confidence 468999999999998887776554343
No 19
>TIGR02052 MerP mercuric transport protein periplasmic component. This model represents the periplasmic mercury (II) binding protein of the bacterial mercury detoxification system which passes mercuric ion to the MerT transporter for subsequent reduction to Hg(0) by the mercuric reductase MerA. MerP contains a distinctive GMTCXXC motif associated with metal binding. MerP is related to a larger family of metal binding proteins (pfam00403).
Probab=24.91 E-value=47 Score=22.77 Aligned_cols=14 Identities=43% Similarity=0.548 Sum_probs=7.5
Q ss_pred CcchhHHHHHHHHHHHh
Q 040351 1 MKTSLVTTLSFLILALT 17 (232)
Q Consensus 1 MK~~~~~~lsfll~a~~ 17 (232)
||..+ +| |+||.++
T Consensus 1 ~~~~~--~~-~~~~~~~ 14 (92)
T TIGR02052 1 MKKLA--TL-LALFVLT 14 (92)
T ss_pred ChhHH--HH-HHHHHHh
Confidence 88764 33 4444444
No 20
>COG5510 Predicted small secreted protein [Function unknown]
Probab=23.04 E-value=64 Score=21.76 Aligned_cols=17 Identities=18% Similarity=0.251 Sum_probs=9.5
Q ss_pred CcchhHHHHHHHHHHHh
Q 040351 1 MKTSLVTTLSFLILALT 17 (232)
Q Consensus 1 MK~~~~~~lsfll~a~~ 17 (232)
||.+.++.+++++.++.
T Consensus 2 mk~t~l~i~~vll~s~l 18 (44)
T COG5510 2 MKKTILLIALVLLASTL 18 (44)
T ss_pred chHHHHHHHHHHHHHHH
Confidence 77766555545555443
No 21
>COG2824 PhnA Uncharacterized Zn-ribbon-containing protein involved in phosphonate metabolism [Inorganic ion transport and metabolism]
Probab=22.84 E-value=1e+02 Score=24.68 Aligned_cols=27 Identities=19% Similarity=0.096 Sum_probs=21.4
Q ss_pred CceecCCCCcccCCCCEEEEecccCCC
Q 040351 28 EPILDVYGNQVDSSHRYYLVSALWGVK 54 (232)
Q Consensus 28 ~~VlDtdG~~L~~G~~YYIlPa~~g~G 54 (232)
..|.|.+||.|..|.+-.|+-...=.|
T Consensus 43 ~~v~DsnGn~L~dGDsV~lIKDLkVKG 69 (112)
T COG2824 43 LIVKDSNGNLLADGDSVTLIKDLKVKG 69 (112)
T ss_pred eEEEcCCCcEeccCCeEEEEEeeeecC
Confidence 589999999999998888776554343
No 22
>PRK10159 outer membrane phosphoporin protein E; Provisional
Probab=22.59 E-value=51 Score=30.66 Aligned_cols=36 Identities=19% Similarity=0.385 Sum_probs=20.9
Q ss_pred CcchhHHHHHHHHHHHhcCCCCCCCCCCceecCCCCcccCCC
Q 040351 1 MKTSLVTTLSFLILALTTKPQLGASESEPILDVYGNQVDSSH 42 (232)
Q Consensus 1 MK~~~~~~lsfll~a~~t~~~l~~a~~~~VlDtdG~~L~~G~ 42 (232)
||.+++ +|. +++++..+ ++.+.+|+|.||.-|...+
T Consensus 1 Mkk~l~-a~~--~~a~~~~~---~a~A~~vy~~d~ssvtlyG 36 (351)
T PRK10159 1 MKKSTL-ALV--VMGIVASA---SVQAAEVYNKDGNKLDVYG 36 (351)
T ss_pred CchhhH-HHH--HHHHHHhc---cccEEEEEECCCCEEEEEE
Confidence 898764 332 23222111 2455789999997777643
No 23
>PF00812 Ephrin: Ephrin; InterPro: IPR001799 Ephrins are a family of proteins [] that are ligands of class V (EPH-related) receptor protein-tyrosine kinases (see IPR001426 from INTERPRO). These receptors and their ligands have been implicated in regulating neuronal axon guidance and in patterning of the developing nervous system and may also serve a patterning and compartmentalisation role outside of the nervous system as well. Ephrins are membrane-attached proteins of 205 to 340 residues. Attachment appears to be crucial for their normal function. Type-A ephrins are linked to the membrane via a glycosylphosphatidylinositol (GPI)-linkage, while type-B ephrins are type-I membrane proteins.; GO: 0016020 membrane; PDB: 3HEI_P 3CZU_B 3MBW_B 1KGY_E 1IKO_P 2WO3_B 2I85_A 2VSK_B 3GXU_B 2VSM_B ....
Probab=22.44 E-value=55 Score=27.23 Aligned_cols=21 Identities=29% Similarity=0.585 Sum_probs=16.5
Q ss_pred CCCcccCCCCEEEEecccCCC
Q 040351 34 YGNQVDSSHRYYLVSALWGVK 54 (232)
Q Consensus 34 dG~~L~~G~~YYIlPa~~g~G 54 (232)
.|-+-++|.+||++....|.-
T Consensus 101 ~G~EF~pG~~YY~ISts~g~~ 121 (145)
T PF00812_consen 101 LGLEFQPGHDYYYISTSTGTQ 121 (145)
T ss_dssp TSSS--TTEEEEEEEEESSSS
T ss_pred CCeeecCCCeEEEEEccCCCC
Confidence 799999999999999887753
No 24
>PLN03207 stomagen; Provisional
Probab=21.86 E-value=43 Score=26.51 Aligned_cols=8 Identities=50% Similarity=0.716 Sum_probs=3.1
Q ss_pred HHHHHHHH
Q 040351 8 TLSFLILA 15 (232)
Q Consensus 8 ~lsfll~a 15 (232)
.|.||||+
T Consensus 14 ~lffLl~~ 21 (113)
T PLN03207 14 TLFFLLFF 21 (113)
T ss_pred HHHHHHHH
Confidence 33344433
No 25
>PF03831 PhnA: PhnA protein; InterPro: IPR013988 The PhnA protein family includes the uncharacterised Escherichia coli protein PhnA and its homologues. The E. coli phnA gene is part of a large operon associated with alkylphosphonate uptake and carbon-phosphorus bond cleavage []. The protein is not related to the characterised phosphonoacetate hydrolase designated PhnA []. This entry represents the C-terminal domain of PhnA.; PDB: 2AKK_A 2AKL_A.
Probab=20.90 E-value=52 Score=23.33 Aligned_cols=21 Identities=24% Similarity=0.431 Sum_probs=11.8
Q ss_pred eecCCCCcccCCCCEEEEecc
Q 040351 30 ILDVYGNQVDSSHRYYLVSAL 50 (232)
Q Consensus 30 VlDtdG~~L~~G~~YYIlPa~ 50 (232)
|.|.+|++|..|.+--++--.
T Consensus 2 v~DsnGn~L~dGDsV~~iKDL 22 (56)
T PF03831_consen 2 VKDSNGNELQDGDSVTLIKDL 22 (56)
T ss_dssp -B-TTS-B--TTEEEEESS-E
T ss_pred eEcCCCCCccCCCEEEEEeee
Confidence 789999999999877765443
Done!