Query 047082
Match_columns 263
No_of_seqs 119 out of 1284
Neff 7.6
Searched_HMMs 46136
Date Fri Mar 29 07:26:30 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/047082.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/047082hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF01453 B_lectin: D-mannose b 100.0 3E-29 6.4E-34 195.8 6.4 87 5-91 1-93 (114)
2 PF00954 S_locus_glycop: S-loc 99.8 4.2E-21 9.1E-26 148.8 9.6 99 120-218 1-108 (110)
3 cd00028 B_lectin Bulb-type man 99.8 8E-21 1.7E-25 148.5 9.8 75 6-80 41-116 (116)
4 smart00108 B_lectin Bulb-type 99.8 8.8E-20 1.9E-24 142.2 9.7 74 6-79 40-114 (114)
5 smart00108 B_lectin Bulb-type 98.4 1.1E-06 2.4E-11 68.1 7.3 55 22-76 22-79 (114)
6 cd00028 B_lectin Bulb-type man 98.4 5.6E-07 1.2E-11 70.1 5.4 56 23-78 23-82 (116)
7 PF08276 PAN_2: PAN-like domai 98.4 2.8E-07 6E-12 64.7 2.9 25 235-259 25-49 (66)
8 PF01453 B_lectin: D-mannose b 98.1 2.6E-05 5.6E-10 60.7 8.8 73 6-80 38-113 (114)
9 cd01098 PAN_AP_plant Plant PAN 98.1 3E-06 6.4E-11 61.5 3.0 27 235-261 30-56 (84)
10 cd00129 PAN_APPLE PAN/APPLE-li 97.8 2.1E-05 4.6E-10 57.3 3.2 25 236-260 24-51 (80)
11 smart00473 PAN_AP divergent su 96.5 0.0029 6.3E-08 44.5 3.6 25 235-259 23-48 (78)
12 cd01100 APPLE_Factor_XI_like S 96.3 0.0045 9.7E-08 43.9 3.3 26 235-260 23-48 (73)
13 PF00024 PAN_1: PAN domain Thi 77.0 2.8 6E-05 29.0 2.9 26 236-261 22-48 (79)
14 smart00223 APPLE APPLE domain. 76.2 3.2 7E-05 29.9 3.1 32 231-262 16-47 (79)
15 PF07354 Sp38: Zona-pellucida- 75.9 4.2 9.2E-05 36.1 4.3 36 2-37 9-44 (271)
16 PF01436 NHL: NHL repeat; Int 74.8 5.9 0.00013 22.4 3.4 21 23-43 6-26 (28)
17 PF14295 PAN_4: PAN domain; PD 73.1 4 8.6E-05 25.9 2.7 25 235-259 14-38 (51)
18 cd05845 Ig2_L1-CAM_like Second 67.0 10 0.00023 28.3 4.1 33 3-35 31-63 (95)
19 cd01099 PAN_AP_HGF Subfamily o 66.5 7.5 0.00016 27.8 3.2 25 236-260 24-50 (80)
20 PF13360 PQQ_2: PQQ-like domai 65.3 14 0.0003 31.1 5.1 73 4-76 11-102 (238)
21 cd00053 EGF Epidermal growth f 62.4 7.9 0.00017 21.9 2.2 29 190-218 2-31 (36)
22 TIGR03066 Gem_osc_para_1 Gemma 61.5 27 0.00058 27.0 5.5 55 18-72 32-104 (111)
23 PRK11138 outer membrane biogen 60.9 35 0.00076 31.6 7.4 56 21-76 120-186 (394)
24 PRK11138 outer membrane biogen 58.9 23 0.00051 32.8 5.9 19 58-76 342-361 (394)
25 PF13360 PQQ_2: PQQ-like domai 50.7 63 0.0014 26.9 6.8 73 4-76 54-148 (238)
26 KOG0640 mRNA cleavage stimulat 45.7 1.1E+02 0.0024 28.2 7.6 71 5-75 247-333 (430)
27 PLN00033 photosystem II stabil 44.7 73 0.0016 30.1 6.7 52 23-74 252-305 (398)
28 PF07974 EGF_2: EGF-like domai 43.7 19 0.0004 21.3 1.7 22 194-216 6-27 (32)
29 TIGR03300 assembly_YfgL outer 43.1 1.4E+02 0.003 27.2 8.3 72 4-75 83-170 (377)
30 PF09064 Tme5_EGF_like: Thromb 41.5 16 0.00035 22.0 1.1 16 201-216 11-26 (34)
31 PF08277 PAN_3: PAN-like domai 41.3 36 0.00078 23.2 3.2 24 235-258 18-41 (71)
32 KOG4649 PQQ (pyrrolo-quinoline 39.8 50 0.0011 29.7 4.4 43 6-48 169-217 (354)
33 TIGR03300 assembly_YfgL outer 39.8 68 0.0015 29.3 5.7 17 58-74 327-344 (377)
34 cd05852 Ig5_Contactin-1 Fifth 39.0 35 0.00076 23.6 2.8 33 3-36 13-45 (73)
35 smart00605 CW CW domain. 39.0 37 0.00081 24.8 3.1 23 236-258 21-43 (94)
36 PF07645 EGF_CA: Calcium-bindi 35.7 9.4 0.0002 23.7 -0.5 28 190-217 5-34 (42)
37 KOG0291 WD40-repeat-containing 35.3 5.3E+02 0.011 26.8 16.7 108 5-120 372-505 (893)
38 smart00179 EGF_CA Calcium-bind 34.9 36 0.00079 19.7 2.1 28 190-217 5-33 (39)
39 cd00216 PQQ_DH Dehydrogenases 34.0 1.4E+02 0.003 28.8 7.0 72 4-75 37-135 (488)
40 PF13570 PQQ_3: PQQ-like domai 31.2 40 0.00088 20.2 1.9 8 40-47 2-9 (40)
41 PF05935 Arylsulfotrans: Aryls 30.3 49 0.0011 31.9 3.2 52 29-81 127-186 (477)
42 PF01683 EB: EB module; Inter 28.4 43 0.00093 21.5 1.7 27 189-218 21-47 (52)
43 smart00564 PQQ beta-propeller 27.4 1.2E+02 0.0026 16.9 3.4 17 28-44 14-31 (33)
44 PF06006 DUF905: Bacterial pro 24.8 80 0.0017 22.2 2.5 17 63-79 35-51 (70)
45 cd00054 EGF_CA Calcium-binding 24.5 70 0.0015 18.0 2.1 28 190-217 5-33 (38)
46 cd00216 PQQ_DH Dehydrogenases 24.2 1.7E+02 0.0037 28.1 5.7 16 60-75 415-431 (488)
47 KOG3881 Uncharacterized conser 23.8 2.9E+02 0.0064 26.1 6.7 61 22-83 218-280 (412)
48 PF10636 hemP: Hemin uptake pr 23.4 90 0.002 19.3 2.3 13 23-35 25-37 (38)
49 PF02035 Coagulin: Coagulin; 22.1 76 0.0016 25.2 2.2 34 172-205 101-142 (174)
50 cd05764 Ig_2 Subgroup of the i 21.8 1.6E+02 0.0035 19.6 3.8 33 3-35 13-45 (74)
51 PF05833 FbpA: Fibronectin-bin 21.8 75 0.0016 30.2 2.7 39 53-91 115-159 (455)
52 KOG3848 Extracellular protein 21.4 3.3E+02 0.0072 26.1 6.6 51 11-67 194-250 (516)
53 TIGR03075 PQQ_enz_alc_DH PQQ-d 20.7 4.1E+02 0.0089 26.0 7.6 75 1-75 84-196 (527)
54 COG1520 FOG: WD40-like repeat 20.5 4.3E+02 0.0093 24.1 7.4 42 5-46 130-180 (370)
55 COG3236 Uncharacterized protei 20.3 55 0.0012 26.5 1.2 21 55-75 116-137 (162)
No 1
>PF01453 B_lectin: D-mannose binding lectin; InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]: Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity. Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=99.95 E-value=3e-29 Score=195.75 Aligned_cols=87 Identities=47% Similarity=0.679 Sum_probs=67.0
Q ss_pred CCcEEEEcCCCCCCC---CCCEEEEecCCcEEEEcCCCcEEEcC-CCCCCC--eeEEEEeeCCCeeEEcCCCceEeeecc
Q 047082 5 KNSCSLNSIFASPVC---GKPTFSLGSDGNLVLAEADGTVVCQS-NTANKG--VVGFKLLPNGNMVLHDSKGNFIWQSFD 78 (263)
Q Consensus 5 ~~~~vWvANr~~Pv~---~~~~l~l~~~G~L~l~~~~g~~~Wst-~~~~~~--~~~a~Lld~GNlvl~~~~~~~~WqSFd 78 (263)
++++||+|||+.||. +..+|.|+.||+|+|++..++.+|++ .+.+.+ ...|+|+|+|||||++..+.+||||||
T Consensus 1 ~~tvvW~an~~~p~~~~s~~~~L~l~~dGnLvl~~~~~~~iWss~~t~~~~~~~~~~~L~~~GNlvl~d~~~~~lW~Sf~ 80 (114)
T PF01453_consen 1 PRTVVWVANRNSPLTSSSGNYTLILQSDGNLVLYDSNGSVIWSSNNTSGRGNSGCYLVLQDDGNLVLYDSSGNVLWQSFD 80 (114)
T ss_dssp ---------TTEEEEECETTEEEEEETTSEEEEEETTTEEEEE--S-TTSS-SSEEEEEETTSEEEEEETTSEEEEESTT
T ss_pred CcccccccccccccccccccccceECCCCeEEEEcCCCCEEEEecccCCccccCeEEEEeCCCCEEEEeecceEEEeecC
Confidence 468999999999994 34899999999999999998999999 666544 689999999999999988999999999
Q ss_pred CCCceeccCcccC
Q 047082 79 CPTDTLLVGQSLL 91 (263)
Q Consensus 79 ~PTDTlLpGq~l~ 91 (263)
|||||+||||+|+
T Consensus 81 ~ptdt~L~~q~l~ 93 (114)
T PF01453_consen 81 YPTDTLLPGQKLG 93 (114)
T ss_dssp SSS-EEEEEET--
T ss_pred CCccEEEeccCcc
Confidence 9999999999986
No 2
>PF00954 S_locus_glycop: S-locus glycoprotein family; InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=99.85 E-value=4.2e-21 Score=148.79 Aligned_cols=99 Identities=15% Similarity=0.192 Sum_probs=85.0
Q ss_pred eecCCCCceeeecccCCCccceeEEEEEecCCCceEEEeecCCCceEEEEEeecCCEEEEEecCCC-CCCC--------C
Q 047082 120 LYDPSVQNLTFNSRPETDEAFAFKLTLDISDSGSDILARPKYNIRSSFLRLGMHGNLKIYTHYDKV-DSQP--------T 190 (263)
Q Consensus 120 ~~sg~w~~~~f~~~p~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~r~~Ld~dG~lr~y~~~~~~-~w~~--------C 190 (263)
|++|+|+|..|++.|+|.....+.+.|+.++.+.++.+.....+.++|++||.+|+|++|.|.+.. .|.. |
T Consensus 1 wrsG~WnG~~f~g~p~~~~~~~~~~~fv~~~~e~~~t~~~~~~s~~~r~~ld~~G~l~~~~w~~~~~~W~~~~~~p~d~C 80 (110)
T PF00954_consen 1 WRSGPWNGQRFSGIPEMSSNSLYNYSFVSNNEEVYYTYSLSNSSVLSRLVLDSDGQLQRYIWNESTQSWSVFWSAPKDQC 80 (110)
T ss_pred CCccccCCeEECCcccccccceeEEEEEECCCeEEEEEecCCCceEEEEEEeeeeEEEEEEEecCCCcEEEEEEecccCC
Confidence 568999999999999987656677778777778888888777778999999999999999998654 4542 9
Q ss_pred CCCCCCCCCcccCCCCCcCCCCCCCCcc
Q 047082 191 QLPERCSKLGVCDDNQCVACPTEKGLLG 218 (263)
Q Consensus 191 ~~~~~CG~~g~C~~~~~~~C~c~~g~~~ 218 (263)
|+|++||+||+|+.+..+.|.|++||..
T Consensus 81 d~y~~CG~~g~C~~~~~~~C~Cl~GF~P 108 (110)
T PF00954_consen 81 DVYGFCGPNGICNSNNSPKCSCLPGFEP 108 (110)
T ss_pred CCccccCCccEeCCCCCCceECCCCcCC
Confidence 9999999999999888889999999964
No 3
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=99.84 E-value=8e-21 Score=148.53 Aligned_cols=75 Identities=45% Similarity=0.678 Sum_probs=68.7
Q ss_pred CcEEEEcCCCCCCCCCCEEEEecCCcEEEEcCCCcEEEcCCCCC-CCeeEEEEeeCCCeeEEcCCCceEeeeccCC
Q 047082 6 NSCSLNSIFASPVCGKPTFSLGSDGNLVLAEADGTVVCQSNTAN-KGVVGFKLLPNGNMVLHDSKGNFIWQSFDCP 80 (263)
Q Consensus 6 ~~~vWvANr~~Pv~~~~~l~l~~~G~L~l~~~~g~~~Wst~~~~-~~~~~a~Lld~GNlvl~~~~~~~~WqSFd~P 80 (263)
.++||+|||+.|....+.|.|+.+|+|+|.|.+|.++|++++.+ .+...|+|+|+|||||++.++++||||||||
T Consensus 41 ~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S~~~~~~~~~~~~L~ddGnlvl~~~~~~~~W~Sf~~P 116 (116)
T cd00028 41 RTVVWVANRDNPSGSSCTLTLQSDGNLVIYDGSGTVVWSSNTTRVNGNYVLVLLDDGNLVLYDSDGNFLWQSFDYP 116 (116)
T ss_pred CeEEEECCCCCCCCCCEEEEEecCCCeEEEcCCCcEEEEecccCCCCceEEEEeCCCCEEEECCCCCEEEcCCCCC
Confidence 57999999999966678999999999999999999999999876 5567899999999999998899999999999
No 4
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=99.81 E-value=8.8e-20 Score=142.16 Aligned_cols=74 Identities=46% Similarity=0.685 Sum_probs=67.8
Q ss_pred CcEEEEcCCCCCCCCCCEEEEecCCcEEEEcCCCcEEEcCCCC-CCCeeEEEEeeCCCeeEEcCCCceEeeeccC
Q 047082 6 NSCSLNSIFASPVCGKPTFSLGSDGNLVLAEADGTVVCQSNTA-NKGVVGFKLLPNGNMVLHDSKGNFIWQSFDC 79 (263)
Q Consensus 6 ~~~vWvANr~~Pv~~~~~l~l~~~G~L~l~~~~g~~~Wst~~~-~~~~~~a~Lld~GNlvl~~~~~~~~WqSFd~ 79 (263)
.++||+|||+.|+.+++.|.|+++|+|+|.+.+|.++|++++. +.+...|+|+|+|||||++..+++|||||||
T Consensus 40 ~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S~t~~~~~~~~~~L~ddGnlvl~~~~~~~~W~Sf~~ 114 (114)
T smart00108 40 RTVVWVANRDNPVSDSCTLTLQSDGNLVLYDGDGRVVWSSNTTGANGNYVLVLLDDGNLVIYDSDGNFLWQSFDY 114 (114)
T ss_pred CcEEEECCCCCCCCCCEEEEEeCCCCEEEEeCCCCEEEEecccCCCCceEEEEeCCCCEEEECCCCCEEeCCCCC
Confidence 5799999999999877899999999999999999999999986 4556789999999999999988999999997
No 5
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=98.39 E-value=1.1e-06 Score=68.12 Aligned_cols=55 Identities=35% Similarity=0.528 Sum_probs=44.9
Q ss_pred CEEEEecCCcEEEEcCC-CcEEEcCCCCCC--CeeEEEEeeCCCeeEEcCCCceEeee
Q 047082 22 PTFSLGSDGNLVLAEAD-GTVVCQSNTANK--GVVGFKLLPNGNMVLHDSKGNFIWQS 76 (263)
Q Consensus 22 ~~l~l~~~G~L~l~~~~-g~~~Wst~~~~~--~~~~a~Lld~GNlvl~~~~~~~~WqS 76 (263)
-.+.++.+|+||+.... +.++|++++... ....+.|.++|||||++.++.++|+|
T Consensus 22 ~~~~~q~dgnlV~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S 79 (114)
T smart00108 22 FTLIMQNDYNLILYKSSSRTVVWVANRDNPVSDSCTLTLQSDGNLVLYDGDGRVVWSS 79 (114)
T ss_pred cccCCCCCEEEEEEECCCCcEEEECCCCCCCCCCEEEEEeCCCCEEEEeCCCCEEEEe
Confidence 35667789999999765 479999998532 23678999999999999888999997
No 6
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=98.38 E-value=5.6e-07 Score=70.06 Aligned_cols=56 Identities=32% Similarity=0.503 Sum_probs=45.1
Q ss_pred EEEEec-CCcEEEEcCC-CcEEEcCCCCC--CCeeEEEEeeCCCeeEEcCCCceEeeecc
Q 047082 23 TFSLGS-DGNLVLAEAD-GTVVCQSNTAN--KGVVGFKLLPNGNMVLHDSKGNFIWQSFD 78 (263)
Q Consensus 23 ~l~l~~-~G~L~l~~~~-g~~~Wst~~~~--~~~~~a~Lld~GNlvl~~~~~~~~WqSFd 78 (263)
.+.++. +|+|++++.. +.++|++++.. .....+.|.++|||||.+.++.++|+|--
T Consensus 23 ~~~~q~~dgnlv~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S~~ 82 (116)
T cd00028 23 KLIMQSRDYNLILYKGSSRTVVWVANRDNPSGSSCTLTLQSDGNLVIYDGSGTVVWSSNT 82 (116)
T ss_pred cCCCCCCeEEEEEEeCCCCeEEEECCCCCCCCCCEEEEEecCCCeEEEcCCCcEEEEecc
Confidence 455666 9999999764 47999999854 24567899999999999988899999643
No 7
>PF08276 PAN_2: PAN-like domain; InterPro: IPR013227 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs
Probab=98.36 E-value=2.8e-07 Score=64.74 Aligned_cols=25 Identities=32% Similarity=1.071 Sum_probs=23.1
Q ss_pred CCCCHHHHHHHhhccCCeEEEEecc
Q 047082 235 TAIKVEDCGRKCTSDCKCSGYFYHQ 259 (263)
Q Consensus 235 ~~~s~~~C~~~Cl~nCsC~a~~y~~ 259 (263)
...++++|+++||+||||+||+|.+
T Consensus 25 ~~~s~~~C~~~Cl~nCsC~Ayay~~ 49 (66)
T PF08276_consen 25 SSVSLEECEKACLSNCSCTAYAYSN 49 (66)
T ss_pred cCCCHHHHHhhcCCCCCEeeEEeec
Confidence 4589999999999999999999985
No 8
>PF01453 B_lectin: D-mannose binding lectin; InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]: Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity. Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=98.08 E-value=2.6e-05 Score=60.65 Aligned_cols=73 Identities=25% Similarity=0.202 Sum_probs=50.1
Q ss_pred CcEEEEc-CCCCCCCCCCEEEEecCCcEEEEcCCCcEEEcCCCCCCCeeEEEEee--CCCeeEEcCCCceEeeeccCC
Q 047082 6 NSCSLNS-IFASPVCGKPTFSLGSDGNLVLAEADGTVVCQSNTANKGVVGFKLLP--NGNMVLHDSKGNFIWQSFDCP 80 (263)
Q Consensus 6 ~~~vWvA-Nr~~Pv~~~~~l~l~~~G~L~l~~~~g~~~Wst~~~~~~~~~a~Lld--~GNlvl~~~~~~~~WqSFd~P 80 (263)
.++||.. +........+.+.|..+|||||.|..+.++|++... ..-+.+.+++ .||++ ......++|.|=..|
T Consensus 38 ~~~iWss~~t~~~~~~~~~~~L~~~GNlvl~d~~~~~lW~Sf~~-ptdt~L~~q~l~~~~~~-~~~~~~~sw~s~~dp 113 (114)
T PF01453_consen 38 GSVIWSSNNTSGRGNSGCYLVLQDDGNLVLYDSSGNVLWQSFDY-PTDTLLPGQKLGDGNVT-GKNDSLTSWSSNTDP 113 (114)
T ss_dssp TEEEEE--S-TTSS-SSEEEEEETTSEEEEEETTSEEEEESTTS-SS-EEEEEET--TSEEE-EESTSSEEEESS---
T ss_pred CCEEEEecccCCccccCeEEEEeCCCCEEEEeecceEEEeecCC-CccEEEeccCcccCCCc-cccceEEeECCCCCC
Confidence 4679999 434333346889999999999999999999999542 2335566677 88888 654567899876665
No 9
>cd01098 PAN_AP_plant Plant PAN/APPLE-like domain; present in plant S-receptor protein kinases and secreted glycoproteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions. S-receptor protein kinases and S-locus glycoproteins are involved in sporophytic self-incompatibility response in Brassica, one of probably many molecular mechanisms, by which hermaphrodite flowering plants avoid self-fertilization.
Probab=98.06 E-value=3e-06 Score=61.50 Aligned_cols=27 Identities=41% Similarity=1.026 Sum_probs=23.9
Q ss_pred CCCCHHHHHHHhhccCCeEEEEeccCC
Q 047082 235 TAIKVEDCGRKCTSDCKCSGYFYHQET 261 (263)
Q Consensus 235 ~~~s~~~C~~~Cl~nCsC~a~~y~~~~ 261 (263)
...++++|+++||+||+|+||+|.+++
T Consensus 30 ~~~s~~~C~~~Cl~nCsC~a~~~~~~~ 56 (84)
T cd01098 30 TAISLEECREACLSNCSCTAYAYNNGS 56 (84)
T ss_pred ccCCHHHHHHHHhcCCCcceeeecCCC
Confidence 457999999999999999999998643
No 10
>cd00129 PAN_APPLE PAN/APPLE-like domain; present in N-terminal (N) domains of plasminogen/ hepatocyte growth factor proteins, plasma prekallikrein/coagulation factor XI and microneme antigen proteins, plant receptor-like protein kinases, and various nematode and leech anti-platelet proteins. Common structural features include two disulfide bonds that link the alpha-helix to the central region of the protein. PAN domains have significant functional versatility, fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=97.77 E-value=2.1e-05 Score=57.32 Aligned_cols=25 Identities=16% Similarity=0.712 Sum_probs=22.4
Q ss_pred CCCHHHHHHHhhc---cCCeEEEEeccC
Q 047082 236 AIKVEDCGRKCTS---DCKCSGYFYHQE 260 (263)
Q Consensus 236 ~~s~~~C~~~Cl~---nCsC~a~~y~~~ 260 (263)
..++++|+++|++ ||||.||+|.+.
T Consensus 24 ~~s~~eC~~~Cl~~~~nCsC~Aya~~~~ 51 (80)
T cd00129 24 ANTADECANRCEKNGLPFSCKAFVFAKA 51 (80)
T ss_pred ccCHHHHHHHHhcCCCCCCceeeeccCC
Confidence 3789999999999 999999999653
No 11
>smart00473 PAN_AP divergent subfamily of APPLE domains. Apple-like domains present in Plasminogen, C. elegans hypothetical ORFs and the extracellular portion of plant receptor-like protein kinases. Predicted to possess protein- and/or carbohydrate-binding functions.
Probab=96.55 E-value=0.0029 Score=44.50 Aligned_cols=25 Identities=28% Similarity=0.975 Sum_probs=22.7
Q ss_pred CCCCHHHHHHHhhc-cCCeEEEEecc
Q 047082 235 TAIKVEDCGRKCTS-DCKCSGYFYHQ 259 (263)
Q Consensus 235 ~~~s~~~C~~~Cl~-nCsC~a~~y~~ 259 (263)
...++++|++.|++ +|+|.||.|..
T Consensus 23 ~~~s~~~C~~~C~~~~~~C~s~~y~~ 48 (78)
T smart00473 23 SVASLEECASKCLNSNCSCRSFTYNN 48 (78)
T ss_pred cCCCHHHHHHHhCCCCCceEEEEEcC
Confidence 35799999999999 99999999975
No 12
>cd01100 APPLE_Factor_XI_like Subfamily of PAN/APPLE-like domains; present in plasma prekallikrein/coagulation factor XI, microneme antigen proteins, and a few prokaryotic proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=96.28 E-value=0.0045 Score=43.94 Aligned_cols=26 Identities=31% Similarity=0.657 Sum_probs=23.0
Q ss_pred CCCCHHHHHHHhhccCCeEEEEeccC
Q 047082 235 TAIKVEDCGRKCTSDCKCSGYFYHQE 260 (263)
Q Consensus 235 ~~~s~~~C~~~Cl~nCsC~a~~y~~~ 260 (263)
...+.++|++.|+.+|+|.||.|..+
T Consensus 23 ~~~s~~~Cq~~C~~~~~C~afT~~~~ 48 (73)
T cd01100 23 FASSAEQCQAACTADPGCLAFTYNTK 48 (73)
T ss_pred ecCCHHHHHHHcCCCCCceEEEEECC
Confidence 34689999999999999999999754
No 13
>PF00024 PAN_1: PAN domain This Prosite entry concerns apple domains, a subset of PAN domains; InterPro: IPR003014 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs It has been shown that, the N-terminal N domains of members of the plasminogen/hepatocyte growth factor family, the apple domains of the plasma prekallikrein/coagulation factor XI family, and domains of various nematode proteins belong to the same module superfamily, the PAN module []. PAN contains a conserved core of three disulphide bridges. In some members of the family there is an additional fourth disulphide bridge that links the N and C termini of the domain.; PDB: 1GP9_C 2QJ2_B 1GMO_H 1NK1_B 3MKP_B 1BHT_B 3HN4_A 1GMN_A 3HMS_A 3HMT_B ....
Probab=76.96 E-value=2.8 Score=29.04 Aligned_cols=26 Identities=19% Similarity=0.695 Sum_probs=23.0
Q ss_pred CCCHHHHHHHhhccCC-eEEEEeccCC
Q 047082 236 AIKVEDCGRKCTSDCK-CSGYFYHQET 261 (263)
Q Consensus 236 ~~s~~~C~~~Cl~nCs-C~a~~y~~~~ 261 (263)
..++++|.+.|+.+=. |.+|.|...+
T Consensus 22 v~s~~~C~~~C~~~~~~C~s~~y~~~~ 48 (79)
T PF00024_consen 22 VPSLEECAQLCLNEPRRCKSFNYDPSS 48 (79)
T ss_dssp ESSHHHHHHHHHHSTT-ESEEEEETTT
T ss_pred CCCHHHHHhhcCcCcccCCeEEEECCC
Confidence 4599999999999999 9999997653
No 14
>smart00223 APPLE APPLE domain. Four-fold repeat in plasma kallikrein and coagulation factor XI. Factor XI apple 3 mediates binding to platelets. Factor XI apple 1 binds high-molecular-mass kininogen. Apple 4 in factor XI mediates dimer formation and binds to factor XIIa. Mutations in apple 4 cause factor XI deficiency, an inherited bleeding disorder.
Probab=76.24 E-value=3.2 Score=29.92 Aligned_cols=32 Identities=16% Similarity=0.396 Sum_probs=26.0
Q ss_pred cccCCCCCHHHHHHHhhccCCeEEEEeccCCC
Q 047082 231 YTSGTAIKVEDCGRKCTSDCKCSGYFYHQETS 262 (263)
Q Consensus 231 ~~~~~~~s~~~C~~~Cl~nCsC~a~~y~~~~~ 262 (263)
.......+.++|++.|..+=.|.+|.|...+.
T Consensus 16 l~~~~~~~~~~Cq~~Ct~~~~C~~FTf~~~~~ 47 (79)
T smart00223 16 INTVYVPSAQVCQKRCTSHPRCLFFTFSTNEP 47 (79)
T ss_pred eeeeecCCHHHHHHhhcCCCCccEEEeeCCCC
Confidence 33344579999999999999999999976654
No 15
>PF07354 Sp38: Zona-pellucida-binding protein (Sp38); InterPro: IPR010857 This family contains a number of zona-pellucida-binding proteins that seem to be restricted to mammals. These are sperm proteins that bind to the 90 kDa family of zona pellucida glycoproteins in a calcium-dependent manner []. These represent some of the specific molecules that mediate the first steps of gamete interaction, allowing fertilisation to occur [].; GO: 0007339 binding of sperm to zona pellucida, 0005576 extracellular region
Probab=75.89 E-value=4.2 Score=36.06 Aligned_cols=36 Identities=14% Similarity=0.243 Sum_probs=33.0
Q ss_pred CCCCCcEEEEcCCCCCCCCCCEEEEecCCcEEEEcC
Q 047082 2 EYPKNSCSLNSIFASPVCGKPTFSLGSDGNLVLAEA 37 (263)
Q Consensus 2 ~~~~~~~vWvANr~~Pv~~~~~l~l~~~G~L~l~~~ 37 (263)
|+...+..|+--.+.++++++.+.|++.|.|++.+-
T Consensus 9 E~iDP~y~W~GP~g~~l~gn~~~nIT~TG~L~~~~F 44 (271)
T PF07354_consen 9 ELIDPTYLWTGPNGKPLSGNSYVNITETGKLMFKNF 44 (271)
T ss_pred ccCCCceEEECCCCcccCCCCeEEEccCceEEeecc
Confidence 677889999999999999999999999999999764
No 16
>PF01436 NHL: NHL repeat; InterPro: IPR001258 The NHL repeat, named after NCL-1, HT2A and Lin-41, is found largely in a large number of eukaryotic and prokaryotic proteins. For example, the repeat is found in a variety of enzymes of the copper type II, ascorbate-dependent monooxygenase family which catalyse the C terminus alpha-amidation of biological peptides []. In many it occurs in tandem arrays, for example in the ringfinger beta-box, coiled-coil (RBCC) eukaryotic growth regulators []. The 'Brain Tumor' protein (Brat) is one such growth regulator that contains a 6-bladed NHL-repeat beta-propeller [, ]. The NHL repeats are also found in serine/threonine protein kinase (STPK) in diverse range of pathogenic bacteria. These STPK are transmembrane receptors with a intracellular N-terminal kinase domain and extracellular C-terminal sensor domain. In the STPK, PknD, from Mycobacterium tuberculosis, the sensor domain forms a rigid, six-bladed b-propeller composed of NHL repeats with a flexible tether to the transmembrane domain.; GO: 0005515 protein binding; PDB: 3FVZ_A 3FW0_A 1RWL_A 1RWI_A 1Q7F_A.
Probab=74.81 E-value=5.9 Score=22.36 Aligned_cols=21 Identities=29% Similarity=0.449 Sum_probs=16.6
Q ss_pred EEEEecCCcEEEEcCCCcEEE
Q 047082 23 TFSLGSDGNLVLAEADGTVVC 43 (263)
Q Consensus 23 ~l~l~~~G~L~l~~~~g~~~W 43 (263)
-+.++.+|+|++.|..+.-||
T Consensus 6 gvav~~~g~i~VaD~~n~rV~ 26 (28)
T PF01436_consen 6 GVAVDSDGNIYVADSGNHRVQ 26 (28)
T ss_dssp EEEEETTSEEEEEECCCTEEE
T ss_pred EEEEeCCCCEEEEECCCCEEE
Confidence 477888899999988776665
No 17
>PF14295 PAN_4: PAN domain; PDB: 2YIL_E 2YIP_C 2YIO_A.
Probab=73.12 E-value=4 Score=25.88 Aligned_cols=25 Identities=28% Similarity=0.695 Sum_probs=17.9
Q ss_pred CCCCHHHHHHHhhccCCeEEEEecc
Q 047082 235 TAIKVEDCGRKCTSDCKCSGYFYHQ 259 (263)
Q Consensus 235 ~~~s~~~C~~~Cl~nCsC~a~~y~~ 259 (263)
...+.++|.++|..+=.|.+|.|..
T Consensus 14 ~~~s~~~C~~~C~~~~~C~~~~~~~ 38 (51)
T PF14295_consen 14 TASSPEECQAACAADPGCQAFTFNP 38 (51)
T ss_dssp ----HHHHHHHHHTSTT--EEEEET
T ss_pred cCCCHHHHHHHccCCCCCCEEEEEC
Confidence 4568999999999999999999976
No 18
>cd05845 Ig2_L1-CAM_like Second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM) and similar proteins. Ig2_L1-CAM_like: domain similar to the second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM). L1 belongs to the L1 subfamily of cell adhesion molecules (CAMs) and is comprised of an extracellular region having six Ig-like domains, five fibronectin type III domains, a transmembrane region and an intracellular domain. L1 is primarily expressed in the nervous system and is involved in its development and function. L1 is associated with an X-linked recessive disorder, X-linked hydrocephalus, MASA syndrome, or spastic paraplegia type 1, that involves abnormalities of axonal growth.
Probab=66.98 E-value=10 Score=28.25 Aligned_cols=33 Identities=18% Similarity=0.100 Sum_probs=22.3
Q ss_pred CCCCcEEEEcCCCCCCCCCCEEEEecCCcEEEE
Q 047082 3 YPKNSCSLNSIFASPVCGKPTFSLGSDGNLVLA 35 (263)
Q Consensus 3 ~~~~~~vWvANr~~Pv~~~~~l~l~~~G~L~l~ 35 (263)
+|+.++.|+-+....+.....+.++.+|+|.+.
T Consensus 31 ~P~P~i~W~~~~~~~i~~~~Ri~~~~~GnL~fs 63 (95)
T cd05845 31 AVPLRIYWMNSDLLHITQDERVSMGQNGNLYFA 63 (95)
T ss_pred CCCCEEEEECCCCccccccccEEECCCceEEEE
Confidence 567788888544444554567777777888774
No 19
>cd01099 PAN_AP_HGF Subfamily of PAN/APPLE-like domains; present in N-terminal (N) domains of plasminogen/hepatocyte growth factor proteins, and various proteins found in Bilateria, such as leech anti-platelet proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=66.50 E-value=7.5 Score=27.78 Aligned_cols=25 Identities=28% Similarity=0.761 Sum_probs=21.9
Q ss_pred CCCHHHHHHHhhc--cCCeEEEEeccC
Q 047082 236 AIKVEDCGRKCTS--DCKCSGYFYHQE 260 (263)
Q Consensus 236 ~~s~~~C~~~Cl~--nCsC~a~~y~~~ 260 (263)
..++++|.++|++ +=.|.+|.|...
T Consensus 24 ~~s~~~C~~~C~~~~~f~CrSf~y~~~ 50 (80)
T cd01099 24 VASLEECLRKCLEETEFTCRSFNYNYK 50 (80)
T ss_pred cCCHHHHHHHhCCCCCceEeEEEEEcC
Confidence 4799999999999 899999988654
No 20
>PF13360 PQQ_2: PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=65.30 E-value=14 Score=31.07 Aligned_cols=73 Identities=18% Similarity=0.236 Sum_probs=43.9
Q ss_pred CCCcEEEEcCC----CCCC----CCCCEEEE-ecCCcEEEEcC-CCcEEEcCCCCCCCe-------eEEEEe-eCCCeeE
Q 047082 4 PKNSCSLNSIF----ASPV----CGKPTFSL-GSDGNLVLAEA-DGTVVCQSNTANKGV-------VGFKLL-PNGNMVL 65 (263)
Q Consensus 4 ~~~~~vWvANr----~~Pv----~~~~~l~l-~~~G~L~l~~~-~g~~~Wst~~~~~~~-------~~a~Ll-d~GNlvl 65 (263)
.....+|..+- ..++ .....|.+ +.+|.|+..|. .|..+|+........ ..+.+. .+|.|+.
T Consensus 11 ~tG~~~W~~~~~~~~~~~~~~~~~~~~~v~~~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~ 90 (238)
T PF13360_consen 11 RTGKELWSYDLGPGIGGPVATAVPDGGRVYVASGDGNLYALDAKTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYA 90 (238)
T ss_dssp TTTEEEEEEECSSSCSSEEETEEEETTEEEEEETTSEEEEEETTTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEE
T ss_pred CCCCEEEEEECCCCCCCccceEEEeCCEEEEEcCCCEEEEEECCCCCEEEEeeccccccceeeecccccccccceeeeEe
Confidence 45678888753 2222 12343444 48899999996 899999987632210 111222 2344666
Q ss_pred Ec-CCCceEeee
Q 047082 66 HD-SKGNFIWQS 76 (263)
Q Consensus 66 ~~-~~~~~~WqS 76 (263)
.| .+++++|+.
T Consensus 91 ~d~~tG~~~W~~ 102 (238)
T PF13360_consen 91 LDAKTGKVLWSI 102 (238)
T ss_dssp EETTTSCEEEEE
T ss_pred cccCCcceeeee
Confidence 66 578999995
No 21
>cd00053 EGF Epidermal growth factor domain, found in epidermal growth factor (EGF) presents in a large number of proteins, mostly animal; the list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied; the functional significance of EGF-like domains in what appear to be unrelated proteins is not yet clear; a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase); the domain includes six cysteine residues which have been shown to be involved in disulfide bonds; the main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet; Subdomains between the conserved cysteines vary in length; the region between the 5th and 6th cysteine contains two conserved glycines of which at least one is present in most EGF-like domains; a subset of these bind calcium.
Probab=62.40 E-value=7.9 Score=21.94 Aligned_cols=29 Identities=24% Similarity=0.505 Sum_probs=20.6
Q ss_pred CCCCCCCCCCcccCCC-CCcCCCCCCCCcc
Q 047082 190 TQLPERCSKLGVCDDN-QCVACPTEKGLLG 218 (263)
Q Consensus 190 C~~~~~CG~~g~C~~~-~~~~C~c~~g~~~ 218 (263)
|.....|...++|... ....|.|+.||..
T Consensus 2 C~~~~~C~~~~~C~~~~~~~~C~C~~g~~g 31 (36)
T cd00053 2 CAASNPCSNGGTCVNTPGSYRCVCPPGYTG 31 (36)
T ss_pred CCCCCCCCCCCEEecCCCCeEeECCCCCcc
Confidence 4435678888999643 4568999998854
No 22
>TIGR03066 Gem_osc_para_1 Gemmata obscuriglobus paralogous family TIGR03066. This model represents an uncharacterized paralogous family in Gemmata obscuriglobus UQM 2246, a member of the Planctomycetes. This family shows sequence similarity to TIGR03067, which is also found in Gemmata obscuriglobus as well as in a few other species.
Probab=61.49 E-value=27 Score=26.97 Aligned_cols=55 Identities=18% Similarity=0.272 Sum_probs=32.8
Q ss_pred CCCCCEEEEecCCcEEEEcCCCcE------EEcCC---------CCCC---CeeEEEEeeCCCeeEEcCCCce
Q 047082 18 VCGKPTFSLGSDGNLVLAEADGTV------VCQSN---------TANK---GVVGFKLLPNGNMVLHDSKGNF 72 (263)
Q Consensus 18 v~~~~~l~l~~~G~L~l~~~~g~~------~Wst~---------~~~~---~~~~a~Lld~GNlvl~~~~~~~ 72 (263)
+.....|.|..+|.|+|+.++++- -|+-. ..+. .-..-.-+++|.|||.|++++.
T Consensus 32 ~~~~~~leF~~dGKL~v~~gnng~~~~~~Gty~L~G~kLtL~~~p~g~t~k~~Vtv~~l~~~~Lvl~d~dg~~ 104 (111)
T TIGR03066 32 TKDDVVIEFAKDGKLVVTIGEKGKEVKADGTYKLDGNKLTLTLKAGGKEKKETLTVKKLTDDELVGKDPDGKK 104 (111)
T ss_pred eCCceEEEEcCCCeEEEecCCCCcEeccCceEEEECCEEEEEEcCCCccccceEEEEEecCCeEEEEcCCCCE
Confidence 334678999999999998775442 12211 0011 1011123688899999887653
No 23
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=60.93 E-value=35 Score=31.64 Aligned_cols=56 Identities=23% Similarity=0.370 Sum_probs=34.3
Q ss_pred CCEEEEe-cCCcEEEEcC-CCcEEEcCCCCCC----Ce----eEEEEeeCCCeeEEcC-CCceEeee
Q 047082 21 KPTFSLG-SDGNLVLAEA-DGTVVCQSNTANK----GV----VGFKLLPNGNMVLHDS-KGNFIWQS 76 (263)
Q Consensus 21 ~~~l~l~-~~G~L~l~~~-~g~~~Wst~~~~~----~~----~~a~Lld~GNlvl~~~-~~~~~WqS 76 (263)
...|.+. .+|.|+-+|. +|..+|+...... ++ ....-..+|.|+-.|. +++.+|+-
T Consensus 120 ~~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~ssP~v~~~~v~v~~~~g~l~ald~~tG~~~W~~ 186 (394)
T PRK11138 120 GGKVYIGSEKGQVYALNAEDGEVAWQTKVAGEALSRPVVSDGLVLVHTSNGMLQALNESDGAVKWTV 186 (394)
T ss_pred CCEEEEEcCCCEEEEEECCCCCCcccccCCCceecCCEEECCEEEEECCCCEEEEEEccCCCEeeee
Confidence 3444444 6688887775 6899999876431 11 1112234566777775 68899964
No 24
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=58.89 E-value=23 Score=32.80 Aligned_cols=19 Identities=21% Similarity=0.338 Sum_probs=10.1
Q ss_pred eeCCCeeEEcC-CCceEeee
Q 047082 58 LPNGNMVLHDS-KGNFIWQS 76 (263)
Q Consensus 58 ld~GNlvl~~~-~~~~~WqS 76 (263)
-++|.|...|. +++++|+-
T Consensus 342 ~~~G~l~~ld~~tG~~~~~~ 361 (394)
T PRK11138 342 DSEGYLHWINREDGRFVAQQ 361 (394)
T ss_pred eCCCEEEEEECCCCCEEEEE
Confidence 34555555553 45666653
No 25
>PF13360 PQQ_2: PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=50.66 E-value=63 Score=26.90 Aligned_cols=73 Identities=18% Similarity=0.321 Sum_probs=44.5
Q ss_pred CCCcEEEEcCCCCCCC-----CCC-EEEEecCCcEEEEc-CCCcEEEcC-CCC----C---CCe-e----EEEE-eeCCC
Q 047082 4 PKNSCSLNSIFASPVC-----GKP-TFSLGSDGNLVLAE-ADGTVVCQS-NTA----N---KGV-V----GFKL-LPNGN 62 (263)
Q Consensus 4 ~~~~~vWvANr~~Pv~-----~~~-~l~l~~~G~L~l~~-~~g~~~Wst-~~~----~---~~~-~----~a~L-ld~GN 62 (263)
..-+.+|....+.++. ... .+..+.+|.|...| .+|.++|.. ... . ... + .+.+ ..+|.
T Consensus 54 ~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~g~ 133 (238)
T PF13360_consen 54 KTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYALDAKTGKVLWSIYLTSSPPAGVRSSSSPAVDGDRLYVGTSSGK 133 (238)
T ss_dssp TTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEEEETTTSCEEEEEEE-SSCTCSTB--SEEEEETTEEEEEETCSE
T ss_pred CCCCEEEEeeccccccceeeecccccccccceeeeEecccCCcceeeeeccccccccccccccCceEecCEEEEEeccCc
Confidence 4567788877655532 233 44455678888888 789999994 321 1 010 0 1222 33788
Q ss_pred eeEEcC-CCceEeee
Q 047082 63 MVLHDS-KGNFIWQS 76 (263)
Q Consensus 63 lvl~~~-~~~~~WqS 76 (263)
|+..|. +++.+|+-
T Consensus 134 l~~~d~~tG~~~w~~ 148 (238)
T PF13360_consen 134 LVALDPKTGKLLWKY 148 (238)
T ss_dssp EEEEETTTTEEEEEE
T ss_pred EEEEecCCCcEEEEe
Confidence 888884 68899964
No 26
>KOG0640 consensus mRNA cleavage stimulating factor complex; subunit 1 [RNA processing and modification]
Probab=45.73 E-value=1.1e+02 Score=28.19 Aligned_cols=71 Identities=18% Similarity=0.350 Sum_probs=49.7
Q ss_pred CCcEEEEcCCCCCCCCC-CEEEEecCCcEEEEcC-CCc-EEEcCCCC-----------CCCeeEEEEeeCCCeeEEcCCC
Q 047082 5 KNSCSLNSIFASPVCGK-PTFSLGSDGNLVLAEA-DGT-VVCQSNTA-----------NKGVVGFKLLPNGNMVLHDSKG 70 (263)
Q Consensus 5 ~~~~vWvANr~~Pv~~~-~~l~l~~~G~L~l~~~-~g~-~~Wst~~~-----------~~~~~~a~Lld~GNlvl~~~~~ 70 (263)
..-+-=.||.+..+++. ..+.-++.|+|+++.+ +|. -+|.-... +..+.+|++-.+|.++|....+
T Consensus 247 T~QcfvsanPd~qht~ai~~V~Ys~t~~lYvTaSkDG~IklwDGVS~rCv~t~~~AH~gsevcSa~Ftkn~kyiLsSG~D 326 (430)
T KOG0640|consen 247 TYQCFVSANPDDQHTGAITQVRYSSTGSLYVTASKDGAIKLWDGVSNRCVRTIGNAHGGSEVCSAVFTKNGKYILSSGKD 326 (430)
T ss_pred ceeEeeecCcccccccceeEEEecCCccEEEEeccCCcEEeeccccHHHHHHHHhhcCCceeeeEEEccCCeEEeecCCc
Confidence 33445568888777775 6788999999999855 344 47874321 2336789999999999986533
Q ss_pred c--eEee
Q 047082 71 N--FIWQ 75 (263)
Q Consensus 71 ~--~~Wq 75 (263)
. -||+
T Consensus 327 S~vkLWE 333 (430)
T KOG0640|consen 327 STVKLWE 333 (430)
T ss_pred ceeeeee
Confidence 3 4786
No 27
>PLN00033 photosystem II stability/assembly factor; Provisional
Probab=44.70 E-value=73 Score=30.14 Aligned_cols=52 Identities=17% Similarity=0.248 Sum_probs=31.7
Q ss_pred EEEEecCCcEEEEcCCCcEEEcCCCCC--CCeeEEEEeeCCCeeEEcCCCceEe
Q 047082 23 TFSLGSDGNLVLAEADGTVVCQSNTAN--KGVVGFKLLPNGNMVLHDSKGNFIW 74 (263)
Q Consensus 23 ~l~l~~~G~L~l~~~~g~~~Wst~~~~--~~~~~a~Lld~GNlvl~~~~~~~~W 74 (263)
.+.+...|++++.+.+|...|...... ...+.+...++|.|+|....+.++|
T Consensus 252 ~~~vg~~G~~~~s~d~G~~~W~~~~~~~~~~l~~v~~~~dg~l~l~g~~G~l~~ 305 (398)
T PLN00033 252 YVAVSSRGNFYLTWEPGQPYWQPHNRASARRIQNMGWRADGGLWLLTRGGGLYV 305 (398)
T ss_pred EEEEECCccEEEecCCCCcceEEecCCCccceeeeeEcCCCCEEEEeCCceEEE
Confidence 444445556555555566667744322 2334556688999999887666555
No 28
>PF07974 EGF_2: EGF-like domain; InterPro: IPR013111 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length. This entry contains EGF domains found in a variety of extracellular and membrane proteins
Probab=43.71 E-value=19 Score=21.25 Aligned_cols=22 Identities=32% Similarity=0.761 Sum_probs=17.3
Q ss_pred CCCCCCcccCCCCCcCCCCCCCC
Q 047082 194 ERCSKLGVCDDNQCVACPTEKGL 216 (263)
Q Consensus 194 ~~CG~~g~C~~~~~~~C~c~~g~ 216 (263)
.+|...|+|+.. ...|.|.+||
T Consensus 6 ~~C~~~G~C~~~-~g~C~C~~g~ 27 (32)
T PF07974_consen 6 NICSGHGTCVSP-CGRCVCDSGY 27 (32)
T ss_pred CccCCCCEEeCC-CCEEECCCCC
Confidence 479999999765 3379999886
No 29
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=43.14 E-value=1.4e+02 Score=27.21 Aligned_cols=72 Identities=11% Similarity=0.197 Sum_probs=42.3
Q ss_pred CCCcEEEEcCCCC-----CCCCCCEEEE-ecCCcEEEEcC-CCcEEEcCCCCCC----Ce----eEEEEeeCCCeeEEcC
Q 047082 4 PKNSCSLNSIFAS-----PVCGKPTFSL-GSDGNLVLAEA-DGTVVCQSNTANK----GV----VGFKLLPNGNMVLHDS 68 (263)
Q Consensus 4 ~~~~~vWvANr~~-----Pv~~~~~l~l-~~~G~L~l~~~-~g~~~Wst~~~~~----~~----~~a~Lld~GNlvl~~~ 68 (263)
..-..+|.-+-.. |+.+...+.+ +.+|.|+.+|. +|..+|+...... .. ....-..+|.|+..|.
T Consensus 83 ~tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~a~d~ 162 (377)
T TIGR03300 83 ETGKRLWRVDLDERLSGGVGADGGLVFVGTEKGEVIALDAEDGKELWRAKLSSEVLSPPLVANGLVVVRTNDGRLTALDA 162 (377)
T ss_pred cCCcEeeeecCCCCcccceEEcCCEEEEEcCCCEEEEEECCCCcEeeeeccCceeecCCEEECCEEEEECCCCeEEEEEc
Confidence 3456788655433 3333444444 46788888886 6889998765321 11 1111234567777775
Q ss_pred -CCceEee
Q 047082 69 -KGNFIWQ 75 (263)
Q Consensus 69 -~~~~~Wq 75 (263)
+++.+|+
T Consensus 163 ~tG~~~W~ 170 (377)
T TIGR03300 163 ATGERLWT 170 (377)
T ss_pred CCCceeeE
Confidence 6788996
No 30
>PF09064 Tme5_EGF_like: Thrombomodulin like fifth domain, EGF-like; InterPro: IPR015149 This domain adopts a fold similar to other EGF domains, with a flat major and a twisted minor beta sheet. Disulphide pairing, however, is not of the usual 1-3, 2-4, 5-6 type; rather 1-2, 3-4, 5-6 pairing is found. Its extended major sheet (strands beta-2 and beta-3 and the connecting loop) projects into thrombin's active site groove. This domain is required for interaction of thrombomodulin with thrombin, and subsequent activation of protein-C []. ; GO: 0004888 transmembrane signaling receptor activity, 0016021 integral to membrane
Probab=41.53 E-value=16 Score=21.99 Aligned_cols=16 Identities=31% Similarity=0.511 Sum_probs=11.7
Q ss_pred ccCCCCCcCCCCCCCC
Q 047082 201 VCDDNQCVACPTEKGL 216 (263)
Q Consensus 201 ~C~~~~~~~C~c~~g~ 216 (263)
.|+.+...+|.||.||
T Consensus 11 ~CDpn~~~~C~CPeGy 26 (34)
T PF09064_consen 11 DCDPNSPGQCFCPEGY 26 (34)
T ss_pred ccCCCCCCceeCCCce
Confidence 4665555589999987
No 31
>PF08277 PAN_3: PAN-like domain; InterPro: IPR006583 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs The PAN-3 or CW is a domain associated with a number of Caenorhabditis elegans hypothetical proteins.
Probab=41.30 E-value=36 Score=23.17 Aligned_cols=24 Identities=29% Similarity=0.652 Sum_probs=21.3
Q ss_pred CCCCHHHHHHHhhccCCeEEEEec
Q 047082 235 TAIKVEDCGRKCTSDCKCSGYFYH 258 (263)
Q Consensus 235 ~~~s~~~C~~~Cl~nCsC~a~~y~ 258 (263)
...+.++|-..|..+=.|+.+.+.
T Consensus 18 ~~~sw~~Cv~~C~~~~~C~la~~~ 41 (71)
T PF08277_consen 18 TNTSWDDCVQKCYNDENCVLAYFD 41 (71)
T ss_pred cCCCHHHHhHHhCCCCEEEEEEeC
Confidence 347889999999999999998876
No 32
>KOG4649 consensus PQQ (pyrrolo-quinoline quinone) repeat protein [Secondary metabolites biosynthesis, transport and catabolism]
Probab=39.79 E-value=50 Score=29.72 Aligned_cols=43 Identities=16% Similarity=0.198 Sum_probs=32.9
Q ss_pred CcEEEEcCCCCCCCCC-----CEEEE-ecCCcEEEEcCCCcEEEcCCCC
Q 047082 6 NSCSLNSIFASPVCGK-----PTFSL-GSDGNLVLAEADGTVVCQSNTA 48 (263)
Q Consensus 6 ~~~vWvANr~~Pv~~~-----~~l~l-~~~G~L~l~~~~g~~~Wst~~~ 48 (263)
.+.+|.|.|..||-.+ ..+.+ +-||+|.-.|+.|+.||.-.+.
T Consensus 169 ~~~~w~~~~~~PiF~splcv~~sv~i~~VdG~l~~f~~sG~qvwr~~t~ 217 (354)
T KOG4649|consen 169 STEFWAATRFGPIFASPLCVGSSVIITTVDGVLTSFDESGRQVWRPATK 217 (354)
T ss_pred cceehhhhcCCccccCceeccceEEEEEeccEEEEEcCCCcEEEeecCC
Confidence 4678999999998763 33444 4789999889989999976554
No 33
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=39.78 E-value=68 Score=29.30 Aligned_cols=17 Identities=18% Similarity=0.223 Sum_probs=9.9
Q ss_pred eeCCCeeEEcC-CCceEe
Q 047082 58 LPNGNMVLHDS-KGNFIW 74 (263)
Q Consensus 58 ld~GNlvl~~~-~~~~~W 74 (263)
-.+|.|.+.+. +++.+|
T Consensus 327 ~~~G~l~~~d~~tG~~~~ 344 (377)
T TIGR03300 327 DFEGYLHWLSREDGSFVA 344 (377)
T ss_pred eCCCEEEEEECCCCCEEE
Confidence 34566666654 456666
No 34
>cd05852 Ig5_Contactin-1 Fifth Ig domain of contactin-1. Ig5_Contactin-1: fifth Ig domain of the neural cell adhesion molecule contactin-1. Contactins are comprised of six Ig domains followed by four fibronectin type III (FnIII) domains anchored to the membrane by glycosylphosphatidylinositol. Contactin-1 is differentially expressed in tumor tissues and may through a RhoA mechanism, facilitate invasion and metastasis of human lung adenocarcinoma.
Probab=39.05 E-value=35 Score=23.57 Aligned_cols=33 Identities=21% Similarity=0.236 Sum_probs=22.3
Q ss_pred CCCCcEEEEcCCCCCCCCCCEEEEecCCcEEEEc
Q 047082 3 YPKNSCSLNSIFASPVCGKPTFSLGSDGNLVLAE 36 (263)
Q Consensus 3 ~~~~~~vWvANr~~Pv~~~~~l~l~~~G~L~l~~ 36 (263)
.|..++.|.=+. .++..+..+.+..+|.|+|.+
T Consensus 13 ~P~p~v~W~k~~-~~l~~~~r~~~~~~g~L~I~~ 45 (73)
T cd05852 13 APKPKFSWSKGT-ELLVNNSRISIWDDGSLEILN 45 (73)
T ss_pred eCCCEEEEEeCC-EecccCCCEEEcCCCEEEECc
Confidence 467788998654 355555567777778887754
No 35
>smart00605 CW CW domain.
Probab=39.00 E-value=37 Score=24.79 Aligned_cols=23 Identities=22% Similarity=0.574 Sum_probs=19.7
Q ss_pred CCCHHHHHHHhhccCCeEEEEec
Q 047082 236 AIKVEDCGRKCTSDCKCSGYFYH 258 (263)
Q Consensus 236 ~~s~~~C~~~Cl~nCsC~a~~y~ 258 (263)
..+.++|...|..+..|+.+...
T Consensus 21 ~~sw~~Ci~~C~~~~~Cvlay~~ 43 (94)
T smart00605 21 TLSWDECIQKCYEDSNCVLAYGN 43 (94)
T ss_pred CCCHHHHHHHHhCCCceEEEecC
Confidence 46889999999999999987544
No 36
>PF07645 EGF_CA: Calcium-binding EGF domain; InterPro: IPR001881 A sequence of about forty amino-acid residues found in epidermal growth factor (EGF) has been shown [, , , , , ] to be present in a large number of membrane-bound and extracellular, mostly animal, proteins. Many of these proteins require calcium for their biological function and a calcium-binding site has been found at the N terminus of some EGF-like domains []. Calcium-binding may be crucial for numerous protein-protein interactions. For human coagulation factor IX it has been shown [] that the calcium-ligands form a pentagonal bipyramid. The first, third and fourth conserved negatively charged or polar residues are side chain ligands. The latter is possibly hydroxylated (see aspartic acid and asparagine hydroxylation site) []. A conserved aromatic residue, as well as the second conserved negative residue, are thought to be involved in stabilising the calcium-binding site. As in non-calcium binding EGF-like domains, there are six conserved cysteines and the structure of both types is very similar as calcium-binding induces only strictly local structural changes []. +------------------+ +---------+ | | | | nxnnC-x(3,14)-C-x(3,7)-CxxbxxxxaxC-x(1,6)-C-x(8,13)-Cx | | +------------------+ 'n': negatively charged or polar residue [DEQN] 'b': possibly beta-hydroxylated residue [DN] 'a': aromatic amino acid 'C': cysteine, involved in disulphide bond 'x': any amino acid. ; GO: 0005509 calcium ion binding; PDB: 2VJ3_A 1TOZ_A 1LMJ_A 1UZQ_A 1UZK_A 1UZJ_B 1UZP_A 1EMO_A 1EMN_A 2RR0_A ....
Probab=35.69 E-value=9.4 Score=23.67 Aligned_cols=28 Identities=18% Similarity=0.418 Sum_probs=20.8
Q ss_pred CCCC-CCCCCCcccC-CCCCcCCCCCCCCc
Q 047082 190 TQLP-ERCSKLGVCD-DNQCVACPTEKGLL 217 (263)
Q Consensus 190 C~~~-~~CG~~g~C~-~~~~~~C~c~~g~~ 217 (263)
|... ..|..++.|. ......|.|++||.
T Consensus 5 C~~~~~~C~~~~~C~N~~Gsy~C~C~~Gy~ 34 (42)
T PF07645_consen 5 CAEGPHNCPENGTCVNTEGSYSCSCPPGYE 34 (42)
T ss_dssp TTTTSSSSSTTSEEEEETTEEEEEESTTEE
T ss_pred cCCCCCcCCCCCEEEcCCCCEEeeCCCCcE
Confidence 5553 5798899995 34566899999985
No 37
>KOG0291 consensus WD40-repeat-containing subunit of the 18S rRNA processing complex [RNA processing and modification]
Probab=35.28 E-value=5.3e+02 Score=26.75 Aligned_cols=108 Identities=19% Similarity=0.246 Sum_probs=62.0
Q ss_pred CCcEEEEcCCCCCC-------CCCCEEEEecCCcEEEEcC-CCcE-EEcCCC--------CCCCeeEEEEee--CCCeeE
Q 047082 5 KNSCSLNSIFASPV-------CGKPTFSLGSDGNLVLAEA-DGTV-VCQSNT--------ANKGVVGFKLLP--NGNMVL 65 (263)
Q Consensus 5 ~~~~vWvANr~~Pv-------~~~~~l~l~~~G~L~l~~~-~g~~-~Wst~~--------~~~~~~~a~Lld--~GNlvl 65 (263)
.+.-||-..+..-+ ++-..+.++..|+.+|... +|++ +|--.. ...+..-..|.. +|.||.
T Consensus 372 gKVKvWn~~SgfC~vTFteHts~Vt~v~f~~~g~~llssSLDGtVRAwDlkRYrNfRTft~P~p~QfscvavD~sGelV~ 451 (893)
T KOG0291|consen 372 GKVKVWNTQSGFCFVTFTEHTSGVTAVQFTARGNVLLSSSLDGTVRAWDLKRYRNFRTFTSPEPIQFSCVAVDPSGELVC 451 (893)
T ss_pred CcEEEEeccCceEEEEeccCCCceEEEEEEecCCEEEEeecCCeEEeeeecccceeeeecCCCceeeeEEEEcCCCCEEE
Confidence 34678888875532 2225789999999888644 6776 787652 222332333444 499999
Q ss_pred EcCCCc---eEeeeccCCCceeccCcccC---CCCCCe-EEEEecCCceEEEecCCCCCcee
Q 047082 66 HDSKGN---FIWQSFDCPTDTLLVGQSLL---SVKENV-SFVMEPKRFTLYYKGSNSPQPVL 120 (263)
Q Consensus 66 ~~~~~~---~~WqSFd~PTDTlLpGq~l~---S~~dps-sl~l~~~~~~~~~~~~~~~~~~~ 120 (263)
...-+. .+| |+. -||-|- .+.-|. .|.+.+.+-.+....|+.+.+.|
T Consensus 452 AG~~d~F~IfvW-S~q-------TGqllDiLsGHEgPVs~l~f~~~~~~LaS~SWDkTVRiW 505 (893)
T KOG0291|consen 452 AGAQDSFEIFVW-SVQ-------TGQLLDILSGHEGPVSGLSFSPDGSLLASGSWDKTVRIW 505 (893)
T ss_pred eeccceEEEEEE-Eee-------cCeeeehhcCCCCcceeeEEccccCeEEeccccceEEEE
Confidence 865433 477 332 355443 344454 67777766555544444443333
No 38
>smart00179 EGF_CA Calcium-binding EGF-like domain.
Probab=34.90 E-value=36 Score=19.73 Aligned_cols=28 Identities=18% Similarity=0.402 Sum_probs=19.0
Q ss_pred CCCCCCCCCCcccCCC-CCcCCCCCCCCc
Q 047082 190 TQLPERCSKLGVCDDN-QCVACPTEKGLL 217 (263)
Q Consensus 190 C~~~~~CG~~g~C~~~-~~~~C~c~~g~~ 217 (263)
|.....|...+.|... ....|.|++||.
T Consensus 5 C~~~~~C~~~~~C~~~~g~~~C~C~~g~~ 33 (39)
T smart00179 5 CASGNPCQNGGTCVNTVGSYRCECPPGYT 33 (39)
T ss_pred CcCCCCcCCCCEeECCCCCeEeECCCCCc
Confidence 5443568878889642 345799999875
No 39
>cd00216 PQQ_DH Dehydrogenases with pyrrolo-quinoline quinone (PQQ) as cofactor, like ethanol, methanol, and membrane bound glucose dehydrogenases. The alignment model contains an 8-bladed beta-propeller.
Probab=34.00 E-value=1.4e+02 Score=28.75 Aligned_cols=72 Identities=19% Similarity=0.290 Sum_probs=42.3
Q ss_pred CCCcEEEEcCCC-------CCCCCCCEEEEe-cCCcEEEEcC-CCcEEEcCCCCCC-----------Cee----EEEE--
Q 047082 4 PKNSCSLNSIFA-------SPVCGKPTFSLG-SDGNLVLAEA-DGTVVCQSNTANK-----------GVV----GFKL-- 57 (263)
Q Consensus 4 ~~~~~vWvANr~-------~Pv~~~~~l~l~-~~G~L~l~~~-~g~~~Wst~~~~~-----------~~~----~a~L-- 57 (263)
++-.++|...-. .|+-....+-+. .+|.|+-+|. .|.++|+...... +++ ..++
T Consensus 37 ~~~~~~W~~~~~~~~~~~~sPvv~~g~vy~~~~~g~l~AlD~~tG~~~W~~~~~~~~~~~~~~~~~~g~~~~~~~~V~v~ 116 (488)
T cd00216 37 KKLKVAWTFSTGDERGQEGTPLVVDGDMYFTTSHSALFALDAATGKVLWRYDPKLPADRGCCDVVNRGVAYWDPRKVFFG 116 (488)
T ss_pred hcceeeEEEECCCCCCcccCCEEECCEEEEeCCCCcEEEEECCCChhhceeCCCCCccccccccccCCcEEccCCeEEEe
Confidence 345678887654 354444444444 5788887775 5889999764211 100 0111
Q ss_pred eeCCCeeEEcC-CCceEee
Q 047082 58 LPNGNMVLHDS-KGNFIWQ 75 (263)
Q Consensus 58 ld~GNlvl~~~-~~~~~Wq 75 (263)
-.+|.|+-.|. +++.+|+
T Consensus 117 ~~~g~v~AlD~~TG~~~W~ 135 (488)
T cd00216 117 TFDGRLVALDAETGKQVWK 135 (488)
T ss_pred cCCCeEEEEECCCCCEeee
Confidence 13566666665 6899997
No 40
>PF13570 PQQ_3: PQQ-like domain; PDB: 3HXJ_B 3Q54_A.
Probab=31.24 E-value=40 Score=20.21 Aligned_cols=8 Identities=25% Similarity=0.305 Sum_probs=2.8
Q ss_pred cEEEcCCC
Q 047082 40 TVVCQSNT 47 (263)
Q Consensus 40 ~~~Wst~~ 47 (263)
.++|+..+
T Consensus 2 ~~~W~~~~ 9 (40)
T PF13570_consen 2 KVLWSYDT 9 (40)
T ss_dssp -EEEEEE-
T ss_pred ceeEEEEC
Confidence 34444443
No 41
>PF05935 Arylsulfotrans: Arylsulfotransferase (ASST); InterPro: IPR010262 This family consists of several bacterial arylsulphotransferase proteins. Arylsulphotransferase (ASST) transfers a sulphate group from phenolic sulphate esters to a phenolic acceptor substrate [].; PDB: 3ETT_B 3ELQ_A 3ETS_A.
Probab=30.30 E-value=49 Score=31.94 Aligned_cols=52 Identities=29% Similarity=0.514 Sum_probs=30.6
Q ss_pred CCcEEEEcCCCcEEEcCCCCCCCeeEEEEeeCCCeeEEc--------CCCceEeeeccCCC
Q 047082 29 DGNLVLAEADGTVVCQSNTANKGVVGFKLLPNGNMVLHD--------SKGNFIWQSFDCPT 81 (263)
Q Consensus 29 ~G~L~l~~~~g~~~Wst~~~~~~~~~a~Lld~GNlvl~~--------~~~~~~WqSFd~PT 81 (263)
.+..++.|.+|.++|..............+++|+|.... -.|+++|+ ++.|.
T Consensus 127 ~~~~~~iD~~G~Vrw~~~~~~~~~~~~~~l~nG~ll~~~~~~~~e~D~~G~v~~~-~~l~~ 186 (477)
T PF05935_consen 127 SSYTYLIDNNGDVRWYLPLDSGSDNSFKQLPNGNLLIGSGNRLYEIDLLGKVIWE-YDLPG 186 (477)
T ss_dssp EEEEEEEETTS-EEEEE-GGGT--SSEEE-TTS-EEEEEBTEEEEE-TT--EEEE-EE--T
T ss_pred CceEEEECCCccEEEEEccCccccceeeEcCCCCEEEecCCceEEEcCCCCEEEe-eecCC
Confidence 467889999999999987543222226789999987653 35789998 77776
No 42
>PF01683 EB: EB module; InterPro: IPR006149 The EB domain has no known function. It is found in several Caenorhabditis sp. and Drosophila sp. proteins. The domain contains 8 conserved cysteines that probably form four disulphide bridges and is found associated with kunitz domains IPR002223 from INTERPRO
Probab=28.40 E-value=43 Score=21.51 Aligned_cols=27 Identities=22% Similarity=0.492 Sum_probs=21.1
Q ss_pred CCCCCCCCCCCcccCCCCCcCCCCCCCCcc
Q 047082 189 PTQLPERCSKLGVCDDNQCVACPTEKGLLG 218 (263)
Q Consensus 189 ~C~~~~~CG~~g~C~~~~~~~C~c~~g~~~ 218 (263)
.|.....|-.+++|..+ .|.|++||..
T Consensus 21 ~C~~~~qC~~~s~C~~g---~C~C~~g~~~ 47 (52)
T PF01683_consen 21 SCESDEQCIGGSVCVNG---RCQCPPGYVE 47 (52)
T ss_pred CCCCcCCCCCcCEEcCC---EeECCCCCEe
Confidence 39988999999999432 5889998743
No 43
>smart00564 PQQ beta-propeller repeat. Beta-propeller repeat occurring in enzymes with pyrrolo-quinoline quinone (PQQ) as cofactor, in Ire1p-like Ser/Thr kinases, and in prokaryotic dehydrogenases.
Probab=27.36 E-value=1.2e+02 Score=16.86 Aligned_cols=17 Identities=29% Similarity=0.575 Sum_probs=8.8
Q ss_pred cCCcEEEEcC-CCcEEEc
Q 047082 28 SDGNLVLAEA-DGTVVCQ 44 (263)
Q Consensus 28 ~~G~L~l~~~-~g~~~Ws 44 (263)
.+|.|+-.|. +|..+|.
T Consensus 14 ~~g~l~a~d~~~G~~~W~ 31 (33)
T smart00564 14 TDGTLYALDAKTGEILWT 31 (33)
T ss_pred CCCEEEEEEcccCcEEEE
Confidence 3455555544 4555664
No 44
>PF06006 DUF905: Bacterial protein of unknown function (DUF905); InterPro: IPR009253 This family consists of several short hypothetical proteobacterial proteins of unknown function.; PDB: 2HJJ_A.
Probab=24.77 E-value=80 Score=22.20 Aligned_cols=17 Identities=24% Similarity=0.976 Sum_probs=10.1
Q ss_pred eeEEcCCCceEeeeccC
Q 047082 63 MVLHDSKGNFIWQSFDC 79 (263)
Q Consensus 63 lvl~~~~~~~~WqSFd~ 79 (263)
||+|+.++.-+|..|.+
T Consensus 35 lvvRd~~g~mvWRaWNF 51 (70)
T PF06006_consen 35 LVVRDTEGQMVWRAWNF 51 (70)
T ss_dssp EEEE-SS--EEEEEESS
T ss_pred EEEEcCCCcEEEEeecc
Confidence 67777777778877654
No 45
>cd00054 EGF_CA Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.
Probab=24.51 E-value=70 Score=18.03 Aligned_cols=28 Identities=18% Similarity=0.413 Sum_probs=18.0
Q ss_pred CCCCCCCCCCcccCCC-CCcCCCCCCCCc
Q 047082 190 TQLPERCSKLGVCDDN-QCVACPTEKGLL 217 (263)
Q Consensus 190 C~~~~~CG~~g~C~~~-~~~~C~c~~g~~ 217 (263)
|.....|...+.|... ....|.|+.||.
T Consensus 5 C~~~~~C~~~~~C~~~~~~~~C~C~~g~~ 33 (38)
T cd00054 5 CASGNPCQNGGTCVNTVGSYRCSCPPGYT 33 (38)
T ss_pred CCCCCCcCCCCEeECCCCCeEeECCCCCc
Confidence 4433467777788543 344799998874
No 46
>cd00216 PQQ_DH Dehydrogenases with pyrrolo-quinoline quinone (PQQ) as cofactor, like ethanol, methanol, and membrane bound glucose dehydrogenases. The alignment model contains an 8-bladed beta-propeller.
Probab=24.19 E-value=1.7e+02 Score=28.14 Aligned_cols=16 Identities=25% Similarity=0.700 Sum_probs=6.9
Q ss_pred CCCeeEEcC-CCceEee
Q 047082 60 NGNMVLHDS-KGNFIWQ 75 (263)
Q Consensus 60 ~GNlvl~~~-~~~~~Wq 75 (263)
+|.|.-.|. +++.+|+
T Consensus 415 dG~l~ald~~tG~~lW~ 431 (488)
T cd00216 415 DGYFRAFDATTGKELWK 431 (488)
T ss_pred CCeEEEEECCCCceeeE
Confidence 344444432 3445553
No 47
>KOG3881 consensus Uncharacterized conserved protein [Function unknown]
Probab=23.82 E-value=2.9e+02 Score=26.10 Aligned_cols=61 Identities=15% Similarity=0.160 Sum_probs=41.9
Q ss_pred CEEEEecCCcEEEEcCCC--cEEEcCCCCCCCeeEEEEeeCCCeeEEcCCCceEeeeccCCCce
Q 047082 22 PTFSLGSDGNLVLAEADG--TVVCQSNTANKGVVGFKLLPNGNMVLHDSKGNFIWQSFDCPTDT 83 (263)
Q Consensus 22 ~~l~l~~~G~L~l~~~~g--~~~Wst~~~~~~~~~a~Lld~GNlvl~~~~~~~~WqSFd~PTDT 83 (263)
--++++.-|.+.|+|... +||=+..-...+.+...|.-+||+|+....... --+||+-+--
T Consensus 218 ~fat~T~~hqvR~YDt~~qRRPV~~fd~~E~~is~~~l~p~gn~Iy~gn~~g~-l~~FD~r~~k 280 (412)
T KOG3881|consen 218 KFATITRYHQVRLYDTRHQRRPVAQFDFLENPISSTGLTPSGNFIYTGNTKGQ-LAKFDLRGGK 280 (412)
T ss_pred eEEEEecceeEEEecCcccCcceeEeccccCcceeeeecCCCcEEEEecccch-hheecccCce
Confidence 467888889999998642 466555544455677889999999988543222 3589976543
No 48
>PF10636 hemP: Hemin uptake protein hemP; InterPro: IPR019600 This entry represents bacterial proteins that are involved in the uptake of the iron source hemin []. ; PDB: 2JRA_B 2LOJ_A.
Probab=23.44 E-value=90 Score=19.29 Aligned_cols=13 Identities=23% Similarity=0.593 Sum_probs=8.5
Q ss_pred EEEEecCCcEEEE
Q 047082 23 TFSLGSDGNLVLA 35 (263)
Q Consensus 23 ~l~l~~~G~L~l~ 35 (263)
.|.++..|.|+|+
T Consensus 25 ~LR~Tr~gKLILT 37 (38)
T PF10636_consen 25 RLRITRQGKLILT 37 (38)
T ss_dssp EEEEETTTEEEEE
T ss_pred EeeEccCCcEEEc
Confidence 5666666666664
No 49
>PF02035 Coagulin: Coagulin; InterPro: IPR000275 Coagulogen is a gel-forming protein of hemolymph that hinders the spread of invaders by immobilising them [, ]. The protein contains a single 175- residue polypeptide chain; this is cleaved after Arg-18 and Arg-46 by a clotting enzyme contained in the hemocyte and activated by a bacterial endotoxin (lipopolysaccharide). Cleavage releases two chains of coagulin, A and B, linked by two disulphide bonds, together with the peptide C [, ]. Gel formation results from interlinking of coagulin molecules. Secondary structure prediction suggests the C peptide forms an alpha- helix, which is released during the proteolytic conversion of coagulogen to coagulin gel []. The beta-sheet structure and 16 half-cystines found in the molecule appear to yield a compact protein stable to acid and heat. Mammalian blood coagulation is based on the proteolytically induced polymerisation of fibrinogens. Initially, fibrin monomers noncovalently interact with each other. The resulting homopolymers are further stabilised when the plasma transglutaminase (TGase) intermolecularly cross-links epsilon-(gamma-glutamyl)lysine bonds. In crustaceans, hemolymph coagulation depends on the TGase-mediated cross-linking of specific plasma-clotting proteins, but without the proteolytic cascade. In horseshoe crabs, the proteolytic coagulation cascade triggered by lipopolysaccharides and beta-1,3-glucans leads to the conversion of coagulogen into coagulin, resulting in noncovalent coagulin homopolymers through head-to-tail interaction. Horseshoe crab TGase, however, does not cross-link coagulins intermolecularly. Recently, we found that coagulins are cross-linked on hemocyte cell surface proteins called proxins. This indicates that a cross-linking reaction at the final stage of hemolymph coagulation is an important innate immune system of horseshoe crabs [].; GO: 0042381 hemolymph coagulation, 0005576 extracellular region; PDB: 1AOC_A.
Probab=22.08 E-value=76 Score=25.21 Aligned_cols=34 Identities=15% Similarity=0.304 Sum_probs=15.0
Q ss_pred ecCCEEEEEecCCC-----CCCC-CCCCCC--CCCCcccCCC
Q 047082 172 MHGNLKIYTHYDKV-----DSQP-TQLPER--CSKLGVCDDN 205 (263)
Q Consensus 172 ~dG~lr~y~~~~~~-----~w~~-C~~~~~--CG~~g~C~~~ 205 (263)
..|.+|+..-.+.. .|+. |..||. ||.+|-|+..
T Consensus 101 ~a~efrvivqapragfrqcvwqhkcraygsn~c~~~grctqq 142 (174)
T PF02035_consen 101 VAGEFRVIVQAPRAGFRQCVWQHKCRAYGSNNCGFNGRCTQQ 142 (174)
T ss_dssp TTS-EEEE--BCCCTB-B---EEEET-TSSSB-SSS-EE--E
T ss_pred ecceEEEEEeCchhhHHHHHHHhhhccccccccCcCceeccc
Confidence 34556655533322 3665 988754 9999999753
No 50
>cd05764 Ig_2 Subgroup of the immunoglobulin (Ig) superfamily. Ig_2: subgroup of the immunoglobulin (Ig) domain found in the Ig superfamily. The Ig superfamily is a heterogenous group of proteins, built on a common fold comprised of a sandwich of two beta sheets. Members of the Ig superfamily are components of immunoglobulin, neuroglia, cell surface glycoproteins, such as T-cell receptors, CD2, CD4, CD8, and membrane glycoproteins, such as butyrophilin and chondroitin sulfate proteoglycan core protein. A predominant feature of most Ig domains is a disulfide bridge connecting the two beta-sheets with a tryptophan residue packed against the disulfide bond.
Probab=21.81 E-value=1.6e+02 Score=19.61 Aligned_cols=33 Identities=12% Similarity=0.050 Sum_probs=20.3
Q ss_pred CCCCcEEEEcCCCCCCCCCCEEEEecCCcEEEE
Q 047082 3 YPKNSCSLNSIFASPVCGKPTFSLGSDGNLVLA 35 (263)
Q Consensus 3 ~~~~~~vWvANr~~Pv~~~~~l~l~~~G~L~l~ 35 (263)
.|...+.|.-+.+.++.......+..+|.|.|.
T Consensus 13 ~P~p~v~W~~~~~~~~~~~~~~~~~~~~~L~i~ 45 (74)
T cd05764 13 DPEPAIHWISPDGKLISNSSRTLVYDNGTLDIL 45 (74)
T ss_pred cCCCEEEEEeCCCEEecCCCeEEEecCCEEEEE
Confidence 366788888655556554444445556666664
No 51
>PF05833 FbpA: Fibronectin-binding protein A N-terminus (FbpA); InterPro: IPR008616 This family consists of the N-terminal region of the prokaryotic fibronectin-binding protein, the C-terminal region is IPR008532 from INTERPRO. Fibronectin binding is considered to be an important virulence factor in streptococcal infections. Fibronectin is a dimeric glycoprotein that is present in a soluble form in plasma and extracellular fluids; it is also present in a fibrillar form on cell surfaces. Both the soluble and cellular forms of fibronectin may be incorporated into the extracellular tissue matrix. While fibronectin has critical roles in eukaryotic cellular processes, such as adhesion, migration and differentiation, it is also a substrate for the attachment of bacteria. The binding of pathogenic Streptococcus pyogenes and Staphylococcus aureus to epithelial cells via fibronectin facilitates their internalisation and systemic spread within the host [].; PDB: 3DOA_A 2ZBK_F 2HKJ_A 1Z5B_A 1Z5C_B 1MX0_F 1Z5A_A 1MU5_A 1Z59_A.
Probab=21.80 E-value=75 Score=30.18 Aligned_cols=39 Identities=18% Similarity=0.336 Sum_probs=22.2
Q ss_pred eEEEEeeC-CCeeEEcCCCceEeeeccCCCc-----eeccCcccC
Q 047082 53 VGFKLLPN-GNMVLHDSKGNFIWQSFDCPTD-----TLLVGQSLL 91 (263)
Q Consensus 53 ~~a~Lld~-GNlvl~~~~~~~~WqSFd~PTD-----TlLpGq~l~ 91 (263)
-.++|... ||++|.|+++.+|+---.++.+ +++||+...
T Consensus 115 Li~El~g~~~NiiL~d~~~~Il~a~~~~~~~~~~~R~i~~G~~Y~ 159 (455)
T PF05833_consen 115 LIIELMGRHSNIILTDEDGKILDALRRVSFSQSRDREILPGEPYI 159 (455)
T ss_dssp EEEE--GGG-EEEEEETT-BEEEESS-B---------BSTTSB--
T ss_pred EEEEEcCCcccEEEEcCCCeEEeehhhcCcccccceeeccCcccc
Confidence 45677777 9999999888877754444554 899999976
No 52
>KOG3848 consensus Extracellular protein TEM7, contains PSI domain (tumor endothelial marker in humans) [Extracellular structures]
Probab=21.42 E-value=3.3e+02 Score=26.06 Aligned_cols=51 Identities=14% Similarity=0.136 Sum_probs=35.8
Q ss_pred EcCCCCCCCCCCEEEEecCCcEEEEcCCCcEEEcCCCC------CCCeeEEEEeeCCCeeEEc
Q 047082 11 NSIFASPVCGKPTFSLGSDGNLVLAEADGTVVCQSNTA------NKGVVGFKLLPNGNMVLHD 67 (263)
Q Consensus 11 vANr~~Pv~~~~~l~l~~~G~L~l~~~~g~~~Wst~~~------~~~~~~a~Lld~GNlvl~~ 67 (263)
.||-+...++++.+..-.+|.+++ +.|..... ++-.-.|.|+.+|.+|..-
T Consensus 194 MANFdts~snnS~V~y~DnGtafv------vqWdnV~Lqd~~d~gsFTFqatL~~dGdIVFaY 250 (516)
T KOG3848|consen 194 MANFDTSYSNNSTVVYFDNGTAFV------VQWDNVQLQDDKDEGSFTFQATLHKDGDIVFAY 250 (516)
T ss_pred hhcCCccccCCceEEEecCCeEEE------EEeeeEEeccCCCCCcEEEEEEeccCCcEEEEE
Confidence 488888788888888888998776 34554321 2223467888889888764
No 53
>TIGR03075 PQQ_enz_alc_DH PQQ-dependent dehydrogenase, methanol/ethanol family. This protein family has a phylogenetic distribution very similar to that coenzyme PQQ biosynthesis enzymes, as shown by partial phylogenetic profiling. Genes in this family often are found adjacent to the PQQ biosynthesis genes themselves. An unusual, strained disulfide bond between adjacent Cys residues contributes to PQQ-binding, as does a Trp residue that is part of a PQQ enzyme repeat (see pfam01011). Characterized members include the dehydrogenase subunit of a membrane-anchored, three subunit alcohol (ethanol) dehydrogenase of Gluconobacter suboxydans, a homodimeric ethanol dehydrogenase in Pseudomonas aeruginosa, and the large subunit of an alpha2/beta2 heterotetrameric methanol dehydrogenase in Methylobacterium extorquens.
Probab=20.70 E-value=4.1e+02 Score=26.00 Aligned_cols=75 Identities=19% Similarity=0.208 Sum_probs=0.0
Q ss_pred CCCCCCcEEEEcCCCCCCCCCC-----------------EEEEecCCcEEEEcC-CCcEEEcCCCC-----CCCeeEEEE
Q 047082 1 MEYPKNSCSLNSIFASPVCGKP-----------------TFSLGSDGNLVLAEA-DGTVVCQSNTA-----NKGVVGFKL 57 (263)
Q Consensus 1 ~~~~~~~~vWvANr~~Pv~~~~-----------------~l~l~~~G~L~l~~~-~g~~~Wst~~~-----~~~~~~a~L 57 (263)
++..+-..+|.-+...|..... .+.-+.+|.|+-+|. .|.++|+.... ....+...+
T Consensus 84 lDa~TGk~lW~~~~~~~~~~~~~~~~~~~~rg~av~~~~v~v~t~dg~l~ALDa~TGk~~W~~~~~~~~~~~~~tssP~v 163 (527)
T TIGR03075 84 LDAKTGKELWKYDPKLPDDVIPVMCCDVVNRGVALYDGKVFFGTLDARLVALDAKTGKVVWSKKNGDYKAGYTITAAPLV 163 (527)
T ss_pred EECCCCceeeEecCCCCcccccccccccccccceEECCEEEEEcCCCEEEEEECCCCCEEeecccccccccccccCCcEE
Q ss_pred ee--------------CCCeeEEcC-CCceEee
Q 047082 58 LP--------------NGNMVLHDS-KGNFIWQ 75 (263)
Q Consensus 58 ld--------------~GNlvl~~~-~~~~~Wq 75 (263)
.+ .|.|+-.|. +++.+|+
T Consensus 164 ~~g~Vivg~~~~~~~~~G~v~AlD~~TG~~lW~ 196 (527)
T TIGR03075 164 VKGKVITGISGGEFGVRGYVTAYDAKTGKLVWR 196 (527)
T ss_pred ECCEEEEeecccccCCCcEEEEEECCCCceeEe
No 54
>COG1520 FOG: WD40-like repeat [Function unknown]
Probab=20.48 E-value=4.3e+02 Score=24.07 Aligned_cols=42 Identities=29% Similarity=0.414 Sum_probs=29.1
Q ss_pred CCcEEEEcCCCC-------CCCCCCEEEEe-cCCcEEEEcCC-CcEEEcCC
Q 047082 5 KNSCSLNSIFAS-------PVCGKPTFSLG-SDGNLVLAEAD-GTVVCQSN 46 (263)
Q Consensus 5 ~~~~vWvANr~~-------Pv~~~~~l~l~-~~G~L~l~~~~-g~~~Wst~ 46 (263)
.-+.+|..+... |+...+.+-+. .+|.|+-.+.+ |..+|...
T Consensus 130 ~G~~~W~~~~~~~~~~~~~~v~~~~~v~~~s~~g~~~al~~~tG~~~W~~~ 180 (370)
T COG1520 130 TGTLVWSRNVGGSPYYASPPVVGDGTVYVGTDDGHLYALNADTGTLKWTYE 180 (370)
T ss_pred CCcEEEEEecCCCeEEecCcEEcCcEEEEecCCCeEEEEEccCCcEEEEEe
Confidence 456788877666 23334556666 57999888877 89999944
No 55
>COG3236 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=20.29 E-value=55 Score=26.52 Aligned_cols=21 Identities=33% Similarity=0.588 Sum_probs=16.1
Q ss_pred EEEeeCCCeeEEcC-CCceEee
Q 047082 55 FKLLPNGNMVLHDS-KGNFIWQ 75 (263)
Q Consensus 55 a~Lld~GNlvl~~~-~~~~~Wq 75 (263)
..||+||+.||... .+..+|-
T Consensus 116 e~LL~Tgd~vLVE~s~~D~~WG 137 (162)
T COG3236 116 ELLLATGDAVLVEASPNDAIWG 137 (162)
T ss_pred HHHHhcCCeeEEecCCCcceee
Confidence 45899999999954 4567884
Done!