Query 043869
Match_columns 450
No_of_seqs 240 out of 1540
Neff 7.6
Searched_HMMs 46136
Date Fri Mar 29 09:04:23 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/043869.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/043869hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF01453 B_lectin: D-mannose b 99.9 1.5E-28 3.2E-33 209.4 5.5 102 70-172 1-114 (114)
2 cd00028 B_lectin Bulb-type man 99.9 1.1E-24 2.5E-29 186.2 14.5 107 36-146 7-116 (116)
3 smart00108 B_lectin Bulb-type 99.9 2.1E-23 4.5E-28 177.9 14.2 106 36-145 7-114 (114)
4 PF00954 S_locus_glycop: S-loc 99.8 3.6E-19 7.9E-24 150.7 10.7 70 226-298 41-110 (110)
5 PF08276 PAN_2: PAN-like domai 99.5 1.8E-14 4E-19 110.4 5.3 58 317-374 4-66 (66)
6 cd01098 PAN_AP_plant Plant PAN 99.5 6.2E-14 1.3E-18 112.2 7.3 71 318-389 12-84 (84)
7 cd00129 PAN_APPLE PAN/APPLE-li 99.3 6.1E-12 1.3E-16 99.8 5.9 66 318-388 9-80 (80)
8 smart00108 B_lectin Bulb-type 98.7 5.4E-08 1.2E-12 82.7 8.8 73 90-189 24-101 (114)
9 smart00473 PAN_AP divergent su 98.7 7.4E-08 1.6E-12 75.3 7.5 70 318-387 4-77 (78)
10 cd00028 B_lectin Bulb-type man 98.6 8.8E-08 1.9E-12 81.7 7.9 72 91-189 25-102 (116)
11 PF01453 B_lectin: D-mannose b 98.5 1.8E-06 4E-11 73.4 10.9 100 37-147 12-114 (114)
12 cd01100 APPLE_Factor_XI_like S 97.9 1.1E-05 2.4E-10 62.9 4.1 48 322-369 8-57 (73)
13 smart00223 APPLE APPLE domain. 95.2 0.026 5.7E-07 44.6 4.0 47 323-369 6-57 (79)
14 PF00024 PAN_1: PAN domain Thi 94.5 0.07 1.5E-06 41.2 4.7 51 319-369 3-56 (79)
15 PF14295 PAN_4: PAN domain; PD 94.3 0.041 8.9E-07 38.9 2.8 41 327-367 2-51 (51)
16 PF08277 PAN_3: PAN-like domai 90.9 1.1 2.5E-05 34.0 6.9 32 337-368 18-49 (71)
17 smart00605 CW CW domain. 89.7 2.1 4.6E-05 34.7 8.0 55 337-391 20-77 (94)
18 PF04478 Mid2: Mid2 like cell 89.0 0.45 9.8E-06 42.1 3.6 38 412-450 50-87 (154)
19 PF08693 SKG6: Transmembrane a 88.7 0.45 9.8E-06 32.3 2.6 31 409-439 8-39 (40)
20 cd01099 PAN_AP_HGF Subfamily o 87.9 1.9 4.2E-05 33.9 6.3 33 337-369 23-59 (80)
21 PF07645 EGF_CA: Calcium-bindi 87.6 0.24 5.2E-06 34.0 0.8 32 265-296 3-35 (42)
22 cd00053 EGF Epidermal growth f 87.3 0.45 9.8E-06 30.2 2.0 30 267-296 2-31 (36)
23 PF15102 TMEM154: TMEM154 prot 86.4 0.92 2E-05 39.9 3.9 27 414-440 61-87 (146)
24 smart00179 EGF_CA Calcium-bind 82.1 1.1 2.4E-05 29.3 2.1 31 265-295 3-33 (39)
25 PTZ00382 Variant-specific surf 80.6 1.9 4.1E-05 35.4 3.3 14 426-439 82-95 (96)
26 cd00054 EGF_CA Calcium-binding 78.4 1.8 3.8E-05 27.9 2.1 32 265-296 3-34 (38)
27 PF01683 EB: EB module; Inter 76.5 2.5 5.3E-05 30.2 2.5 33 262-298 17-49 (52)
28 PF02009 Rifin_STEVOR: Rifin/s 75.6 0.51 1.1E-05 46.8 -1.7 32 415-446 258-291 (299)
29 PF12947 EGF_3: EGF domain; I 74.9 0.65 1.4E-05 30.9 -0.8 27 270-296 5-31 (36)
30 PF01102 Glycophorin_A: Glycop 74.1 0.57 1.2E-05 40.1 -1.5 28 413-442 68-95 (122)
31 PF01299 Lamp: Lysosome-associ 73.1 2.5 5.5E-05 42.1 2.6 19 428-446 287-306 (306)
32 PF02439 Adeno_E3_CR2: Adenovi 70.8 0.97 2.1E-05 30.2 -0.7 29 415-443 8-36 (38)
33 PF06697 DUF1191: Protein of u 69.1 3.6 7.9E-05 40.2 2.5 36 413-448 214-250 (278)
34 PF00008 EGF: EGF-like domain 69.0 1.1 2.5E-05 28.7 -0.7 24 272-295 5-29 (32)
35 PF12661 hEGF: Human growth fa 68.8 1.1 2.4E-05 22.9 -0.6 9 287-295 1-9 (13)
36 smart00181 EGF Epidermal growt 67.6 4.4 9.5E-05 25.9 1.9 25 271-296 6-30 (35)
37 PF14575 EphA2_TM: Ephrin type 66.6 1.5 3.1E-05 34.4 -0.6 25 415-439 3-27 (75)
38 PTZ00046 rifin; Provisional 63.8 1.7 3.8E-05 43.9 -0.8 24 415-438 317-340 (358)
39 TIGR01477 RIFIN variant surfac 63.3 1.8 3.9E-05 43.7 -0.9 33 415-447 312-346 (353)
40 PF12662 cEGF: Complement Clr- 62.3 4 8.6E-05 24.6 0.8 11 287-297 3-13 (24)
41 PF13908 Shisa: Wnt and FGF in 60.5 7.2 0.00016 35.6 2.7 20 411-430 77-96 (179)
42 PF06024 DUF912: Nucleopolyhed 58.8 7.7 0.00017 32.1 2.3 9 382-390 15-23 (101)
43 PF06365 CD34_antigen: CD34/Po 58.4 9.4 0.0002 35.7 3.0 36 413-448 101-139 (202)
44 PF07974 EGF_2: EGF-like domai 57.8 7.7 0.00017 25.0 1.7 24 271-296 6-29 (32)
45 PF05454 DAG1: Dystroglycan (D 54.3 4.2 9.1E-05 40.2 0.0 29 410-438 145-173 (290)
46 PRK11138 outer membrane biogen 53.9 49 0.0011 33.9 7.9 71 70-142 88-186 (394)
47 KOG1219 Uncharacterized conser 53.2 15 0.00032 45.7 4.1 32 265-297 3865-3897(4289)
48 PF13360 PQQ_2: PQQ-like domai 50.3 1.5E+02 0.0031 27.3 9.9 75 68-142 53-148 (238)
49 PF09064 Tme5_EGF_like: Thromb 49.8 9.4 0.0002 25.0 1.1 18 278-296 11-28 (34)
50 PF15102 TMEM154: TMEM154 prot 48.1 16 0.00034 32.3 2.6 31 415-445 58-89 (146)
51 PF12768 Rax2: Cortical protei 46.5 12 0.00025 37.1 1.7 27 408-434 226-252 (281)
52 PF14670 FXa_inhibition: Coagu 45.3 5.6 0.00012 26.4 -0.5 21 278-298 11-31 (36)
53 PRK11138 outer membrane biogen 45.2 1.2E+02 0.0027 30.9 9.2 48 94-142 302-361 (394)
54 cd05845 Ig2_L1-CAM_like Second 44.8 35 0.00075 27.9 4.0 35 68-103 31-65 (95)
55 PF12690 BsuPI: Intracellular 44.7 28 0.00061 27.5 3.4 16 97-113 27-42 (82)
56 KOG0291 WD40-repeat-containing 43.4 4.7E+02 0.01 29.5 13.1 56 86-141 352-420 (893)
57 TIGR03300 assembly_YfgL outer 43.3 1.7E+02 0.0038 29.4 9.9 50 93-142 248-305 (377)
58 PF01436 NHL: NHL repeat; Int 41.3 49 0.0011 20.2 3.4 21 89-110 6-26 (28)
59 PF13360 PQQ_2: PQQ-like domai 41.2 1.3E+02 0.0028 27.7 8.0 73 70-142 12-102 (238)
60 PF03302 VSP: Giardia variant- 41.2 22 0.00048 36.9 2.9 15 425-439 382-396 (397)
61 PF12877 DUF3827: Domain of un 40.3 21 0.00045 38.9 2.5 30 410-439 267-297 (684)
62 PF12191 stn_TNFRSF12A: Tumour 40.1 8 0.00017 33.1 -0.4 37 413-450 79-118 (129)
63 PHA02887 EGF-like protein; Pro 39.8 32 0.00069 29.2 3.0 35 260-295 79-117 (126)
64 PF05337 CSF-1: Macrophage col 39.4 9.8 0.00021 37.0 0.0 30 419-448 231-260 (285)
65 PHA03099 epidermal growth fact 38.8 16 0.00034 31.5 1.1 31 417-447 104-134 (139)
66 TIGR03300 assembly_YfgL outer 37.8 2.1E+02 0.0046 28.7 9.6 18 125-142 249-267 (377)
67 PF15065 NCU-G1: Lysosomal tra 37.6 12 0.00026 38.1 0.2 32 411-442 318-349 (350)
68 COG3763 Uncharacterized protei 37.3 4.3 9.4E-05 31.0 -2.2 28 416-443 5-32 (71)
69 PF01034 Syndecan: Syndecan do 35.7 11 0.00024 28.4 -0.2 8 432-439 30-37 (64)
70 PF14991 MLANA: Protein melan- 35.1 12 0.00026 31.5 -0.1 17 424-440 36-52 (118)
71 PF07354 Sp38: Zona-pellucida- 34.3 52 0.0011 32.0 4.0 35 68-103 10-44 (271)
72 KOG0640 mRNA cleavage stimulat 32.5 1.6E+02 0.0034 29.6 6.9 65 76-141 253-333 (430)
73 KOG3637 Vitronectin receptor, 32.3 31 0.00067 40.3 2.5 22 412-433 978-1000(1030)
74 TIGR03066 Gem_osc_para_1 Gemma 31.6 1.8E+02 0.0039 24.6 6.3 52 86-138 34-104 (111)
75 PF06006 DUF905: Bacterial pro 31.3 53 0.0011 25.1 2.8 20 128-147 34-53 (70)
76 KOG4649 PQQ (pyrrolo-quinoline 29.3 2E+02 0.0043 28.3 6.9 40 71-111 169-213 (354)
77 PF12946 EGF_MSP1_1: MSP1 EGF 29.2 11 0.00025 25.2 -1.0 27 271-297 5-32 (37)
78 PHA03290 envelope glycoprotein 29.1 52 0.0011 33.0 3.1 35 12-55 11-45 (357)
79 TIGR03075 PQQ_enz_alc_DH PQQ-d 28.4 7.7E+02 0.017 26.5 13.2 121 50-178 73-250 (527)
80 PF08114 PMP1_2: ATPase proteo 28.1 6.4 0.00014 26.8 -2.3 19 429-447 23-41 (43)
81 PF02480 Herpes_gE: Alphaherpe 28.1 20 0.00042 37.8 0.0 30 302-331 237-267 (439)
82 cd00216 PQQ_DH Dehydrogenases 27.6 1.9E+02 0.0042 30.7 7.5 71 69-141 37-135 (488)
83 PF05545 FixQ: Cbb3-type cytoc 27.6 21 0.00045 25.2 0.0 14 432-445 27-40 (49)
84 PF11403 Yeast_MT: Yeast metal 26.3 47 0.001 21.6 1.5 19 271-294 12-30 (40)
85 PF12458 DUF3686: ATPase invol 25.9 1.7E+02 0.0037 30.6 6.2 55 38-103 312-367 (448)
86 PF00558 Vpu: Vpu protein; In 24.7 36 0.00078 27.0 0.9 6 441-446 34-39 (81)
87 PTZ00382 Variant-specific surf 24.6 27 0.00059 28.6 0.3 28 415-442 68-95 (96)
88 KOG1214 Nidogen and related ba 24.0 55 0.0012 36.8 2.5 31 265-296 828-858 (1289)
89 PRK12785 fliL flagellar basal 23.1 95 0.0021 28.0 3.5 18 421-438 32-49 (166)
90 PF02237 BPL_C: Biotin protein 22.1 1.1E+02 0.0024 21.3 3.0 14 90-103 20-33 (48)
91 PF06247 Plasmod_Pvs28: Plasmo 22.1 23 0.00049 32.7 -0.7 41 264-309 39-88 (197)
92 PF05568 ASFV_J13L: African sw 21.8 17 0.00036 31.9 -1.6 16 432-447 49-64 (189)
93 PRK01844 hypothetical protein; 20.5 11 0.00024 29.1 -2.6 25 418-442 7-31 (72)
94 TIGR03503 conserved hypothetic 20.4 52 0.0011 33.8 1.3 21 419-440 353-373 (374)
95 PF14316 DUF4381: Domain of un 20.4 23 0.0005 31.2 -1.1 11 435-445 42-52 (146)
96 KOG4289 Cadherin EGF LAG seven 20.2 50 0.0011 39.5 1.3 40 270-309 1244-1285(2531)
97 cd05852 Ig5_Contactin-1 Fifth 20.0 1.1E+02 0.0023 23.1 2.8 34 68-103 13-46 (73)
98 PTZ00208 65 kDa invariant surf 20.0 48 0.001 34.1 1.0 30 410-439 384-414 (436)
No 1
>PF01453 B_lectin: D-mannose binding lectin; InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]: Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity. Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=99.95 E-value=1.5e-28 Score=209.35 Aligned_cols=102 Identities=46% Similarity=0.740 Sum_probs=74.1
Q ss_pred CCcEEEEecCCCCCCC--CccEEEEecCCcEEEEeCCCCceEEeecCC--C--CccEEEEecCCCeEEEecCCeeEEeec
Q 043869 70 EKTVVWTANRDNPPVS--SNATLMFNSEGRIVLRSGEQGQNSIIADNS--Q--SASSASMLDSGSFVLHNSDGKVIWQTF 143 (450)
Q Consensus 70 ~~tvVW~ANr~~Pv~~--~~~~L~l~~~G~LvL~d~~~g~~vW~st~~--~--~~~~a~LldsGNlVL~~~~~~~lWQSF 143 (450)
++||||+|||+.|+.+ ...+|.|+.||+|+|.+. .++.+|.+..+ . .+..|.|+|+|||||+|..+.+|||||
T Consensus 1 ~~tvvW~an~~~p~~~~s~~~~L~l~~dGnLvl~~~-~~~~iWss~~t~~~~~~~~~~~L~~~GNlvl~d~~~~~lW~Sf 79 (114)
T PF01453_consen 1 PRTVVWVANRNSPLTSSSGNYTLILQSDGNLVLYDS-NGSVIWSSNNTSGRGNSGCYLVLQDDGNLVLYDSSGNVLWQSF 79 (114)
T ss_dssp ---------TTEEEEECETTEEEEEETTSEEEEEET-TTEEEEE--S-TTSS-SSEEEEEETTSEEEEEETTSEEEEEST
T ss_pred CcccccccccccccccccccccceECCCCeEEEEcC-CCCEEEEecccCCccccCeEEEEeCCCCEEEEeecceEEEeec
Confidence 3699999999999953 348999999999999998 88899944244 2 478999999999999999999999999
Q ss_pred CCCCCccCCCcccCC----C--CeEEeccCCCCCC
Q 043869 144 DHPTDTLLPTQRLSA----G--TELCSGISETDPS 172 (450)
Q Consensus 144 d~PTDTlLpgq~L~~----~--~~L~S~~s~~dps 172 (450)
||||||+||+|+|+. + ..|+||++.+|||
T Consensus 80 ~~ptdt~L~~q~l~~~~~~~~~~~~~sw~s~~dps 114 (114)
T PF01453_consen 80 DYPTDTLLPGQKLGDGNVTGKNDSLTSWSSNTDPS 114 (114)
T ss_dssp TSSS-EEEEEET--TSEEEEESTSSEEEESS----
T ss_pred CCCccEEEeccCcccCCCccccceEEeECCCCCCC
Confidence 999999999999987 3 3599999999996
No 2
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=99.92 E-value=1.1e-24 Score=186.23 Aligned_cols=107 Identities=37% Similarity=0.669 Sum_probs=93.2
Q ss_pred CCeEEeCCCeEEEEEEeCCCCCeeEEEEEEeecCCCcEEEEecCCCCCCCCccEEEEecCCcEEEEeCCCCceEEeecCC
Q 043869 36 NSSWRSPSGLYAFGFYPQRNGSRYYVGVFLAGIPEKTVVWTANRDNPPVSSNATLMFNSEGRIVLRSGEQGQNSIIADNS 115 (450)
Q Consensus 36 ~~~l~S~~g~F~lGF~~~~~~~~~~lgIw~~~~~~~tvVW~ANr~~Pv~~~~~~L~l~~~G~LvL~d~~~g~~vW~st~~ 115 (450)
+++|+|+++.|++|||.......++++|||.+.+ +++||.||++.|.. ..++|.|+.||+|+|.|. +|.++| ++++
T Consensus 7 ~~~l~s~~~~f~~G~~~~~~q~~dgnlv~~~~~~-~~~vW~snt~~~~~-~~~~l~l~~dGnLvl~~~-~g~~vW-~S~~ 82 (116)
T cd00028 7 GQTLVSSGSLFELGFFKLIMQSRDYNLILYKGSS-RTVVWVANRDNPSG-SSCTLTLQSDGNLVIYDG-SGTVVW-SSNT 82 (116)
T ss_pred CCEEEeCCCcEEEecccCCCCCCeEEEEEEeCCC-CeEEEECCCCCCCC-CCEEEEEecCCCeEEEcC-CCcEEE-Eecc
Confidence 3899999999999999854332389999998876 78999999999844 678999999999999998 899999 5554
Q ss_pred ---CCccEEEEecCCCeEEEecCCeeEEeecCCC
Q 043869 116 ---QSASSASMLDSGSFVLHNSDGKVIWQTFDHP 146 (450)
Q Consensus 116 ---~~~~~a~LldsGNlVL~~~~~~~lWQSFd~P 146 (450)
....+|.|+|+|||||++.++.+||||||||
T Consensus 83 ~~~~~~~~~~L~ddGnlvl~~~~~~~~W~Sf~~P 116 (116)
T cd00028 83 TRVNGNYVLVLLDDGNLVLYDSDGNFLWQSFDYP 116 (116)
T ss_pred cCCCCceEEEEeCCCCEEEECCCCCEEEcCCCCC
Confidence 3467899999999999999999999999999
No 3
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=99.90 E-value=2.1e-23 Score=177.86 Aligned_cols=106 Identities=32% Similarity=0.600 Sum_probs=91.6
Q ss_pred CCeEEeCCCeEEEEEEeCCCCCeeEEEEEEeecCCCcEEEEecCCCCCCCCccEEEEecCCcEEEEeCCCCceEEeecCC
Q 043869 36 NSSWRSPSGLYAFGFYPQRNGSRYYVGVFLAGIPEKTVVWTANRDNPPVSSNATLMFNSEGRIVLRSGEQGQNSIIADNS 115 (450)
Q Consensus 36 ~~~l~S~~g~F~lGF~~~~~~~~~~lgIw~~~~~~~tvVW~ANr~~Pv~~~~~~L~l~~~G~LvL~d~~~g~~vW~st~~ 115 (450)
++.|+|+++.|++|||..... .++++|||...+ +++||+|||+.|+. .+++|.|++||+|+|.+. +|.++|++...
T Consensus 7 ~~~l~s~~~~f~~G~~~~~~q-~dgnlV~~~~~~-~~~vW~snt~~~~~-~~~~l~l~~dGnLvl~~~-~g~~vW~S~t~ 82 (114)
T smart00108 7 GQTLVSGNSLFELGFFTLIMQ-NDYNLILYKSSS-RTVVWVANRDNPVS-DSCTLTLQSDGNLVLYDG-DGRVVWSSNTT 82 (114)
T ss_pred CCEEecCCCcEeeeccccCCC-CCEEEEEEECCC-CcEEEECCCCCCCC-CCEEEEEeCCCCEEEEeC-CCCEEEEeccc
Confidence 389999999999999985433 488999999876 78999999999976 458999999999999998 89999944332
Q ss_pred --CCccEEEEecCCCeEEEecCCeeEEeecCC
Q 043869 116 --QSASSASMLDSGSFVLHNSDGKVIWQTFDH 145 (450)
Q Consensus 116 --~~~~~a~LldsGNlVL~~~~~~~lWQSFd~ 145 (450)
....+|+|+|+|||||++.++.++||||||
T Consensus 83 ~~~~~~~~~L~ddGnlvl~~~~~~~~W~Sf~~ 114 (114)
T smart00108 83 GANGNYVLVLLDDGNLVIYDSDGNFLWQSFDY 114 (114)
T ss_pred CCCCceEEEEeCCCCEEEECCCCCEEeCCCCC
Confidence 345689999999999999999999999997
No 4
>PF00954 S_locus_glycop: S-locus glycoprotein family; InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=99.80 E-value=3.6e-19 Score=150.74 Aligned_cols=70 Identities=33% Similarity=0.835 Sum_probs=63.7
Q ss_pred CCCceEEEEEEccCCcEEEEEecccCCCCceEEEccccCCCCCCcCCCCCCcccccCCCCCCCcCCCCCeecc
Q 043869 226 PIQGMMYLMKIDSDGIFRLYSYNLRWQNSTWSEVWPSTSEKCDPIGLCGFNSFCVLNDQTPNCTCLPGFVAIS 298 (450)
Q Consensus 226 ~~~~~~~r~~Ld~dG~lr~y~~~~~~~~~~W~~~w~~p~~~C~~~g~CG~~g~C~~~~~~~~C~C~~GF~~~~ 298 (450)
.+...+.|++||.||++++|.|. +..+.|..+|++|.++||+|+.||+||+|+. .+.+.|+||+||+|++
T Consensus 41 ~~~s~~~r~~ld~~G~l~~~~w~--~~~~~W~~~~~~p~d~Cd~y~~CG~~g~C~~-~~~~~C~Cl~GF~P~n 110 (110)
T PF00954_consen 41 SNSSVLSRLVLDSDGQLQRYIWN--ESTQSWSVFWSAPKDQCDVYGFCGPNGICNS-NNSPKCSCLPGFEPKN 110 (110)
T ss_pred CCCceEEEEEEeeeeEEEEEEEe--cCCCcEEEEEEecccCCCCccccCCccEeCC-CCCCceECCCCcCCCc
Confidence 35567899999999999999998 7788999999999999999999999999987 4578999999999974
No 5
>PF08276 PAN_2: PAN-like domain; InterPro: IPR013227 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs
Probab=99.51 E-value=1.8e-14 Score=110.41 Aligned_cols=58 Identities=26% Similarity=0.496 Sum_probs=50.6
Q ss_pred CcceEEeccccccCccccee-cccCHHHHHHHHhcCCCeEEEEec----CCceEeeeccccce
Q 043869 317 NKAIQELENTNWEDVSYNVL-SEITKEKCKQACLEDCNCEAALYK----NEECKMQRLPLRFG 374 (450)
Q Consensus 317 ~~~f~~l~~v~~p~~~~~~~-~~~s~~~C~~~CL~nCsC~A~~y~----~g~C~~~~~~L~~~ 374 (450)
+++|++|++|++|+++...+ .++++++|+++||+||||+||+|. +++|++|.++|+|.
T Consensus 4 ~d~F~~l~~~~~p~~~~~~~~~~~s~~~C~~~Cl~nCsC~Ayay~~~~~~~~C~lW~~~L~d~ 66 (66)
T PF08276_consen 4 GDGFLKLPNMKLPDFDNAIVDSSVSLEECEKACLSNCSCTAYAYSNLSGGGGCLLWYGDLVDL 66 (66)
T ss_pred CCEEEEECCeeCCCCcceeeecCCCHHHHHhhcCCCCCEeeEEeeccCCCCEEEEEcCEeecC
Confidence 38999999999999876654 568999999999999999999995 26799999998873
No 6
>cd01098 PAN_AP_plant Plant PAN/APPLE-like domain; present in plant S-receptor protein kinases and secreted glycoproteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions. S-receptor protein kinases and S-locus glycoproteins are involved in sporophytic self-incompatibility response in Brassica, one of probably many molecular mechanisms, by which hermaphrodite flowering plants avoid self-fertilization.
Probab=99.49 E-value=6.2e-14 Score=112.24 Aligned_cols=71 Identities=20% Similarity=0.447 Sum_probs=61.0
Q ss_pred cceEEeccccccCcccceecccCHHHHHHHHhcCCCeEEEEec--CCceEeeeccccceeeeCCCCceEEEEec
Q 043869 318 KAIQELENTNWEDVSYNVLSEITKEKCKQACLEDCNCEAALYK--NEECKMQRLPLRFGKRNLRDSDITFVKVD 389 (450)
Q Consensus 318 ~~f~~l~~v~~p~~~~~~~~~~s~~~C~~~CL~nCsC~A~~y~--~g~C~~~~~~L~~~~~~~~~~~~~yiKv~ 389 (450)
++|++++++++|+..+.. ...++++|+++||+||+|+||+|. +++|++|..++.+.+.....+.++||||+
T Consensus 12 ~~f~~~~~~~~~~~~~~~-~~~s~~~C~~~Cl~nCsC~a~~~~~~~~~C~~~~~~~~~~~~~~~~~~~~yiKv~ 84 (84)
T cd01098 12 DGFLKLPDVKLPDNASAI-TAISLEECREACLSNCSCTAYAYNNGSGGCLLWNGLLNNLRSLSSGGGTLYLRLA 84 (84)
T ss_pred CEEEEeCCeeCCCchhhh-ccCCHHHHHHHHhcCCCcceeeecCCCCeEEEEeceecceEeecCCCcEEEEEeC
Confidence 689999999999876654 667999999999999999999994 47899999999987765444589999985
No 7
>cd00129 PAN_APPLE PAN/APPLE-like domain; present in N-terminal (N) domains of plasminogen/ hepatocyte growth factor proteins, plasma prekallikrein/coagulation factor XI and microneme antigen proteins, plant receptor-like protein kinases, and various nematode and leech anti-platelet proteins. Common structural features include two disulfide bonds that link the alpha-helix to the central region of the protein. PAN domains have significant functional versatility, fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=99.27 E-value=6.1e-12 Score=99.83 Aligned_cols=66 Identities=9% Similarity=0.150 Sum_probs=56.9
Q ss_pred cceEEeccccccCcccceecccCHHHHHHHHhc---CCCeEEEEec--CCceEeeeccc-cceeeeCCCCceEEEEe
Q 043869 318 KAIQELENTNWEDVSYNVLSEITKEKCKQACLE---DCNCEAALYK--NEECKMQRLPL-RFGKRNLRDSDITFVKV 388 (450)
Q Consensus 318 ~~f~~l~~v~~p~~~~~~~~~~s~~~C~~~CL~---nCsC~A~~y~--~g~C~~~~~~L-~~~~~~~~~~~~~yiKv 388 (450)
..|+++.++++|+... .+++||++.|++ ||||.||+|. +++|++|.++| .+.+....++.++|+|.
T Consensus 9 g~fl~~~~~klpd~~~-----~s~~eC~~~Cl~~~~nCsC~Aya~~~~~~gC~~W~~~l~~d~~~~~~~g~~Ly~r~ 80 (80)
T cd00129 9 GTTLIKIALKIKTTKA-----NTADECANRCEKNGLPFSCKAFVFAKARKQCLWFPFNSMSGVRKEFSHGFDLYENK 80 (80)
T ss_pred CeEEEeecccCCcccc-----cCHHHHHHHHhcCCCCCCceeeeccCCCCCeEEecCcchhhHHhccCCCceeEeEC
Confidence 6799999999998644 689999999999 9999999994 35899999999 88887766569999983
No 8
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=98.73 E-value=5.4e-08 Score=82.75 Aligned_cols=73 Identities=26% Similarity=0.536 Sum_probs=56.4
Q ss_pred EEEecCCcEEEEeCCC-CceEEeecCC----CCccEEEEecCCCeEEEecCCeeEEeecCCCCCccCCCcccCCCCeEEe
Q 043869 90 LMFNSEGRIVLRSGEQ-GQNSIIADNS----QSASSASMLDSGSFVLHNSDGKVIWQTFDHPTDTLLPTQRLSAGTELCS 164 (450)
Q Consensus 90 L~l~~~G~LvL~d~~~-g~~vW~st~~----~~~~~a~LldsGNlVL~~~~~~~lWQSFd~PTDTlLpgq~L~~~~~L~S 164 (450)
+.+..||+||+.+. . +.++| ++++ .....+.|+++|||||++.++.++|+|= |-
T Consensus 24 ~~~q~dgnlV~~~~-~~~~~vW-~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S~-----t~-------------- 82 (114)
T smart00108 24 LIMQNDYNLILYKS-SSRTVVW-VANRDNPVSDSCTLTLQSDGNLVLYDGDGRVVWSSN-----TT-------------- 82 (114)
T ss_pred cCCCCCEEEEEEEC-CCCcEEE-ECCCCCCCCCCEEEEEeCCCCEEEEeCCCCEEEEec-----cc--------------
Confidence 44567999999986 4 47999 6654 1236789999999999999899999971 10
Q ss_pred ccCCCCCCCCceEEEecCCCceeEc
Q 043869 165 GISETDPSTGKFRLKMQNDGNLVQY 189 (450)
Q Consensus 165 ~~s~~dps~G~f~l~~~~~g~~~l~ 189 (450)
...|.+.+.|+++|++++|
T Consensus 83 ------~~~~~~~~~L~ddGnlvl~ 101 (114)
T smart00108 83 ------GANGNYVLVLLDDGNLVIY 101 (114)
T ss_pred ------CCCCceEEEEeCCCCEEEE
Confidence 1356688999999999998
No 9
>smart00473 PAN_AP divergent subfamily of APPLE domains. Apple-like domains present in Plasminogen, C. elegans hypothetical ORFs and the extracellular portion of plant receptor-like protein kinases. Predicted to possess protein- and/or carbohydrate-binding functions.
Probab=98.67 E-value=7.4e-08 Score=75.28 Aligned_cols=70 Identities=23% Similarity=0.378 Sum_probs=55.3
Q ss_pred cceEEeccccccCcccceecccCHHHHHHHHhc-CCCeEEEEec--CCceEeee-ccccceeeeCCCCceEEEE
Q 043869 318 KAIQELENTNWEDVSYNVLSEITKEKCKQACLE-DCNCEAALYK--NEECKMQR-LPLRFGKRNLRDSDITFVK 387 (450)
Q Consensus 318 ~~f~~l~~v~~p~~~~~~~~~~s~~~C~~~CL~-nCsC~A~~y~--~g~C~~~~-~~L~~~~~~~~~~~~~yiK 387 (450)
..|++++++.+++.........++++|++.|++ +|+|.|+.|. ++.|++|. .++.+.+.....+.++|.|
T Consensus 4 ~~f~~~~~~~l~~~~~~~~~~~s~~~C~~~C~~~~~~C~s~~y~~~~~~C~l~~~~~~~~~~~~~~~~~~~y~~ 77 (78)
T smart00473 4 DCFVRLPNTKLPGFSRIVISVASLEECASKCLNSNCSCRSFTYNNGTKGCLLWSESSLGDARLFPSGGVDLYEK 77 (78)
T ss_pred ceeEEecCccCCCCcceeEcCCCHHHHHHHhCCCCCceEEEEEcCCCCEEEEeeCCccccceecccCCceeEEe
Confidence 568999999998654433456799999999999 9999999994 57899998 7777776444444677776
No 10
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=98.65 E-value=8.8e-08 Score=81.71 Aligned_cols=72 Identities=25% Similarity=0.513 Sum_probs=55.9
Q ss_pred EEec-CCcEEEEeCCC-CceEEeecCC----CCccEEEEecCCCeEEEecCCeeEEeecCCCCCccCCCcccCCCCeEEe
Q 043869 91 MFNS-EGRIVLRSGEQ-GQNSIIADNS----QSASSASMLDSGSFVLHNSDGKVIWQTFDHPTDTLLPTQRLSAGTELCS 164 (450)
Q Consensus 91 ~l~~-~G~LvL~d~~~-g~~vW~st~~----~~~~~a~LldsGNlVL~~~~~~~lWQSFd~PTDTlLpgq~L~~~~~L~S 164 (450)
.... ||+|++.+. . +.++| ++++ .....+.|+++|||||+|.++.++|+|=..
T Consensus 25 ~~q~~dgnlv~~~~-~~~~~vW-~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S~~~------------------- 83 (116)
T cd00028 25 IMQSRDYNLILYKG-SSRTVVW-VANRDNPSGSSCTLTLQSDGNLVIYDGSGTVVWSSNTT------------------- 83 (116)
T ss_pred CCCCCeEEEEEEeC-CCCeEEE-ECCCCCCCCCCEEEEEecCCCeEEEcCCCcEEEEeccc-------------------
Confidence 3454 899999976 4 47899 6654 245678999999999999999999996211
Q ss_pred ccCCCCCCCCceEEEecCCCceeEc
Q 043869 165 GISETDPSTGKFRLKMQNDGNLVQY 189 (450)
Q Consensus 165 ~~s~~dps~G~f~l~~~~~g~~~l~ 189 (450)
...+.+.+.|++||++++|
T Consensus 84 ------~~~~~~~~~L~ddGnlvl~ 102 (116)
T cd00028 84 ------RVNGNYVLVLLDDGNLVLY 102 (116)
T ss_pred ------CCCCceEEEEeCCCCEEEE
Confidence 0256788999999999998
No 11
>PF01453 B_lectin: D-mannose binding lectin; InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]: Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity. Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=98.46 E-value=1.8e-06 Score=73.44 Aligned_cols=100 Identities=20% Similarity=0.328 Sum_probs=69.1
Q ss_pred CeEEeCCCeEEEEEEeCCCCCeeEEEEEEeecCCCcEEEEe-cCCCCCCCCccEEEEecCCcEEEEeCCCCceEEeecCC
Q 043869 37 SSWRSPSGLYAFGFYPQRNGSRYYVGVFLAGIPEKTVVWTA-NRDNPPVSSNATLMFNSEGRIVLRSGEQGQNSIIADNS 115 (450)
Q Consensus 37 ~~l~S~~g~F~lGF~~~~~~~~~~lgIw~~~~~~~tvVW~A-Nr~~Pv~~~~~~L~l~~~G~LvL~d~~~g~~vW~st~~ 115 (450)
+.+.+.+|.+.|-|..+|+- +.|.. ..+++|.. +...... ..+.+.|.++|||||+|. .+.++|+|...
T Consensus 12 ~p~~~~s~~~~L~l~~dGnL------vl~~~--~~~~iWss~~t~~~~~-~~~~~~L~~~GNlvl~d~-~~~~lW~Sf~~ 81 (114)
T PF01453_consen 12 SPLTSSSGNYTLILQSDGNL------VLYDS--NGSVIWSSNNTSGRGN-SGCYLVLQDDGNLVLYDS-SGNVLWQSFDY 81 (114)
T ss_dssp EEEEECETTEEEEEETTSEE------EEEET--TTEEEEE--S-TTSS--SSEEEEEETTSEEEEEET-TSEEEEESTTS
T ss_pred cccccccccccceECCCCeE------EEEcC--CCCEEEEecccCCccc-cCeEEEEeCCCCEEEEee-cceEEEeecCC
Confidence 45656558999999987764 44443 45789999 4443332 478899999999999998 89999955433
Q ss_pred CCccEEEEec--CCCeEEEecCCeeEEeecCCCC
Q 043869 116 QSASSASMLD--SGSFVLHNSDGKVIWQTFDHPT 147 (450)
Q Consensus 116 ~~~~~a~Lld--sGNlVL~~~~~~~lWQSFd~PT 147 (450)
...+.+..++ .||++ +.....+.|.|=+.|+
T Consensus 82 ptdt~L~~q~l~~~~~~-~~~~~~~sw~s~~dps 114 (114)
T PF01453_consen 82 PTDTLLPGQKLGDGNVT-GKNDSLTSWSSNTDPS 114 (114)
T ss_dssp SS-EEEEEET--TSEEE-EESTSSEEEESS----
T ss_pred CccEEEeccCcccCCCc-cccceEEeECCCCCCC
Confidence 5566677777 88888 7666778999876663
No 12
>cd01100 APPLE_Factor_XI_like Subfamily of PAN/APPLE-like domains; present in plasma prekallikrein/coagulation factor XI, microneme antigen proteins, and a few prokaryotic proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=97.92 E-value=1.1e-05 Score=62.88 Aligned_cols=48 Identities=21% Similarity=0.460 Sum_probs=39.3
Q ss_pred EeccccccCcccceecccCHHHHHHHHhcCCCeEEEEec--CCceEeeec
Q 043869 322 ELENTNWEDVSYNVLSEITKEKCKQACLEDCNCEAALYK--NEECKMQRL 369 (450)
Q Consensus 322 ~l~~v~~p~~~~~~~~~~s~~~C~~~CL~nCsC~A~~y~--~g~C~~~~~ 369 (450)
.++++++++.+.......+.++|++.|+.+|+|.||.|. .+.|+++..
T Consensus 8 ~~~~~~~~g~d~~~~~~~s~~~Cq~~C~~~~~C~afT~~~~~~~C~lk~~ 57 (73)
T cd01100 8 QGSNVDFRGGDLSTVFASSAEQCQAACTADPGCLAFTYNTKSKKCFLKSS 57 (73)
T ss_pred ccCCCccccCCcceeecCCHHHHHHHcCCCCCceEEEEECCCCeEEcccC
Confidence 346888888777655566899999999999999999993 478998865
No 13
>smart00223 APPLE APPLE domain. Four-fold repeat in plasma kallikrein and coagulation factor XI. Factor XI apple 3 mediates binding to platelets. Factor XI apple 1 binds high-molecular-mass kininogen. Apple 4 in factor XI mediates dimer formation and binds to factor XIIa. Mutations in apple 4 cause factor XI deficiency, an inherited bleeding disorder.
Probab=95.22 E-value=0.026 Score=44.65 Aligned_cols=47 Identities=13% Similarity=0.383 Sum_probs=40.1
Q ss_pred eccccccCcccceecccCHHHHHHHHhcCCCeEEEEe--cCC---ceEeeec
Q 043869 323 LENTNWEDVSYNVLSEITKEKCKQACLEDCNCEAALY--KNE---ECKMQRL 369 (450)
Q Consensus 323 l~~v~~p~~~~~~~~~~s~~~C~~~CL~nCsC~A~~y--~~g---~C~~~~~ 369 (450)
.+|++|++.+...+...+.++|++.|..+=.|.+|.| .+. .|+++..
T Consensus 6 ~~~~df~G~Dl~~~~~~~~~~Cq~~Ct~~~~C~~FTf~~~~~~~~~C~LK~s 57 (79)
T smart00223 6 YKNVDFRGSDINTVYVPSAQVCQKRCTSHPRCLFFTFSTNEPPEEKCLLKDS 57 (79)
T ss_pred ccCccccCceeeeeecCCHHHHHHhhcCCCCccEEEeeCCCCCCCEeEeCcC
Confidence 4689999988887777899999999999999999999 345 7998754
No 14
>PF00024 PAN_1: PAN domain This Prosite entry concerns apple domains, a subset of PAN domains; InterPro: IPR003014 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs It has been shown that, the N-terminal N domains of members of the plasminogen/hepatocyte growth factor family, the apple domains of the plasma prekallikrein/coagulation factor XI family, and domains of various nematode proteins belong to the same module superfamily, the PAN module []. PAN contains a conserved core of three disulphide bridges. In some members of the family there is an additional fourth disulphide bridge that links the N and C termini of the domain.; PDB: 1GP9_C 2QJ2_B 1GMO_H 1NK1_B 3MKP_B 1BHT_B 3HN4_A 1GMN_A 3HMS_A 3HMT_B ....
Probab=94.46 E-value=0.07 Score=41.20 Aligned_cols=51 Identities=20% Similarity=0.429 Sum_probs=40.2
Q ss_pred ceEEeccccccCcccceecccCHHHHHHHHhcCCC-eEEEEe--cCCceEeeec
Q 043869 319 AIQELENTNWEDVSYNVLSEITKEKCKQACLEDCN-CEAALY--KNEECKMQRL 369 (450)
Q Consensus 319 ~f~~l~~v~~p~~~~~~~~~~s~~~C~~~CL~nCs-C~A~~y--~~g~C~~~~~ 369 (450)
.|..+++..+...........++++|.+.|+++=. |.++.| ..+.|.+...
T Consensus 3 ~f~~~~~~~l~~~~~~~~~v~s~~~C~~~C~~~~~~C~s~~y~~~~~~C~L~~~ 56 (79)
T PF00024_consen 3 AFERIPGYRLSGHSIKEINVPSLEECAQLCLNEPRRCKSFNYDPSSKTCYLSSS 56 (79)
T ss_dssp TEEEEEEEEEESCEEEEEEESSHHHHHHHHHHSTT-ESEEEEETTTTEEEEECS
T ss_pred CeEEECCEEEeCCcceEEcCCCHHHHHhhcCcCcccCCeEEEECCCCEEEEcCC
Confidence 47778888877755554444599999999999999 999999 3468998754
No 15
>PF14295 PAN_4: PAN domain; PDB: 2YIL_E 2YIP_C 2YIO_A.
Probab=94.29 E-value=0.041 Score=38.95 Aligned_cols=41 Identities=20% Similarity=0.549 Sum_probs=17.2
Q ss_pred cccCcccce--ecccCHHHHHHHHhcCCCeEEEEecC-------CceEee
Q 043869 327 NWEDVSYNV--LSEITKEKCKQACLEDCNCEAALYKN-------EECKMQ 367 (450)
Q Consensus 327 ~~p~~~~~~--~~~~s~~~C~~~CL~nCsC~A~~y~~-------g~C~~~ 367 (450)
++++.++.. ....+.++|.++|.++=.|.++.|.. +.|++|
T Consensus 2 d~~G~dl~~~~~~~~s~~~C~~~C~~~~~C~~~~~~~~~~~~~~~~C~LK 51 (51)
T PF14295_consen 2 DYPGGDLRSFPVTASSPEECQAACAADPGCQAFTFNPPGCPSSSGRCYLK 51 (51)
T ss_dssp ----------------HHHHHHHHHTSTT--EEEEETTEE----------
T ss_pred cccccccccccccCCCHHHHHHHccCCCCCCEEEEECCCcccccccccCC
Confidence 455544443 24568999999999999999999932 457764
No 16
>PF08277 PAN_3: PAN-like domain; InterPro: IPR006583 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs The PAN-3 or CW is a domain associated with a number of Caenorhabditis elegans hypothetical proteins.
Probab=90.91 E-value=1.1 Score=33.99 Aligned_cols=32 Identities=25% Similarity=0.641 Sum_probs=27.9
Q ss_pred cccCHHHHHHHHhcCCCeEEEEecCCceEeee
Q 043869 337 SEITKEKCKQACLEDCNCEAALYKNEECKMQR 368 (450)
Q Consensus 337 ~~~s~~~C~~~CL~nCsC~A~~y~~g~C~~~~ 368 (450)
...+.++|-+.|.++=+|.++.+.++.|.++.
T Consensus 18 ~~~sw~~Cv~~C~~~~~C~la~~~~~~C~~y~ 49 (71)
T PF08277_consen 18 TNTSWDDCVQKCYNDENCVLAYFDSGKCYLYN 49 (71)
T ss_pred cCCCHHHHhHHhCCCCEEEEEEeCCCCEEEEE
Confidence 45678999999999999999998777899865
No 17
>smart00605 CW CW domain.
Probab=89.74 E-value=2.1 Score=34.67 Aligned_cols=55 Identities=24% Similarity=0.372 Sum_probs=39.5
Q ss_pred cccCHHHHHHHHhcCCCeEEEEecC-CceEeeec-cccceeee-CCCCceEEEEecCC
Q 043869 337 SEITKEKCKQACLEDCNCEAALYKN-EECKMQRL-PLRFGKRN-LRDSDITFVKVDDA 391 (450)
Q Consensus 337 ~~~s~~~C~~~CL~nCsC~A~~y~~-g~C~~~~~-~L~~~~~~-~~~~~~~yiKv~~~ 391 (450)
...+-++|.+.|.++..|..+...+ ..|.+... .+...++. ...+..+=||+..+
T Consensus 20 ~~~sw~~Ci~~C~~~~~Cvlay~~~~~~C~~f~~~~~~~v~~~~~~~~~~VAfK~~~~ 77 (94)
T smart00605 20 ATLSWDECIQKCYEDSNCVLAYGNSSETCYLFSYGTVLTVKKLSSSSGKKVAFKVSTD 77 (94)
T ss_pred cCCCHHHHHHHHhCCCceEEEecCCCCceEEEEcCCeEEEEEccCCCCcEEEEEEeCC
Confidence 3467899999999999999887643 78987643 34555554 33446788888754
No 18
>PF04478 Mid2: Mid2 like cell wall stress sensor; InterPro: IPR007567 This family represents a region near the C terminus of Mid2, which contains a transmembrane region. The remainder of the protein sequence is serine-rich and of low complexity, and is therefore impossible to align accurately. Mid2 is thought to act as a mechanosensor of cell wall stress. The C-terminal cytoplasmic region of Mid2 is known to interact with Rom2, a guanine nucleotide exchange factor (GEF) for Rho1, which is part of the cell wall integrity signalling pathway [].
Probab=88.98 E-value=0.45 Score=42.08 Aligned_cols=38 Identities=16% Similarity=0.194 Sum_probs=15.2
Q ss_pred EEEEehhhHHHHHHHHHHhheeeeEeeccceeecccCCC
Q 043869 412 NIVIICLFVTVVILISVVTFGIFIYRYRVGSYRRIQGNG 450 (450)
Q Consensus 412 ~i~i~~~~~~~~~l~~~~~~~~~~~r~~~~~~~~~~~~~ 450 (450)
++|.+.+-+|+.+|+.+++.+|++++|+ +|..=|.++|
T Consensus 50 IVIGvVVGVGg~ill~il~lvf~~c~r~-kktdfidSdG 87 (154)
T PF04478_consen 50 IVIGVVVGVGGPILLGILALVFIFCIRR-KKTDFIDSDG 87 (154)
T ss_pred EEEEEEecccHHHHHHHHHhheeEEEec-ccCccccCCC
Confidence 4444433334444444433334333333 2334454443
No 19
>PF08693 SKG6: Transmembrane alpha-helix domain; InterPro: IPR014805 SKG6 and AXL2 are membrane proteins that show polarised intracellular localisation [, ]. This entry represents the highly conserved transmembrane alpha-helical domain found in these proteins [, ]. The full-length AXL2 protein has a negative regulatory function in cytokinesis [].
Probab=88.67 E-value=0.45 Score=32.33 Aligned_cols=31 Identities=23% Similarity=0.285 Sum_probs=16.0
Q ss_pred cceEEEEehhhHHHHHHHHHHhheee-eEeec
Q 043869 409 LWKNIVIICLFVTVVILISVVTFGIF-IYRYR 439 (450)
Q Consensus 409 ~~~~i~i~~~~~~~~~l~~~~~~~~~-~~r~~ 439 (450)
++..-+.+++++.+..++++++.++| +|||+
T Consensus 8 ~~~vaIa~~VvVPV~vI~~vl~~~l~~~~rR~ 39 (40)
T PF08693_consen 8 SNTVAIAVGVVVPVGVIIIVLGAFLFFWYRRK 39 (40)
T ss_pred CceEEEEEEEEechHHHHHHHHHHhheEEecc
Confidence 34455666666665555444433344 45544
No 20
>cd01099 PAN_AP_HGF Subfamily of PAN/APPLE-like domains; present in N-terminal (N) domains of plasminogen/hepatocyte growth factor proteins, and various proteins found in Bilateria, such as leech anti-platelet proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=87.89 E-value=1.9 Score=33.86 Aligned_cols=33 Identities=30% Similarity=0.694 Sum_probs=27.3
Q ss_pred cccCHHHHHHHHhc--CCCeEEEEe--cCCceEeeec
Q 043869 337 SEITKEKCKQACLE--DCNCEAALY--KNEECKMQRL 369 (450)
Q Consensus 337 ~~~s~~~C~~~CL~--nCsC~A~~y--~~g~C~~~~~ 369 (450)
...++++|.+.|++ +=.|.++.| .++.|.+-..
T Consensus 23 ~~~s~~~C~~~C~~~~~f~CrSf~y~~~~~~C~L~~~ 59 (80)
T cd01099 23 TVASLEECLRKCLEETEFTCRSFNYNYKSKECILSDE 59 (80)
T ss_pred ecCCHHHHHHHhCCCCCceEeEEEEEcCCCEEEEeCC
Confidence 34789999999999 889999998 4678987543
No 21
>PF07645 EGF_CA: Calcium-binding EGF domain; InterPro: IPR001881 A sequence of about forty amino-acid residues found in epidermal growth factor (EGF) has been shown [, , , , , ] to be present in a large number of membrane-bound and extracellular, mostly animal, proteins. Many of these proteins require calcium for their biological function and a calcium-binding site has been found at the N terminus of some EGF-like domains []. Calcium-binding may be crucial for numerous protein-protein interactions. For human coagulation factor IX it has been shown [] that the calcium-ligands form a pentagonal bipyramid. The first, third and fourth conserved negatively charged or polar residues are side chain ligands. The latter is possibly hydroxylated (see aspartic acid and asparagine hydroxylation site) []. A conserved aromatic residue, as well as the second conserved negative residue, are thought to be involved in stabilising the calcium-binding site. As in non-calcium binding EGF-like domains, there are six conserved cysteines and the structure of both types is very similar as calcium-binding induces only strictly local structural changes []. +------------------+ +---------+ | | | | nxnnC-x(3,14)-C-x(3,7)-CxxbxxxxaxC-x(1,6)-C-x(8,13)-Cx | | +------------------+ 'n': negatively charged or polar residue [DEQN] 'b': possibly beta-hydroxylated residue [DN] 'a': aromatic amino acid 'C': cysteine, involved in disulphide bond 'x': any amino acid. ; GO: 0005509 calcium ion binding; PDB: 2VJ3_A 1TOZ_A 1LMJ_A 1UZQ_A 1UZK_A 1UZJ_B 1UZP_A 1EMO_A 1EMN_A 2RR0_A ....
Probab=87.59 E-value=0.24 Score=33.99 Aligned_cols=32 Identities=28% Similarity=0.682 Sum_probs=25.3
Q ss_pred CCCCCc-CCCCCCcccccCCCCCCCcCCCCCee
Q 043869 265 EKCDPI-GLCGFNSFCVLNDQTPNCTCLPGFVA 296 (450)
Q Consensus 265 ~~C~~~-g~CG~~g~C~~~~~~~~C~C~~GF~~ 296 (450)
|+|... ..|..++.|......-.|.|++||+.
T Consensus 3 dEC~~~~~~C~~~~~C~N~~Gsy~C~C~~Gy~~ 35 (42)
T PF07645_consen 3 DECAEGPHNCPENGTCVNTEGSYSCSCPPGYEL 35 (42)
T ss_dssp STTTTTSSSSSTTSEEEEETTEEEEEESTTEEE
T ss_pred cccCCCCCcCCCCCEEEcCCCCEEeeCCCCcEE
Confidence 567764 47999999986444678999999994
No 22
>cd00053 EGF Epidermal growth factor domain, found in epidermal growth factor (EGF) presents in a large number of proteins, mostly animal; the list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied; the functional significance of EGF-like domains in what appear to be unrelated proteins is not yet clear; a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase); the domain includes six cysteine residues which have been shown to be involved in disulfide bonds; the main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet; Subdomains between the conserved cysteines vary in length; the region between the 5th and 6th cysteine contains two conserved glycines of which at least one is present in most EGF-like domains; a subset of these bind calcium.
Probab=87.34 E-value=0.45 Score=30.25 Aligned_cols=30 Identities=27% Similarity=0.723 Sum_probs=22.7
Q ss_pred CCCcCCCCCCcccccCCCCCCCcCCCCCee
Q 043869 267 CDPIGLCGFNSFCVLNDQTPNCTCLPGFVA 296 (450)
Q Consensus 267 C~~~g~CG~~g~C~~~~~~~~C~C~~GF~~ 296 (450)
|.....|..++.|........|.|++||..
T Consensus 2 C~~~~~C~~~~~C~~~~~~~~C~C~~g~~g 31 (36)
T cd00053 2 CAASNPCSNGGTCVNTPGSYRCVCPPGYTG 31 (36)
T ss_pred CCCCCCCCCCCEEecCCCCeEeECCCCCcc
Confidence 443467888999986444788999999964
No 23
>PF15102 TMEM154: TMEM154 protein family
Probab=86.35 E-value=0.92 Score=39.91 Aligned_cols=27 Identities=37% Similarity=0.581 Sum_probs=12.1
Q ss_pred EEehhhHHHHHHHHHHhheeeeEeecc
Q 043869 414 VIICLFVTVVILISVVTFGIFIYRYRV 440 (450)
Q Consensus 414 ~i~~~~~~~~~l~~~~~~~~~~~r~~~ 440 (450)
+++.+++.+++|+++++.+++.+|||.
T Consensus 61 IlIP~VLLvlLLl~vV~lv~~~kRkr~ 87 (146)
T PF15102_consen 61 ILIPLVLLVLLLLSVVCLVIYYKRKRT 87 (146)
T ss_pred EeHHHHHHHHHHHHHHHheeEEeeccc
Confidence 333335555555554444443444443
No 24
>smart00179 EGF_CA Calcium-binding EGF-like domain.
Probab=82.08 E-value=1.1 Score=29.28 Aligned_cols=31 Identities=26% Similarity=0.669 Sum_probs=22.9
Q ss_pred CCCCCcCCCCCCcccccCCCCCCCcCCCCCe
Q 043869 265 EKCDPIGLCGFNSFCVLNDQTPNCTCLPGFV 295 (450)
Q Consensus 265 ~~C~~~g~CG~~g~C~~~~~~~~C~C~~GF~ 295 (450)
+.|.....|...+.|........|.|++||.
T Consensus 3 ~~C~~~~~C~~~~~C~~~~g~~~C~C~~g~~ 33 (39)
T smart00179 3 DECASGNPCQNGGTCVNTVGSYRCECPPGYT 33 (39)
T ss_pred ccCcCCCCcCCCCEeECCCCCeEeECCCCCc
Confidence 4565545788888998643356799999997
No 25
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=80.63 E-value=1.9 Score=35.40 Aligned_cols=14 Identities=14% Similarity=0.090 Sum_probs=8.2
Q ss_pred HHHHhheeeeEeec
Q 043869 426 ISVVTFGIFIYRYR 439 (450)
Q Consensus 426 ~~~~~~~~~~~r~~ 439 (450)
|+.++.|||++|||
T Consensus 82 lv~~l~w~f~~r~k 95 (96)
T PTZ00382 82 LVGFLCWWFVCRGK 95 (96)
T ss_pred HHHHHhheeEEeec
Confidence 33455567777765
No 26
>cd00054 EGF_CA Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.
Probab=78.41 E-value=1.8 Score=27.85 Aligned_cols=32 Identities=25% Similarity=0.670 Sum_probs=22.7
Q ss_pred CCCCCcCCCCCCcccccCCCCCCCcCCCCCee
Q 043869 265 EKCDPIGLCGFNSFCVLNDQTPNCTCLPGFVA 296 (450)
Q Consensus 265 ~~C~~~g~CG~~g~C~~~~~~~~C~C~~GF~~ 296 (450)
+.|.....|...+.|........|.|++||.-
T Consensus 3 ~~C~~~~~C~~~~~C~~~~~~~~C~C~~g~~g 34 (38)
T cd00054 3 DECASGNPCQNGGTCVNTVGSYRCSCPPGYTG 34 (38)
T ss_pred ccCCCCCCcCCCCEeECCCCCeEeECCCCCcC
Confidence 45654356877888976444567999999863
No 27
>PF01683 EB: EB module; InterPro: IPR006149 The EB domain has no known function. It is found in several Caenorhabditis sp. and Drosophila sp. proteins. The domain contains 8 conserved cysteines that probably form four disulphide bridges and is found associated with kunitz domains IPR002223 from INTERPRO
Probab=76.48 E-value=2.5 Score=30.23 Aligned_cols=33 Identities=33% Similarity=0.755 Sum_probs=27.3
Q ss_pred ccCCCCCCcCCCCCCcccccCCCCCCCcCCCCCeecc
Q 043869 262 STSEKCDPIGLCGFNSFCVLNDQTPNCTCLPGFVAIS 298 (450)
Q Consensus 262 ~p~~~C~~~g~CG~~g~C~~~~~~~~C~C~~GF~~~~ 298 (450)
.|-+.|+...-|-.++.|.. ..|.|++||.+..
T Consensus 17 ~~g~~C~~~~qC~~~s~C~~----g~C~C~~g~~~~~ 49 (52)
T PF01683_consen 17 QPGESCESDEQCIGGSVCVN----GRCQCPPGYVEVG 49 (52)
T ss_pred CCCCCCCCcCCCCCcCEEcC----CEeECCCCCEecC
Confidence 35678999999999999953 6899999998753
No 28
>PF02009 Rifin_STEVOR: Rifin/stevor family; InterPro: IPR002858 Malaria is still a major cause of mortality in many areas of the world. Plasmodium falciparum causes the most severe human form of the disease and is responsible for most fatalities. Severe cases of malaria can occur when the parasite invades and then proliferates within red blood cell erythrocytes. The parasite produces many variant antigenic proteins, encoded by multigene families, which are present on the surface of the infected erythrocyte and play important roles in virulence. A crucial survival mechanism for the malaria parasite is its ability to evade the immune response by switching these variant surface antigens. The high virulence of P. falciparum relative to other malarial parasites is in large part due to the fact that in this organism many of these surface antigens mediate the binding of infected erythrocytes to the vascular endothelium (cytoadherence) and non-infected erythrocytes (rosetting). This can lead to the accumulation of infected cells in the vasculature of a variety of organs, blocking the blood flow and reducing the oxygen supply. Clinical symptoms of severe infection can include fever, progressive anaemia, multi-organ dysfunction and coma. For more information see []. Several multicopy gene families have been described in Plasmodium falciparum, including the stevor family of subtelomeric open reading frames and the rif interspersed repetitive elements. Both families contain three predicted transmembrane segments. It has been proposed that stevor and rif are members of a larger superfamily that code for variant surface antigens [].
Probab=75.64 E-value=0.51 Score=46.84 Aligned_cols=32 Identities=22% Similarity=0.515 Sum_probs=18.9
Q ss_pred EehhhHHHHHHHHHHhheeeeEeec--cceeecc
Q 043869 415 IICLFVTVVILISVVTFGIFIYRYR--VGSYRRI 446 (450)
Q Consensus 415 i~~~~~~~~~l~~~~~~~~~~~r~~--~~~~~~~ 446 (450)
|+++++.+++++.+++++|+|+|.| ++..+|+
T Consensus 258 I~aSiiaIliIVLIMvIIYLILRYRRKKKmkKKl 291 (299)
T PF02009_consen 258 IIASIIAILIIVLIMVIIYLILRYRRKKKMKKKL 291 (299)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHHHhhhhHHH
Confidence 4456666666666777778876533 3444444
No 29
>PF12947 EGF_3: EGF domain; InterPro: IPR024731 This entry represents an EGF domain found in the the C terminus of malarial parasite merozoite surface protein 1 [], as well as other proteins.; PDB: 2NPR_A 1N1I_C 1B9W_A 1YO8_A 2RHP_A.
Probab=74.95 E-value=0.65 Score=30.86 Aligned_cols=27 Identities=33% Similarity=0.700 Sum_probs=19.3
Q ss_pred cCCCCCCcccccCCCCCCCcCCCCCee
Q 043869 270 IGLCGFNSFCVLNDQTPNCTCLPGFVA 296 (450)
Q Consensus 270 ~g~CG~~g~C~~~~~~~~C~C~~GF~~ 296 (450)
.+-|.++..|+.......|.|.+||+-
T Consensus 5 ~~~C~~nA~C~~~~~~~~C~C~~Gy~G 31 (36)
T PF12947_consen 5 NGGCHPNATCTNTGGSYTCTCKPGYEG 31 (36)
T ss_dssp GGGS-TTCEEEE-TTSEEEEE-CEEEC
T ss_pred CCCCCCCcEeecCCCCEEeECCCCCcc
Confidence 356889999987555778999999963
No 30
>PF01102 Glycophorin_A: Glycophorin A; InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=74.06 E-value=0.57 Score=40.15 Aligned_cols=28 Identities=21% Similarity=0.279 Sum_probs=11.9
Q ss_pred EEEehhhHHHHHHHHHHhheeeeEeeccce
Q 043869 413 IVIICLFVTVVILISVVTFGIFIYRYRVGS 442 (450)
Q Consensus 413 i~i~~~~~~~~~l~~~~~~~~~~~r~~~~~ 442 (450)
.||++++.|++.+ |+++.|+++|+|++.
T Consensus 68 ~Ii~gv~aGvIg~--Illi~y~irR~~Kk~ 95 (122)
T PF01102_consen 68 GIIFGVMAGVIGI--ILLISYCIRRLRKKS 95 (122)
T ss_dssp HHHHHHHHHHHHH--HHHHHHHHHHHS---
T ss_pred ehhHHHHHHHHHH--HHHHHHHHHHHhccC
Confidence 3444444444333 234456666655443
No 31
>PF01299 Lamp: Lysosome-associated membrane glycoprotein (Lamp); InterPro: IPR002000 Lysosome-associated membrane glycoproteins (lamp) [] are integral membrane proteins, specific to lysosomes, and whose exact biological function is not yet clear. Structurally, the lamp proteins consist of two internally homologous lysosome-luminal domains separated by a proline-rich hinge region; at the C-terminal extremity there is a transmembrane region (TM) followed by a very short cytoplasmic tail (C). In each of the duplicated domains, there are two conserved disulphide bonds. This structure is schematically represented in the figure below. +-----+ +-----+ +-----+ +-----+ | | | | | | | | xCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxxxCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxx +--------------------------++Hinge++--------------------------++TM++C+ In mammals, there are two closely related types of lamp: lamp-1 and lamp-2, which form major components of the lysosome membrane. In chicken lamp-1 is known as LEP100. Also included in this entry is the macrophage protein CD68 (or macrosialin) [] is a heavily glycosylated integral membrane protein whose structure consists of a mucin-like domain followed by a proline-rich hinge; a single lamp-like domain; a transmembrane region and a short cytoplasmic tail. Similar to CD68, mammalian lamp-3, which is expressed in lymphoid organs, dendritic cells and in lung, contains all the C-terminal regions but lacks the N-terminal lamp-like region []. In a lamp-family protein from nematodes [] only the part C-terminal to the hinge is conserved. ; GO: 0016020 membrane
Probab=73.13 E-value=2.5 Score=42.14 Aligned_cols=19 Identities=32% Similarity=0.488 Sum_probs=13.5
Q ss_pred HHhheeeeEeeccce-eecc
Q 043869 428 VVTFGIFIYRYRVGS-YRRI 446 (450)
Q Consensus 428 ~~~~~~~~~r~~~~~-~~~~ 446 (450)
|++++|+|.|||.++ |+.|
T Consensus 287 ivLiaYli~Rrr~~~gYq~~ 306 (306)
T PF01299_consen 287 IVLIAYLIGRRRSRAGYQSI 306 (306)
T ss_pred HHHHhheeEecccccccccC
Confidence 455578888888766 7764
No 32
>PF02439 Adeno_E3_CR2: Adenovirus E3 region protein CR2; InterPro: IPR003470 Early region 3 (E3) of human adenoviruses (Ads) codes for proteins that appear to control viral interactions with the host []. This region called CR1 (conserved region 1) [] is found three times in Human adenovirus 19 (a subgroup D adenovirus) 49 kDa protein in the E3 region. CR1 is also found in the 20.1 Kd protein of subgroup B adenoviruses. The function of this 80 amino acid region is unknown. This region is probably a divergent immunoglobulin domain.
Probab=70.81 E-value=0.97 Score=30.25 Aligned_cols=29 Identities=17% Similarity=0.089 Sum_probs=13.0
Q ss_pred EehhhHHHHHHHHHHhheeeeEeecccee
Q 043869 415 IICLFVTVVILISVVTFGIFIYRYRVGSY 443 (450)
Q Consensus 415 i~~~~~~~~~l~~~~~~~~~~~r~~~~~~ 443 (450)
|+++|+..++++++++..|-..+||.+++
T Consensus 8 IIv~V~vg~~iiii~~~~YaCcykk~~~~ 36 (38)
T PF02439_consen 8 IIVAVVVGMAIIIICMFYYACCYKKHRRQ 36 (38)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHcccccc
Confidence 33344444455555555554434443333
No 33
>PF06697 DUF1191: Protein of unknown function (DUF1191); InterPro: IPR010605 This family contains hypothetical plant proteins of unknown function.
Probab=69.10 E-value=3.6 Score=40.23 Aligned_cols=36 Identities=11% Similarity=0.108 Sum_probs=19.4
Q ss_pred EEEehhhHHHHHHHHHHhheee-eEeeccceeecccC
Q 043869 413 IVIICLFVTVVILISVVTFGIF-IYRYRVGSYRRIQG 448 (450)
Q Consensus 413 i~i~~~~~~~~~l~~~~~~~~~-~~r~~~~~~~~~~~ 448 (450)
.+|.++++|+++|.++.+.+.. .+-+|++|-++|+.
T Consensus 214 ~iv~g~~~G~~~L~ll~~lv~~~vr~krk~k~~eMEr 250 (278)
T PF06697_consen 214 KIVVGVVGGVVLLGLLSLLVAMLVRYKRKKKIEEMER 250 (278)
T ss_pred EEEEEehHHHHHHHHHHHHHHhhhhhhHHHHHHHHHH
Confidence 3455667777776655443333 33345555555543
No 34
>PF00008 EGF: EGF-like domain This is a sub-family of the Pfam entry This is a sub-family of the Pfam entry; InterPro: IPR006209 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length.; GO: 0005515 protein binding; PDB: 1WHE_A 1CCF_A 1APO_A 1WHF_A 2VJ3_A 1TOZ_A 4D90_B 3CFW_A 1EDM_B 1IXA_A ....
Probab=68.96 E-value=1.1 Score=28.73 Aligned_cols=24 Identities=25% Similarity=0.636 Sum_probs=18.9
Q ss_pred CCCCCcccccCC-CCCCCcCCCCCe
Q 043869 272 LCGFNSFCVLND-QTPNCTCLPGFV 295 (450)
Q Consensus 272 ~CG~~g~C~~~~-~~~~C~C~~GF~ 295 (450)
.|...|.|.... ....|.|++||.
T Consensus 5 ~C~n~g~C~~~~~~~y~C~C~~G~~ 29 (32)
T PF00008_consen 5 PCQNGGTCIDLPGGGYTCECPPGYT 29 (32)
T ss_dssp SSTTTEEEEEESTSEEEEEEBTTEE
T ss_pred cCCCCeEEEeCCCCCEEeECCCCCc
Confidence 677788887644 467899999986
No 35
>PF12661 hEGF: Human growth factor-like EGF; PDB: 2YGQ_A 2E26_A 3A7Q_A 2YGP_A 2YGO_A 1HRE_A 1HAE_A 1HAF_A 1HRF_A.
Probab=68.76 E-value=1.1 Score=22.89 Aligned_cols=9 Identities=44% Similarity=1.442 Sum_probs=6.6
Q ss_pred CCcCCCCCe
Q 043869 287 NCTCLPGFV 295 (450)
Q Consensus 287 ~C~C~~GF~ 295 (450)
.|.|++||.
T Consensus 1 ~C~C~~G~~ 9 (13)
T PF12661_consen 1 TCQCPPGWT 9 (13)
T ss_dssp EEEE-TTEE
T ss_pred CccCcCCCc
Confidence 489999986
No 36
>smart00181 EGF Epidermal growth factor-like domain.
Probab=67.61 E-value=4.4 Score=25.89 Aligned_cols=25 Identities=28% Similarity=0.834 Sum_probs=18.8
Q ss_pred CCCCCCcccccCCCCCCCcCCCCCee
Q 043869 271 GLCGFNSFCVLNDQTPNCTCLPGFVA 296 (450)
Q Consensus 271 g~CG~~g~C~~~~~~~~C~C~~GF~~ 296 (450)
..|... .|........|.|++||..
T Consensus 6 ~~C~~~-~C~~~~~~~~C~C~~g~~g 30 (35)
T smart00181 6 GPCSNG-TCINTPGSYTCSCPPGYTG 30 (35)
T ss_pred CCCCCC-EEECCCCCeEeECCCCCcc
Confidence 456666 8876445788999999975
No 37
>PF14575 EphA2_TM: Ephrin type-A receptor 2 transmembrane domain; PDB: 3KUL_A 2XVD_A 2VX1_A 2VWV_A 2VX0_A 2VWY_A 2VWZ_A 2VWW_A 2VWU_A 2VWX_A ....
Probab=66.58 E-value=1.5 Score=34.36 Aligned_cols=25 Identities=28% Similarity=0.458 Sum_probs=12.9
Q ss_pred EehhhHHHHHHHHHHhheeeeEeec
Q 043869 415 IICLFVTVVILISVVTFGIFIYRYR 439 (450)
Q Consensus 415 i~~~~~~~~~l~~~~~~~~~~~r~~ 439 (450)
+.++++|+++++++++..++++||+
T Consensus 3 i~~~~~g~~~ll~~v~~~~~~~rr~ 27 (75)
T PF14575_consen 3 IASIIVGVLLLLVLVIIVIVCFRRC 27 (75)
T ss_dssp HHHHHHHHHHHHHHHHHHHCCCTT-
T ss_pred EehHHHHHHHHHHhheeEEEEEeeE
Confidence 4455666666655555444444443
No 38
>PTZ00046 rifin; Provisional
Probab=63.78 E-value=1.7 Score=43.89 Aligned_cols=24 Identities=29% Similarity=0.609 Sum_probs=15.6
Q ss_pred EehhhHHHHHHHHHHhheeeeEee
Q 043869 415 IICLFVTVVILISVVTFGIFIYRY 438 (450)
Q Consensus 415 i~~~~~~~~~l~~~~~~~~~~~r~ 438 (450)
|++++++.++++.+++++|++.|.
T Consensus 317 IiaSiiAIvVIVLIMvIIYLILRY 340 (358)
T PTZ00046 317 IIASIVAIVVIVLIMVIIYLILRY 340 (358)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHh
Confidence 455666666666667777886553
No 39
>TIGR01477 RIFIN variant surface antigen, rifin family. This model represents the rifin branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of rifin sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 20 bits.
Probab=63.28 E-value=1.8 Score=43.67 Aligned_cols=33 Identities=18% Similarity=0.410 Sum_probs=18.9
Q ss_pred EehhhHHHHHHHHHHhheeeeEe--eccceeeccc
Q 043869 415 IICLFVTVVILISVVTFGIFIYR--YRVGSYRRIQ 447 (450)
Q Consensus 415 i~~~~~~~~~l~~~~~~~~~~~r--~~~~~~~~~~ 447 (450)
|++++++.++++.+++++|++.| ||++-.+|||
T Consensus 312 IiaSiIAIvvIVLIMvIIYLILRYRRKKKMkKKLQ 346 (353)
T TIGR01477 312 IIASIIAILIIVLIMVIIYLILRYRRKKKMKKKLQ 346 (353)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHhhhcchhHHHHH
Confidence 45566666666666777788655 3333334443
No 40
>PF12662 cEGF: Complement Clr-like EGF-like
Probab=62.25 E-value=4 Score=24.60 Aligned_cols=11 Identities=36% Similarity=1.135 Sum_probs=9.5
Q ss_pred CCcCCCCCeec
Q 043869 287 NCTCLPGFVAI 297 (450)
Q Consensus 287 ~C~C~~GF~~~ 297 (450)
.|+|++||+..
T Consensus 3 ~C~C~~Gy~l~ 13 (24)
T PF12662_consen 3 TCSCPPGYQLS 13 (24)
T ss_pred EeeCCCCCcCC
Confidence 69999999864
No 41
>PF13908 Shisa: Wnt and FGF inhibitory regulator
Probab=60.50 E-value=7.2 Score=35.62 Aligned_cols=20 Identities=10% Similarity=0.323 Sum_probs=10.7
Q ss_pred eEEEEehhhHHHHHHHHHHh
Q 043869 411 KNIVIICLFVTVVILISVVT 430 (450)
Q Consensus 411 ~~i~i~~~~~~~~~l~~~~~ 430 (450)
...||+++++++++++++++
T Consensus 77 ~~~iivgvi~~Vi~Iv~~Iv 96 (179)
T PF13908_consen 77 ITGIIVGVICGVIAIVVLIV 96 (179)
T ss_pred eeeeeeehhhHHHHHHHhHh
Confidence 34556666666655544333
No 42
>PF06024 DUF912: Nucleopolyhedrovirus protein of unknown function (DUF912); InterPro: IPR009261 This entry is represented by Autographa californica nuclear polyhedrosis virus (AcMNPV), Orf78; it is a family of uncharacterised viral proteins.
Probab=58.82 E-value=7.7 Score=32.09 Aligned_cols=9 Identities=22% Similarity=0.080 Sum_probs=5.5
Q ss_pred ceEEEEecC
Q 043869 382 DITFVKVDD 390 (450)
Q Consensus 382 ~~~yiKv~~ 390 (450)
.-+-+||+-
T Consensus 15 ~yIPLKLal 23 (101)
T PF06024_consen 15 DYIPLKLAL 23 (101)
T ss_pred cceeeeeec
Confidence 346677774
No 43
>PF06365 CD34_antigen: CD34/Podocalyxin family; InterPro: IPR013836 This family consists of several mammalian CD34 antigen proteins. The CD34 antigen is a human leukocyte membrane protein expressed specifically by lymphohematopoietic progenitor cells. CD34 is a phosphoprotein. Activation of protein kinase C (PKC) has been found to enhance CD34 phosphorylation [, ]. This family contains several eukaryotic podocalyxin proteins. Podocalyxin is a major membrane protein of the glomerular epithelium and is thought to be involved in maintenance of the architecture of the foot processes and filtration slits characteristic of this unique epithelium by virtue of its high negative charge. Podocalyxin functions as an anti-adhesin that maintains an open filtration pathway between neighbouring foot processes in the glomerular epithelium by charge repulsion [].
Probab=58.44 E-value=9.4 Score=35.68 Aligned_cols=36 Identities=14% Similarity=0.140 Sum_probs=19.8
Q ss_pred EEEehhhHHHHHHHH-HHhheeeeEeeccc--eeecccC
Q 043869 413 IVIICLFVTVVILIS-VVTFGIFIYRYRVG--SYRRIQG 448 (450)
Q Consensus 413 i~i~~~~~~~~~l~~-~~~~~~~~~r~~~~--~~~~~~~ 448 (450)
++|..++.|.++|++ +.+++||++.||.+ +-.|+.|
T Consensus 101 ~lI~lv~~g~~lLla~~~~~~Y~~~~Rrs~~~~~~rl~E 139 (202)
T PF06365_consen 101 TLIALVTSGSFLLLAILLGAGYCCHQRRSWSKKGQRLGE 139 (202)
T ss_pred EEEehHHhhHHHHHHHHHHHHHHhhhhccCCcchhhhcc
Confidence 555566666444444 45556777666543 3444443
No 44
>PF07974 EGF_2: EGF-like domain; InterPro: IPR013111 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length. This entry contains EGF domains found in a variety of extracellular and membrane proteins
Probab=57.84 E-value=7.7 Score=25.01 Aligned_cols=24 Identities=25% Similarity=0.708 Sum_probs=19.0
Q ss_pred CCCCCCcccccCCCCCCCcCCCCCee
Q 043869 271 GLCGFNSFCVLNDQTPNCTCLPGFVA 296 (450)
Q Consensus 271 g~CG~~g~C~~~~~~~~C~C~~GF~~ 296 (450)
..|...|.|... ...|.|.+||.-
T Consensus 6 ~~C~~~G~C~~~--~g~C~C~~g~~G 29 (32)
T PF07974_consen 6 NICSGHGTCVSP--CGRCVCDSGYTG 29 (32)
T ss_pred CccCCCCEEeCC--CCEEECCCCCcC
Confidence 469999999852 478999999863
No 45
>PF05454 DAG1: Dystroglycan (Dystrophin-associated glycoprotein 1); InterPro: IPR008465 Dystroglycan is one of the dystrophin-associated glycoproteins, which is encoded by a 5.5 kb transcript in Homo sapiens. The protein product is cleaved into two non-covalently associated subunits, [alpha] (N-terminal) and [beta] (C-terminal). In skeletal muscle the dystroglycan complex works as a transmembrane linkage between the extracellular matrix and the cytoskeleton [alpha]-dystroglycan is extracellular and binds to merosin ([alpha]-2 laminin) in the basement membrane, while [beta]-dystroglycan is a transmembrane protein and binds to dystrophin, which is a large rod-like cytoskeletal protein, absent in Duchenne muscular dystrophy patients. Dystrophin binds to intracellular actin cables. In this way, the dystroglycan complex, which links the extracellular matrix to the intracellular actin cables, is thought to provide structural integrity in muscle tissues. The dystroglycan complex is also known to serve as an agrin receptor in muscle, where it may regulate agrin-induced acetylcholine receptor clustering at the neuromuscular junction. There is also evidence which suggests the function of dystroglycan as a part of the signal transduction pathway because it is shown that Grb2, a mediator of the Ras-related signal pathway, can interact with the cytoplasmic domain of dystroglycan. In general, aberrant expression of dystrophin-associated protein complex underlies the pathogenesis of Duchenne muscular dystrophy, Becker muscular dystrophy and severe childhood autosomal recessive muscular dystrophy. Interestingly, no genetic disease has been described for either [alpha]- or [beta]-dystroglycan. Dystroglycan is widely distributed in non-muscle tissues as well as in muscle tissues. During epithelial morphogenesis of kidney, the dystroglycan complex is shown to act as a receptor for the basement membrane. Dystroglycan expression in Mus musculus brain and neural retina has also been reported. However, the physiological role of dystroglycan in non-muscle tissues has remained unclear [].; PDB: 1EG4_P.
Probab=54.29 E-value=4.2 Score=40.22 Aligned_cols=29 Identities=17% Similarity=0.272 Sum_probs=0.0
Q ss_pred ceEEEEehhhHHHHHHHHHHhheeeeEee
Q 043869 410 WKNIVIICLFVTVVILISVVTFGIFIYRY 438 (450)
Q Consensus 410 ~~~i~i~~~~~~~~~l~~~~~~~~~~~r~ 438 (450)
....+|.++|+.+++|+++++++++.+||
T Consensus 145 yL~T~IpaVVI~~iLLIA~iIa~icyrrk 173 (290)
T PF05454_consen 145 YLHTFIPAVVIAAILLIAGIIACICYRRK 173 (290)
T ss_dssp -----------------------------
T ss_pred hHHHHHHHHHHHHHHHHHHHHHHHhhhhh
Confidence 33444566676666666555554443333
No 46
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=53.91 E-value=49 Score=33.88 Aligned_cols=71 Identities=15% Similarity=0.330 Sum_probs=42.7
Q ss_pred CCcEEEEecCCC----------------CCCCCccEEEE-ecCCcEEEEeCCCCceEEeecCCC-----Cc-----cEEE
Q 043869 70 EKTVVWTANRDN----------------PPVSSNATLMF-NSEGRIVLRSGEQGQNSIIADNSQ-----SA-----SSAS 122 (450)
Q Consensus 70 ~~tvVW~ANr~~----------------Pv~~~~~~L~l-~~~G~LvL~d~~~g~~vW~st~~~-----~~-----~~a~ 122 (450)
...++|...-.. |+. .+..+-+ +.+|.|+-.|..+|.++| +.... .. ....
T Consensus 88 tG~~~W~~~~~~~~~~~~~~~~~~~~~~~~v-~~~~v~v~~~~g~l~ald~~tG~~~W-~~~~~~~~~ssP~v~~~~v~v 165 (394)
T PRK11138 88 TGKEIWSVDLSEKDGWFSKNKSALLSGGVTV-AGGKVYIGSEKGQVYALNAEDGEVAW-QTKVAGEALSRPVVSDGLVLV 165 (394)
T ss_pred CCcEeeEEcCCCcccccccccccccccccEE-ECCEEEEEcCCCEEEEEECCCCCCcc-cccCCCceecCCEEECCEEEE
Confidence 467899865432 222 2344545 467888877754799999 44321 11 1112
Q ss_pred EecCCCeEEEec-CCeeEEee
Q 043869 123 MLDSGSFVLHNS-DGKVIWQT 142 (450)
Q Consensus 123 LldsGNlVL~~~-~~~~lWQS 142 (450)
...+|.|+-.|. +|+++|+-
T Consensus 166 ~~~~g~l~ald~~tG~~~W~~ 186 (394)
T PRK11138 166 HTSNGMLQALNESDGAVKWTV 186 (394)
T ss_pred ECCCCEEEEEEccCCCEeeee
Confidence 245677888885 78899984
No 47
>KOG1219 consensus Uncharacterized conserved protein, contains laminin, cadherin and EGF domains [Signal transduction mechanisms]
Probab=53.18 E-value=15 Score=45.68 Aligned_cols=32 Identities=16% Similarity=0.539 Sum_probs=21.1
Q ss_pred CCCCCcCCCCCCcccccCCC-CCCCcCCCCCeec
Q 043869 265 EKCDPIGLCGFNSFCVLNDQ-TPNCTCLPGFVAI 297 (450)
Q Consensus 265 ~~C~~~g~CG~~g~C~~~~~-~~~C~C~~GF~~~ 297 (450)
+.|. ...|---|.|+.... .-.|.||+-|.-.
T Consensus 3865 d~C~-~npCqhgG~C~~~~~ggy~CkCpsqysG~ 3897 (4289)
T KOG1219|consen 3865 DPCN-DNPCQHGGTCISQPKGGYKCKCPSQYSGN 3897 (4289)
T ss_pred cccc-cCcccCCCEecCCCCCceEEeCcccccCc
Confidence 4454 356666677775433 5689999988754
No 48
>PF13360 PQQ_2: PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=50.35 E-value=1.5e+02 Score=27.32 Aligned_cols=75 Identities=17% Similarity=0.289 Sum_probs=46.5
Q ss_pred cCCCcEEEEecCCCCCCC----CccE-EEEecCCcEEEEeCCCCceEEeecC-C---CC----------ccEEEE-ecCC
Q 043869 68 IPEKTVVWTANRDNPPVS----SNAT-LMFNSEGRIVLRSGEQGQNSIIADN-S---QS----------ASSASM-LDSG 127 (450)
Q Consensus 68 ~~~~tvVW~ANr~~Pv~~----~~~~-L~l~~~G~LvL~d~~~g~~vW~st~-~---~~----------~~~a~L-ldsG 127 (450)
+...+++|....+.++.. .... +..+.+|.|...|..+|.++|.... . .. ...+.+ ..+|
T Consensus 53 ~~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~g 132 (238)
T PF13360_consen 53 AKTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYALDAKTGKVLWSIYLTSSPPAGVRSSSSPAVDGDRLYVGTSSG 132 (238)
T ss_dssp TTTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEEEETTTSCEEEEEEE-SSCTCSTB--SEEEEETTEEEEEETCS
T ss_pred CCCCCEEEEeeccccccceeeecccccccccceeeeEecccCCcceeeeeccccccccccccccCceEecCEEEEEeccC
Confidence 346789999886655331 2333 4445678888888438999995211 1 10 111222 3488
Q ss_pred CeEEEe-cCCeeEEee
Q 043869 128 SFVLHN-SDGKVIWQT 142 (450)
Q Consensus 128 NlVL~~-~~~~~lWQS 142 (450)
.++..| .+|+.+|+-
T Consensus 133 ~l~~~d~~tG~~~w~~ 148 (238)
T PF13360_consen 133 KLVALDPKTGKLLWKY 148 (238)
T ss_dssp EEEEEETTTTEEEEEE
T ss_pred cEEEEecCCCcEEEEe
Confidence 999998 478999984
No 49
>PF09064 Tme5_EGF_like: Thrombomodulin like fifth domain, EGF-like; InterPro: IPR015149 This domain adopts a fold similar to other EGF domains, with a flat major and a twisted minor beta sheet. Disulphide pairing, however, is not of the usual 1-3, 2-4, 5-6 type; rather 1-2, 3-4, 5-6 pairing is found. Its extended major sheet (strands beta-2 and beta-3 and the connecting loop) projects into thrombin's active site groove. This domain is required for interaction of thrombomodulin with thrombin, and subsequent activation of protein-C []. ; GO: 0004888 transmembrane signaling receptor activity, 0016021 integral to membrane
Probab=49.77 E-value=9.4 Score=24.96 Aligned_cols=18 Identities=22% Similarity=0.648 Sum_probs=12.7
Q ss_pred ccccCCCCCCCcCCCCCee
Q 043869 278 FCVLNDQTPNCTCLPGFVA 296 (450)
Q Consensus 278 ~C~~~~~~~~C~C~~GF~~ 296 (450)
.|+. ++...|.||.||..
T Consensus 11 ~CDp-n~~~~C~CPeGyIl 28 (34)
T PF09064_consen 11 DCDP-NSPGQCFCPEGYIL 28 (34)
T ss_pred ccCC-CCCCceeCCCceEe
Confidence 4544 23568999999976
No 50
>PF15102 TMEM154: TMEM154 protein family
Probab=48.12 E-value=16 Score=32.31 Aligned_cols=31 Identities=13% Similarity=0.426 Sum_probs=20.0
Q ss_pred EehhhHHHHHH-HHHHhheeeeEeeccceeec
Q 043869 415 IICLFVTVVIL-ISVVTFGIFIYRYRVGSYRR 445 (450)
Q Consensus 415 i~~~~~~~~~l-~~~~~~~~~~~r~~~~~~~~ 445 (450)
|+-+++..++| +.++++++++.+.|+||.++
T Consensus 58 iLmIlIP~VLLvlLLl~vV~lv~~~kRkr~K~ 89 (146)
T PF15102_consen 58 ILMILIPLVLLVLLLLSVVCLVIYYKRKRTKQ 89 (146)
T ss_pred EEEEeHHHHHHHHHHHHHHHheeEEeecccCC
Confidence 44555564444 55566666688888888765
No 51
>PF12768 Rax2: Cortical protein marker for cell polarity
Probab=46.47 E-value=12 Score=37.06 Aligned_cols=27 Identities=7% Similarity=0.173 Sum_probs=17.3
Q ss_pred CcceEEEEehhhHHHHHHHHHHhheee
Q 043869 408 GLWKNIVIICLFVTVVILISVVTFGIF 434 (450)
Q Consensus 408 ~~~~~i~i~~~~~~~~~l~~~~~~~~~ 434 (450)
+-.+++|.+++.+|.+.|+.++.+++.
T Consensus 226 ~G~VVlIslAiALG~v~ll~l~Gii~~ 252 (281)
T PF12768_consen 226 RGFVVLISLAIALGTVFLLVLIGIILA 252 (281)
T ss_pred ceEEEEEehHHHHHHHHHHHHHHHHHH
Confidence 334567777888888777766544443
No 52
>PF14670 FXa_inhibition: Coagulation Factor Xa inhibitory site; PDB: 3Q3K_B 1NFY_B 1LQD_A 1G2L_B 1IQF_L 2UWP_B 2VH6_B 3KQC_L 2P93_L 2BQW_A ....
Probab=45.26 E-value=5.6 Score=26.41 Aligned_cols=21 Identities=29% Similarity=0.768 Sum_probs=13.3
Q ss_pred ccccCCCCCCCcCCCCCeecc
Q 043869 278 FCVLNDQTPNCTCLPGFVAIS 298 (450)
Q Consensus 278 ~C~~~~~~~~C~C~~GF~~~~ 298 (450)
+|........|+|++||....
T Consensus 11 ~C~~~~g~~~C~C~~Gy~L~~ 31 (36)
T PF14670_consen 11 ICVNTPGSYRCSCPPGYKLAE 31 (36)
T ss_dssp EEEEETTSEEEE-STTEEE-T
T ss_pred CCccCCCceEeECCCCCEECc
Confidence 455433367899999998753
No 53
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=45.17 E-value=1.2e+02 Score=30.86 Aligned_cols=48 Identities=15% Similarity=0.144 Sum_probs=27.7
Q ss_pred cCCcEEEEeCCCCceEEeecCC------CC-----ccEEEEecCCCeEEEec-CCeeEEee
Q 043869 94 SEGRIVLRSGEQGQNSIIADNS------QS-----ASSASMLDSGSFVLHNS-DGKVIWQT 142 (450)
Q Consensus 94 ~~G~LvL~d~~~g~~vW~st~~------~~-----~~~a~LldsGNlVL~~~-~~~~lWQS 142 (450)
.+|.|+..|..+|..+| +.+. .+ .......++|.|...|. +|+++|+-
T Consensus 302 ~~g~l~ald~~tG~~~W-~~~~~~~~~~~sp~v~~g~l~v~~~~G~l~~ld~~tG~~~~~~ 361 (394)
T PRK11138 302 QNDRVYALDTRGGVELW-SQSDLLHRLLTAPVLYNGYLVVGDSEGYLHWINREDGRFVAQQ 361 (394)
T ss_pred CCCeEEEEECCCCcEEE-cccccCCCcccCCEEECCEEEEEeCCCEEEEEECCCCCEEEEE
Confidence 46777766654677788 3321 11 11122356777777775 67788875
No 54
>cd05845 Ig2_L1-CAM_like Second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM) and similar proteins. Ig2_L1-CAM_like: domain similar to the second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM). L1 belongs to the L1 subfamily of cell adhesion molecules (CAMs) and is comprised of an extracellular region having six Ig-like domains, five fibronectin type III domains, a transmembrane region and an intracellular domain. L1 is primarily expressed in the nervous system and is involved in its development and function. L1 is associated with an X-linked recessive disorder, X-linked hydrocephalus, MASA syndrome, or spastic paraplegia type 1, that involves abnormalities of axonal growth.
Probab=44.77 E-value=35 Score=27.86 Aligned_cols=35 Identities=6% Similarity=0.127 Sum_probs=23.1
Q ss_pred cCCCcEEEEecCCCCCCCCccEEEEecCCcEEEEeC
Q 043869 68 IPEKTVVWTANRDNPPVSSNATLMFNSEGRIVLRSG 103 (450)
Q Consensus 68 ~~~~tvVW~ANr~~Pv~~~~~~L~l~~~G~LvL~d~ 103 (450)
.|..++.|+-+....+. ......++.+|+|.+.+-
T Consensus 31 ~P~P~i~W~~~~~~~i~-~~~Ri~~~~~GnL~fs~v 65 (95)
T cd05845 31 AVPLRIYWMNSDLLHIT-QDERVSMGQNGNLYFANV 65 (95)
T ss_pred CCCCEEEEECCCCcccc-ccccEEECCCceEEEEEE
Confidence 45678889844433344 466777877888887653
No 55
>PF12690 BsuPI: Intracellular proteinase inhibitor; InterPro: IPR020481 BsuPI is a intracellular proteinase inhibitor that directly regulates the major intracellular proteinase (ISP-1) activity in vivo. It inhibits ISP-1 in the early stages of sporulation and then may be inactivated by a membrane-bound proteinase [].; PDB: 3ISY_A.
Probab=44.68 E-value=28 Score=27.55 Aligned_cols=16 Identities=13% Similarity=0.219 Sum_probs=9.2
Q ss_pred cEEEEeCCCCceEEeec
Q 043869 97 RIVLRSGEQGQNSIIAD 113 (450)
Q Consensus 97 ~LvL~d~~~g~~vW~st 113 (450)
+++|.|. +|..||.-+
T Consensus 27 D~~v~d~-~g~~vwrwS 42 (82)
T PF12690_consen 27 DFVVKDK-EGKEVWRWS 42 (82)
T ss_dssp EEEEE-T-T--EEEETT
T ss_pred EEEEECC-CCCEEEEec
Confidence 4778887 888888543
No 56
>KOG0291 consensus WD40-repeat-containing subunit of the 18S rRNA processing complex [RNA processing and modification]
Probab=43.37 E-value=4.7e+02 Score=29.52 Aligned_cols=56 Identities=20% Similarity=0.492 Sum_probs=39.4
Q ss_pred CccEEEEecCCcEEEEeCCCCce-EEeecC----------CCCccEEEEecCCCeEEEec-CCee-EEe
Q 043869 86 SNATLMFNSEGRIVLRSGEQGQN-SIIADN----------SQSASSASMLDSGSFVLHNS-DGKV-IWQ 141 (450)
Q Consensus 86 ~~~~L~l~~~G~LvL~d~~~g~~-vW~st~----------~~~~~~a~LldsGNlVL~~~-~~~~-lWQ 141 (450)
.-..+.-+.||.++.+.+.+|.+ ||.... +++++.....-+||.+|-.. +|+| .|.
T Consensus 352 ~i~~l~YSpDgq~iaTG~eDgKVKvWn~~SgfC~vTFteHts~Vt~v~f~~~g~~llssSLDGtVRAwD 420 (893)
T KOG0291|consen 352 RITSLAYSPDGQLIATGAEDGKVKVWNTQSGFCFVTFTEHTSGVTAVQFTARGNVLLSSSLDGTVRAWD 420 (893)
T ss_pred ceeeEEECCCCcEEEeccCCCcEEEEeccCceEEEEeccCCCceEEEEEEecCCEEEEeecCCeEEeee
Confidence 34668888999998887645554 894322 13566777889999999876 6765 665
No 57
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=43.29 E-value=1.7e+02 Score=29.38 Aligned_cols=50 Identities=20% Similarity=0.266 Sum_probs=24.9
Q ss_pred ecCCcEEEEeCCCCceEEeecCC--C-----CccEEEEecCCCeEEEec-CCeeEEee
Q 043869 93 NSEGRIVLRSGEQGQNSIIADNS--Q-----SASSASMLDSGSFVLHNS-DGKVIWQT 142 (450)
Q Consensus 93 ~~~G~LvL~d~~~g~~vW~st~~--~-----~~~~a~LldsGNlVL~~~-~~~~lWQS 142 (450)
+.+|.|+..|..+|.++|..... . ....-...++|.++..|. +++.+|+.
T Consensus 248 ~~~g~l~a~d~~tG~~~W~~~~~~~~~p~~~~~~vyv~~~~G~l~~~d~~tG~~~W~~ 305 (377)
T TIGR03300 248 SYQGRVAALDLRSGRVLWKRDASSYQGPAVDDNRLYVTDADGVVVALDRRSGSELWKN 305 (377)
T ss_pred EcCCEEEEEECCCCcEEEeeccCCccCceEeCCEEEEECCCCeEEEEECCCCcEEEcc
Confidence 45666666665356677732211 0 111112234566666664 45667753
No 58
>PF01436 NHL: NHL repeat; InterPro: IPR001258 The NHL repeat, named after NCL-1, HT2A and Lin-41, is found largely in a large number of eukaryotic and prokaryotic proteins. For example, the repeat is found in a variety of enzymes of the copper type II, ascorbate-dependent monooxygenase family which catalyse the C terminus alpha-amidation of biological peptides []. In many it occurs in tandem arrays, for example in the ringfinger beta-box, coiled-coil (RBCC) eukaryotic growth regulators []. The 'Brain Tumor' protein (Brat) is one such growth regulator that contains a 6-bladed NHL-repeat beta-propeller [, ]. The NHL repeats are also found in serine/threonine protein kinase (STPK) in diverse range of pathogenic bacteria. These STPK are transmembrane receptors with a intracellular N-terminal kinase domain and extracellular C-terminal sensor domain. In the STPK, PknD, from Mycobacterium tuberculosis, the sensor domain forms a rigid, six-bladed b-propeller composed of NHL repeats with a flexible tether to the transmembrane domain.; GO: 0005515 protein binding; PDB: 3FVZ_A 3FW0_A 1RWL_A 1RWI_A 1Q7F_A.
Probab=41.33 E-value=49 Score=20.23 Aligned_cols=21 Identities=14% Similarity=0.273 Sum_probs=14.6
Q ss_pred EEEEecCCcEEEEeCCCCceEE
Q 043869 89 TLMFNSEGRIVLRSGEQGQNSI 110 (450)
Q Consensus 89 ~L~l~~~G~LvL~d~~~g~~vW 110 (450)
-+.++.+|++++.|. ....||
T Consensus 6 gvav~~~g~i~VaD~-~n~rV~ 26 (28)
T PF01436_consen 6 GVAVDSDGNIYVADS-GNHRVQ 26 (28)
T ss_dssp EEEEETTSEEEEEEC-CCTEEE
T ss_pred EEEEeCCCCEEEEEC-CCCEEE
Confidence 366777888888886 555555
No 59
>PF13360 PQQ_2: PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=41.19 E-value=1.3e+02 Score=27.70 Aligned_cols=73 Identities=16% Similarity=0.335 Sum_probs=42.0
Q ss_pred CCcEEEEecC----CCCC--C-CCccEEEE-ecCCcEEEEeCCCCceEEeecCCCC---c-----cEEEE-ecCCCeEEE
Q 043869 70 EKTVVWTANR----DNPP--V-SSNATLMF-NSEGRIVLRSGEQGQNSIIADNSQS---A-----SSASM-LDSGSFVLH 132 (450)
Q Consensus 70 ~~tvVW~ANr----~~Pv--~-~~~~~L~l-~~~G~LvL~d~~~g~~vW~st~~~~---~-----~~a~L-ldsGNlVL~ 132 (450)
....+|..+- ..++ . ..+..+.+ +.+|.|+..|..+|..+|....... . ....+ ..+|.|+..
T Consensus 12 tG~~~W~~~~~~~~~~~~~~~~~~~~~v~~~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~~ 91 (238)
T PF13360_consen 12 TGKELWSYDLGPGIGGPVATAVPDGGRVYVASGDGNLYALDAKTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYAL 91 (238)
T ss_dssp TTEEEEEEECSSSCSSEEETEEEETTEEEEEETTSEEEEEETTTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEEE
T ss_pred CCCEEEEEECCCCCCCccceEEEeCCEEEEEcCCCEEEEEECCCCCEEEEeeccccccceeeecccccccccceeeeEec
Confidence 5677888753 2222 1 02333444 4888999888547899995432211 1 11222 234556667
Q ss_pred e-cCCeeEEee
Q 043869 133 N-SDGKVIWQT 142 (450)
Q Consensus 133 ~-~~~~~lWQS 142 (450)
| .+|+++|+.
T Consensus 92 d~~tG~~~W~~ 102 (238)
T PF13360_consen 92 DAKTGKVLWSI 102 (238)
T ss_dssp ETTTSCEEEEE
T ss_pred ccCCcceeeee
Confidence 7 678999995
No 60
>PF03302 VSP: Giardia variant-specific surface protein; InterPro: IPR005127 During infection, the intestinal protozoan parasite Giardia lamblia virus undergoes continuous antigenic variation which is determined by diversification of the parasite's major surface antigen, named VSP (variant surface protein).
Probab=41.19 E-value=22 Score=36.88 Aligned_cols=15 Identities=20% Similarity=0.007 Sum_probs=11.2
Q ss_pred HHHHHhheeeeEeec
Q 043869 425 LISVVTFGIFIYRYR 439 (450)
Q Consensus 425 l~~~~~~~~~~~r~~ 439 (450)
-|+-++.||||.|.|
T Consensus 382 glvGfLcWwf~crgk 396 (397)
T PF03302_consen 382 GLVGFLCWWFICRGK 396 (397)
T ss_pred HHHHHHhhheeeccc
Confidence 345577799999876
No 61
>PF12877 DUF3827: Domain of unknown function (DUF3827); InterPro: IPR024606 The function of the proteins in this entry is not currently known, but one of the human proteins (Q9HCM3 from SWISSPROT) has been implicated in pilocytic astrocytomas [, , ]. In the majority of cases of pilocytic astrocytomas a tandem duplication produces an in-frame fusion of the gene encoding this protein and the BRAF oncogene. The resulting fusion protein has constitutive BRAF kinase activity and is capable of transforming cells.
Probab=40.27 E-value=21 Score=38.88 Aligned_cols=30 Identities=13% Similarity=0.280 Sum_probs=17.0
Q ss_pred ceEEEEehhhHHHHHHHHHHhheee-eEeec
Q 043869 410 WKNIVIICLFVTVVILISVVTFGIF-IYRYR 439 (450)
Q Consensus 410 ~~~i~i~~~~~~~~~l~~~~~~~~~-~~r~~ 439 (450)
..++||+++++.++++++|++++++ ++|++
T Consensus 267 ~NlWII~gVlvPv~vV~~Iiiil~~~LCRk~ 297 (684)
T PF12877_consen 267 NNLWIIAGVLVPVLVVLLIIIILYWKLCRKN 297 (684)
T ss_pred CCeEEEehHhHHHHHHHHHHHHHHHHHhccc
Confidence 4456677777666665555444444 55543
No 62
>PF12191 stn_TNFRSF12A: Tumour necrosis factor receptor stn_TNFRSF12A_TNFR domain; InterPro: IPR022316 The tumour necrosis factor (TNF) receptor (TNFR) superfamily comprises more than 20 type-I transmembrane proteins. Family members are defined based on similarity in their extracellular domain - a region that contains many cysteine residues arranged in a specific repetitive pattern []. The cysteines allow formation of an extended rod-like structure, responsible for ligand binding []. Upon receptor activation, different intracellular signalling complexes are assembled for different members of the TNFR superfamily, depending on their intracellular domains and sequences []. Activation of TNFRs can therefore induce a range of disparate effects, including cell proliferation, differentiation, survival, or apoptotic cell death, depending upon the receptor involved []. TNFRs are widely distributed and play important roles in many crucial biological processes, such as lymphoid and neuronal development, innate and adaptive immunity, and maintenance of cellular homeostasis []. Drugs that manipulate their signalling have potential roles in the prevention and treatment of many diseases, such as viral infections, coronary heart disease, transplant rejection, and immune disease []. TNF receptor 12 (also known as TWEAK receptor, and fibroblast growth factor-inducible-14 (Fn14)) has been implicated in endothelial cell growth and migration []. The receptor may also play a role in cell-matrix interactions [].; PDB: 2KN0_A 2RPJ_A 2KMZ_A 2EQP_A.
Probab=40.14 E-value=8 Score=33.14 Aligned_cols=37 Identities=22% Similarity=0.434 Sum_probs=0.0
Q ss_pred EEEehhhHHHHHHHHHHhheeeeEe---eccceeecccCCC
Q 043869 413 IVIICLFVTVVILISVVTFGIFIYR---YRVGSYRRIQGNG 450 (450)
Q Consensus 413 i~i~~~~~~~~~l~~~~~~~~~~~r---~~~~~~~~~~~~~ 450 (450)
..|..+++++++++.++. .++++| ||.+-..-|+|+|
T Consensus 79 ~pi~~sal~v~lVl~lls-g~lv~rrcrrr~~~ttPIeeTg 118 (129)
T PF12191_consen 79 WPILGSALSVVLVLALLS-GFLVWRRCRRREKFTTPIEETG 118 (129)
T ss_dssp -----------------------------------------
T ss_pred hhhhhhHHHHHHHHHHHH-HHHHHhhhhccccCCCcccccC
Confidence 334445555554433322 233322 3334444677765
No 63
>PHA02887 EGF-like protein; Provisional
Probab=39.80 E-value=32 Score=29.17 Aligned_cols=35 Identities=26% Similarity=0.412 Sum_probs=25.1
Q ss_pred ccccCCCCC--CcCCCCCCcccccCC--CCCCCcCCCCCe
Q 043869 260 WPSTSEKCD--PIGLCGFNSFCVLND--QTPNCTCLPGFV 295 (450)
Q Consensus 260 w~~p~~~C~--~~g~CG~~g~C~~~~--~~~~C~C~~GF~ 295 (450)
++..-++|. ..++|= +|.|.+-. +.|.|.|++||.
T Consensus 79 ~~~hf~pC~~eyk~YCi-HG~C~yI~dL~epsCrC~~GYt 117 (126)
T PHA02887 79 NSMFFEKCKNDFNDFCI-NGECMNIIDLDEKFCICNKGYT 117 (126)
T ss_pred cccCccccChHhhCEee-CCEEEccccCCCceeECCCCcc
Confidence 344456785 367787 78997643 378999999985
No 64
>PF05337 CSF-1: Macrophage colony stimulating factor-1 (CSF-1); InterPro: IPR008001 Colony stimulating factor 1 (CSF-1) is a homodimeric polypeptide growth factor whose primary function is to regulate the survival, proliferation, differentiation, and function of cells of the mononuclear phagocytic lineage. This lineage includes mononuclear phagocytic precursors, blood monocytes, tissue macrophages, osteoclasts, and microglia of the brain, all of which possess cell surface receptors for CSF-1. The protein has also been linked with male fertility [] and mutations in the Csf-1 gene have been found to cause osteopetrosis and failure of tooth eruption [].; GO: 0005125 cytokine activity, 0008083 growth factor activity, 0016021 integral to membrane; PDB: 3EJJ_A.
Probab=39.42 E-value=9.8 Score=37.04 Aligned_cols=30 Identities=33% Similarity=0.470 Sum_probs=0.0
Q ss_pred hHHHHHHHHHHhheeeeEeeccceeecccC
Q 043869 419 FVTVVILISVVTFGIFIYRYRVGSYRRIQG 448 (450)
Q Consensus 419 ~~~~~~l~~~~~~~~~~~r~~~~~~~~~~~ 448 (450)
.+..|+++.++++.+++||+|++..++-|.
T Consensus 231 LVPSiILVLLaVGGLLfYr~rrRs~~e~q~ 260 (285)
T PF05337_consen 231 LVPSIILVLLAVGGLLFYRRRRRSHREPQT 260 (285)
T ss_dssp ------------------------------
T ss_pred cccchhhhhhhccceeeecccccccccccc
Confidence 344555666777778888877666665543
No 65
>PHA03099 epidermal growth factor-like protein (EGF-like protein); Provisional
Probab=38.84 E-value=16 Score=31.47 Aligned_cols=31 Identities=19% Similarity=0.200 Sum_probs=13.5
Q ss_pred hhhHHHHHHHHHHhheeeeEeeccceeeccc
Q 043869 417 CLFVTVVILISVVTFGIFIYRYRVGSYRRIQ 447 (450)
Q Consensus 417 ~~~~~~~~l~~~~~~~~~~~r~~~~~~~~~~ 447 (450)
.+++++++++++..+.++++|.-++|+..+|
T Consensus 104 ~~il~il~~i~is~~~~~~yr~~r~~~~~~~ 134 (139)
T PHA03099 104 PGIVLVLVGIIITCCLLSVYRFTRRTKLPLQ 134 (139)
T ss_pred hHHHHHHHHHHHHHHHHhhheeeecccCchh
Confidence 3444444444444444444444333343343
No 66
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=37.83 E-value=2.1e+02 Score=28.74 Aligned_cols=18 Identities=22% Similarity=0.602 Sum_probs=12.1
Q ss_pred cCCCeEEEec-CCeeEEee
Q 043869 125 DSGSFVLHNS-DGKVIWQT 142 (450)
Q Consensus 125 dsGNlVL~~~-~~~~lWQS 142 (450)
.+|+++.+|. +++.+|+.
T Consensus 249 ~~g~l~a~d~~tG~~~W~~ 267 (377)
T TIGR03300 249 YQGRVAALDLRSGRVLWKR 267 (377)
T ss_pred cCCEEEEEECCCCcEEEee
Confidence 4667777775 56778864
No 67
>PF15065 NCU-G1: Lysosomal transcription factor, NCU-G1
Probab=37.60 E-value=12 Score=38.13 Aligned_cols=32 Identities=19% Similarity=0.122 Sum_probs=21.6
Q ss_pred eEEEEehhhHHHHHHHHHHhheeeeEeeccce
Q 043869 411 KNIVIICLFVTVVILISVVTFGIFIYRYRVGS 442 (450)
Q Consensus 411 ~~i~i~~~~~~~~~l~~~~~~~~~~~r~~~~~ 442 (450)
.+|+|+++.+|+-++++++.++|.+.||+++|
T Consensus 318 lvi~i~~vgLG~P~l~li~Ggl~v~~~r~r~~ 349 (350)
T PF15065_consen 318 LVIMIMAVGLGVPLLLLILGGLYVCLRRRRKR 349 (350)
T ss_pred HHHHHHHHHhhHHHHHHHHhhheEEEeccccC
Confidence 35667777788877777777777666555444
No 68
>COG3763 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=37.28 E-value=4.3 Score=31.02 Aligned_cols=28 Identities=21% Similarity=0.402 Sum_probs=17.3
Q ss_pred ehhhHHHHHHHHHHhheeeeEeecccee
Q 043869 416 ICLFVTVVILISVVTFGIFIYRYRVGSY 443 (450)
Q Consensus 416 ~~~~~~~~~l~~~~~~~~~~~r~~~~~~ 443 (450)
+++++.++.+++.++++||+.||.-+++
T Consensus 5 lail~ivl~ll~G~~~G~fiark~~~k~ 32 (71)
T COG3763 5 LAILLIVLALLAGLIGGFFIARKQMKKQ 32 (71)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 3445556666666777788877654443
No 69
>PF01034 Syndecan: Syndecan domain; InterPro: IPR001050 The syndecans are transmembrane proteoglycans which are involved in the organisation of cytoskeleton and/or actin microfilaments, and have important roles as cell surface receptors during cell-cell and/or cell-matrix interactions [, ]. Structurally, these proteins consist of four separate domains: A signal sequence; An extracellular domain (ectodomain) of variable length whose sequence is not evolutionary conserved in the various forms of syndecans. The ectodomain contains the sites of attachment of the heparan sulphate glycosaminoglycan side chains; A transmembrane region; A highly conserved cytoplasmic domain of about 30 to 35 residues, which could interact with cytoskeletal proteins. The proteins known to belong to this family are: Syndecan 1. Syndecan 2 or fibroglycan. Syndecan 3 or neuroglycan or N-syndecan. Syndecan 4 or amphiglycan or ryudocan. Drosophila syndecan. Caenorhabditis elegans probable syndecan (F57C7.3). Syndecan-4, a transmembrane heparan sulphate proteoglycan, is a coreceptor with integrins in cell adhesion. It has been suggested to form a ternary signalling complex with protein kinase Calpha and phosphatidylinositol 4,5-bisphosphate (PIP2). Structural studies have demonstrated that the cytoplasmic domain undergoes a conformational transition and forms a symmetric dimer in the presence of phospholipid activator PIP2, and whose overall structure in solution exhibits a twisted clamp shape having a cavity in the centre of dimeric interface. In addition, it has been observed that the syndecan-4 variable domain interacts, strongly, not only with fatty acyl groups but also the anionic head group of PIP2. These findings indicate that PIP2 promotes oligomerisation of the syndecan-4 cytoplasmic domain for transmembrane signalling and cell-matrix adhesion [, ].; GO: 0008092 cytoskeletal protein binding, 0016020 membrane; PDB: 1EJQ_B 1EJP_B 1YBO_C 1OBY_Q.
Probab=35.69 E-value=11 Score=28.36 Aligned_cols=8 Identities=50% Similarity=0.816 Sum_probs=0.4
Q ss_pred eeeeEeec
Q 043869 432 GIFIYRYR 439 (450)
Q Consensus 432 ~~~~~r~~ 439 (450)
.++++|.|
T Consensus 30 lf~iyR~r 37 (64)
T PF01034_consen 30 LFLIYRMR 37 (64)
T ss_dssp -------S
T ss_pred HHHHHHHH
Confidence 34466644
No 70
>PF14991 MLANA: Protein melan-A; PDB: 2GTZ_F 2GT9_F 3MRO_P 2GUO_C 3MRQ_P 2GTW_C 3L6F_C 3MRP_P.
Probab=35.08 E-value=12 Score=31.46 Aligned_cols=17 Identities=12% Similarity=-0.083 Sum_probs=0.0
Q ss_pred HHHHHHhheeeeEeecc
Q 043869 424 ILISVVTFGIFIYRYRV 440 (450)
Q Consensus 424 ~l~~~~~~~~~~~r~~~ 440 (450)
+.+.+++++|+.+||..
T Consensus 36 LgiLLliGCWYckRRSG 52 (118)
T PF14991_consen 36 LGILLLIGCWYCKRRSG 52 (118)
T ss_dssp -----------------
T ss_pred HHHHHHHhheeeeecch
Confidence 33344556777766643
No 71
>PF07354 Sp38: Zona-pellucida-binding protein (Sp38); InterPro: IPR010857 This family contains a number of zona-pellucida-binding proteins that seem to be restricted to mammals. These are sperm proteins that bind to the 90 kDa family of zona pellucida glycoproteins in a calcium-dependent manner []. These represent some of the specific molecules that mediate the first steps of gamete interaction, allowing fertilisation to occur [].; GO: 0007339 binding of sperm to zona pellucida, 0005576 extracellular region
Probab=34.25 E-value=52 Score=32.03 Aligned_cols=35 Identities=17% Similarity=0.439 Sum_probs=26.7
Q ss_pred cCCCcEEEEecCCCCCCCCccEEEEecCCcEEEEeC
Q 043869 68 IPEKTVVWTANRDNPPVSSNATLMFNSEGRIVLRSG 103 (450)
Q Consensus 68 ~~~~tvVW~ANr~~Pv~~~~~~L~l~~~G~LvL~d~ 103 (450)
+.+.+..|+--.+.++. .++.++||+.|.|++.|=
T Consensus 10 ~iDP~y~W~GP~g~~l~-gn~~~nIT~TG~L~~~~F 44 (271)
T PF07354_consen 10 LIDPTYLWTGPNGKPLS-GNSYVNITETGKLMFKNF 44 (271)
T ss_pred cCCCceEEECCCCcccC-CCCeEEEccCceEEeecc
Confidence 34567788887777777 677888888888888653
No 72
>KOG0640 consensus mRNA cleavage stimulating factor complex; subunit 1 [RNA processing and modification]
Probab=32.45 E-value=1.6e+02 Score=29.61 Aligned_cols=65 Identities=18% Similarity=0.411 Sum_probs=44.9
Q ss_pred EecCCCCCCCCccEEEEecCCcEEEEeCCCCc-eEEeecCC-------------CCccEEEEecCCCeEEEecCCe--eE
Q 043869 76 TANRDNPPVSSNATLMFNSEGRIVLRSGEQGQ-NSIIADNS-------------QSASSASMLDSGSFVLHNSDGK--VI 139 (450)
Q Consensus 76 ~ANr~~Pv~~~~~~L~l~~~G~LvL~d~~~g~-~vW~st~~-------------~~~~~a~LldsGNlVL~~~~~~--~l 139 (450)
.||.+....+.-..+.-++.|+|+++...+|. -+| -.-+ +.+.+|+...+|..+|.....+ -+
T Consensus 253 sanPd~qht~ai~~V~Ys~t~~lYvTaSkDG~Iklw-DGVS~rCv~t~~~AH~gsevcSa~Ftkn~kyiLsSG~DS~vkL 331 (430)
T KOG0640|consen 253 SANPDDQHTGAITQVRYSSTGSLYVTASKDGAIKLW-DGVSNRCVRTIGNAHGGSEVCSAVFTKNGKYILSSGKDSTVKL 331 (430)
T ss_pred ecCcccccccceeEEEecCCccEEEEeccCCcEEee-ccccHHHHHHHHhhcCCceeeeEEEccCCeEEeecCCcceeee
Confidence 56766655544456777889999998663554 488 3321 2467899999999999876433 38
Q ss_pred Ee
Q 043869 140 WQ 141 (450)
Q Consensus 140 WQ 141 (450)
|+
T Consensus 332 WE 333 (430)
T KOG0640|consen 332 WE 333 (430)
T ss_pred ee
Confidence 98
No 73
>KOG3637 consensus Vitronectin receptor, alpha subunit [Extracellular structures]
Probab=32.29 E-value=31 Score=40.35 Aligned_cols=22 Identities=18% Similarity=0.190 Sum_probs=10.7
Q ss_pred EEEEehhhHHHHHHHHH-Hhhee
Q 043869 412 NIVIICLFVTVVILISV-VTFGI 433 (450)
Q Consensus 412 ~i~i~~~~~~~~~l~~~-~~~~~ 433 (450)
+.||+.++++.++||++ +++.|
T Consensus 978 ~wiIi~svl~GLLlL~llv~~Lw 1000 (1030)
T KOG3637|consen 978 LWIIILSVLGGLLLLALLVLLLW 1000 (1030)
T ss_pred eeeehHHHHHHHHHHHHHHHHHH
Confidence 44455555555555444 44433
No 74
>TIGR03066 Gem_osc_para_1 Gemmata obscuriglobus paralogous family TIGR03066. This model represents an uncharacterized paralogous family in Gemmata obscuriglobus UQM 2246, a member of the Planctomycetes. This family shows sequence similarity to TIGR03067, which is also found in Gemmata obscuriglobus as well as in a few other species.
Probab=31.61 E-value=1.8e+02 Score=24.58 Aligned_cols=52 Identities=19% Similarity=0.366 Sum_probs=31.2
Q ss_pred CccEEEEecCCcEEEEeCCCCce------EEe-ecCC-------C----CccE-EEEecCCCeEEEecCCee
Q 043869 86 SNATLMFNSEGRIVLRSGEQGQN------SII-ADNS-------Q----SASS-ASMLDSGSFVLHNSDGKV 138 (450)
Q Consensus 86 ~~~~L~l~~~G~LvL~d~~~g~~------vW~-st~~-------~----~~~~-a~LldsGNlVL~~~~~~~ 138 (450)
+...|+|..||.|+|+.+ ++.- -|. +.+. . .... ..=+++|-|||.|+++.+
T Consensus 34 ~~~~leF~~dGKL~v~~g-nng~~~~~~Gty~L~G~kLtL~~~p~g~t~k~~Vtv~~l~~~~Lvl~d~dg~~ 104 (111)
T TIGR03066 34 DDVVIEFAKDGKLVVTIG-EKGKEVKADGTYKLDGNKLTLTLKAGGKEKKETLTVKKLTDDELVGKDPDGKK 104 (111)
T ss_pred CceEEEEcCCCeEEEecC-CCCcEeccCceEEEECCEEEEEEcCCCccccceEEEEEecCCeEEEEcCCCCE
Confidence 456799999999998776 4321 131 1110 0 0011 123688999999988764
No 75
>PF06006 DUF905: Bacterial protein of unknown function (DUF905); InterPro: IPR009253 This family consists of several short hypothetical proteobacterial proteins of unknown function.; PDB: 2HJJ_A.
Probab=31.35 E-value=53 Score=25.14 Aligned_cols=20 Identities=15% Similarity=0.806 Sum_probs=11.7
Q ss_pred CeEEEecCCeeEEeecCCCC
Q 043869 128 SFVLHNSDGKVIWQTFDHPT 147 (450)
Q Consensus 128 NlVL~~~~~~~lWQSFd~PT 147 (450)
.||++|.+|..+|..|.+-.
T Consensus 34 RlvvRd~~g~mvWRaWNFEp 53 (70)
T PF06006_consen 34 RLVVRDTEGQMVWRAWNFEP 53 (70)
T ss_dssp EEEEE-SS--EEEEEESSST
T ss_pred EEEEEcCCCcEEEEeeccCC
Confidence 46777777788888776543
No 76
>KOG4649 consensus PQQ (pyrrolo-quinoline quinone) repeat protein [Secondary metabolites biosynthesis, transport and catabolism]
Probab=29.33 E-value=2e+02 Score=28.32 Aligned_cols=40 Identities=20% Similarity=0.365 Sum_probs=28.3
Q ss_pred CcEEEEecCCCCCCCC----ccEEEE-ecCCcEEEEeCCCCceEEe
Q 043869 71 KTVVWTANRDNPPVSS----NATLMF-NSEGRIVLRSGEQGQNSII 111 (450)
Q Consensus 71 ~tvVW~ANr~~Pv~~~----~~~L~l-~~~G~LvL~d~~~g~~vW~ 111 (450)
.+..|.|.|..|+-.+ +....+ +-||+|.-.++ .|+.||+
T Consensus 169 ~~~~w~~~~~~PiF~splcv~~sv~i~~VdG~l~~f~~-sG~qvwr 213 (354)
T KOG4649|consen 169 STEFWAATRFGPIFASPLCVGSSVIITTVDGVLTSFDE-SGRQVWR 213 (354)
T ss_pred cceehhhhcCCccccCceeccceEEEEEeccEEEEEcC-CCcEEEe
Confidence 4788999998887633 233333 57888877777 8888884
No 77
>PF12946 EGF_MSP1_1: MSP1 EGF domain 1; InterPro: IPR024730 This EGF-like domain is found at the C terminus of the malaria parasite MSP1 protein. MSP1 is the merozoite surface protein 1. This domain is part of the C-terminal fragment that is proteolytically processed from the the rest of the protein and is left attached to the surface of the invading parasite [].; PDB: 1N1I_C 2FLG_A 1CEJ_A 2NPR_A 1B9W_A 1OB1_F.
Probab=29.18 E-value=11 Score=25.16 Aligned_cols=27 Identities=30% Similarity=0.701 Sum_probs=17.3
Q ss_pred CCCCCCcccccCCC-CCCCcCCCCCeec
Q 043869 271 GLCGFNSFCVLNDQ-TPNCTCLPGFVAI 297 (450)
Q Consensus 271 g~CG~~g~C~~~~~-~~~C~C~~GF~~~ 297 (450)
..|-.|+-|....+ ...|.|++||+..
T Consensus 5 ~~cP~NA~C~~~~dG~eecrCllgyk~~ 32 (37)
T PF12946_consen 5 TKCPANAGCFRYDDGSEECRCLLGYKKV 32 (37)
T ss_dssp S---TTEEEEEETTSEEEEEE-TTEEEE
T ss_pred ccCCCCcccEEcCCCCEEEEeeCCcccc
Confidence 45777888865443 6789999999864
No 78
>PHA03290 envelope glycoprotein I; Provisional
Probab=29.13 E-value=52 Score=33.03 Aligned_cols=35 Identities=17% Similarity=0.070 Sum_probs=17.9
Q ss_pred hhhhccccccccccCCCCcccCCCCCeEEeCCCeEEEEEEeCCC
Q 043869 12 GCFTAAAQKKHSNISIGSSLSPTGNSSWRSPSGLYAFGFYPQRN 55 (450)
Q Consensus 12 ~~~~~~~~~~~~~i~~g~~l~~~~~~~l~S~~g~F~lGF~~~~~ 55 (450)
++|....++ .-+-.|..++ +.+ ++.-+.||...++
T Consensus 11 ~~~~i~~~~--gIVyRG~~VS------L~v-DsSa~v~f~~~g~ 45 (357)
T PHA03290 11 MIFGIQCAA--AIIFKGDHIS------LQV-NSSATSIFIKMGN 45 (357)
T ss_pred HHHHhhhee--EEEEECCeEE------EEE-CCccceeeecCCC
Confidence 555433322 2466677653 333 4455677876433
No 79
>TIGR03075 PQQ_enz_alc_DH PQQ-dependent dehydrogenase, methanol/ethanol family. This protein family has a phylogenetic distribution very similar to that coenzyme PQQ biosynthesis enzymes, as shown by partial phylogenetic profiling. Genes in this family often are found adjacent to the PQQ biosynthesis genes themselves. An unusual, strained disulfide bond between adjacent Cys residues contributes to PQQ-binding, as does a Trp residue that is part of a PQQ enzyme repeat (see pfam01011). Characterized members include the dehydrogenase subunit of a membrane-anchored, three subunit alcohol (ethanol) dehydrogenase of Gluconobacter suboxydans, a homodimeric ethanol dehydrogenase in Pseudomonas aeruginosa, and the large subunit of an alpha2/beta2 heterotetrameric methanol dehydrogenase in Methylobacterium extorquens.
Probab=28.43 E-value=7.7e+02 Score=26.53 Aligned_cols=121 Identities=15% Similarity=0.313 Sum_probs=0.0
Q ss_pred EEeCCCCCeeEEEEEEeecCCCcEEEEecCCCCCCCCc---------------cEEEE-ecCCcEEEEeCCCCceEEeec
Q 043869 50 FYPQRNGSRYYVGVFLAGIPEKTVVWTANRDNPPVSSN---------------ATLMF-NSEGRIVLRSGEQGQNSIIAD 113 (450)
Q Consensus 50 F~~~~~~~~~~lgIw~~~~~~~tvVW~ANr~~Pv~~~~---------------~~L~l-~~~G~LvL~d~~~g~~vW~st 113 (450)
|+..... ...+| ......++|..+...|..... .++.+ +.+|.|+-.|..+|.++| +.
T Consensus 73 yv~s~~g--~v~Al---Da~TGk~lW~~~~~~~~~~~~~~~~~~~~rg~av~~~~v~v~t~dg~l~ALDa~TGk~~W-~~ 146 (527)
T TIGR03075 73 YVTTSYS--RVYAL---DAKTGKELWKYDPKLPDDVIPVMCCDVVNRGVALYDGKVFFGTLDARLVALDAKTGKVVW-SK 146 (527)
T ss_pred EEECCCC--cEEEE---ECCCCceeeEecCCCCcccccccccccccccceEECCEEEEEcCCCEEEEEECCCCCEEe-ec
Q ss_pred CCC-------CccEEEEec--------------CCCeEEEec-CCeeEEeecCCCCC------------------ccCCC
Q 043869 114 NSQ-------SASSASMLD--------------SGSFVLHNS-DGKVIWQTFDHPTD------------------TLLPT 153 (450)
Q Consensus 114 ~~~-------~~~~a~Lld--------------sGNlVL~~~-~~~~lWQSFd~PTD------------------TlLpg 153 (450)
... ..+...+.+ +|.++-+|. +|+.+|+--.-|.+ |+ +|
T Consensus 147 ~~~~~~~~~~~tssP~v~~g~Vivg~~~~~~~~~G~v~AlD~~TG~~lW~~~~~p~~~~~~~~~~~~~~~~~~~~tw-~~ 225 (527)
T TIGR03075 147 KNGDYKAGYTITAAPLVVKGKVITGISGGEFGVRGYVTAYDAKTGKLVWRRYTVPGDMGYLDKADKPVGGEPGAKTW-PG 225 (527)
T ss_pred ccccccccccccCCcEEECCEEEEeecccccCCCcEEEEEECCCCceeEeccCcCCCcccccccccccccccccCCC-CC
Q ss_pred cccCCCCeEEeccCCC-CCCCCceEE
Q 043869 154 QRLSAGTELCSGISET-DPSTGKFRL 178 (450)
Q Consensus 154 q~L~~~~~L~S~~s~~-dps~G~f~l 178 (450)
+....+.- ..|-..+ |+..|...+
T Consensus 226 ~~~~~gg~-~~W~~~s~D~~~~lvy~ 250 (527)
T TIGR03075 226 DAWKTGGG-ATWGTGSYDPETNLIYF 250 (527)
T ss_pred CccccCCC-CccCceeEcCCCCeEEE
No 80
>PF08114 PMP1_2: ATPase proteolipid family; InterPro: IPR012589 This family consists of small proteolipids associated with the plasma membrane H+ ATPase. Two proteolipids (PMP1 and PMP2) are associated with the ATPase and both genes are similarly expressed in the wild-type strain of yeast. No modification of the level of transcription of one PMP gene is detected in a strain deleted of the other. Though both proteolipids show similarity with other small proteolipids associated with other cation -transporting ATPases, their functions remain unclear [].
Probab=28.12 E-value=6.4 Score=26.76 Aligned_cols=19 Identities=32% Similarity=0.450 Sum_probs=9.4
Q ss_pred HhheeeeEeeccceeeccc
Q 043869 429 VTFGIFIYRYRVGSYRRIQ 447 (450)
Q Consensus 429 ~~~~~~~~r~~~~~~~~~~ 447 (450)
.+...|+|||-..|.+-+|
T Consensus 23 ~iva~~iYRKw~aRkr~l~ 41 (43)
T PF08114_consen 23 GIVALFIYRKWQARKRALQ 41 (43)
T ss_pred HHHHHHHHHHHHHHHHHHh
Confidence 3334456666555544443
No 81
>PF02480 Herpes_gE: Alphaherpesvirus glycoprotein E; InterPro: IPR003404 Glycoprotein E (gE) of Alphaherpesvirus forms a complex with glycoprotein I (gI), functioning as an immunoglobulin G (IgG) Fc binding protein. gE is involved in virus spread but is not essential for propagation [].; GO: 0016020 membrane; PDB: 2GJ7_F 2GIY_B.
Probab=28.10 E-value=20 Score=37.80 Aligned_cols=30 Identities=17% Similarity=0.363 Sum_probs=14.6
Q ss_pred CCCCCccCCcccccCC-cceEEeccccccCc
Q 043869 302 WTAGCERNYTAESCGN-KAIQELENTNWEDV 331 (450)
Q Consensus 302 ~s~GC~r~~~l~~C~~-~~f~~l~~v~~p~~ 331 (450)
.+.+|.+......|.+ ..+.+..++.+.++
T Consensus 237 ~y~~C~~~~~~~~C~~~~~~~~~~~~~~~~~ 267 (439)
T PF02480_consen 237 RYANCSPSGWPRRCPSTSHIEPVPGLRWASN 267 (439)
T ss_dssp EEEEEBTTC-TTTTEEEEEE---TTEEE-TT
T ss_pred hhcCCCCCCCcCCCCchhccCcCccccccCC
Confidence 3678888644445754 34444556666543
No 82
>cd00216 PQQ_DH Dehydrogenases with pyrrolo-quinoline quinone (PQQ) as cofactor, like ethanol, methanol, and membrane bound glucose dehydrogenases. The alignment model contains an 8-bladed beta-propeller.
Probab=27.61 E-value=1.9e+02 Score=30.66 Aligned_cols=71 Identities=18% Similarity=0.348 Sum_probs=42.6
Q ss_pred CCCcEEEEecCC-------CCCCCCccEEEE-ecCCcEEEEeCCCCceEEeecCC-CC--------c---------cEEE
Q 043869 69 PEKTVVWTANRD-------NPPVSSNATLMF-NSEGRIVLRSGEQGQNSIIADNS-QS--------A---------SSAS 122 (450)
Q Consensus 69 ~~~tvVW~ANr~-------~Pv~~~~~~L~l-~~~G~LvL~d~~~g~~vW~st~~-~~--------~---------~~a~ 122 (450)
...+++|..+-. .|+. .+.++.+ +.+|.|+-.|..+|.++| +... .. . ....
T Consensus 37 ~~~~~~W~~~~~~~~~~~~sPvv-~~g~vy~~~~~g~l~AlD~~tG~~~W-~~~~~~~~~~~~~~~~~~g~~~~~~~~V~ 114 (488)
T cd00216 37 KKLKVAWTFSTGDERGQEGTPLV-VDGDMYFTTSHSALFALDAATGKVLW-RYDPKLPADRGCCDVVNRGVAYWDPRKVF 114 (488)
T ss_pred hcceeeEEEECCCCCCcccCCEE-ECCEEEEeCCCCcEEEEECCCChhhc-eeCCCCCccccccccccCCcEEccCCeEE
Confidence 345678887643 3555 3455555 457988877764789999 4322 10 0 0111
Q ss_pred E-ecCCCeEEEec-CCeeEEe
Q 043869 123 M-LDSGSFVLHNS-DGKVIWQ 141 (450)
Q Consensus 123 L-ldsGNlVL~~~-~~~~lWQ 141 (450)
+ ..+|.++-+|. +++.+|+
T Consensus 115 v~~~~g~v~AlD~~TG~~~W~ 135 (488)
T cd00216 115 FGTFDGRLVALDAETGKQVWK 135 (488)
T ss_pred EecCCCeEEEEECCCCCEeee
Confidence 1 23677777776 6889999
No 83
>PF05545 FixQ: Cbb3-type cytochrome oxidase component FixQ; InterPro: IPR008621 This family consists of several Cbb3-type cytochrome oxidase components (FixQ/CcoQ). FixQ is found in nitrogen fixing bacteria. Since nitrogen fixation is an energy-consuming process, effective symbioses depend on operation of a respiratory chain with a high affinity for O2, closely coupled to ATP production. This requirement is fulfilled by a special three-subunit terminal oxidase (cytochrome terminal oxidase cbb3), which was first identified in Bradyrhizobium japonicum as the product of the fixNOQP operon [].
Probab=27.59 E-value=21 Score=25.24 Aligned_cols=14 Identities=0% Similarity=-0.249 Sum_probs=6.7
Q ss_pred eeeeEeeccceeec
Q 043869 432 GIFIYRYRVGSYRR 445 (450)
Q Consensus 432 ~~~~~r~~~~~~~~ 445 (450)
+|..++++++++.+
T Consensus 27 ~w~~~~~~k~~~e~ 40 (49)
T PF05545_consen 27 IWAYRPRNKKRFEE 40 (49)
T ss_pred HHHHcccchhhHHH
Confidence 33344455555544
No 84
>PF11403 Yeast_MT: Yeast metallothionein; InterPro: IPR022710 Metallothioneins are characterised by an abundance of cysteine residues and a lack of generic secondary structure motifs. This protein functions in primary metal storage, transport and detoxification []. For the first 40 residues in the protein the polypeptide wraps around the metal by forming two large parallel loops separated by a deep cleft containing the metal cluster []. ; PDB: 1AQS_A 1AQR_A 1RJU_V 1FMY_A 1AOO_A 1AQQ_A.
Probab=26.30 E-value=47 Score=21.58 Aligned_cols=19 Identities=37% Similarity=0.931 Sum_probs=9.2
Q ss_pred CCCCCCcccccCCCCCCCcCCCCC
Q 043869 271 GLCGFNSFCVLNDQTPNCTCLPGF 294 (450)
Q Consensus 271 g~CG~~g~C~~~~~~~~C~C~~GF 294 (450)
|.|-.|.-| ...|+||.|-
T Consensus 12 gscknneqc-----qkscscptgc 30 (40)
T PF11403_consen 12 GSCKNNEQC-----QKSCSCPTGC 30 (40)
T ss_dssp STTTT-TTS-----TTS-SS-TTT
T ss_pred CCccChHHH-----hhcCCCCCCC
Confidence 444444444 4579998764
No 85
>PF12458 DUF3686: ATPase involved in DNA repair ; InterPro: IPR020958 This entry represents an N-terminal domain associated with ATPases and some uncharacterised proteins; it is approximately 450 amino acids in length and contains two conserved sequence motifs: DVF and SPNGED.
Probab=25.91 E-value=1.7e+02 Score=30.60 Aligned_cols=55 Identities=24% Similarity=0.392 Sum_probs=30.4
Q ss_pred eEEeCCC-eEEEEEEeCCCCCeeEEEEEEeecCCCcEEEEecCCCCCCCCccEEEEecCCcEEEEeC
Q 043869 38 SWRSPSG-LYAFGFYPQRNGSRYYVGVFLAGIPEKTVVWTANRDNPPVSSNATLMFNSEGRIVLRSG 103 (450)
Q Consensus 38 ~l~S~~g-~F~lGF~~~~~~~~~~lgIw~~~~~~~tvVW~ANr~~Pv~~~~~~L~l~~~G~LvL~d~ 103 (450)
.+.|||| ++-.-||.+..+ .|+=+-|+-|... + .+|+... -..+.+||.|++..+
T Consensus 312 ~vrSPNGEDvLYvF~~~~~g--~~~Ll~YN~I~k~----v---~tPi~ch--G~alf~DG~l~~fra 367 (448)
T PF12458_consen 312 KVRSPNGEDVLYVFYAREEG--RYLLLPYNLIRKE----V---ATPIICH--GYALFEDGRLVYFRA 367 (448)
T ss_pred EecCCCCceEEEEEEECCCC--cEEEEechhhhhh----h---cCCeecc--ceeEecCCEEEEEec
Confidence 4567777 455556665555 4555556554322 1 2466522 245667777777654
No 86
>PF00558 Vpu: Vpu protein; InterPro: IPR008187 The Human immunodeficiency virus 1 (HIV-1) Vpu protein acts in the degradation of CD4 in the endoplasmic reticulum and in the enhancement of virion release from the plasma membrane of infected cells [].; GO: 0019076 release of virus from host; PDB: 2JPX_A 1PI8_A 2GOH_A 2GOF_A 1PI7_A 1PJE_A 1VPU_A 2K7Y_A.
Probab=24.71 E-value=36 Score=27.01 Aligned_cols=6 Identities=33% Similarity=0.468 Sum_probs=0.4
Q ss_pred ceeecc
Q 043869 441 GSYRRI 446 (450)
Q Consensus 441 ~~~~~~ 446 (450)
+|.+||
T Consensus 34 ~rqrkI 39 (81)
T PF00558_consen 34 KRQRKI 39 (81)
T ss_dssp -----C
T ss_pred HHHHhH
Confidence 333443
No 87
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=24.65 E-value=27 Score=28.58 Aligned_cols=28 Identities=18% Similarity=0.087 Sum_probs=18.1
Q ss_pred EehhhHHHHHHHHHHhheeeeEeeccce
Q 043869 415 IICLFVTVVILISVVTFGIFIYRYRVGS 442 (450)
Q Consensus 415 i~~~~~~~~~l~~~~~~~~~~~r~~~~~ 442 (450)
|.++++++++++.+++++.+++..+++|
T Consensus 68 iagi~vg~~~~v~~lv~~l~w~f~~r~k 95 (96)
T PTZ00382 68 IAGISVAVVAVVGGLVGFLCWWFVCRGK 95 (96)
T ss_pred EEEEEeehhhHHHHHHHHHhheeEEeec
Confidence 4556667777666677666677777543
No 88
>KOG1214 consensus Nidogen and related basement membrane protein proteins [Cell wall/membrane/envelope biogenesis; Extracellular structures]
Probab=23.99 E-value=55 Score=36.83 Aligned_cols=31 Identities=23% Similarity=0.657 Sum_probs=24.4
Q ss_pred CCCCCcCCCCCCcccccCCCCCCCcCCCCCee
Q 043869 265 EKCDPIGLCGFNSFCVLNDQTPNCTCLPGFVA 296 (450)
Q Consensus 265 ~~C~~~g~CG~~g~C~~~~~~~~C~C~~GF~~ 296 (450)
|+|. +..|-++..|.+..+.-.|.|-|||.-
T Consensus 828 DeC~-psrChp~A~CyntpgsfsC~C~pGy~G 858 (1289)
T KOG1214|consen 828 DECS-PSRCHPAATCYNTPGSFSCRCQPGYYG 858 (1289)
T ss_pred cccC-ccccCCCceEecCCCcceeecccCccC
Confidence 5666 788999999987555778999998863
No 89
>PRK12785 fliL flagellar basal body-associated protein FliL; Reviewed
Probab=23.13 E-value=95 Score=28.02 Aligned_cols=18 Identities=22% Similarity=0.312 Sum_probs=9.0
Q ss_pred HHHHHHHHHhheeeeEee
Q 043869 421 TVVILISVVTFGIFIYRY 438 (450)
Q Consensus 421 ~~~~l~~~~~~~~~~~r~ 438 (450)
.+++++.+..+.||+...
T Consensus 32 ~~lll~~~g~g~~f~~~~ 49 (166)
T PRK12785 32 AAVLLLGGGGGGFFFFFS 49 (166)
T ss_pred HHHHHHhcchheEEEEEe
Confidence 344444444556665553
No 90
>PF02237 BPL_C: Biotin protein ligase C terminal domain; InterPro: IPR003142 This C-terminal domain has an SH3-like barrel fold, the function of which is unknown. It is found associated with prokaryotic bifunctional transcriptional repressors [] and eukaryotic enzymes involved in biotin utilization [, ]. In Escherichia coli the biotin operon repressor (BirA) is a bifunctional protein. BirA acts both as the acetyl-coA carboxylase biotin holoenzyme synthetase (6.3.4.15 from EC) and as the biotin operon repressor. DNA sequence analysis of mutations indicates that the helix-turn-helix DNA binding region is located at the N terminus while mutations affecting enzyme function, although mapping over a large region, are found mainly in the central part of the protein's primary sequence [].; GO: 0006464 protein modification process; PDB: 3RUX_A 2CGH_A 3L1A_B 3L2Z_A 1HXD_A 1BIB_A 2EWN_B 1BIA_A 2EJ9_A 3FJP_A ....
Probab=22.13 E-value=1.1e+02 Score=21.31 Aligned_cols=14 Identities=14% Similarity=0.427 Sum_probs=6.9
Q ss_pred EEEecCCcEEEEeC
Q 043869 90 LMFNSEGRIVLRSG 103 (450)
Q Consensus 90 L~l~~~G~LvL~d~ 103 (450)
.-++++|.|+|...
T Consensus 20 ~gId~~G~L~v~~~ 33 (48)
T PF02237_consen 20 EGIDDDGALLVRTE 33 (48)
T ss_dssp EEEETTSEEEEEET
T ss_pred EEECCCCEEEEEEC
Confidence 33455555555444
No 91
>PF06247 Plasmod_Pvs28: Plasmodium ookinete surface protein Pvs28; InterPro: IPR010423 This family consists of several ookinete surface protein (Pvs28) from several species of Plasmodium. Pvs25 and Pvs28 are expressed on the surface of ookinetes. These proteins are potential candidates for vaccine and induce antibodies that block the infectivity of Plasmodium vivax in immunised animals [].; GO: 0009986 cell surface, 0016020 membrane; PDB: 1Z3G_B 1Z1Y_B 1Z27_A.
Probab=22.09 E-value=23 Score=32.66 Aligned_cols=41 Identities=24% Similarity=0.619 Sum_probs=25.5
Q ss_pred CCCCCC----cCCCCCCcccccCCC-----CCCCcCCCCCeeccCCCCCCCCccC
Q 043869 264 SEKCDP----IGLCGFNSFCVLNDQ-----TPNCTCLPGFVAISKGNWTAGCERN 309 (450)
Q Consensus 264 ~~~C~~----~g~CG~~g~C~~~~~-----~~~C~C~~GF~~~~~~~~s~GC~r~ 309 (450)
...|+. .-.||.|+.|....+ .-.|.|.+||.... .-|+|.
T Consensus 39 kv~C~~~e~~~K~Cgdya~C~~~~~~~~~~~~~C~C~~gY~~~~-----~vCvp~ 88 (197)
T PF06247_consen 39 KVECDKLENVNKPCGDYAKCINQANKGEERAYKCDCINGYILKQ-----GVCVPN 88 (197)
T ss_dssp ----SG-GGTTSEEETTEEEEE-SSTTSSTSEEEEE-TTEEESS-----SSEEEG
T ss_pred ceecCcccccCccccchhhhhcCCCcccceeEEEecccCceeeC-----CeEchh
Confidence 345654 568999999976432 34699999999863 347764
No 92
>PF05568 ASFV_J13L: African swine fever virus J13L protein; InterPro: IPR008385 This family consists of several African swine fever virus (ASFV) j13L proteins [, , ].
Probab=21.84 E-value=17 Score=31.95 Aligned_cols=16 Identities=6% Similarity=-0.099 Sum_probs=7.1
Q ss_pred eeeeEeeccceeeccc
Q 043869 432 GIFIYRYRVGSYRRIQ 447 (450)
Q Consensus 432 ~~~~~r~~~~~~~~~~ 447 (450)
++++.+||+|+-.-|.
T Consensus 49 i~lcssRKkKaaAAi~ 64 (189)
T PF05568_consen 49 IYLCSSRKKKAAAAIE 64 (189)
T ss_pred HHHHhhhhHHHHhhhh
Confidence 3444444445444443
No 93
>PRK01844 hypothetical protein; Provisional
Probab=20.46 E-value=11 Score=29.07 Aligned_cols=25 Identities=36% Similarity=0.537 Sum_probs=13.7
Q ss_pred hhHHHHHHHHHHhheeeeEeeccce
Q 043869 418 LFVTVVILISVVTFGIFIYRYRVGS 442 (450)
Q Consensus 418 ~~~~~~~l~~~~~~~~~~~r~~~~~ 442 (450)
+++.++.+++.++++||+.||.-++
T Consensus 7 I~l~I~~li~G~~~Gff~ark~~~k 31 (72)
T PRK01844 7 ILVGVVALVAGVALGFFIARKYMMN 31 (72)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 3444445555556667776665433
No 94
>TIGR03503 conserved hypothetical protein TIGR03503. This set of conserved hypothetical protein has a phylogenetic range that closely matches that of TIGR03501, a putative C-terminal protein targeting signal.
Probab=20.39 E-value=52 Score=33.82 Aligned_cols=21 Identities=29% Similarity=0.518 Sum_probs=11.5
Q ss_pred hHHHHHHHHHHhheeeeEeecc
Q 043869 419 FVTVVILISVVTFGIFIYRYRV 440 (450)
Q Consensus 419 ~~~~~~l~~~~~~~~~~~r~~~ 440 (450)
+.++++ +++.+.+|+++|||+
T Consensus 353 ~~N~v~-lllg~~~~~~~rk~k 373 (374)
T TIGR03503 353 VGNVVI-LLLGGIGFFVWRKKK 373 (374)
T ss_pred hhhhhh-hhhheeeEEEEEEee
Confidence 434444 444556677777664
No 95
>PF14316 DUF4381: Domain of unknown function (DUF4381)
Probab=20.36 E-value=23 Score=31.17 Aligned_cols=11 Identities=45% Similarity=0.767 Sum_probs=5.3
Q ss_pred eEeeccceeec
Q 043869 435 IYRYRVGSYRR 445 (450)
Q Consensus 435 ~~r~~~~~~~~ 445 (450)
.+|+|+.+|+|
T Consensus 42 ~r~~~~~~yrr 52 (146)
T PF14316_consen 42 WRRWRRNRYRR 52 (146)
T ss_pred HHHHHccHHHH
Confidence 33444455654
No 96
>KOG4289 consensus Cadherin EGF LAG seven-pass G-type receptor [Signal transduction mechanisms]
Probab=20.21 E-value=50 Score=39.45 Aligned_cols=40 Identities=28% Similarity=0.562 Sum_probs=28.4
Q ss_pred cCCCCCCcccccCCCCCCCcCCCCCeeccCC--CCCCCCccC
Q 043869 270 IGLCGFNSFCVLNDQTPNCTCLPGFVAISKG--NWTAGCERN 309 (450)
Q Consensus 270 ~g~CG~~g~C~~~~~~~~C~C~~GF~~~~~~--~~s~GC~r~ 309 (450)
.+.||++|-|..-....+|.|-|||.-..=+ ..+.-|++.
T Consensus 1244 s~pC~nng~C~srEggYtCeCrpg~tGehCEvs~~agrCvpG 1285 (2531)
T KOG4289|consen 1244 SGPCGNNGRCRSREGGYTCECRPGFTGEHCEVSARAGRCVPG 1285 (2531)
T ss_pred cCCCCCCCceEEecCceeEEecCCccccceeeecccCccccc
Confidence 6899999999874557889999999643211 134556654
No 97
>cd05852 Ig5_Contactin-1 Fifth Ig domain of contactin-1. Ig5_Contactin-1: fifth Ig domain of the neural cell adhesion molecule contactin-1. Contactins are comprised of six Ig domains followed by four fibronectin type III (FnIII) domains anchored to the membrane by glycosylphosphatidylinositol. Contactin-1 is differentially expressed in tumor tissues and may through a RhoA mechanism, facilitate invasion and metastasis of human lung adenocarcinoma.
Probab=20.05 E-value=1.1e+02 Score=23.14 Aligned_cols=34 Identities=15% Similarity=0.332 Sum_probs=22.8
Q ss_pred cCCCcEEEEecCCCCCCCCccEEEEecCCcEEEEeC
Q 043869 68 IPEKTVVWTANRDNPPVSSNATLMFNSEGRIVLRSG 103 (450)
Q Consensus 68 ~~~~tvVW~ANr~~Pv~~~~~~L~l~~~G~LvL~d~ 103 (450)
.|..++.|.=+. .++. .+....+..+|.|.|.+.
T Consensus 13 ~P~p~v~W~k~~-~~l~-~~~r~~~~~~g~L~I~~v 46 (73)
T cd05852 13 APKPKFSWSKGT-ELLV-NNSRISIWDDGSLEILNI 46 (73)
T ss_pred eCCCEEEEEeCC-Eecc-cCCCEEEcCCCEEEECcC
Confidence 466788898654 3444 345677777888888654
No 98
>PTZ00208 65 kDa invariant surface glycoprotein; Provisional
Probab=20.02 E-value=48 Score=34.10 Aligned_cols=30 Identities=20% Similarity=0.371 Sum_probs=16.8
Q ss_pred ceEEEEehhhHHHHHHHHHHhheee-eEeec
Q 043869 410 WKNIVIICLFVTVVILISVVTFGIF-IYRYR 439 (450)
Q Consensus 410 ~~~i~i~~~~~~~~~l~~~~~~~~~-~~r~~ 439 (450)
+..+||.++.+-+++|.+++...|. ++|||
T Consensus 384 ~~~~i~~avl~p~~il~~~~~~~~~~v~rrr 414 (436)
T PTZ00208 384 RTAMIILAVLVPAIILAIIAVAFFIMVKRRR 414 (436)
T ss_pred hhHHHHHHHHHHHHHHHHHHHHhheeeeecc
Confidence 3456677777777776655443333 44444
Done!