Query 042843
Match_columns 484
No_of_seqs 193 out of 1495
Neff 7.7
Searched_HMMs 46136
Date Fri Mar 29 10:07:29 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/042843.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/042843hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF01453 B_lectin: D-mannose b 100.0 6.7E-30 1.5E-34 219.6 3.9 111 77-190 1-114 (114)
2 PF00954 S_locus_glycop: S-loc 99.9 1.2E-27 2.6E-32 204.7 11.1 110 217-329 1-110 (110)
3 cd00028 B_lectin Bulb-type man 99.9 5.7E-26 1.2E-30 196.2 14.9 113 36-158 2-116 (116)
4 smart00108 B_lectin Bulb-type 99.9 7.5E-25 1.6E-29 188.6 14.3 112 36-157 2-114 (114)
5 PF08276 PAN_2: PAN-like domai 99.5 6.3E-15 1.4E-19 114.0 5.0 58 359-416 3-66 (66)
6 cd01098 PAN_AP_plant Plant PAN 99.5 1.1E-13 2.3E-18 112.0 8.2 81 347-431 2-84 (84)
7 cd00129 PAN_APPLE PAN/APPLE-li 99.4 6.6E-13 1.4E-17 106.2 6.1 67 360-430 8-80 (80)
8 smart00108 B_lectin Bulb-type 98.6 1.8E-07 3.8E-12 80.4 9.7 87 97-219 23-111 (114)
9 cd00028 B_lectin Bulb-type man 98.6 4.2E-07 9.1E-12 78.3 10.3 83 102-220 30-113 (116)
10 smart00473 PAN_AP divergent su 98.6 2.1E-07 4.6E-12 73.4 7.5 70 360-429 3-77 (78)
11 PF01453 B_lectin: D-mannose b 98.1 5.7E-05 1.2E-09 64.9 11.4 100 41-159 11-114 (114)
12 cd01100 APPLE_Factor_XI_like S 97.4 0.0002 4.3E-09 56.3 3.8 47 366-412 9-58 (73)
13 PF04478 Mid2: Mid2 like cell 93.7 0.12 2.6E-06 46.1 5.0 33 441-474 47-79 (154)
14 PF08693 SKG6: Transmembrane a 91.9 0.29 6.3E-06 33.5 3.8 10 464-473 30-39 (40)
15 PF15102 TMEM154: TMEM154 prot 91.6 0.26 5.7E-06 43.6 4.4 25 445-469 58-82 (146)
16 smart00223 APPLE APPLE domain. 91.2 0.27 5.8E-06 39.2 3.7 45 367-411 7-57 (79)
17 PF00024 PAN_1: PAN domain Thi 91.0 0.33 7.1E-06 37.7 4.1 54 362-415 3-60 (79)
18 PF08277 PAN_3: PAN-like domai 90.1 1.5 3.3E-05 33.6 7.0 39 380-418 19-58 (71)
19 smart00605 CW CW domain. 87.0 2.1 4.6E-05 35.0 6.3 53 380-432 21-76 (94)
20 PF02009 Rifin_STEVOR: Rifin/s 85.4 0.14 3E-06 51.3 -1.9 20 459-478 269-288 (299)
21 PTZ00382 Variant-specific surf 84.9 1.4 3.1E-05 36.5 4.2 29 444-472 67-95 (96)
22 PF07645 EGF_CA: Calcium-bindi 84.4 0.51 1.1E-05 32.6 1.1 31 297-327 3-35 (42)
23 PF01034 Syndecan: Syndecan do 84.2 0.32 6.9E-06 36.8 0.0 16 462-477 27-42 (64)
24 PF14295 PAN_4: PAN domain; PD 83.7 0.9 2E-05 32.2 2.3 24 380-403 15-38 (51)
25 cd00053 EGF Epidermal growth f 81.1 1.2 2.7E-05 28.4 2.0 29 299-327 2-31 (36)
26 PF01102 Glycophorin_A: Glycop 80.8 0.37 8.1E-06 41.7 -0.8 32 445-477 66-97 (122)
27 smart00179 EGF_CA Calcium-bind 75.8 2.2 4.8E-05 28.1 2.0 30 297-326 3-33 (39)
28 KOG4649 PQQ (pyrrolo-quinoline 75.6 10 0.00023 37.1 7.2 46 77-122 168-218 (354)
29 PF01299 Lamp: Lysosome-associ 74.8 1.8 3.9E-05 43.7 2.0 31 444-475 271-301 (306)
30 cd01099 PAN_AP_HGF Subfamily o 74.1 5.8 0.00013 31.4 4.4 33 380-412 24-60 (80)
31 PF01683 EB: EB module; Inter 73.8 2.9 6.2E-05 30.1 2.3 33 293-328 16-48 (52)
32 PTZ00046 rifin; Provisional 72.2 0.81 1.8E-05 46.6 -1.2 18 461-478 330-347 (358)
33 TIGR01477 RIFIN variant surfac 72.2 0.81 1.7E-05 46.5 -1.2 18 461-478 325-342 (353)
34 PF09064 Tme5_EGF_like: Thromb 71.5 2.2 4.9E-05 28.0 1.1 18 310-327 11-28 (34)
35 PF12661 hEGF: Human growth fa 70.3 1.2 2.6E-05 22.9 -0.2 9 318-326 1-9 (13)
36 cd00054 EGF_CA Calcium-binding 69.7 3.9 8.4E-05 26.4 2.1 30 297-326 3-33 (38)
37 PF07974 EGF_2: EGF-like domai 69.1 3.8 8.2E-05 26.7 1.8 23 303-326 6-28 (32)
38 PHA03265 envelope glycoprotein 67.4 2.3 5.1E-05 42.9 0.8 39 443-482 347-385 (402)
39 PF00008 EGF: EGF-like domain 64.1 2.5 5.4E-05 27.3 0.2 23 304-326 5-29 (32)
40 PF12662 cEGF: Complement Clr- 60.7 3.7 8E-05 24.9 0.5 11 318-328 3-13 (24)
41 PF02439 Adeno_E3_CR2: Adenovi 58.3 2.1 4.5E-05 28.9 -0.9 8 446-453 6-13 (38)
42 PF12947 EGF_3: EGF domain; I 57.8 3.3 7.2E-05 27.7 -0.0 24 303-326 6-30 (36)
43 smart00181 EGF Epidermal growt 52.2 12 0.00025 24.1 1.9 24 303-327 6-30 (35)
44 PF06024 DUF912: Nucleopolyhed 52.1 9.3 0.0002 31.9 1.8 12 465-476 83-94 (101)
45 PF12877 DUF3827: Domain of un 51.1 10 0.00023 41.4 2.4 15 443-457 269-283 (684)
46 PF03302 VSP: Giardia variant- 49.9 18 0.00039 37.9 3.9 30 443-472 367-396 (397)
47 PF14610 DUF4448: Protein of u 49.5 17 0.00036 33.9 3.2 22 448-470 161-182 (189)
48 PF13360 PQQ_2: PQQ-like domai 48.4 1.6E+02 0.0035 27.4 10.0 77 76-154 54-148 (238)
49 KOG0291 WD40-repeat-containing 47.4 4.8E+02 0.01 29.7 15.7 85 95-203 353-448 (893)
50 PRK11138 outer membrane biogen 46.9 76 0.0016 32.8 8.1 57 96-154 121-186 (394)
51 PF08374 Protocadherin: Protoc 45.3 13 0.00028 35.2 1.8 11 443-453 38-48 (221)
52 TIGR01478 STEVOR variant surfa 44.6 7.1 0.00015 38.5 -0.1 7 317-323 143-149 (295)
53 PTZ00370 STEVOR; Provisional 40.4 8 0.00017 38.2 -0.4 7 317-323 143-149 (296)
54 TIGR03300 assembly_YfgL outer 39.7 1.1E+02 0.0023 31.3 7.8 55 97-153 67-130 (377)
55 PF02480 Herpes_gE: Alphaherpe 39.5 9.8 0.00021 40.4 0.0 13 339-351 238-251 (439)
56 PF06365 CD34_antigen: CD34/Po 39.3 17 0.00038 34.2 1.6 18 392-409 36-54 (202)
57 PF01436 NHL: NHL repeat; Int 38.4 56 0.0012 20.2 3.4 20 97-116 6-26 (28)
58 cd05845 Ig2_L1-CAM_like Second 38.2 46 0.001 27.4 3.8 31 76-108 32-63 (95)
59 PRK11138 outer membrane biogen 37.6 4.7E+02 0.01 26.9 17.2 75 77-153 139-230 (394)
60 PF06697 DUF1191: Protein of u 36.7 15 0.00031 36.5 0.7 8 425-432 187-194 (278)
61 PF00954 S_locus_glycop: S-loc 36.1 84 0.0018 26.2 5.3 58 249-312 42-101 (110)
62 PF14991 MLANA: Protein melan- 34.6 12 0.00027 31.6 -0.1 13 463-475 42-54 (118)
63 PF12191 stn_TNFRSF12A: Tumour 32.9 12 0.00025 32.4 -0.6 10 468-477 101-110 (129)
64 PF13360 PQQ_2: PQQ-like domai 32.1 97 0.0021 28.9 5.6 50 102-154 2-62 (238)
65 TIGR03300 assembly_YfgL outer 31.1 3.9E+02 0.0084 27.2 10.3 75 76-154 83-171 (377)
66 PF15330 SIT: SHP2-interacting 30.9 13 0.00028 31.4 -0.6 9 464-472 17-25 (107)
67 PTZ00382 Variant-specific surf 28.6 18 0.0004 29.9 -0.1 32 444-475 63-95 (96)
68 PF05393 Hum_adeno_E3A: Human 28.5 14 0.00031 29.7 -0.7 26 450-476 37-62 (94)
69 PF14670 FXa_inhibition: Coagu 28.0 19 0.00042 24.0 -0.0 14 315-328 17-30 (36)
70 PF02009 Rifin_STEVOR: Rifin/s 27.9 6.5 0.00014 39.5 -3.4 28 449-476 262-289 (299)
71 PHA02887 EGF-like protein; Pro 27.8 62 0.0013 27.7 2.9 30 296-326 83-117 (126)
72 PF12946 EGF_MSP1_1: MSP1 EGF 27.8 18 0.0004 24.4 -0.2 25 303-327 5-31 (37)
73 PF15102 TMEM154: TMEM154 prot 27.7 48 0.001 29.6 2.4 15 463-477 73-87 (146)
74 KOG1219 Uncharacterized conser 27.2 51 0.0011 41.8 3.1 24 303-326 3870-3895(4289)
75 PF14870 PSII_BNR: Photosynthe 27.0 6.6E+02 0.014 25.3 10.7 98 94-218 114-212 (302)
76 KOG1214 Nidogen and related ba 25.6 51 0.0011 37.4 2.6 30 297-327 828-858 (1289)
77 PF14575 EphA2_TM: Ephrin type 24.1 22 0.00049 27.9 -0.3 9 465-473 20-28 (75)
78 PF06247 Plasmod_Pvs28: Plasmo 21.2 33 0.00072 31.9 0.1 27 302-328 49-81 (197)
79 smart00564 PQQ beta-propeller 20.5 1.9E+02 0.0041 17.8 3.6 16 102-117 15-31 (33)
No 1
>PF01453 B_lectin: D-mannose binding lectin; InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]: Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity. Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=99.96 E-value=6.7e-30 Score=219.61 Aligned_cols=111 Identities=41% Similarity=0.709 Sum_probs=78.5
Q ss_pred CCcEEEEcCCCCCCCCC-CceEEEEe-CCeEEEEcCCCceEEEe-ccCCCCCCceEEEEecCCCEEEEeccCCCCcceee
Q 042843 77 ERTIVWVANREQPVSDR-FSSVLRIS-DGNLVLFNESQLPIWST-NLTATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQ 153 (484)
Q Consensus 77 ~~tvVW~ANr~~Pv~~~-~~~~l~l~-~G~LvL~~~~~~~vWst-~~~~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQ 153 (484)
++|+||+|||+.|+... ...+|.|+ ||+|+|+|..++.+|++ ++.+....+..|+|+|+|||||++.. +.+|||
T Consensus 1 ~~tvvW~an~~~p~~~~s~~~~L~l~~dGnLvl~~~~~~~iWss~~t~~~~~~~~~~~L~~~GNlvl~d~~---~~~lW~ 77 (114)
T PF01453_consen 1 PRTVVWVANRNSPLTSSSGNYTLILQSDGNLVLYDSNGSVIWSSNNTSGRGNSGCYLVLQDDGNLVLYDSS---GNVLWQ 77 (114)
T ss_dssp ---------TTEEEEECETTEEEEEETTSEEEEEETTTEEEEE--S-TTSS-SSEEEEEETTSEEEEEETT---SEEEEE
T ss_pred CcccccccccccccccccccccceECCCCeEEEEcCCCCEEEEecccCCccccCeEEEEeCCCCEEEEeec---ceEEEe
Confidence 36899999999999531 25899999 99999999998899999 55443114789999999999999964 479999
Q ss_pred ecccCceeccCCceeeeecCCCCceEEEecCCCCCCC
Q 042843 154 SFDHPAHTWIPGMKLTFNKRNNVSQLITSWKNKENPA 190 (484)
Q Consensus 154 SFd~PTDTlLpgq~l~~n~~~g~~~~L~Sw~s~~dps 190 (484)
||||||||+||||+|+.+..+|....|+||++.+|||
T Consensus 78 Sf~~ptdt~L~~q~l~~~~~~~~~~~~~sw~s~~dps 114 (114)
T PF01453_consen 78 SFDYPTDTLLPGQKLGDGNVTGKNDSLTSWSSNTDPS 114 (114)
T ss_dssp STTSSS-EEEEEET--TSEEEEESTSSEEEESS----
T ss_pred ecCCCccEEEeccCcccCCCccccceEEeECCCCCCC
Confidence 9999999999999999866666556799999999996
No 2
>PF00954 S_locus_glycop: S-locus glycoprotein family; InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=99.95 E-value=1.2e-27 Score=204.68 Aligned_cols=110 Identities=45% Similarity=0.994 Sum_probs=104.2
Q ss_pred EecCCcCCCCceeeeeeccccceeeEEEEEecCCeeEEEEeecCCceeEEEEEccCCcEEEEeeCCCCCCCeEEEeecCC
Q 042843 217 WSSGPWDENAKIFSMVPEMNQNYIYNFSYVSNENESYFTYNVKDSTYTSRAFMDVSGQDKQMNWLPLPTNSWFLFWSQPR 296 (484)
Q Consensus 217 w~sg~w~~~g~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~rl~Ld~dG~l~~y~w~~~~~~~W~~~w~~p~ 296 (484)
||+|+| ||..|+++|++.....+.+.|+.++++.+++|.+.+.+.++|++||++|++++|.|++.. +.|.+.|++|.
T Consensus 1 wrsG~W--nG~~f~g~p~~~~~~~~~~~fv~~~~e~~~t~~~~~~s~~~r~~ld~~G~l~~~~w~~~~-~~W~~~~~~p~ 77 (110)
T PF00954_consen 1 WRSGPW--NGQRFSGIPEMSSNSLYNYSFVSNNEEVYYTYSLSNSSVLSRLVLDSDGQLQRYIWNEST-QSWSVFWSAPK 77 (110)
T ss_pred CCcccc--CCeEECCcccccccceeEEEEEECCCeEEEEEecCCCceEEEEEEeeeeEEEEEEEecCC-CcEEEEEEecc
Confidence 899999 999999999998777889999999999999999988889999999999999999999887 99999999999
Q ss_pred CCCcccccCCCCccccCCCCccccccCCCccCC
Q 042843 297 QQCEVYALCGQFSTCNQQTERFCSCLKGFQQKS 329 (484)
Q Consensus 297 d~C~~~~~CG~~giC~~~~~~~C~C~~GF~p~~ 329 (484)
|+||+|+.||+||+|+.+..+.|+||+||+|++
T Consensus 78 d~Cd~y~~CG~~g~C~~~~~~~C~Cl~GF~P~n 110 (110)
T PF00954_consen 78 DQCDVYGFCGPNGICNSNNSPKCSCLPGFEPKN 110 (110)
T ss_pred cCCCCccccCCccEeCCCCCCceECCCCcCCCc
Confidence 999999999999999988788999999999963
No 3
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=99.94 E-value=5.7e-26 Score=196.20 Aligned_cols=113 Identities=46% Similarity=0.747 Sum_probs=99.5
Q ss_pred CcccCCCeEEecCCeEEEEeecCCCCCCCc-EEEEEEEEecCCCcEEEEcCCCCCCCCCCceEEEEe-CCeEEEEcCCCc
Q 042843 36 QSLSGDQTIVSKGGVFAFGFFNPAPGKSSN-YYIGMWYNKVSERTIVWVANREQPVSDRFSSVLRIS-DGNLVLFNESQL 113 (484)
Q Consensus 36 ~~l~~~~~l~S~~g~F~lGFf~~~~~~~~~-~~lgIw~~~~~~~tvVW~ANr~~Pv~~~~~~~l~l~-~G~LvL~~~~~~ 113 (484)
+.|..|++|+|+++.|++|||.+... . .+.+|||.+.+ .++||.|||+.|.. ..++|.|+ ||+|+|+|.++.
T Consensus 2 ~~l~~~~~l~s~~~~f~~G~~~~~~q---~~dgnlv~~~~~~-~~~vW~snt~~~~~--~~~~l~l~~dGnLvl~~~~g~ 75 (116)
T cd00028 2 NPLSSGQTLVSSGSLFELGFFKLIMQ---SRDYNLILYKGSS-RTVVWVANRDNPSG--SSCTLTLQSDGNLVIYDGSGT 75 (116)
T ss_pred cCcCCCCEEEeCCCcEEEecccCCCC---CCeEEEEEEeCCC-CeEEEECCCCCCCC--CCEEEEEecCCCeEEEcCCCc
Confidence 56889999999999999999998754 4 89999998766 78999999999854 47899999 999999999999
Q ss_pred eEEEeccCCCCCCceEEEEecCCCEEEEeccCCCCcceeeecccC
Q 042843 114 PIWSTNLTATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQSFDHP 158 (484)
Q Consensus 114 ~vWst~~~~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQSFd~P 158 (484)
++|++++.+. .....|+|+|+|||||++.++ ++||||||||
T Consensus 76 ~vW~S~~~~~-~~~~~~~L~ddGnlvl~~~~~---~~~W~Sf~~P 116 (116)
T cd00028 76 VVWSSNTTRV-NGNYVLVLLDDGNLVLYDSDG---NFLWQSFDYP 116 (116)
T ss_pred EEEEecccCC-CCceEEEEeCCCCEEEECCCC---CEEEcCCCCC
Confidence 9999998752 267899999999999999753 7899999999
No 4
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=99.92 E-value=7.5e-25 Score=188.61 Aligned_cols=112 Identities=47% Similarity=0.761 Sum_probs=98.0
Q ss_pred CcccCCCeEEecCCeEEEEeecCCCCCCCcEEEEEEEEecCCCcEEEEcCCCCCCCCCCceEEEEe-CCeEEEEcCCCce
Q 042843 36 QSLSGDQTIVSKGGVFAFGFFNPAPGKSSNYYIGMWYNKVSERTIVWVANREQPVSDRFSSVLRIS-DGNLVLFNESQLP 114 (484)
Q Consensus 36 ~~l~~~~~l~S~~g~F~lGFf~~~~~~~~~~~lgIw~~~~~~~tvVW~ANr~~Pv~~~~~~~l~l~-~G~LvL~~~~~~~ 114 (484)
+.|..|+.|+|+++.|++|||.+... .++.+|||...+ .++||+|||+.|+.. ++.|.|+ ||+|+|+|.++.+
T Consensus 2 ~~l~~~~~l~s~~~~f~~G~~~~~~q---~dgnlV~~~~~~-~~~vW~snt~~~~~~--~~~l~l~~dGnLvl~~~~g~~ 75 (114)
T smart00108 2 NTLSSGQTLVSGNSLFELGFFTLIMQ---NDYNLILYKSSS-RTVVWVANRDNPVSD--SCTLTLQSDGNLVLYDGDGRV 75 (114)
T ss_pred cccCCCCEEecCCCcEeeeccccCCC---CCEEEEEEECCC-CcEEEECCCCCCCCC--CEEEEEeCCCCEEEEeCCCCE
Confidence 56888999999999999999998653 688999998876 789999999999873 5889999 9999999999999
Q ss_pred EEEeccCCCCCCceEEEEecCCCEEEEeccCCCCcceeeeccc
Q 042843 115 IWSTNLTATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQSFDH 157 (484)
Q Consensus 115 vWst~~~~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQSFd~ 157 (484)
+|++++... .+...|+|+|+|||||++..+ +++||||||
T Consensus 76 vW~S~t~~~-~~~~~~~L~ddGnlvl~~~~~---~~~W~Sf~~ 114 (114)
T smart00108 76 VWSSNTTGA-NGNYVLVLLDDGNLVIYDSDG---NFLWQSFDY 114 (114)
T ss_pred EEEecccCC-CCceEEEEeCCCCEEEECCCC---CEEeCCCCC
Confidence 999988622 257889999999999998753 799999997
No 5
>PF08276 PAN_2: PAN-like domain; InterPro: IPR013227 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs
Probab=99.54 E-value=6.3e-15 Score=113.99 Aligned_cols=58 Identities=43% Similarity=0.995 Sum_probs=50.4
Q ss_pred CCccEEEEecccCCCCCccc--ccCCHHHHHHHHhhcCCeEEEEeCC----CceEEecccccce
Q 042843 359 KSDQFFQYSNMKLPKHPQSV--AVGGIRECETHCMNNCSCTAYAYKD----NACSIWVGSFVGL 416 (484)
Q Consensus 359 ~~~~F~~l~~v~~p~~~~~~--~~~~~~~C~~~CL~nCSC~Ay~y~~----~~C~~w~~~l~~~ 416 (484)
.+|+|++|++|++|++...+ ...++++|+++||+||||+||+|.+ ++|++|+++|+|+
T Consensus 3 ~~d~F~~l~~~~~p~~~~~~~~~~~s~~~C~~~Cl~nCsC~Ayay~~~~~~~~C~lW~~~L~d~ 66 (66)
T PF08276_consen 3 SGDGFLKLPNMKLPDFDNAIVDSSVSLEECEKACLSNCSCTAYAYSNLSGGGGCLLWYGDLVDL 66 (66)
T ss_pred CCCEEEEECCeeCCCCcceeeecCCCHHHHHhhcCCCCCEeeEEeeccCCCCEEEEEcCEeecC
Confidence 46899999999999985544 3489999999999999999999973 5799999999874
No 6
>cd01098 PAN_AP_plant Plant PAN/APPLE-like domain; present in plant S-receptor protein kinases and secreted glycoproteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions. S-receptor protein kinases and S-locus glycoproteins are involved in sporophytic self-incompatibility response in Brassica, one of probably many molecular mechanisms, by which hermaphrodite flowering plants avoid self-fertilization.
Probab=99.48 E-value=1.1e-13 Score=112.00 Aligned_cols=81 Identities=37% Similarity=0.909 Sum_probs=63.5
Q ss_pred CCCCCCCCCCCcCCccEEEEecccCCCCCcccccCCHHHHHHHHhhcCCeEEEEeCC--CceEEecccccceEEecCCCe
Q 042843 347 PLQCENISPANRKSDQFFQYSNMKLPKHPQSVAVGGIRECETHCMNNCSCTAYAYKD--NACSIWVGSFVGLQQLQGGGD 424 (484)
Q Consensus 347 ~l~C~~~~~~~~~~~~F~~l~~v~~p~~~~~~~~~~~~~C~~~CL~nCSC~Ay~y~~--~~C~~w~~~l~~~~~~~~~~~ 424 (484)
+++|..+. ..+.|++++++++|+........++++|++.||+||+|+||+|.+ ++|++|..++.+.+.....+.
T Consensus 2 ~~~C~~~~----~~~~f~~~~~~~~~~~~~~~~~~s~~~C~~~Cl~nCsC~a~~~~~~~~~C~~~~~~~~~~~~~~~~~~ 77 (84)
T cd01098 2 PLNCGGDG----STDGFLKLPDVKLPDNASAITAISLEECREACLSNCSCTAYAYNNGSGGCLLWNGLLNNLRSLSSGGG 77 (84)
T ss_pred CcccCCCC----CCCEEEEeCCeeCCCchhhhccCCHHHHHHHHhcCCCcceeeecCCCCeEEEEeceecceEeecCCCc
Confidence 45675321 136899999999998754434589999999999999999999974 679999999998876543335
Q ss_pred EEEEEec
Q 042843 425 IIYIKLA 431 (484)
Q Consensus 425 ~~yikv~ 431 (484)
++||||+
T Consensus 78 ~~yiKv~ 84 (84)
T cd01098 78 TLYLRLA 84 (84)
T ss_pred EEEEEeC
Confidence 9999985
No 7
>cd00129 PAN_APPLE PAN/APPLE-like domain; present in N-terminal (N) domains of plasminogen/ hepatocyte growth factor proteins, plasma prekallikrein/coagulation factor XI and microneme antigen proteins, plant receptor-like protein kinases, and various nematode and leech anti-platelet proteins. Common structural features include two disulfide bonds that link the alpha-helix to the central region of the protein. PAN domains have significant functional versatility, fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=99.38 E-value=6.6e-13 Score=106.24 Aligned_cols=67 Identities=16% Similarity=0.313 Sum_probs=57.4
Q ss_pred CccEEEEecccCCCCCcccccCCHHHHHHHHhh---cCCeEEEEeC--CCceEEecccc-cceEEecCCCeEEEEEe
Q 042843 360 SDQFFQYSNMKLPKHPQSVAVGGIRECETHCMN---NCSCTAYAYK--DNACSIWVGSF-VGLQQLQGGGDIIYIKL 430 (484)
Q Consensus 360 ~~~F~~l~~v~~p~~~~~~~~~~~~~C~~~CL~---nCSC~Ay~y~--~~~C~~w~~~l-~~~~~~~~~~~~~yikv 430 (484)
+..|+++.+|++|++.. .++++|+++|++ ||||+||+|. +.+|++|.++| .++++..+.+.++|||.
T Consensus 8 ~g~fl~~~~~klpd~~~----~s~~eC~~~Cl~~~~nCsC~Aya~~~~~~gC~~W~~~l~~d~~~~~~~g~~Ly~r~ 80 (80)
T cd00129 8 AGTTLIKIALKIKTTKA----NTADECANRCEKNGLPFSCKAFVFAKARKQCLWFPFNSMSGVRKEFSHGFDLYENK 80 (80)
T ss_pred CCeEEEeecccCCcccc----cCHHHHHHHHhcCCCCCCceeeeccCCCCCeEEecCcchhhHHhccCCCceeEeEC
Confidence 56799999999998754 678999999999 9999999995 35899999999 99887765555999983
No 8
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=98.65 E-value=1.8e-07 Score=80.38 Aligned_cols=87 Identities=30% Similarity=0.470 Sum_probs=62.2
Q ss_pred EEEEe-CCeEEEEcCC-CceEEEeccCCCCCCceEEEEecCCCEEEEeccCCCCcceeeecccCceeccCCceeeeecCC
Q 042843 97 VLRIS-DGNLVLFNES-QLPIWSTNLTATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQSFDHPAHTWIPGMKLTFNKRN 174 (484)
Q Consensus 97 ~l~l~-~G~LvL~~~~-~~~vWst~~~~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQSFd~PTDTlLpgq~l~~n~~~ 174 (484)
++.++ ||+||+++.. +.++|++++..+......+.|.++|||||++.++ .++|+|=. +
T Consensus 23 ~~~~q~dgnlV~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g---~~vW~S~t--~--------------- 82 (114)
T smart00108 23 TLIMQNDYNLILYKSSSRTVVWVANRDNPVSDSCTLTLQSDGNLVLYDGDG---RVVWSSNT--T--------------- 82 (114)
T ss_pred ccCCCCCEEEEEEECCCCcEEEECCCCCCCCCCEEEEEeCCCCEEEEeCCC---CEEEEecc--c---------------
Confidence 35567 9999999865 4789999986542133788999999999998753 68999810 0
Q ss_pred CCceEEEecCCCCCCCCceEEEEEcCCCCcEEEEEeeCCeeEEec
Q 042843 175 NVSQLITSWKNKENPAPGLFSLERAPDGSNQYVMLWNRSEQYWSS 219 (484)
Q Consensus 175 g~~~~L~Sw~s~~dps~G~y~l~~~~~g~~~~~l~~~~~~~Yw~s 219 (484)
...+.|.+.|+++|+ |+++-...++.|.+
T Consensus 83 --------------~~~~~~~~~L~ddGn--lvl~~~~~~~~W~S 111 (114)
T smart00108 83 --------------GANGNYVLVLLDDGN--LVIYDSDGNFLWQS 111 (114)
T ss_pred --------------CCCCceEEEEeCCCC--EEEECCCCCEEeCC
Confidence 124568899999998 66532334578875
No 9
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=98.58 E-value=4.2e-07 Score=78.30 Aligned_cols=83 Identities=30% Similarity=0.443 Sum_probs=60.2
Q ss_pred CCeEEEEcCC-CceEEEeccCCCCCCceEEEEecCCCEEEEeccCCCCcceeeecccCceeccCCceeeeecCCCCceEE
Q 042843 102 DGNLVLFNES-QLPIWSTNLTATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQSFDHPAHTWIPGMKLTFNKRNNVSQLI 180 (484)
Q Consensus 102 ~G~LvL~~~~-~~~vWst~~~~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQSFd~PTDTlLpgq~l~~n~~~g~~~~L 180 (484)
||+||+++.. ++++|++++..+......+.|.++|||||++.++ .++|+|=-.
T Consensus 30 dgnlv~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g---~~vW~S~~~----------------------- 83 (116)
T cd00028 30 DYNLILYKGSSRTVVWVANRDNPSGSSCTLTLQSDGNLVIYDGSG---TVVWSSNTT----------------------- 83 (116)
T ss_pred eEEEEEEeCCCCeEEEECCCCCCCCCCEEEEEecCCCeEEEcCCC---cEEEEeccc-----------------------
Confidence 7899998765 4789999986532246789999999999998754 689987310
Q ss_pred EecCCCCCCCCceEEEEEcCCCCcEEEEEeeCCeeEEecC
Q 042843 181 TSWKNKENPAPGLFSLERAPDGSNQYVMLWNRSEQYWSSG 220 (484)
Q Consensus 181 ~Sw~s~~dps~G~y~l~~~~~g~~~~~l~~~~~~~Yw~sg 220 (484)
...+.+.+.|+++|+ |+++-....+.|.+.
T Consensus 84 --------~~~~~~~~~L~ddGn--lvl~~~~~~~~W~Sf 113 (116)
T cd00028 84 --------RVNGNYVLVLLDDGN--LVLYDSDGNFLWQSF 113 (116)
T ss_pred --------CCCCceEEEEeCCCC--EEEECCCCCEEEcCC
Confidence 024568999999998 665322346788864
No 10
>smart00473 PAN_AP divergent subfamily of APPLE domains. Apple-like domains present in Plasminogen, C. elegans hypothetical ORFs and the extracellular portion of plant receptor-like protein kinases. Predicted to possess protein- and/or carbohydrate-binding functions.
Probab=98.57 E-value=2.1e-07 Score=73.40 Aligned_cols=70 Identities=33% Similarity=0.793 Sum_probs=53.3
Q ss_pred CccEEEEecccCCCCCcc-cccCCHHHHHHHHhh-cCCeEEEEeC--CCceEEec-ccccceEEecCCCeEEEEE
Q 042843 360 SDQFFQYSNMKLPKHPQS-VAVGGIRECETHCMN-NCSCTAYAYK--DNACSIWV-GSFVGLQQLQGGGDIIYIK 429 (484)
Q Consensus 360 ~~~F~~l~~v~~p~~~~~-~~~~~~~~C~~~CL~-nCSC~Ay~y~--~~~C~~w~-~~l~~~~~~~~~~~~~yik 429 (484)
.+.|..++++.+++.... ....++++|++.|++ +|+|.||.|. +++|.+|. +++.+.+.....+.++|.|
T Consensus 3 ~~~f~~~~~~~l~~~~~~~~~~~s~~~C~~~C~~~~~~C~s~~y~~~~~~C~l~~~~~~~~~~~~~~~~~~~y~~ 77 (78)
T smart00473 3 DDCFVRLPNTKLPGFSRIVISVASLEECASKCLNSNCSCRSFTYNNGTKGCLLWSESSLGDARLFPSGGVDLYEK 77 (78)
T ss_pred CceeEEecCccCCCCcceeEcCCCHHHHHHHhCCCCCceEEEEEcCCCCEEEEeeCCccccceecccCCceeEEe
Confidence 357999999999854332 234799999999999 9999999997 46799998 7777776433333377766
No 11
>PF01453 B_lectin: D-mannose binding lectin; InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]: Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity. Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=98.06 E-value=5.7e-05 Score=64.85 Aligned_cols=100 Identities=21% Similarity=0.368 Sum_probs=65.6
Q ss_pred CCeEEecCCeEEEEeecCCCCCCCcEEEEEEEEecCCCcEEEEc-CCCCCCCCCCceEEEEe-CCeEEEEcCCCceEEEe
Q 042843 41 DQTIVSKGGVFAFGFFNPAPGKSSNYYIGMWYNKVSERTIVWVA-NREQPVSDRFSSVLRIS-DGNLVLFNESQLPIWST 118 (484)
Q Consensus 41 ~~~l~S~~g~F~lGFf~~~~~~~~~~~lgIw~~~~~~~tvVW~A-Nr~~Pv~~~~~~~l~l~-~G~LvL~~~~~~~vWst 118 (484)
.+.+.+.+|.+.|-|...++ .. .|. ...++||.. +...... ..+.+.|. ||||||+|..+.++|++
T Consensus 11 ~~p~~~~s~~~~L~l~~dGn-----Lv---l~~--~~~~~iWss~~t~~~~~--~~~~~~L~~~GNlvl~d~~~~~lW~S 78 (114)
T PF01453_consen 11 NSPLTSSSGNYTLILQSDGN-----LV---LYD--SNGSVIWSSNNTSGRGN--SGCYLVLQDDGNLVLYDSSGNVLWQS 78 (114)
T ss_dssp TEEEEECETTEEEEEETTSE-----EE---EEE--TTTEEEEE--S-TTSS---SSEEEEEETTSEEEEEETTSEEEEES
T ss_pred ccccccccccccceECCCCe-----EE---EEc--CCCCEEEEecccCCccc--cCeEEEEeCCCCEEEEeecceEEEee
Confidence 45666655888998887542 22 243 245779999 4444432 26789999 99999999999999999
Q ss_pred ccCCCCCCceEEEEec--CCCEEEEeccCCCCcceeeecccCc
Q 042843 119 NLTATSRRSVEAVLLD--EGNLVLRDLSNNLSKPLWQSFDHPA 159 (484)
Q Consensus 119 ~~~~~~~~~~~a~LlD--sGNLVL~~~~~~~~~~lWQSFd~PT 159 (484)
... + ....+.+++ .||++ +... ..+.|.|=+.|+
T Consensus 79 f~~-p--tdt~L~~q~l~~~~~~-~~~~---~~~sw~s~~dps 114 (114)
T PF01453_consen 79 FDY-P--TDTLLPGQKLGDGNVT-GKND---SLTSWSSNTDPS 114 (114)
T ss_dssp TTS-S--S-EEEEEET--TSEEE-EEST---SSEEEESS----
T ss_pred cCC-C--ccEEEeccCcccCCCc-cccc---eEEeECCCCCCC
Confidence 432 2 467777777 89998 5432 358999877764
No 12
>cd01100 APPLE_Factor_XI_like Subfamily of PAN/APPLE-like domains; present in plasma prekallikrein/coagulation factor XI, microneme antigen proteins, and a few prokaryotic proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=97.35 E-value=0.0002 Score=56.26 Aligned_cols=47 Identities=19% Similarity=0.430 Sum_probs=34.3
Q ss_pred EecccCCCCCccc-ccCCHHHHHHHHhhcCCeEEEEeCC--CceEEeccc
Q 042843 366 YSNMKLPKHPQSV-AVGGIRECETHCMNNCSCTAYAYKD--NACSIWVGS 412 (484)
Q Consensus 366 l~~v~~p~~~~~~-~~~~~~~C~~~CL~nCSC~Ay~y~~--~~C~~w~~~ 412 (484)
+++++++...... ...+.++|++.|+.+|+|.||.|.. +.|+++...
T Consensus 9 ~~~~~~~g~d~~~~~~~s~~~Cq~~C~~~~~C~afT~~~~~~~C~lk~~~ 58 (73)
T cd01100 9 GSNVDFRGGDLSTVFASSAEQCQAACTADPGCLAFTYNTKSKKCFLKSSE 58 (73)
T ss_pred cCCCccccCCcceeecCCHHHHHHHcCCCCCceEEEEECCCCeEEcccCC
Confidence 3566666543222 2368999999999999999999963 569997643
No 13
>PF04478 Mid2: Mid2 like cell wall stress sensor; InterPro: IPR007567 This family represents a region near the C terminus of Mid2, which contains a transmembrane region. The remainder of the protein sequence is serine-rich and of low complexity, and is therefore impossible to align accurately. Mid2 is thought to act as a mechanosensor of cell wall stress. The C-terminal cytoplasmic region of Mid2 is known to interact with Rom2, a guanine nucleotide exchange factor (GEF) for Rho1, which is part of the cell wall integrity signalling pathway [].
Probab=93.67 E-value=0.12 Score=46.07 Aligned_cols=33 Identities=30% Similarity=0.405 Sum_probs=16.9
Q ss_pred ccceEEEEEehHHHHHHHHHHHHheeeeeeeccC
Q 042843 441 KKGVVIGGVVGSVAVVALIGLIMLVYLGRRKTAT 474 (484)
Q Consensus 441 ~~~~~i~~~v~~~~~~~~~~~~~~~~~~~r~~~~ 474 (484)
.++++|+++||+-++++|+ +++++|++++|+++
T Consensus 47 nknIVIGvVVGVGg~ill~-il~lvf~~c~r~kk 79 (154)
T PF04478_consen 47 NKNIVIGVVVGVGGPILLG-ILALVFIFCIRRKK 79 (154)
T ss_pred CccEEEEEEecccHHHHHH-HHHhheeEEEeccc
Confidence 4468899998754333322 22333444444433
No 14
>PF08693 SKG6: Transmembrane alpha-helix domain; InterPro: IPR014805 SKG6 and AXL2 are membrane proteins that show polarised intracellular localisation [, ]. This entry represents the highly conserved transmembrane alpha-helical domain found in these proteins [, ]. The full-length AXL2 protein has a negative regulatory function in cytokinesis [].
Probab=91.86 E-value=0.29 Score=33.50 Aligned_cols=10 Identities=0% Similarity=-0.184 Sum_probs=4.5
Q ss_pred heeeeeeecc
Q 042843 464 LVYLGRRKTA 473 (484)
Q Consensus 464 ~~~~~~r~~~ 473 (484)
+++++|+||+
T Consensus 30 ~~l~~~~rR~ 39 (40)
T PF08693_consen 30 AFLFFWYRRK 39 (40)
T ss_pred HHhheEEecc
Confidence 3444455544
No 15
>PF15102 TMEM154: TMEM154 protein family
Probab=91.65 E-value=0.26 Score=43.64 Aligned_cols=25 Identities=12% Similarity=0.089 Sum_probs=9.5
Q ss_pred EEEEEehHHHHHHHHHHHHheeeee
Q 042843 445 VIGGVVGSVAVVALIGLIMLVYLGR 469 (484)
Q Consensus 445 ~i~~~v~~~~~~~~~~~~~~~~~~~ 469 (484)
++.++|..++++++++++++++++.
T Consensus 58 iLmIlIP~VLLvlLLl~vV~lv~~~ 82 (146)
T PF15102_consen 58 ILMILIPLVLLVLLLLSVVCLVIYY 82 (146)
T ss_pred EEEEeHHHHHHHHHHHHHHHheeEE
Confidence 3333333233333333334444433
No 16
>smart00223 APPLE APPLE domain. Four-fold repeat in plasma kallikrein and coagulation factor XI. Factor XI apple 3 mediates binding to platelets. Factor XI apple 1 binds high-molecular-mass kininogen. Apple 4 in factor XI mediates dimer formation and binds to factor XIIa. Mutations in apple 4 cause factor XI deficiency, an inherited bleeding disorder.
Probab=91.17 E-value=0.27 Score=39.23 Aligned_cols=45 Identities=16% Similarity=0.398 Sum_probs=33.9
Q ss_pred ecccCCCCCcc-cccCCHHHHHHHHhhcCCeEEEEeCC--C---ceEEecc
Q 042843 367 SNMKLPKHPQS-VAVGGIRECETHCMNNCSCTAYAYKD--N---ACSIWVG 411 (484)
Q Consensus 367 ~~v~~p~~~~~-~~~~~~~~C~~~CL~nCSC~Ay~y~~--~---~C~~w~~ 411 (484)
+|++++..... +...+.++|++.|..+=.|.||.|.. . .|+++..
T Consensus 7 ~~~df~G~Dl~~~~~~~~~~Cq~~Ct~~~~C~~FTf~~~~~~~~~C~LK~s 57 (79)
T smart00223 7 KNVDFRGSDINTVYVPSAQVCQKRCTSHPRCLFFTFSTNEPPEEKCLLKDS 57 (79)
T ss_pred cCccccCceeeeeecCCHHHHHHhhcCCCCccEEEeeCCCCCCCEeEeCcC
Confidence 46777665332 33478999999999999999999953 3 6998743
No 17
>PF00024 PAN_1: PAN domain This Prosite entry concerns apple domains, a subset of PAN domains; InterPro: IPR003014 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs It has been shown that, the N-terminal N domains of members of the plasminogen/hepatocyte growth factor family, the apple domains of the plasma prekallikrein/coagulation factor XI family, and domains of various nematode proteins belong to the same module superfamily, the PAN module []. PAN contains a conserved core of three disulphide bridges. In some members of the family there is an additional fourth disulphide bridge that links the N and C termini of the domain.; PDB: 1GP9_C 2QJ2_B 1GMO_H 1NK1_B 3MKP_B 1BHT_B 3HN4_A 1GMN_A 3HMS_A 3HMT_B ....
Probab=91.00 E-value=0.33 Score=37.75 Aligned_cols=54 Identities=19% Similarity=0.449 Sum_probs=39.0
Q ss_pred cEEEEecccCCCCCccccc-CCHHHHHHHHhhcCC-eEEEEeCC--CceEEecccccc
Q 042843 362 QFFQYSNMKLPKHPQSVAV-GGIRECETHCMNNCS-CTAYAYKD--NACSIWVGSFVG 415 (484)
Q Consensus 362 ~F~~l~~v~~p~~~~~~~~-~~~~~C~~~CL~nCS-C~Ay~y~~--~~C~~w~~~l~~ 415 (484)
.|..+++..+......... .++++|.+.|+.+=. |.+|.|.. ..|.+....-..
T Consensus 3 ~f~~~~~~~l~~~~~~~~~v~s~~~C~~~C~~~~~~C~s~~y~~~~~~C~L~~~~~~~ 60 (79)
T PF00024_consen 3 AFERIPGYRLSGHSIKEINVPSLEECAQLCLNEPRRCKSFNYDPSSKTCYLSSSDRSS 60 (79)
T ss_dssp TEEEEEEEEEESCEEEEEEESSHHHHHHHHHHSTT-ESEEEEETTTTEEEEECSSSSS
T ss_pred CeEEECCEEEeCCcceEEcCCCHHHHHhhcCcCcccCCeEEEECCCCEEEEcCCCCCc
Confidence 4777777776654222223 589999999999999 99999964 469997654433
No 18
>PF08277 PAN_3: PAN-like domain; InterPro: IPR006583 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs The PAN-3 or CW is a domain associated with a number of Caenorhabditis elegans hypothetical proteins.
Probab=90.07 E-value=1.5 Score=33.62 Aligned_cols=39 Identities=21% Similarity=0.540 Sum_probs=30.3
Q ss_pred cCCHHHHHHHHhhcCCeEEEEeCCCceEEec-ccccceEE
Q 042843 380 VGGIRECETHCMNNCSCTAYAYKDNACSIWV-GSFVGLQQ 418 (484)
Q Consensus 380 ~~~~~~C~~~CL~nCSC~Ay~y~~~~C~~w~-~~l~~~~~ 418 (484)
..+.++|-+.|..+=.|.++.+..+.|.++. +++..+++
T Consensus 19 ~~sw~~Cv~~C~~~~~C~la~~~~~~C~~y~~~~i~~v~~ 58 (71)
T PF08277_consen 19 NTSWDDCVQKCYNDENCVLAYFDSGKCYLYNYGSISTVQK 58 (71)
T ss_pred CCCHHHHhHHhCCCCEEEEEEeCCCCEEEEEcCCEEEEEE
Confidence 4778999999999999999988866799864 33334443
No 19
>smart00605 CW CW domain.
Probab=87.03 E-value=2.1 Score=35.05 Aligned_cols=53 Identities=15% Similarity=0.429 Sum_probs=38.2
Q ss_pred cCCHHHHHHHHhhcCCeEEEEeCC-CceEEec-ccccceEEecCCC-eEEEEEecc
Q 042843 380 VGGIRECETHCMNNCSCTAYAYKD-NACSIWV-GSFVGLQQLQGGG-DIIYIKLAA 432 (484)
Q Consensus 380 ~~~~~~C~~~CL~nCSC~Ay~y~~-~~C~~w~-~~l~~~~~~~~~~-~~~yikv~~ 432 (484)
..+.++|.+.|..+..|+.+.... ..|.+.. +.+..+++..... ..+=+|+..
T Consensus 21 ~~sw~~Ci~~C~~~~~Cvlay~~~~~~C~~f~~~~~~~v~~~~~~~~~~VAfK~~~ 76 (94)
T smart00605 21 TLSWDECIQKCYEDSNCVLAYGNSSETCYLFSYGTVLTVKKLSSSSGKKVAFKVST 76 (94)
T ss_pred CCCHHHHHHHHhCCCceEEEecCCCCceEEEEcCCeEEEEEccCCCCcEEEEEEeC
Confidence 467899999999999999876654 6798765 3466666654323 367777753
No 20
>PF02009 Rifin_STEVOR: Rifin/stevor family; InterPro: IPR002858 Malaria is still a major cause of mortality in many areas of the world. Plasmodium falciparum causes the most severe human form of the disease and is responsible for most fatalities. Severe cases of malaria can occur when the parasite invades and then proliferates within red blood cell erythrocytes. The parasite produces many variant antigenic proteins, encoded by multigene families, which are present on the surface of the infected erythrocyte and play important roles in virulence. A crucial survival mechanism for the malaria parasite is its ability to evade the immune response by switching these variant surface antigens. The high virulence of P. falciparum relative to other malarial parasites is in large part due to the fact that in this organism many of these surface antigens mediate the binding of infected erythrocytes to the vascular endothelium (cytoadherence) and non-infected erythrocytes (rosetting). This can lead to the accumulation of infected cells in the vasculature of a variety of organs, blocking the blood flow and reducing the oxygen supply. Clinical symptoms of severe infection can include fever, progressive anaemia, multi-organ dysfunction and coma. For more information see []. Several multicopy gene families have been described in Plasmodium falciparum, including the stevor family of subtelomeric open reading frames and the rif interspersed repetitive elements. Both families contain three predicted transmembrane segments. It has been proposed that stevor and rif are members of a larger superfamily that code for variant surface antigens [].
Probab=85.41 E-value=0.14 Score=51.33 Aligned_cols=20 Identities=20% Similarity=0.326 Sum_probs=13.0
Q ss_pred HHHHHheeeeeeeccCcccc
Q 042843 459 IGLIMLVYLGRRKTATVTTK 478 (484)
Q Consensus 459 ~~~~~~~~~~~r~~~~~~~~ 478 (484)
+|++++.|++||+|||+|.+
T Consensus 269 VLIMvIIYLILRYRRKKKmk 288 (299)
T PF02009_consen 269 VLIMVIIYLILRYRRKKKMK 288 (299)
T ss_pred HHHHHHHHHHHHHHHHhhhh
Confidence 33445677778888776665
No 21
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=84.94 E-value=1.4 Score=36.50 Aligned_cols=29 Identities=24% Similarity=0.165 Sum_probs=12.9
Q ss_pred eEEEEEehHHHHHHHHHHHHheeeeeeec
Q 042843 444 VVIGGVVGSVAVVALIGLIMLVYLGRRKT 472 (484)
Q Consensus 444 ~~i~~~v~~~~~~~~~~~~~~~~~~~r~~ 472 (484)
.|.+++|++++++.+++.++++++++|+|
T Consensus 67 aiagi~vg~~~~v~~lv~~l~w~f~~r~k 95 (96)
T PTZ00382 67 AIAGISVAVVAVVGGLVGFLCWWFVCRGK 95 (96)
T ss_pred cEEEEEeehhhHHHHHHHHHhheeEEeec
Confidence 45566665543322222333444445443
No 22
>PF07645 EGF_CA: Calcium-binding EGF domain; InterPro: IPR001881 A sequence of about forty amino-acid residues found in epidermal growth factor (EGF) has been shown [, , , , , ] to be present in a large number of membrane-bound and extracellular, mostly animal, proteins. Many of these proteins require calcium for their biological function and a calcium-binding site has been found at the N terminus of some EGF-like domains []. Calcium-binding may be crucial for numerous protein-protein interactions. For human coagulation factor IX it has been shown [] that the calcium-ligands form a pentagonal bipyramid. The first, third and fourth conserved negatively charged or polar residues are side chain ligands. The latter is possibly hydroxylated (see aspartic acid and asparagine hydroxylation site) []. A conserved aromatic residue, as well as the second conserved negative residue, are thought to be involved in stabilising the calcium-binding site. As in non-calcium binding EGF-like domains, there are six conserved cysteines and the structure of both types is very similar as calcium-binding induces only strictly local structural changes []. +------------------+ +---------+ | | | | nxnnC-x(3,14)-C-x(3,7)-CxxbxxxxaxC-x(1,6)-C-x(8,13)-Cx | | +------------------+ 'n': negatively charged or polar residue [DEQN] 'b': possibly beta-hydroxylated residue [DN] 'a': aromatic amino acid 'C': cysteine, involved in disulphide bond 'x': any amino acid. ; GO: 0005509 calcium ion binding; PDB: 2VJ3_A 1TOZ_A 1LMJ_A 1UZQ_A 1UZK_A 1UZJ_B 1UZP_A 1EMO_A 1EMN_A 2RR0_A ....
Probab=84.44 E-value=0.51 Score=32.64 Aligned_cols=31 Identities=26% Similarity=0.604 Sum_probs=24.6
Q ss_pred CCCccc-ccCCCCccccC-CCCccccccCCCcc
Q 042843 297 QQCEVY-ALCGQFSTCNQ-QTERFCSCLKGFQQ 327 (484)
Q Consensus 297 d~C~~~-~~CG~~giC~~-~~~~~C~C~~GF~p 327 (484)
|+|... ..|..++.|.. ..+-.|.|++||+.
T Consensus 3 dEC~~~~~~C~~~~~C~N~~Gsy~C~C~~Gy~~ 35 (42)
T PF07645_consen 3 DECAEGPHNCPENGTCVNTEGSYSCSCPPGYEL 35 (42)
T ss_dssp STTTTTSSSSSTTSEEEEETTEEEEEESTTEEE
T ss_pred cccCCCCCcCCCCCEEEcCCCCEEeeCCCCcEE
Confidence 678775 47999999974 34568999999984
No 23
>PF01034 Syndecan: Syndecan domain; InterPro: IPR001050 The syndecans are transmembrane proteoglycans which are involved in the organisation of cytoskeleton and/or actin microfilaments, and have important roles as cell surface receptors during cell-cell and/or cell-matrix interactions [, ]. Structurally, these proteins consist of four separate domains: A signal sequence; An extracellular domain (ectodomain) of variable length whose sequence is not evolutionary conserved in the various forms of syndecans. The ectodomain contains the sites of attachment of the heparan sulphate glycosaminoglycan side chains; A transmembrane region; A highly conserved cytoplasmic domain of about 30 to 35 residues, which could interact with cytoskeletal proteins. The proteins known to belong to this family are: Syndecan 1. Syndecan 2 or fibroglycan. Syndecan 3 or neuroglycan or N-syndecan. Syndecan 4 or amphiglycan or ryudocan. Drosophila syndecan. Caenorhabditis elegans probable syndecan (F57C7.3). Syndecan-4, a transmembrane heparan sulphate proteoglycan, is a coreceptor with integrins in cell adhesion. It has been suggested to form a ternary signalling complex with protein kinase Calpha and phosphatidylinositol 4,5-bisphosphate (PIP2). Structural studies have demonstrated that the cytoplasmic domain undergoes a conformational transition and forms a symmetric dimer in the presence of phospholipid activator PIP2, and whose overall structure in solution exhibits a twisted clamp shape having a cavity in the centre of dimeric interface. In addition, it has been observed that the syndecan-4 variable domain interacts, strongly, not only with fatty acyl groups but also the anionic head group of PIP2. These findings indicate that PIP2 promotes oligomerisation of the syndecan-4 cytoplasmic domain for transmembrane signalling and cell-matrix adhesion [, ].; GO: 0008092 cytoskeletal protein binding, 0016020 membrane; PDB: 1EJQ_B 1EJP_B 1YBO_C 1OBY_Q.
Probab=84.20 E-value=0.32 Score=36.79 Aligned_cols=16 Identities=13% Similarity=0.245 Sum_probs=0.6
Q ss_pred HHheeeeeeeccCccc
Q 042843 462 IMLVYLGRRKTATVTT 477 (484)
Q Consensus 462 ~~~~~~~~r~~~~~~~ 477 (484)
++++++++|.|+|.++
T Consensus 27 lLIlf~iyR~rkkdEG 42 (64)
T PF01034_consen 27 LLILFLIYRMRKKDEG 42 (64)
T ss_dssp ----------S-----
T ss_pred HHHHHHHHHHHhcCCC
Confidence 3445556666655544
No 24
>PF14295 PAN_4: PAN domain; PDB: 2YIL_E 2YIP_C 2YIO_A.
Probab=83.66 E-value=0.9 Score=32.16 Aligned_cols=24 Identities=21% Similarity=0.696 Sum_probs=17.5
Q ss_pred cCCHHHHHHHHhhcCCeEEEEeCC
Q 042843 380 VGGIRECETHCMNNCSCTAYAYKD 403 (484)
Q Consensus 380 ~~~~~~C~~~CL~nCSC~Ay~y~~ 403 (484)
..+.++|.+.|..+=.|.+|.|..
T Consensus 15 ~~s~~~C~~~C~~~~~C~~~~~~~ 38 (51)
T PF14295_consen 15 ASSPEECQAACAADPGCQAFTFNP 38 (51)
T ss_dssp ---HHHHHHHHHTSTT--EEEEET
T ss_pred CCCHHHHHHHccCCCCCCEEEEEC
Confidence 368899999999999999999854
No 25
>cd00053 EGF Epidermal growth factor domain, found in epidermal growth factor (EGF) presents in a large number of proteins, mostly animal; the list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied; the functional significance of EGF-like domains in what appear to be unrelated proteins is not yet clear; a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase); the domain includes six cysteine residues which have been shown to be involved in disulfide bonds; the main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet; Subdomains between the conserved cysteines vary in length; the region between the 5th and 6th cysteine contains two conserved glycines of which at least one is present in most EGF-like domains; a subset of these bind calcium.
Probab=81.08 E-value=1.2 Score=28.43 Aligned_cols=29 Identities=24% Similarity=0.590 Sum_probs=21.3
Q ss_pred CcccccCCCCccccCC-CCccccccCCCcc
Q 042843 299 CEVYALCGQFSTCNQQ-TERFCSCLKGFQQ 327 (484)
Q Consensus 299 C~~~~~CG~~giC~~~-~~~~C~C~~GF~p 327 (484)
|.....|..++.|... ....|.|++||..
T Consensus 2 C~~~~~C~~~~~C~~~~~~~~C~C~~g~~g 31 (36)
T cd00053 2 CAASNPCSNGGTCVNTPGSYRCVCPPGYTG 31 (36)
T ss_pred CCCCCCCCCCCEEecCCCCeEeECCCCCcc
Confidence 4434678888999753 4578999999964
No 26
>PF01102 Glycophorin_A: Glycophorin A; InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=80.83 E-value=0.37 Score=41.66 Aligned_cols=32 Identities=19% Similarity=0.325 Sum_probs=13.9
Q ss_pred EEEEEehHHHHHHHHHHHHheeeeeeeccCccc
Q 042843 445 VIGGVVGSVAVVALIGLIMLVYLGRRKTATVTT 477 (484)
Q Consensus 445 ~i~~~v~~~~~~~~~~~~~~~~~~~r~~~~~~~ 477 (484)
+++++++++ +++++++++++|++||+|||...
T Consensus 66 i~~Ii~gv~-aGvIg~Illi~y~irR~~Kk~~~ 97 (122)
T PF01102_consen 66 IIGIIFGVM-AGVIGIILLISYCIRRLRKKSSS 97 (122)
T ss_dssp HHHHHHHHH-HHHHHHHHHHHHHHHHHS-----
T ss_pred eeehhHHHH-HHHHHHHHHHHHHHHHHhccCCC
Confidence 333444444 33334444566777777666543
No 27
>smart00179 EGF_CA Calcium-binding EGF-like domain.
Probab=75.76 E-value=2.2 Score=28.08 Aligned_cols=30 Identities=23% Similarity=0.573 Sum_probs=22.3
Q ss_pred CCCcccccCCCCccccCC-CCccccccCCCc
Q 042843 297 QQCEVYALCGQFSTCNQQ-TERFCSCLKGFQ 326 (484)
Q Consensus 297 d~C~~~~~CG~~giC~~~-~~~~C~C~~GF~ 326 (484)
++|.....|...+.|... ....|.|++||.
T Consensus 3 ~~C~~~~~C~~~~~C~~~~g~~~C~C~~g~~ 33 (39)
T smart00179 3 DECASGNPCQNGGTCVNTVGSYRCECPPGYT 33 (39)
T ss_pred ccCcCCCCcCCCCEeECCCCCeEeECCCCCc
Confidence 567655678888899743 345799999986
No 28
>KOG4649 consensus PQQ (pyrrolo-quinoline quinone) repeat protein [Secondary metabolites biosynthesis, transport and catabolism]
Probab=75.65 E-value=10 Score=37.12 Aligned_cols=46 Identities=30% Similarity=0.506 Sum_probs=34.9
Q ss_pred CCcEEEEcCCCCCCCCC---CceEEEEe--CCeEEEEcCCCceEEEeccCC
Q 042843 77 ERTIVWVANREQPVSDR---FSSVLRIS--DGNLVLFNESQLPIWSTNLTA 122 (484)
Q Consensus 77 ~~tvVW~ANr~~Pv~~~---~~~~l~l~--~G~LvL~~~~~~~vWst~~~~ 122 (484)
+.+..|.|.|..|+-.+ -+..+.++ ||+|.-.|+.|+.||.-.+.+
T Consensus 168 ~~~~~w~~~~~~PiF~splcv~~sv~i~~VdG~l~~f~~sG~qvwr~~t~G 218 (354)
T KOG4649|consen 168 SSTEFWAATRFGPIFASPLCVGSSVIITTVDGVLTSFDESGRQVWRPATKG 218 (354)
T ss_pred CcceehhhhcCCccccCceeccceEEEEEeccEEEEEcCCCcEEEeecCCC
Confidence 45889999999998743 12345565 999999999999999765543
No 29
>PF01299 Lamp: Lysosome-associated membrane glycoprotein (Lamp); InterPro: IPR002000 Lysosome-associated membrane glycoproteins (lamp) [] are integral membrane proteins, specific to lysosomes, and whose exact biological function is not yet clear. Structurally, the lamp proteins consist of two internally homologous lysosome-luminal domains separated by a proline-rich hinge region; at the C-terminal extremity there is a transmembrane region (TM) followed by a very short cytoplasmic tail (C). In each of the duplicated domains, there are two conserved disulphide bonds. This structure is schematically represented in the figure below. +-----+ +-----+ +-----+ +-----+ | | | | | | | | xCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxxxCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxx +--------------------------++Hinge++--------------------------++TM++C+ In mammals, there are two closely related types of lamp: lamp-1 and lamp-2, which form major components of the lysosome membrane. In chicken lamp-1 is known as LEP100. Also included in this entry is the macrophage protein CD68 (or macrosialin) [] is a heavily glycosylated integral membrane protein whose structure consists of a mucin-like domain followed by a proline-rich hinge; a single lamp-like domain; a transmembrane region and a short cytoplasmic tail. Similar to CD68, mammalian lamp-3, which is expressed in lymphoid organs, dendritic cells and in lung, contains all the C-terminal regions but lacks the N-terminal lamp-like region []. In a lamp-family protein from nematodes [] only the part C-terminal to the hinge is conserved. ; GO: 0016020 membrane
Probab=74.81 E-value=1.8 Score=43.71 Aligned_cols=31 Identities=16% Similarity=0.285 Sum_probs=18.2
Q ss_pred eEEEEEehHHHHHHHHHHHHheeeeeeeccCc
Q 042843 444 VVIGGVVGSVAVVALIGLIMLVYLGRRKTATV 475 (484)
Q Consensus 444 ~~i~~~v~~~~~~~~~~~~~~~~~~~r~~~~~ 475 (484)
.+|-++||++++++ +++++++|++.|||++.
T Consensus 271 ~~vPIaVG~~La~l-vlivLiaYli~Rrr~~~ 301 (306)
T PF01299_consen 271 DLVPIAVGAALAGL-VLIVLIAYLIGRRRSRA 301 (306)
T ss_pred chHHHHHHHHHHHH-HHHHHHhheeEeccccc
Confidence 44445666665544 43455677777776554
No 30
>cd01099 PAN_AP_HGF Subfamily of PAN/APPLE-like domains; present in N-terminal (N) domains of plasminogen/hepatocyte growth factor proteins, and various proteins found in Bilateria, such as leech anti-platelet proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=74.12 E-value=5.8 Score=31.43 Aligned_cols=33 Identities=21% Similarity=0.708 Sum_probs=26.6
Q ss_pred cCCHHHHHHHHhh--cCCeEEEEeC--CCceEEeccc
Q 042843 380 VGGIRECETHCMN--NCSCTAYAYK--DNACSIWVGS 412 (484)
Q Consensus 380 ~~~~~~C~~~CL~--nCSC~Ay~y~--~~~C~~w~~~ 412 (484)
..++++|.+.|++ +=.|.++.|. ...|.+-..+
T Consensus 24 ~~s~~~C~~~C~~~~~f~CrSf~y~~~~~~C~L~~~~ 60 (80)
T cd01099 24 VASLEECLRKCLEETEFTCRSFNYNYKSKECILSDED 60 (80)
T ss_pred cCCHHHHHHHhCCCCCceEeEEEEEcCCCEEEEeCCC
Confidence 3789999999999 8899999984 3569985433
No 31
>PF01683 EB: EB module; InterPro: IPR006149 The EB domain has no known function. It is found in several Caenorhabditis sp. and Drosophila sp. proteins. The domain contains 8 conserved cysteines that probably form four disulphide bridges and is found associated with kunitz domains IPR002223 from INTERPRO
Probab=73.81 E-value=2.9 Score=30.14 Aligned_cols=33 Identities=30% Similarity=0.598 Sum_probs=27.3
Q ss_pred ecCCCCCcccccCCCCccccCCCCccccccCCCccC
Q 042843 293 SQPRQQCEVYALCGQFSTCNQQTERFCSCLKGFQQK 328 (484)
Q Consensus 293 ~~p~d~C~~~~~CG~~giC~~~~~~~C~C~~GF~p~ 328 (484)
..|.+.|....-|-.++.|.. ..|.|++||.+.
T Consensus 16 ~~~g~~C~~~~qC~~~s~C~~---g~C~C~~g~~~~ 48 (52)
T PF01683_consen 16 VQPGESCESDEQCIGGSVCVN---GRCQCPPGYVEV 48 (52)
T ss_pred CCCCCCCCCcCCCCCcCEEcC---CEeECCCCCEec
Confidence 456678999999999999953 589999999864
No 32
>PTZ00046 rifin; Provisional
Probab=72.21 E-value=0.81 Score=46.65 Aligned_cols=18 Identities=22% Similarity=0.386 Sum_probs=11.7
Q ss_pred HHHheeeeeeeccCcccc
Q 042843 461 LIMLVYLGRRKTATVTTK 478 (484)
Q Consensus 461 ~~~~~~~~~r~~~~~~~~ 478 (484)
++++.|++.|+|||++.|
T Consensus 330 IMvIIYLILRYRRKKKMk 347 (358)
T PTZ00046 330 IMVIIYLILRYRRKKKMK 347 (358)
T ss_pred HHHHHHHHHHhhhcchhH
Confidence 445666677777777665
No 33
>TIGR01477 RIFIN variant surface antigen, rifin family. This model represents the rifin branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of rifin sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 20 bits.
Probab=72.19 E-value=0.81 Score=46.53 Aligned_cols=18 Identities=22% Similarity=0.386 Sum_probs=11.6
Q ss_pred HHHheeeeeeeccCcccc
Q 042843 461 LIMLVYLGRRKTATVTTK 478 (484)
Q Consensus 461 ~~~~~~~~~r~~~~~~~~ 478 (484)
++++.|++.|+|||++.|
T Consensus 325 IMvIIYLILRYRRKKKMk 342 (353)
T TIGR01477 325 IMVIIYLILRYRRKKKMK 342 (353)
T ss_pred HHHHHHHHHHhhhcchhH
Confidence 445666677777776665
No 34
>PF09064 Tme5_EGF_like: Thrombomodulin like fifth domain, EGF-like; InterPro: IPR015149 This domain adopts a fold similar to other EGF domains, with a flat major and a twisted minor beta sheet. Disulphide pairing, however, is not of the usual 1-3, 2-4, 5-6 type; rather 1-2, 3-4, 5-6 pairing is found. Its extended major sheet (strands beta-2 and beta-3 and the connecting loop) projects into thrombin's active site groove. This domain is required for interaction of thrombomodulin with thrombin, and subsequent activation of protein-C []. ; GO: 0004888 transmembrane signaling receptor activity, 0016021 integral to membrane
Probab=71.54 E-value=2.2 Score=28.00 Aligned_cols=18 Identities=22% Similarity=0.720 Sum_probs=13.7
Q ss_pred cccCCCCccccccCCCcc
Q 042843 310 TCNQQTERFCSCLKGFQQ 327 (484)
Q Consensus 310 iC~~~~~~~C~C~~GF~p 327 (484)
.|+.+...+|.||.||..
T Consensus 11 ~CDpn~~~~C~CPeGyIl 28 (34)
T PF09064_consen 11 DCDPNSPGQCFCPEGYIL 28 (34)
T ss_pred ccCCCCCCceeCCCceEe
Confidence 455555668999999975
No 35
>PF12661 hEGF: Human growth factor-like EGF; PDB: 2YGQ_A 2E26_A 3A7Q_A 2YGP_A 2YGO_A 1HRE_A 1HAE_A 1HAF_A 1HRF_A.
Probab=70.30 E-value=1.2 Score=22.87 Aligned_cols=9 Identities=33% Similarity=1.047 Sum_probs=6.7
Q ss_pred cccccCCCc
Q 042843 318 FCSCLKGFQ 326 (484)
Q Consensus 318 ~C~C~~GF~ 326 (484)
.|.|++||.
T Consensus 1 ~C~C~~G~~ 9 (13)
T PF12661_consen 1 TCQCPPGWT 9 (13)
T ss_dssp EEEE-TTEE
T ss_pred CccCcCCCc
Confidence 489999986
No 36
>cd00054 EGF_CA Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.
Probab=69.65 E-value=3.9 Score=26.42 Aligned_cols=30 Identities=27% Similarity=0.595 Sum_probs=21.6
Q ss_pred CCCcccccCCCCccccCC-CCccccccCCCc
Q 042843 297 QQCEVYALCGQFSTCNQQ-TERFCSCLKGFQ 326 (484)
Q Consensus 297 d~C~~~~~CG~~giC~~~-~~~~C~C~~GF~ 326 (484)
++|.....|...+.|... ....|.|++||.
T Consensus 3 ~~C~~~~~C~~~~~C~~~~~~~~C~C~~g~~ 33 (38)
T cd00054 3 DECASGNPCQNGGTCVNTVGSYRCSCPPGYT 33 (38)
T ss_pred ccCCCCCCcCCCCEeECCCCCeEeECCCCCc
Confidence 567654568878889743 345799999985
No 37
>PF07974 EGF_2: EGF-like domain; InterPro: IPR013111 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length. This entry contains EGF domains found in a variety of extracellular and membrane proteins
Probab=69.10 E-value=3.8 Score=26.67 Aligned_cols=23 Identities=26% Similarity=0.697 Sum_probs=18.5
Q ss_pred ccCCCCccccCCCCccccccCCCc
Q 042843 303 ALCGQFSTCNQQTERFCSCLKGFQ 326 (484)
Q Consensus 303 ~~CG~~giC~~~~~~~C~C~~GF~ 326 (484)
..|...|.|... ...|.|.+||.
T Consensus 6 ~~C~~~G~C~~~-~g~C~C~~g~~ 28 (32)
T PF07974_consen 6 NICSGHGTCVSP-CGRCVCDSGYT 28 (32)
T ss_pred CccCCCCEEeCC-CCEEECCCCCc
Confidence 479999999743 46899999986
No 38
>PHA03265 envelope glycoprotein D; Provisional
Probab=67.39 E-value=2.3 Score=42.86 Aligned_cols=39 Identities=15% Similarity=0.256 Sum_probs=19.6
Q ss_pred ceEEEEEehHHHHHHHHHHHHheeeeeeeccCcccccccC
Q 042843 443 GVVIGGVVGSVAVVALIGLIMLVYLGRRKTATVTTKTVEG 482 (484)
Q Consensus 443 ~~~i~~~v~~~~~~~~~~~~~~~~~~~r~~~~~~~~~~~~ 482 (484)
...++++|+..|+.++++.+ ++|.+||||+..++.+..|
T Consensus 347 ~~~~g~~ig~~i~glv~vg~-il~~~~rr~k~~~k~~~~~ 385 (402)
T PHA03265 347 STFVGISVGLGIAGLVLVGV-ILYVCLRRKKELKKSAQNG 385 (402)
T ss_pred CcccceEEccchhhhhhhhH-HHHHHhhhhhhhhhhhhcC
Confidence 35566777665554433232 3444566655444435544
No 39
>PF00008 EGF: EGF-like domain This is a sub-family of the Pfam entry This is a sub-family of the Pfam entry; InterPro: IPR006209 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length.; GO: 0005515 protein binding; PDB: 1WHE_A 1CCF_A 1APO_A 1WHF_A 2VJ3_A 1TOZ_A 4D90_B 3CFW_A 1EDM_B 1IXA_A ....
Probab=64.05 E-value=2.5 Score=27.34 Aligned_cols=23 Identities=26% Similarity=0.610 Sum_probs=17.9
Q ss_pred cCCCCccccC-C-CCccccccCCCc
Q 042843 304 LCGQFSTCNQ-Q-TERFCSCLKGFQ 326 (484)
Q Consensus 304 ~CG~~giC~~-~-~~~~C~C~~GF~ 326 (484)
.|...|.|.. . ....|.|++||.
T Consensus 5 ~C~n~g~C~~~~~~~y~C~C~~G~~ 29 (32)
T PF00008_consen 5 PCQNGGTCIDLPGGGYTCECPPGYT 29 (32)
T ss_dssp SSTTTEEEEEESTSEEEEEEBTTEE
T ss_pred cCCCCeEEEeCCCCCEEeECCCCCc
Confidence 6888888873 2 457899999986
No 40
>PF12662 cEGF: Complement Clr-like EGF-like
Probab=60.75 E-value=3.7 Score=24.93 Aligned_cols=11 Identities=45% Similarity=0.999 Sum_probs=9.4
Q ss_pred cccccCCCccC
Q 042843 318 FCSCLKGFQQK 328 (484)
Q Consensus 318 ~C~C~~GF~p~ 328 (484)
.|+|++||+..
T Consensus 3 ~C~C~~Gy~l~ 13 (24)
T PF12662_consen 3 TCSCPPGYQLS 13 (24)
T ss_pred EeeCCCCCcCC
Confidence 69999999864
No 41
>PF02439 Adeno_E3_CR2: Adenovirus E3 region protein CR2; InterPro: IPR003470 Early region 3 (E3) of human adenoviruses (Ads) codes for proteins that appear to control viral interactions with the host []. This region called CR1 (conserved region 1) [] is found three times in Human adenovirus 19 (a subgroup D adenovirus) 49 kDa protein in the E3 region. CR1 is also found in the 20.1 Kd protein of subgroup B adenoviruses. The function of this 80 amino acid region is unknown. This region is probably a divergent immunoglobulin domain.
Probab=58.26 E-value=2.1 Score=28.93 Aligned_cols=8 Identities=38% Similarity=0.393 Sum_probs=3.0
Q ss_pred EEEEehHH
Q 042843 446 IGGVVGSV 453 (484)
Q Consensus 446 i~~~v~~~ 453 (484)
|+++++++
T Consensus 6 IaIIv~V~ 13 (38)
T PF02439_consen 6 IAIIVAVV 13 (38)
T ss_pred hhHHHHHH
Confidence 33334333
No 42
>PF12947 EGF_3: EGF domain; InterPro: IPR024731 This entry represents an EGF domain found in the the C terminus of malarial parasite merozoite surface protein 1 [], as well as other proteins.; PDB: 2NPR_A 1N1I_C 1B9W_A 1YO8_A 2RHP_A.
Probab=57.85 E-value=3.3 Score=27.70 Aligned_cols=24 Identities=25% Similarity=0.688 Sum_probs=16.9
Q ss_pred ccCCCCccccCC-CCccccccCCCc
Q 042843 303 ALCGQFSTCNQQ-TERFCSCLKGFQ 326 (484)
Q Consensus 303 ~~CG~~giC~~~-~~~~C~C~~GF~ 326 (484)
+-|.++..|... .+..|.|.+||.
T Consensus 6 ~~C~~nA~C~~~~~~~~C~C~~Gy~ 30 (36)
T PF12947_consen 6 GGCHPNATCTNTGGSYTCTCKPGYE 30 (36)
T ss_dssp GGS-TTCEEEE-TTSEEEEE-CEEE
T ss_pred CCCCCCcEeecCCCCEEeECCCCCc
Confidence 568889999743 457899999996
No 43
>smart00181 EGF Epidermal growth factor-like domain.
Probab=52.15 E-value=12 Score=24.07 Aligned_cols=24 Identities=29% Similarity=0.634 Sum_probs=17.2
Q ss_pred ccCCCCccccCC-CCccccccCCCcc
Q 042843 303 ALCGQFSTCNQQ-TERFCSCLKGFQQ 327 (484)
Q Consensus 303 ~~CG~~giC~~~-~~~~C~C~~GF~p 327 (484)
..|... .|... ....|.|++||.-
T Consensus 6 ~~C~~~-~C~~~~~~~~C~C~~g~~g 30 (35)
T smart00181 6 GPCSNG-TCINTPGSYTCSCPPGYTG 30 (35)
T ss_pred CCCCCC-EEECCCCCeEeECCCCCcc
Confidence 456666 78643 4578999999964
No 44
>PF06024 DUF912: Nucleopolyhedrovirus protein of unknown function (DUF912); InterPro: IPR009261 This entry is represented by Autographa californica nuclear polyhedrosis virus (AcMNPV), Orf78; it is a family of uncharacterised viral proteins.
Probab=52.12 E-value=9.3 Score=31.91 Aligned_cols=12 Identities=8% Similarity=0.088 Sum_probs=5.3
Q ss_pred eeeeeeeccCcc
Q 042843 465 VYLGRRKTATVT 476 (484)
Q Consensus 465 ~~~~~r~~~~~~ 476 (484)
.+++.|.|+++.
T Consensus 83 YFVILRer~~~~ 94 (101)
T PF06024_consen 83 YFVILRERQKSI 94 (101)
T ss_pred EEEEEecccccc
Confidence 333455544443
No 45
>PF12877 DUF3827: Domain of unknown function (DUF3827); InterPro: IPR024606 The function of the proteins in this entry is not currently known, but one of the human proteins (Q9HCM3 from SWISSPROT) has been implicated in pilocytic astrocytomas [, , ]. In the majority of cases of pilocytic astrocytomas a tandem duplication produces an in-frame fusion of the gene encoding this protein and the BRAF oncogene. The resulting fusion protein has constitutive BRAF kinase activity and is capable of transforming cells.
Probab=51.10 E-value=10 Score=41.44 Aligned_cols=15 Identities=40% Similarity=0.386 Sum_probs=6.7
Q ss_pred ceEEEEEehHHHHHH
Q 042843 443 GVVIGGVVGSVAVVA 457 (484)
Q Consensus 443 ~~~i~~~v~~~~~~~ 457 (484)
-|||+.|++++++++
T Consensus 269 lWII~gVlvPv~vV~ 283 (684)
T PF12877_consen 269 LWIIAGVLVPVLVVL 283 (684)
T ss_pred eEEEehHhHHHHHHH
Confidence 344444444544433
No 46
>PF03302 VSP: Giardia variant-specific surface protein; InterPro: IPR005127 During infection, the intestinal protozoan parasite Giardia lamblia virus undergoes continuous antigenic variation which is determined by diversification of the parasite's major surface antigen, named VSP (variant surface protein).
Probab=49.90 E-value=18 Score=37.94 Aligned_cols=30 Identities=23% Similarity=0.194 Sum_probs=16.7
Q ss_pred ceEEEEEehHHHHHHHHHHHHheeeeeeec
Q 042843 443 GVVIGGVVGSVAVVALIGLIMLVYLGRRKT 472 (484)
Q Consensus 443 ~~~i~~~v~~~~~~~~~~~~~~~~~~~r~~ 472 (484)
..|.+|.|+++|++-.|+.++++|++.|.|
T Consensus 367 gaIaGIsvavvvvVgglvGfLcWwf~crgk 396 (397)
T PF03302_consen 367 GAIAGISVAVVVVVGGLVGFLCWWFICRGK 396 (397)
T ss_pred cceeeeeehhHHHHHHHHHHHhhheeeccc
Confidence 466667776554433233455666666654
No 47
>PF14610 DUF4448: Protein of unknown function (DUF4448)
Probab=49.50 E-value=17 Score=33.93 Aligned_cols=22 Identities=32% Similarity=0.348 Sum_probs=9.1
Q ss_pred EEehHHHHHHHHHHHHheeeeee
Q 042843 448 GVVGSVAVVALIGLIMLVYLGRR 470 (484)
Q Consensus 448 ~~v~~~~~~~~~~~~~~~~~~~r 470 (484)
+|++++++++++ ++++++++|+
T Consensus 161 aI~lPvvv~~~~-~~~~~~~~~~ 182 (189)
T PF14610_consen 161 AIALPVVVVVLA-LIMYGFFFWN 182 (189)
T ss_pred EEEccHHHHHHH-HHHHhhheee
Confidence 444455444433 2333444443
No 48
>PF13360 PQQ_2: PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=48.39 E-value=1.6e+02 Score=27.36 Aligned_cols=77 Identities=22% Similarity=0.335 Sum_probs=43.1
Q ss_pred CCCcEEEEcCCCCCCCCC---C-ceEEEEe-CCeEEEEc-CCCceEEEe-ccCCC-CC--C--------ceEEEEecCCC
Q 042843 76 SERTIVWVANREQPVSDR---F-SSVLRIS-DGNLVLFN-ESQLPIWST-NLTAT-SR--R--------SVEAVLLDEGN 137 (484)
Q Consensus 76 ~~~tvVW~ANr~~Pv~~~---~-~~~l~l~-~G~LvL~~-~~~~~vWst-~~~~~-~~--~--------~~~a~LlDsGN 137 (484)
....++|...-+.++... . +..+..+ +|.|..+| .+|.++|.. ..... .. . ........+|.
T Consensus 54 ~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~g~ 133 (238)
T PF13360_consen 54 KTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYALDAKTGKVLWSIYLTSSPPAGVRSSSSPAVDGDRLYVGTSSGK 133 (238)
T ss_dssp TTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEEEETTTSCEEEEEEE-SSCTCSTB--SEEEEETTEEEEEETCSE
T ss_pred CCCCEEEEeeccccccceeeecccccccccceeeeEecccCCcceeeeeccccccccccccccCceEecCEEEEEeccCc
Confidence 456789988876554321 2 3334455 78888888 788899984 32210 00 0 11122233777
Q ss_pred EEEEeccCCCCcceeee
Q 042843 138 LVLRDLSNNLSKPLWQS 154 (484)
Q Consensus 138 LVL~~~~~~~~~~lWQS 154 (484)
|+..|..+ ++.+|+=
T Consensus 134 l~~~d~~t--G~~~w~~ 148 (238)
T PF13360_consen 134 LVALDPKT--GKLLWKY 148 (238)
T ss_dssp EEEEETTT--TEEEEEE
T ss_pred EEEEecCC--CcEEEEe
Confidence 77777542 4677765
No 49
>KOG0291 consensus WD40-repeat-containing subunit of the 18S rRNA processing complex [RNA processing and modification]
Probab=47.38 E-value=4.8e+02 Score=29.71 Aligned_cols=85 Identities=20% Similarity=0.351 Sum_probs=54.5
Q ss_pred ceEEEEe-CCeEEEEcC-CCce-EEEeccC-------CCCCCceEEEEecCCCEEEEeccCCCCcceeeecccCceeccC
Q 042843 95 SSVLRIS-DGNLVLFNE-SQLP-IWSTNLT-------ATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQSFDHPAHTWIP 164 (484)
Q Consensus 95 ~~~l~l~-~G~LvL~~~-~~~~-vWst~~~-------~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQSFd~PTDTlLp 164 (484)
-..+..+ ||.++.+.. +|.+ ||.+... ..+.+++..+..-+||.+|..+- -
T Consensus 353 i~~l~YSpDgq~iaTG~eDgKVKvWn~~SgfC~vTFteHts~Vt~v~f~~~g~~llssSL-------------------D 413 (893)
T KOG0291|consen 353 ITSLAYSPDGQLIATGAEDGKVKVWNTQSGFCFVTFTEHTSGVTAVQFTARGNVLLSSSL-------------------D 413 (893)
T ss_pred eeeEEECCCCcEEEeccCCCcEEEEeccCceEEEEeccCCCceEEEEEEecCCEEEEeec-------------------C
Confidence 4568899 999998854 4666 9988641 12236788899999999997542 2
Q ss_pred CceeeeecCCCCceEEEecCCCCCCCCceEE-EEEcCCCC
Q 042843 165 GMKLTFNKRNNVSQLITSWKNKENPAPGLFS-LERAPDGS 203 (484)
Q Consensus 165 gq~l~~n~~~g~~~~L~Sw~s~~dps~G~y~-l~~~~~g~ 203 (484)
|-.=-||.+.+. -.|+-+-|.+-.|+ +..||.|.
T Consensus 414 GtVRAwDlkRYr-----NfRTft~P~p~QfscvavD~sGe 448 (893)
T KOG0291|consen 414 GTVRAWDLKRYR-----NFRTFTSPEPIQFSCVAVDPSGE 448 (893)
T ss_pred CeEEeeeecccc-----eeeeecCCCceeeeEEEEcCCCC
Confidence 222233334331 22333456777776 77888885
No 50
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=46.86 E-value=76 Score=32.84 Aligned_cols=57 Identities=18% Similarity=0.239 Sum_probs=34.8
Q ss_pred eEEEEe--CCeEEEEcC-CCceEEEeccCCCCCC------ceEEEEecCCCEEEEeccCCCCcceeee
Q 042843 96 SVLRIS--DGNLVLFNE-SQLPIWSTNLTATSRR------SVEAVLLDEGNLVLRDLSNNLSKPLWQS 154 (484)
Q Consensus 96 ~~l~l~--~G~LvL~~~-~~~~vWst~~~~~~~~------~~~a~LlDsGNLVL~~~~~~~~~~lWQS 154 (484)
..+.+. +|.|+-+|. +|+++|+.+..+.... .....-..+|.|+-.|.. +++++|+-
T Consensus 121 ~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~ssP~v~~~~v~v~~~~g~l~ald~~--tG~~~W~~ 186 (394)
T PRK11138 121 GKVYIGSEKGQVYALNAEDGEVAWQTKVAGEALSRPVVSDGLVLVHTSNGMLQALNES--DGAVKWTV 186 (394)
T ss_pred CEEEEEcCCCEEEEEECCCCCCcccccCCCceecCCEEECCEEEEECCCCEEEEEEcc--CCCEeeee
Confidence 455555 788888885 6899998875432101 111222346667777764 35788975
No 51
>PF08374 Protocadherin: Protocadherin; InterPro: IPR013585 The structure of protocadherins is similar to that of classic cadherins (IPR002126 from INTERPRO), but they also have some unique features associated with the cytoplasmic domains. They are expressed in a variety of organisms and are found in high concentrations in the brain where they seem to be localised mainly at cell-cell contact sites. Their expression seems to be developmentally regulated [].
Probab=45.33 E-value=13 Score=35.18 Aligned_cols=11 Identities=27% Similarity=0.410 Sum_probs=4.9
Q ss_pred ceEEEEEehHH
Q 042843 443 GVVIGGVVGSV 453 (484)
Q Consensus 443 ~~~i~~~v~~~ 453 (484)
+++|++|.|++
T Consensus 38 ~I~iaiVAG~~ 48 (221)
T PF08374_consen 38 KIMIAIVAGIM 48 (221)
T ss_pred eeeeeeecchh
Confidence 44554444443
No 52
>TIGR01478 STEVOR variant surface antigen, stevor family. This model represents the stevor branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of stevor sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 8 bits.
Probab=44.62 E-value=7.1 Score=38.51 Aligned_cols=7 Identities=29% Similarity=1.075 Sum_probs=4.4
Q ss_pred ccccccC
Q 042843 317 RFCSCLK 323 (484)
Q Consensus 317 ~~C~C~~ 323 (484)
..|+|-+
T Consensus 143 s~cectd 149 (295)
T TIGR01478 143 KSCECTN 149 (295)
T ss_pred Cceeeec
Confidence 5677753
No 53
>PTZ00370 STEVOR; Provisional
Probab=40.39 E-value=8 Score=38.22 Aligned_cols=7 Identities=29% Similarity=0.952 Sum_probs=4.5
Q ss_pred ccccccC
Q 042843 317 RFCSCLK 323 (484)
Q Consensus 317 ~~C~C~~ 323 (484)
..|+|-+
T Consensus 143 s~cectd 149 (296)
T PTZ00370 143 STCECTD 149 (296)
T ss_pred Cceeeee
Confidence 4688753
No 54
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=39.69 E-value=1.1e+02 Score=31.29 Aligned_cols=55 Identities=22% Similarity=0.427 Sum_probs=28.7
Q ss_pred EEEEe--CCeEEEEc-CCCceEEEeccCCCCCC------ceEEEEecCCCEEEEeccCCCCcceee
Q 042843 97 VLRIS--DGNLVLFN-ESQLPIWSTNLTATSRR------SVEAVLLDEGNLVLRDLSNNLSKPLWQ 153 (484)
Q Consensus 97 ~l~l~--~G~LvL~~-~~~~~vWst~~~~~~~~------~~~a~LlDsGNLVL~~~~~~~~~~lWQ 153 (484)
.+.+. +|.|.-+| .+|+++|+.+....... .....-..+|+|+-.|.. +++.+|+
T Consensus 67 ~v~v~~~~g~v~a~d~~tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~ald~~--tG~~~W~ 130 (377)
T TIGR03300 67 KVYAADADGTVVALDAETGKRLWRVDLDERLSGGVGADGGLVFVGTEKGEVIALDAE--DGKELWR 130 (377)
T ss_pred EEEEECCCCeEEEEEccCCcEeeeecCCCCcccceEEcCCEEEEEcCCCEEEEEECC--CCcEeee
Confidence 44444 67777777 56778887654322100 011111245666666643 2467786
No 55
>PF02480 Herpes_gE: Alphaherpesvirus glycoprotein E; InterPro: IPR003404 Glycoprotein E (gE) of Alphaherpesvirus forms a complex with glycoprotein I (gI), functioning as an immunoglobulin G (IgG) Fc binding protein. gE is involved in virus spread but is not essential for propagation [].; GO: 0016020 membrane; PDB: 2GJ7_F 2GIY_B.
Probab=39.52 E-value=9.8 Score=40.44 Aligned_cols=13 Identities=23% Similarity=0.593 Sum_probs=8.0
Q ss_pred CCCcccC-CCCCCC
Q 042843 339 SGGCVRK-TPLQCE 351 (484)
Q Consensus 339 s~GC~r~-~~l~C~ 351 (484)
..+|.+. -+..|.
T Consensus 238 y~~C~~~~~~~~C~ 251 (439)
T PF02480_consen 238 YANCSPSGWPRRCP 251 (439)
T ss_dssp EEEEBTTC-TTTTE
T ss_pred hcCCCCCCCcCCCC
Confidence 3588876 344784
No 56
>PF06365 CD34_antigen: CD34/Podocalyxin family; InterPro: IPR013836 This family consists of several mammalian CD34 antigen proteins. The CD34 antigen is a human leukocyte membrane protein expressed specifically by lymphohematopoietic progenitor cells. CD34 is a phosphoprotein. Activation of protein kinase C (PKC) has been found to enhance CD34 phosphorylation [, ]. This family contains several eukaryotic podocalyxin proteins. Podocalyxin is a major membrane protein of the glomerular epithelium and is thought to be involved in maintenance of the architecture of the foot processes and filtration slits characteristic of this unique epithelium by virtue of its high negative charge. Podocalyxin functions as an anti-adhesin that maintains an open filtration pathway between neighbouring foot processes in the glomerular epithelium by charge repulsion [].
Probab=39.31 E-value=17 Score=34.24 Aligned_cols=18 Identities=17% Similarity=0.519 Sum_probs=9.9
Q ss_pred hcCCeEEEEeC-CCceEEe
Q 042843 392 NNCSCTAYAYK-DNACSIW 409 (484)
Q Consensus 392 ~nCSC~Ay~y~-~~~C~~w 409 (484)
.+|+..-+.-. +..|.++
T Consensus 36 ~~C~l~Laq~~~~~q~Lll 54 (202)
T PF06365_consen 36 DDCSLSLAQSEENQQCLLL 54 (202)
T ss_pred CCcEEEEecCCCCcceEEE
Confidence 46666655432 3457765
No 57
>PF01436 NHL: NHL repeat; InterPro: IPR001258 The NHL repeat, named after NCL-1, HT2A and Lin-41, is found largely in a large number of eukaryotic and prokaryotic proteins. For example, the repeat is found in a variety of enzymes of the copper type II, ascorbate-dependent monooxygenase family which catalyse the C terminus alpha-amidation of biological peptides []. In many it occurs in tandem arrays, for example in the ringfinger beta-box, coiled-coil (RBCC) eukaryotic growth regulators []. The 'Brain Tumor' protein (Brat) is one such growth regulator that contains a 6-bladed NHL-repeat beta-propeller [, ]. The NHL repeats are also found in serine/threonine protein kinase (STPK) in diverse range of pathogenic bacteria. These STPK are transmembrane receptors with a intracellular N-terminal kinase domain and extracellular C-terminal sensor domain. In the STPK, PknD, from Mycobacterium tuberculosis, the sensor domain forms a rigid, six-bladed b-propeller composed of NHL repeats with a flexible tether to the transmembrane domain.; GO: 0005515 protein binding; PDB: 3FVZ_A 3FW0_A 1RWL_A 1RWI_A 1Q7F_A.
Probab=38.43 E-value=56 Score=20.16 Aligned_cols=20 Identities=15% Similarity=0.317 Sum_probs=13.9
Q ss_pred EEEEe-CCeEEEEcCCCceEE
Q 042843 97 VLRIS-DGNLVLFNESQLPIW 116 (484)
Q Consensus 97 ~l~l~-~G~LvL~~~~~~~vW 116 (484)
-+.++ +|++++.|.....||
T Consensus 6 gvav~~~g~i~VaD~~n~rV~ 26 (28)
T PF01436_consen 6 GVAVDSDGNIYVADSGNHRVQ 26 (28)
T ss_dssp EEEEETTSEEEEEECCCTEEE
T ss_pred EEEEeCCCCEEEEECCCCEEE
Confidence 36777 888888887655554
No 58
>cd05845 Ig2_L1-CAM_like Second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM) and similar proteins. Ig2_L1-CAM_like: domain similar to the second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM). L1 belongs to the L1 subfamily of cell adhesion molecules (CAMs) and is comprised of an extracellular region having six Ig-like domains, five fibronectin type III domains, a transmembrane region and an intracellular domain. L1 is primarily expressed in the nervous system and is involved in its development and function. L1 is associated with an X-linked recessive disorder, X-linked hydrocephalus, MASA syndrome, or spastic paraplegia type 1, that involves abnormalities of axonal growth.
Probab=38.21 E-value=46 Score=27.41 Aligned_cols=31 Identities=16% Similarity=0.324 Sum_probs=21.2
Q ss_pred CCCcEEEEcCCCCCCCCCCceEEEEe-CCeEEEE
Q 042843 76 SERTIVWVANREQPVSDRFSSVLRIS-DGNLVLF 108 (484)
Q Consensus 76 ~~~tvVW~ANr~~Pv~~~~~~~l~l~-~G~LvL~ 108 (484)
|..++.|+-+....+. ...++.++ +|+|.+.
T Consensus 32 P~P~i~W~~~~~~~i~--~~~Ri~~~~~GnL~fs 63 (95)
T cd05845 32 VPLRIYWMNSDLLHIT--QDERVSMGQNGNLYFA 63 (95)
T ss_pred CCCEEEEECCCCcccc--ccccEEECCCceEEEE
Confidence 5667888844434454 35678888 8999874
No 59
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=37.59 E-value=4.7e+02 Score=26.86 Aligned_cols=75 Identities=20% Similarity=0.344 Sum_probs=43.1
Q ss_pred CCcEEEEcCCCCCCCCC---CceEEEEe--CCeEEEEcC-CCceEEEeccCCC-------CCC----ceEEEEecCCCEE
Q 042843 77 ERTIVWVANREQPVSDR---FSSVLRIS--DGNLVLFNE-SQLPIWSTNLTAT-------SRR----SVEAVLLDEGNLV 139 (484)
Q Consensus 77 ~~tvVW~ANr~~Pv~~~---~~~~l~l~--~G~LvL~~~-~~~~vWst~~~~~-------~~~----~~~a~LlDsGNLV 139 (484)
...++|..+...++..+ ....+.+. +|.|+-+|. +|+++|+.+.... ... .....-..+|.++
T Consensus 139 tG~~~W~~~~~~~~~ssP~v~~~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~~~~~~~sP~v~~~~v~~~~~~g~v~ 218 (394)
T PRK11138 139 DGEVAWQTKVAGEALSRPVVSDGLVLVHTSNGMLQALNESDGAVKWTVNLDVPSLTLRGESAPATAFGGAIVGGDNGRVS 218 (394)
T ss_pred CCCCcccccCCCceecCCEEECCEEEEECCCCEEEEEEccCCCEeeeecCCCCcccccCCCCCEEECCEEEEEcCCCEEE
Confidence 45678988765433210 13344444 788888885 6889998764311 000 0111224566777
Q ss_pred EEeccCCCCcceee
Q 042843 140 LRDLSNNLSKPLWQ 153 (484)
Q Consensus 140 L~~~~~~~~~~lWQ 153 (484)
-.+..+ ++.+|+
T Consensus 219 a~d~~~--G~~~W~ 230 (394)
T PRK11138 219 AVLMEQ--GQLIWQ 230 (394)
T ss_pred EEEccC--Chhhhe
Confidence 666543 578997
No 60
>PF06697 DUF1191: Protein of unknown function (DUF1191); InterPro: IPR010605 This family contains hypothetical plant proteins of unknown function.
Probab=36.70 E-value=15 Score=36.47 Aligned_cols=8 Identities=13% Similarity=0.127 Sum_probs=3.9
Q ss_pred EEEEEecc
Q 042843 425 IIYIKLAA 432 (484)
Q Consensus 425 ~~yikv~~ 432 (484)
.+-++...
T Consensus 187 slVV~~~~ 194 (278)
T PF06697_consen 187 SLVVPSPA 194 (278)
T ss_pred EEEEcCCC
Confidence 44555443
No 61
>PF00954 S_locus_glycop: S-locus glycoprotein family; InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=36.08 E-value=84 Score=26.21 Aligned_cols=58 Identities=12% Similarity=0.220 Sum_probs=36.0
Q ss_pred CCeeEEEEeecCCceeEEEEEccCCcEEEEeeCCCCCCCeEEEeecCCCCCcccccCCCC--cccc
Q 042843 249 ENESYFTYNVKDSTYTSRAFMDVSGQDKQMNWLPLPTNSWFLFWSQPRQQCEVYALCGQF--STCN 312 (484)
Q Consensus 249 ~~~~~~~~~~~~~~~~~rl~Ld~dG~l~~y~w~~~~~~~W~~~w~~p~d~C~~~~~CG~~--giC~ 312 (484)
+...+.++.+.....+++++.+.+.+-....|.... .. +..+ ..|..+++|-.+ ..|.
T Consensus 42 ~~s~~~r~~ld~~G~l~~~~w~~~~~~W~~~~~~p~-d~-Cd~y----~~CG~~g~C~~~~~~~C~ 101 (110)
T PF00954_consen 42 NSSVLSRLVLDSDGQLQRYIWNESTQSWSVFWSAPK-DQ-CDVY----GFCGPNGICNSNNSPKCS 101 (110)
T ss_pred CCceEEEEEEeeeeEEEEEEEecCCCcEEEEEEecc-cC-CCCc----cccCCccEeCCCCCCceE
Confidence 334444455555557888888777777776775543 22 2222 689999999654 3575
No 62
>PF14991 MLANA: Protein melan-A; PDB: 2GTZ_F 2GT9_F 3MRO_P 2GUO_C 3MRQ_P 2GTW_C 3L6F_C 3MRP_P.
Probab=34.63 E-value=12 Score=31.63 Aligned_cols=13 Identities=8% Similarity=0.225 Sum_probs=0.0
Q ss_pred HheeeeeeeccCc
Q 042843 463 MLVYLGRRKTATV 475 (484)
Q Consensus 463 ~~~~~~~r~~~~~ 475 (484)
+..|.++||...+
T Consensus 42 iGCWYckRRSGYk 54 (118)
T PF14991_consen 42 IGCWYCKRRSGYK 54 (118)
T ss_dssp -------------
T ss_pred Hhheeeeecchhh
Confidence 3444444444433
No 63
>PF12191 stn_TNFRSF12A: Tumour necrosis factor receptor stn_TNFRSF12A_TNFR domain; InterPro: IPR022316 The tumour necrosis factor (TNF) receptor (TNFR) superfamily comprises more than 20 type-I transmembrane proteins. Family members are defined based on similarity in their extracellular domain - a region that contains many cysteine residues arranged in a specific repetitive pattern []. The cysteines allow formation of an extended rod-like structure, responsible for ligand binding []. Upon receptor activation, different intracellular signalling complexes are assembled for different members of the TNFR superfamily, depending on their intracellular domains and sequences []. Activation of TNFRs can therefore induce a range of disparate effects, including cell proliferation, differentiation, survival, or apoptotic cell death, depending upon the receptor involved []. TNFRs are widely distributed and play important roles in many crucial biological processes, such as lymphoid and neuronal development, innate and adaptive immunity, and maintenance of cellular homeostasis []. Drugs that manipulate their signalling have potential roles in the prevention and treatment of many diseases, such as viral infections, coronary heart disease, transplant rejection, and immune disease []. TNF receptor 12 (also known as TWEAK receptor, and fibroblast growth factor-inducible-14 (Fn14)) has been implicated in endothelial cell growth and migration []. The receptor may also play a role in cell-matrix interactions [].; PDB: 2KN0_A 2RPJ_A 2KMZ_A 2EQP_A.
Probab=32.94 E-value=12 Score=32.45 Aligned_cols=10 Identities=20% Similarity=-0.077 Sum_probs=0.0
Q ss_pred eeeeccCccc
Q 042843 468 GRRKTATVTT 477 (484)
Q Consensus 468 ~~r~~~~~~~ 477 (484)
+||-|++++.
T Consensus 101 ~rrcrrr~~~ 110 (129)
T PF12191_consen 101 WRRCRRREKF 110 (129)
T ss_dssp ----------
T ss_pred HhhhhccccC
Confidence 3544555554
No 64
>PF13360 PQQ_2: PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=32.07 E-value=97 Score=28.89 Aligned_cols=50 Identities=28% Similarity=0.445 Sum_probs=28.9
Q ss_pred CCeEEEEcC-CCceEEEeccCCCCCCceEE-EE---------ecCCCEEEEeccCCCCcceeee
Q 042843 102 DGNLVLFNE-SQLPIWSTNLTATSRRSVEA-VL---------LDEGNLVLRDLSNNLSKPLWQS 154 (484)
Q Consensus 102 ~G~LvL~~~-~~~~vWst~~~~~~~~~~~a-~L---------lDsGNLVL~~~~~~~~~~lWQS 154 (484)
+|.|...|. +|+.+|+.+..... ....+ .+ ..+|+|+..|.. +++++|+-
T Consensus 2 ~g~l~~~d~~tG~~~W~~~~~~~~-~~~~~~~~~~~~~v~~~~~~~~l~~~d~~--tG~~~W~~ 62 (238)
T PF13360_consen 2 DGTLSALDPRTGKELWSYDLGPGI-GGPVATAVPDGGRVYVASGDGNLYALDAK--TGKVLWRF 62 (238)
T ss_dssp TSEEEEEETTTTEEEEEEECSSSC-SSEEETEEEETTEEEEEETTSEEEEEETT--TSEEEEEE
T ss_pred CCEEEEEECCCCCEEEEEECCCCC-CCccceEEEeCCEEEEEcCCCEEEEEECC--CCCEEEEe
Confidence 477777886 78889988642111 11111 12 366666666653 25778875
No 65
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=31.14 E-value=3.9e+02 Score=27.16 Aligned_cols=75 Identities=13% Similarity=0.353 Sum_probs=44.6
Q ss_pred CCCcEEEEcCCCC-----CCCCCCceEEEEe--CCeEEEEcC-CCceEEEeccCCCCC------CceEEEEecCCCEEEE
Q 042843 76 SERTIVWVANREQ-----PVSDRFSSVLRIS--DGNLVLFNE-SQLPIWSTNLTATSR------RSVEAVLLDEGNLVLR 141 (484)
Q Consensus 76 ~~~tvVW~ANr~~-----Pv~~~~~~~l~l~--~G~LvL~~~-~~~~vWst~~~~~~~------~~~~a~LlDsGNLVL~ 141 (484)
....++|.-+-.. |+. ....+.+. +|.|+-+|. +|+++|+....+... ......-..+|.|+..
T Consensus 83 ~tG~~~W~~~~~~~~~~~p~v--~~~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~a~ 160 (377)
T TIGR03300 83 ETGKRLWRVDLDERLSGGVGA--DGGLVFVGTEKGEVIALDAEDGKELWRAKLSSEVLSPPLVANGLVVVRTNDGRLTAL 160 (377)
T ss_pred cCCcEeeeecCCCCcccceEE--cCCEEEEEcCCCEEEEEECCCCcEeeeeccCceeecCCEEECCEEEEECCCCeEEEE
Confidence 3567888755433 333 23455554 899998886 689999876533210 1111222356777777
Q ss_pred eccCCCCcceeee
Q 042843 142 DLSNNLSKPLWQS 154 (484)
Q Consensus 142 ~~~~~~~~~lWQS 154 (484)
|.. +++++|+-
T Consensus 161 d~~--tG~~~W~~ 171 (377)
T TIGR03300 161 DAA--TGERLWTY 171 (377)
T ss_pred EcC--CCceeeEE
Confidence 764 35789984
No 66
>PF15330 SIT: SHP2-interacting transmembrane adaptor protein, SIT
Probab=30.91 E-value=13 Score=31.44 Aligned_cols=9 Identities=22% Similarity=0.154 Sum_probs=4.1
Q ss_pred heeeeeeec
Q 042843 464 LVYLGRRKT 472 (484)
Q Consensus 464 ~~~~~~r~~ 472 (484)
+.++.||.+
T Consensus 17 asl~~wr~~ 25 (107)
T PF15330_consen 17 ASLLAWRMK 25 (107)
T ss_pred HHHHHHHHH
Confidence 344445544
No 67
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=28.62 E-value=18 Score=29.88 Aligned_cols=32 Identities=9% Similarity=-0.111 Sum_probs=17.9
Q ss_pred eEEEEEehHHHHHHHHHHH-HheeeeeeeccCc
Q 042843 444 VVIGGVVGSVAVVALIGLI-MLVYLGRRKTATV 475 (484)
Q Consensus 444 ~~i~~~v~~~~~~~~~~~~-~~~~~~~r~~~~~ 475 (484)
.-.+++++++|.+++++.+ ++++++|..+|+|
T Consensus 63 ls~gaiagi~vg~~~~v~~lv~~l~w~f~~r~k 95 (96)
T PTZ00382 63 LSTGAIAGISVAVVAVVGGLVGFLCWWFVCRGK 95 (96)
T ss_pred cccccEEEEEeehhhHHHHHHHHHhheeEEeec
Confidence 4466777776665544433 4455556555543
No 68
>PF05393 Hum_adeno_E3A: Human adenovirus early E3A glycoprotein; InterPro: IPR008652 This family consists of several early glycoproteins (E3A), from human adenovirus type 2.; GO: 0016021 integral to membrane
Probab=28.53 E-value=14 Score=29.73 Aligned_cols=26 Identities=12% Similarity=0.245 Sum_probs=11.1
Q ss_pred ehHHHHHHHHHHHHheeeeeeeccCcc
Q 042843 450 VGSVAVVALIGLIMLVYLGRRKTATVT 476 (484)
Q Consensus 450 v~~~~~~~~~~~~~~~~~~~r~~~~~~ 476 (484)
.+++.+.|++ ++++.+..|++|+|.+
T Consensus 37 ~lvI~~iFil-~VilwfvCC~kRkrsR 62 (94)
T PF05393_consen 37 FLVICGIFIL-LVILWFVCCKKRKRSR 62 (94)
T ss_pred HHHHHHHHHH-HHHHHHHHHHHhhhcc
Confidence 3344343433 3334444455554443
No 69
>PF14670 FXa_inhibition: Coagulation Factor Xa inhibitory site; PDB: 3Q3K_B 1NFY_B 1LQD_A 1G2L_B 1IQF_L 2UWP_B 2VH6_B 3KQC_L 2P93_L 2BQW_A ....
Probab=28.02 E-value=19 Score=24.05 Aligned_cols=14 Identities=29% Similarity=0.657 Sum_probs=9.9
Q ss_pred CCccccccCCCccC
Q 042843 315 TERFCSCLKGFQQK 328 (484)
Q Consensus 315 ~~~~C~C~~GF~p~ 328 (484)
....|+|++||...
T Consensus 17 g~~~C~C~~Gy~L~ 30 (36)
T PF14670_consen 17 GSYRCSCPPGYKLA 30 (36)
T ss_dssp TSEEEE-STTEEE-
T ss_pred CceEeECCCCCEEC
Confidence 45689999999875
No 70
>PF02009 Rifin_STEVOR: Rifin/stevor family; InterPro: IPR002858 Malaria is still a major cause of mortality in many areas of the world. Plasmodium falciparum causes the most severe human form of the disease and is responsible for most fatalities. Severe cases of malaria can occur when the parasite invades and then proliferates within red blood cell erythrocytes. The parasite produces many variant antigenic proteins, encoded by multigene families, which are present on the surface of the infected erythrocyte and play important roles in virulence. A crucial survival mechanism for the malaria parasite is its ability to evade the immune response by switching these variant surface antigens. The high virulence of P. falciparum relative to other malarial parasites is in large part due to the fact that in this organism many of these surface antigens mediate the binding of infected erythrocytes to the vascular endothelium (cytoadherence) and non-infected erythrocytes (rosetting). This can lead to the accumulation of infected cells in the vasculature of a variety of organs, blocking the blood flow and reducing the oxygen supply. Clinical symptoms of severe infection can include fever, progressive anaemia, multi-organ dysfunction and coma. For more information see []. Several multicopy gene families have been described in Plasmodium falciparum, including the stevor family of subtelomeric open reading frames and the rif interspersed repetitive elements. Both families contain three predicted transmembrane segments. It has been proposed that stevor and rif are members of a larger superfamily that code for variant surface antigens [].
Probab=27.92 E-value=6.5 Score=39.49 Aligned_cols=28 Identities=21% Similarity=0.368 Sum_probs=15.9
Q ss_pred EehHHHHHHHHHHHHheeeeeeeccCcc
Q 042843 449 VVGSVAVVALIGLIMLVYLGRRKTATVT 476 (484)
Q Consensus 449 ~v~~~~~~~~~~~~~~~~~~~r~~~~~~ 476 (484)
+++++|.+++++++.+.+.+||+|+-++
T Consensus 262 iiaIliIVLIMvIIYLILRYRRKKKmkK 289 (299)
T PF02009_consen 262 IIAILIIVLIMVIIYLILRYRRKKKMKK 289 (299)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHhhhhH
Confidence 3333333444666777777888665443
No 71
>PHA02887 EGF-like protein; Provisional
Probab=27.85 E-value=62 Score=27.73 Aligned_cols=30 Identities=30% Similarity=0.775 Sum_probs=22.3
Q ss_pred CCCCcc--cccCCCCccccC---CCCccccccCCCc
Q 042843 296 RQQCEV--YALCGQFSTCNQ---QTERFCSCLKGFQ 326 (484)
Q Consensus 296 ~d~C~~--~~~CG~~giC~~---~~~~~C~C~~GF~ 326 (484)
-++|.- .++|= +|.|.. -+.+.|.|++||.
T Consensus 83 f~pC~~eyk~YCi-HG~C~yI~dL~epsCrC~~GYt 117 (126)
T PHA02887 83 FEKCKNDFNDFCI-NGECMNIIDLDEKFCICNKGYT 117 (126)
T ss_pred ccccChHhhCEee-CCEEEccccCCCceeECCCCcc
Confidence 367853 57787 789974 2458999999985
No 72
>PF12946 EGF_MSP1_1: MSP1 EGF domain 1; InterPro: IPR024730 This EGF-like domain is found at the C terminus of the malaria parasite MSP1 protein. MSP1 is the merozoite surface protein 1. This domain is part of the C-terminal fragment that is proteolytically processed from the the rest of the protein and is left attached to the surface of the invading parasite [].; PDB: 1N1I_C 2FLG_A 1CEJ_A 2NPR_A 1B9W_A 1OB1_F.
Probab=27.84 E-value=18 Score=24.38 Aligned_cols=25 Identities=24% Similarity=0.679 Sum_probs=16.3
Q ss_pred ccCCCCccccC--CCCccccccCCCcc
Q 042843 303 ALCGQFSTCNQ--QTERFCSCLKGFQQ 327 (484)
Q Consensus 303 ~~CG~~giC~~--~~~~~C~C~~GF~p 327 (484)
..|-.++-|-. +.+..|.|++||..
T Consensus 5 ~~cP~NA~C~~~~dG~eecrCllgyk~ 31 (37)
T PF12946_consen 5 TKCPANAGCFRYDDGSEECRCLLGYKK 31 (37)
T ss_dssp S---TTEEEEEETTSEEEEEE-TTEEE
T ss_pred ccCCCCcccEEcCCCCEEEEeeCCccc
Confidence 45778888963 35678999999975
No 73
>PF15102 TMEM154: TMEM154 protein family
Probab=27.67 E-value=48 Score=29.57 Aligned_cols=15 Identities=13% Similarity=-0.248 Sum_probs=6.6
Q ss_pred HheeeeeeeccCccc
Q 042843 463 MLVYLGRRKTATVTT 477 (484)
Q Consensus 463 ~~~~~~~r~~~~~~~ 477 (484)
+..+++-.+.|++|.
T Consensus 73 l~vV~lv~~~kRkr~ 87 (146)
T PF15102_consen 73 LSVVCLVIYYKRKRT 87 (146)
T ss_pred HHHHHheeEEeeccc
Confidence 334444444444444
No 74
>KOG1219 consensus Uncharacterized conserved protein, contains laminin, cadherin and EGF domains [Signal transduction mechanisms]
Probab=27.20 E-value=51 Score=41.80 Aligned_cols=24 Identities=25% Similarity=0.509 Sum_probs=17.9
Q ss_pred ccCCCCccccCC--CCccccccCCCc
Q 042843 303 ALCGQFSTCNQQ--TERFCSCLKGFQ 326 (484)
Q Consensus 303 ~~CG~~giC~~~--~~~~C~C~~GF~ 326 (484)
..|---|.|+.. +.-.|.||+-|.
T Consensus 3870 npCqhgG~C~~~~~ggy~CkCpsqys 3895 (4289)
T KOG1219|consen 3870 NPCQHGGTCISQPKGGYKCKCPSQYS 3895 (4289)
T ss_pred CcccCCCEecCCCCCceEEeCccccc
Confidence 567778899853 335799999875
No 75
>PF14870 PSII_BNR: Photosynthesis system II assembly factor YCF48; PDB: 2XBG_A.
Probab=26.96 E-value=6.6e+02 Score=25.31 Aligned_cols=98 Identities=21% Similarity=0.289 Sum_probs=0.0
Q ss_pred CceEEEEe-CCeEEEEcCCCceEEEeccCCCCCCceEEEEecCCCEEEEeccCCCCcceeeecccCceeccCCceeeeec
Q 042843 94 FSSVLRIS-DGNLVLFNESQLPIWSTNLTATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQSFDHPAHTWIPGMKLTFNK 172 (484)
Q Consensus 94 ~~~~l~l~-~G~LvL~~~~~~~vWst~~~~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQSFd~PTDTlLpgq~l~~n~ 172 (484)
....+..+ .|++......|.. |.............+.-+++|..|+....+ .+.+|+|.--+++.|-++..
T Consensus 114 ~~~~~l~~~~G~iy~T~DgG~t-W~~~~~~~~gs~~~~~r~~dG~~vavs~~G----~~~~s~~~G~~~w~~~~r~~--- 185 (302)
T PF14870_consen 114 DGSAELAGDRGAIYRTTDGGKT-WQAVVSETSGSINDITRSSDGRYVAVSSRG----NFYSSWDPGQTTWQPHNRNS--- 185 (302)
T ss_dssp TTEEEEEETT--EEEESSTTSS-EEEEE-S----EEEEEE-TTS-EEEEETTS----SEEEEE-TT-SS-EEEE--S---
T ss_pred CCcEEEEcCCCcEEEeCCCCCC-eeEcccCCcceeEeEEECCCCcEEEEECcc----cEEEEecCCCccceEEccCc---
Q ss_pred CCCCceEEEecCCCCCCCCceEEEEEcCCCCcEEEEEeeCCeeEEe
Q 042843 173 RNNVSQLITSWKNKENPAPGLFSLERAPDGSNQYVMLWNRSEQYWS 218 (484)
Q Consensus 173 ~~g~~~~L~Sw~s~~dps~G~y~l~~~~~g~~~~~l~~~~~~~Yw~ 218 (484)
.++|. .+.+.+++. +.+.-+|.+.+..
T Consensus 186 ----~~riq-------------~~gf~~~~~--lw~~~~Gg~~~~s 212 (302)
T PF14870_consen 186 ----SRRIQ-------------SMGFSPDGN--LWMLARGGQIQFS 212 (302)
T ss_dssp ----SS-EE-------------EEEE-TTS---EEEEETTTEEEEE
T ss_pred ----cceeh-------------hceecCCCC--EEEEeCCcEEEEc
No 76
>KOG1214 consensus Nidogen and related basement membrane protein proteins [Cell wall/membrane/envelope biogenesis; Extracellular structures]
Probab=25.57 E-value=51 Score=37.40 Aligned_cols=30 Identities=23% Similarity=0.609 Sum_probs=23.6
Q ss_pred CCCcccccCCCCccccCC-CCccccccCCCcc
Q 042843 297 QQCEVYALCGQFSTCNQQ-TERFCSCLKGFQQ 327 (484)
Q Consensus 297 d~C~~~~~CG~~giC~~~-~~~~C~C~~GF~p 327 (484)
|+|. +..|-+...|-+. .+..|.|.|||.-
T Consensus 828 DeC~-psrChp~A~CyntpgsfsC~C~pGy~G 858 (1289)
T KOG1214|consen 828 DECS-PSRCHPAATCYNTPGSFSCRCQPGYYG 858 (1289)
T ss_pred cccC-ccccCCCceEecCCCcceeecccCccC
Confidence 6777 7889999999754 3567999999963
No 77
>PF14575 EphA2_TM: Ephrin type-A receptor 2 transmembrane domain; PDB: 3KUL_A 2XVD_A 2VX1_A 2VWV_A 2VX0_A 2VWY_A 2VWZ_A 2VWW_A 2VWU_A 2VWX_A ....
Probab=24.13 E-value=22 Score=27.93 Aligned_cols=9 Identities=22% Similarity=0.217 Sum_probs=3.0
Q ss_pred eeeeeeecc
Q 042843 465 VYLGRRKTA 473 (484)
Q Consensus 465 ~~~~~r~~~ 473 (484)
+++++|+++
T Consensus 20 ~~~~~rr~~ 28 (75)
T PF14575_consen 20 VIVCFRRCK 28 (75)
T ss_dssp HHCCCTT--
T ss_pred EEEEEeeEc
Confidence 334444443
No 78
>PF06247 Plasmod_Pvs28: Plasmodium ookinete surface protein Pvs28; InterPro: IPR010423 This family consists of several ookinete surface protein (Pvs28) from several species of Plasmodium. Pvs25 and Pvs28 are expressed on the surface of ookinetes. These proteins are potential candidates for vaccine and induce antibodies that block the infectivity of Plasmodium vivax in immunised animals [].; GO: 0009986 cell surface, 0016020 membrane; PDB: 1Z3G_B 1Z1Y_B 1Z27_A.
Probab=21.21 E-value=33 Score=31.89 Aligned_cols=27 Identities=30% Similarity=0.830 Sum_probs=19.6
Q ss_pred cccCCCCccccCCC------CccccccCCCccC
Q 042843 302 YALCGQFSTCNQQT------ERFCSCLKGFQQK 328 (484)
Q Consensus 302 ~~~CG~~giC~~~~------~~~C~C~~GF~p~ 328 (484)
.-.||.|+.|...+ .-.|.|.+||...
T Consensus 49 ~K~Cgdya~C~~~~~~~~~~~~~C~C~~gY~~~ 81 (197)
T PF06247_consen 49 NKPCGDYAKCINQANKGEERAYKCDCINGYILK 81 (197)
T ss_dssp TSEEETTEEEEE-SSTTSSTSEEEEE-TTEEES
T ss_pred CccccchhhhhcCCCcccceeEEEecccCceee
Confidence 45799999998422 2469999999874
No 79
>smart00564 PQQ beta-propeller repeat. Beta-propeller repeat occurring in enzymes with pyrrolo-quinoline quinone (PQQ) as cofactor, in Ire1p-like Ser/Thr kinases, and in prokaryotic dehydrogenases.
Probab=20.54 E-value=1.9e+02 Score=17.80 Aligned_cols=16 Identities=25% Similarity=0.652 Sum_probs=9.4
Q ss_pred CCeEEEEcC-CCceEEE
Q 042843 102 DGNLVLFNE-SQLPIWS 117 (484)
Q Consensus 102 ~G~LvL~~~-~~~~vWs 117 (484)
+|.|+-.|. +|..+|.
T Consensus 15 ~g~l~a~d~~~G~~~W~ 31 (33)
T smart00564 15 DGTLYALDAKTGEILWT 31 (33)
T ss_pred CCEEEEEEcccCcEEEE
Confidence 566665554 4566665
Done!