Query 037760
Match_columns 471
No_of_seqs 217 out of 1680
Neff 7.4
Searched_HMMs 46136
Date Fri Mar 29 04:07:20 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/037760.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/037760hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF01453 B_lectin: D-mannose b 99.9 1.5E-27 3.3E-32 205.3 3.3 110 70-183 1-114 (114)
2 PF00954 S_locus_glycop: S-loc 99.9 8.5E-26 1.9E-30 193.4 11.7 110 209-321 1-110 (110)
3 cd00028 B_lectin Bulb-type man 99.9 4.6E-24 1E-28 184.4 15.3 114 32-151 2-116 (116)
4 smart00108 B_lectin Bulb-type 99.9 1.1E-22 2.3E-27 175.3 14.6 112 32-150 2-114 (114)
5 PF08276 PAN_2: PAN-like domai 99.7 1E-16 2.3E-21 124.3 5.7 64 340-404 1-66 (66)
6 cd00129 PAN_APPLE PAN/APPLE-li 99.5 2.6E-14 5.7E-19 114.7 6.2 72 340-421 5-80 (80)
7 cd01098 PAN_AP_plant Plant PAN 99.5 6.9E-14 1.5E-18 113.1 7.9 78 338-422 3-84 (84)
8 PF01453 B_lectin: D-mannose b 98.9 2.6E-08 5.7E-13 85.7 12.2 100 38-152 12-114 (114)
9 smart00473 PAN_AP divergent su 98.6 1.9E-07 4.1E-12 73.7 7.5 71 344-420 4-77 (78)
10 smart00108 B_lectin Bulb-type 98.5 9.9E-07 2.2E-11 75.8 9.6 86 89-211 23-111 (114)
11 cd00028 B_lectin Bulb-type man 98.3 3.9E-06 8.4E-11 72.3 8.6 86 90-212 24-113 (116)
12 cd01100 APPLE_Factor_XI_like S 97.3 0.00029 6.3E-09 55.5 4.1 49 349-400 9-58 (73)
13 PF00954 S_locus_glycop: S-loc 97.2 0.0017 3.7E-08 55.3 8.3 68 232-304 32-101 (110)
14 PF00024 PAN_1: PAN domain Thi 92.4 0.15 3.3E-06 39.7 3.5 55 345-402 3-59 (79)
15 PF04478 Mid2: Mid2 like cell 92.3 0.068 1.5E-06 47.8 1.4 32 439-470 45-79 (154)
16 smart00223 APPLE APPLE domain. 89.9 0.55 1.2E-05 37.6 4.4 51 349-399 6-57 (79)
17 PF08693 SKG6: Transmembrane a 86.9 0.95 2E-05 31.2 3.3 29 442-470 11-40 (40)
18 smart00605 CW CW domain. 84.3 3.8 8.3E-05 33.6 6.6 58 362-425 20-78 (94)
19 PF08277 PAN_3: PAN-like domai 82.8 4.3 9.4E-05 31.1 6.0 32 362-398 18-49 (71)
20 PF02009 Rifin_STEVOR: Rifin/s 82.4 0.4 8.7E-06 48.1 -0.1 25 446-470 259-284 (299)
21 PF02439 Adeno_E3_CR2: Adenovi 81.6 0.43 9.2E-06 32.4 -0.1 23 446-468 10-32 (38)
22 PF01102 Glycophorin_A: Glycop 80.1 0.54 1.2E-05 40.8 -0.1 16 455-470 79-94 (122)
23 cd00053 EGF Epidermal growth f 79.0 1.8 4E-05 27.6 2.3 29 291-319 2-31 (36)
24 PF14295 PAN_4: PAN domain; PD 77.6 2.4 5.3E-05 29.9 2.8 25 362-386 14-38 (51)
25 cd01099 PAN_AP_HGF Subfamily o 76.4 7.7 0.00017 30.8 5.7 36 363-401 24-61 (80)
26 PF07645 EGF_CA: Calcium-bindi 75.5 1.2 2.5E-05 30.9 0.6 31 289-319 3-35 (42)
27 PHA03265 envelope glycoprotein 74.6 2.5 5.4E-05 42.8 2.8 28 441-469 349-376 (402)
28 smart00179 EGF_CA Calcium-bind 72.1 3.6 7.8E-05 27.1 2.4 30 289-318 3-33 (39)
29 PTZ00382 Variant-specific surf 71.6 3.1 6.7E-05 34.6 2.3 8 313-320 9-16 (96)
30 PF14610 DUF4448: Protein of u 71.4 3.4 7.3E-05 38.6 2.8 28 444-471 160-187 (189)
31 PF09064 Tme5_EGF_like: Thromb 68.0 3.2 6.9E-05 27.5 1.3 18 302-319 11-28 (34)
32 PRK11138 outer membrane biogen 66.7 1.2E+02 0.0026 31.3 13.6 53 93-148 127-187 (394)
33 PF12877 DUF3827: Domain of un 66.3 6.5 0.00014 43.1 3.9 18 438-455 265-282 (684)
34 cd00054 EGF_CA Calcium-binding 66.1 5.6 0.00012 25.7 2.3 31 289-319 3-34 (38)
35 PTZ00046 rifin; Provisional 65.6 2.3 4.9E-05 43.6 0.3 27 445-471 317-344 (358)
36 TIGR01477 RIFIN variant surfac 64.4 2.4 5.3E-05 43.2 0.3 27 445-471 312-339 (353)
37 PF12661 hEGF: Human growth fa 63.6 2 4.4E-05 22.2 -0.2 9 310-318 1-9 (13)
38 PF01683 EB: EB module; Inter 62.2 7.4 0.00016 28.1 2.5 33 286-321 17-49 (52)
39 PF07974 EGF_2: EGF-like domai 60.7 7.6 0.00016 25.4 2.0 23 295-318 6-28 (32)
40 PF12947 EGF_3: EGF domain; I 59.2 2.7 5.9E-05 28.3 -0.3 26 294-319 5-31 (36)
41 PF01034 Syndecan: Syndecan do 58.6 3.2 6.9E-05 31.7 -0.0 14 457-470 26-39 (64)
42 PF05454 DAG1: Dystroglycan (D 56.4 3.7 8E-05 41.1 0.0 24 445-468 150-173 (290)
43 PF01299 Lamp: Lysosome-associ 53.4 7.9 0.00017 39.0 1.8 27 444-470 271-300 (306)
44 PF00008 EGF: EGF-like domain 53.4 4.1 9E-05 26.4 -0.1 23 296-318 5-29 (32)
45 PF06697 DUF1191: Protein of u 53.0 12 0.00027 37.0 3.0 22 445-466 216-238 (278)
46 PF12662 cEGF: Complement Clr- 52.9 6.5 0.00014 24.0 0.7 11 310-320 3-13 (24)
47 PF13908 Shisa: Wnt and FGF in 50.9 12 0.00026 34.5 2.5 17 443-459 79-95 (179)
48 PTZ00382 Variant-specific surf 50.9 14 0.00031 30.6 2.7 10 443-452 66-75 (96)
49 PF08374 Protocadherin: Protoc 49.7 14 0.0003 35.1 2.7 15 439-453 34-48 (221)
50 PF03302 VSP: Giardia variant- 46.0 17 0.00036 38.2 2.9 24 440-463 364-387 (397)
51 PF01436 NHL: NHL repeat; Int 45.1 34 0.00074 21.3 3.2 21 89-109 6-26 (28)
52 smart00181 EGF Epidermal growt 44.3 21 0.00045 22.9 2.2 24 295-319 6-30 (35)
53 TIGR01478 STEVOR variant surfa 42.8 11 0.00024 37.3 1.0 22 330-351 169-190 (295)
54 PTZ00370 STEVOR; Provisional 40.7 13 0.00027 37.0 1.0 22 330-351 169-190 (296)
55 TIGR01167 LPXTG_anchor LPXTG-m 36.1 36 0.00078 22.0 2.3 10 460-469 24-33 (34)
56 PF13360 PQQ_2: PQQ-like domai 35.7 81 0.0018 29.3 5.7 51 95-148 2-63 (238)
57 PRK11138 outer membrane biogen 35.0 1.7E+02 0.0037 30.1 8.5 20 93-112 263-283 (394)
58 cd05845 Ig2_L1-CAM_like Second 32.7 77 0.0017 26.2 4.3 33 69-102 32-64 (95)
59 PF15345 TMEM51: Transmembrane 32.6 50 0.0011 31.9 3.5 15 454-468 71-85 (233)
60 PF13360 PQQ_2: PQQ-like domai 32.5 2.3E+02 0.005 26.2 8.3 75 70-147 12-102 (238)
61 PF10681 Rot1: Chaperone for p 30.6 2.7E+02 0.0058 26.5 7.9 74 106-197 64-138 (212)
62 PF01102 Glycophorin_A: Glycop 30.0 15 0.00032 32.0 -0.4 27 445-471 66-92 (122)
63 TIGR03300 assembly_YfgL outer 29.9 2.1E+02 0.0046 29.0 8.1 20 94-113 73-93 (377)
64 PF06365 CD34_antigen: CD34/Po 29.5 61 0.0013 30.7 3.5 26 444-469 101-129 (202)
65 TIGR03300 assembly_YfgL outer 29.0 2.7E+02 0.0058 28.3 8.6 75 69-147 83-171 (377)
66 KOG3637 Vitronectin receptor, 26.8 48 0.001 39.2 2.9 15 452-466 987-1001(1030)
67 KOG4649 PQQ (pyrrolo-quinoline 26.0 2.3E+02 0.0051 28.2 6.9 46 69-114 167-217 (354)
68 PF10661 EssA: WXG100 protein 26.0 20 0.00043 32.1 -0.3 23 445-467 120-142 (145)
69 KOG1219 Uncharacterized conser 25.3 98 0.0021 39.6 4.9 25 296-320 3871-3897(4289)
70 PF05545 FixQ: Cbb3-type cytoc 25.0 64 0.0014 23.0 2.2 13 458-470 24-36 (49)
71 PF05393 Hum_adeno_E3A: Human 25.0 59 0.0013 26.5 2.2 8 456-463 45-52 (94)
72 TIGR03503 conserved hypothetic 24.6 12 0.00025 38.8 -2.3 21 449-469 353-373 (374)
73 PF02480 Herpes_gE: Alphaherpe 21.6 31 0.00067 36.8 0.0 13 329-341 237-251 (439)
74 PF12301 CD99L2: CD99 antigen 21.4 65 0.0014 29.7 2.0 24 441-464 113-136 (169)
75 PF05092 PIF: Per os infectivi 20.5 77 0.0017 34.2 2.6 47 287-336 147-195 (522)
76 PF05283 MGC-24: Multi-glycosy 20.1 67 0.0014 30.1 1.9 24 446-469 160-185 (186)
No 1
>PF01453 B_lectin: D-mannose binding lectin; InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]: Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity. Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=99.93 E-value=1.5e-27 Score=205.34 Aligned_cols=110 Identities=50% Similarity=0.801 Sum_probs=80.2
Q ss_pred CCeEEEEecCCCCCCC--CCceEEEccCCcEEEEeCCCceEEee-cCCCCc-cceEEEEecCCCEEEEeCcCCCCCceee
Q 037760 70 PRTVVWVANRYKPITD--KNGVLTLSNNGSILLLNQERSTIWSS-NSSRVL-ETAVVRLLDSGNLVLRDNVSRSSDEYMW 145 (471)
Q Consensus 70 ~~~~vW~an~~~pv~~--~~~~l~l~~~GnLvl~d~~~~~vWss-~~~~~~-~~~~~~Lld~GNlvl~~~~~~~~~~~~W 145 (471)
++++||+|||+.|+.. ...+|.|++||||||+|..+..+|++ .+.+.. ....++|+|+|||||++. .+.++|
T Consensus 1 ~~tvvW~an~~~p~~~~s~~~~L~l~~dGnLvl~~~~~~~iWss~~t~~~~~~~~~~~L~~~GNlvl~d~----~~~~lW 76 (114)
T PF01453_consen 1 PRTVVWVANRNSPLTSSSGNYTLILQSDGNLVLYDSNGSVIWSSNNTSGRGNSGCYLVLQDDGNLVLYDS----SGNVLW 76 (114)
T ss_dssp ---------TTEEEEECETTEEEEEETTSEEEEEETTTEEEEE--S-TTSS-SSEEEEEETTSEEEEEET----TSEEEE
T ss_pred CcccccccccccccccccccccceECCCCeEEEEcCCCCEEEEecccCCccccCeEEEEeCCCCEEEEee----cceEEE
Confidence 3689999999999854 24789999999999999998899999 555443 468899999999999996 678999
Q ss_pred eeccCCCCCCCCCCeeeeeccCCceeEEEEecCCCCCC
Q 037760 146 QSFDYPSDTLLPGMKLGWNLRTRFERYLTAWRNADDPT 183 (471)
Q Consensus 146 qSFd~PTDTLLPGq~L~~~~~tg~~~~L~Sw~s~~dps 183 (471)
|||||||||+||||+|+.+..+|.+..++||++.+|||
T Consensus 77 ~Sf~~ptdt~L~~q~l~~~~~~~~~~~~~sw~s~~dps 114 (114)
T PF01453_consen 77 QSFDYPTDTLLPGQKLGDGNVTGKNDSLTSWSSNTDPS 114 (114)
T ss_dssp ESTTSSS-EEEEEET--TSEEEEESTSSEEEESS----
T ss_pred eecCCCccEEEeccCcccCCCccccceEEeECCCCCCC
Confidence 99999999999999999876666666799999999996
No 2
>PF00954 S_locus_glycop: S-locus glycoprotein family; InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=99.93 E-value=8.5e-26 Score=193.40 Aligned_cols=110 Identities=45% Similarity=0.945 Sum_probs=102.9
Q ss_pred eeeCCCCCceeeccccccccCCCceeeeEEEecCCeeEEEEEecCCCceeEEEEccCCcEEEEEeecCCCCeeEeeeccC
Q 037760 209 VRSGPWNGQQFVGIPMFFPRLKNKVYIPMLVRTEDEAYYTYKPINDKVIPRLYLDQSGKLQRFVWNQTSSEWRMSYSWPF 288 (471)
Q Consensus 209 w~sg~w~~~~~~~~p~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~rl~ld~dG~l~~y~w~~~~~~W~~~~~~p~ 288 (471)
||+|+|+|..|++.|+|.. ...+.+.|+.++++++++|.+.+...++|++||++|+++++.|.+..++|.+.|.+|.
T Consensus 1 wrsG~WnG~~f~g~p~~~~---~~~~~~~fv~~~~e~~~t~~~~~~s~~~r~~ld~~G~l~~~~w~~~~~~W~~~~~~p~ 77 (110)
T PF00954_consen 1 WRSGPWNGQRFSGIPEMSS---NSLYNYSFVSNNEEVYYTYSLSNSSVLSRLVLDSDGQLQRYIWNESTQSWSVFWSAPK 77 (110)
T ss_pred CCccccCCeEECCcccccc---cceeEEEEEECCCeEEEEEecCCCceEEEEEEeeeeEEEEEEEecCCCcEEEEEEecc
Confidence 8999999999999999875 5678889999999999999998888899999999999999999999999999999999
Q ss_pred CCCcccCCCCCCccccCCCCCccccCCCCccCC
Q 037760 289 DACDNYAQCGANSNCRISKTPICECLAGFISKP 321 (471)
Q Consensus 289 ~~C~~~g~CG~~g~C~~~~~~~C~C~~GF~~~~ 321 (471)
+.|++|+.||+||+|+.+..+.|+|++||+|++
T Consensus 78 d~Cd~y~~CG~~g~C~~~~~~~C~Cl~GF~P~n 110 (110)
T PF00954_consen 78 DQCDVYGFCGPNGICNSNNSPKCSCLPGFEPKN 110 (110)
T ss_pred cCCCCccccCCccEeCCCCCCceECCCCcCCCc
Confidence 999999999999999987778999999999964
No 3
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=99.92 E-value=4.6e-24 Score=184.41 Aligned_cols=114 Identities=48% Similarity=0.813 Sum_probs=98.6
Q ss_pred CccCCCCeEEeCCCeeEEEEECCCCCCceEEEEEEeC-CCCeEEEEecCCCCCCCCCceEEEccCCcEEEEeCCCceEEe
Q 037760 32 QSISDGETLVSSSLRFELGFFSPGNSNNRYLGIWYKS-SPRTVVWVANRYKPITDKNGVLTLSNNGSILLLNQERSTIWS 110 (471)
Q Consensus 32 ~~L~~~~~L~S~~g~F~lgf~~~~~~~~~yl~i~~~~-~~~~~vW~an~~~pv~~~~~~l~l~~~GnLvl~d~~~~~vWs 110 (471)
+.|..|++|+|+++.|++|||.+......+++|||.. + .++||+||++.| ....+.|.|++||||+|+|.++.++|+
T Consensus 2 ~~l~~~~~l~s~~~~f~~G~~~~~~q~~dgnlv~~~~~~-~~~vW~snt~~~-~~~~~~l~l~~dGnLvl~~~~g~~vW~ 79 (116)
T cd00028 2 NPLSSGQTLVSSGSLFELGFFKLIMQSRDYNLILYKGSS-RTVVWVANRDNP-SGSSCTLTLQSDGNLVIYDGSGTVVWS 79 (116)
T ss_pred cCcCCCCEEEeCCCcEEEecccCCCCCCeEEEEEEeCCC-CeEEEECCCCCC-CCCCEEEEEecCCCeEEEcCCCcEEEE
Confidence 5688999999999999999999875433788999986 4 789999999988 345688999999999999999999999
Q ss_pred ecCCCCccceEEEEecCCCEEEEeCcCCCCCceeeeeccCC
Q 037760 111 SNSSRVLETAVVRLLDSGNLVLRDNVSRSSDEYMWQSFDYP 151 (471)
Q Consensus 111 s~~~~~~~~~~~~Lld~GNlvl~~~~~~~~~~~~WqSFd~P 151 (471)
|++.+......++|+|+|||||++. .+.++|||||||
T Consensus 80 S~~~~~~~~~~~~L~ddGnlvl~~~----~~~~~W~Sf~~P 116 (116)
T cd00028 80 SNTTRVNGNYVLVLLDDGNLVLYDS----DGNFLWQSFDYP 116 (116)
T ss_pred ecccCCCCceEEEEeCCCCEEEECC----CCCEEEcCCCCC
Confidence 9987523356899999999999997 577999999999
No 4
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=99.89 E-value=1.1e-22 Score=175.28 Aligned_cols=112 Identities=48% Similarity=0.819 Sum_probs=96.7
Q ss_pred CccCCCCeEEeCCCeeEEEEECCCCCCceEEEEEEeC-CCCeEEEEecCCCCCCCCCceEEEccCCcEEEEeCCCceEEe
Q 037760 32 QSISDGETLVSSSLRFELGFFSPGNSNNRYLGIWYKS-SPRTVVWVANRYKPITDKNGVLTLSNNGSILLLNQERSTIWS 110 (471)
Q Consensus 32 ~~L~~~~~L~S~~g~F~lgf~~~~~~~~~yl~i~~~~-~~~~~vW~an~~~pv~~~~~~l~l~~~GnLvl~d~~~~~vWs 110 (471)
+.|..|++|+|+++.|++|||.+... ..+++|||.. + .++||+||++.|+.. ++.|.|++||||||+|.++.++|+
T Consensus 2 ~~l~~~~~l~s~~~~f~~G~~~~~~q-~dgnlV~~~~~~-~~~vW~snt~~~~~~-~~~l~l~~dGnLvl~~~~g~~vW~ 78 (114)
T smart00108 2 NTLSSGQTLVSGNSLFELGFFTLIMQ-NDYNLILYKSSS-RTVVWVANRDNPVSD-SCTLTLQSDGNLVLYDGDGRVVWS 78 (114)
T ss_pred cccCCCCEEecCCCcEeeeccccCCC-CCEEEEEEECCC-CcEEEECCCCCCCCC-CEEEEEeCCCCEEEEeCCCCEEEE
Confidence 56788999999999999999998653 4788899987 5 789999999988754 488999999999999998999999
Q ss_pred ecCCCCccceEEEEecCCCEEEEeCcCCCCCceeeeeccC
Q 037760 111 SNSSRVLETAVVRLLDSGNLVLRDNVSRSSDEYMWQSFDY 150 (471)
Q Consensus 111 s~~~~~~~~~~~~Lld~GNlvl~~~~~~~~~~~~WqSFd~ 150 (471)
|++........++|+|+|||||++. .+.++||||||
T Consensus 79 S~t~~~~~~~~~~L~ddGnlvl~~~----~~~~~W~Sf~~ 114 (114)
T smart00108 79 SNTTGANGNYVLVLLDDGNLVIYDS----DGNFLWQSFDY 114 (114)
T ss_pred ecccCCCCceEEEEeCCCCEEEECC----CCCEEeCCCCC
Confidence 9986323356799999999999987 56799999997
No 5
>PF08276 PAN_2: PAN-like domain; InterPro: IPR013227 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs
Probab=99.66 E-value=1e-16 Score=124.32 Aligned_cols=64 Identities=56% Similarity=1.231 Sum_probs=54.7
Q ss_pred CCCCCceEEeccCCCCCc--ccccCCCCHHHHHHHHhhCCCeEEEEeccCCCCCCceEeecccccce
Q 037760 340 CPSGEGFLKLQRMKLPEN--YWSNKSMNLKECEAECIRNCSCRAYANSDITGGGDGCLMWFGDLIDI 404 (471)
Q Consensus 340 C~~~~~f~~~~~~~~p~~--~~~~~~~~~~~C~~~Cl~nCsC~a~~y~~~~~~g~gC~~w~~~l~~~ 404 (471)
|+.+|+|+++++|++|++ +.++.++++++|++.||+||||+||+|.++. ++++|++|.++|+|+
T Consensus 1 C~~~d~F~~l~~~~~p~~~~~~~~~~~s~~~C~~~Cl~nCsC~Ayay~~~~-~~~~C~lW~~~L~d~ 66 (66)
T PF08276_consen 1 CGSGDGFLKLPNMKLPDFDNAIVDSSVSLEECEKACLSNCSCTAYAYSNLS-GGGGCLLWYGDLVDL 66 (66)
T ss_pred CcCCCEEEEECCeeCCCCcceeeecCCCHHHHHhhcCCCCCEeeEEeeccC-CCCEEEEEcCEeecC
Confidence 545789999999999998 4444668999999999999999999998654 467899999999885
No 6
>cd00129 PAN_APPLE PAN/APPLE-like domain; present in N-terminal (N) domains of plasminogen/ hepatocyte growth factor proteins, plasma prekallikrein/coagulation factor XI and microneme antigen proteins, plant receptor-like protein kinases, and various nematode and leech anti-platelet proteins. Common structural features include two disulfide bonds that link the alpha-helix to the central region of the protein. PAN domains have significant functional versatility, fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=99.50 E-value=2.6e-14 Score=114.70 Aligned_cols=72 Identities=24% Similarity=0.407 Sum_probs=61.3
Q ss_pred CCCCCceEEeccCCCCCcccccCCCCHHHHHHHHhh---CCCeEEEEeccCCCCCCceEeecccc-cceeeeccccCCce
Q 037760 340 CPSGEGFLKLQRMKLPENYWSNKSMNLKECEAECIR---NCSCRAYANSDITGGGDGCLMWFGDL-IDIRECTEEFSWGQ 415 (471)
Q Consensus 340 C~~~~~f~~~~~~~~p~~~~~~~~~~~~~C~~~Cl~---nCsC~a~~y~~~~~~g~gC~~w~~~l-~~~~~~~~~~~~~~ 415 (471)
|..++.|+++.++++|++ ..++++||+++|++ ||||.||+|.+. +.||++|.++| +|+++.. ..+.
T Consensus 5 ~~~~g~fl~~~~~klpd~----~~~s~~eC~~~Cl~~~~nCsC~Aya~~~~---~~gC~~W~~~l~~d~~~~~---~~g~ 74 (80)
T cd00129 5 CKSAGTTLIKIALKIKTT----KANTADECANRCEKNGLPFSCKAFVFAKA---RKQCLWFPFNSMSGVRKEF---SHGF 74 (80)
T ss_pred eecCCeEEEeecccCCcc----cccCHHHHHHHHhcCCCCCCceeeeccCC---CCCeEEecCcchhhHHhcc---CCCc
Confidence 444578999999999998 22689999999999 999999999752 45899999999 9998877 6789
Q ss_pred eEEEEe
Q 037760 416 DIFIRV 421 (471)
Q Consensus 416 ~~yirv 421 (471)
++|||.
T Consensus 75 ~Ly~r~ 80 (80)
T cd00129 75 DLYENK 80 (80)
T ss_pred eeEeEC
Confidence 999983
No 7
>cd01098 PAN_AP_plant Plant PAN/APPLE-like domain; present in plant S-receptor protein kinases and secreted glycoproteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions. S-receptor protein kinases and S-locus glycoproteins are involved in sporophytic self-incompatibility response in Brassica, one of probably many molecular mechanisms, by which hermaphrodite flowering plants avoid self-fertilization.
Probab=99.49 E-value=6.9e-14 Score=113.10 Aligned_cols=78 Identities=41% Similarity=0.913 Sum_probs=63.5
Q ss_pred CCCCCC---CceEEeccCCCCCc-ccccCCCCHHHHHHHHhhCCCeEEEEeccCCCCCCceEeecccccceeeeccccCC
Q 037760 338 SDCPSG---EGFLKLQRMKLPEN-YWSNKSMNLKECEAECIRNCSCRAYANSDITGGGDGCLMWFGDLIDIRECTEEFSW 413 (471)
Q Consensus 338 ~~C~~~---~~f~~~~~~~~p~~-~~~~~~~~~~~C~~~Cl~nCsC~a~~y~~~~~~g~gC~~w~~~l~~~~~~~~~~~~ 413 (471)
++|... +.|+++.++++|+. ... ...++++|++.||+||+|+||+|.+ ++++|++|...+.+.+... ..
T Consensus 3 ~~C~~~~~~~~f~~~~~~~~~~~~~~~-~~~s~~~C~~~Cl~nCsC~a~~~~~---~~~~C~~~~~~~~~~~~~~---~~ 75 (84)
T cd01098 3 LNCGGDGSTDGFLKLPDVKLPDNASAI-TAISLEECREACLSNCSCTAYAYNN---GSGGCLLWNGLLNNLRSLS---SG 75 (84)
T ss_pred cccCCCCCCCEEEEeCCeeCCCchhhh-ccCCHHHHHHHHhcCCCcceeeecC---CCCeEEEEeceecceEeec---CC
Confidence 467543 68999999999987 333 6679999999999999999999985 2457999999999877654 44
Q ss_pred ceeEEEEee
Q 037760 414 GQDIFIRVP 422 (471)
Q Consensus 414 ~~~~yirv~ 422 (471)
+..+||||+
T Consensus 76 ~~~~yiKv~ 84 (84)
T cd01098 76 GGTLYLRLA 84 (84)
T ss_pred CcEEEEEeC
Confidence 688999985
No 8
>PF01453 B_lectin: D-mannose binding lectin; InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]: Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity. Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=98.89 E-value=2.6e-08 Score=85.71 Aligned_cols=100 Identities=27% Similarity=0.416 Sum_probs=68.3
Q ss_pred CeEEeCCCeeEEEEECCCCCCceEEEEEEeCCCCeEEEEe-cCCCCCCCCCceEEEccCCcEEEEeCCCceEEeecCCCC
Q 037760 38 ETLVSSSLRFELGFFSPGNSNNRYLGIWYKSSPRTVVWVA-NRYKPITDKNGVLTLSNNGSILLLNQERSTIWSSNSSRV 116 (471)
Q Consensus 38 ~~L~S~~g~F~lgf~~~~~~~~~yl~i~~~~~~~~~vW~a-n~~~pv~~~~~~l~l~~~GnLvl~d~~~~~vWss~~~~~ 116 (471)
+.+.+.+|.+.|-|+.+++ |.| |. ...+++|.+ +...... ..+.|+|+++|||||+|..+.+||+|+....
T Consensus 12 ~p~~~~s~~~~L~l~~dGn-----Lvl-~~-~~~~~iWss~~t~~~~~-~~~~~~L~~~GNlvl~d~~~~~lW~Sf~~pt 83 (114)
T PF01453_consen 12 SPLTSSSGNYTLILQSDGN-----LVL-YD-SNGSVIWSSNNTSGRGN-SGCYLVLQDDGNLVLYDSSGNVLWQSFDYPT 83 (114)
T ss_dssp EEEEECETTEEEEEETTSE-----EEE-EE-TTTEEEEE--S-TTSS--SSEEEEEETTSEEEEEETTSEEEEESTTSSS
T ss_pred cccccccccccceECCCCe-----EEE-Ec-CCCCEEEEecccCCccc-cCeEEEEeCCCCEEEEeecceEEEeecCCCc
Confidence 4565656999999999875 433 44 345789999 4443321 4689999999999999999999999976322
Q ss_pred ccceEEEEec--CCCEEEEeCcCCCCCceeeeeccCCC
Q 037760 117 LETAVVRLLD--SGNLVLRDNVSRSSDEYMWQSFDYPS 152 (471)
Q Consensus 117 ~~~~~~~Lld--~GNlvl~~~~~~~~~~~~WqSFd~PT 152 (471)
...+..++ .||++ +.. ...++|.|-++|+
T Consensus 84 --dt~L~~q~l~~~~~~-~~~----~~~~sw~s~~dps 114 (114)
T PF01453_consen 84 --DTLLPGQKLGDGNVT-GKN----DSLTSWSSNTDPS 114 (114)
T ss_dssp ---EEEEEET--TSEEE-EES----TSSEEEESS----
T ss_pred --cEEEeccCcccCCCc-ccc----ceEEeECCCCCCC
Confidence 34566666 78888 553 4569999999885
No 9
>smart00473 PAN_AP divergent subfamily of APPLE domains. Apple-like domains present in Plasminogen, C. elegans hypothetical ORFs and the extracellular portion of plant receptor-like protein kinases. Predicted to possess protein- and/or carbohydrate-binding functions.
Probab=98.58 E-value=1.9e-07 Score=73.66 Aligned_cols=71 Identities=35% Similarity=0.759 Sum_probs=55.0
Q ss_pred CceEEeccCCCCCc-ccccCCCCHHHHHHHHhh-CCCeEEEEeccCCCCCCceEeec-ccccceeeeccccCCceeEEEE
Q 037760 344 EGFLKLQRMKLPEN-YWSNKSMNLKECEAECIR-NCSCRAYANSDITGGGDGCLMWF-GDLIDIRECTEEFSWGQDIFIR 420 (471)
Q Consensus 344 ~~f~~~~~~~~p~~-~~~~~~~~~~~C~~~Cl~-nCsC~a~~y~~~~~~g~gC~~w~-~~l~~~~~~~~~~~~~~~~yir 420 (471)
..|.+++++.+++. .......++++|++.|++ +|+|.||.|.. .+.+|.+|. +.+.+.+... ..+.++|.|
T Consensus 4 ~~f~~~~~~~l~~~~~~~~~~~s~~~C~~~C~~~~~~C~s~~y~~---~~~~C~l~~~~~~~~~~~~~---~~~~~~y~~ 77 (78)
T smart00473 4 DCFVRLPNTKLPGFSRIVISVASLEECASKCLNSNCSCRSFTYNN---GTKGCLLWSESSLGDARLFP---SGGVDLYEK 77 (78)
T ss_pred ceeEEecCccCCCCcceeEcCCCHHHHHHHhCCCCCceEEEEEcC---CCCEEEEeeCCccccceecc---cCCceeEEe
Confidence 46899999999865 322345699999999999 99999999974 245699999 7777776444 456677776
No 10
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=98.47 E-value=9.9e-07 Score=75.76 Aligned_cols=86 Identities=20% Similarity=0.365 Sum_probs=61.0
Q ss_pred eEEEccCCcEEEEeCC-CceEEeecCCCCcc-ceEEEEecCCCEEEEeCcCCCCCceeeeeccCCCCCCCCCCeeeeecc
Q 037760 89 VLTLSNNGSILLLNQE-RSTIWSSNSSRVLE-TAVVRLLDSGNLVLRDNVSRSSDEYMWQSFDYPSDTLLPGMKLGWNLR 166 (471)
Q Consensus 89 ~l~l~~~GnLvl~d~~-~~~vWss~~~~~~~-~~~~~Lld~GNlvl~~~~~~~~~~~~WqSFd~PTDTLLPGq~L~~~~~ 166 (471)
.+.++.|||||+++.. ..++|++++..+.. ...+.|+++|||||++. .+.++|+|-..
T Consensus 23 ~~~~q~dgnlV~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~----~g~~vW~S~t~---------------- 82 (114)
T smart00108 23 TLIMQNDYNLILYKSSSRTVVWVANRDNPVSDSCTLTLQSDGNLVLYDG----DGRVVWSSNTT---------------- 82 (114)
T ss_pred ccCCCCCEEEEEEECCCCcEEEECCCCCCCCCCEEEEEeCCCCEEEEeC----CCCEEEEeccc----------------
Confidence 3556789999999865 57999999864422 36789999999999987 46789998221
Q ss_pred CCceeEEEEecCCCCCCCceEEEEEccCCCeeEEEee-Cceeeeee
Q 037760 167 TRFERYLTAWRNADDPTPGEFSFRFDISTMAELVTVT-GSKIEVRS 211 (471)
Q Consensus 167 tg~~~~L~Sw~s~~dps~G~f~l~l~~~g~~~~~~~~-~~~~Yw~s 211 (471)
...|.+.+.|+++|..++ ++ ..++.|.+
T Consensus 83 ---------------~~~~~~~~~L~ddGnlvl--~~~~~~~~W~S 111 (114)
T smart00108 83 ---------------GANGNYVLVLLDDGNLVI--YDSDGNFLWQS 111 (114)
T ss_pred ---------------CCCCceEEEEeCCCCEEE--ECCCCCEEeCC
Confidence 023456778888888544 43 23577865
No 11
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=98.27 E-value=3.9e-06 Score=72.32 Aligned_cols=86 Identities=20% Similarity=0.310 Sum_probs=60.7
Q ss_pred EEEcc-CCcEEEEeCC-CceEEeecCCCC-ccceEEEEecCCCEEEEeCcCCCCCceeeeeccCCCCCCCCCCeeeeecc
Q 037760 90 LTLSN-NGSILLLNQE-RSTIWSSNSSRV-LETAVVRLLDSGNLVLRDNVSRSSDEYMWQSFDYPSDTLLPGMKLGWNLR 166 (471)
Q Consensus 90 l~l~~-~GnLvl~d~~-~~~vWss~~~~~-~~~~~~~Lld~GNlvl~~~~~~~~~~~~WqSFd~PTDTLLPGq~L~~~~~ 166 (471)
+.++. +|+||+++.. ..++|++++..+ .....+.|+++|||||++. ++.++|+|-...
T Consensus 24 ~~~q~~dgnlv~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~----~g~~vW~S~~~~--------------- 84 (116)
T cd00028 24 LIMQSRDYNLILYKGSSRTVVWVANRDNPSGSSCTLTLQSDGNLVIYDG----SGTVVWSSNTTR--------------- 84 (116)
T ss_pred CCCCCCeEEEEEEeCCCCeEEEECCCCCCCCCCEEEEEecCCCeEEEcC----CCcEEEEecccC---------------
Confidence 44565 9999999764 479999998653 2346789999999999987 467899874321
Q ss_pred CCceeEEEEecCCCCCCCceEEEEEccCCCeeEEEee-CceeeeeeC
Q 037760 167 TRFERYLTAWRNADDPTPGEFSFRFDISTMAELVTVT-GSKIEVRSG 212 (471)
Q Consensus 167 tg~~~~L~Sw~s~~dps~G~f~l~l~~~g~~~~~~~~-~~~~Yw~sg 212 (471)
..+.+.+.|+++|...+ ++ ...+.|.+.
T Consensus 85 ----------------~~~~~~~~L~ddGnlvl--~~~~~~~~W~Sf 113 (116)
T cd00028 85 ----------------VNGNYVLVLLDDGNLVL--YDSDGNFLWQSF 113 (116)
T ss_pred ----------------CCCceEEEEeCCCCEEE--ECCCCCEEEcCC
Confidence 13456778888887444 43 245778764
No 12
>cd01100 APPLE_Factor_XI_like Subfamily of PAN/APPLE-like domains; present in plasma prekallikrein/coagulation factor XI, microneme antigen proteins, and a few prokaryotic proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=97.30 E-value=0.00029 Score=55.46 Aligned_cols=49 Identities=12% Similarity=0.328 Sum_probs=34.5
Q ss_pred eccCCCCCc-ccccCCCCHHHHHHHHhhCCCeEEEEeccCCCCCCceEeeccc
Q 037760 349 LQRMKLPEN-YWSNKSMNLKECEAECIRNCSCRAYANSDITGGGDGCLMWFGD 400 (471)
Q Consensus 349 ~~~~~~p~~-~~~~~~~~~~~C~~~Cl~nCsC~a~~y~~~~~~g~gC~~w~~~ 400 (471)
+++++++.. .......+.++|++.|+.+|+|.||.|.. +...|+++...
T Consensus 9 ~~~~~~~g~d~~~~~~~s~~~Cq~~C~~~~~C~afT~~~---~~~~C~lk~~~ 58 (73)
T cd01100 9 GSNVDFRGGDLSTVFASSAEQCQAACTADPGCLAFTYNT---KSKKCFLKSSE 58 (73)
T ss_pred cCCCccccCCcceeecCCHHHHHHHcCCCCCceEEEEEC---CCCeEEcccCC
Confidence 346666554 21222458999999999999999999974 23359997653
No 13
>PF00954 S_locus_glycop: S-locus glycoprotein family; InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=97.22 E-value=0.0017 Score=55.34 Aligned_cols=68 Identities=12% Similarity=0.222 Sum_probs=56.0
Q ss_pred ceeeeEEEecCCeeEEEEEecCCCceeEEEEccCCcEEEEEeecCCCCeeEeeeccCCCCcccCCCCCC--cccc
Q 037760 232 KVYIPMLVRTEDEAYYTYKPINDKVIPRLYLDQSGKLQRFVWNQTSSEWRMSYSWPFDACDNYAQCGAN--SNCR 304 (471)
Q Consensus 232 ~~~~~~~~~~~~~~~~~~~~~~~~~~~rl~ld~dG~l~~y~w~~~~~~W~~~~~~p~~~C~~~g~CG~~--g~C~ 304 (471)
....++|...+..++.++.+...+.+++++++.+.+.|...|..+.+. |+.++.|+.+|+|..+ ..|.
T Consensus 32 ~e~~~t~~~~~~s~~~r~~ld~~G~l~~~~w~~~~~~W~~~~~~p~d~-----Cd~y~~CG~~g~C~~~~~~~C~ 101 (110)
T PF00954_consen 32 EEVYYTYSLSNSSVLSRLVLDSDGQLQRYIWNESTQSWSVFWSAPKDQ-----CDVYGFCGPNGICNSNNSPKCS 101 (110)
T ss_pred CeEEEEEecCCCceEEEEEEeeeeEEEEEEEecCCCcEEEEEEecccC-----CCCccccCCccEeCCCCCCceE
Confidence 345567776666777788888888999999999999999999988776 9999999999999765 4575
No 14
>PF00024 PAN_1: PAN domain This Prosite entry concerns apple domains, a subset of PAN domains; InterPro: IPR003014 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs It has been shown that, the N-terminal N domains of members of the plasminogen/hepatocyte growth factor family, the apple domains of the plasma prekallikrein/coagulation factor XI family, and domains of various nematode proteins belong to the same module superfamily, the PAN module []. PAN contains a conserved core of three disulphide bridges. In some members of the family there is an additional fourth disulphide bridge that links the N and C termini of the domain.; PDB: 1GP9_C 2QJ2_B 1GMO_H 1NK1_B 3MKP_B 1BHT_B 3HN4_A 1GMN_A 3HMS_A 3HMT_B ....
Probab=92.37 E-value=0.15 Score=39.70 Aligned_cols=55 Identities=16% Similarity=0.397 Sum_probs=39.5
Q ss_pred ceEEeccCCCCCc-ccccCCCCHHHHHHHHhhCCC-eEEEEeccCCCCCCceEeeccccc
Q 037760 345 GFLKLQRMKLPEN-YWSNKSMNLKECEAECIRNCS-CRAYANSDITGGGDGCLMWFGDLI 402 (471)
Q Consensus 345 ~f~~~~~~~~p~~-~~~~~~~~~~~C~~~Cl~nCs-C~a~~y~~~~~~g~gC~~w~~~l~ 402 (471)
.|.++.+..+... .......++++|.+.|+.+=. |.+|.|.. ....|.+...+-.
T Consensus 3 ~f~~~~~~~l~~~~~~~~~v~s~~~C~~~C~~~~~~C~s~~y~~---~~~~C~L~~~~~~ 59 (79)
T PF00024_consen 3 AFERIPGYRLSGHSIKEINVPSLEECAQLCLNEPRRCKSFNYDP---SSKTCYLSSSDRS 59 (79)
T ss_dssp TEEEEEEEEEESCEEEEEEESSHHHHHHHHHHSTT-ESEEEEET---TTTEEEEECSSSS
T ss_pred CeEEECCEEEeCCcceEEcCCCHHHHHhhcCcCcccCCeEEEEC---CCCEEEEcCCCCC
Confidence 3777777776665 222234589999999999999 99999985 2335998765443
No 15
>PF04478 Mid2: Mid2 like cell wall stress sensor; InterPro: IPR007567 This family represents a region near the C terminus of Mid2, which contains a transmembrane region. The remainder of the protein sequence is serine-rich and of low complexity, and is therefore impossible to align accurately. Mid2 is thought to act as a mechanosensor of cell wall stress. The C-terminal cytoplasmic region of Mid2 is known to interact with Rom2, a guanine nucleotide exchange factor (GEF) for Rho1, which is part of the cell wall integrity signalling pathway [].
Probab=92.26 E-value=0.068 Score=47.78 Aligned_cols=32 Identities=19% Similarity=0.232 Sum_probs=16.7
Q ss_pred ccceeEEEEehhHHHH--HH-HHHHHHhhhhhhcC
Q 037760 439 KKRLKIIVAMSIISGM--LI-LGLLLGMAWKKAKN 470 (471)
Q Consensus 439 ~~~~~~ii~~~v~~~~--~~-~~~~~~~~~~~~~~ 470 (471)
.+.+++||+++||+.+ ++ +++++|++++|+||
T Consensus 45 ~knknIVIGvVVGVGg~ill~il~lvf~~c~r~kk 79 (154)
T PF04478_consen 45 SKNKNIVIGVVVGVGGPILLGILALVFIFCIRRKK 79 (154)
T ss_pred cCCccEEEEEEecccHHHHHHHHHhheeEEEeccc
Confidence 3344688998887543 22 23334444444443
No 16
>smart00223 APPLE APPLE domain. Four-fold repeat in plasma kallikrein and coagulation factor XI. Factor XI apple 3 mediates binding to platelets. Factor XI apple 1 binds high-molecular-mass kininogen. Apple 4 in factor XI mediates dimer formation and binds to factor XIIa. Mutations in apple 4 cause factor XI deficiency, an inherited bleeding disorder.
Probab=89.89 E-value=0.55 Score=37.58 Aligned_cols=51 Identities=12% Similarity=0.217 Sum_probs=34.4
Q ss_pred eccCCCCCc-ccccCCCCHHHHHHHHhhCCCeEEEEeccCCCCCCceEeecc
Q 037760 349 LQRMKLPEN-YWSNKSMNLKECEAECIRNCSCRAYANSDITGGGDGCLMWFG 399 (471)
Q Consensus 349 ~~~~~~p~~-~~~~~~~~~~~C~~~Cl~nCsC~a~~y~~~~~~g~gC~~w~~ 399 (471)
+++++++.. .......+.++|++.|..+=.|.+|.|.........|+++..
T Consensus 6 ~~~~df~G~Dl~~~~~~~~~~Cq~~Ct~~~~C~~FTf~~~~~~~~~C~LK~s 57 (79)
T smart00223 6 YKNVDFRGSDINTVYVPSAQVCQKRCTSHPRCLFFTFSTNEPPEEKCLLKDS 57 (79)
T ss_pred ccCccccCceeeeeecCCHHHHHHhhcCCCCccEEEeeCCCCCCCEeEeCcC
Confidence 345555554 222234589999999999999999999753221226998643
No 17
>PF08693 SKG6: Transmembrane alpha-helix domain; InterPro: IPR014805 SKG6 and AXL2 are membrane proteins that show polarised intracellular localisation [, ]. This entry represents the highly conserved transmembrane alpha-helical domain found in these proteins [, ]. The full-length AXL2 protein has a negative regulatory function in cytokinesis [].
Probab=86.86 E-value=0.95 Score=31.24 Aligned_cols=29 Identities=24% Similarity=0.472 Sum_probs=12.8
Q ss_pred eeEEEEehhHHHHHHHHH-HHHhhhhhhcC
Q 037760 442 LKIIVAMSIISGMLILGL-LLGMAWKKAKN 470 (471)
Q Consensus 442 ~~~ii~~~v~~~~~~~~~-~~~~~~~~~~~ 470 (471)
.-+-++++|.+.++++.+ +.+++|+||+|
T Consensus 11 vaIa~~VvVPV~vI~~vl~~~l~~~~rR~k 40 (40)
T PF08693_consen 11 VAIAVGVVVPVGVIIIVLGAFLFFWYRRKK 40 (40)
T ss_pred EEEEEEEEechHHHHHHHHHHhheEEeccC
Confidence 334455555544433222 33344555543
No 18
>smart00605 CW CW domain.
Probab=84.32 E-value=3.8 Score=33.56 Aligned_cols=58 Identities=14% Similarity=0.461 Sum_probs=39.6
Q ss_pred CCCCHHHHHHHHhhCCCeEEEEeccCCCCCCceEeecc-cccceeeeccccCCceeEEEEeecCc
Q 037760 362 KSMNLKECEAECIRNCSCRAYANSDITGGGDGCLMWFG-DLIDIRECTEEFSWGQDIFIRVPAAD 425 (471)
Q Consensus 362 ~~~~~~~C~~~Cl~nCsC~a~~y~~~~~~g~gC~~w~~-~l~~~~~~~~~~~~~~~~yirv~~s~ 425 (471)
...+.++|...|..+..|+.+.... ...|.++.- ++..+++... ..+..+=+|+..+.
T Consensus 20 ~~~sw~~Ci~~C~~~~~Cvlay~~~----~~~C~~f~~~~~~~v~~~~~--~~~~~VAfK~~~~~ 78 (94)
T smart00605 20 ATLSWDECIQKCYEDSNCVLAYGNS----SETCYLFSYGTVLTVKKLSS--SSGKKVAFKVSTDQ 78 (94)
T ss_pred cCCCHHHHHHHHhCCCceEEEecCC----CCceEEEEcCCeEEEEEccC--CCCcEEEEEEeCCC
Confidence 3467899999999999999876541 246987643 4556666541 34566778876443
No 19
>PF08277 PAN_3: PAN-like domain; InterPro: IPR006583 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs The PAN-3 or CW is a domain associated with a number of Caenorhabditis elegans hypothetical proteins.
Probab=82.81 E-value=4.3 Score=31.08 Aligned_cols=32 Identities=13% Similarity=0.434 Sum_probs=26.5
Q ss_pred CCCCHHHHHHHHhhCCCeEEEEeccCCCCCCceEeec
Q 037760 362 KSMNLKECEAECIRNCSCRAYANSDITGGGDGCLMWF 398 (471)
Q Consensus 362 ~~~~~~~C~~~Cl~nCsC~a~~y~~~~~~g~gC~~w~ 398 (471)
...+.++|-..|..+=.|.++.+. ...|.++.
T Consensus 18 ~~~sw~~Cv~~C~~~~~C~la~~~-----~~~C~~y~ 49 (71)
T PF08277_consen 18 TNTSWDDCVQKCYNDENCVLAYFD-----SGKCYLYN 49 (71)
T ss_pred cCCCHHHHhHHhCCCCEEEEEEeC-----CCCEEEEE
Confidence 345789999999999999998886 24799874
No 20
>PF02009 Rifin_STEVOR: Rifin/stevor family; InterPro: IPR002858 Malaria is still a major cause of mortality in many areas of the world. Plasmodium falciparum causes the most severe human form of the disease and is responsible for most fatalities. Severe cases of malaria can occur when the parasite invades and then proliferates within red blood cell erythrocytes. The parasite produces many variant antigenic proteins, encoded by multigene families, which are present on the surface of the infected erythrocyte and play important roles in virulence. A crucial survival mechanism for the malaria parasite is its ability to evade the immune response by switching these variant surface antigens. The high virulence of P. falciparum relative to other malarial parasites is in large part due to the fact that in this organism many of these surface antigens mediate the binding of infected erythrocytes to the vascular endothelium (cytoadherence) and non-infected erythrocytes (rosetting). This can lead to the accumulation of infected cells in the vasculature of a variety of organs, blocking the blood flow and reducing the oxygen supply. Clinical symptoms of severe infection can include fever, progressive anaemia, multi-organ dysfunction and coma. For more information see []. Several multicopy gene families have been described in Plasmodium falciparum, including the stevor family of subtelomeric open reading frames and the rif interspersed repetitive elements. Both families contain three predicted transmembrane segments. It has been proposed that stevor and rif are members of a larger superfamily that code for variant surface antigens [].
Probab=82.36 E-value=0.4 Score=48.10 Aligned_cols=25 Identities=8% Similarity=0.187 Sum_probs=13.8
Q ss_pred EEehhHHHH-HHHHHHHHhhhhhhcC
Q 037760 446 VAMSIISGM-LILGLLLGMAWKKAKN 470 (471)
Q Consensus 446 i~~~v~~~~-~~~~~~~~~~~~~~~~ 470 (471)
++++|+++| +++.+++|++||.|||
T Consensus 259 ~aSiiaIliIVLIMvIIYLILRYRRK 284 (299)
T PF02009_consen 259 IASIIAILIIVLIMVIIYLILRYRRK 284 (299)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 334444333 3356677777776654
No 21
>PF02439 Adeno_E3_CR2: Adenovirus E3 region protein CR2; InterPro: IPR003470 Early region 3 (E3) of human adenoviruses (Ads) codes for proteins that appear to control viral interactions with the host []. This region called CR1 (conserved region 1) [] is found three times in Human adenovirus 19 (a subgroup D adenovirus) 49 kDa protein in the E3 region. CR1 is also found in the 20.1 Kd protein of subgroup B adenoviruses. The function of this 80 amino acid region is unknown. This region is probably a divergent immunoglobulin domain.
Probab=81.60 E-value=0.43 Score=32.38 Aligned_cols=23 Identities=17% Similarity=0.212 Sum_probs=11.3
Q ss_pred EEehhHHHHHHHHHHHHhhhhhh
Q 037760 446 VAMSIISGMLILGLLLGMAWKKA 468 (471)
Q Consensus 446 i~~~v~~~~~~~~~~~~~~~~~~ 468 (471)
++++++.++++++++.|..++||
T Consensus 10 v~V~vg~~iiii~~~~YaCcykk 32 (38)
T PF02439_consen 10 VAVVVGMAIIIICMFYYACCYKK 32 (38)
T ss_pred HHHHHHHHHHHHHHHHHHHHHcc
Confidence 33444444455555555455444
No 22
>PF01102 Glycophorin_A: Glycophorin A; InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=80.07 E-value=0.54 Score=40.82 Aligned_cols=16 Identities=19% Similarity=-0.010 Sum_probs=7.3
Q ss_pred HHHHHHHHhhhhhhcC
Q 037760 455 LILGLLLGMAWKKAKN 470 (471)
Q Consensus 455 ~~~~~~~~~~~~~~~~ 470 (471)
++++++.|+++|+|||
T Consensus 79 g~Illi~y~irR~~Kk 94 (122)
T PF01102_consen 79 GIILLISYCIRRLRKK 94 (122)
T ss_dssp HHHHHHHHHHHHHS--
T ss_pred HHHHHHHHHHHHHhcc
Confidence 3344555555555554
No 23
>cd00053 EGF Epidermal growth factor domain, found in epidermal growth factor (EGF) presents in a large number of proteins, mostly animal; the list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied; the functional significance of EGF-like domains in what appear to be unrelated proteins is not yet clear; a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase); the domain includes six cysteine residues which have been shown to be involved in disulfide bonds; the main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet; Subdomains between the conserved cysteines vary in length; the region between the 5th and 6th cysteine contains two conserved glycines of which at least one is present in most EGF-like domains; a subset of these bind calcium.
Probab=79.05 E-value=1.8 Score=27.62 Aligned_cols=29 Identities=21% Similarity=0.639 Sum_probs=20.7
Q ss_pred CcccCCCCCCccccCCC-CCccccCCCCcc
Q 037760 291 CDNYAQCGANSNCRISK-TPICECLAGFIS 319 (471)
Q Consensus 291 C~~~g~CG~~g~C~~~~-~~~C~C~~GF~~ 319 (471)
|.....|..++.|.... ...|.|++||..
T Consensus 2 C~~~~~C~~~~~C~~~~~~~~C~C~~g~~g 31 (36)
T cd00053 2 CAASNPCSNGGTCVNTPGSYRCVCPPGYTG 31 (36)
T ss_pred CCCCCCCCCCCEEecCCCCeEeECCCCCcc
Confidence 34346788888897543 358999999964
No 24
>PF14295 PAN_4: PAN domain; PDB: 2YIL_E 2YIP_C 2YIO_A.
Probab=77.57 E-value=2.4 Score=29.94 Aligned_cols=25 Identities=24% Similarity=0.622 Sum_probs=17.8
Q ss_pred CCCCHHHHHHHHhhCCCeEEEEecc
Q 037760 362 KSMNLKECEAECIRNCSCRAYANSD 386 (471)
Q Consensus 362 ~~~~~~~C~~~Cl~nCsC~a~~y~~ 386 (471)
...+.++|.+.|..+=.|.+|.|..
T Consensus 14 ~~~s~~~C~~~C~~~~~C~~~~~~~ 38 (51)
T PF14295_consen 14 TASSPEECQAACAADPGCQAFTFNP 38 (51)
T ss_dssp ----HHHHHHHHHTSTT--EEEEET
T ss_pred cCCCHHHHHHHccCCCCCCEEEEEC
Confidence 3458999999999999999999974
No 25
>cd01099 PAN_AP_HGF Subfamily of PAN/APPLE-like domains; present in N-terminal (N) domains of plasminogen/hepatocyte growth factor proteins, and various proteins found in Bilateria, such as leech anti-platelet proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=76.42 E-value=7.7 Score=30.79 Aligned_cols=36 Identities=22% Similarity=0.578 Sum_probs=27.8
Q ss_pred CCCHHHHHHHHhh--CCCeEEEEeccCCCCCCceEeecccc
Q 037760 363 SMNLKECEAECIR--NCSCRAYANSDITGGGDGCLMWFGDL 401 (471)
Q Consensus 363 ~~~~~~C~~~Cl~--nCsC~a~~y~~~~~~g~gC~~w~~~l 401 (471)
..++++|.+.|++ +=.|.+|.|.. ....|.+-..+.
T Consensus 24 ~~s~~~C~~~C~~~~~f~CrSf~y~~---~~~~C~L~~~~~ 61 (80)
T cd01099 24 VASLEECLRKCLEETEFTCRSFNYNY---KSKECILSDEDR 61 (80)
T ss_pred cCCHHHHHHHhCCCCCceEeEEEEEc---CCCEEEEeCCCc
Confidence 4689999999999 89999999974 133598754443
No 26
>PF07645 EGF_CA: Calcium-binding EGF domain; InterPro: IPR001881 A sequence of about forty amino-acid residues found in epidermal growth factor (EGF) has been shown [, , , , , ] to be present in a large number of membrane-bound and extracellular, mostly animal, proteins. Many of these proteins require calcium for their biological function and a calcium-binding site has been found at the N terminus of some EGF-like domains []. Calcium-binding may be crucial for numerous protein-protein interactions. For human coagulation factor IX it has been shown [] that the calcium-ligands form a pentagonal bipyramid. The first, third and fourth conserved negatively charged or polar residues are side chain ligands. The latter is possibly hydroxylated (see aspartic acid and asparagine hydroxylation site) []. A conserved aromatic residue, as well as the second conserved negative residue, are thought to be involved in stabilising the calcium-binding site. As in non-calcium binding EGF-like domains, there are six conserved cysteines and the structure of both types is very similar as calcium-binding induces only strictly local structural changes []. +------------------+ +---------+ | | | | nxnnC-x(3,14)-C-x(3,7)-CxxbxxxxaxC-x(1,6)-C-x(8,13)-Cx | | +------------------+ 'n': negatively charged or polar residue [DEQN] 'b': possibly beta-hydroxylated residue [DN] 'a': aromatic amino acid 'C': cysteine, involved in disulphide bond 'x': any amino acid. ; GO: 0005509 calcium ion binding; PDB: 2VJ3_A 1TOZ_A 1LMJ_A 1UZQ_A 1UZK_A 1UZJ_B 1UZP_A 1EMO_A 1EMN_A 2RR0_A ....
Probab=75.45 E-value=1.2 Score=30.91 Aligned_cols=31 Identities=26% Similarity=0.636 Sum_probs=23.3
Q ss_pred CCCccc-CCCCCCccccCC-CCCccccCCCCcc
Q 037760 289 DACDNY-AQCGANSNCRIS-KTPICECLAGFIS 319 (471)
Q Consensus 289 ~~C~~~-g~CG~~g~C~~~-~~~~C~C~~GF~~ 319 (471)
|+|... ..|..++.|... ++-.|.|++||+.
T Consensus 3 dEC~~~~~~C~~~~~C~N~~Gsy~C~C~~Gy~~ 35 (42)
T PF07645_consen 3 DECAEGPHNCPENGTCVNTEGSYSCSCPPGYEL 35 (42)
T ss_dssp STTTTTSSSSSTTSEEEEETTEEEEEESTTEEE
T ss_pred cccCCCCCcCCCCCEEEcCCCCEEeeCCCCcEE
Confidence 567764 479889999743 3348999999984
No 27
>PHA03265 envelope glycoprotein D; Provisional
Probab=74.59 E-value=2.5 Score=42.81 Aligned_cols=28 Identities=21% Similarity=0.482 Sum_probs=17.9
Q ss_pred ceeEEEEehhHHHHHHHHHHHHhhhhhhc
Q 037760 441 RLKIIVAMSIISGMLILGLLLGMAWKKAK 469 (471)
Q Consensus 441 ~~~~ii~~~v~~~~~~~~~~~~~~~~~~~ 469 (471)
.+.++|+..|+. ++++|+++|++|||||
T Consensus 349 ~~g~~ig~~i~g-lv~vg~il~~~~rr~k 376 (402)
T PHA03265 349 FVGISVGLGIAG-LVLVGVILYVCLRRKK 376 (402)
T ss_pred ccceEEccchhh-hhhhhHHHHHHhhhhh
Confidence 345556554432 4567888898888774
No 28
>smart00179 EGF_CA Calcium-binding EGF-like domain.
Probab=72.10 E-value=3.6 Score=27.11 Aligned_cols=30 Identities=27% Similarity=0.689 Sum_probs=21.1
Q ss_pred CCCcccCCCCCCccccCCC-CCccccCCCCc
Q 037760 289 DACDNYAQCGANSNCRISK-TPICECLAGFI 318 (471)
Q Consensus 289 ~~C~~~g~CG~~g~C~~~~-~~~C~C~~GF~ 318 (471)
+.|.....|...+.|.... ...|.|++||.
T Consensus 3 ~~C~~~~~C~~~~~C~~~~g~~~C~C~~g~~ 33 (39)
T smart00179 3 DECASGNPCQNGGTCVNTVGSYRCECPPGYT 33 (39)
T ss_pred ccCcCCCCcCCCCEeECCCCCeEeECCCCCc
Confidence 4565545687778897443 34799999996
No 29
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=71.64 E-value=3.1 Score=34.60 Aligned_cols=8 Identities=13% Similarity=0.189 Sum_probs=4.6
Q ss_pred cCCCCccC
Q 037760 313 CLAGFISK 320 (471)
Q Consensus 313 C~~GF~~~ 320 (471)
|.+|+.|.
T Consensus 9 C~~g~~~~ 16 (96)
T PTZ00382 9 CDSDKKPN 16 (96)
T ss_pred CCCCCccC
Confidence 55666553
No 30
>PF14610 DUF4448: Protein of unknown function (DUF4448)
Probab=71.44 E-value=3.4 Score=38.61 Aligned_cols=28 Identities=18% Similarity=0.393 Sum_probs=15.1
Q ss_pred EEEEehhHHHHHHHHHHHHhhhhhhcCC
Q 037760 444 IIVAMSIISGMLILGLLLGMAWKKAKNK 471 (471)
Q Consensus 444 ~ii~~~v~~~~~~~~~~~~~~~~~~~~~ 471 (471)
+.|++-+++++++++++++++|+||+||
T Consensus 160 laI~lPvvv~~~~~~~~~~~~~~R~~Rr 187 (189)
T PF14610_consen 160 LAIALPVVVVVLALIMYGFFFWNRKKRR 187 (189)
T ss_pred EEEEccHHHHHHHHHHHhhheeecccee
Confidence 3444444444455566666667666554
No 31
>PF09064 Tme5_EGF_like: Thrombomodulin like fifth domain, EGF-like; InterPro: IPR015149 This domain adopts a fold similar to other EGF domains, with a flat major and a twisted minor beta sheet. Disulphide pairing, however, is not of the usual 1-3, 2-4, 5-6 type; rather 1-2, 3-4, 5-6 pairing is found. Its extended major sheet (strands beta-2 and beta-3 and the connecting loop) projects into thrombin's active site groove. This domain is required for interaction of thrombomodulin with thrombin, and subsequent activation of protein-C []. ; GO: 0004888 transmembrane signaling receptor activity, 0016021 integral to membrane
Probab=68.03 E-value=3.2 Score=27.53 Aligned_cols=18 Identities=28% Similarity=0.719 Sum_probs=13.3
Q ss_pred cccCCCCCccccCCCCcc
Q 037760 302 NCRISKTPICECLAGFIS 319 (471)
Q Consensus 302 ~C~~~~~~~C~C~~GF~~ 319 (471)
.|+.+...+|.||.||..
T Consensus 11 ~CDpn~~~~C~CPeGyIl 28 (34)
T PF09064_consen 11 DCDPNSPGQCFCPEGYIL 28 (34)
T ss_pred ccCCCCCCceeCCCceEe
Confidence 465555568999999964
No 32
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=66.72 E-value=1.2e+02 Score=31.29 Aligned_cols=53 Identities=17% Similarity=0.313 Sum_probs=32.6
Q ss_pred ccCCcEEEEeC-CCceEEeecCCCCc-cce------EEEEecCCCEEEEeCcCCCCCceeeeec
Q 037760 93 SNNGSILLLNQ-ERSTIWSSNSSRVL-ETA------VVRLLDSGNLVLRDNVSRSSDEYMWQSF 148 (471)
Q Consensus 93 ~~~GnLvl~d~-~~~~vWss~~~~~~-~~~------~~~Lld~GNlvl~~~~~~~~~~~~WqSF 148 (471)
..+|.|+-+|. .|..+|+....+.. ..+ ......+|.|+-.|.. +++++|+--
T Consensus 127 ~~~g~l~ald~~tG~~~W~~~~~~~~~ssP~v~~~~v~v~~~~g~l~ald~~---tG~~~W~~~ 187 (394)
T PRK11138 127 SEKGQVYALNAEDGEVAWQTKVAGEALSRPVVSDGLVLVHTSNGMLQALNES---DGAVKWTVN 187 (394)
T ss_pred cCCCEEEEEECCCCCCcccccCCCceecCCEEECCEEEEECCCCEEEEEEcc---CCCEeeeec
Confidence 45788887886 68899998754321 111 1122345666666652 577899863
No 33
>PF12877 DUF3827: Domain of unknown function (DUF3827); InterPro: IPR024606 The function of the proteins in this entry is not currently known, but one of the human proteins (Q9HCM3 from SWISSPROT) has been implicated in pilocytic astrocytomas [, , ]. In the majority of cases of pilocytic astrocytomas a tandem duplication produces an in-frame fusion of the gene encoding this protein and the BRAF oncogene. The resulting fusion protein has constitutive BRAF kinase activity and is capable of transforming cells.
Probab=66.29 E-value=6.5 Score=43.05 Aligned_cols=18 Identities=17% Similarity=0.132 Sum_probs=11.2
Q ss_pred cccceeEEEEehhHHHHH
Q 037760 438 KKKRLKIIVAMSIISGML 455 (471)
Q Consensus 438 ~~~~~~~ii~~~v~~~~~ 455 (471)
..++.||||||++.++++
T Consensus 265 ~~~NlWII~gVlvPv~vV 282 (684)
T PF12877_consen 265 PPNNLWIIAGVLVPVLVV 282 (684)
T ss_pred CCCCeEEEehHhHHHHHH
Confidence 345667778777665543
No 34
>cd00054 EGF_CA Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.
Probab=66.14 E-value=5.6 Score=25.69 Aligned_cols=31 Identities=23% Similarity=0.635 Sum_probs=20.8
Q ss_pred CCCcccCCCCCCccccCCC-CCccccCCCCcc
Q 037760 289 DACDNYAQCGANSNCRISK-TPICECLAGFIS 319 (471)
Q Consensus 289 ~~C~~~g~CG~~g~C~~~~-~~~C~C~~GF~~ 319 (471)
+.|.....|...+.|.... ...|.|++||.-
T Consensus 3 ~~C~~~~~C~~~~~C~~~~~~~~C~C~~g~~g 34 (38)
T cd00054 3 DECASGNPCQNGGTCVNTVGSYRCSCPPGYTG 34 (38)
T ss_pred ccCCCCCCcCCCCEeECCCCCeEeECCCCCcC
Confidence 4565435687778887433 347999999853
No 35
>PTZ00046 rifin; Provisional
Probab=65.57 E-value=2.3 Score=43.59 Aligned_cols=27 Identities=11% Similarity=0.270 Sum_probs=14.1
Q ss_pred EEEehhHHHHHH-HHHHHHhhhhhhcCC
Q 037760 445 IVAMSIISGMLI-LGLLLGMAWKKAKNK 471 (471)
Q Consensus 445 ii~~~v~~~~~~-~~~~~~~~~~~~~~~ 471 (471)
|++++|+++|+| +.+++|++.|.||||
T Consensus 317 IiaSiiAIvVIVLIMvIIYLILRYRRKK 344 (358)
T PTZ00046 317 IIASIVAIVVIVLIMVIIYLILRYRRKK 344 (358)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHhhhcc
Confidence 444444444433 455667766655443
No 36
>TIGR01477 RIFIN variant surface antigen, rifin family. This model represents the rifin branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of rifin sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 20 bits.
Probab=64.44 E-value=2.4 Score=43.24 Aligned_cols=27 Identities=15% Similarity=0.278 Sum_probs=13.9
Q ss_pred EEEehhHHHHHH-HHHHHHhhhhhhcCC
Q 037760 445 IVAMSIISGMLI-LGLLLGMAWKKAKNK 471 (471)
Q Consensus 445 ii~~~v~~~~~~-~~~~~~~~~~~~~~~ 471 (471)
|++.+|+++|+| +.+++|++.|.||||
T Consensus 312 IiaSiIAIvvIVLIMvIIYLILRYRRKK 339 (353)
T TIGR01477 312 IIASIIAILIIVLIMVIIYLILRYRRKK 339 (353)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHhhhcc
Confidence 344444444433 455667666655443
No 37
>PF12661 hEGF: Human growth factor-like EGF; PDB: 2YGQ_A 2E26_A 3A7Q_A 2YGP_A 2YGO_A 1HRE_A 1HAE_A 1HAF_A 1HRF_A.
Probab=63.55 E-value=2 Score=22.24 Aligned_cols=9 Identities=33% Similarity=1.173 Sum_probs=6.6
Q ss_pred ccccCCCCc
Q 037760 310 ICECLAGFI 318 (471)
Q Consensus 310 ~C~C~~GF~ 318 (471)
.|.|++||.
T Consensus 1 ~C~C~~G~~ 9 (13)
T PF12661_consen 1 TCQCPPGWT 9 (13)
T ss_dssp EEEE-TTEE
T ss_pred CccCcCCCc
Confidence 499999985
No 38
>PF01683 EB: EB module; InterPro: IPR006149 The EB domain has no known function. It is found in several Caenorhabditis sp. and Drosophila sp. proteins. The domain contains 8 conserved cysteines that probably form four disulphide bridges and is found associated with kunitz domains IPR002223 from INTERPRO
Probab=62.19 E-value=7.4 Score=28.06 Aligned_cols=33 Identities=27% Similarity=0.642 Sum_probs=26.3
Q ss_pred ccCCCCcccCCCCCCccccCCCCCccccCCCCccCC
Q 037760 286 WPFDACDNYAQCGANSNCRISKTPICECLAGFISKP 321 (471)
Q Consensus 286 ~p~~~C~~~g~CG~~g~C~~~~~~~C~C~~GF~~~~ 321 (471)
.|-+.|.....|-.++.|.. ..|.|++||.+..
T Consensus 17 ~~g~~C~~~~qC~~~s~C~~---g~C~C~~g~~~~~ 49 (52)
T PF01683_consen 17 QPGESCESDEQCIGGSVCVN---GRCQCPPGYVEVG 49 (52)
T ss_pred CCCCCCCCcCCCCCcCEEcC---CEeECCCCCEecC
Confidence 35567998889999999953 5899999997654
No 39
>PF07974 EGF_2: EGF-like domain; InterPro: IPR013111 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length. This entry contains EGF domains found in a variety of extracellular and membrane proteins
Probab=60.71 E-value=7.6 Score=25.40 Aligned_cols=23 Identities=22% Similarity=0.655 Sum_probs=17.9
Q ss_pred CCCCCCccccCCCCCccccCCCCc
Q 037760 295 AQCGANSNCRISKTPICECLAGFI 318 (471)
Q Consensus 295 g~CG~~g~C~~~~~~~C~C~~GF~ 318 (471)
..|...|.|... ..+|.|.+||.
T Consensus 6 ~~C~~~G~C~~~-~g~C~C~~g~~ 28 (32)
T PF07974_consen 6 NICSGHGTCVSP-CGRCVCDSGYT 28 (32)
T ss_pred CccCCCCEEeCC-CCEEECCCCCc
Confidence 468888899743 45899999985
No 40
>PF12947 EGF_3: EGF domain; InterPro: IPR024731 This entry represents an EGF domain found in the the C terminus of malarial parasite merozoite surface protein 1 [], as well as other proteins.; PDB: 2NPR_A 1N1I_C 1B9W_A 1YO8_A 2RHP_A.
Probab=59.18 E-value=2.7 Score=28.29 Aligned_cols=26 Identities=23% Similarity=0.655 Sum_probs=17.2
Q ss_pred cCCCCCCccccCCC-CCccccCCCCcc
Q 037760 294 YAQCGANSNCRISK-TPICECLAGFIS 319 (471)
Q Consensus 294 ~g~CG~~g~C~~~~-~~~C~C~~GF~~ 319 (471)
.+-|.++..|.... .-.|.|.+||.-
T Consensus 5 ~~~C~~nA~C~~~~~~~~C~C~~Gy~G 31 (36)
T PF12947_consen 5 NGGCHPNATCTNTGGSYTCTCKPGYEG 31 (36)
T ss_dssp GGGS-TTCEEEE-TTSEEEEE-CEEEC
T ss_pred CCCCCCCcEeecCCCCEEeECCCCCcc
Confidence 35788899998543 348999999963
No 41
>PF01034 Syndecan: Syndecan domain; InterPro: IPR001050 The syndecans are transmembrane proteoglycans which are involved in the organisation of cytoskeleton and/or actin microfilaments, and have important roles as cell surface receptors during cell-cell and/or cell-matrix interactions [, ]. Structurally, these proteins consist of four separate domains: A signal sequence; An extracellular domain (ectodomain) of variable length whose sequence is not evolutionary conserved in the various forms of syndecans. The ectodomain contains the sites of attachment of the heparan sulphate glycosaminoglycan side chains; A transmembrane region; A highly conserved cytoplasmic domain of about 30 to 35 residues, which could interact with cytoskeletal proteins. The proteins known to belong to this family are: Syndecan 1. Syndecan 2 or fibroglycan. Syndecan 3 or neuroglycan or N-syndecan. Syndecan 4 or amphiglycan or ryudocan. Drosophila syndecan. Caenorhabditis elegans probable syndecan (F57C7.3). Syndecan-4, a transmembrane heparan sulphate proteoglycan, is a coreceptor with integrins in cell adhesion. It has been suggested to form a ternary signalling complex with protein kinase Calpha and phosphatidylinositol 4,5-bisphosphate (PIP2). Structural studies have demonstrated that the cytoplasmic domain undergoes a conformational transition and forms a symmetric dimer in the presence of phospholipid activator PIP2, and whose overall structure in solution exhibits a twisted clamp shape having a cavity in the centre of dimeric interface. In addition, it has been observed that the syndecan-4 variable domain interacts, strongly, not only with fatty acyl groups but also the anionic head group of PIP2. These findings indicate that PIP2 promotes oligomerisation of the syndecan-4 cytoplasmic domain for transmembrane signalling and cell-matrix adhesion [, ].; GO: 0008092 cytoskeletal protein binding, 0016020 membrane; PDB: 1EJQ_B 1EJP_B 1YBO_C 1OBY_Q.
Probab=58.60 E-value=3.2 Score=31.68 Aligned_cols=14 Identities=21% Similarity=0.354 Sum_probs=0.5
Q ss_pred HHHHHHhhhhhhcC
Q 037760 457 LGLLLGMAWKKAKN 470 (471)
Q Consensus 457 ~~~~~~~~~~~~~~ 470 (471)
+.++++++.|.|||
T Consensus 26 ilLIlf~iyR~rkk 39 (64)
T PF01034_consen 26 ILLILFLIYRMRKK 39 (64)
T ss_dssp -----------S--
T ss_pred HHHHHHHHHHHHhc
Confidence 33444455554443
No 42
>PF05454 DAG1: Dystroglycan (Dystrophin-associated glycoprotein 1); InterPro: IPR008465 Dystroglycan is one of the dystrophin-associated glycoproteins, which is encoded by a 5.5 kb transcript in Homo sapiens. The protein product is cleaved into two non-covalently associated subunits, [alpha] (N-terminal) and [beta] (C-terminal). In skeletal muscle the dystroglycan complex works as a transmembrane linkage between the extracellular matrix and the cytoskeleton [alpha]-dystroglycan is extracellular and binds to merosin ([alpha]-2 laminin) in the basement membrane, while [beta]-dystroglycan is a transmembrane protein and binds to dystrophin, which is a large rod-like cytoskeletal protein, absent in Duchenne muscular dystrophy patients. Dystrophin binds to intracellular actin cables. In this way, the dystroglycan complex, which links the extracellular matrix to the intracellular actin cables, is thought to provide structural integrity in muscle tissues. The dystroglycan complex is also known to serve as an agrin receptor in muscle, where it may regulate agrin-induced acetylcholine receptor clustering at the neuromuscular junction. There is also evidence which suggests the function of dystroglycan as a part of the signal transduction pathway because it is shown that Grb2, a mediator of the Ras-related signal pathway, can interact with the cytoplasmic domain of dystroglycan. In general, aberrant expression of dystrophin-associated protein complex underlies the pathogenesis of Duchenne muscular dystrophy, Becker muscular dystrophy and severe childhood autosomal recessive muscular dystrophy. Interestingly, no genetic disease has been described for either [alpha]- or [beta]-dystroglycan. Dystroglycan is widely distributed in non-muscle tissues as well as in muscle tissues. During epithelial morphogenesis of kidney, the dystroglycan complex is shown to act as a receptor for the basement membrane. Dystroglycan expression in Mus musculus brain and neural retina has also been reported. However, the physiological role of dystroglycan in non-muscle tissues has remained unclear [].; PDB: 1EG4_P.
Probab=56.41 E-value=3.7 Score=41.07 Aligned_cols=24 Identities=25% Similarity=0.497 Sum_probs=0.0
Q ss_pred EEEehhHHHHHHHHHHHHhhhhhh
Q 037760 445 IVAMSIISGMLILGLLLGMAWKKA 468 (471)
Q Consensus 445 ii~~~v~~~~~~~~~~~~~~~~~~ 468 (471)
|++++|+++++|.++++++++|||
T Consensus 150 IpaVVI~~iLLIA~iIa~icyrrk 173 (290)
T PF05454_consen 150 IPAVVIAAILLIAGIIACICYRRK 173 (290)
T ss_dssp ------------------------
T ss_pred HHHHHHHHHHHHHHHHHHHhhhhh
Confidence 344444443444444444444433
No 43
>PF01299 Lamp: Lysosome-associated membrane glycoprotein (Lamp); InterPro: IPR002000 Lysosome-associated membrane glycoproteins (lamp) [] are integral membrane proteins, specific to lysosomes, and whose exact biological function is not yet clear. Structurally, the lamp proteins consist of two internally homologous lysosome-luminal domains separated by a proline-rich hinge region; at the C-terminal extremity there is a transmembrane region (TM) followed by a very short cytoplasmic tail (C). In each of the duplicated domains, there are two conserved disulphide bonds. This structure is schematically represented in the figure below. +-----+ +-----+ +-----+ +-----+ | | | | | | | | xCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxxxCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxx +--------------------------++Hinge++--------------------------++TM++C+ In mammals, there are two closely related types of lamp: lamp-1 and lamp-2, which form major components of the lysosome membrane. In chicken lamp-1 is known as LEP100. Also included in this entry is the macrophage protein CD68 (or macrosialin) [] is a heavily glycosylated integral membrane protein whose structure consists of a mucin-like domain followed by a proline-rich hinge; a single lamp-like domain; a transmembrane region and a short cytoplasmic tail. Similar to CD68, mammalian lamp-3, which is expressed in lymphoid organs, dendritic cells and in lung, contains all the C-terminal regions but lacks the N-terminal lamp-like region []. In a lamp-family protein from nematodes [] only the part C-terminal to the hinge is conserved. ; GO: 0016020 membrane
Probab=53.45 E-value=7.9 Score=39.04 Aligned_cols=27 Identities=7% Similarity=0.226 Sum_probs=14.7
Q ss_pred EEEEehhHHHH---HHHHHHHHhhhhhhcC
Q 037760 444 IIVAMSIISGM---LILGLLLGMAWKKAKN 470 (471)
Q Consensus 444 ~ii~~~v~~~~---~~~~~~~~~~~~~~~~ 470 (471)
.+|.++||+++ +++.++.|++.|||++
T Consensus 271 ~~vPIaVG~~La~lvlivLiaYli~Rrr~~ 300 (306)
T PF01299_consen 271 DLVPIAVGAALAGLVLIVLIAYLIGRRRSR 300 (306)
T ss_pred chHHHHHHHHHHHHHHHHHHhheeEecccc
Confidence 45555555443 2344556666676654
No 44
>PF00008 EGF: EGF-like domain This is a sub-family of the Pfam entry This is a sub-family of the Pfam entry; InterPro: IPR006209 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length.; GO: 0005515 protein binding; PDB: 1WHE_A 1CCF_A 1APO_A 1WHF_A 2VJ3_A 1TOZ_A 4D90_B 3CFW_A 1EDM_B 1IXA_A ....
Probab=53.43 E-value=4.1 Score=26.45 Aligned_cols=23 Identities=26% Similarity=0.642 Sum_probs=17.3
Q ss_pred CCCCCccccCC--CCCccccCCCCc
Q 037760 296 QCGANSNCRIS--KTPICECLAGFI 318 (471)
Q Consensus 296 ~CG~~g~C~~~--~~~~C~C~~GF~ 318 (471)
.|...|.|... ....|.|++||.
T Consensus 5 ~C~n~g~C~~~~~~~y~C~C~~G~~ 29 (32)
T PF00008_consen 5 PCQNGGTCIDLPGGGYTCECPPGYT 29 (32)
T ss_dssp SSTTTEEEEEESTSEEEEEEBTTEE
T ss_pred cCCCCeEEEeCCCCCEEeECCCCCc
Confidence 67778888743 335899999985
No 45
>PF06697 DUF1191: Protein of unknown function (DUF1191); InterPro: IPR010605 This family contains hypothetical plant proteins of unknown function.
Probab=52.99 E-value=12 Score=37.01 Aligned_cols=22 Identities=27% Similarity=0.221 Sum_probs=9.4
Q ss_pred EEEehhHHHHHH-HHHHHHhhhh
Q 037760 445 IVAMSIISGMLI-LGLLLGMAWK 466 (471)
Q Consensus 445 ii~~~v~~~~~~-~~~~~~~~~~ 466 (471)
++++++|+++++ +++++++..|
T Consensus 216 v~g~~~G~~~L~ll~~lv~~~vr 238 (278)
T PF06697_consen 216 VVGVVGGVVLLGLLSLLVAMLVR 238 (278)
T ss_pred EEEehHHHHHHHHHHHHHHhhhh
Confidence 444455554433 3333333333
No 46
>PF12662 cEGF: Complement Clr-like EGF-like
Probab=52.92 E-value=6.5 Score=24.04 Aligned_cols=11 Identities=27% Similarity=0.860 Sum_probs=9.4
Q ss_pred ccccCCCCccC
Q 037760 310 ICECLAGFISK 320 (471)
Q Consensus 310 ~C~C~~GF~~~ 320 (471)
.|+|++||+..
T Consensus 3 ~C~C~~Gy~l~ 13 (24)
T PF12662_consen 3 TCSCPPGYQLS 13 (24)
T ss_pred EeeCCCCCcCC
Confidence 69999999864
No 47
>PF13908 Shisa: Wnt and FGF inhibitory regulator
Probab=50.89 E-value=12 Score=34.49 Aligned_cols=17 Identities=18% Similarity=0.098 Sum_probs=8.2
Q ss_pred eEEEEehhHHHHHHHHH
Q 037760 443 KIIVAMSIISGMLILGL 459 (471)
Q Consensus 443 ~~ii~~~v~~~~~~~~~ 459 (471)
.++++|+++++++|+++
T Consensus 79 ~iivgvi~~Vi~Iv~~I 95 (179)
T PF13908_consen 79 GIIVGVICGVIAIVVLI 95 (179)
T ss_pred eeeeehhhHHHHHHHhH
Confidence 45555555544444333
No 48
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=50.85 E-value=14 Score=30.61 Aligned_cols=10 Identities=20% Similarity=0.292 Sum_probs=4.7
Q ss_pred eEEEEehhHH
Q 037760 443 KIIVAMSIIS 452 (471)
Q Consensus 443 ~~ii~~~v~~ 452 (471)
..|++++|++
T Consensus 66 gaiagi~vg~ 75 (96)
T PTZ00382 66 GAIAGISVAV 75 (96)
T ss_pred ccEEEEEeeh
Confidence 3455555443
No 49
>PF08374 Protocadherin: Protocadherin; InterPro: IPR013585 The structure of protocadherins is similar to that of classic cadherins (IPR002126 from INTERPRO), but they also have some unique features associated with the cytoplasmic domains. They are expressed in a variety of organisms and are found in high concentrations in the brain where they seem to be localised mainly at cell-cell contact sites. Their expression seems to be developmentally regulated [].
Probab=49.71 E-value=14 Score=35.12 Aligned_cols=15 Identities=27% Similarity=0.219 Sum_probs=9.4
Q ss_pred ccceeEEEEehhHHH
Q 037760 439 KKRLKIIVAMSIISG 453 (471)
Q Consensus 439 ~~~~~~ii~~~v~~~ 453 (471)
+.+.+|+|+++.|++
T Consensus 34 ~d~~~I~iaiVAG~~ 48 (221)
T PF08374_consen 34 KDYVKIMIAIVAGIM 48 (221)
T ss_pred ccceeeeeeeecchh
Confidence 445667777776544
No 50
>PF03302 VSP: Giardia variant-specific surface protein; InterPro: IPR005127 During infection, the intestinal protozoan parasite Giardia lamblia virus undergoes continuous antigenic variation which is determined by diversification of the parasite's major surface antigen, named VSP (variant surface protein).
Probab=45.97 E-value=17 Score=38.19 Aligned_cols=24 Identities=17% Similarity=0.134 Sum_probs=13.7
Q ss_pred cceeEEEEehhHHHHHHHHHHHHh
Q 037760 440 KRLKIIVAMSIISGMLILGLLLGM 463 (471)
Q Consensus 440 ~~~~~ii~~~v~~~~~~~~~~~~~ 463 (471)
-....|.+|+|+++|+|-+++-||
T Consensus 364 LstgaIaGIsvavvvvVgglvGfL 387 (397)
T PF03302_consen 364 LSTGAIAGISVAVVVVVGGLVGFL 387 (397)
T ss_pred ccccceeeeeehhHHHHHHHHHHH
Confidence 345677777777665553343333
No 51
>PF01436 NHL: NHL repeat; InterPro: IPR001258 The NHL repeat, named after NCL-1, HT2A and Lin-41, is found largely in a large number of eukaryotic and prokaryotic proteins. For example, the repeat is found in a variety of enzymes of the copper type II, ascorbate-dependent monooxygenase family which catalyse the C terminus alpha-amidation of biological peptides []. In many it occurs in tandem arrays, for example in the ringfinger beta-box, coiled-coil (RBCC) eukaryotic growth regulators []. The 'Brain Tumor' protein (Brat) is one such growth regulator that contains a 6-bladed NHL-repeat beta-propeller [, ]. The NHL repeats are also found in serine/threonine protein kinase (STPK) in diverse range of pathogenic bacteria. These STPK are transmembrane receptors with a intracellular N-terminal kinase domain and extracellular C-terminal sensor domain. In the STPK, PknD, from Mycobacterium tuberculosis, the sensor domain forms a rigid, six-bladed b-propeller composed of NHL repeats with a flexible tether to the transmembrane domain.; GO: 0005515 protein binding; PDB: 3FVZ_A 3FW0_A 1RWL_A 1RWI_A 1Q7F_A.
Probab=45.14 E-value=34 Score=21.26 Aligned_cols=21 Identities=10% Similarity=0.295 Sum_probs=15.2
Q ss_pred eEEEccCCcEEEEeCCCceEE
Q 037760 89 VLTLSNNGSILLLNQERSTIW 109 (471)
Q Consensus 89 ~l~l~~~GnLvl~d~~~~~vW 109 (471)
-+.++.+|+|++.|.++..||
T Consensus 6 gvav~~~g~i~VaD~~n~rV~ 26 (28)
T PF01436_consen 6 GVAVDSDGNIYVADSGNHRVQ 26 (28)
T ss_dssp EEEEETTSEEEEEECCCTEEE
T ss_pred EEEEeCCCCEEEEECCCCEEE
Confidence 366678888888887665555
No 52
>smart00181 EGF Epidermal growth factor-like domain.
Probab=44.33 E-value=21 Score=22.88 Aligned_cols=24 Identities=21% Similarity=0.695 Sum_probs=16.7
Q ss_pred CCCCCCccccCC-CCCccccCCCCcc
Q 037760 295 AQCGANSNCRIS-KTPICECLAGFIS 319 (471)
Q Consensus 295 g~CG~~g~C~~~-~~~~C~C~~GF~~ 319 (471)
..|... .|... ....|.|++||.-
T Consensus 6 ~~C~~~-~C~~~~~~~~C~C~~g~~g 30 (35)
T smart00181 6 GPCSNG-TCINTPGSYTCSCPPGYTG 30 (35)
T ss_pred CCCCCC-EEECCCCCeEeECCCCCcc
Confidence 456666 78643 3458999999964
No 53
>TIGR01478 STEVOR variant surface antigen, stevor family. This model represents the stevor branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of stevor sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 8 bits.
Probab=42.77 E-value=11 Score=37.25 Aligned_cols=22 Identities=14% Similarity=0.305 Sum_probs=10.1
Q ss_pred CCCCCCCCCCCCCCCceEEecc
Q 037760 330 SRRCDRKPSDCPSGEGFLKLQR 351 (471)
Q Consensus 330 s~GC~~~~~~C~~~~~f~~~~~ 351 (471)
-.+|.+.--.|.-+.-|+.+-|
T Consensus 169 K~rC~~gi~~CsvGSA~LT~IG 190 (295)
T TIGR01478 169 KKGCTAGVGTCALSSALLGNIG 190 (295)
T ss_pred hccCCCeeEeeccHHHHHHHHH
Confidence 3577665223543333444333
No 54
>PTZ00370 STEVOR; Provisional
Probab=40.74 E-value=13 Score=36.98 Aligned_cols=22 Identities=27% Similarity=0.563 Sum_probs=10.2
Q ss_pred CCCCCCCCCCCCCCCceEEecc
Q 037760 330 SRRCDRKPSDCPSGEGFLKLQR 351 (471)
Q Consensus 330 s~GC~~~~~~C~~~~~f~~~~~ 351 (471)
-.+|.+.--.|.-+.-|+.+-|
T Consensus 169 K~rC~~gi~~CsVGSafLT~IG 190 (296)
T PTZ00370 169 KHRCTGGICSCSLGSALLTLIG 190 (296)
T ss_pred hccCCCeeEeeccHHHHHHHHH
Confidence 3577665223543334444434
No 55
>TIGR01167 LPXTG_anchor LPXTG-motif cell wall anchor domain. A common feature of this proteins containing this domain appears to be a high proportion of charged and zwitterionic residues immediatedly upstream of the LPXTG motif. This model differs from other descriptions of the LPXTG region by including a portion of that upstream charged region.
Probab=36.08 E-value=36 Score=21.96 Aligned_cols=10 Identities=20% Similarity=-0.147 Sum_probs=4.4
Q ss_pred HHHhhhhhhc
Q 037760 460 LLGMAWKKAK 469 (471)
Q Consensus 460 ~~~~~~~~~~ 469 (471)
..++++|||+
T Consensus 24 ~~~~~~~rk~ 33 (34)
T TIGR01167 24 GGLLLRKRKK 33 (34)
T ss_pred HHHHheeccc
Confidence 3344444444
No 56
>PF13360 PQQ_2: PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=35.68 E-value=81 Score=29.35 Aligned_cols=51 Identities=20% Similarity=0.382 Sum_probs=26.9
Q ss_pred CCcEEEEeC-CCceEEeecCCCCccceEE-EE---------ecCCCEEEEeCcCCCCCceeeeec
Q 037760 95 NGSILLLNQ-ERSTIWSSNSSRVLETAVV-RL---------LDSGNLVLRDNVSRSSDEYMWQSF 148 (471)
Q Consensus 95 ~GnLvl~d~-~~~~vWss~~~~~~~~~~~-~L---------ld~GNlvl~~~~~~~~~~~~WqSF 148 (471)
+|.|...|. +|..+|+...........+ .+ ..+|+|+..|.. +++++|+--
T Consensus 2 ~g~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~~~~v~~~~~~~~l~~~d~~---tG~~~W~~~ 63 (238)
T PF13360_consen 2 DGTLSALDPRTGKELWSYDLGPGIGGPVATAVPDGGRVYVASGDGNLYALDAK---TGKVLWRFD 63 (238)
T ss_dssp TSEEEEEETTTTEEEEEEECSSSCSSEEETEEEETTEEEEEETTSEEEEEETT---TSEEEEEEE
T ss_pred CCEEEEEECCCCCEEEEEECCCCCCCccceEEEeCCEEEEEcCCCEEEEEECC---CCCEEEEee
Confidence 566777775 6777787754110111111 22 255555555532 567788754
No 57
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=35.01 E-value=1.7e+02 Score=30.13 Aligned_cols=20 Identities=20% Similarity=0.599 Sum_probs=10.2
Q ss_pred ccCCcEEEEeC-CCceEEeec
Q 037760 93 SNNGSILLLNQ-ERSTIWSSN 112 (471)
Q Consensus 93 ~~~GnLvl~d~-~~~~vWss~ 112 (471)
..+|.|+-+|. +|+++|...
T Consensus 263 ~~~g~l~ald~~tG~~~W~~~ 283 (394)
T PRK11138 263 AYNGNLVALDLRSGQIVWKRE 283 (394)
T ss_pred EcCCeEEEEECCCCCEEEeec
Confidence 34555555554 345566543
No 58
>cd05845 Ig2_L1-CAM_like Second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM) and similar proteins. Ig2_L1-CAM_like: domain similar to the second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM). L1 belongs to the L1 subfamily of cell adhesion molecules (CAMs) and is comprised of an extracellular region having six Ig-like domains, five fibronectin type III domains, a transmembrane region and an intracellular domain. L1 is primarily expressed in the nervous system and is involved in its development and function. L1 is associated with an X-linked recessive disorder, X-linked hydrocephalus, MASA syndrome, or spastic paraplegia type 1, that involves abnormalities of axonal growth.
Probab=32.72 E-value=77 Score=26.16 Aligned_cols=33 Identities=21% Similarity=0.530 Sum_probs=22.3
Q ss_pred CCCeEEEEecCCCCCCCCCceEEEccCCcEEEEe
Q 037760 69 SPRTVVWVANRYKPITDKNGVLTLSNNGSILLLN 102 (471)
Q Consensus 69 ~~~~~vW~an~~~pv~~~~~~l~l~~~GnLvl~d 102 (471)
|..++.|.-+....+. ....+.++.+|||.+.+
T Consensus 32 P~P~i~W~~~~~~~i~-~~~Ri~~~~~GnL~fs~ 64 (95)
T cd05845 32 VPLRIYWMNSDLLHIT-QDERVSMGQNGNLYFAN 64 (95)
T ss_pred CCCEEEEECCCCcccc-ccccEEECCCceEEEEE
Confidence 5667889855434443 35678888889998854
No 59
>PF15345 TMEM51: Transmembrane protein 51
Probab=32.62 E-value=50 Score=31.85 Aligned_cols=15 Identities=27% Similarity=0.483 Sum_probs=7.1
Q ss_pred HHHHHHHHHhhhhhh
Q 037760 454 MLILGLLLGMAWKKA 468 (471)
Q Consensus 454 ~~~~~~~~~~~~~~~ 468 (471)
++++++|+-+.-|||
T Consensus 71 LLLLSICL~IR~KRr 85 (233)
T PF15345_consen 71 LLLLSICLSIRDKRR 85 (233)
T ss_pred HHHHHHHHHHHHHHH
Confidence 344566554433333
No 60
>PF13360 PQQ_2: PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=32.49 E-value=2.3e+02 Score=26.22 Aligned_cols=75 Identities=17% Similarity=0.363 Sum_probs=43.0
Q ss_pred CCeEEEEecC----CCCC---CCCCceEEE-ccCCcEEEEeC-CCceEEeecCCCCccce-------EEEEecCCCEEEE
Q 037760 70 PRTVVWVANR----YKPI---TDKNGVLTL-SNNGSILLLNQ-ERSTIWSSNSSRVLETA-------VVRLLDSGNLVLR 133 (471)
Q Consensus 70 ~~~~vW~an~----~~pv---~~~~~~l~l-~~~GnLvl~d~-~~~~vWss~~~~~~~~~-------~~~Lld~GNlvl~ 133 (471)
....+|..+- ..++ ...+..+.+ +.+|.|+.+|. .|..+|+....+....+ ......+|.|...
T Consensus 12 tG~~~W~~~~~~~~~~~~~~~~~~~~~v~~~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~~ 91 (238)
T PF13360_consen 12 TGKELWSYDLGPGIGGPVATAVPDGGRVYVASGDGNLYALDAKTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYAL 91 (238)
T ss_dssp TTEEEEEEECSSSCSSEEETEEEETTEEEEEETTSEEEEEETTTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEEE
T ss_pred CCCEEEEEECCCCCCCccceEEEeCCEEEEEcCCCEEEEEECCCCCEEEEeeccccccceeeecccccccccceeeeEec
Confidence 5678887653 2221 101233333 47889999996 88999998764321111 1122234456666
Q ss_pred eCcCCCCCceeeee
Q 037760 134 DNVSRSSDEYMWQS 147 (471)
Q Consensus 134 ~~~~~~~~~~~WqS 147 (471)
|. .++.++|+.
T Consensus 92 d~---~tG~~~W~~ 102 (238)
T PF13360_consen 92 DA---KTGKVLWSI 102 (238)
T ss_dssp ET---TTSCEEEEE
T ss_pred cc---CCcceeeee
Confidence 63 267899995
No 61
>PF10681 Rot1: Chaperone for protein-folding within the ER, fungal; InterPro: IPR019623 This conserved fungal family is an essential molecular chaperone in the endoplasmic reticulum. Molecular chaperones transiently interact with unfolded proteins to inhibit their self-aggregation and to support their folding and/or assembly. Rot1 is a general chaperone with some substrate specificity, its substrates being the structurally unrelated Kre5 Kre6 Big1 Atg22, which are type I, type II, and polytopic membrane proteins. The dependencies of each for Rot1 do not share similarities. However, their folding does require BiP, and one of these proteins was simultaneously associated with both Rot1 and BiP. In addition, Rot1 may cooperate with BiP/Kar2 in the folding of Kre6 [].
Probab=30.64 E-value=2.7e+02 Score=26.54 Aligned_cols=74 Identities=16% Similarity=0.264 Sum_probs=44.8
Q ss_pred ceEEeecCCCCccceEEEEecCCCEEEEeCcCCCCCceeeeeccCCCCCCCCCCeeeeeccCCceeEEEEecCCCCCCCc
Q 037760 106 STIWSSNSSRVLETAVVRLLDSGNLVLRDNVSRSSDEYMWQSFDYPSDTLLPGMKLGWNLRTRFERYLTAWRNADDPTPG 185 (471)
Q Consensus 106 ~~vWss~~~~~~~~~~~~Lld~GNlvl~~~~~~~~~~~~WqSFd~PTDTLLPGq~L~~~~~tg~~~~L~Sw~s~~dps~G 185 (471)
..+|+-. .-+|+++|.|+|.-- ..+++ |-+.+|...= -..--..+ +...+.+|.-..|+-.|
T Consensus 64 ~l~wQHG--------tY~l~~nGsl~L~P~--~~DGr---Ql~sdPC~~~-~s~y~rYn----q~e~f~~~~v~~D~y~~ 125 (212)
T PF10681_consen 64 VLIWQHG--------TYELNSNGSLTLTPF--AVDGR---QLVSDPCADD-SSTYTRYN----QTELFKSFDVYVDPYHG 125 (212)
T ss_pred EEEEecc--------eEEECCCCcEEEeec--CCCCc---eeccCCCCCC-cccEEEEc----ceEEEEEEEEEEeCCCC
Confidence 3567743 347788999998754 12444 4456776522 11111222 33567778777899899
Q ss_pred eEEEEE-ccCCCe
Q 037760 186 EFSFRF-DISTMA 197 (471)
Q Consensus 186 ~f~l~l-~~~g~~ 197 (471)
.|+|.| +.+|.|
T Consensus 126 ~~~L~L~~fDGsp 138 (212)
T PF10681_consen 126 RYRLQLYQFDGSP 138 (212)
T ss_pred eeEEEEEccCCCc
Confidence 999988 356643
No 62
>PF01102 Glycophorin_A: Glycophorin A; InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=29.99 E-value=15 Score=31.96 Aligned_cols=27 Identities=11% Similarity=0.218 Sum_probs=18.5
Q ss_pred EEEehhHHHHHHHHHHHHhhhhhhcCC
Q 037760 445 IVAMSIISGMLILGLLLGMAWKKAKNK 471 (471)
Q Consensus 445 ii~~~v~~~~~~~~~~~~~~~~~~~~~ 471 (471)
|+++++|+++-++++++++.+..||+|
T Consensus 66 i~~Ii~gv~aGvIg~Illi~y~irR~~ 92 (122)
T PF01102_consen 66 IIGIIFGVMAGVIGIILLISYCIRRLR 92 (122)
T ss_dssp HHHHHHHHHHHHHHHHHHHHHHHHHHS
T ss_pred eeehhHHHHHHHHHHHHHHHHHHHHHh
Confidence 444555555557788887788888875
No 63
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=29.90 E-value=2.1e+02 Score=29.03 Aligned_cols=20 Identities=15% Similarity=0.606 Sum_probs=10.4
Q ss_pred cCCcEEEEe-CCCceEEeecC
Q 037760 94 NNGSILLLN-QERSTIWSSNS 113 (471)
Q Consensus 94 ~~GnLvl~d-~~~~~vWss~~ 113 (471)
.+|.|+-.| .+|.++|+...
T Consensus 73 ~~g~v~a~d~~tG~~~W~~~~ 93 (377)
T TIGR03300 73 ADGTVVALDAETGKRLWRVDL 93 (377)
T ss_pred CCCeEEEEEccCCcEeeeecC
Confidence 345555555 35556665443
No 64
>PF06365 CD34_antigen: CD34/Podocalyxin family; InterPro: IPR013836 This family consists of several mammalian CD34 antigen proteins. The CD34 antigen is a human leukocyte membrane protein expressed specifically by lymphohematopoietic progenitor cells. CD34 is a phosphoprotein. Activation of protein kinase C (PKC) has been found to enhance CD34 phosphorylation [, ]. This family contains several eukaryotic podocalyxin proteins. Podocalyxin is a major membrane protein of the glomerular epithelium and is thought to be involved in maintenance of the architecture of the foot processes and filtration slits characteristic of this unique epithelium by virtue of its high negative charge. Podocalyxin functions as an anti-adhesin that maintains an open filtration pathway between neighbouring foot processes in the glomerular epithelium by charge repulsion [].
Probab=29.54 E-value=61 Score=30.74 Aligned_cols=26 Identities=12% Similarity=-0.038 Sum_probs=13.1
Q ss_pred EEEEehhHH--H-HHHHHHHHHhhhhhhc
Q 037760 444 IIVAMSIIS--G-MLILGLLLGMAWKKAK 469 (471)
Q Consensus 444 ~ii~~~v~~--~-~~~~~~~~~~~~~~~~ 469 (471)
.+|++++.. . +++++..+|++|.||.
T Consensus 101 ~lI~lv~~g~~lLla~~~~~~Y~~~~Rrs 129 (202)
T PF06365_consen 101 TLIALVTSGSFLLLAILLGAGYCCHQRRS 129 (202)
T ss_pred EEEehHHhhHHHHHHHHHHHHHHhhhhcc
Confidence 555544433 2 2234556677776653
No 65
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=29.05 E-value=2.7e+02 Score=28.32 Aligned_cols=75 Identities=19% Similarity=0.416 Sum_probs=44.0
Q ss_pred CCCeEEEEecCCCC-----CCCCCceEEE-ccCCcEEEEeC-CCceEEeecCCCCc-c------ceEEEEecCCCEEEEe
Q 037760 69 SPRTVVWVANRYKP-----ITDKNGVLTL-SNNGSILLLNQ-ERSTIWSSNSSRVL-E------TAVVRLLDSGNLVLRD 134 (471)
Q Consensus 69 ~~~~~vW~an~~~p-----v~~~~~~l~l-~~~GnLvl~d~-~~~~vWss~~~~~~-~------~~~~~Lld~GNlvl~~ 134 (471)
....++|..+-..+ +.+ +..+.+ +.+|.|+-+|. +|..+|+....+.. . .....-..+|.|+..|
T Consensus 83 ~tG~~~W~~~~~~~~~~~p~v~-~~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~a~d 161 (377)
T TIGR03300 83 ETGKRLWRVDLDERLSGGVGAD-GGLVFVGTEKGEVIALDAEDGKELWRAKLSSEVLSPPLVANGLVVVRTNDGRLTALD 161 (377)
T ss_pred cCCcEeeeecCCCCcccceEEc-CCEEEEEcCCCEEEEEECCCCcEeeeeccCceeecCCEEECCEEEEECCCCeEEEEE
Confidence 45678997654433 222 233434 56788888886 68899987653311 0 1112223566677776
Q ss_pred CcCCCCCceeeee
Q 037760 135 NVSRSSDEYMWQS 147 (471)
Q Consensus 135 ~~~~~~~~~~WqS 147 (471)
.. +++++|+-
T Consensus 162 ~~---tG~~~W~~ 171 (377)
T TIGR03300 162 AA---TGERLWTY 171 (377)
T ss_pred cC---CCceeeEE
Confidence 52 56788984
No 66
>KOG3637 consensus Vitronectin receptor, alpha subunit [Extracellular structures]
Probab=26.80 E-value=48 Score=39.21 Aligned_cols=15 Identities=47% Similarity=1.010 Sum_probs=9.8
Q ss_pred HHHHHHHHHHHhhhh
Q 037760 452 SGMLILGLLLGMAWK 466 (471)
Q Consensus 452 ~~~~~~~~~~~~~~~ 466 (471)
+.+|+++++++++||
T Consensus 987 ~GLLlL~llv~~LwK 1001 (1030)
T KOG3637|consen 987 GGLLLLALLVLLLWK 1001 (1030)
T ss_pred HHHHHHHHHHHHHHh
Confidence 336666777777776
No 67
>KOG4649 consensus PQQ (pyrrolo-quinoline quinone) repeat protein [Secondary metabolites biosynthesis, transport and catabolism]
Probab=25.98 E-value=2.3e+02 Score=28.16 Aligned_cols=46 Identities=17% Similarity=0.485 Sum_probs=33.6
Q ss_pred CCCeEEEEecCCCCCCCCC----ceEEE-ccCCcEEEEeCCCceEEeecCC
Q 037760 69 SPRTVVWVANRYKPITDKN----GVLTL-SNNGSILLLNQERSTIWSSNSS 114 (471)
Q Consensus 69 ~~~~~vW~an~~~pv~~~~----~~l~l-~~~GnLvl~d~~~~~vWss~~~ 114 (471)
.+.+.+|.+.+..|+-.+. ..+.+ +-||+|.-.|..|+.||+-.+.
T Consensus 167 ~~~~~~w~~~~~~PiF~splcv~~sv~i~~VdG~l~~f~~sG~qvwr~~t~ 217 (354)
T KOG4649|consen 167 YSSTEFWAATRFGPIFASPLCVGSSVIITTVDGVLTSFDESGRQVWRPATK 217 (354)
T ss_pred CCcceehhhhcCCccccCceeccceEEEEEeccEEEEEcCCCcEEEeecCC
Confidence 3457899999999987542 23333 4689998888888889976554
No 68
>PF10661 EssA: WXG100 protein secretion system (Wss), protein EssA; InterPro: IPR018920 The Wss (WXG100 protein secretion system) in Staphylococcus aureus seems to be encoded by a locus of eight ORFs, called ess (eSAT-6 secretion system) []. This locus encodes, amongst several other proteins, EssA, a protein predicted to possess one transmembrane domain. Due to its predicted membrane location and its absolute requirement for WXG100 protein secretion, it has been speculated that EssA could form a secretion apparatus in conjunction with YukC and YukAB. Proteins homologous to EssA, YukC, EsaA and YukD were absent from mycobacteria []. Members of this family are associated with type VII secretion of WXG100 family targets in the Firmicutes, but not in the Actinobacteria. This highly divergent protein family consists largely of a central region of highly polar low-complexity sequence containing occasional LF motifs in weak repeats about 17 residues in length, flanked by hydrophobic N- and C-terminal regions.
Probab=25.95 E-value=20 Score=32.13 Aligned_cols=23 Identities=17% Similarity=0.081 Sum_probs=14.2
Q ss_pred EEEehhHHHHHHHHHHHHhhhhh
Q 037760 445 IVAMSIISGMLILGLLLGMAWKK 467 (471)
Q Consensus 445 ii~~~v~~~~~~~~~~~~~~~~~ 467 (471)
+|+++|+++|+++++.+|.+.|+
T Consensus 120 ~i~~~i~g~ll~i~~giy~~~r~ 142 (145)
T PF10661_consen 120 TILLSIGGILLAICGGIYVVLRK 142 (145)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHH
Confidence 44455555566677777776664
No 69
>KOG1219 consensus Uncharacterized conserved protein, contains laminin, cadherin and EGF domains [Signal transduction mechanisms]
Probab=25.27 E-value=98 Score=39.61 Aligned_cols=25 Identities=16% Similarity=0.522 Sum_probs=16.6
Q ss_pred CCCCCccccCCCC--CccccCCCCccC
Q 037760 296 QCGANSNCRISKT--PICECLAGFISK 320 (471)
Q Consensus 296 ~CG~~g~C~~~~~--~~C~C~~GF~~~ 320 (471)
.|-.-|.|+.... -.|.||+.|.-.
T Consensus 3871 pCqhgG~C~~~~~ggy~CkCpsqysG~ 3897 (4289)
T KOG1219|consen 3871 PCQHGGTCISQPKGGYKCKCPSQYSGN 3897 (4289)
T ss_pred cccCCCEecCCCCCceEEeCcccccCc
Confidence 3444578875433 389999988643
No 70
>PF05545 FixQ: Cbb3-type cytochrome oxidase component FixQ; InterPro: IPR008621 This family consists of several Cbb3-type cytochrome oxidase components (FixQ/CcoQ). FixQ is found in nitrogen fixing bacteria. Since nitrogen fixation is an energy-consuming process, effective symbioses depend on operation of a respiratory chain with a high affinity for O2, closely coupled to ATP production. This requirement is fulfilled by a special three-subunit terminal oxidase (cytochrome terminal oxidase cbb3), which was first identified in Bradyrhizobium japonicum as the product of the fixNOQP operon [].
Probab=25.02 E-value=64 Score=22.99 Aligned_cols=13 Identities=15% Similarity=0.291 Sum_probs=5.5
Q ss_pred HHHHHhhhhhhcC
Q 037760 458 GLLLGMAWKKAKN 470 (471)
Q Consensus 458 ~~~~~~~~~~~~~ 470 (471)
+++.+..++++|+
T Consensus 24 gi~~w~~~~~~k~ 36 (49)
T PF05545_consen 24 GIVIWAYRPRNKK 36 (49)
T ss_pred HHHHHHHcccchh
Confidence 4444444444443
No 71
>PF05393 Hum_adeno_E3A: Human adenovirus early E3A glycoprotein; InterPro: IPR008652 This family consists of several early glycoproteins (E3A), from human adenovirus type 2.; GO: 0016021 integral to membrane
Probab=25.00 E-value=59 Score=26.45 Aligned_cols=8 Identities=38% Similarity=0.409 Sum_probs=3.3
Q ss_pred HHHHHHHh
Q 037760 456 ILGLLLGM 463 (471)
Q Consensus 456 ~~~~~~~~ 463 (471)
++.+++|+
T Consensus 45 il~Vilwf 52 (94)
T PF05393_consen 45 ILLVILWF 52 (94)
T ss_pred HHHHHHHH
Confidence 33444444
No 72
>TIGR03503 conserved hypothetical protein TIGR03503. This set of conserved hypothetical protein has a phylogenetic range that closely matches that of TIGR03501, a putative C-terminal protein targeting signal.
Probab=24.65 E-value=12 Score=38.85 Aligned_cols=21 Identities=29% Similarity=0.330 Sum_probs=13.3
Q ss_pred hhHHHHHHHHHHHHhhhhhhc
Q 037760 449 SIISGMLILGLLLGMAWKKAK 469 (471)
Q Consensus 449 ~v~~~~~~~~~~~~~~~~~~~ 469 (471)
++.+++++++++.+++|||||
T Consensus 353 ~~N~v~lllg~~~~~~~rk~k 373 (374)
T TIGR03503 353 VGNVVILLLGGIGFFVWRKKK 373 (374)
T ss_pred hhhhhhhhhheeeEEEEEEee
Confidence 334445667777777787776
No 73
>PF02480 Herpes_gE: Alphaherpesvirus glycoprotein E; InterPro: IPR003404 Glycoprotein E (gE) of Alphaherpesvirus forms a complex with glycoprotein I (gI), functioning as an immunoglobulin G (IgG) Fc binding protein. gE is involved in virus spread but is not essential for propagation [].; GO: 0016020 membrane; PDB: 2GJ7_F 2GIY_B.
Probab=21.62 E-value=31 Score=36.75 Aligned_cols=13 Identities=31% Similarity=0.856 Sum_probs=9.6
Q ss_pred CCCCCCCC--CCCCC
Q 037760 329 YSRRCDRK--PSDCP 341 (471)
Q Consensus 329 ~s~GC~~~--~~~C~ 341 (471)
...+|.+. +..|.
T Consensus 237 ~y~~C~~~~~~~~C~ 251 (439)
T PF02480_consen 237 RYANCSPSGWPRRCP 251 (439)
T ss_dssp EEEEEBTTC-TTTTE
T ss_pred hhcCCCCCCCcCCCC
Confidence 46789886 67894
No 74
>PF12301 CD99L2: CD99 antigen like protein 2; InterPro: IPR022078 This family of proteins is found in eukaryotes. Proteins in this family are typically between 165 and 237 amino acids in length. CD99L2 and CD99 are involved in trans-endothelial migration of neutrophils in vitro and in the recruitment of neutrophils into inflamed peritoneum.
Probab=21.41 E-value=65 Score=29.68 Aligned_cols=24 Identities=8% Similarity=0.088 Sum_probs=12.7
Q ss_pred ceeEEEEehhHHHHHHHHHHHHhh
Q 037760 441 RLKIIVAMSIISGMLILGLLLGMA 464 (471)
Q Consensus 441 ~~~~ii~~~v~~~~~~~~~~~~~~ 464 (471)
..++|.+|+-+++++|++.+.-++
T Consensus 113 ~~g~IaGIvsav~valvGAvsSyi 136 (169)
T PF12301_consen 113 EAGTIAGIVSAVVVALVGAVSSYI 136 (169)
T ss_pred ccchhhhHHHHHHHHHHHHHHHHH
Confidence 344555555555566665544443
No 75
>PF05092 PIF: Per os infectivity; InterPro: IPR007784 This entry represents a group of dsDNA Baculovirus proteins. It is required for the infectivity of the OBs or occlusion bodies. It is a structural protein of the ODV envelope required only in the first steps of per os larva infection, as viruses being produced in cells expressing the gene for this protein but not containing it in their genomes are able to produce successful infections. Baculoviruses are large DNA viruses that infect arthropods, mainly members of the order Lepidoptera. In their life cycle, they produce two kinds of particles, a budded, non-occluded virus (BV), which buds out of the infected cell and is responsible for the cell-to-cell transmission of the virus, and an occluded form, the occlusion body (OB), which is responsible for protecting the virus between encounters with larvae. A variable number of virions are included in the para-crystalline structure of the OB, mainly constituted by the virus-encoded polyhedrin protein; these virions are called occlusion body-derived virions or ODVs [].
Probab=20.47 E-value=77 Score=34.25 Aligned_cols=47 Identities=23% Similarity=0.600 Sum_probs=34.0
Q ss_pred cCCCCcccCCCCCCcccc-CCCCC-ccccCCCCccCCCCCCCCCCCCCCCCC
Q 037760 287 PFDACDNYAQCGANSNCR-ISKTP-ICECLAGFISKPQDDWDSPYSRRCDRK 336 (471)
Q Consensus 287 p~~~C~~~g~CG~~g~C~-~~~~~-~C~C~~GF~~~~~~~w~~~~s~GC~~~ 336 (471)
.++.|+++--|.|+|.=. .+..| .|.|.+||.+..-.+ ....=|++.
T Consensus 147 iy~DC~vpVGC~PhG~I~din~~pi~C~Cd~GyVsd~~~~---t~tP~CRp~ 195 (522)
T PF05092_consen 147 IYEDCDVPVGCQPHGRIADINESPIRCVCDDGYVSDFDSD---TETPYCRPR 195 (522)
T ss_pred hhccCCCcEecCCCCEEeeecCCceEeECCCCcccccccC---CCCcceece
Confidence 457799999999998754 55556 899999998763111 246778765
No 76
>PF05283 MGC-24: Multi-glycosylated core protein 24 (MGC-24); InterPro: IPR007947 CD164 is a mucin-like receptor, or sialomucin, with specificity in receptor/ ligand interactions that depends on the structural characteristics of the mucin-like receptor. Its functions include mediating, or regulating, haematopoietic progenitor cell adhesion and the negative regulation of their growth and/or-differentiation. It exists in the native state as a disulphide- linked homodimer of two 80-85kDa subunits. It is usually expressed by CD34+ and CD341o/- haematopoietic stem cells and associated microenvironmental cells. It contains, in its extracellular region, two mucin domains (I and II) linked by a non-mucin domain, which has been predicted to contain intra- disulphide bridges. This receptor may play a key role in haematopoiesis by facilitating the adhesion of human CD34+ cells to bone marrow stroma and by negatively regulating CD34+ CD341o/- haematopoietic progenitor cell proliferation. These effects involve the CD164 class I and/or II epitopes recognised by the monoclonal antibodies (mAbs) 105A5 and 103B2/9E10. These epitopes are carbohydrate-dependent and are located on the N-terminal mucin domain I [, ]. It has been found that murine MGC-24v and rat endolyn share significant sequence similarities with human CD164. However, CD164 lacks the consensus glycosaminoglycan (GAG)-attachment site found in MGC-24; it is possible that GAG-association is responsible for the high molecular weight of the epithelial-derived MGC-24 glycoprotein []. Genomic structure studies have placed CD164 within the mucin-subgroup that comprises multiple exons, and demonstrate the diverse chromosomal distribution of this family of molecules. Molecules with such multiple exons may have sophisticated regulatory mechanisms that involve not only post-translational modifications of the oligosaccharide side chains, but also differential exon usage. Although differences in the intron and exon sizes are seen between the mouse and human genes, the predicted proteins are similar in size and structure, maintaining functionally important motifs that regulate cell proliferation or subcellular distribution []. CD164 is a gene whose expression depends on differential usage of poly- adenylation sites within the 3'-UTR. The conserved distribution of the 3.2- and 1.2-kb CD164 transcripts between mouse and human suggests that (i) a mechanism may exist to regulate tissue-specific polyadenylation, and (ii) differences in polyadenylation are important for the expression and function of CD164 in different tissues. Two other aspects of the structure of CD164 are of particular interest. First, it shares one of several conserved features of a cytokine-binding pocket - in this respect, it is notable that evidence exists for a class of cell-surface sialomucin modulators that directly interact with growth factor receptors to regulate their response to physiological ligands. Second, its cytoplasmic tail contains a C-terminal YHTL motif found in many endocytic membrane proteins or receptors. These Tyr-based motifs bind to adaptor proteins, which mediate the sorting of membrane proteins into transport vesicles from the plasma membrane to the endosomes, and between intracellular compartments.
Probab=20.10 E-value=67 Score=30.06 Aligned_cols=24 Identities=29% Similarity=0.274 Sum_probs=13.9
Q ss_pred EEehhHHHHHHHHH--HHHhhhhhhc
Q 037760 446 VAMSIISGMLILGL--LLGMAWKKAK 469 (471)
Q Consensus 446 i~~~v~~~~~~~~~--~~~~~~~~~~ 469 (471)
.+..||.+||++|+ ++|+++|..|
T Consensus 160 ~~SFiGGIVL~LGv~aI~ff~~KF~k 185 (186)
T PF05283_consen 160 AASFIGGIVLTLGVLAIIFFLYKFCK 185 (186)
T ss_pred hhhhhhHHHHHHHHHHHHHHHhhhcc
Confidence 44666666666554 5556666544
Done!