Query         037760
Match_columns 471
No_of_seqs    217 out of 1680
Neff          7.4 
Searched_HMMs 46136
Date          Fri Mar 29 04:07:20 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/037760.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/037760hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF01453 B_lectin:  D-mannose b  99.9 1.5E-27 3.3E-32  205.3   3.3  110   70-183     1-114 (114)
  2 PF00954 S_locus_glycop:  S-loc  99.9 8.5E-26 1.9E-30  193.4  11.7  110  209-321     1-110 (110)
  3 cd00028 B_lectin Bulb-type man  99.9 4.6E-24   1E-28  184.4  15.3  114   32-151     2-116 (116)
  4 smart00108 B_lectin Bulb-type   99.9 1.1E-22 2.3E-27  175.3  14.6  112   32-150     2-114 (114)
  5 PF08276 PAN_2:  PAN-like domai  99.7   1E-16 2.3E-21  124.3   5.7   64  340-404     1-66  (66)
  6 cd00129 PAN_APPLE PAN/APPLE-li  99.5 2.6E-14 5.7E-19  114.7   6.2   72  340-421     5-80  (80)
  7 cd01098 PAN_AP_plant Plant PAN  99.5 6.9E-14 1.5E-18  113.1   7.9   78  338-422     3-84  (84)
  8 PF01453 B_lectin:  D-mannose b  98.9 2.6E-08 5.7E-13   85.7  12.2  100   38-152    12-114 (114)
  9 smart00473 PAN_AP divergent su  98.6 1.9E-07 4.1E-12   73.7   7.5   71  344-420     4-77  (78)
 10 smart00108 B_lectin Bulb-type   98.5 9.9E-07 2.2E-11   75.8   9.6   86   89-211    23-111 (114)
 11 cd00028 B_lectin Bulb-type man  98.3 3.9E-06 8.4E-11   72.3   8.6   86   90-212    24-113 (116)
 12 cd01100 APPLE_Factor_XI_like S  97.3 0.00029 6.3E-09   55.5   4.1   49  349-400     9-58  (73)
 13 PF00954 S_locus_glycop:  S-loc  97.2  0.0017 3.7E-08   55.3   8.3   68  232-304    32-101 (110)
 14 PF00024 PAN_1:  PAN domain Thi  92.4    0.15 3.3E-06   39.7   3.5   55  345-402     3-59  (79)
 15 PF04478 Mid2:  Mid2 like cell   92.3   0.068 1.5E-06   47.8   1.4   32  439-470    45-79  (154)
 16 smart00223 APPLE APPLE domain.  89.9    0.55 1.2E-05   37.6   4.4   51  349-399     6-57  (79)
 17 PF08693 SKG6:  Transmembrane a  86.9    0.95   2E-05   31.2   3.3   29  442-470    11-40  (40)
 18 smart00605 CW CW domain.        84.3     3.8 8.3E-05   33.6   6.6   58  362-425    20-78  (94)
 19 PF08277 PAN_3:  PAN-like domai  82.8     4.3 9.4E-05   31.1   6.0   32  362-398    18-49  (71)
 20 PF02009 Rifin_STEVOR:  Rifin/s  82.4     0.4 8.7E-06   48.1  -0.1   25  446-470   259-284 (299)
 21 PF02439 Adeno_E3_CR2:  Adenovi  81.6    0.43 9.2E-06   32.4  -0.1   23  446-468    10-32  (38)
 22 PF01102 Glycophorin_A:  Glycop  80.1    0.54 1.2E-05   40.8  -0.1   16  455-470    79-94  (122)
 23 cd00053 EGF Epidermal growth f  79.0     1.8   4E-05   27.6   2.3   29  291-319     2-31  (36)
 24 PF14295 PAN_4:  PAN domain; PD  77.6     2.4 5.3E-05   29.9   2.8   25  362-386    14-38  (51)
 25 cd01099 PAN_AP_HGF Subfamily o  76.4     7.7 0.00017   30.8   5.7   36  363-401    24-61  (80)
 26 PF07645 EGF_CA:  Calcium-bindi  75.5     1.2 2.5E-05   30.9   0.6   31  289-319     3-35  (42)
 27 PHA03265 envelope glycoprotein  74.6     2.5 5.4E-05   42.8   2.8   28  441-469   349-376 (402)
 28 smart00179 EGF_CA Calcium-bind  72.1     3.6 7.8E-05   27.1   2.4   30  289-318     3-33  (39)
 29 PTZ00382 Variant-specific surf  71.6     3.1 6.7E-05   34.6   2.3    8  313-320     9-16  (96)
 30 PF14610 DUF4448:  Protein of u  71.4     3.4 7.3E-05   38.6   2.8   28  444-471   160-187 (189)
 31 PF09064 Tme5_EGF_like:  Thromb  68.0     3.2 6.9E-05   27.5   1.3   18  302-319    11-28  (34)
 32 PRK11138 outer membrane biogen  66.7 1.2E+02  0.0026   31.3  13.6   53   93-148   127-187 (394)
 33 PF12877 DUF3827:  Domain of un  66.3     6.5 0.00014   43.1   3.9   18  438-455   265-282 (684)
 34 cd00054 EGF_CA Calcium-binding  66.1     5.6 0.00012   25.7   2.3   31  289-319     3-34  (38)
 35 PTZ00046 rifin; Provisional     65.6     2.3 4.9E-05   43.6   0.3   27  445-471   317-344 (358)
 36 TIGR01477 RIFIN variant surfac  64.4     2.4 5.3E-05   43.2   0.3   27  445-471   312-339 (353)
 37 PF12661 hEGF:  Human growth fa  63.6       2 4.4E-05   22.2  -0.2    9  310-318     1-9   (13)
 38 PF01683 EB:  EB module;  Inter  62.2     7.4 0.00016   28.1   2.5   33  286-321    17-49  (52)
 39 PF07974 EGF_2:  EGF-like domai  60.7     7.6 0.00016   25.4   2.0   23  295-318     6-28  (32)
 40 PF12947 EGF_3:  EGF domain;  I  59.2     2.7 5.9E-05   28.3  -0.3   26  294-319     5-31  (36)
 41 PF01034 Syndecan:  Syndecan do  58.6     3.2 6.9E-05   31.7  -0.0   14  457-470    26-39  (64)
 42 PF05454 DAG1:  Dystroglycan (D  56.4     3.7   8E-05   41.1   0.0   24  445-468   150-173 (290)
 43 PF01299 Lamp:  Lysosome-associ  53.4     7.9 0.00017   39.0   1.8   27  444-470   271-300 (306)
 44 PF00008 EGF:  EGF-like domain   53.4     4.1   9E-05   26.4  -0.1   23  296-318     5-29  (32)
 45 PF06697 DUF1191:  Protein of u  53.0      12 0.00027   37.0   3.0   22  445-466   216-238 (278)
 46 PF12662 cEGF:  Complement Clr-  52.9     6.5 0.00014   24.0   0.7   11  310-320     3-13  (24)
 47 PF13908 Shisa:  Wnt and FGF in  50.9      12 0.00026   34.5   2.5   17  443-459    79-95  (179)
 48 PTZ00382 Variant-specific surf  50.9      14 0.00031   30.6   2.7   10  443-452    66-75  (96)
 49 PF08374 Protocadherin:  Protoc  49.7      14  0.0003   35.1   2.7   15  439-453    34-48  (221)
 50 PF03302 VSP:  Giardia variant-  46.0      17 0.00036   38.2   2.9   24  440-463   364-387 (397)
 51 PF01436 NHL:  NHL repeat;  Int  45.1      34 0.00074   21.3   3.2   21   89-109     6-26  (28)
 52 smart00181 EGF Epidermal growt  44.3      21 0.00045   22.9   2.2   24  295-319     6-30  (35)
 53 TIGR01478 STEVOR variant surfa  42.8      11 0.00024   37.3   1.0   22  330-351   169-190 (295)
 54 PTZ00370 STEVOR; Provisional    40.7      13 0.00027   37.0   1.0   22  330-351   169-190 (296)
 55 TIGR01167 LPXTG_anchor LPXTG-m  36.1      36 0.00078   22.0   2.3   10  460-469    24-33  (34)
 56 PF13360 PQQ_2:  PQQ-like domai  35.7      81  0.0018   29.3   5.7   51   95-148     2-63  (238)
 57 PRK11138 outer membrane biogen  35.0 1.7E+02  0.0037   30.1   8.5   20   93-112   263-283 (394)
 58 cd05845 Ig2_L1-CAM_like Second  32.7      77  0.0017   26.2   4.3   33   69-102    32-64  (95)
 59 PF15345 TMEM51:  Transmembrane  32.6      50  0.0011   31.9   3.5   15  454-468    71-85  (233)
 60 PF13360 PQQ_2:  PQQ-like domai  32.5 2.3E+02   0.005   26.2   8.3   75   70-147    12-102 (238)
 61 PF10681 Rot1:  Chaperone for p  30.6 2.7E+02  0.0058   26.5   7.9   74  106-197    64-138 (212)
 62 PF01102 Glycophorin_A:  Glycop  30.0      15 0.00032   32.0  -0.4   27  445-471    66-92  (122)
 63 TIGR03300 assembly_YfgL outer   29.9 2.1E+02  0.0046   29.0   8.1   20   94-113    73-93  (377)
 64 PF06365 CD34_antigen:  CD34/Po  29.5      61  0.0013   30.7   3.5   26  444-469   101-129 (202)
 65 TIGR03300 assembly_YfgL outer   29.0 2.7E+02  0.0058   28.3   8.6   75   69-147    83-171 (377)
 66 KOG3637 Vitronectin receptor,   26.8      48   0.001   39.2   2.9   15  452-466   987-1001(1030)
 67 KOG4649 PQQ (pyrrolo-quinoline  26.0 2.3E+02  0.0051   28.2   6.9   46   69-114   167-217 (354)
 68 PF10661 EssA:  WXG100 protein   26.0      20 0.00043   32.1  -0.3   23  445-467   120-142 (145)
 69 KOG1219 Uncharacterized conser  25.3      98  0.0021   39.6   4.9   25  296-320  3871-3897(4289)
 70 PF05545 FixQ:  Cbb3-type cytoc  25.0      64  0.0014   23.0   2.2   13  458-470    24-36  (49)
 71 PF05393 Hum_adeno_E3A:  Human   25.0      59  0.0013   26.5   2.2    8  456-463    45-52  (94)
 72 TIGR03503 conserved hypothetic  24.6      12 0.00025   38.8  -2.3   21  449-469   353-373 (374)
 73 PF02480 Herpes_gE:  Alphaherpe  21.6      31 0.00067   36.8   0.0   13  329-341   237-251 (439)
 74 PF12301 CD99L2:  CD99 antigen   21.4      65  0.0014   29.7   2.0   24  441-464   113-136 (169)
 75 PF05092 PIF:  Per os infectivi  20.5      77  0.0017   34.2   2.6   47  287-336   147-195 (522)
 76 PF05283 MGC-24:  Multi-glycosy  20.1      67  0.0014   30.1   1.9   24  446-469   160-185 (186)

No 1  
>PF01453 B_lectin:  D-mannose binding lectin;  InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]:  Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein   This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity.  Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=99.93  E-value=1.5e-27  Score=205.34  Aligned_cols=110  Identities=50%  Similarity=0.801  Sum_probs=80.2

Q ss_pred             CCeEEEEecCCCCCCC--CCceEEEccCCcEEEEeCCCceEEee-cCCCCc-cceEEEEecCCCEEEEeCcCCCCCceee
Q 037760           70 PRTVVWVANRYKPITD--KNGVLTLSNNGSILLLNQERSTIWSS-NSSRVL-ETAVVRLLDSGNLVLRDNVSRSSDEYMW  145 (471)
Q Consensus        70 ~~~~vW~an~~~pv~~--~~~~l~l~~~GnLvl~d~~~~~vWss-~~~~~~-~~~~~~Lld~GNlvl~~~~~~~~~~~~W  145 (471)
                      ++++||+|||+.|+..  ...+|.|++||||||+|..+..+|++ .+.+.. ....++|+|+|||||++.    .+.++|
T Consensus         1 ~~tvvW~an~~~p~~~~s~~~~L~l~~dGnLvl~~~~~~~iWss~~t~~~~~~~~~~~L~~~GNlvl~d~----~~~~lW   76 (114)
T PF01453_consen    1 PRTVVWVANRNSPLTSSSGNYTLILQSDGNLVLYDSNGSVIWSSNNTSGRGNSGCYLVLQDDGNLVLYDS----SGNVLW   76 (114)
T ss_dssp             ---------TTEEEEECETTEEEEEETTSEEEEEETTTEEEEE--S-TTSS-SSEEEEEETTSEEEEEET----TSEEEE
T ss_pred             CcccccccccccccccccccccceECCCCeEEEEcCCCCEEEEecccCCccccCeEEEEeCCCCEEEEee----cceEEE
Confidence            3689999999999854  24789999999999999998899999 555443 468899999999999996    678999


Q ss_pred             eeccCCCCCCCCCCeeeeeccCCceeEEEEecCCCCCC
Q 037760          146 QSFDYPSDTLLPGMKLGWNLRTRFERYLTAWRNADDPT  183 (471)
Q Consensus       146 qSFd~PTDTLLPGq~L~~~~~tg~~~~L~Sw~s~~dps  183 (471)
                      |||||||||+||||+|+.+..+|.+..++||++.+|||
T Consensus        77 ~Sf~~ptdt~L~~q~l~~~~~~~~~~~~~sw~s~~dps  114 (114)
T PF01453_consen   77 QSFDYPTDTLLPGQKLGDGNVTGKNDSLTSWSSNTDPS  114 (114)
T ss_dssp             ESTTSSS-EEEEEET--TSEEEEESTSSEEEESS----
T ss_pred             eecCCCccEEEeccCcccCCCccccceEEeECCCCCCC
Confidence            99999999999999999876666666799999999996


No 2  
>PF00954 S_locus_glycop:  S-locus glycoprotein family;  InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=99.93  E-value=8.5e-26  Score=193.40  Aligned_cols=110  Identities=45%  Similarity=0.945  Sum_probs=102.9

Q ss_pred             eeeCCCCCceeeccccccccCCCceeeeEEEecCCeeEEEEEecCCCceeEEEEccCCcEEEEEeecCCCCeeEeeeccC
Q 037760          209 VRSGPWNGQQFVGIPMFFPRLKNKVYIPMLVRTEDEAYYTYKPINDKVIPRLYLDQSGKLQRFVWNQTSSEWRMSYSWPF  288 (471)
Q Consensus       209 w~sg~w~~~~~~~~p~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~rl~ld~dG~l~~y~w~~~~~~W~~~~~~p~  288 (471)
                      ||+|+|+|..|++.|+|..   ...+.+.|+.++++++++|.+.+...++|++||++|+++++.|.+..++|.+.|.+|.
T Consensus         1 wrsG~WnG~~f~g~p~~~~---~~~~~~~fv~~~~e~~~t~~~~~~s~~~r~~ld~~G~l~~~~w~~~~~~W~~~~~~p~   77 (110)
T PF00954_consen    1 WRSGPWNGQRFSGIPEMSS---NSLYNYSFVSNNEEVYYTYSLSNSSVLSRLVLDSDGQLQRYIWNESTQSWSVFWSAPK   77 (110)
T ss_pred             CCccccCCeEECCcccccc---cceeEEEEEECCCeEEEEEecCCCceEEEEEEeeeeEEEEEEEecCCCcEEEEEEecc
Confidence            8999999999999999875   5678889999999999999998888899999999999999999999999999999999


Q ss_pred             CCCcccCCCCCCccccCCCCCccccCCCCccCC
Q 037760          289 DACDNYAQCGANSNCRISKTPICECLAGFISKP  321 (471)
Q Consensus       289 ~~C~~~g~CG~~g~C~~~~~~~C~C~~GF~~~~  321 (471)
                      +.|++|+.||+||+|+.+..+.|+|++||+|++
T Consensus        78 d~Cd~y~~CG~~g~C~~~~~~~C~Cl~GF~P~n  110 (110)
T PF00954_consen   78 DQCDVYGFCGPNGICNSNNSPKCSCLPGFEPKN  110 (110)
T ss_pred             cCCCCccccCCccEeCCCCCCceECCCCcCCCc
Confidence            999999999999999987778999999999964


No 3  
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=99.92  E-value=4.6e-24  Score=184.41  Aligned_cols=114  Identities=48%  Similarity=0.813  Sum_probs=98.6

Q ss_pred             CccCCCCeEEeCCCeeEEEEECCCCCCceEEEEEEeC-CCCeEEEEecCCCCCCCCCceEEEccCCcEEEEeCCCceEEe
Q 037760           32 QSISDGETLVSSSLRFELGFFSPGNSNNRYLGIWYKS-SPRTVVWVANRYKPITDKNGVLTLSNNGSILLLNQERSTIWS  110 (471)
Q Consensus        32 ~~L~~~~~L~S~~g~F~lgf~~~~~~~~~yl~i~~~~-~~~~~vW~an~~~pv~~~~~~l~l~~~GnLvl~d~~~~~vWs  110 (471)
                      +.|..|++|+|+++.|++|||.+......+++|||.. + .++||+||++.| ....+.|.|++||||+|+|.++.++|+
T Consensus         2 ~~l~~~~~l~s~~~~f~~G~~~~~~q~~dgnlv~~~~~~-~~~vW~snt~~~-~~~~~~l~l~~dGnLvl~~~~g~~vW~   79 (116)
T cd00028           2 NPLSSGQTLVSSGSLFELGFFKLIMQSRDYNLILYKGSS-RTVVWVANRDNP-SGSSCTLTLQSDGNLVIYDGSGTVVWS   79 (116)
T ss_pred             cCcCCCCEEEeCCCcEEEecccCCCCCCeEEEEEEeCCC-CeEEEECCCCCC-CCCCEEEEEecCCCeEEEcCCCcEEEE
Confidence            5688999999999999999999875433788999986 4 789999999988 345688999999999999999999999


Q ss_pred             ecCCCCccceEEEEecCCCEEEEeCcCCCCCceeeeeccCC
Q 037760          111 SNSSRVLETAVVRLLDSGNLVLRDNVSRSSDEYMWQSFDYP  151 (471)
Q Consensus       111 s~~~~~~~~~~~~Lld~GNlvl~~~~~~~~~~~~WqSFd~P  151 (471)
                      |++.+......++|+|+|||||++.    .+.++|||||||
T Consensus        80 S~~~~~~~~~~~~L~ddGnlvl~~~----~~~~~W~Sf~~P  116 (116)
T cd00028          80 SNTTRVNGNYVLVLLDDGNLVLYDS----DGNFLWQSFDYP  116 (116)
T ss_pred             ecccCCCCceEEEEeCCCCEEEECC----CCCEEEcCCCCC
Confidence            9987523356899999999999997    577999999999


No 4  
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=99.89  E-value=1.1e-22  Score=175.28  Aligned_cols=112  Identities=48%  Similarity=0.819  Sum_probs=96.7

Q ss_pred             CccCCCCeEEeCCCeeEEEEECCCCCCceEEEEEEeC-CCCeEEEEecCCCCCCCCCceEEEccCCcEEEEeCCCceEEe
Q 037760           32 QSISDGETLVSSSLRFELGFFSPGNSNNRYLGIWYKS-SPRTVVWVANRYKPITDKNGVLTLSNNGSILLLNQERSTIWS  110 (471)
Q Consensus        32 ~~L~~~~~L~S~~g~F~lgf~~~~~~~~~yl~i~~~~-~~~~~vW~an~~~pv~~~~~~l~l~~~GnLvl~d~~~~~vWs  110 (471)
                      +.|..|++|+|+++.|++|||.+... ..+++|||.. + .++||+||++.|+.. ++.|.|++||||||+|.++.++|+
T Consensus         2 ~~l~~~~~l~s~~~~f~~G~~~~~~q-~dgnlV~~~~~~-~~~vW~snt~~~~~~-~~~l~l~~dGnLvl~~~~g~~vW~   78 (114)
T smart00108        2 NTLSSGQTLVSGNSLFELGFFTLIMQ-NDYNLILYKSSS-RTVVWVANRDNPVSD-SCTLTLQSDGNLVLYDGDGRVVWS   78 (114)
T ss_pred             cccCCCCEEecCCCcEeeeccccCCC-CCEEEEEEECCC-CcEEEECCCCCCCCC-CEEEEEeCCCCEEEEeCCCCEEEE
Confidence            56788999999999999999998653 4788899987 5 789999999988754 488999999999999998999999


Q ss_pred             ecCCCCccceEEEEecCCCEEEEeCcCCCCCceeeeeccC
Q 037760          111 SNSSRVLETAVVRLLDSGNLVLRDNVSRSSDEYMWQSFDY  150 (471)
Q Consensus       111 s~~~~~~~~~~~~Lld~GNlvl~~~~~~~~~~~~WqSFd~  150 (471)
                      |++........++|+|+|||||++.    .+.++||||||
T Consensus        79 S~t~~~~~~~~~~L~ddGnlvl~~~----~~~~~W~Sf~~  114 (114)
T smart00108       79 SNTTGANGNYVLVLLDDGNLVIYDS----DGNFLWQSFDY  114 (114)
T ss_pred             ecccCCCCceEEEEeCCCCEEEECC----CCCEEeCCCCC
Confidence            9986323356799999999999987    56799999997


No 5  
>PF08276 PAN_2:  PAN-like domain;  InterPro: IPR013227 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs
Probab=99.66  E-value=1e-16  Score=124.32  Aligned_cols=64  Identities=56%  Similarity=1.231  Sum_probs=54.7

Q ss_pred             CCCCCceEEeccCCCCCc--ccccCCCCHHHHHHHHhhCCCeEEEEeccCCCCCCceEeecccccce
Q 037760          340 CPSGEGFLKLQRMKLPEN--YWSNKSMNLKECEAECIRNCSCRAYANSDITGGGDGCLMWFGDLIDI  404 (471)
Q Consensus       340 C~~~~~f~~~~~~~~p~~--~~~~~~~~~~~C~~~Cl~nCsC~a~~y~~~~~~g~gC~~w~~~l~~~  404 (471)
                      |+.+|+|+++++|++|++  +.++.++++++|++.||+||||+||+|.++. ++++|++|.++|+|+
T Consensus         1 C~~~d~F~~l~~~~~p~~~~~~~~~~~s~~~C~~~Cl~nCsC~Ayay~~~~-~~~~C~lW~~~L~d~   66 (66)
T PF08276_consen    1 CGSGDGFLKLPNMKLPDFDNAIVDSSVSLEECEKACLSNCSCTAYAYSNLS-GGGGCLLWYGDLVDL   66 (66)
T ss_pred             CcCCCEEEEECCeeCCCCcceeeecCCCHHHHHhhcCCCCCEeeEEeeccC-CCCEEEEEcCEeecC
Confidence            545789999999999998  4444668999999999999999999998654 467899999999885


No 6  
>cd00129 PAN_APPLE PAN/APPLE-like domain; present in N-terminal (N) domains of plasminogen/ hepatocyte growth factor proteins,  plasma prekallikrein/coagulation factor XI and microneme antigen proteins, plant receptor-like protein kinases, and various nematode and leech anti-platelet proteins. Common structural features include two disulfide bonds that link the alpha-helix to the central region of the protein. PAN domains have significant functional versatility, fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=99.50  E-value=2.6e-14  Score=114.70  Aligned_cols=72  Identities=24%  Similarity=0.407  Sum_probs=61.3

Q ss_pred             CCCCCceEEeccCCCCCcccccCCCCHHHHHHHHhh---CCCeEEEEeccCCCCCCceEeecccc-cceeeeccccCCce
Q 037760          340 CPSGEGFLKLQRMKLPENYWSNKSMNLKECEAECIR---NCSCRAYANSDITGGGDGCLMWFGDL-IDIRECTEEFSWGQ  415 (471)
Q Consensus       340 C~~~~~f~~~~~~~~p~~~~~~~~~~~~~C~~~Cl~---nCsC~a~~y~~~~~~g~gC~~w~~~l-~~~~~~~~~~~~~~  415 (471)
                      |..++.|+++.++++|++    ..++++||+++|++   ||||.||+|.+.   +.||++|.++| +|+++..   ..+.
T Consensus         5 ~~~~g~fl~~~~~klpd~----~~~s~~eC~~~Cl~~~~nCsC~Aya~~~~---~~gC~~W~~~l~~d~~~~~---~~g~   74 (80)
T cd00129           5 CKSAGTTLIKIALKIKTT----KANTADECANRCEKNGLPFSCKAFVFAKA---RKQCLWFPFNSMSGVRKEF---SHGF   74 (80)
T ss_pred             eecCCeEEEeecccCCcc----cccCHHHHHHHHhcCCCCCCceeeeccCC---CCCeEEecCcchhhHHhcc---CCCc
Confidence            444578999999999998    22689999999999   999999999752   45899999999 9998877   6789


Q ss_pred             eEEEEe
Q 037760          416 DIFIRV  421 (471)
Q Consensus       416 ~~yirv  421 (471)
                      ++|||.
T Consensus        75 ~Ly~r~   80 (80)
T cd00129          75 DLYENK   80 (80)
T ss_pred             eeEeEC
Confidence            999983


No 7  
>cd01098 PAN_AP_plant Plant PAN/APPLE-like domain; present in plant S-receptor protein kinases and secreted glycoproteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions. S-receptor protein kinases and S-locus glycoproteins are involved in sporophytic self-incompatibility response in Brassica, one of probably many molecular mechanisms, by which hermaphrodite flowering plants avoid self-fertilization.
Probab=99.49  E-value=6.9e-14  Score=113.10  Aligned_cols=78  Identities=41%  Similarity=0.913  Sum_probs=63.5

Q ss_pred             CCCCCC---CceEEeccCCCCCc-ccccCCCCHHHHHHHHhhCCCeEEEEeccCCCCCCceEeecccccceeeeccccCC
Q 037760          338 SDCPSG---EGFLKLQRMKLPEN-YWSNKSMNLKECEAECIRNCSCRAYANSDITGGGDGCLMWFGDLIDIRECTEEFSW  413 (471)
Q Consensus       338 ~~C~~~---~~f~~~~~~~~p~~-~~~~~~~~~~~C~~~Cl~nCsC~a~~y~~~~~~g~gC~~w~~~l~~~~~~~~~~~~  413 (471)
                      ++|...   +.|+++.++++|+. ... ...++++|++.||+||+|+||+|.+   ++++|++|...+.+.+...   ..
T Consensus         3 ~~C~~~~~~~~f~~~~~~~~~~~~~~~-~~~s~~~C~~~Cl~nCsC~a~~~~~---~~~~C~~~~~~~~~~~~~~---~~   75 (84)
T cd01098           3 LNCGGDGSTDGFLKLPDVKLPDNASAI-TAISLEECREACLSNCSCTAYAYNN---GSGGCLLWNGLLNNLRSLS---SG   75 (84)
T ss_pred             cccCCCCCCCEEEEeCCeeCCCchhhh-ccCCHHHHHHHHhcCCCcceeeecC---CCCeEEEEeceecceEeec---CC
Confidence            467543   68999999999987 333 6679999999999999999999985   2457999999999877654   44


Q ss_pred             ceeEEEEee
Q 037760          414 GQDIFIRVP  422 (471)
Q Consensus       414 ~~~~yirv~  422 (471)
                      +..+||||+
T Consensus        76 ~~~~yiKv~   84 (84)
T cd01098          76 GGTLYLRLA   84 (84)
T ss_pred             CcEEEEEeC
Confidence            688999985


No 8  
>PF01453 B_lectin:  D-mannose binding lectin;  InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]:  Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein   This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity.  Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=98.89  E-value=2.6e-08  Score=85.71  Aligned_cols=100  Identities=27%  Similarity=0.416  Sum_probs=68.3

Q ss_pred             CeEEeCCCeeEEEEECCCCCCceEEEEEEeCCCCeEEEEe-cCCCCCCCCCceEEEccCCcEEEEeCCCceEEeecCCCC
Q 037760           38 ETLVSSSLRFELGFFSPGNSNNRYLGIWYKSSPRTVVWVA-NRYKPITDKNGVLTLSNNGSILLLNQERSTIWSSNSSRV  116 (471)
Q Consensus        38 ~~L~S~~g~F~lgf~~~~~~~~~yl~i~~~~~~~~~vW~a-n~~~pv~~~~~~l~l~~~GnLvl~d~~~~~vWss~~~~~  116 (471)
                      +.+.+.+|.+.|-|+.+++     |.| |. ...+++|.+ +...... ..+.|+|+++|||||+|..+.+||+|+....
T Consensus        12 ~p~~~~s~~~~L~l~~dGn-----Lvl-~~-~~~~~iWss~~t~~~~~-~~~~~~L~~~GNlvl~d~~~~~lW~Sf~~pt   83 (114)
T PF01453_consen   12 SPLTSSSGNYTLILQSDGN-----LVL-YD-SNGSVIWSSNNTSGRGN-SGCYLVLQDDGNLVLYDSSGNVLWQSFDYPT   83 (114)
T ss_dssp             EEEEECETTEEEEEETTSE-----EEE-EE-TTTEEEEE--S-TTSS--SSEEEEEETTSEEEEEETTSEEEEESTTSSS
T ss_pred             cccccccccccceECCCCe-----EEE-Ec-CCCCEEEEecccCCccc-cCeEEEEeCCCCEEEEeecceEEEeecCCCc
Confidence            4565656999999999875     433 44 345789999 4443321 4689999999999999999999999976322


Q ss_pred             ccceEEEEec--CCCEEEEeCcCCCCCceeeeeccCCC
Q 037760          117 LETAVVRLLD--SGNLVLRDNVSRSSDEYMWQSFDYPS  152 (471)
Q Consensus       117 ~~~~~~~Lld--~GNlvl~~~~~~~~~~~~WqSFd~PT  152 (471)
                        ...+..++  .||++ +..    ...++|.|-++|+
T Consensus        84 --dt~L~~q~l~~~~~~-~~~----~~~~sw~s~~dps  114 (114)
T PF01453_consen   84 --DTLLPGQKLGDGNVT-GKN----DSLTSWSSNTDPS  114 (114)
T ss_dssp             ---EEEEEET--TSEEE-EES----TSSEEEESS----
T ss_pred             --cEEEeccCcccCCCc-ccc----ceEEeECCCCCCC
Confidence              34566666  78888 553    4569999999885


No 9  
>smart00473 PAN_AP divergent subfamily of APPLE domains. Apple-like domains present in Plasminogen, C. elegans hypothetical ORFs and the extracellular portion of plant receptor-like protein kinases. Predicted to possess protein- and/or carbohydrate-binding functions.
Probab=98.58  E-value=1.9e-07  Score=73.66  Aligned_cols=71  Identities=35%  Similarity=0.759  Sum_probs=55.0

Q ss_pred             CceEEeccCCCCCc-ccccCCCCHHHHHHHHhh-CCCeEEEEeccCCCCCCceEeec-ccccceeeeccccCCceeEEEE
Q 037760          344 EGFLKLQRMKLPEN-YWSNKSMNLKECEAECIR-NCSCRAYANSDITGGGDGCLMWF-GDLIDIRECTEEFSWGQDIFIR  420 (471)
Q Consensus       344 ~~f~~~~~~~~p~~-~~~~~~~~~~~C~~~Cl~-nCsC~a~~y~~~~~~g~gC~~w~-~~l~~~~~~~~~~~~~~~~yir  420 (471)
                      ..|.+++++.+++. .......++++|++.|++ +|+|.||.|..   .+.+|.+|. +.+.+.+...   ..+.++|.|
T Consensus         4 ~~f~~~~~~~l~~~~~~~~~~~s~~~C~~~C~~~~~~C~s~~y~~---~~~~C~l~~~~~~~~~~~~~---~~~~~~y~~   77 (78)
T smart00473        4 DCFVRLPNTKLPGFSRIVISVASLEECASKCLNSNCSCRSFTYNN---GTKGCLLWSESSLGDARLFP---SGGVDLYEK   77 (78)
T ss_pred             ceeEEecCccCCCCcceeEcCCCHHHHHHHhCCCCCceEEEEEcC---CCCEEEEeeCCccccceecc---cCCceeEEe
Confidence            46899999999865 322345699999999999 99999999974   245699999 7777776444   456677776


No 10 
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=98.47  E-value=9.9e-07  Score=75.76  Aligned_cols=86  Identities=20%  Similarity=0.365  Sum_probs=61.0

Q ss_pred             eEEEccCCcEEEEeCC-CceEEeecCCCCcc-ceEEEEecCCCEEEEeCcCCCCCceeeeeccCCCCCCCCCCeeeeecc
Q 037760           89 VLTLSNNGSILLLNQE-RSTIWSSNSSRVLE-TAVVRLLDSGNLVLRDNVSRSSDEYMWQSFDYPSDTLLPGMKLGWNLR  166 (471)
Q Consensus        89 ~l~l~~~GnLvl~d~~-~~~vWss~~~~~~~-~~~~~Lld~GNlvl~~~~~~~~~~~~WqSFd~PTDTLLPGq~L~~~~~  166 (471)
                      .+.++.|||||+++.. ..++|++++..+.. ...+.|+++|||||++.    .+.++|+|-..                
T Consensus        23 ~~~~q~dgnlV~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~----~g~~vW~S~t~----------------   82 (114)
T smart00108       23 TLIMQNDYNLILYKSSSRTVVWVANRDNPVSDSCTLTLQSDGNLVLYDG----DGRVVWSSNTT----------------   82 (114)
T ss_pred             ccCCCCCEEEEEEECCCCcEEEECCCCCCCCCCEEEEEeCCCCEEEEeC----CCCEEEEeccc----------------
Confidence            3556789999999865 57999999864422 36789999999999987    46789998221                


Q ss_pred             CCceeEEEEecCCCCCCCceEEEEEccCCCeeEEEee-Cceeeeee
Q 037760          167 TRFERYLTAWRNADDPTPGEFSFRFDISTMAELVTVT-GSKIEVRS  211 (471)
Q Consensus       167 tg~~~~L~Sw~s~~dps~G~f~l~l~~~g~~~~~~~~-~~~~Yw~s  211 (471)
                                     ...|.+.+.|+++|..++  ++ ..++.|.+
T Consensus        83 ---------------~~~~~~~~~L~ddGnlvl--~~~~~~~~W~S  111 (114)
T smart00108       83 ---------------GANGNYVLVLLDDGNLVI--YDSDGNFLWQS  111 (114)
T ss_pred             ---------------CCCCceEEEEeCCCCEEE--ECCCCCEEeCC
Confidence                           023456778888888544  43 23577865


No 11 
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=98.27  E-value=3.9e-06  Score=72.32  Aligned_cols=86  Identities=20%  Similarity=0.310  Sum_probs=60.7

Q ss_pred             EEEcc-CCcEEEEeCC-CceEEeecCCCC-ccceEEEEecCCCEEEEeCcCCCCCceeeeeccCCCCCCCCCCeeeeecc
Q 037760           90 LTLSN-NGSILLLNQE-RSTIWSSNSSRV-LETAVVRLLDSGNLVLRDNVSRSSDEYMWQSFDYPSDTLLPGMKLGWNLR  166 (471)
Q Consensus        90 l~l~~-~GnLvl~d~~-~~~vWss~~~~~-~~~~~~~Lld~GNlvl~~~~~~~~~~~~WqSFd~PTDTLLPGq~L~~~~~  166 (471)
                      +.++. +|+||+++.. ..++|++++..+ .....+.|+++|||||++.    ++.++|+|-...               
T Consensus        24 ~~~q~~dgnlv~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~----~g~~vW~S~~~~---------------   84 (116)
T cd00028          24 LIMQSRDYNLILYKGSSRTVVWVANRDNPSGSSCTLTLQSDGNLVIYDG----SGTVVWSSNTTR---------------   84 (116)
T ss_pred             CCCCCCeEEEEEEeCCCCeEEEECCCCCCCCCCEEEEEecCCCeEEEcC----CCcEEEEecccC---------------
Confidence            44565 9999999764 479999998653 2346789999999999987    467899874321               


Q ss_pred             CCceeEEEEecCCCCCCCceEEEEEccCCCeeEEEee-CceeeeeeC
Q 037760          167 TRFERYLTAWRNADDPTPGEFSFRFDISTMAELVTVT-GSKIEVRSG  212 (471)
Q Consensus       167 tg~~~~L~Sw~s~~dps~G~f~l~l~~~g~~~~~~~~-~~~~Yw~sg  212 (471)
                                      ..+.+.+.|+++|...+  ++ ...+.|.+.
T Consensus        85 ----------------~~~~~~~~L~ddGnlvl--~~~~~~~~W~Sf  113 (116)
T cd00028          85 ----------------VNGNYVLVLLDDGNLVL--YDSDGNFLWQSF  113 (116)
T ss_pred             ----------------CCCceEEEEeCCCCEEE--ECCCCCEEEcCC
Confidence                            13456778888887444  43 245778764


No 12 
>cd01100 APPLE_Factor_XI_like Subfamily of PAN/APPLE-like domains; present in plasma prekallikrein/coagulation factor XI, microneme antigen proteins, and a few prokaryotic proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=97.30  E-value=0.00029  Score=55.46  Aligned_cols=49  Identities=12%  Similarity=0.328  Sum_probs=34.5

Q ss_pred             eccCCCCCc-ccccCCCCHHHHHHHHhhCCCeEEEEeccCCCCCCceEeeccc
Q 037760          349 LQRMKLPEN-YWSNKSMNLKECEAECIRNCSCRAYANSDITGGGDGCLMWFGD  400 (471)
Q Consensus       349 ~~~~~~p~~-~~~~~~~~~~~C~~~Cl~nCsC~a~~y~~~~~~g~gC~~w~~~  400 (471)
                      +++++++.. .......+.++|++.|+.+|+|.||.|..   +...|+++...
T Consensus         9 ~~~~~~~g~d~~~~~~~s~~~Cq~~C~~~~~C~afT~~~---~~~~C~lk~~~   58 (73)
T cd01100           9 GSNVDFRGGDLSTVFASSAEQCQAACTADPGCLAFTYNT---KSKKCFLKSSE   58 (73)
T ss_pred             cCCCccccCCcceeecCCHHHHHHHcCCCCCceEEEEEC---CCCeEEcccCC
Confidence            346666554 21222458999999999999999999974   23359997653


No 13 
>PF00954 S_locus_glycop:  S-locus glycoprotein family;  InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=97.22  E-value=0.0017  Score=55.34  Aligned_cols=68  Identities=12%  Similarity=0.222  Sum_probs=56.0

Q ss_pred             ceeeeEEEecCCeeEEEEEecCCCceeEEEEccCCcEEEEEeecCCCCeeEeeeccCCCCcccCCCCCC--cccc
Q 037760          232 KVYIPMLVRTEDEAYYTYKPINDKVIPRLYLDQSGKLQRFVWNQTSSEWRMSYSWPFDACDNYAQCGAN--SNCR  304 (471)
Q Consensus       232 ~~~~~~~~~~~~~~~~~~~~~~~~~~~rl~ld~dG~l~~y~w~~~~~~W~~~~~~p~~~C~~~g~CG~~--g~C~  304 (471)
                      ....++|...+..++.++.+...+.+++++++.+.+.|...|..+.+.     |+.++.|+.+|+|..+  ..|.
T Consensus        32 ~e~~~t~~~~~~s~~~r~~ld~~G~l~~~~w~~~~~~W~~~~~~p~d~-----Cd~y~~CG~~g~C~~~~~~~C~  101 (110)
T PF00954_consen   32 EEVYYTYSLSNSSVLSRLVLDSDGQLQRYIWNESTQSWSVFWSAPKDQ-----CDVYGFCGPNGICNSNNSPKCS  101 (110)
T ss_pred             CeEEEEEecCCCceEEEEEEeeeeEEEEEEEecCCCcEEEEEEecccC-----CCCccccCCccEeCCCCCCceE
Confidence            345567776666777788888888999999999999999999988776     9999999999999765  4575


No 14 
>PF00024 PAN_1:  PAN domain This Prosite entry concerns apple domains, a subset of PAN domains;  InterPro: IPR003014 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs It has been shown that, the N-terminal N domains of members of the plasminogen/hepatocyte growth factor family, the apple domains of the plasma prekallikrein/coagulation factor XI family, and domains of various nematode proteins belong to the same module superfamily, the PAN module []. PAN contains a conserved core of three disulphide bridges. In some members of the family there is an additional fourth disulphide bridge that links the N and C termini of the domain.; PDB: 1GP9_C 2QJ2_B 1GMO_H 1NK1_B 3MKP_B 1BHT_B 3HN4_A 1GMN_A 3HMS_A 3HMT_B ....
Probab=92.37  E-value=0.15  Score=39.70  Aligned_cols=55  Identities=16%  Similarity=0.397  Sum_probs=39.5

Q ss_pred             ceEEeccCCCCCc-ccccCCCCHHHHHHHHhhCCC-eEEEEeccCCCCCCceEeeccccc
Q 037760          345 GFLKLQRMKLPEN-YWSNKSMNLKECEAECIRNCS-CRAYANSDITGGGDGCLMWFGDLI  402 (471)
Q Consensus       345 ~f~~~~~~~~p~~-~~~~~~~~~~~C~~~Cl~nCs-C~a~~y~~~~~~g~gC~~w~~~l~  402 (471)
                      .|.++.+..+... .......++++|.+.|+.+=. |.+|.|..   ....|.+...+-.
T Consensus         3 ~f~~~~~~~l~~~~~~~~~v~s~~~C~~~C~~~~~~C~s~~y~~---~~~~C~L~~~~~~   59 (79)
T PF00024_consen    3 AFERIPGYRLSGHSIKEINVPSLEECAQLCLNEPRRCKSFNYDP---SSKTCYLSSSDRS   59 (79)
T ss_dssp             TEEEEEEEEEESCEEEEEEESSHHHHHHHHHHSTT-ESEEEEET---TTTEEEEECSSSS
T ss_pred             CeEEECCEEEeCCcceEEcCCCHHHHHhhcCcCcccCCeEEEEC---CCCEEEEcCCCCC
Confidence            3777777776665 222234589999999999999 99999985   2335998765443


No 15 
>PF04478 Mid2:  Mid2 like cell wall stress sensor;  InterPro: IPR007567 This family represents a region near the C terminus of Mid2, which contains a transmembrane region. The remainder of the protein sequence is serine-rich and of low complexity, and is therefore impossible to align accurately. Mid2 is thought to act as a mechanosensor of cell wall stress. The C-terminal cytoplasmic region of Mid2 is known to interact with Rom2, a guanine nucleotide exchange factor (GEF) for Rho1, which is part of the cell wall integrity signalling pathway [].
Probab=92.26  E-value=0.068  Score=47.78  Aligned_cols=32  Identities=19%  Similarity=0.232  Sum_probs=16.7

Q ss_pred             ccceeEEEEehhHHHH--HH-HHHHHHhhhhhhcC
Q 037760          439 KKRLKIIVAMSIISGM--LI-LGLLLGMAWKKAKN  470 (471)
Q Consensus       439 ~~~~~~ii~~~v~~~~--~~-~~~~~~~~~~~~~~  470 (471)
                      .+.+++||+++||+.+  ++ +++++|++++|+||
T Consensus        45 ~knknIVIGvVVGVGg~ill~il~lvf~~c~r~kk   79 (154)
T PF04478_consen   45 SKNKNIVIGVVVGVGGPILLGILALVFIFCIRRKK   79 (154)
T ss_pred             cCCccEEEEEEecccHHHHHHHHHhheeEEEeccc
Confidence            3344688998887543  22 23334444444443


No 16 
>smart00223 APPLE APPLE domain. Four-fold repeat in plasma kallikrein and coagulation factor XI. Factor XI apple 3 mediates binding to platelets. Factor XI apple 1 binds high-molecular-mass kininogen. Apple 4 in factor XI mediates dimer formation and binds to factor XIIa. Mutations in apple 4 cause factor XI deficiency, an inherited bleeding disorder.
Probab=89.89  E-value=0.55  Score=37.58  Aligned_cols=51  Identities=12%  Similarity=0.217  Sum_probs=34.4

Q ss_pred             eccCCCCCc-ccccCCCCHHHHHHHHhhCCCeEEEEeccCCCCCCceEeecc
Q 037760          349 LQRMKLPEN-YWSNKSMNLKECEAECIRNCSCRAYANSDITGGGDGCLMWFG  399 (471)
Q Consensus       349 ~~~~~~p~~-~~~~~~~~~~~C~~~Cl~nCsC~a~~y~~~~~~g~gC~~w~~  399 (471)
                      +++++++.. .......+.++|++.|..+=.|.+|.|.........|+++..
T Consensus         6 ~~~~df~G~Dl~~~~~~~~~~Cq~~Ct~~~~C~~FTf~~~~~~~~~C~LK~s   57 (79)
T smart00223        6 YKNVDFRGSDINTVYVPSAQVCQKRCTSHPRCLFFTFSTNEPPEEKCLLKDS   57 (79)
T ss_pred             ccCccccCceeeeeecCCHHHHHHhhcCCCCccEEEeeCCCCCCCEeEeCcC
Confidence            345555554 222234589999999999999999999753221226998643


No 17 
>PF08693 SKG6:  Transmembrane alpha-helix domain;  InterPro: IPR014805 SKG6 and AXL2 are membrane proteins that show polarised intracellular localisation [, ]. This entry represents the highly conserved transmembrane alpha-helical domain found in these proteins [, ]. The full-length AXL2 protein has a negative regulatory function in cytokinesis [].
Probab=86.86  E-value=0.95  Score=31.24  Aligned_cols=29  Identities=24%  Similarity=0.472  Sum_probs=12.8

Q ss_pred             eeEEEEehhHHHHHHHHH-HHHhhhhhhcC
Q 037760          442 LKIIVAMSIISGMLILGL-LLGMAWKKAKN  470 (471)
Q Consensus       442 ~~~ii~~~v~~~~~~~~~-~~~~~~~~~~~  470 (471)
                      .-+-++++|.+.++++.+ +.+++|+||+|
T Consensus        11 vaIa~~VvVPV~vI~~vl~~~l~~~~rR~k   40 (40)
T PF08693_consen   11 VAIAVGVVVPVGVIIIVLGAFLFFWYRRKK   40 (40)
T ss_pred             EEEEEEEEechHHHHHHHHHHhheEEeccC
Confidence            334455555544433222 33344555543


No 18 
>smart00605 CW CW domain.
Probab=84.32  E-value=3.8  Score=33.56  Aligned_cols=58  Identities=14%  Similarity=0.461  Sum_probs=39.6

Q ss_pred             CCCCHHHHHHHHhhCCCeEEEEeccCCCCCCceEeecc-cccceeeeccccCCceeEEEEeecCc
Q 037760          362 KSMNLKECEAECIRNCSCRAYANSDITGGGDGCLMWFG-DLIDIRECTEEFSWGQDIFIRVPAAD  425 (471)
Q Consensus       362 ~~~~~~~C~~~Cl~nCsC~a~~y~~~~~~g~gC~~w~~-~l~~~~~~~~~~~~~~~~yirv~~s~  425 (471)
                      ...+.++|...|..+..|+.+....    ...|.++.- ++..+++...  ..+..+=+|+..+.
T Consensus        20 ~~~sw~~Ci~~C~~~~~Cvlay~~~----~~~C~~f~~~~~~~v~~~~~--~~~~~VAfK~~~~~   78 (94)
T smart00605       20 ATLSWDECIQKCYEDSNCVLAYGNS----SETCYLFSYGTVLTVKKLSS--SSGKKVAFKVSTDQ   78 (94)
T ss_pred             cCCCHHHHHHHHhCCCceEEEecCC----CCceEEEEcCCeEEEEEccC--CCCcEEEEEEeCCC
Confidence            3467899999999999999876541    246987643 4556666541  34566778876443


No 19 
>PF08277 PAN_3:  PAN-like domain;  InterPro: IPR006583 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs The PAN-3 or CW is a domain associated with a number of Caenorhabditis elegans hypothetical proteins.
Probab=82.81  E-value=4.3  Score=31.08  Aligned_cols=32  Identities=13%  Similarity=0.434  Sum_probs=26.5

Q ss_pred             CCCCHHHHHHHHhhCCCeEEEEeccCCCCCCceEeec
Q 037760          362 KSMNLKECEAECIRNCSCRAYANSDITGGGDGCLMWF  398 (471)
Q Consensus       362 ~~~~~~~C~~~Cl~nCsC~a~~y~~~~~~g~gC~~w~  398 (471)
                      ...+.++|-..|..+=.|.++.+.     ...|.++.
T Consensus        18 ~~~sw~~Cv~~C~~~~~C~la~~~-----~~~C~~y~   49 (71)
T PF08277_consen   18 TNTSWDDCVQKCYNDENCVLAYFD-----SGKCYLYN   49 (71)
T ss_pred             cCCCHHHHhHHhCCCCEEEEEEeC-----CCCEEEEE
Confidence            345789999999999999998886     24799874


No 20 
>PF02009 Rifin_STEVOR:  Rifin/stevor family;  InterPro: IPR002858 Malaria is still a major cause of mortality in many areas of the world. Plasmodium falciparum causes the most severe human form of the disease and is responsible for most fatalities. Severe cases of malaria can occur when the parasite invades and then proliferates within red blood cell erythrocytes. The parasite produces many variant antigenic proteins, encoded by multigene families, which are present on the surface of the infected erythrocyte and play important roles in virulence. A crucial survival mechanism for the malaria parasite is its ability to evade the immune response by switching these variant surface antigens. The high virulence of P. falciparum relative to other malarial parasites is in large part due to the fact that in this organism many of these surface antigens mediate the binding of infected erythrocytes to the vascular endothelium (cytoadherence) and non-infected erythrocytes (rosetting). This can lead to the accumulation of infected cells in the vasculature of a variety of organs, blocking the blood flow and reducing the oxygen supply. Clinical symptoms of severe infection can include fever, progressive anaemia, multi-organ dysfunction and coma. For more information see []. Several multicopy gene families have been described in Plasmodium falciparum, including the stevor family of subtelomeric open reading frames and the rif interspersed repetitive elements. Both families contain three predicted transmembrane segments. It has been proposed that stevor and rif are members of a larger superfamily that code for variant surface antigens [].
Probab=82.36  E-value=0.4  Score=48.10  Aligned_cols=25  Identities=8%  Similarity=0.187  Sum_probs=13.8

Q ss_pred             EEehhHHHH-HHHHHHHHhhhhhhcC
Q 037760          446 VAMSIISGM-LILGLLLGMAWKKAKN  470 (471)
Q Consensus       446 i~~~v~~~~-~~~~~~~~~~~~~~~~  470 (471)
                      ++++|+++| +++.+++|++||.|||
T Consensus       259 ~aSiiaIliIVLIMvIIYLILRYRRK  284 (299)
T PF02009_consen  259 IASIIAILIIVLIMVIIYLILRYRRK  284 (299)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence            334444333 3356677777776654


No 21 
>PF02439 Adeno_E3_CR2:  Adenovirus E3 region protein CR2;  InterPro: IPR003470 Early region 3 (E3) of human adenoviruses (Ads) codes for proteins that appear to control viral interactions with the host []. This region called CR1 (conserved region 1) [] is found three times in Human adenovirus 19 (a subgroup D adenovirus) 49 kDa protein in the E3 region. CR1 is also found in the 20.1 Kd protein of subgroup B adenoviruses. The function of this 80 amino acid region is unknown. This region is probably a divergent immunoglobulin domain.
Probab=81.60  E-value=0.43  Score=32.38  Aligned_cols=23  Identities=17%  Similarity=0.212  Sum_probs=11.3

Q ss_pred             EEehhHHHHHHHHHHHHhhhhhh
Q 037760          446 VAMSIISGMLILGLLLGMAWKKA  468 (471)
Q Consensus       446 i~~~v~~~~~~~~~~~~~~~~~~  468 (471)
                      ++++++.++++++++.|..++||
T Consensus        10 v~V~vg~~iiii~~~~YaCcykk   32 (38)
T PF02439_consen   10 VAVVVGMAIIIICMFYYACCYKK   32 (38)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHcc
Confidence            33444444455555555455444


No 22 
>PF01102 Glycophorin_A:  Glycophorin A;  InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=80.07  E-value=0.54  Score=40.82  Aligned_cols=16  Identities=19%  Similarity=-0.010  Sum_probs=7.3

Q ss_pred             HHHHHHHHhhhhhhcC
Q 037760          455 LILGLLLGMAWKKAKN  470 (471)
Q Consensus       455 ~~~~~~~~~~~~~~~~  470 (471)
                      ++++++.|+++|+|||
T Consensus        79 g~Illi~y~irR~~Kk   94 (122)
T PF01102_consen   79 GIILLISYCIRRLRKK   94 (122)
T ss_dssp             HHHHHHHHHHHHHS--
T ss_pred             HHHHHHHHHHHHHhcc
Confidence            3344555555555554


No 23 
>cd00053 EGF Epidermal growth factor domain, found in epidermal growth factor (EGF) presents in a large number of proteins, mostly animal; the list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied; the functional significance of EGF-like domains in what appear to be unrelated proteins is not yet clear; a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase); the domain includes six cysteine residues which have been shown to be involved in disulfide bonds; the main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet; Subdomains between the conserved cysteines vary in length; the region between the 5th and 6th cysteine contains two conserved glycines of which at  least  one  is  present  in  most EGF-like domains; a subset of these bind calcium.
Probab=79.05  E-value=1.8  Score=27.62  Aligned_cols=29  Identities=21%  Similarity=0.639  Sum_probs=20.7

Q ss_pred             CcccCCCCCCccccCCC-CCccccCCCCcc
Q 037760          291 CDNYAQCGANSNCRISK-TPICECLAGFIS  319 (471)
Q Consensus       291 C~~~g~CG~~g~C~~~~-~~~C~C~~GF~~  319 (471)
                      |.....|..++.|.... ...|.|++||..
T Consensus         2 C~~~~~C~~~~~C~~~~~~~~C~C~~g~~g   31 (36)
T cd00053           2 CAASNPCSNGGTCVNTPGSYRCVCPPGYTG   31 (36)
T ss_pred             CCCCCCCCCCCEEecCCCCeEeECCCCCcc
Confidence            34346788888897543 358999999964


No 24 
>PF14295 PAN_4:  PAN domain; PDB: 2YIL_E 2YIP_C 2YIO_A.
Probab=77.57  E-value=2.4  Score=29.94  Aligned_cols=25  Identities=24%  Similarity=0.622  Sum_probs=17.8

Q ss_pred             CCCCHHHHHHHHhhCCCeEEEEecc
Q 037760          362 KSMNLKECEAECIRNCSCRAYANSD  386 (471)
Q Consensus       362 ~~~~~~~C~~~Cl~nCsC~a~~y~~  386 (471)
                      ...+.++|.+.|..+=.|.+|.|..
T Consensus        14 ~~~s~~~C~~~C~~~~~C~~~~~~~   38 (51)
T PF14295_consen   14 TASSPEECQAACAADPGCQAFTFNP   38 (51)
T ss_dssp             ----HHHHHHHHHTSTT--EEEEET
T ss_pred             cCCCHHHHHHHccCCCCCCEEEEEC
Confidence            3458999999999999999999974


No 25 
>cd01099 PAN_AP_HGF Subfamily of PAN/APPLE-like domains; present in N-terminal (N) domains of plasminogen/hepatocyte growth factor proteins, and various proteins found in Bilateria, such as leech anti-platelet proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=76.42  E-value=7.7  Score=30.79  Aligned_cols=36  Identities=22%  Similarity=0.578  Sum_probs=27.8

Q ss_pred             CCCHHHHHHHHhh--CCCeEEEEeccCCCCCCceEeecccc
Q 037760          363 SMNLKECEAECIR--NCSCRAYANSDITGGGDGCLMWFGDL  401 (471)
Q Consensus       363 ~~~~~~C~~~Cl~--nCsC~a~~y~~~~~~g~gC~~w~~~l  401 (471)
                      ..++++|.+.|++  +=.|.+|.|..   ....|.+-..+.
T Consensus        24 ~~s~~~C~~~C~~~~~f~CrSf~y~~---~~~~C~L~~~~~   61 (80)
T cd01099          24 VASLEECLRKCLEETEFTCRSFNYNY---KSKECILSDEDR   61 (80)
T ss_pred             cCCHHHHHHHhCCCCCceEeEEEEEc---CCCEEEEeCCCc
Confidence            4689999999999  89999999974   133598754443


No 26 
>PF07645 EGF_CA:  Calcium-binding EGF domain;  InterPro: IPR001881 A sequence of about forty amino-acid residues found in epidermal growth factor (EGF) has been shown [, , , , , ] to be present in a large number of membrane-bound and extracellular, mostly animal, proteins. Many of these proteins require calcium for their biological function and a calcium-binding site has been found at the N terminus of some EGF-like domains []. Calcium-binding may be crucial for numerous protein-protein interactions. For human coagulation factor IX it has been shown [] that the calcium-ligands form a pentagonal bipyramid. The first, third and fourth conserved negatively charged or polar residues are side chain ligands. The latter is possibly hydroxylated (see aspartic acid and asparagine hydroxylation site) []. A conserved aromatic residue, as well as the second conserved negative residue, are thought to be involved in stabilising the calcium-binding site. As in non-calcium binding EGF-like domains, there are six conserved cysteines and the structure of both types is very similar as calcium-binding induces only strictly local structural changes [].  +------------------+ +---------+ | | | | nxnnC-x(3,14)-C-x(3,7)-CxxbxxxxaxC-x(1,6)-C-x(8,13)-Cx | | +------------------+ 'n': negatively charged or polar residue [DEQN] 'b': possibly beta-hydroxylated residue [DN] 'a': aromatic amino acid 'C': cysteine, involved in disulphide bond 'x': any amino acid. ; GO: 0005509 calcium ion binding; PDB: 2VJ3_A 1TOZ_A 1LMJ_A 1UZQ_A 1UZK_A 1UZJ_B 1UZP_A 1EMO_A 1EMN_A 2RR0_A ....
Probab=75.45  E-value=1.2  Score=30.91  Aligned_cols=31  Identities=26%  Similarity=0.636  Sum_probs=23.3

Q ss_pred             CCCccc-CCCCCCccccCC-CCCccccCCCCcc
Q 037760          289 DACDNY-AQCGANSNCRIS-KTPICECLAGFIS  319 (471)
Q Consensus       289 ~~C~~~-g~CG~~g~C~~~-~~~~C~C~~GF~~  319 (471)
                      |+|... ..|..++.|... ++-.|.|++||+.
T Consensus         3 dEC~~~~~~C~~~~~C~N~~Gsy~C~C~~Gy~~   35 (42)
T PF07645_consen    3 DECAEGPHNCPENGTCVNTEGSYSCSCPPGYEL   35 (42)
T ss_dssp             STTTTTSSSSSTTSEEEEETTEEEEEESTTEEE
T ss_pred             cccCCCCCcCCCCCEEEcCCCCEEeeCCCCcEE
Confidence            567764 479889999743 3348999999984


No 27 
>PHA03265 envelope glycoprotein D; Provisional
Probab=74.59  E-value=2.5  Score=42.81  Aligned_cols=28  Identities=21%  Similarity=0.482  Sum_probs=17.9

Q ss_pred             ceeEEEEehhHHHHHHHHHHHHhhhhhhc
Q 037760          441 RLKIIVAMSIISGMLILGLLLGMAWKKAK  469 (471)
Q Consensus       441 ~~~~ii~~~v~~~~~~~~~~~~~~~~~~~  469 (471)
                      .+.++|+..|+. ++++|+++|++|||||
T Consensus       349 ~~g~~ig~~i~g-lv~vg~il~~~~rr~k  376 (402)
T PHA03265        349 FVGISVGLGIAG-LVLVGVILYVCLRRKK  376 (402)
T ss_pred             ccceEEccchhh-hhhhhHHHHHHhhhhh
Confidence            345556554432 4567888898888774


No 28 
>smart00179 EGF_CA Calcium-binding EGF-like domain.
Probab=72.10  E-value=3.6  Score=27.11  Aligned_cols=30  Identities=27%  Similarity=0.689  Sum_probs=21.1

Q ss_pred             CCCcccCCCCCCccccCCC-CCccccCCCCc
Q 037760          289 DACDNYAQCGANSNCRISK-TPICECLAGFI  318 (471)
Q Consensus       289 ~~C~~~g~CG~~g~C~~~~-~~~C~C~~GF~  318 (471)
                      +.|.....|...+.|.... ...|.|++||.
T Consensus         3 ~~C~~~~~C~~~~~C~~~~g~~~C~C~~g~~   33 (39)
T smart00179        3 DECASGNPCQNGGTCVNTVGSYRCECPPGYT   33 (39)
T ss_pred             ccCcCCCCcCCCCEeECCCCCeEeECCCCCc
Confidence            4565545687778897443 34799999996


No 29 
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=71.64  E-value=3.1  Score=34.60  Aligned_cols=8  Identities=13%  Similarity=0.189  Sum_probs=4.6

Q ss_pred             cCCCCccC
Q 037760          313 CLAGFISK  320 (471)
Q Consensus       313 C~~GF~~~  320 (471)
                      |.+|+.|.
T Consensus         9 C~~g~~~~   16 (96)
T PTZ00382          9 CDSDKKPN   16 (96)
T ss_pred             CCCCCccC
Confidence            55666553


No 30 
>PF14610 DUF4448:  Protein of unknown function (DUF4448)
Probab=71.44  E-value=3.4  Score=38.61  Aligned_cols=28  Identities=18%  Similarity=0.393  Sum_probs=15.1

Q ss_pred             EEEEehhHHHHHHHHHHHHhhhhhhcCC
Q 037760          444 IIVAMSIISGMLILGLLLGMAWKKAKNK  471 (471)
Q Consensus       444 ~ii~~~v~~~~~~~~~~~~~~~~~~~~~  471 (471)
                      +.|++-+++++++++++++++|+||+||
T Consensus       160 laI~lPvvv~~~~~~~~~~~~~~R~~Rr  187 (189)
T PF14610_consen  160 LAIALPVVVVVLALIMYGFFFWNRKKRR  187 (189)
T ss_pred             EEEEccHHHHHHHHHHHhhheeecccee
Confidence            3444444444455566666667666554


No 31 
>PF09064 Tme5_EGF_like:  Thrombomodulin like fifth domain, EGF-like;  InterPro: IPR015149 This domain adopts a fold similar to other EGF domains, with a flat major and a twisted minor beta sheet. Disulphide pairing, however, is not of the usual 1-3, 2-4, 5-6 type; rather 1-2, 3-4, 5-6 pairing is found. Its extended major sheet (strands beta-2 and beta-3 and the connecting loop) projects into thrombin's active site groove. This domain is required for interaction of thrombomodulin with thrombin, and subsequent activation of protein-C []. ; GO: 0004888 transmembrane signaling receptor activity, 0016021 integral to membrane
Probab=68.03  E-value=3.2  Score=27.53  Aligned_cols=18  Identities=28%  Similarity=0.719  Sum_probs=13.3

Q ss_pred             cccCCCCCccccCCCCcc
Q 037760          302 NCRISKTPICECLAGFIS  319 (471)
Q Consensus       302 ~C~~~~~~~C~C~~GF~~  319 (471)
                      .|+.+...+|.||.||..
T Consensus        11 ~CDpn~~~~C~CPeGyIl   28 (34)
T PF09064_consen   11 DCDPNSPGQCFCPEGYIL   28 (34)
T ss_pred             ccCCCCCCceeCCCceEe
Confidence            465555568999999964


No 32 
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=66.72  E-value=1.2e+02  Score=31.29  Aligned_cols=53  Identities=17%  Similarity=0.313  Sum_probs=32.6

Q ss_pred             ccCCcEEEEeC-CCceEEeecCCCCc-cce------EEEEecCCCEEEEeCcCCCCCceeeeec
Q 037760           93 SNNGSILLLNQ-ERSTIWSSNSSRVL-ETA------VVRLLDSGNLVLRDNVSRSSDEYMWQSF  148 (471)
Q Consensus        93 ~~~GnLvl~d~-~~~~vWss~~~~~~-~~~------~~~Lld~GNlvl~~~~~~~~~~~~WqSF  148 (471)
                      ..+|.|+-+|. .|..+|+....+.. ..+      ......+|.|+-.|..   +++++|+--
T Consensus       127 ~~~g~l~ald~~tG~~~W~~~~~~~~~ssP~v~~~~v~v~~~~g~l~ald~~---tG~~~W~~~  187 (394)
T PRK11138        127 SEKGQVYALNAEDGEVAWQTKVAGEALSRPVVSDGLVLVHTSNGMLQALNES---DGAVKWTVN  187 (394)
T ss_pred             cCCCEEEEEECCCCCCcccccCCCceecCCEEECCEEEEECCCCEEEEEEcc---CCCEeeeec
Confidence            45788887886 68899998754321 111      1122345666666652   577899863


No 33 
>PF12877 DUF3827:  Domain of unknown function (DUF3827);  InterPro: IPR024606 The function of the proteins in this entry is not currently known, but one of the human proteins (Q9HCM3 from SWISSPROT) has been implicated in pilocytic astrocytomas [, , ]. In the majority of cases of pilocytic astrocytomas a tandem duplication produces an in-frame fusion of the gene encoding this protein and the BRAF oncogene. The resulting fusion protein has constitutive BRAF kinase activity and is capable of transforming cells. 
Probab=66.29  E-value=6.5  Score=43.05  Aligned_cols=18  Identities=17%  Similarity=0.132  Sum_probs=11.2

Q ss_pred             cccceeEEEEehhHHHHH
Q 037760          438 KKKRLKIIVAMSIISGML  455 (471)
Q Consensus       438 ~~~~~~~ii~~~v~~~~~  455 (471)
                      ..++.||||||++.++++
T Consensus       265 ~~~NlWII~gVlvPv~vV  282 (684)
T PF12877_consen  265 PPNNLWIIAGVLVPVLVV  282 (684)
T ss_pred             CCCCeEEEehHhHHHHHH
Confidence            345667778777665543


No 34 
>cd00054 EGF_CA Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.
Probab=66.14  E-value=5.6  Score=25.69  Aligned_cols=31  Identities=23%  Similarity=0.635  Sum_probs=20.8

Q ss_pred             CCCcccCCCCCCccccCCC-CCccccCCCCcc
Q 037760          289 DACDNYAQCGANSNCRISK-TPICECLAGFIS  319 (471)
Q Consensus       289 ~~C~~~g~CG~~g~C~~~~-~~~C~C~~GF~~  319 (471)
                      +.|.....|...+.|.... ...|.|++||.-
T Consensus         3 ~~C~~~~~C~~~~~C~~~~~~~~C~C~~g~~g   34 (38)
T cd00054           3 DECASGNPCQNGGTCVNTVGSYRCSCPPGYTG   34 (38)
T ss_pred             ccCCCCCCcCCCCEeECCCCCeEeECCCCCcC
Confidence            4565435687778887433 347999999853


No 35 
>PTZ00046 rifin; Provisional
Probab=65.57  E-value=2.3  Score=43.59  Aligned_cols=27  Identities=11%  Similarity=0.270  Sum_probs=14.1

Q ss_pred             EEEehhHHHHHH-HHHHHHhhhhhhcCC
Q 037760          445 IVAMSIISGMLI-LGLLLGMAWKKAKNK  471 (471)
Q Consensus       445 ii~~~v~~~~~~-~~~~~~~~~~~~~~~  471 (471)
                      |++++|+++|+| +.+++|++.|.||||
T Consensus       317 IiaSiiAIvVIVLIMvIIYLILRYRRKK  344 (358)
T PTZ00046        317 IIASIVAIVVIVLIMVIIYLILRYRRKK  344 (358)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHHhhhcc
Confidence            444444444433 455667766655443


No 36 
>TIGR01477 RIFIN variant surface antigen, rifin family. This model represents the rifin branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of rifin sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 20 bits.
Probab=64.44  E-value=2.4  Score=43.24  Aligned_cols=27  Identities=15%  Similarity=0.278  Sum_probs=13.9

Q ss_pred             EEEehhHHHHHH-HHHHHHhhhhhhcCC
Q 037760          445 IVAMSIISGMLI-LGLLLGMAWKKAKNK  471 (471)
Q Consensus       445 ii~~~v~~~~~~-~~~~~~~~~~~~~~~  471 (471)
                      |++.+|+++|+| +.+++|++.|.||||
T Consensus       312 IiaSiIAIvvIVLIMvIIYLILRYRRKK  339 (353)
T TIGR01477       312 IIASIIAILIIVLIMVIIYLILRYRRKK  339 (353)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHHhhhcc
Confidence            344444444433 455667666655443


No 37 
>PF12661 hEGF:  Human growth factor-like EGF; PDB: 2YGQ_A 2E26_A 3A7Q_A 2YGP_A 2YGO_A 1HRE_A 1HAE_A 1HAF_A 1HRF_A.
Probab=63.55  E-value=2  Score=22.24  Aligned_cols=9  Identities=33%  Similarity=1.173  Sum_probs=6.6

Q ss_pred             ccccCCCCc
Q 037760          310 ICECLAGFI  318 (471)
Q Consensus       310 ~C~C~~GF~  318 (471)
                      .|.|++||.
T Consensus         1 ~C~C~~G~~    9 (13)
T PF12661_consen    1 TCQCPPGWT    9 (13)
T ss_dssp             EEEE-TTEE
T ss_pred             CccCcCCCc
Confidence            499999985


No 38 
>PF01683 EB:  EB module;  InterPro: IPR006149  The EB domain has no known function. It is found in several Caenorhabditis sp. and Drosophila sp. proteins. The domain contains 8 conserved cysteines that probably form four disulphide bridges and is found associated with kunitz domains IPR002223 from INTERPRO 
Probab=62.19  E-value=7.4  Score=28.06  Aligned_cols=33  Identities=27%  Similarity=0.642  Sum_probs=26.3

Q ss_pred             ccCCCCcccCCCCCCccccCCCCCccccCCCCccCC
Q 037760          286 WPFDACDNYAQCGANSNCRISKTPICECLAGFISKP  321 (471)
Q Consensus       286 ~p~~~C~~~g~CG~~g~C~~~~~~~C~C~~GF~~~~  321 (471)
                      .|-+.|.....|-.++.|..   ..|.|++||.+..
T Consensus        17 ~~g~~C~~~~qC~~~s~C~~---g~C~C~~g~~~~~   49 (52)
T PF01683_consen   17 QPGESCESDEQCIGGSVCVN---GRCQCPPGYVEVG   49 (52)
T ss_pred             CCCCCCCCcCCCCCcCEEcC---CEeECCCCCEecC
Confidence            35567998889999999953   5899999997654


No 39 
>PF07974 EGF_2:  EGF-like domain;  InterPro: IPR013111 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length. This entry contains EGF domains found in a variety of extracellular and membrane proteins
Probab=60.71  E-value=7.6  Score=25.40  Aligned_cols=23  Identities=22%  Similarity=0.655  Sum_probs=17.9

Q ss_pred             CCCCCCccccCCCCCccccCCCCc
Q 037760          295 AQCGANSNCRISKTPICECLAGFI  318 (471)
Q Consensus       295 g~CG~~g~C~~~~~~~C~C~~GF~  318 (471)
                      ..|...|.|... ..+|.|.+||.
T Consensus         6 ~~C~~~G~C~~~-~g~C~C~~g~~   28 (32)
T PF07974_consen    6 NICSGHGTCVSP-CGRCVCDSGYT   28 (32)
T ss_pred             CccCCCCEEeCC-CCEEECCCCCc
Confidence            468888899743 45899999985


No 40 
>PF12947 EGF_3:  EGF domain;  InterPro: IPR024731 This entry represents an EGF domain found in the the C terminus of malarial parasite merozoite surface protein 1 [], as well as other proteins.; PDB: 2NPR_A 1N1I_C 1B9W_A 1YO8_A 2RHP_A.
Probab=59.18  E-value=2.7  Score=28.29  Aligned_cols=26  Identities=23%  Similarity=0.655  Sum_probs=17.2

Q ss_pred             cCCCCCCccccCCC-CCccccCCCCcc
Q 037760          294 YAQCGANSNCRISK-TPICECLAGFIS  319 (471)
Q Consensus       294 ~g~CG~~g~C~~~~-~~~C~C~~GF~~  319 (471)
                      .+-|.++..|.... .-.|.|.+||.-
T Consensus         5 ~~~C~~nA~C~~~~~~~~C~C~~Gy~G   31 (36)
T PF12947_consen    5 NGGCHPNATCTNTGGSYTCTCKPGYEG   31 (36)
T ss_dssp             GGGS-TTCEEEE-TTSEEEEE-CEEEC
T ss_pred             CCCCCCCcEeecCCCCEEeECCCCCcc
Confidence            35788899998543 348999999963


No 41 
>PF01034 Syndecan:  Syndecan domain;  InterPro: IPR001050 The syndecans are transmembrane proteoglycans which are involved in the organisation of cytoskeleton and/or actin microfilaments, and have important roles as cell surface receptors during cell-cell and/or cell-matrix interactions [, ]. Structurally, these proteins consist of four separate domains:   A signal sequence; An extracellular domain (ectodomain) of variable length whose sequence is not evolutionary conserved in the various forms of syndecans. The ectodomain contains the sites of attachment of the heparan sulphate glycosaminoglycan side chains;  A transmembrane region;  A highly conserved cytoplasmic domain of about 30 to 35 residues, which could interact with cytoskeletal proteins.    The proteins known to belong to this family are:    Syndecan 1.  Syndecan 2 or fibroglycan.  Syndecan 3 or neuroglycan or N-syndecan.  Syndecan 4 or amphiglycan or ryudocan.  Drosophila syndecan.   Caenorhabditis elegans probable syndecan (F57C7.3).    Syndecan-4, a transmembrane heparan sulphate proteoglycan, is a coreceptor with integrins in cell adhesion. It has been suggested to form a ternary signalling complex with protein kinase Calpha and phosphatidylinositol 4,5-bisphosphate (PIP2). Structural studies have demonstrated that the cytoplasmic domain undergoes a conformational transition and forms a symmetric dimer in the presence of phospholipid activator PIP2, and whose overall structure in solution exhibits a twisted clamp shape having a cavity in the centre of dimeric interface. In addition, it has been observed that the syndecan-4 variable domain interacts, strongly, not only with fatty acyl groups but also the anionic head group of PIP2. These findings indicate that PIP2 promotes oligomerisation of the syndecan-4 cytoplasmic domain for transmembrane signalling and cell-matrix adhesion [, ].; GO: 0008092 cytoskeletal protein binding, 0016020 membrane; PDB: 1EJQ_B 1EJP_B 1YBO_C 1OBY_Q.
Probab=58.60  E-value=3.2  Score=31.68  Aligned_cols=14  Identities=21%  Similarity=0.354  Sum_probs=0.5

Q ss_pred             HHHHHHhhhhhhcC
Q 037760          457 LGLLLGMAWKKAKN  470 (471)
Q Consensus       457 ~~~~~~~~~~~~~~  470 (471)
                      +.++++++.|.|||
T Consensus        26 ilLIlf~iyR~rkk   39 (64)
T PF01034_consen   26 ILLILFLIYRMRKK   39 (64)
T ss_dssp             -----------S--
T ss_pred             HHHHHHHHHHHHhc
Confidence            33444455554443


No 42 
>PF05454 DAG1:  Dystroglycan (Dystrophin-associated glycoprotein 1);  InterPro: IPR008465 Dystroglycan is one of the dystrophin-associated glycoproteins, which is encoded by a 5.5 kb transcript in Homo sapiens. The protein product is cleaved into two non-covalently associated subunits, [alpha] (N-terminal) and [beta] (C-terminal). In skeletal muscle the dystroglycan complex works as a transmembrane linkage between the extracellular matrix and the cytoskeleton [alpha]-dystroglycan is extracellular and binds to merosin ([alpha]-2 laminin) in the basement membrane, while [beta]-dystroglycan is a transmembrane protein and binds to dystrophin, which is a large rod-like cytoskeletal protein, absent in Duchenne muscular dystrophy patients. Dystrophin binds to intracellular actin cables. In this way, the dystroglycan complex, which links the extracellular matrix to the intracellular actin cables, is thought to provide structural integrity in muscle tissues. The dystroglycan complex is also known to serve as an agrin receptor in muscle, where it may regulate agrin-induced acetylcholine receptor clustering at the neuromuscular junction. There is also evidence which suggests the function of dystroglycan as a part of the signal transduction pathway because it is shown that Grb2, a mediator of the Ras-related signal pathway, can interact with the cytoplasmic domain of dystroglycan. In general, aberrant expression of dystrophin-associated protein complex underlies the pathogenesis of Duchenne muscular dystrophy, Becker muscular dystrophy and severe childhood autosomal recessive muscular dystrophy. Interestingly, no genetic disease has been described for either [alpha]- or [beta]-dystroglycan. Dystroglycan is widely distributed in non-muscle tissues as well as in muscle tissues. During epithelial morphogenesis of kidney, the dystroglycan complex is shown to act as a receptor for the basement membrane. Dystroglycan expression in Mus musculus brain and neural retina has also been reported. However, the physiological role of dystroglycan in non-muscle tissues has remained unclear [].; PDB: 1EG4_P.
Probab=56.41  E-value=3.7  Score=41.07  Aligned_cols=24  Identities=25%  Similarity=0.497  Sum_probs=0.0

Q ss_pred             EEEehhHHHHHHHHHHHHhhhhhh
Q 037760          445 IVAMSIISGMLILGLLLGMAWKKA  468 (471)
Q Consensus       445 ii~~~v~~~~~~~~~~~~~~~~~~  468 (471)
                      |++++|+++++|.++++++++|||
T Consensus       150 IpaVVI~~iLLIA~iIa~icyrrk  173 (290)
T PF05454_consen  150 IPAVVIAAILLIAGIIACICYRRK  173 (290)
T ss_dssp             ------------------------
T ss_pred             HHHHHHHHHHHHHHHHHHHhhhhh
Confidence            344444443444444444444433


No 43 
>PF01299 Lamp:  Lysosome-associated membrane glycoprotein (Lamp);  InterPro: IPR002000 Lysosome-associated membrane glycoproteins (lamp) [] are integral membrane proteins, specific to lysosomes, and whose exact biological function is not yet clear. Structurally, the lamp proteins consist of two internally homologous lysosome-luminal domains separated by a proline-rich hinge region; at the C-terminal extremity there is a transmembrane region (TM) followed by a very short cytoplasmic tail (C). In each of the duplicated domains, there are two conserved disulphide bonds. This structure is schematically represented in the figure below.   +-----+ +-----+ +-----+ +-----+ | | | | | | | | xCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxxxCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxx +--------------------------++Hinge++--------------------------++TM++C+  In mammals, there are two closely related types of lamp: lamp-1 and lamp-2, which form major components of the lysosome membrane. In chicken lamp-1 is known as LEP100.  Also included in this entry is the macrophage protein CD68 (or macrosialin) [] is a heavily glycosylated integral membrane protein whose structure consists of a mucin-like domain followed by a proline-rich hinge; a single lamp-like domain; a transmembrane region and a short cytoplasmic tail.   Similar to CD68, mammalian lamp-3, which is expressed in lymphoid organs, dendritic cells and in lung, contains all the C-terminal regions but lacks the N-terminal lamp-like region []. In a lamp-family protein from nematodes [] only the part C-terminal to the hinge is conserved. ; GO: 0016020 membrane
Probab=53.45  E-value=7.9  Score=39.04  Aligned_cols=27  Identities=7%  Similarity=0.226  Sum_probs=14.7

Q ss_pred             EEEEehhHHHH---HHHHHHHHhhhhhhcC
Q 037760          444 IIVAMSIISGM---LILGLLLGMAWKKAKN  470 (471)
Q Consensus       444 ~ii~~~v~~~~---~~~~~~~~~~~~~~~~  470 (471)
                      .+|.++||+++   +++.++.|++.|||++
T Consensus       271 ~~vPIaVG~~La~lvlivLiaYli~Rrr~~  300 (306)
T PF01299_consen  271 DLVPIAVGAALAGLVLIVLIAYLIGRRRSR  300 (306)
T ss_pred             chHHHHHHHHHHHHHHHHHHhheeEecccc
Confidence            45555555443   2344556666676654


No 44 
>PF00008 EGF:  EGF-like domain This is a sub-family of the Pfam entry This is a sub-family of the Pfam entry;  InterPro: IPR006209 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length.; GO: 0005515 protein binding; PDB: 1WHE_A 1CCF_A 1APO_A 1WHF_A 2VJ3_A 1TOZ_A 4D90_B 3CFW_A 1EDM_B 1IXA_A ....
Probab=53.43  E-value=4.1  Score=26.45  Aligned_cols=23  Identities=26%  Similarity=0.642  Sum_probs=17.3

Q ss_pred             CCCCCccccCC--CCCccccCCCCc
Q 037760          296 QCGANSNCRIS--KTPICECLAGFI  318 (471)
Q Consensus       296 ~CG~~g~C~~~--~~~~C~C~~GF~  318 (471)
                      .|...|.|...  ....|.|++||.
T Consensus         5 ~C~n~g~C~~~~~~~y~C~C~~G~~   29 (32)
T PF00008_consen    5 PCQNGGTCIDLPGGGYTCECPPGYT   29 (32)
T ss_dssp             SSTTTEEEEEESTSEEEEEEBTTEE
T ss_pred             cCCCCeEEEeCCCCCEEeECCCCCc
Confidence            67778888743  335899999985


No 45 
>PF06697 DUF1191:  Protein of unknown function (DUF1191);  InterPro: IPR010605 This family contains hypothetical plant proteins of unknown function.
Probab=52.99  E-value=12  Score=37.01  Aligned_cols=22  Identities=27%  Similarity=0.221  Sum_probs=9.4

Q ss_pred             EEEehhHHHHHH-HHHHHHhhhh
Q 037760          445 IVAMSIISGMLI-LGLLLGMAWK  466 (471)
Q Consensus       445 ii~~~v~~~~~~-~~~~~~~~~~  466 (471)
                      ++++++|+++++ +++++++..|
T Consensus       216 v~g~~~G~~~L~ll~~lv~~~vr  238 (278)
T PF06697_consen  216 VVGVVGGVVLLGLLSLLVAMLVR  238 (278)
T ss_pred             EEEehHHHHHHHHHHHHHHhhhh
Confidence            444455554433 3333333333


No 46 
>PF12662 cEGF:  Complement Clr-like EGF-like
Probab=52.92  E-value=6.5  Score=24.04  Aligned_cols=11  Identities=27%  Similarity=0.860  Sum_probs=9.4

Q ss_pred             ccccCCCCccC
Q 037760          310 ICECLAGFISK  320 (471)
Q Consensus       310 ~C~C~~GF~~~  320 (471)
                      .|+|++||+..
T Consensus         3 ~C~C~~Gy~l~   13 (24)
T PF12662_consen    3 TCSCPPGYQLS   13 (24)
T ss_pred             EeeCCCCCcCC
Confidence            69999999864


No 47 
>PF13908 Shisa:  Wnt and FGF inhibitory regulator
Probab=50.89  E-value=12  Score=34.49  Aligned_cols=17  Identities=18%  Similarity=0.098  Sum_probs=8.2

Q ss_pred             eEEEEehhHHHHHHHHH
Q 037760          443 KIIVAMSIISGMLILGL  459 (471)
Q Consensus       443 ~~ii~~~v~~~~~~~~~  459 (471)
                      .++++|+++++++|+++
T Consensus        79 ~iivgvi~~Vi~Iv~~I   95 (179)
T PF13908_consen   79 GIIVGVICGVIAIVVLI   95 (179)
T ss_pred             eeeeehhhHHHHHHHhH
Confidence            45555555544444333


No 48 
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=50.85  E-value=14  Score=30.61  Aligned_cols=10  Identities=20%  Similarity=0.292  Sum_probs=4.7

Q ss_pred             eEEEEehhHH
Q 037760          443 KIIVAMSIIS  452 (471)
Q Consensus       443 ~~ii~~~v~~  452 (471)
                      ..|++++|++
T Consensus        66 gaiagi~vg~   75 (96)
T PTZ00382         66 GAIAGISVAV   75 (96)
T ss_pred             ccEEEEEeeh
Confidence            3455555443


No 49 
>PF08374 Protocadherin:  Protocadherin;  InterPro: IPR013585 The structure of protocadherins is similar to that of classic cadherins (IPR002126 from INTERPRO), but they also have some unique features associated with the cytoplasmic domains. They are expressed in a variety of organisms and are found in high concentrations in the brain where they seem to be localised mainly at cell-cell contact sites. Their expression seems to be developmentally regulated []. 
Probab=49.71  E-value=14  Score=35.12  Aligned_cols=15  Identities=27%  Similarity=0.219  Sum_probs=9.4

Q ss_pred             ccceeEEEEehhHHH
Q 037760          439 KKRLKIIVAMSIISG  453 (471)
Q Consensus       439 ~~~~~~ii~~~v~~~  453 (471)
                      +.+.+|+|+++.|++
T Consensus        34 ~d~~~I~iaiVAG~~   48 (221)
T PF08374_consen   34 KDYVKIMIAIVAGIM   48 (221)
T ss_pred             ccceeeeeeeecchh
Confidence            445667777776544


No 50 
>PF03302 VSP:  Giardia variant-specific surface protein;  InterPro: IPR005127 During infection, the intestinal protozoan parasite Giardia lamblia virus undergoes continuous antigenic variation which is determined by diversification of the parasite's major surface antigen, named VSP (variant surface protein).
Probab=45.97  E-value=17  Score=38.19  Aligned_cols=24  Identities=17%  Similarity=0.134  Sum_probs=13.7

Q ss_pred             cceeEEEEehhHHHHHHHHHHHHh
Q 037760          440 KRLKIIVAMSIISGMLILGLLLGM  463 (471)
Q Consensus       440 ~~~~~ii~~~v~~~~~~~~~~~~~  463 (471)
                      -....|.+|+|+++|+|-+++-||
T Consensus       364 LstgaIaGIsvavvvvVgglvGfL  387 (397)
T PF03302_consen  364 LSTGAIAGISVAVVVVVGGLVGFL  387 (397)
T ss_pred             ccccceeeeeehhHHHHHHHHHHH
Confidence            345677777777665553343333


No 51 
>PF01436 NHL:  NHL repeat;  InterPro: IPR001258 The NHL repeat, named after NCL-1, HT2A and Lin-41, is found largely in a large number of eukaryotic and prokaryotic proteins. For example, the repeat is found in a variety of enzymes of the copper type II, ascorbate-dependent monooxygenase family which catalyse the C terminus alpha-amidation of biological peptides []. In many it occurs in tandem arrays, for example in the ringfinger beta-box, coiled-coil (RBCC) eukaryotic growth regulators []. The 'Brain Tumor' protein (Brat) is one such growth regulator that contains a 6-bladed NHL-repeat beta-propeller [, ].  The NHL repeats are also found in serine/threonine protein kinase (STPK) in diverse range of pathogenic bacteria. These STPK are transmembrane receptors with a intracellular N-terminal kinase domain and extracellular C-terminal sensor domain. In the STPK, PknD, from Mycobacterium tuberculosis, the sensor domain forms a rigid, six-bladed b-propeller composed of NHL repeats with a flexible tether to the transmembrane domain.; GO: 0005515 protein binding; PDB: 3FVZ_A 3FW0_A 1RWL_A 1RWI_A 1Q7F_A.
Probab=45.14  E-value=34  Score=21.26  Aligned_cols=21  Identities=10%  Similarity=0.295  Sum_probs=15.2

Q ss_pred             eEEEccCCcEEEEeCCCceEE
Q 037760           89 VLTLSNNGSILLLNQERSTIW  109 (471)
Q Consensus        89 ~l~l~~~GnLvl~d~~~~~vW  109 (471)
                      -+.++.+|+|++.|.++..||
T Consensus         6 gvav~~~g~i~VaD~~n~rV~   26 (28)
T PF01436_consen    6 GVAVDSDGNIYVADSGNHRVQ   26 (28)
T ss_dssp             EEEEETTSEEEEEECCCTEEE
T ss_pred             EEEEeCCCCEEEEECCCCEEE
Confidence            366678888888887665555


No 52 
>smart00181 EGF Epidermal growth factor-like domain.
Probab=44.33  E-value=21  Score=22.88  Aligned_cols=24  Identities=21%  Similarity=0.695  Sum_probs=16.7

Q ss_pred             CCCCCCccccCC-CCCccccCCCCcc
Q 037760          295 AQCGANSNCRIS-KTPICECLAGFIS  319 (471)
Q Consensus       295 g~CG~~g~C~~~-~~~~C~C~~GF~~  319 (471)
                      ..|... .|... ....|.|++||.-
T Consensus         6 ~~C~~~-~C~~~~~~~~C~C~~g~~g   30 (35)
T smart00181        6 GPCSNG-TCINTPGSYTCSCPPGYTG   30 (35)
T ss_pred             CCCCCC-EEECCCCCeEeECCCCCcc
Confidence            456666 78643 3458999999964


No 53 
>TIGR01478 STEVOR variant surface antigen, stevor family. This model represents the stevor branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of stevor sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 8 bits.
Probab=42.77  E-value=11  Score=37.25  Aligned_cols=22  Identities=14%  Similarity=0.305  Sum_probs=10.1

Q ss_pred             CCCCCCCCCCCCCCCceEEecc
Q 037760          330 SRRCDRKPSDCPSGEGFLKLQR  351 (471)
Q Consensus       330 s~GC~~~~~~C~~~~~f~~~~~  351 (471)
                      -.+|.+.--.|.-+.-|+.+-|
T Consensus       169 K~rC~~gi~~CsvGSA~LT~IG  190 (295)
T TIGR01478       169 KKGCTAGVGTCALSSALLGNIG  190 (295)
T ss_pred             hccCCCeeEeeccHHHHHHHHH
Confidence            3577665223543333444333


No 54 
>PTZ00370 STEVOR; Provisional
Probab=40.74  E-value=13  Score=36.98  Aligned_cols=22  Identities=27%  Similarity=0.563  Sum_probs=10.2

Q ss_pred             CCCCCCCCCCCCCCCceEEecc
Q 037760          330 SRRCDRKPSDCPSGEGFLKLQR  351 (471)
Q Consensus       330 s~GC~~~~~~C~~~~~f~~~~~  351 (471)
                      -.+|.+.--.|.-+.-|+.+-|
T Consensus       169 K~rC~~gi~~CsVGSafLT~IG  190 (296)
T PTZ00370        169 KHRCTGGICSCSLGSALLTLIG  190 (296)
T ss_pred             hccCCCeeEeeccHHHHHHHHH
Confidence            3577665223543334444434


No 55 
>TIGR01167 LPXTG_anchor LPXTG-motif cell wall anchor domain. A common feature of this proteins containing this domain appears to be a high proportion of charged and zwitterionic residues immediatedly upstream of the LPXTG motif. This model differs from other descriptions of the LPXTG region by including a portion of that upstream charged region.
Probab=36.08  E-value=36  Score=21.96  Aligned_cols=10  Identities=20%  Similarity=-0.147  Sum_probs=4.4

Q ss_pred             HHHhhhhhhc
Q 037760          460 LLGMAWKKAK  469 (471)
Q Consensus       460 ~~~~~~~~~~  469 (471)
                      ..++++|||+
T Consensus        24 ~~~~~~~rk~   33 (34)
T TIGR01167        24 GGLLLRKRKK   33 (34)
T ss_pred             HHHHheeccc
Confidence            3344444444


No 56 
>PF13360 PQQ_2:  PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=35.68  E-value=81  Score=29.35  Aligned_cols=51  Identities=20%  Similarity=0.382  Sum_probs=26.9

Q ss_pred             CCcEEEEeC-CCceEEeecCCCCccceEE-EE---------ecCCCEEEEeCcCCCCCceeeeec
Q 037760           95 NGSILLLNQ-ERSTIWSSNSSRVLETAVV-RL---------LDSGNLVLRDNVSRSSDEYMWQSF  148 (471)
Q Consensus        95 ~GnLvl~d~-~~~~vWss~~~~~~~~~~~-~L---------ld~GNlvl~~~~~~~~~~~~WqSF  148 (471)
                      +|.|...|. +|..+|+...........+ .+         ..+|+|+..|..   +++++|+--
T Consensus         2 ~g~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~~~~v~~~~~~~~l~~~d~~---tG~~~W~~~   63 (238)
T PF13360_consen    2 DGTLSALDPRTGKELWSYDLGPGIGGPVATAVPDGGRVYVASGDGNLYALDAK---TGKVLWRFD   63 (238)
T ss_dssp             TSEEEEEETTTTEEEEEEECSSSCSSEEETEEEETTEEEEEETTSEEEEEETT---TSEEEEEEE
T ss_pred             CCEEEEEECCCCCEEEEEECCCCCCCccceEEEeCCEEEEEcCCCEEEEEECC---CCCEEEEee
Confidence            566777775 6777787754110111111 22         255555555532   567788754


No 57 
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=35.01  E-value=1.7e+02  Score=30.13  Aligned_cols=20  Identities=20%  Similarity=0.599  Sum_probs=10.2

Q ss_pred             ccCCcEEEEeC-CCceEEeec
Q 037760           93 SNNGSILLLNQ-ERSTIWSSN  112 (471)
Q Consensus        93 ~~~GnLvl~d~-~~~~vWss~  112 (471)
                      ..+|.|+-+|. +|+++|...
T Consensus       263 ~~~g~l~ald~~tG~~~W~~~  283 (394)
T PRK11138        263 AYNGNLVALDLRSGQIVWKRE  283 (394)
T ss_pred             EcCCeEEEEECCCCCEEEeec
Confidence            34555555554 345566543


No 58 
>cd05845 Ig2_L1-CAM_like Second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM) and similar proteins. Ig2_L1-CAM_like: domain similar to the second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM). L1 belongs to the L1 subfamily of cell adhesion molecules (CAMs) and is comprised of an extracellular region having six Ig-like domains, five fibronectin type III domains, a transmembrane region and an intracellular domain. L1 is primarily expressed in the nervous system and is involved in its development and function. L1 is associated with an X-linked recessive disorder, X-linked hydrocephalus, MASA syndrome, or spastic paraplegia type 1, that involves abnormalities of axonal growth.
Probab=32.72  E-value=77  Score=26.16  Aligned_cols=33  Identities=21%  Similarity=0.530  Sum_probs=22.3

Q ss_pred             CCCeEEEEecCCCCCCCCCceEEEccCCcEEEEe
Q 037760           69 SPRTVVWVANRYKPITDKNGVLTLSNNGSILLLN  102 (471)
Q Consensus        69 ~~~~~vW~an~~~pv~~~~~~l~l~~~GnLvl~d  102 (471)
                      |..++.|.-+....+. ....+.++.+|||.+.+
T Consensus        32 P~P~i~W~~~~~~~i~-~~~Ri~~~~~GnL~fs~   64 (95)
T cd05845          32 VPLRIYWMNSDLLHIT-QDERVSMGQNGNLYFAN   64 (95)
T ss_pred             CCCEEEEECCCCcccc-ccccEEECCCceEEEEE
Confidence            5667889855434443 35678888889998854


No 59 
>PF15345 TMEM51:  Transmembrane protein 51
Probab=32.62  E-value=50  Score=31.85  Aligned_cols=15  Identities=27%  Similarity=0.483  Sum_probs=7.1

Q ss_pred             HHHHHHHHHhhhhhh
Q 037760          454 MLILGLLLGMAWKKA  468 (471)
Q Consensus       454 ~~~~~~~~~~~~~~~  468 (471)
                      ++++++|+-+.-|||
T Consensus        71 LLLLSICL~IR~KRr   85 (233)
T PF15345_consen   71 LLLLSICLSIRDKRR   85 (233)
T ss_pred             HHHHHHHHHHHHHHH
Confidence            344566554433333


No 60 
>PF13360 PQQ_2:  PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=32.49  E-value=2.3e+02  Score=26.22  Aligned_cols=75  Identities=17%  Similarity=0.363  Sum_probs=43.0

Q ss_pred             CCeEEEEecC----CCCC---CCCCceEEE-ccCCcEEEEeC-CCceEEeecCCCCccce-------EEEEecCCCEEEE
Q 037760           70 PRTVVWVANR----YKPI---TDKNGVLTL-SNNGSILLLNQ-ERSTIWSSNSSRVLETA-------VVRLLDSGNLVLR  133 (471)
Q Consensus        70 ~~~~vW~an~----~~pv---~~~~~~l~l-~~~GnLvl~d~-~~~~vWss~~~~~~~~~-------~~~Lld~GNlvl~  133 (471)
                      ....+|..+-    ..++   ...+..+.+ +.+|.|+.+|. .|..+|+....+....+       ......+|.|...
T Consensus        12 tG~~~W~~~~~~~~~~~~~~~~~~~~~v~~~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~~   91 (238)
T PF13360_consen   12 TGKELWSYDLGPGIGGPVATAVPDGGRVYVASGDGNLYALDAKTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYAL   91 (238)
T ss_dssp             TTEEEEEEECSSSCSSEEETEEEETTEEEEEETTSEEEEEETTTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEEE
T ss_pred             CCCEEEEEECCCCCCCccceEEEeCCEEEEEcCCCEEEEEECCCCCEEEEeeccccccceeeecccccccccceeeeEec
Confidence            5678887653    2221   101233333 47889999996 88999998764321111       1122234456666


Q ss_pred             eCcCCCCCceeeee
Q 037760          134 DNVSRSSDEYMWQS  147 (471)
Q Consensus       134 ~~~~~~~~~~~WqS  147 (471)
                      |.   .++.++|+.
T Consensus        92 d~---~tG~~~W~~  102 (238)
T PF13360_consen   92 DA---KTGKVLWSI  102 (238)
T ss_dssp             ET---TTSCEEEEE
T ss_pred             cc---CCcceeeee
Confidence            63   267899995


No 61 
>PF10681 Rot1:  Chaperone for protein-folding within the ER, fungal;  InterPro: IPR019623  This conserved fungal family is an essential molecular chaperone in the endoplasmic reticulum. Molecular chaperones transiently interact with unfolded proteins to inhibit their self-aggregation and to support their folding and/or assembly. Rot1 is a general chaperone with some substrate specificity, its substrates being the structurally unrelated Kre5 Kre6 Big1 Atg22, which are type I, type II, and polytopic membrane proteins. The dependencies of each for Rot1 do not share similarities. However, their folding does require BiP, and one of these proteins was simultaneously associated with both Rot1 and BiP. In addition, Rot1 may cooperate with BiP/Kar2 in the folding of Kre6 []. 
Probab=30.64  E-value=2.7e+02  Score=26.54  Aligned_cols=74  Identities=16%  Similarity=0.264  Sum_probs=44.8

Q ss_pred             ceEEeecCCCCccceEEEEecCCCEEEEeCcCCCCCceeeeeccCCCCCCCCCCeeeeeccCCceeEEEEecCCCCCCCc
Q 037760          106 STIWSSNSSRVLETAVVRLLDSGNLVLRDNVSRSSDEYMWQSFDYPSDTLLPGMKLGWNLRTRFERYLTAWRNADDPTPG  185 (471)
Q Consensus       106 ~~vWss~~~~~~~~~~~~Lld~GNlvl~~~~~~~~~~~~WqSFd~PTDTLLPGq~L~~~~~tg~~~~L~Sw~s~~dps~G  185 (471)
                      ..+|+-.        .-+|+++|.|+|.--  ..+++   |-+.+|...= -..--..+    +...+.+|.-..|+-.|
T Consensus        64 ~l~wQHG--------tY~l~~nGsl~L~P~--~~DGr---Ql~sdPC~~~-~s~y~rYn----q~e~f~~~~v~~D~y~~  125 (212)
T PF10681_consen   64 VLIWQHG--------TYELNSNGSLTLTPF--AVDGR---QLVSDPCADD-SSTYTRYN----QTELFKSFDVYVDPYHG  125 (212)
T ss_pred             EEEEecc--------eEEECCCCcEEEeec--CCCCc---eeccCCCCCC-cccEEEEc----ceEEEEEEEEEEeCCCC
Confidence            3567743        347788999998754  12444   4456776522 11111222    33567778777899899


Q ss_pred             eEEEEE-ccCCCe
Q 037760          186 EFSFRF-DISTMA  197 (471)
Q Consensus       186 ~f~l~l-~~~g~~  197 (471)
                      .|+|.| +.+|.|
T Consensus       126 ~~~L~L~~fDGsp  138 (212)
T PF10681_consen  126 RYRLQLYQFDGSP  138 (212)
T ss_pred             eeEEEEEccCCCc
Confidence            999988 356643


No 62 
>PF01102 Glycophorin_A:  Glycophorin A;  InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=29.99  E-value=15  Score=31.96  Aligned_cols=27  Identities=11%  Similarity=0.218  Sum_probs=18.5

Q ss_pred             EEEehhHHHHHHHHHHHHhhhhhhcCC
Q 037760          445 IVAMSIISGMLILGLLLGMAWKKAKNK  471 (471)
Q Consensus       445 ii~~~v~~~~~~~~~~~~~~~~~~~~~  471 (471)
                      |+++++|+++-++++++++.+..||+|
T Consensus        66 i~~Ii~gv~aGvIg~Illi~y~irR~~   92 (122)
T PF01102_consen   66 IIGIIFGVMAGVIGIILLISYCIRRLR   92 (122)
T ss_dssp             HHHHHHHHHHHHHHHHHHHHHHHHHHS
T ss_pred             eeehhHHHHHHHHHHHHHHHHHHHHHh
Confidence            444555555557788887788888875


No 63 
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=29.90  E-value=2.1e+02  Score=29.03  Aligned_cols=20  Identities=15%  Similarity=0.606  Sum_probs=10.4

Q ss_pred             cCCcEEEEe-CCCceEEeecC
Q 037760           94 NNGSILLLN-QERSTIWSSNS  113 (471)
Q Consensus        94 ~~GnLvl~d-~~~~~vWss~~  113 (471)
                      .+|.|+-.| .+|.++|+...
T Consensus        73 ~~g~v~a~d~~tG~~~W~~~~   93 (377)
T TIGR03300        73 ADGTVVALDAETGKRLWRVDL   93 (377)
T ss_pred             CCCeEEEEEccCCcEeeeecC
Confidence            345555555 35556665443


No 64 
>PF06365 CD34_antigen:  CD34/Podocalyxin family;  InterPro: IPR013836 This family consists of several mammalian CD34 antigen proteins. The CD34 antigen is a human leukocyte membrane protein expressed specifically by lymphohematopoietic progenitor cells. CD34 is a phosphoprotein. Activation of protein kinase C (PKC) has been found to enhance CD34 phosphorylation [, ]. This family contains several eukaryotic podocalyxin proteins. Podocalyxin is a major membrane protein of the glomerular epithelium and is thought to be involved in maintenance of the architecture of the foot processes and filtration slits characteristic of this unique epithelium by virtue of its high negative charge. Podocalyxin functions as an anti-adhesin that maintains an open filtration pathway between neighbouring foot processes in the glomerular epithelium by charge repulsion [].
Probab=29.54  E-value=61  Score=30.74  Aligned_cols=26  Identities=12%  Similarity=-0.038  Sum_probs=13.1

Q ss_pred             EEEEehhHH--H-HHHHHHHHHhhhhhhc
Q 037760          444 IIVAMSIIS--G-MLILGLLLGMAWKKAK  469 (471)
Q Consensus       444 ~ii~~~v~~--~-~~~~~~~~~~~~~~~~  469 (471)
                      .+|++++..  . +++++..+|++|.||.
T Consensus       101 ~lI~lv~~g~~lLla~~~~~~Y~~~~Rrs  129 (202)
T PF06365_consen  101 TLIALVTSGSFLLLAILLGAGYCCHQRRS  129 (202)
T ss_pred             EEEehHHhhHHHHHHHHHHHHHHhhhhcc
Confidence            555544433  2 2234556677776653


No 65 
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=29.05  E-value=2.7e+02  Score=28.32  Aligned_cols=75  Identities=19%  Similarity=0.416  Sum_probs=44.0

Q ss_pred             CCCeEEEEecCCCC-----CCCCCceEEE-ccCCcEEEEeC-CCceEEeecCCCCc-c------ceEEEEecCCCEEEEe
Q 037760           69 SPRTVVWVANRYKP-----ITDKNGVLTL-SNNGSILLLNQ-ERSTIWSSNSSRVL-E------TAVVRLLDSGNLVLRD  134 (471)
Q Consensus        69 ~~~~~vW~an~~~p-----v~~~~~~l~l-~~~GnLvl~d~-~~~~vWss~~~~~~-~------~~~~~Lld~GNlvl~~  134 (471)
                      ....++|..+-..+     +.+ +..+.+ +.+|.|+-+|. +|..+|+....+.. .      .....-..+|.|+..|
T Consensus        83 ~tG~~~W~~~~~~~~~~~p~v~-~~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~a~d  161 (377)
T TIGR03300        83 ETGKRLWRVDLDERLSGGVGAD-GGLVFVGTEKGEVIALDAEDGKELWRAKLSSEVLSPPLVANGLVVVRTNDGRLTALD  161 (377)
T ss_pred             cCCcEeeeecCCCCcccceEEc-CCEEEEEcCCCEEEEEECCCCcEeeeeccCceeecCCEEECCEEEEECCCCeEEEEE
Confidence            45678997654433     222 233434 56788888886 68899987653311 0      1112223566677776


Q ss_pred             CcCCCCCceeeee
Q 037760          135 NVSRSSDEYMWQS  147 (471)
Q Consensus       135 ~~~~~~~~~~WqS  147 (471)
                      ..   +++++|+-
T Consensus       162 ~~---tG~~~W~~  171 (377)
T TIGR03300       162 AA---TGERLWTY  171 (377)
T ss_pred             cC---CCceeeEE
Confidence            52   56788984


No 66 
>KOG3637 consensus Vitronectin receptor, alpha subunit [Extracellular structures]
Probab=26.80  E-value=48  Score=39.21  Aligned_cols=15  Identities=47%  Similarity=1.010  Sum_probs=9.8

Q ss_pred             HHHHHHHHHHHhhhh
Q 037760          452 SGMLILGLLLGMAWK  466 (471)
Q Consensus       452 ~~~~~~~~~~~~~~~  466 (471)
                      +.+|+++++++++||
T Consensus       987 ~GLLlL~llv~~LwK 1001 (1030)
T KOG3637|consen  987 GGLLLLALLVLLLWK 1001 (1030)
T ss_pred             HHHHHHHHHHHHHHh
Confidence            336666777777776


No 67 
>KOG4649 consensus PQQ (pyrrolo-quinoline quinone) repeat protein [Secondary metabolites biosynthesis, transport and catabolism]
Probab=25.98  E-value=2.3e+02  Score=28.16  Aligned_cols=46  Identities=17%  Similarity=0.485  Sum_probs=33.6

Q ss_pred             CCCeEEEEecCCCCCCCCC----ceEEE-ccCCcEEEEeCCCceEEeecCC
Q 037760           69 SPRTVVWVANRYKPITDKN----GVLTL-SNNGSILLLNQERSTIWSSNSS  114 (471)
Q Consensus        69 ~~~~~vW~an~~~pv~~~~----~~l~l-~~~GnLvl~d~~~~~vWss~~~  114 (471)
                      .+.+.+|.+.+..|+-.+.    ..+.+ +-||+|.-.|..|+.||+-.+.
T Consensus       167 ~~~~~~w~~~~~~PiF~splcv~~sv~i~~VdG~l~~f~~sG~qvwr~~t~  217 (354)
T KOG4649|consen  167 YSSTEFWAATRFGPIFASPLCVGSSVIITTVDGVLTSFDESGRQVWRPATK  217 (354)
T ss_pred             CCcceehhhhcCCccccCceeccceEEEEEeccEEEEEcCCCcEEEeecCC
Confidence            3457899999999987542    23333 4689998888888889976554


No 68 
>PF10661 EssA:  WXG100 protein secretion system (Wss), protein EssA;  InterPro: IPR018920  The Wss (WXG100 protein secretion system) in Staphylococcus aureus seems to be encoded by a locus of eight ORFs, called ess (eSAT-6 secretion system) []. This locus encodes, amongst several other proteins, EssA, a protein predicted to possess one transmembrane domain. Due to its predicted membrane location and its absolute requirement for WXG100 protein secretion, it has been speculated that EssA could form a secretion apparatus in conjunction with YukC and YukAB. Proteins homologous to EssA, YukC, EsaA and YukD were absent from mycobacteria [].   Members of this family are associated with type VII secretion of WXG100 family targets in the Firmicutes, but not in the Actinobacteria. This highly divergent protein family consists largely of a central region of highly polar low-complexity sequence containing occasional LF motifs in weak repeats about 17 residues in length, flanked by hydrophobic N- and C-terminal regions. 
Probab=25.95  E-value=20  Score=32.13  Aligned_cols=23  Identities=17%  Similarity=0.081  Sum_probs=14.2

Q ss_pred             EEEehhHHHHHHHHHHHHhhhhh
Q 037760          445 IVAMSIISGMLILGLLLGMAWKK  467 (471)
Q Consensus       445 ii~~~v~~~~~~~~~~~~~~~~~  467 (471)
                      +|+++|+++|+++++.+|.+.|+
T Consensus       120 ~i~~~i~g~ll~i~~giy~~~r~  142 (145)
T PF10661_consen  120 TILLSIGGILLAICGGIYVVLRK  142 (145)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHH
Confidence            44455555566677777776664


No 69 
>KOG1219 consensus Uncharacterized conserved protein, contains laminin, cadherin and EGF domains [Signal transduction mechanisms]
Probab=25.27  E-value=98  Score=39.61  Aligned_cols=25  Identities=16%  Similarity=0.522  Sum_probs=16.6

Q ss_pred             CCCCCccccCCCC--CccccCCCCccC
Q 037760          296 QCGANSNCRISKT--PICECLAGFISK  320 (471)
Q Consensus       296 ~CG~~g~C~~~~~--~~C~C~~GF~~~  320 (471)
                      .|-.-|.|+....  -.|.||+.|.-.
T Consensus      3871 pCqhgG~C~~~~~ggy~CkCpsqysG~ 3897 (4289)
T KOG1219|consen 3871 PCQHGGTCISQPKGGYKCKCPSQYSGN 3897 (4289)
T ss_pred             cccCCCEecCCCCCceEEeCcccccCc
Confidence            3444578875433  389999988643


No 70 
>PF05545 FixQ:  Cbb3-type cytochrome oxidase component FixQ;  InterPro: IPR008621 This family consists of several Cbb3-type cytochrome oxidase components (FixQ/CcoQ). FixQ is found in nitrogen fixing bacteria. Since nitrogen fixation is an energy-consuming process, effective symbioses depend on operation of a respiratory chain with a high affinity for O2, closely coupled to ATP production. This requirement is fulfilled by a special three-subunit terminal oxidase (cytochrome terminal oxidase cbb3), which was first identified in Bradyrhizobium japonicum as the product of the fixNOQP operon [].
Probab=25.02  E-value=64  Score=22.99  Aligned_cols=13  Identities=15%  Similarity=0.291  Sum_probs=5.5

Q ss_pred             HHHHHhhhhhhcC
Q 037760          458 GLLLGMAWKKAKN  470 (471)
Q Consensus       458 ~~~~~~~~~~~~~  470 (471)
                      +++.+..++++|+
T Consensus        24 gi~~w~~~~~~k~   36 (49)
T PF05545_consen   24 GIVIWAYRPRNKK   36 (49)
T ss_pred             HHHHHHHcccchh
Confidence            4444444444443


No 71 
>PF05393 Hum_adeno_E3A:  Human adenovirus early E3A glycoprotein;  InterPro: IPR008652 This family consists of several early glycoproteins (E3A), from human adenovirus type 2.; GO: 0016021 integral to membrane
Probab=25.00  E-value=59  Score=26.45  Aligned_cols=8  Identities=38%  Similarity=0.409  Sum_probs=3.3

Q ss_pred             HHHHHHHh
Q 037760          456 ILGLLLGM  463 (471)
Q Consensus       456 ~~~~~~~~  463 (471)
                      ++.+++|+
T Consensus        45 il~Vilwf   52 (94)
T PF05393_consen   45 ILLVILWF   52 (94)
T ss_pred             HHHHHHHH
Confidence            33444444


No 72 
>TIGR03503 conserved hypothetical protein TIGR03503. This set of conserved hypothetical protein has a phylogenetic range that closely matches that of TIGR03501, a putative C-terminal protein targeting signal.
Probab=24.65  E-value=12  Score=38.85  Aligned_cols=21  Identities=29%  Similarity=0.330  Sum_probs=13.3

Q ss_pred             hhHHHHHHHHHHHHhhhhhhc
Q 037760          449 SIISGMLILGLLLGMAWKKAK  469 (471)
Q Consensus       449 ~v~~~~~~~~~~~~~~~~~~~  469 (471)
                      ++.+++++++++.+++|||||
T Consensus       353 ~~N~v~lllg~~~~~~~rk~k  373 (374)
T TIGR03503       353 VGNVVILLLGGIGFFVWRKKK  373 (374)
T ss_pred             hhhhhhhhhheeeEEEEEEee
Confidence            334445667777777787776


No 73 
>PF02480 Herpes_gE:  Alphaherpesvirus glycoprotein E;  InterPro: IPR003404 Glycoprotein E (gE) of Alphaherpesvirus forms a complex with glycoprotein I (gI), functioning as an immunoglobulin G (IgG) Fc binding protein. gE is involved in virus spread but is not essential for propagation [].; GO: 0016020 membrane; PDB: 2GJ7_F 2GIY_B.
Probab=21.62  E-value=31  Score=36.75  Aligned_cols=13  Identities=31%  Similarity=0.856  Sum_probs=9.6

Q ss_pred             CCCCCCCC--CCCCC
Q 037760          329 YSRRCDRK--PSDCP  341 (471)
Q Consensus       329 ~s~GC~~~--~~~C~  341 (471)
                      ...+|.+.  +..|.
T Consensus       237 ~y~~C~~~~~~~~C~  251 (439)
T PF02480_consen  237 RYANCSPSGWPRRCP  251 (439)
T ss_dssp             EEEEEBTTC-TTTTE
T ss_pred             hhcCCCCCCCcCCCC
Confidence            46789886  67894


No 74 
>PF12301 CD99L2:  CD99 antigen like protein 2;  InterPro: IPR022078  This family of proteins is found in eukaryotes. Proteins in this family are typically between 165 and 237 amino acids in length. CD99L2 and CD99 are involved in trans-endothelial migration of neutrophils in vitro and in the recruitment of neutrophils into inflamed peritoneum. 
Probab=21.41  E-value=65  Score=29.68  Aligned_cols=24  Identities=8%  Similarity=0.088  Sum_probs=12.7

Q ss_pred             ceeEEEEehhHHHHHHHHHHHHhh
Q 037760          441 RLKIIVAMSIISGMLILGLLLGMA  464 (471)
Q Consensus       441 ~~~~ii~~~v~~~~~~~~~~~~~~  464 (471)
                      ..++|.+|+-+++++|++.+.-++
T Consensus       113 ~~g~IaGIvsav~valvGAvsSyi  136 (169)
T PF12301_consen  113 EAGTIAGIVSAVVVALVGAVSSYI  136 (169)
T ss_pred             ccchhhhHHHHHHHHHHHHHHHHH
Confidence            344555555555566665544443


No 75 
>PF05092 PIF:  Per os infectivity;  InterPro: IPR007784 This entry represents a group of dsDNA Baculovirus proteins. It is required for the infectivity of the OBs or occlusion bodies. It is a structural protein of the ODV envelope required only in the first steps of per os larva infection, as viruses being produced in cells expressing the gene for this protein but not containing it in their genomes are able to produce successful infections. Baculoviruses are large DNA viruses that infect arthropods, mainly members of the order Lepidoptera. In their life cycle, they produce two kinds of particles, a budded, non-occluded virus (BV), which buds out of the infected cell and is responsible for the cell-to-cell transmission of the virus, and an occluded form, the occlusion body (OB), which is responsible for protecting the virus between encounters with larvae. A variable number of virions are included in the para-crystalline structure of the OB, mainly constituted by the virus-encoded polyhedrin protein; these virions are called occlusion body-derived virions or ODVs [].
Probab=20.47  E-value=77  Score=34.25  Aligned_cols=47  Identities=23%  Similarity=0.600  Sum_probs=34.0

Q ss_pred             cCCCCcccCCCCCCcccc-CCCCC-ccccCCCCccCCCCCCCCCCCCCCCCC
Q 037760          287 PFDACDNYAQCGANSNCR-ISKTP-ICECLAGFISKPQDDWDSPYSRRCDRK  336 (471)
Q Consensus       287 p~~~C~~~g~CG~~g~C~-~~~~~-~C~C~~GF~~~~~~~w~~~~s~GC~~~  336 (471)
                      .++.|+++--|.|+|.=. .+..| .|.|.+||.+..-.+   ....=|++.
T Consensus       147 iy~DC~vpVGC~PhG~I~din~~pi~C~Cd~GyVsd~~~~---t~tP~CRp~  195 (522)
T PF05092_consen  147 IYEDCDVPVGCQPHGRIADINESPIRCVCDDGYVSDFDSD---TETPYCRPR  195 (522)
T ss_pred             hhccCCCcEecCCCCEEeeecCCceEeECCCCcccccccC---CCCcceece
Confidence            457799999999998754 55556 899999998763111   246778765


No 76 
>PF05283 MGC-24:  Multi-glycosylated core protein 24 (MGC-24);  InterPro: IPR007947 CD164 is a mucin-like receptor, or sialomucin, with specificity in receptor/ ligand interactions that depends on the structural characteristics of the mucin-like receptor. Its functions include mediating, or regulating, haematopoietic progenitor cell adhesion and the negative regulation of their growth and/or-differentiation. It exists in the native state as a disulphide- linked homodimer of two 80-85kDa subunits. It is usually expressed by CD34+ and CD341o/- haematopoietic stem cells and associated microenvironmental cells. It contains, in its extracellular region, two mucin domains (I and II) linked by a non-mucin domain, which has been predicted to contain intra- disulphide bridges. This receptor may play a key role in haematopoiesis by facilitating the adhesion of human CD34+ cells to bone marrow stroma and by negatively regulating CD34+ CD341o/- haematopoietic progenitor cell proliferation. These effects involve the CD164 class I and/or II epitopes recognised by the monoclonal antibodies (mAbs) 105A5 and 103B2/9E10. These epitopes are carbohydrate-dependent and are located on the N-terminal mucin domain I [, ]. It has been found that murine MGC-24v and rat endolyn share significant sequence similarities with human CD164. However, CD164 lacks the consensus glycosaminoglycan (GAG)-attachment site found in MGC-24; it is possible that GAG-association is responsible for the high molecular weight of the epithelial-derived MGC-24 glycoprotein [].  Genomic structure studies have placed CD164 within the mucin-subgroup that comprises multiple exons, and demonstrate the diverse chromosomal distribution of this family of molecules. Molecules with such multiple exons may have sophisticated regulatory mechanisms that involve not only post-translational modifications of the oligosaccharide side chains, but also differential exon usage. Although differences in the intron and exon sizes are seen between the mouse and human genes, the predicted proteins are similar in size and structure, maintaining functionally important motifs that regulate cell proliferation or subcellular distribution [].  CD164 is a gene whose expression depends on differential usage of poly- adenylation sites within the 3'-UTR. The conserved distribution of the 3.2- and 1.2-kb CD164 transcripts between mouse and human suggests that (i) a mechanism may exist to regulate tissue-specific polyadenylation, and (ii) differences in polyadenylation are important for the expression and function of CD164 in different tissues. Two other aspects of the structure of CD164 are of particular interest. First, it shares one of several conserved features of a cytokine-binding pocket - in this respect, it is notable that evidence exists for a class of cell-surface sialomucin modulators that directly interact with growth factor receptors to regulate their response to physiological ligands. Second, its cytoplasmic tail contains a C-terminal YHTL motif found in many endocytic membrane proteins or receptors. These Tyr-based motifs bind to adaptor proteins, which mediate the sorting of membrane proteins into transport vesicles from the plasma membrane to the endosomes, and between intracellular compartments. 
Probab=20.10  E-value=67  Score=30.06  Aligned_cols=24  Identities=29%  Similarity=0.274  Sum_probs=13.9

Q ss_pred             EEehhHHHHHHHHH--HHHhhhhhhc
Q 037760          446 VAMSIISGMLILGL--LLGMAWKKAK  469 (471)
Q Consensus       446 i~~~v~~~~~~~~~--~~~~~~~~~~  469 (471)
                      .+..||.+||++|+  ++|+++|..|
T Consensus       160 ~~SFiGGIVL~LGv~aI~ff~~KF~k  185 (186)
T PF05283_consen  160 AASFIGGIVLTLGVLAIIFFLYKFCK  185 (186)
T ss_pred             hhhhhhHHHHHHHHHHHHHHHhhhcc
Confidence            44666666666554  5556666544


Done!