Query         042843
Match_columns 484
No_of_seqs    193 out of 1495
Neff          7.7 
Searched_HMMs 46136
Date          Fri Mar 29 10:07:29 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/042843.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/042843hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF01453 B_lectin:  D-mannose b 100.0 6.7E-30 1.5E-34  219.6   3.9  111   77-190     1-114 (114)
  2 PF00954 S_locus_glycop:  S-loc  99.9 1.2E-27 2.6E-32  204.7  11.1  110  217-329     1-110 (110)
  3 cd00028 B_lectin Bulb-type man  99.9 5.7E-26 1.2E-30  196.2  14.9  113   36-158     2-116 (116)
  4 smart00108 B_lectin Bulb-type   99.9 7.5E-25 1.6E-29  188.6  14.3  112   36-157     2-114 (114)
  5 PF08276 PAN_2:  PAN-like domai  99.5 6.3E-15 1.4E-19  114.0   5.0   58  359-416     3-66  (66)
  6 cd01098 PAN_AP_plant Plant PAN  99.5 1.1E-13 2.3E-18  112.0   8.2   81  347-431     2-84  (84)
  7 cd00129 PAN_APPLE PAN/APPLE-li  99.4 6.6E-13 1.4E-17  106.2   6.1   67  360-430     8-80  (80)
  8 smart00108 B_lectin Bulb-type   98.6 1.8E-07 3.8E-12   80.4   9.7   87   97-219    23-111 (114)
  9 cd00028 B_lectin Bulb-type man  98.6 4.2E-07 9.1E-12   78.3  10.3   83  102-220    30-113 (116)
 10 smart00473 PAN_AP divergent su  98.6 2.1E-07 4.6E-12   73.4   7.5   70  360-429     3-77  (78)
 11 PF01453 B_lectin:  D-mannose b  98.1 5.7E-05 1.2E-09   64.9  11.4  100   41-159    11-114 (114)
 12 cd01100 APPLE_Factor_XI_like S  97.4  0.0002 4.3E-09   56.3   3.8   47  366-412     9-58  (73)
 13 PF04478 Mid2:  Mid2 like cell   93.7    0.12 2.6E-06   46.1   5.0   33  441-474    47-79  (154)
 14 PF08693 SKG6:  Transmembrane a  91.9    0.29 6.3E-06   33.5   3.8   10  464-473    30-39  (40)
 15 PF15102 TMEM154:  TMEM154 prot  91.6    0.26 5.7E-06   43.6   4.4   25  445-469    58-82  (146)
 16 smart00223 APPLE APPLE domain.  91.2    0.27 5.8E-06   39.2   3.7   45  367-411     7-57  (79)
 17 PF00024 PAN_1:  PAN domain Thi  91.0    0.33 7.1E-06   37.7   4.1   54  362-415     3-60  (79)
 18 PF08277 PAN_3:  PAN-like domai  90.1     1.5 3.3E-05   33.6   7.0   39  380-418    19-58  (71)
 19 smart00605 CW CW domain.        87.0     2.1 4.6E-05   35.0   6.3   53  380-432    21-76  (94)
 20 PF02009 Rifin_STEVOR:  Rifin/s  85.4    0.14   3E-06   51.3  -1.9   20  459-478   269-288 (299)
 21 PTZ00382 Variant-specific surf  84.9     1.4 3.1E-05   36.5   4.2   29  444-472    67-95  (96)
 22 PF07645 EGF_CA:  Calcium-bindi  84.4    0.51 1.1E-05   32.6   1.1   31  297-327     3-35  (42)
 23 PF01034 Syndecan:  Syndecan do  84.2    0.32 6.9E-06   36.8   0.0   16  462-477    27-42  (64)
 24 PF14295 PAN_4:  PAN domain; PD  83.7     0.9   2E-05   32.2   2.3   24  380-403    15-38  (51)
 25 cd00053 EGF Epidermal growth f  81.1     1.2 2.7E-05   28.4   2.0   29  299-327     2-31  (36)
 26 PF01102 Glycophorin_A:  Glycop  80.8    0.37 8.1E-06   41.7  -0.8   32  445-477    66-97  (122)
 27 smart00179 EGF_CA Calcium-bind  75.8     2.2 4.8E-05   28.1   2.0   30  297-326     3-33  (39)
 28 KOG4649 PQQ (pyrrolo-quinoline  75.6      10 0.00023   37.1   7.2   46   77-122   168-218 (354)
 29 PF01299 Lamp:  Lysosome-associ  74.8     1.8 3.9E-05   43.7   2.0   31  444-475   271-301 (306)
 30 cd01099 PAN_AP_HGF Subfamily o  74.1     5.8 0.00013   31.4   4.4   33  380-412    24-60  (80)
 31 PF01683 EB:  EB module;  Inter  73.8     2.9 6.2E-05   30.1   2.3   33  293-328    16-48  (52)
 32 PTZ00046 rifin; Provisional     72.2    0.81 1.8E-05   46.6  -1.2   18  461-478   330-347 (358)
 33 TIGR01477 RIFIN variant surfac  72.2    0.81 1.7E-05   46.5  -1.2   18  461-478   325-342 (353)
 34 PF09064 Tme5_EGF_like:  Thromb  71.5     2.2 4.9E-05   28.0   1.1   18  310-327    11-28  (34)
 35 PF12661 hEGF:  Human growth fa  70.3     1.2 2.6E-05   22.9  -0.2    9  318-326     1-9   (13)
 36 cd00054 EGF_CA Calcium-binding  69.7     3.9 8.4E-05   26.4   2.1   30  297-326     3-33  (38)
 37 PF07974 EGF_2:  EGF-like domai  69.1     3.8 8.2E-05   26.7   1.8   23  303-326     6-28  (32)
 38 PHA03265 envelope glycoprotein  67.4     2.3 5.1E-05   42.9   0.8   39  443-482   347-385 (402)
 39 PF00008 EGF:  EGF-like domain   64.1     2.5 5.4E-05   27.3   0.2   23  304-326     5-29  (32)
 40 PF12662 cEGF:  Complement Clr-  60.7     3.7   8E-05   24.9   0.5   11  318-328     3-13  (24)
 41 PF02439 Adeno_E3_CR2:  Adenovi  58.3     2.1 4.5E-05   28.9  -0.9    8  446-453     6-13  (38)
 42 PF12947 EGF_3:  EGF domain;  I  57.8     3.3 7.2E-05   27.7  -0.0   24  303-326     6-30  (36)
 43 smart00181 EGF Epidermal growt  52.2      12 0.00025   24.1   1.9   24  303-327     6-30  (35)
 44 PF06024 DUF912:  Nucleopolyhed  52.1     9.3  0.0002   31.9   1.8   12  465-476    83-94  (101)
 45 PF12877 DUF3827:  Domain of un  51.1      10 0.00023   41.4   2.4   15  443-457   269-283 (684)
 46 PF03302 VSP:  Giardia variant-  49.9      18 0.00039   37.9   3.9   30  443-472   367-396 (397)
 47 PF14610 DUF4448:  Protein of u  49.5      17 0.00036   33.9   3.2   22  448-470   161-182 (189)
 48 PF13360 PQQ_2:  PQQ-like domai  48.4 1.6E+02  0.0035   27.4  10.0   77   76-154    54-148 (238)
 49 KOG0291 WD40-repeat-containing  47.4 4.8E+02    0.01   29.7  15.7   85   95-203   353-448 (893)
 50 PRK11138 outer membrane biogen  46.9      76  0.0016   32.8   8.1   57   96-154   121-186 (394)
 51 PF08374 Protocadherin:  Protoc  45.3      13 0.00028   35.2   1.8   11  443-453    38-48  (221)
 52 TIGR01478 STEVOR variant surfa  44.6     7.1 0.00015   38.5  -0.1    7  317-323   143-149 (295)
 53 PTZ00370 STEVOR; Provisional    40.4       8 0.00017   38.2  -0.4    7  317-323   143-149 (296)
 54 TIGR03300 assembly_YfgL outer   39.7 1.1E+02  0.0023   31.3   7.8   55   97-153    67-130 (377)
 55 PF02480 Herpes_gE:  Alphaherpe  39.5     9.8 0.00021   40.4   0.0   13  339-351   238-251 (439)
 56 PF06365 CD34_antigen:  CD34/Po  39.3      17 0.00038   34.2   1.6   18  392-409    36-54  (202)
 57 PF01436 NHL:  NHL repeat;  Int  38.4      56  0.0012   20.2   3.4   20   97-116     6-26  (28)
 58 cd05845 Ig2_L1-CAM_like Second  38.2      46   0.001   27.4   3.8   31   76-108    32-63  (95)
 59 PRK11138 outer membrane biogen  37.6 4.7E+02    0.01   26.9  17.2   75   77-153   139-230 (394)
 60 PF06697 DUF1191:  Protein of u  36.7      15 0.00031   36.5   0.7    8  425-432   187-194 (278)
 61 PF00954 S_locus_glycop:  S-loc  36.1      84  0.0018   26.2   5.3   58  249-312    42-101 (110)
 62 PF14991 MLANA:  Protein melan-  34.6      12 0.00027   31.6  -0.1   13  463-475    42-54  (118)
 63 PF12191 stn_TNFRSF12A:  Tumour  32.9      12 0.00025   32.4  -0.6   10  468-477   101-110 (129)
 64 PF13360 PQQ_2:  PQQ-like domai  32.1      97  0.0021   28.9   5.6   50  102-154     2-62  (238)
 65 TIGR03300 assembly_YfgL outer   31.1 3.9E+02  0.0084   27.2  10.3   75   76-154    83-171 (377)
 66 PF15330 SIT:  SHP2-interacting  30.9      13 0.00028   31.4  -0.6    9  464-472    17-25  (107)
 67 PTZ00382 Variant-specific surf  28.6      18  0.0004   29.9  -0.1   32  444-475    63-95  (96)
 68 PF05393 Hum_adeno_E3A:  Human   28.5      14 0.00031   29.7  -0.7   26  450-476    37-62  (94)
 69 PF14670 FXa_inhibition:  Coagu  28.0      19 0.00042   24.0  -0.0   14  315-328    17-30  (36)
 70 PF02009 Rifin_STEVOR:  Rifin/s  27.9     6.5 0.00014   39.5  -3.4   28  449-476   262-289 (299)
 71 PHA02887 EGF-like protein; Pro  27.8      62  0.0013   27.7   2.9   30  296-326    83-117 (126)
 72 PF12946 EGF_MSP1_1:  MSP1 EGF   27.8      18  0.0004   24.4  -0.2   25  303-327     5-31  (37)
 73 PF15102 TMEM154:  TMEM154 prot  27.7      48   0.001   29.6   2.4   15  463-477    73-87  (146)
 74 KOG1219 Uncharacterized conser  27.2      51  0.0011   41.8   3.1   24  303-326  3870-3895(4289)
 75 PF14870 PSII_BNR:  Photosynthe  27.0 6.6E+02   0.014   25.3  10.7   98   94-218   114-212 (302)
 76 KOG1214 Nidogen and related ba  25.6      51  0.0011   37.4   2.6   30  297-327   828-858 (1289)
 77 PF14575 EphA2_TM:  Ephrin type  24.1      22 0.00049   27.9  -0.3    9  465-473    20-28  (75)
 78 PF06247 Plasmod_Pvs28:  Plasmo  21.2      33 0.00072   31.9   0.1   27  302-328    49-81  (197)
 79 smart00564 PQQ beta-propeller   20.5 1.9E+02  0.0041   17.8   3.6   16  102-117    15-31  (33)

No 1  
>PF01453 B_lectin:  D-mannose binding lectin;  InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]:  Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein   This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity.  Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=99.96  E-value=6.7e-30  Score=219.61  Aligned_cols=111  Identities=41%  Similarity=0.709  Sum_probs=78.5

Q ss_pred             CCcEEEEcCCCCCCCCC-CceEEEEe-CCeEEEEcCCCceEEEe-ccCCCCCCceEEEEecCCCEEEEeccCCCCcceee
Q 042843           77 ERTIVWVANREQPVSDR-FSSVLRIS-DGNLVLFNESQLPIWST-NLTATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQ  153 (484)
Q Consensus        77 ~~tvVW~ANr~~Pv~~~-~~~~l~l~-~G~LvL~~~~~~~vWst-~~~~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQ  153 (484)
                      ++|+||+|||+.|+... ...+|.|+ ||+|+|+|..++.+|++ ++.+....+..|+|+|+|||||++..   +.+|||
T Consensus         1 ~~tvvW~an~~~p~~~~s~~~~L~l~~dGnLvl~~~~~~~iWss~~t~~~~~~~~~~~L~~~GNlvl~d~~---~~~lW~   77 (114)
T PF01453_consen    1 PRTVVWVANRNSPLTSSSGNYTLILQSDGNLVLYDSNGSVIWSSNNTSGRGNSGCYLVLQDDGNLVLYDSS---GNVLWQ   77 (114)
T ss_dssp             ---------TTEEEEECETTEEEEEETTSEEEEEETTTEEEEE--S-TTSS-SSEEEEEETTSEEEEEETT---SEEEEE
T ss_pred             CcccccccccccccccccccccceECCCCeEEEEcCCCCEEEEecccCCccccCeEEEEeCCCCEEEEeec---ceEEEe
Confidence            36899999999999531 25899999 99999999998899999 55443114789999999999999964   479999


Q ss_pred             ecccCceeccCCceeeeecCCCCceEEEecCCCCCCC
Q 042843          154 SFDHPAHTWIPGMKLTFNKRNNVSQLITSWKNKENPA  190 (484)
Q Consensus       154 SFd~PTDTlLpgq~l~~n~~~g~~~~L~Sw~s~~dps  190 (484)
                      ||||||||+||||+|+.+..+|....|+||++.+|||
T Consensus        78 Sf~~ptdt~L~~q~l~~~~~~~~~~~~~sw~s~~dps  114 (114)
T PF01453_consen   78 SFDYPTDTLLPGQKLGDGNVTGKNDSLTSWSSNTDPS  114 (114)
T ss_dssp             STTSSS-EEEEEET--TSEEEEESTSSEEEESS----
T ss_pred             ecCCCccEEEeccCcccCCCccccceEEeECCCCCCC
Confidence            9999999999999999866666556799999999996


No 2  
>PF00954 S_locus_glycop:  S-locus glycoprotein family;  InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=99.95  E-value=1.2e-27  Score=204.68  Aligned_cols=110  Identities=45%  Similarity=0.994  Sum_probs=104.2

Q ss_pred             EecCCcCCCCceeeeeeccccceeeEEEEEecCCeeEEEEeecCCceeEEEEEccCCcEEEEeeCCCCCCCeEEEeecCC
Q 042843          217 WSSGPWDENAKIFSMVPEMNQNYIYNFSYVSNENESYFTYNVKDSTYTSRAFMDVSGQDKQMNWLPLPTNSWFLFWSQPR  296 (484)
Q Consensus       217 w~sg~w~~~g~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~rl~Ld~dG~l~~y~w~~~~~~~W~~~w~~p~  296 (484)
                      ||+|+|  ||..|+++|++.....+.+.|+.++++.+++|.+.+.+.++|++||++|++++|.|++.. +.|.+.|++|.
T Consensus         1 wrsG~W--nG~~f~g~p~~~~~~~~~~~fv~~~~e~~~t~~~~~~s~~~r~~ld~~G~l~~~~w~~~~-~~W~~~~~~p~   77 (110)
T PF00954_consen    1 WRSGPW--NGQRFSGIPEMSSNSLYNYSFVSNNEEVYYTYSLSNSSVLSRLVLDSDGQLQRYIWNEST-QSWSVFWSAPK   77 (110)
T ss_pred             CCcccc--CCeEECCcccccccceeEEEEEECCCeEEEEEecCCCceEEEEEEeeeeEEEEEEEecCC-CcEEEEEEecc
Confidence            899999  999999999998777889999999999999999988889999999999999999999887 99999999999


Q ss_pred             CCCcccccCCCCccccCCCCccccccCCCccCC
Q 042843          297 QQCEVYALCGQFSTCNQQTERFCSCLKGFQQKS  329 (484)
Q Consensus       297 d~C~~~~~CG~~giC~~~~~~~C~C~~GF~p~~  329 (484)
                      |+||+|+.||+||+|+.+..+.|+||+||+|++
T Consensus        78 d~Cd~y~~CG~~g~C~~~~~~~C~Cl~GF~P~n  110 (110)
T PF00954_consen   78 DQCDVYGFCGPNGICNSNNSPKCSCLPGFEPKN  110 (110)
T ss_pred             cCCCCccccCCccEeCCCCCCceECCCCcCCCc
Confidence            999999999999999988788999999999963


No 3  
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=99.94  E-value=5.7e-26  Score=196.20  Aligned_cols=113  Identities=46%  Similarity=0.747  Sum_probs=99.5

Q ss_pred             CcccCCCeEEecCCeEEEEeecCCCCCCCc-EEEEEEEEecCCCcEEEEcCCCCCCCCCCceEEEEe-CCeEEEEcCCCc
Q 042843           36 QSLSGDQTIVSKGGVFAFGFFNPAPGKSSN-YYIGMWYNKVSERTIVWVANREQPVSDRFSSVLRIS-DGNLVLFNESQL  113 (484)
Q Consensus        36 ~~l~~~~~l~S~~g~F~lGFf~~~~~~~~~-~~lgIw~~~~~~~tvVW~ANr~~Pv~~~~~~~l~l~-~G~LvL~~~~~~  113 (484)
                      +.|..|++|+|+++.|++|||.+...   . .+.+|||.+.+ .++||.|||+.|..  ..++|.|+ ||+|+|+|.++.
T Consensus         2 ~~l~~~~~l~s~~~~f~~G~~~~~~q---~~dgnlv~~~~~~-~~~vW~snt~~~~~--~~~~l~l~~dGnLvl~~~~g~   75 (116)
T cd00028           2 NPLSSGQTLVSSGSLFELGFFKLIMQ---SRDYNLILYKGSS-RTVVWVANRDNPSG--SSCTLTLQSDGNLVIYDGSGT   75 (116)
T ss_pred             cCcCCCCEEEeCCCcEEEecccCCCC---CCeEEEEEEeCCC-CeEEEECCCCCCCC--CCEEEEEecCCCeEEEcCCCc
Confidence            56889999999999999999998754   4 89999998766 78999999999854  47899999 999999999999


Q ss_pred             eEEEeccCCCCCCceEEEEecCCCEEEEeccCCCCcceeeecccC
Q 042843          114 PIWSTNLTATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQSFDHP  158 (484)
Q Consensus       114 ~vWst~~~~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQSFd~P  158 (484)
                      ++|++++.+. .....|+|+|+|||||++.++   ++||||||||
T Consensus        76 ~vW~S~~~~~-~~~~~~~L~ddGnlvl~~~~~---~~~W~Sf~~P  116 (116)
T cd00028          76 VVWSSNTTRV-NGNYVLVLLDDGNLVLYDSDG---NFLWQSFDYP  116 (116)
T ss_pred             EEEEecccCC-CCceEEEEeCCCCEEEECCCC---CEEEcCCCCC
Confidence            9999998752 267899999999999999753   7899999999


No 4  
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=99.92  E-value=7.5e-25  Score=188.61  Aligned_cols=112  Identities=47%  Similarity=0.761  Sum_probs=98.0

Q ss_pred             CcccCCCeEEecCCeEEEEeecCCCCCCCcEEEEEEEEecCCCcEEEEcCCCCCCCCCCceEEEEe-CCeEEEEcCCCce
Q 042843           36 QSLSGDQTIVSKGGVFAFGFFNPAPGKSSNYYIGMWYNKVSERTIVWVANREQPVSDRFSSVLRIS-DGNLVLFNESQLP  114 (484)
Q Consensus        36 ~~l~~~~~l~S~~g~F~lGFf~~~~~~~~~~~lgIw~~~~~~~tvVW~ANr~~Pv~~~~~~~l~l~-~G~LvL~~~~~~~  114 (484)
                      +.|..|+.|+|+++.|++|||.+...   .++.+|||...+ .++||+|||+.|+..  ++.|.|+ ||+|+|+|.++.+
T Consensus         2 ~~l~~~~~l~s~~~~f~~G~~~~~~q---~dgnlV~~~~~~-~~~vW~snt~~~~~~--~~~l~l~~dGnLvl~~~~g~~   75 (114)
T smart00108        2 NTLSSGQTLVSGNSLFELGFFTLIMQ---NDYNLILYKSSS-RTVVWVANRDNPVSD--SCTLTLQSDGNLVLYDGDGRV   75 (114)
T ss_pred             cccCCCCEEecCCCcEeeeccccCCC---CCEEEEEEECCC-CcEEEECCCCCCCCC--CEEEEEeCCCCEEEEeCCCCE
Confidence            56888999999999999999998653   688999998876 789999999999873  5889999 9999999999999


Q ss_pred             EEEeccCCCCCCceEEEEecCCCEEEEeccCCCCcceeeeccc
Q 042843          115 IWSTNLTATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQSFDH  157 (484)
Q Consensus       115 vWst~~~~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQSFd~  157 (484)
                      +|++++... .+...|+|+|+|||||++..+   +++||||||
T Consensus        76 vW~S~t~~~-~~~~~~~L~ddGnlvl~~~~~---~~~W~Sf~~  114 (114)
T smart00108       76 VWSSNTTGA-NGNYVLVLLDDGNLVIYDSDG---NFLWQSFDY  114 (114)
T ss_pred             EEEecccCC-CCceEEEEeCCCCEEEECCCC---CEEeCCCCC
Confidence            999988622 257889999999999998753   799999997


No 5  
>PF08276 PAN_2:  PAN-like domain;  InterPro: IPR013227 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs
Probab=99.54  E-value=6.3e-15  Score=113.99  Aligned_cols=58  Identities=43%  Similarity=0.995  Sum_probs=50.4

Q ss_pred             CCccEEEEecccCCCCCccc--ccCCHHHHHHHHhhcCCeEEEEeCC----CceEEecccccce
Q 042843          359 KSDQFFQYSNMKLPKHPQSV--AVGGIRECETHCMNNCSCTAYAYKD----NACSIWVGSFVGL  416 (484)
Q Consensus       359 ~~~~F~~l~~v~~p~~~~~~--~~~~~~~C~~~CL~nCSC~Ay~y~~----~~C~~w~~~l~~~  416 (484)
                      .+|+|++|++|++|++...+  ...++++|+++||+||||+||+|.+    ++|++|+++|+|+
T Consensus         3 ~~d~F~~l~~~~~p~~~~~~~~~~~s~~~C~~~Cl~nCsC~Ayay~~~~~~~~C~lW~~~L~d~   66 (66)
T PF08276_consen    3 SGDGFLKLPNMKLPDFDNAIVDSSVSLEECEKACLSNCSCTAYAYSNLSGGGGCLLWYGDLVDL   66 (66)
T ss_pred             CCCEEEEECCeeCCCCcceeeecCCCHHHHHhhcCCCCCEeeEEeeccCCCCEEEEEcCEeecC
Confidence            46899999999999985544  3489999999999999999999973    5799999999874


No 6  
>cd01098 PAN_AP_plant Plant PAN/APPLE-like domain; present in plant S-receptor protein kinases and secreted glycoproteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions. S-receptor protein kinases and S-locus glycoproteins are involved in sporophytic self-incompatibility response in Brassica, one of probably many molecular mechanisms, by which hermaphrodite flowering plants avoid self-fertilization.
Probab=99.48  E-value=1.1e-13  Score=112.00  Aligned_cols=81  Identities=37%  Similarity=0.909  Sum_probs=63.5

Q ss_pred             CCCCCCCCCCCcCCccEEEEecccCCCCCcccccCCHHHHHHHHhhcCCeEEEEeCC--CceEEecccccceEEecCCCe
Q 042843          347 PLQCENISPANRKSDQFFQYSNMKLPKHPQSVAVGGIRECETHCMNNCSCTAYAYKD--NACSIWVGSFVGLQQLQGGGD  424 (484)
Q Consensus       347 ~l~C~~~~~~~~~~~~F~~l~~v~~p~~~~~~~~~~~~~C~~~CL~nCSC~Ay~y~~--~~C~~w~~~l~~~~~~~~~~~  424 (484)
                      +++|..+.    ..+.|++++++++|+........++++|++.||+||+|+||+|.+  ++|++|..++.+.+.....+.
T Consensus         2 ~~~C~~~~----~~~~f~~~~~~~~~~~~~~~~~~s~~~C~~~Cl~nCsC~a~~~~~~~~~C~~~~~~~~~~~~~~~~~~   77 (84)
T cd01098           2 PLNCGGDG----STDGFLKLPDVKLPDNASAITAISLEECREACLSNCSCTAYAYNNGSGGCLLWNGLLNNLRSLSSGGG   77 (84)
T ss_pred             CcccCCCC----CCCEEEEeCCeeCCCchhhhccCCHHHHHHHHhcCCCcceeeecCCCCeEEEEeceecceEeecCCCc
Confidence            45675321    136899999999998754434589999999999999999999974  679999999998876543335


Q ss_pred             EEEEEec
Q 042843          425 IIYIKLA  431 (484)
Q Consensus       425 ~~yikv~  431 (484)
                      ++||||+
T Consensus        78 ~~yiKv~   84 (84)
T cd01098          78 TLYLRLA   84 (84)
T ss_pred             EEEEEeC
Confidence            9999985


No 7  
>cd00129 PAN_APPLE PAN/APPLE-like domain; present in N-terminal (N) domains of plasminogen/ hepatocyte growth factor proteins,  plasma prekallikrein/coagulation factor XI and microneme antigen proteins, plant receptor-like protein kinases, and various nematode and leech anti-platelet proteins. Common structural features include two disulfide bonds that link the alpha-helix to the central region of the protein. PAN domains have significant functional versatility, fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=99.38  E-value=6.6e-13  Score=106.24  Aligned_cols=67  Identities=16%  Similarity=0.313  Sum_probs=57.4

Q ss_pred             CccEEEEecccCCCCCcccccCCHHHHHHHHhh---cCCeEEEEeC--CCceEEecccc-cceEEecCCCeEEEEEe
Q 042843          360 SDQFFQYSNMKLPKHPQSVAVGGIRECETHCMN---NCSCTAYAYK--DNACSIWVGSF-VGLQQLQGGGDIIYIKL  430 (484)
Q Consensus       360 ~~~F~~l~~v~~p~~~~~~~~~~~~~C~~~CL~---nCSC~Ay~y~--~~~C~~w~~~l-~~~~~~~~~~~~~yikv  430 (484)
                      +..|+++.+|++|++..    .++++|+++|++   ||||+||+|.  +.+|++|.++| .++++..+.+.++|||.
T Consensus         8 ~g~fl~~~~~klpd~~~----~s~~eC~~~Cl~~~~nCsC~Aya~~~~~~gC~~W~~~l~~d~~~~~~~g~~Ly~r~   80 (80)
T cd00129           8 AGTTLIKIALKIKTTKA----NTADECANRCEKNGLPFSCKAFVFAKARKQCLWFPFNSMSGVRKEFSHGFDLYENK   80 (80)
T ss_pred             CCeEEEeecccCCcccc----cCHHHHHHHHhcCCCCCCceeeeccCCCCCeEEecCcchhhHHhccCCCceeEeEC
Confidence            56799999999998754    678999999999   9999999995  35899999999 99887765555999983


No 8  
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=98.65  E-value=1.8e-07  Score=80.38  Aligned_cols=87  Identities=30%  Similarity=0.470  Sum_probs=62.2

Q ss_pred             EEEEe-CCeEEEEcCC-CceEEEeccCCCCCCceEEEEecCCCEEEEeccCCCCcceeeecccCceeccCCceeeeecCC
Q 042843           97 VLRIS-DGNLVLFNES-QLPIWSTNLTATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQSFDHPAHTWIPGMKLTFNKRN  174 (484)
Q Consensus        97 ~l~l~-~G~LvL~~~~-~~~vWst~~~~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQSFd~PTDTlLpgq~l~~n~~~  174 (484)
                      ++.++ ||+||+++.. +.++|++++..+......+.|.++|||||++.++   .++|+|=.  +               
T Consensus        23 ~~~~q~dgnlV~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g---~~vW~S~t--~---------------   82 (114)
T smart00108       23 TLIMQNDYNLILYKSSSRTVVWVANRDNPVSDSCTLTLQSDGNLVLYDGDG---RVVWSSNT--T---------------   82 (114)
T ss_pred             ccCCCCCEEEEEEECCCCcEEEECCCCCCCCCCEEEEEeCCCCEEEEeCCC---CEEEEecc--c---------------
Confidence            35567 9999999865 4789999986542133788999999999998753   68999810  0               


Q ss_pred             CCceEEEecCCCCCCCCceEEEEEcCCCCcEEEEEeeCCeeEEec
Q 042843          175 NVSQLITSWKNKENPAPGLFSLERAPDGSNQYVMLWNRSEQYWSS  219 (484)
Q Consensus       175 g~~~~L~Sw~s~~dps~G~y~l~~~~~g~~~~~l~~~~~~~Yw~s  219 (484)
                                    ...+.|.+.|+++|+  |+++-...++.|.+
T Consensus        83 --------------~~~~~~~~~L~ddGn--lvl~~~~~~~~W~S  111 (114)
T smart00108       83 --------------GANGNYVLVLLDDGN--LVIYDSDGNFLWQS  111 (114)
T ss_pred             --------------CCCCceEEEEeCCCC--EEEECCCCCEEeCC
Confidence                          124568899999998  66532334578875


No 9  
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=98.58  E-value=4.2e-07  Score=78.30  Aligned_cols=83  Identities=30%  Similarity=0.443  Sum_probs=60.2

Q ss_pred             CCeEEEEcCC-CceEEEeccCCCCCCceEEEEecCCCEEEEeccCCCCcceeeecccCceeccCCceeeeecCCCCceEE
Q 042843          102 DGNLVLFNES-QLPIWSTNLTATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQSFDHPAHTWIPGMKLTFNKRNNVSQLI  180 (484)
Q Consensus       102 ~G~LvL~~~~-~~~vWst~~~~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQSFd~PTDTlLpgq~l~~n~~~g~~~~L  180 (484)
                      ||+||+++.. ++++|++++..+......+.|.++|||||++.++   .++|+|=-.                       
T Consensus        30 dgnlv~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g---~~vW~S~~~-----------------------   83 (116)
T cd00028          30 DYNLILYKGSSRTVVWVANRDNPSGSSCTLTLQSDGNLVIYDGSG---TVVWSSNTT-----------------------   83 (116)
T ss_pred             eEEEEEEeCCCCeEEEECCCCCCCCCCEEEEEecCCCeEEEcCCC---cEEEEeccc-----------------------
Confidence            7899998765 4789999986532246789999999999998754   689987310                       


Q ss_pred             EecCCCCCCCCceEEEEEcCCCCcEEEEEeeCCeeEEecC
Q 042843          181 TSWKNKENPAPGLFSLERAPDGSNQYVMLWNRSEQYWSSG  220 (484)
Q Consensus       181 ~Sw~s~~dps~G~y~l~~~~~g~~~~~l~~~~~~~Yw~sg  220 (484)
                              ...+.+.+.|+++|+  |+++-....+.|.+.
T Consensus        84 --------~~~~~~~~~L~ddGn--lvl~~~~~~~~W~Sf  113 (116)
T cd00028          84 --------RVNGNYVLVLLDDGN--LVLYDSDGNFLWQSF  113 (116)
T ss_pred             --------CCCCceEEEEeCCCC--EEEECCCCCEEEcCC
Confidence                    024568999999998  665322346788864


No 10 
>smart00473 PAN_AP divergent subfamily of APPLE domains. Apple-like domains present in Plasminogen, C. elegans hypothetical ORFs and the extracellular portion of plant receptor-like protein kinases. Predicted to possess protein- and/or carbohydrate-binding functions.
Probab=98.57  E-value=2.1e-07  Score=73.40  Aligned_cols=70  Identities=33%  Similarity=0.793  Sum_probs=53.3

Q ss_pred             CccEEEEecccCCCCCcc-cccCCHHHHHHHHhh-cCCeEEEEeC--CCceEEec-ccccceEEecCCCeEEEEE
Q 042843          360 SDQFFQYSNMKLPKHPQS-VAVGGIRECETHCMN-NCSCTAYAYK--DNACSIWV-GSFVGLQQLQGGGDIIYIK  429 (484)
Q Consensus       360 ~~~F~~l~~v~~p~~~~~-~~~~~~~~C~~~CL~-nCSC~Ay~y~--~~~C~~w~-~~l~~~~~~~~~~~~~yik  429 (484)
                      .+.|..++++.+++.... ....++++|++.|++ +|+|.||.|.  +++|.+|. +++.+.+.....+.++|.|
T Consensus         3 ~~~f~~~~~~~l~~~~~~~~~~~s~~~C~~~C~~~~~~C~s~~y~~~~~~C~l~~~~~~~~~~~~~~~~~~~y~~   77 (78)
T smart00473        3 DDCFVRLPNTKLPGFSRIVISVASLEECASKCLNSNCSCRSFTYNNGTKGCLLWSESSLGDARLFPSGGVDLYEK   77 (78)
T ss_pred             CceeEEecCccCCCCcceeEcCCCHHHHHHHhCCCCCceEEEEEcCCCCEEEEeeCCccccceecccCCceeEEe
Confidence            357999999999854332 234799999999999 9999999997  46799998 7777776433333377766


No 11 
>PF01453 B_lectin:  D-mannose binding lectin;  InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]:  Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein   This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity.  Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=98.06  E-value=5.7e-05  Score=64.85  Aligned_cols=100  Identities=21%  Similarity=0.368  Sum_probs=65.6

Q ss_pred             CCeEEecCCeEEEEeecCCCCCCCcEEEEEEEEecCCCcEEEEc-CCCCCCCCCCceEEEEe-CCeEEEEcCCCceEEEe
Q 042843           41 DQTIVSKGGVFAFGFFNPAPGKSSNYYIGMWYNKVSERTIVWVA-NREQPVSDRFSSVLRIS-DGNLVLFNESQLPIWST  118 (484)
Q Consensus        41 ~~~l~S~~g~F~lGFf~~~~~~~~~~~lgIw~~~~~~~tvVW~A-Nr~~Pv~~~~~~~l~l~-~G~LvL~~~~~~~vWst  118 (484)
                      .+.+.+.+|.+.|-|...++     ..   .|.  ...++||.. +......  ..+.+.|. ||||||+|..+.++|++
T Consensus        11 ~~p~~~~s~~~~L~l~~dGn-----Lv---l~~--~~~~~iWss~~t~~~~~--~~~~~~L~~~GNlvl~d~~~~~lW~S   78 (114)
T PF01453_consen   11 NSPLTSSSGNYTLILQSDGN-----LV---LYD--SNGSVIWSSNNTSGRGN--SGCYLVLQDDGNLVLYDSSGNVLWQS   78 (114)
T ss_dssp             TEEEEECETTEEEEEETTSE-----EE---EEE--TTTEEEEE--S-TTSS---SSEEEEEETTSEEEEEETTSEEEEES
T ss_pred             ccccccccccccceECCCCe-----EE---EEc--CCCCEEEEecccCCccc--cCeEEEEeCCCCEEEEeecceEEEee
Confidence            45666655888998887542     22   243  245779999 4444432  26789999 99999999999999999


Q ss_pred             ccCCCCCCceEEEEec--CCCEEEEeccCCCCcceeeecccCc
Q 042843          119 NLTATSRRSVEAVLLD--EGNLVLRDLSNNLSKPLWQSFDHPA  159 (484)
Q Consensus       119 ~~~~~~~~~~~a~LlD--sGNLVL~~~~~~~~~~lWQSFd~PT  159 (484)
                      ... +  ....+.+++  .||++ +...   ..+.|.|=+.|+
T Consensus        79 f~~-p--tdt~L~~q~l~~~~~~-~~~~---~~~sw~s~~dps  114 (114)
T PF01453_consen   79 FDY-P--TDTLLPGQKLGDGNVT-GKND---SLTSWSSNTDPS  114 (114)
T ss_dssp             TTS-S--S-EEEEEET--TSEEE-EEST---SSEEEESS----
T ss_pred             cCC-C--ccEEEeccCcccCCCc-cccc---eEEeECCCCCCC
Confidence            432 2  467777777  89998 5432   358999877764


No 12 
>cd01100 APPLE_Factor_XI_like Subfamily of PAN/APPLE-like domains; present in plasma prekallikrein/coagulation factor XI, microneme antigen proteins, and a few prokaryotic proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=97.35  E-value=0.0002  Score=56.26  Aligned_cols=47  Identities=19%  Similarity=0.430  Sum_probs=34.3

Q ss_pred             EecccCCCCCccc-ccCCHHHHHHHHhhcCCeEEEEeCC--CceEEeccc
Q 042843          366 YSNMKLPKHPQSV-AVGGIRECETHCMNNCSCTAYAYKD--NACSIWVGS  412 (484)
Q Consensus       366 l~~v~~p~~~~~~-~~~~~~~C~~~CL~nCSC~Ay~y~~--~~C~~w~~~  412 (484)
                      +++++++...... ...+.++|++.|+.+|+|.||.|..  +.|+++...
T Consensus         9 ~~~~~~~g~d~~~~~~~s~~~Cq~~C~~~~~C~afT~~~~~~~C~lk~~~   58 (73)
T cd01100           9 GSNVDFRGGDLSTVFASSAEQCQAACTADPGCLAFTYNTKSKKCFLKSSE   58 (73)
T ss_pred             cCCCccccCCcceeecCCHHHHHHHcCCCCCceEEEEECCCCeEEcccCC
Confidence            3566666543222 2368999999999999999999963  569997643


No 13 
>PF04478 Mid2:  Mid2 like cell wall stress sensor;  InterPro: IPR007567 This family represents a region near the C terminus of Mid2, which contains a transmembrane region. The remainder of the protein sequence is serine-rich and of low complexity, and is therefore impossible to align accurately. Mid2 is thought to act as a mechanosensor of cell wall stress. The C-terminal cytoplasmic region of Mid2 is known to interact with Rom2, a guanine nucleotide exchange factor (GEF) for Rho1, which is part of the cell wall integrity signalling pathway [].
Probab=93.67  E-value=0.12  Score=46.07  Aligned_cols=33  Identities=30%  Similarity=0.405  Sum_probs=16.9

Q ss_pred             ccceEEEEEehHHHHHHHHHHHHheeeeeeeccC
Q 042843          441 KKGVVIGGVVGSVAVVALIGLIMLVYLGRRKTAT  474 (484)
Q Consensus       441 ~~~~~i~~~v~~~~~~~~~~~~~~~~~~~r~~~~  474 (484)
                      .++++|+++||+-++++|+ +++++|++++|+++
T Consensus        47 nknIVIGvVVGVGg~ill~-il~lvf~~c~r~kk   79 (154)
T PF04478_consen   47 NKNIVIGVVVGVGGPILLG-ILALVFIFCIRRKK   79 (154)
T ss_pred             CccEEEEEEecccHHHHHH-HHHhheeEEEeccc
Confidence            4468899998754333322 22333444444433


No 14 
>PF08693 SKG6:  Transmembrane alpha-helix domain;  InterPro: IPR014805 SKG6 and AXL2 are membrane proteins that show polarised intracellular localisation [, ]. This entry represents the highly conserved transmembrane alpha-helical domain found in these proteins [, ]. The full-length AXL2 protein has a negative regulatory function in cytokinesis [].
Probab=91.86  E-value=0.29  Score=33.50  Aligned_cols=10  Identities=0%  Similarity=-0.184  Sum_probs=4.5

Q ss_pred             heeeeeeecc
Q 042843          464 LVYLGRRKTA  473 (484)
Q Consensus       464 ~~~~~~r~~~  473 (484)
                      +++++|+||+
T Consensus        30 ~~l~~~~rR~   39 (40)
T PF08693_consen   30 AFLFFWYRRK   39 (40)
T ss_pred             HHhheEEecc
Confidence            3444455544


No 15 
>PF15102 TMEM154:  TMEM154 protein family
Probab=91.65  E-value=0.26  Score=43.64  Aligned_cols=25  Identities=12%  Similarity=0.089  Sum_probs=9.5

Q ss_pred             EEEEEehHHHHHHHHHHHHheeeee
Q 042843          445 VIGGVVGSVAVVALIGLIMLVYLGR  469 (484)
Q Consensus       445 ~i~~~v~~~~~~~~~~~~~~~~~~~  469 (484)
                      ++.++|..++++++++++++++++.
T Consensus        58 iLmIlIP~VLLvlLLl~vV~lv~~~   82 (146)
T PF15102_consen   58 ILMILIPLVLLVLLLLSVVCLVIYY   82 (146)
T ss_pred             EEEEeHHHHHHHHHHHHHHHheeEE
Confidence            3333333233333333334444433


No 16 
>smart00223 APPLE APPLE domain. Four-fold repeat in plasma kallikrein and coagulation factor XI. Factor XI apple 3 mediates binding to platelets. Factor XI apple 1 binds high-molecular-mass kininogen. Apple 4 in factor XI mediates dimer formation and binds to factor XIIa. Mutations in apple 4 cause factor XI deficiency, an inherited bleeding disorder.
Probab=91.17  E-value=0.27  Score=39.23  Aligned_cols=45  Identities=16%  Similarity=0.398  Sum_probs=33.9

Q ss_pred             ecccCCCCCcc-cccCCHHHHHHHHhhcCCeEEEEeCC--C---ceEEecc
Q 042843          367 SNMKLPKHPQS-VAVGGIRECETHCMNNCSCTAYAYKD--N---ACSIWVG  411 (484)
Q Consensus       367 ~~v~~p~~~~~-~~~~~~~~C~~~CL~nCSC~Ay~y~~--~---~C~~w~~  411 (484)
                      +|++++..... +...+.++|++.|..+=.|.||.|..  .   .|+++..
T Consensus         7 ~~~df~G~Dl~~~~~~~~~~Cq~~Ct~~~~C~~FTf~~~~~~~~~C~LK~s   57 (79)
T smart00223        7 KNVDFRGSDINTVYVPSAQVCQKRCTSHPRCLFFTFSTNEPPEEKCLLKDS   57 (79)
T ss_pred             cCccccCceeeeeecCCHHHHHHhhcCCCCccEEEeeCCCCCCCEeEeCcC
Confidence            46777665332 33478999999999999999999953  3   6998743


No 17 
>PF00024 PAN_1:  PAN domain This Prosite entry concerns apple domains, a subset of PAN domains;  InterPro: IPR003014 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs It has been shown that, the N-terminal N domains of members of the plasminogen/hepatocyte growth factor family, the apple domains of the plasma prekallikrein/coagulation factor XI family, and domains of various nematode proteins belong to the same module superfamily, the PAN module []. PAN contains a conserved core of three disulphide bridges. In some members of the family there is an additional fourth disulphide bridge that links the N and C termini of the domain.; PDB: 1GP9_C 2QJ2_B 1GMO_H 1NK1_B 3MKP_B 1BHT_B 3HN4_A 1GMN_A 3HMS_A 3HMT_B ....
Probab=91.00  E-value=0.33  Score=37.75  Aligned_cols=54  Identities=19%  Similarity=0.449  Sum_probs=39.0

Q ss_pred             cEEEEecccCCCCCccccc-CCHHHHHHHHhhcCC-eEEEEeCC--CceEEecccccc
Q 042843          362 QFFQYSNMKLPKHPQSVAV-GGIRECETHCMNNCS-CTAYAYKD--NACSIWVGSFVG  415 (484)
Q Consensus       362 ~F~~l~~v~~p~~~~~~~~-~~~~~C~~~CL~nCS-C~Ay~y~~--~~C~~w~~~l~~  415 (484)
                      .|..+++..+......... .++++|.+.|+.+=. |.+|.|..  ..|.+....-..
T Consensus         3 ~f~~~~~~~l~~~~~~~~~v~s~~~C~~~C~~~~~~C~s~~y~~~~~~C~L~~~~~~~   60 (79)
T PF00024_consen    3 AFERIPGYRLSGHSIKEINVPSLEECAQLCLNEPRRCKSFNYDPSSKTCYLSSSDRSS   60 (79)
T ss_dssp             TEEEEEEEEEESCEEEEEEESSHHHHHHHHHHSTT-ESEEEEETTTTEEEEECSSSSS
T ss_pred             CeEEECCEEEeCCcceEEcCCCHHHHHhhcCcCcccCCeEEEECCCCEEEEcCCCCCc
Confidence            4777777776654222223 589999999999999 99999964  469997654433


No 18 
>PF08277 PAN_3:  PAN-like domain;  InterPro: IPR006583 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs The PAN-3 or CW is a domain associated with a number of Caenorhabditis elegans hypothetical proteins.
Probab=90.07  E-value=1.5  Score=33.62  Aligned_cols=39  Identities=21%  Similarity=0.540  Sum_probs=30.3

Q ss_pred             cCCHHHHHHHHhhcCCeEEEEeCCCceEEec-ccccceEE
Q 042843          380 VGGIRECETHCMNNCSCTAYAYKDNACSIWV-GSFVGLQQ  418 (484)
Q Consensus       380 ~~~~~~C~~~CL~nCSC~Ay~y~~~~C~~w~-~~l~~~~~  418 (484)
                      ..+.++|-+.|..+=.|.++.+..+.|.++. +++..+++
T Consensus        19 ~~sw~~Cv~~C~~~~~C~la~~~~~~C~~y~~~~i~~v~~   58 (71)
T PF08277_consen   19 NTSWDDCVQKCYNDENCVLAYFDSGKCYLYNYGSISTVQK   58 (71)
T ss_pred             CCCHHHHhHHhCCCCEEEEEEeCCCCEEEEEcCCEEEEEE
Confidence            4778999999999999999988866799864 33334443


No 19 
>smart00605 CW CW domain.
Probab=87.03  E-value=2.1  Score=35.05  Aligned_cols=53  Identities=15%  Similarity=0.429  Sum_probs=38.2

Q ss_pred             cCCHHHHHHHHhhcCCeEEEEeCC-CceEEec-ccccceEEecCCC-eEEEEEecc
Q 042843          380 VGGIRECETHCMNNCSCTAYAYKD-NACSIWV-GSFVGLQQLQGGG-DIIYIKLAA  432 (484)
Q Consensus       380 ~~~~~~C~~~CL~nCSC~Ay~y~~-~~C~~w~-~~l~~~~~~~~~~-~~~yikv~~  432 (484)
                      ..+.++|.+.|..+..|+.+.... ..|.+.. +.+..+++..... ..+=+|+..
T Consensus        21 ~~sw~~Ci~~C~~~~~Cvlay~~~~~~C~~f~~~~~~~v~~~~~~~~~~VAfK~~~   76 (94)
T smart00605       21 TLSWDECIQKCYEDSNCVLAYGNSSETCYLFSYGTVLTVKKLSSSSGKKVAFKVST   76 (94)
T ss_pred             CCCHHHHHHHHhCCCceEEEecCCCCceEEEEcCCeEEEEEccCCCCcEEEEEEeC
Confidence            467899999999999999876654 6798765 3466666654323 367777753


No 20 
>PF02009 Rifin_STEVOR:  Rifin/stevor family;  InterPro: IPR002858 Malaria is still a major cause of mortality in many areas of the world. Plasmodium falciparum causes the most severe human form of the disease and is responsible for most fatalities. Severe cases of malaria can occur when the parasite invades and then proliferates within red blood cell erythrocytes. The parasite produces many variant antigenic proteins, encoded by multigene families, which are present on the surface of the infected erythrocyte and play important roles in virulence. A crucial survival mechanism for the malaria parasite is its ability to evade the immune response by switching these variant surface antigens. The high virulence of P. falciparum relative to other malarial parasites is in large part due to the fact that in this organism many of these surface antigens mediate the binding of infected erythrocytes to the vascular endothelium (cytoadherence) and non-infected erythrocytes (rosetting). This can lead to the accumulation of infected cells in the vasculature of a variety of organs, blocking the blood flow and reducing the oxygen supply. Clinical symptoms of severe infection can include fever, progressive anaemia, multi-organ dysfunction and coma. For more information see []. Several multicopy gene families have been described in Plasmodium falciparum, including the stevor family of subtelomeric open reading frames and the rif interspersed repetitive elements. Both families contain three predicted transmembrane segments. It has been proposed that stevor and rif are members of a larger superfamily that code for variant surface antigens [].
Probab=85.41  E-value=0.14  Score=51.33  Aligned_cols=20  Identities=20%  Similarity=0.326  Sum_probs=13.0

Q ss_pred             HHHHHheeeeeeeccCcccc
Q 042843          459 IGLIMLVYLGRRKTATVTTK  478 (484)
Q Consensus       459 ~~~~~~~~~~~r~~~~~~~~  478 (484)
                      +|++++.|++||+|||+|.+
T Consensus       269 VLIMvIIYLILRYRRKKKmk  288 (299)
T PF02009_consen  269 VLIMVIIYLILRYRRKKKMK  288 (299)
T ss_pred             HHHHHHHHHHHHHHHHhhhh
Confidence            33445677778888776665


No 21 
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=84.94  E-value=1.4  Score=36.50  Aligned_cols=29  Identities=24%  Similarity=0.165  Sum_probs=12.9

Q ss_pred             eEEEEEehHHHHHHHHHHHHheeeeeeec
Q 042843          444 VVIGGVVGSVAVVALIGLIMLVYLGRRKT  472 (484)
Q Consensus       444 ~~i~~~v~~~~~~~~~~~~~~~~~~~r~~  472 (484)
                      .|.+++|++++++.+++.++++++++|+|
T Consensus        67 aiagi~vg~~~~v~~lv~~l~w~f~~r~k   95 (96)
T PTZ00382         67 AIAGISVAVVAVVGGLVGFLCWWFVCRGK   95 (96)
T ss_pred             cEEEEEeehhhHHHHHHHHHhheeEEeec
Confidence            45566665543322222333444445443


No 22 
>PF07645 EGF_CA:  Calcium-binding EGF domain;  InterPro: IPR001881 A sequence of about forty amino-acid residues found in epidermal growth factor (EGF) has been shown [, , , , , ] to be present in a large number of membrane-bound and extracellular, mostly animal, proteins. Many of these proteins require calcium for their biological function and a calcium-binding site has been found at the N terminus of some EGF-like domains []. Calcium-binding may be crucial for numerous protein-protein interactions. For human coagulation factor IX it has been shown [] that the calcium-ligands form a pentagonal bipyramid. The first, third and fourth conserved negatively charged or polar residues are side chain ligands. The latter is possibly hydroxylated (see aspartic acid and asparagine hydroxylation site) []. A conserved aromatic residue, as well as the second conserved negative residue, are thought to be involved in stabilising the calcium-binding site. As in non-calcium binding EGF-like domains, there are six conserved cysteines and the structure of both types is very similar as calcium-binding induces only strictly local structural changes [].  +------------------+ +---------+ | | | | nxnnC-x(3,14)-C-x(3,7)-CxxbxxxxaxC-x(1,6)-C-x(8,13)-Cx | | +------------------+ 'n': negatively charged or polar residue [DEQN] 'b': possibly beta-hydroxylated residue [DN] 'a': aromatic amino acid 'C': cysteine, involved in disulphide bond 'x': any amino acid. ; GO: 0005509 calcium ion binding; PDB: 2VJ3_A 1TOZ_A 1LMJ_A 1UZQ_A 1UZK_A 1UZJ_B 1UZP_A 1EMO_A 1EMN_A 2RR0_A ....
Probab=84.44  E-value=0.51  Score=32.64  Aligned_cols=31  Identities=26%  Similarity=0.604  Sum_probs=24.6

Q ss_pred             CCCccc-ccCCCCccccC-CCCccccccCCCcc
Q 042843          297 QQCEVY-ALCGQFSTCNQ-QTERFCSCLKGFQQ  327 (484)
Q Consensus       297 d~C~~~-~~CG~~giC~~-~~~~~C~C~~GF~p  327 (484)
                      |+|... ..|..++.|.. ..+-.|.|++||+.
T Consensus         3 dEC~~~~~~C~~~~~C~N~~Gsy~C~C~~Gy~~   35 (42)
T PF07645_consen    3 DECAEGPHNCPENGTCVNTEGSYSCSCPPGYEL   35 (42)
T ss_dssp             STTTTTSSSSSTTSEEEEETTEEEEEESTTEEE
T ss_pred             cccCCCCCcCCCCCEEEcCCCCEEeeCCCCcEE
Confidence            678775 47999999974 34568999999984


No 23 
>PF01034 Syndecan:  Syndecan domain;  InterPro: IPR001050 The syndecans are transmembrane proteoglycans which are involved in the organisation of cytoskeleton and/or actin microfilaments, and have important roles as cell surface receptors during cell-cell and/or cell-matrix interactions [, ]. Structurally, these proteins consist of four separate domains:   A signal sequence; An extracellular domain (ectodomain) of variable length whose sequence is not evolutionary conserved in the various forms of syndecans. The ectodomain contains the sites of attachment of the heparan sulphate glycosaminoglycan side chains;  A transmembrane region;  A highly conserved cytoplasmic domain of about 30 to 35 residues, which could interact with cytoskeletal proteins.    The proteins known to belong to this family are:    Syndecan 1.  Syndecan 2 or fibroglycan.  Syndecan 3 or neuroglycan or N-syndecan.  Syndecan 4 or amphiglycan or ryudocan.  Drosophila syndecan.   Caenorhabditis elegans probable syndecan (F57C7.3).    Syndecan-4, a transmembrane heparan sulphate proteoglycan, is a coreceptor with integrins in cell adhesion. It has been suggested to form a ternary signalling complex with protein kinase Calpha and phosphatidylinositol 4,5-bisphosphate (PIP2). Structural studies have demonstrated that the cytoplasmic domain undergoes a conformational transition and forms a symmetric dimer in the presence of phospholipid activator PIP2, and whose overall structure in solution exhibits a twisted clamp shape having a cavity in the centre of dimeric interface. In addition, it has been observed that the syndecan-4 variable domain interacts, strongly, not only with fatty acyl groups but also the anionic head group of PIP2. These findings indicate that PIP2 promotes oligomerisation of the syndecan-4 cytoplasmic domain for transmembrane signalling and cell-matrix adhesion [, ].; GO: 0008092 cytoskeletal protein binding, 0016020 membrane; PDB: 1EJQ_B 1EJP_B 1YBO_C 1OBY_Q.
Probab=84.20  E-value=0.32  Score=36.79  Aligned_cols=16  Identities=13%  Similarity=0.245  Sum_probs=0.6

Q ss_pred             HHheeeeeeeccCccc
Q 042843          462 IMLVYLGRRKTATVTT  477 (484)
Q Consensus       462 ~~~~~~~~r~~~~~~~  477 (484)
                      ++++++++|.|+|.++
T Consensus        27 lLIlf~iyR~rkkdEG   42 (64)
T PF01034_consen   27 LLILFLIYRMRKKDEG   42 (64)
T ss_dssp             ----------S-----
T ss_pred             HHHHHHHHHHHhcCCC
Confidence            3445556666655544


No 24 
>PF14295 PAN_4:  PAN domain; PDB: 2YIL_E 2YIP_C 2YIO_A.
Probab=83.66  E-value=0.9  Score=32.16  Aligned_cols=24  Identities=21%  Similarity=0.696  Sum_probs=17.5

Q ss_pred             cCCHHHHHHHHhhcCCeEEEEeCC
Q 042843          380 VGGIRECETHCMNNCSCTAYAYKD  403 (484)
Q Consensus       380 ~~~~~~C~~~CL~nCSC~Ay~y~~  403 (484)
                      ..+.++|.+.|..+=.|.+|.|..
T Consensus        15 ~~s~~~C~~~C~~~~~C~~~~~~~   38 (51)
T PF14295_consen   15 ASSPEECQAACAADPGCQAFTFNP   38 (51)
T ss_dssp             ---HHHHHHHHHTSTT--EEEEET
T ss_pred             CCCHHHHHHHccCCCCCCEEEEEC
Confidence            368899999999999999999854


No 25 
>cd00053 EGF Epidermal growth factor domain, found in epidermal growth factor (EGF) presents in a large number of proteins, mostly animal; the list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied; the functional significance of EGF-like domains in what appear to be unrelated proteins is not yet clear; a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase); the domain includes six cysteine residues which have been shown to be involved in disulfide bonds; the main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet; Subdomains between the conserved cysteines vary in length; the region between the 5th and 6th cysteine contains two conserved glycines of which at  least  one  is  present  in  most EGF-like domains; a subset of these bind calcium.
Probab=81.08  E-value=1.2  Score=28.43  Aligned_cols=29  Identities=24%  Similarity=0.590  Sum_probs=21.3

Q ss_pred             CcccccCCCCccccCC-CCccccccCCCcc
Q 042843          299 CEVYALCGQFSTCNQQ-TERFCSCLKGFQQ  327 (484)
Q Consensus       299 C~~~~~CG~~giC~~~-~~~~C~C~~GF~p  327 (484)
                      |.....|..++.|... ....|.|++||..
T Consensus         2 C~~~~~C~~~~~C~~~~~~~~C~C~~g~~g   31 (36)
T cd00053           2 CAASNPCSNGGTCVNTPGSYRCVCPPGYTG   31 (36)
T ss_pred             CCCCCCCCCCCEEecCCCCeEeECCCCCcc
Confidence            4434678888999753 4578999999964


No 26 
>PF01102 Glycophorin_A:  Glycophorin A;  InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=80.83  E-value=0.37  Score=41.66  Aligned_cols=32  Identities=19%  Similarity=0.325  Sum_probs=13.9

Q ss_pred             EEEEEehHHHHHHHHHHHHheeeeeeeccCccc
Q 042843          445 VIGGVVGSVAVVALIGLIMLVYLGRRKTATVTT  477 (484)
Q Consensus       445 ~i~~~v~~~~~~~~~~~~~~~~~~~r~~~~~~~  477 (484)
                      +++++++++ +++++++++++|++||+|||...
T Consensus        66 i~~Ii~gv~-aGvIg~Illi~y~irR~~Kk~~~   97 (122)
T PF01102_consen   66 IIGIIFGVM-AGVIGIILLISYCIRRLRKKSSS   97 (122)
T ss_dssp             HHHHHHHHH-HHHHHHHHHHHHHHHHHS-----
T ss_pred             eeehhHHHH-HHHHHHHHHHHHHHHHHhccCCC
Confidence            333444444 33334444566777777666543


No 27 
>smart00179 EGF_CA Calcium-binding EGF-like domain.
Probab=75.76  E-value=2.2  Score=28.08  Aligned_cols=30  Identities=23%  Similarity=0.573  Sum_probs=22.3

Q ss_pred             CCCcccccCCCCccccCC-CCccccccCCCc
Q 042843          297 QQCEVYALCGQFSTCNQQ-TERFCSCLKGFQ  326 (484)
Q Consensus       297 d~C~~~~~CG~~giC~~~-~~~~C~C~~GF~  326 (484)
                      ++|.....|...+.|... ....|.|++||.
T Consensus         3 ~~C~~~~~C~~~~~C~~~~g~~~C~C~~g~~   33 (39)
T smart00179        3 DECASGNPCQNGGTCVNTVGSYRCECPPGYT   33 (39)
T ss_pred             ccCcCCCCcCCCCEeECCCCCeEeECCCCCc
Confidence            567655678888899743 345799999986


No 28 
>KOG4649 consensus PQQ (pyrrolo-quinoline quinone) repeat protein [Secondary metabolites biosynthesis, transport and catabolism]
Probab=75.65  E-value=10  Score=37.12  Aligned_cols=46  Identities=30%  Similarity=0.506  Sum_probs=34.9

Q ss_pred             CCcEEEEcCCCCCCCCC---CceEEEEe--CCeEEEEcCCCceEEEeccCC
Q 042843           77 ERTIVWVANREQPVSDR---FSSVLRIS--DGNLVLFNESQLPIWSTNLTA  122 (484)
Q Consensus        77 ~~tvVW~ANr~~Pv~~~---~~~~l~l~--~G~LvL~~~~~~~vWst~~~~  122 (484)
                      +.+..|.|.|..|+-.+   -+..+.++  ||+|.-.|+.|+.||.-.+.+
T Consensus       168 ~~~~~w~~~~~~PiF~splcv~~sv~i~~VdG~l~~f~~sG~qvwr~~t~G  218 (354)
T KOG4649|consen  168 SSTEFWAATRFGPIFASPLCVGSSVIITTVDGVLTSFDESGRQVWRPATKG  218 (354)
T ss_pred             CcceehhhhcCCccccCceeccceEEEEEeccEEEEEcCCCcEEEeecCCC
Confidence            45889999999998743   12345565  999999999999999765543


No 29 
>PF01299 Lamp:  Lysosome-associated membrane glycoprotein (Lamp);  InterPro: IPR002000 Lysosome-associated membrane glycoproteins (lamp) [] are integral membrane proteins, specific to lysosomes, and whose exact biological function is not yet clear. Structurally, the lamp proteins consist of two internally homologous lysosome-luminal domains separated by a proline-rich hinge region; at the C-terminal extremity there is a transmembrane region (TM) followed by a very short cytoplasmic tail (C). In each of the duplicated domains, there are two conserved disulphide bonds. This structure is schematically represented in the figure below.   +-----+ +-----+ +-----+ +-----+ | | | | | | | | xCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxxxCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxx +--------------------------++Hinge++--------------------------++TM++C+  In mammals, there are two closely related types of lamp: lamp-1 and lamp-2, which form major components of the lysosome membrane. In chicken lamp-1 is known as LEP100.  Also included in this entry is the macrophage protein CD68 (or macrosialin) [] is a heavily glycosylated integral membrane protein whose structure consists of a mucin-like domain followed by a proline-rich hinge; a single lamp-like domain; a transmembrane region and a short cytoplasmic tail.   Similar to CD68, mammalian lamp-3, which is expressed in lymphoid organs, dendritic cells and in lung, contains all the C-terminal regions but lacks the N-terminal lamp-like region []. In a lamp-family protein from nematodes [] only the part C-terminal to the hinge is conserved. ; GO: 0016020 membrane
Probab=74.81  E-value=1.8  Score=43.71  Aligned_cols=31  Identities=16%  Similarity=0.285  Sum_probs=18.2

Q ss_pred             eEEEEEehHHHHHHHHHHHHheeeeeeeccCc
Q 042843          444 VVIGGVVGSVAVVALIGLIMLVYLGRRKTATV  475 (484)
Q Consensus       444 ~~i~~~v~~~~~~~~~~~~~~~~~~~r~~~~~  475 (484)
                      .+|-++||++++++ +++++++|++.|||++.
T Consensus       271 ~~vPIaVG~~La~l-vlivLiaYli~Rrr~~~  301 (306)
T PF01299_consen  271 DLVPIAVGAALAGL-VLIVLIAYLIGRRRSRA  301 (306)
T ss_pred             chHHHHHHHHHHHH-HHHHHHhheeEeccccc
Confidence            44445666665544 43455677777776554


No 30 
>cd01099 PAN_AP_HGF Subfamily of PAN/APPLE-like domains; present in N-terminal (N) domains of plasminogen/hepatocyte growth factor proteins, and various proteins found in Bilateria, such as leech anti-platelet proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=74.12  E-value=5.8  Score=31.43  Aligned_cols=33  Identities=21%  Similarity=0.708  Sum_probs=26.6

Q ss_pred             cCCHHHHHHHHhh--cCCeEEEEeC--CCceEEeccc
Q 042843          380 VGGIRECETHCMN--NCSCTAYAYK--DNACSIWVGS  412 (484)
Q Consensus       380 ~~~~~~C~~~CL~--nCSC~Ay~y~--~~~C~~w~~~  412 (484)
                      ..++++|.+.|++  +=.|.++.|.  ...|.+-..+
T Consensus        24 ~~s~~~C~~~C~~~~~f~CrSf~y~~~~~~C~L~~~~   60 (80)
T cd01099          24 VASLEECLRKCLEETEFTCRSFNYNYKSKECILSDED   60 (80)
T ss_pred             cCCHHHHHHHhCCCCCceEeEEEEEcCCCEEEEeCCC
Confidence            3789999999999  8899999984  3569985433


No 31 
>PF01683 EB:  EB module;  InterPro: IPR006149  The EB domain has no known function. It is found in several Caenorhabditis sp. and Drosophila sp. proteins. The domain contains 8 conserved cysteines that probably form four disulphide bridges and is found associated with kunitz domains IPR002223 from INTERPRO 
Probab=73.81  E-value=2.9  Score=30.14  Aligned_cols=33  Identities=30%  Similarity=0.598  Sum_probs=27.3

Q ss_pred             ecCCCCCcccccCCCCccccCCCCccccccCCCccC
Q 042843          293 SQPRQQCEVYALCGQFSTCNQQTERFCSCLKGFQQK  328 (484)
Q Consensus       293 ~~p~d~C~~~~~CG~~giC~~~~~~~C~C~~GF~p~  328 (484)
                      ..|.+.|....-|-.++.|..   ..|.|++||.+.
T Consensus        16 ~~~g~~C~~~~qC~~~s~C~~---g~C~C~~g~~~~   48 (52)
T PF01683_consen   16 VQPGESCESDEQCIGGSVCVN---GRCQCPPGYVEV   48 (52)
T ss_pred             CCCCCCCCCcCCCCCcCEEcC---CEeECCCCCEec
Confidence            456678999999999999953   589999999864


No 32 
>PTZ00046 rifin; Provisional
Probab=72.21  E-value=0.81  Score=46.65  Aligned_cols=18  Identities=22%  Similarity=0.386  Sum_probs=11.7

Q ss_pred             HHHheeeeeeeccCcccc
Q 042843          461 LIMLVYLGRRKTATVTTK  478 (484)
Q Consensus       461 ~~~~~~~~~r~~~~~~~~  478 (484)
                      ++++.|++.|+|||++.|
T Consensus       330 IMvIIYLILRYRRKKKMk  347 (358)
T PTZ00046        330 IMVIIYLILRYRRKKKMK  347 (358)
T ss_pred             HHHHHHHHHHhhhcchhH
Confidence            445666677777777665


No 33 
>TIGR01477 RIFIN variant surface antigen, rifin family. This model represents the rifin branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of rifin sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 20 bits.
Probab=72.19  E-value=0.81  Score=46.53  Aligned_cols=18  Identities=22%  Similarity=0.386  Sum_probs=11.6

Q ss_pred             HHHheeeeeeeccCcccc
Q 042843          461 LIMLVYLGRRKTATVTTK  478 (484)
Q Consensus       461 ~~~~~~~~~r~~~~~~~~  478 (484)
                      ++++.|++.|+|||++.|
T Consensus       325 IMvIIYLILRYRRKKKMk  342 (353)
T TIGR01477       325 IMVIIYLILRYRRKKKMK  342 (353)
T ss_pred             HHHHHHHHHHhhhcchhH
Confidence            445666677777776665


No 34 
>PF09064 Tme5_EGF_like:  Thrombomodulin like fifth domain, EGF-like;  InterPro: IPR015149 This domain adopts a fold similar to other EGF domains, with a flat major and a twisted minor beta sheet. Disulphide pairing, however, is not of the usual 1-3, 2-4, 5-6 type; rather 1-2, 3-4, 5-6 pairing is found. Its extended major sheet (strands beta-2 and beta-3 and the connecting loop) projects into thrombin's active site groove. This domain is required for interaction of thrombomodulin with thrombin, and subsequent activation of protein-C []. ; GO: 0004888 transmembrane signaling receptor activity, 0016021 integral to membrane
Probab=71.54  E-value=2.2  Score=28.00  Aligned_cols=18  Identities=22%  Similarity=0.720  Sum_probs=13.7

Q ss_pred             cccCCCCccccccCCCcc
Q 042843          310 TCNQQTERFCSCLKGFQQ  327 (484)
Q Consensus       310 iC~~~~~~~C~C~~GF~p  327 (484)
                      .|+.+...+|.||.||..
T Consensus        11 ~CDpn~~~~C~CPeGyIl   28 (34)
T PF09064_consen   11 DCDPNSPGQCFCPEGYIL   28 (34)
T ss_pred             ccCCCCCCceeCCCceEe
Confidence            455555668999999975


No 35 
>PF12661 hEGF:  Human growth factor-like EGF; PDB: 2YGQ_A 2E26_A 3A7Q_A 2YGP_A 2YGO_A 1HRE_A 1HAE_A 1HAF_A 1HRF_A.
Probab=70.30  E-value=1.2  Score=22.87  Aligned_cols=9  Identities=33%  Similarity=1.047  Sum_probs=6.7

Q ss_pred             cccccCCCc
Q 042843          318 FCSCLKGFQ  326 (484)
Q Consensus       318 ~C~C~~GF~  326 (484)
                      .|.|++||.
T Consensus         1 ~C~C~~G~~    9 (13)
T PF12661_consen    1 TCQCPPGWT    9 (13)
T ss_dssp             EEEE-TTEE
T ss_pred             CccCcCCCc
Confidence            489999986


No 36 
>cd00054 EGF_CA Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.
Probab=69.65  E-value=3.9  Score=26.42  Aligned_cols=30  Identities=27%  Similarity=0.595  Sum_probs=21.6

Q ss_pred             CCCcccccCCCCccccCC-CCccccccCCCc
Q 042843          297 QQCEVYALCGQFSTCNQQ-TERFCSCLKGFQ  326 (484)
Q Consensus       297 d~C~~~~~CG~~giC~~~-~~~~C~C~~GF~  326 (484)
                      ++|.....|...+.|... ....|.|++||.
T Consensus         3 ~~C~~~~~C~~~~~C~~~~~~~~C~C~~g~~   33 (38)
T cd00054           3 DECASGNPCQNGGTCVNTVGSYRCSCPPGYT   33 (38)
T ss_pred             ccCCCCCCcCCCCEeECCCCCeEeECCCCCc
Confidence            567654568878889743 345799999985


No 37 
>PF07974 EGF_2:  EGF-like domain;  InterPro: IPR013111 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length. This entry contains EGF domains found in a variety of extracellular and membrane proteins
Probab=69.10  E-value=3.8  Score=26.67  Aligned_cols=23  Identities=26%  Similarity=0.697  Sum_probs=18.5

Q ss_pred             ccCCCCccccCCCCccccccCCCc
Q 042843          303 ALCGQFSTCNQQTERFCSCLKGFQ  326 (484)
Q Consensus       303 ~~CG~~giC~~~~~~~C~C~~GF~  326 (484)
                      ..|...|.|... ...|.|.+||.
T Consensus         6 ~~C~~~G~C~~~-~g~C~C~~g~~   28 (32)
T PF07974_consen    6 NICSGHGTCVSP-CGRCVCDSGYT   28 (32)
T ss_pred             CccCCCCEEeCC-CCEEECCCCCc
Confidence            479999999743 46899999986


No 38 
>PHA03265 envelope glycoprotein D; Provisional
Probab=67.39  E-value=2.3  Score=42.86  Aligned_cols=39  Identities=15%  Similarity=0.256  Sum_probs=19.6

Q ss_pred             ceEEEEEehHHHHHHHHHHHHheeeeeeeccCcccccccC
Q 042843          443 GVVIGGVVGSVAVVALIGLIMLVYLGRRKTATVTTKTVEG  482 (484)
Q Consensus       443 ~~~i~~~v~~~~~~~~~~~~~~~~~~~r~~~~~~~~~~~~  482 (484)
                      ...++++|+..|+.++++.+ ++|.+||||+..++.+..|
T Consensus       347 ~~~~g~~ig~~i~glv~vg~-il~~~~rr~k~~~k~~~~~  385 (402)
T PHA03265        347 STFVGISVGLGIAGLVLVGV-ILYVCLRRKKELKKSAQNG  385 (402)
T ss_pred             CcccceEEccchhhhhhhhH-HHHHHhhhhhhhhhhhhcC
Confidence            35566777665554433232 3444566655444435544


No 39 
>PF00008 EGF:  EGF-like domain This is a sub-family of the Pfam entry This is a sub-family of the Pfam entry;  InterPro: IPR006209 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length.; GO: 0005515 protein binding; PDB: 1WHE_A 1CCF_A 1APO_A 1WHF_A 2VJ3_A 1TOZ_A 4D90_B 3CFW_A 1EDM_B 1IXA_A ....
Probab=64.05  E-value=2.5  Score=27.34  Aligned_cols=23  Identities=26%  Similarity=0.610  Sum_probs=17.9

Q ss_pred             cCCCCccccC-C-CCccccccCCCc
Q 042843          304 LCGQFSTCNQ-Q-TERFCSCLKGFQ  326 (484)
Q Consensus       304 ~CG~~giC~~-~-~~~~C~C~~GF~  326 (484)
                      .|...|.|.. . ....|.|++||.
T Consensus         5 ~C~n~g~C~~~~~~~y~C~C~~G~~   29 (32)
T PF00008_consen    5 PCQNGGTCIDLPGGGYTCECPPGYT   29 (32)
T ss_dssp             SSTTTEEEEEESTSEEEEEEBTTEE
T ss_pred             cCCCCeEEEeCCCCCEEeECCCCCc
Confidence            6888888873 2 457899999986


No 40 
>PF12662 cEGF:  Complement Clr-like EGF-like
Probab=60.75  E-value=3.7  Score=24.93  Aligned_cols=11  Identities=45%  Similarity=0.999  Sum_probs=9.4

Q ss_pred             cccccCCCccC
Q 042843          318 FCSCLKGFQQK  328 (484)
Q Consensus       318 ~C~C~~GF~p~  328 (484)
                      .|+|++||+..
T Consensus         3 ~C~C~~Gy~l~   13 (24)
T PF12662_consen    3 TCSCPPGYQLS   13 (24)
T ss_pred             EeeCCCCCcCC
Confidence            69999999864


No 41 
>PF02439 Adeno_E3_CR2:  Adenovirus E3 region protein CR2;  InterPro: IPR003470 Early region 3 (E3) of human adenoviruses (Ads) codes for proteins that appear to control viral interactions with the host []. This region called CR1 (conserved region 1) [] is found three times in Human adenovirus 19 (a subgroup D adenovirus) 49 kDa protein in the E3 region. CR1 is also found in the 20.1 Kd protein of subgroup B adenoviruses. The function of this 80 amino acid region is unknown. This region is probably a divergent immunoglobulin domain.
Probab=58.26  E-value=2.1  Score=28.93  Aligned_cols=8  Identities=38%  Similarity=0.393  Sum_probs=3.0

Q ss_pred             EEEEehHH
Q 042843          446 IGGVVGSV  453 (484)
Q Consensus       446 i~~~v~~~  453 (484)
                      |+++++++
T Consensus         6 IaIIv~V~   13 (38)
T PF02439_consen    6 IAIIVAVV   13 (38)
T ss_pred             hhHHHHHH
Confidence            33334333


No 42 
>PF12947 EGF_3:  EGF domain;  InterPro: IPR024731 This entry represents an EGF domain found in the the C terminus of malarial parasite merozoite surface protein 1 [], as well as other proteins.; PDB: 2NPR_A 1N1I_C 1B9W_A 1YO8_A 2RHP_A.
Probab=57.85  E-value=3.3  Score=27.70  Aligned_cols=24  Identities=25%  Similarity=0.688  Sum_probs=16.9

Q ss_pred             ccCCCCccccCC-CCccccccCCCc
Q 042843          303 ALCGQFSTCNQQ-TERFCSCLKGFQ  326 (484)
Q Consensus       303 ~~CG~~giC~~~-~~~~C~C~~GF~  326 (484)
                      +-|.++..|... .+..|.|.+||.
T Consensus         6 ~~C~~nA~C~~~~~~~~C~C~~Gy~   30 (36)
T PF12947_consen    6 GGCHPNATCTNTGGSYTCTCKPGYE   30 (36)
T ss_dssp             GGS-TTCEEEE-TTSEEEEE-CEEE
T ss_pred             CCCCCCcEeecCCCCEEeECCCCCc
Confidence            568889999743 457899999996


No 43 
>smart00181 EGF Epidermal growth factor-like domain.
Probab=52.15  E-value=12  Score=24.07  Aligned_cols=24  Identities=29%  Similarity=0.634  Sum_probs=17.2

Q ss_pred             ccCCCCccccCC-CCccccccCCCcc
Q 042843          303 ALCGQFSTCNQQ-TERFCSCLKGFQQ  327 (484)
Q Consensus       303 ~~CG~~giC~~~-~~~~C~C~~GF~p  327 (484)
                      ..|... .|... ....|.|++||.-
T Consensus         6 ~~C~~~-~C~~~~~~~~C~C~~g~~g   30 (35)
T smart00181        6 GPCSNG-TCINTPGSYTCSCPPGYTG   30 (35)
T ss_pred             CCCCCC-EEECCCCCeEeECCCCCcc
Confidence            456666 78643 4578999999964


No 44 
>PF06024 DUF912:  Nucleopolyhedrovirus protein of unknown function (DUF912);  InterPro: IPR009261 This entry is represented by Autographa californica nuclear polyhedrosis virus (AcMNPV), Orf78; it is a family of uncharacterised viral proteins.
Probab=52.12  E-value=9.3  Score=31.91  Aligned_cols=12  Identities=8%  Similarity=0.088  Sum_probs=5.3

Q ss_pred             eeeeeeeccCcc
Q 042843          465 VYLGRRKTATVT  476 (484)
Q Consensus       465 ~~~~~r~~~~~~  476 (484)
                      .+++.|.|+++.
T Consensus        83 YFVILRer~~~~   94 (101)
T PF06024_consen   83 YFVILRERQKSI   94 (101)
T ss_pred             EEEEEecccccc
Confidence            333455544443


No 45 
>PF12877 DUF3827:  Domain of unknown function (DUF3827);  InterPro: IPR024606 The function of the proteins in this entry is not currently known, but one of the human proteins (Q9HCM3 from SWISSPROT) has been implicated in pilocytic astrocytomas [, , ]. In the majority of cases of pilocytic astrocytomas a tandem duplication produces an in-frame fusion of the gene encoding this protein and the BRAF oncogene. The resulting fusion protein has constitutive BRAF kinase activity and is capable of transforming cells. 
Probab=51.10  E-value=10  Score=41.44  Aligned_cols=15  Identities=40%  Similarity=0.386  Sum_probs=6.7

Q ss_pred             ceEEEEEehHHHHHH
Q 042843          443 GVVIGGVVGSVAVVA  457 (484)
Q Consensus       443 ~~~i~~~v~~~~~~~  457 (484)
                      -|||+.|++++++++
T Consensus       269 lWII~gVlvPv~vV~  283 (684)
T PF12877_consen  269 LWIIAGVLVPVLVVL  283 (684)
T ss_pred             eEEEehHhHHHHHHH
Confidence            344444444544433


No 46 
>PF03302 VSP:  Giardia variant-specific surface protein;  InterPro: IPR005127 During infection, the intestinal protozoan parasite Giardia lamblia virus undergoes continuous antigenic variation which is determined by diversification of the parasite's major surface antigen, named VSP (variant surface protein).
Probab=49.90  E-value=18  Score=37.94  Aligned_cols=30  Identities=23%  Similarity=0.194  Sum_probs=16.7

Q ss_pred             ceEEEEEehHHHHHHHHHHHHheeeeeeec
Q 042843          443 GVVIGGVVGSVAVVALIGLIMLVYLGRRKT  472 (484)
Q Consensus       443 ~~~i~~~v~~~~~~~~~~~~~~~~~~~r~~  472 (484)
                      ..|.+|.|+++|++-.|+.++++|++.|.|
T Consensus       367 gaIaGIsvavvvvVgglvGfLcWwf~crgk  396 (397)
T PF03302_consen  367 GAIAGISVAVVVVVGGLVGFLCWWFICRGK  396 (397)
T ss_pred             cceeeeeehhHHHHHHHHHHHhhheeeccc
Confidence            466667776554433233455666666654


No 47 
>PF14610 DUF4448:  Protein of unknown function (DUF4448)
Probab=49.50  E-value=17  Score=33.93  Aligned_cols=22  Identities=32%  Similarity=0.348  Sum_probs=9.1

Q ss_pred             EEehHHHHHHHHHHHHheeeeee
Q 042843          448 GVVGSVAVVALIGLIMLVYLGRR  470 (484)
Q Consensus       448 ~~v~~~~~~~~~~~~~~~~~~~r  470 (484)
                      +|++++++++++ ++++++++|+
T Consensus       161 aI~lPvvv~~~~-~~~~~~~~~~  182 (189)
T PF14610_consen  161 AIALPVVVVVLA-LIMYGFFFWN  182 (189)
T ss_pred             EEEccHHHHHHH-HHHHhhheee
Confidence            444455444433 2333444443


No 48 
>PF13360 PQQ_2:  PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=48.39  E-value=1.6e+02  Score=27.36  Aligned_cols=77  Identities=22%  Similarity=0.335  Sum_probs=43.1

Q ss_pred             CCCcEEEEcCCCCCCCCC---C-ceEEEEe-CCeEEEEc-CCCceEEEe-ccCCC-CC--C--------ceEEEEecCCC
Q 042843           76 SERTIVWVANREQPVSDR---F-SSVLRIS-DGNLVLFN-ESQLPIWST-NLTAT-SR--R--------SVEAVLLDEGN  137 (484)
Q Consensus        76 ~~~tvVW~ANr~~Pv~~~---~-~~~l~l~-~G~LvL~~-~~~~~vWst-~~~~~-~~--~--------~~~a~LlDsGN  137 (484)
                      ....++|...-+.++...   . +..+..+ +|.|..+| .+|.++|.. ..... ..  .        ........+|.
T Consensus        54 ~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~g~  133 (238)
T PF13360_consen   54 KTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYALDAKTGKVLWSIYLTSSPPAGVRSSSSPAVDGDRLYVGTSSGK  133 (238)
T ss_dssp             TTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEEEETTTSCEEEEEEE-SSCTCSTB--SEEEEETTEEEEEETCSE
T ss_pred             CCCCEEEEeeccccccceeeecccccccccceeeeEecccCCcceeeeeccccccccccccccCceEecCEEEEEeccCc
Confidence            456789988876554321   2 3334455 78888888 788899984 32210 00  0        11122233777


Q ss_pred             EEEEeccCCCCcceeee
Q 042843          138 LVLRDLSNNLSKPLWQS  154 (484)
Q Consensus       138 LVL~~~~~~~~~~lWQS  154 (484)
                      |+..|..+  ++.+|+=
T Consensus       134 l~~~d~~t--G~~~w~~  148 (238)
T PF13360_consen  134 LVALDPKT--GKLLWKY  148 (238)
T ss_dssp             EEEEETTT--TEEEEEE
T ss_pred             EEEEecCC--CcEEEEe
Confidence            77777542  4677765


No 49 
>KOG0291 consensus WD40-repeat-containing subunit of the 18S rRNA processing complex [RNA processing and modification]
Probab=47.38  E-value=4.8e+02  Score=29.71  Aligned_cols=85  Identities=20%  Similarity=0.351  Sum_probs=54.5

Q ss_pred             ceEEEEe-CCeEEEEcC-CCce-EEEeccC-------CCCCCceEEEEecCCCEEEEeccCCCCcceeeecccCceeccC
Q 042843           95 SSVLRIS-DGNLVLFNE-SQLP-IWSTNLT-------ATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQSFDHPAHTWIP  164 (484)
Q Consensus        95 ~~~l~l~-~G~LvL~~~-~~~~-vWst~~~-------~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQSFd~PTDTlLp  164 (484)
                      -..+..+ ||.++.+.. +|.+ ||.+...       ..+.+++..+..-+||.+|..+-                   -
T Consensus       353 i~~l~YSpDgq~iaTG~eDgKVKvWn~~SgfC~vTFteHts~Vt~v~f~~~g~~llssSL-------------------D  413 (893)
T KOG0291|consen  353 ITSLAYSPDGQLIATGAEDGKVKVWNTQSGFCFVTFTEHTSGVTAVQFTARGNVLLSSSL-------------------D  413 (893)
T ss_pred             eeeEEECCCCcEEEeccCCCcEEEEeccCceEEEEeccCCCceEEEEEEecCCEEEEeec-------------------C
Confidence            4568899 999998854 4666 9988641       12236788899999999997542                   2


Q ss_pred             CceeeeecCCCCceEEEecCCCCCCCCceEE-EEEcCCCC
Q 042843          165 GMKLTFNKRNNVSQLITSWKNKENPAPGLFS-LERAPDGS  203 (484)
Q Consensus       165 gq~l~~n~~~g~~~~L~Sw~s~~dps~G~y~-l~~~~~g~  203 (484)
                      |-.=-||.+.+.     -.|+-+-|.+-.|+ +..||.|.
T Consensus       414 GtVRAwDlkRYr-----NfRTft~P~p~QfscvavD~sGe  448 (893)
T KOG0291|consen  414 GTVRAWDLKRYR-----NFRTFTSPEPIQFSCVAVDPSGE  448 (893)
T ss_pred             CeEEeeeecccc-----eeeeecCCCceeeeEEEEcCCCC
Confidence            222233334331     22333456777776 77888885


No 50 
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=46.86  E-value=76  Score=32.84  Aligned_cols=57  Identities=18%  Similarity=0.239  Sum_probs=34.8

Q ss_pred             eEEEEe--CCeEEEEcC-CCceEEEeccCCCCCC------ceEEEEecCCCEEEEeccCCCCcceeee
Q 042843           96 SVLRIS--DGNLVLFNE-SQLPIWSTNLTATSRR------SVEAVLLDEGNLVLRDLSNNLSKPLWQS  154 (484)
Q Consensus        96 ~~l~l~--~G~LvL~~~-~~~~vWst~~~~~~~~------~~~a~LlDsGNLVL~~~~~~~~~~lWQS  154 (484)
                      ..+.+.  +|.|+-+|. +|+++|+.+..+....      .....-..+|.|+-.|..  +++++|+-
T Consensus       121 ~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~ssP~v~~~~v~v~~~~g~l~ald~~--tG~~~W~~  186 (394)
T PRK11138        121 GKVYIGSEKGQVYALNAEDGEVAWQTKVAGEALSRPVVSDGLVLVHTSNGMLQALNES--DGAVKWTV  186 (394)
T ss_pred             CEEEEEcCCCEEEEEECCCCCCcccccCCCceecCCEEECCEEEEECCCCEEEEEEcc--CCCEeeee
Confidence            455555  788888885 6899998875432101      111222346667777764  35788975


No 51 
>PF08374 Protocadherin:  Protocadherin;  InterPro: IPR013585 The structure of protocadherins is similar to that of classic cadherins (IPR002126 from INTERPRO), but they also have some unique features associated with the cytoplasmic domains. They are expressed in a variety of organisms and are found in high concentrations in the brain where they seem to be localised mainly at cell-cell contact sites. Their expression seems to be developmentally regulated []. 
Probab=45.33  E-value=13  Score=35.18  Aligned_cols=11  Identities=27%  Similarity=0.410  Sum_probs=4.9

Q ss_pred             ceEEEEEehHH
Q 042843          443 GVVIGGVVGSV  453 (484)
Q Consensus       443 ~~~i~~~v~~~  453 (484)
                      +++|++|.|++
T Consensus        38 ~I~iaiVAG~~   48 (221)
T PF08374_consen   38 KIMIAIVAGIM   48 (221)
T ss_pred             eeeeeeecchh
Confidence            44554444443


No 52 
>TIGR01478 STEVOR variant surface antigen, stevor family. This model represents the stevor branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of stevor sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 8 bits.
Probab=44.62  E-value=7.1  Score=38.51  Aligned_cols=7  Identities=29%  Similarity=1.075  Sum_probs=4.4

Q ss_pred             ccccccC
Q 042843          317 RFCSCLK  323 (484)
Q Consensus       317 ~~C~C~~  323 (484)
                      ..|+|-+
T Consensus       143 s~cectd  149 (295)
T TIGR01478       143 KSCECTN  149 (295)
T ss_pred             Cceeeec
Confidence            5677753


No 53 
>PTZ00370 STEVOR; Provisional
Probab=40.39  E-value=8  Score=38.22  Aligned_cols=7  Identities=29%  Similarity=0.952  Sum_probs=4.5

Q ss_pred             ccccccC
Q 042843          317 RFCSCLK  323 (484)
Q Consensus       317 ~~C~C~~  323 (484)
                      ..|+|-+
T Consensus       143 s~cectd  149 (296)
T PTZ00370        143 STCECTD  149 (296)
T ss_pred             Cceeeee
Confidence            4688753


No 54 
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=39.69  E-value=1.1e+02  Score=31.29  Aligned_cols=55  Identities=22%  Similarity=0.427  Sum_probs=28.7

Q ss_pred             EEEEe--CCeEEEEc-CCCceEEEeccCCCCCC------ceEEEEecCCCEEEEeccCCCCcceee
Q 042843           97 VLRIS--DGNLVLFN-ESQLPIWSTNLTATSRR------SVEAVLLDEGNLVLRDLSNNLSKPLWQ  153 (484)
Q Consensus        97 ~l~l~--~G~LvL~~-~~~~~vWst~~~~~~~~------~~~a~LlDsGNLVL~~~~~~~~~~lWQ  153 (484)
                      .+.+.  +|.|.-+| .+|+++|+.+.......      .....-..+|+|+-.|..  +++.+|+
T Consensus        67 ~v~v~~~~g~v~a~d~~tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~ald~~--tG~~~W~  130 (377)
T TIGR03300        67 KVYAADADGTVVALDAETGKRLWRVDLDERLSGGVGADGGLVFVGTEKGEVIALDAE--DGKELWR  130 (377)
T ss_pred             EEEEECCCCeEEEEEccCCcEeeeecCCCCcccceEEcCCEEEEEcCCCEEEEEECC--CCcEeee
Confidence            44444  67777777 56778887654322100      011111245666666643  2467786


No 55 
>PF02480 Herpes_gE:  Alphaherpesvirus glycoprotein E;  InterPro: IPR003404 Glycoprotein E (gE) of Alphaherpesvirus forms a complex with glycoprotein I (gI), functioning as an immunoglobulin G (IgG) Fc binding protein. gE is involved in virus spread but is not essential for propagation [].; GO: 0016020 membrane; PDB: 2GJ7_F 2GIY_B.
Probab=39.52  E-value=9.8  Score=40.44  Aligned_cols=13  Identities=23%  Similarity=0.593  Sum_probs=8.0

Q ss_pred             CCCcccC-CCCCCC
Q 042843          339 SGGCVRK-TPLQCE  351 (484)
Q Consensus       339 s~GC~r~-~~l~C~  351 (484)
                      ..+|.+. -+..|.
T Consensus       238 y~~C~~~~~~~~C~  251 (439)
T PF02480_consen  238 YANCSPSGWPRRCP  251 (439)
T ss_dssp             EEEEBTTC-TTTTE
T ss_pred             hcCCCCCCCcCCCC
Confidence            3588876 344784


No 56 
>PF06365 CD34_antigen:  CD34/Podocalyxin family;  InterPro: IPR013836 This family consists of several mammalian CD34 antigen proteins. The CD34 antigen is a human leukocyte membrane protein expressed specifically by lymphohematopoietic progenitor cells. CD34 is a phosphoprotein. Activation of protein kinase C (PKC) has been found to enhance CD34 phosphorylation [, ]. This family contains several eukaryotic podocalyxin proteins. Podocalyxin is a major membrane protein of the glomerular epithelium and is thought to be involved in maintenance of the architecture of the foot processes and filtration slits characteristic of this unique epithelium by virtue of its high negative charge. Podocalyxin functions as an anti-adhesin that maintains an open filtration pathway between neighbouring foot processes in the glomerular epithelium by charge repulsion [].
Probab=39.31  E-value=17  Score=34.24  Aligned_cols=18  Identities=17%  Similarity=0.519  Sum_probs=9.9

Q ss_pred             hcCCeEEEEeC-CCceEEe
Q 042843          392 NNCSCTAYAYK-DNACSIW  409 (484)
Q Consensus       392 ~nCSC~Ay~y~-~~~C~~w  409 (484)
                      .+|+..-+.-. +..|.++
T Consensus        36 ~~C~l~Laq~~~~~q~Lll   54 (202)
T PF06365_consen   36 DDCSLSLAQSEENQQCLLL   54 (202)
T ss_pred             CCcEEEEecCCCCcceEEE
Confidence            46666655432 3457765


No 57 
>PF01436 NHL:  NHL repeat;  InterPro: IPR001258 The NHL repeat, named after NCL-1, HT2A and Lin-41, is found largely in a large number of eukaryotic and prokaryotic proteins. For example, the repeat is found in a variety of enzymes of the copper type II, ascorbate-dependent monooxygenase family which catalyse the C terminus alpha-amidation of biological peptides []. In many it occurs in tandem arrays, for example in the ringfinger beta-box, coiled-coil (RBCC) eukaryotic growth regulators []. The 'Brain Tumor' protein (Brat) is one such growth regulator that contains a 6-bladed NHL-repeat beta-propeller [, ].  The NHL repeats are also found in serine/threonine protein kinase (STPK) in diverse range of pathogenic bacteria. These STPK are transmembrane receptors with a intracellular N-terminal kinase domain and extracellular C-terminal sensor domain. In the STPK, PknD, from Mycobacterium tuberculosis, the sensor domain forms a rigid, six-bladed b-propeller composed of NHL repeats with a flexible tether to the transmembrane domain.; GO: 0005515 protein binding; PDB: 3FVZ_A 3FW0_A 1RWL_A 1RWI_A 1Q7F_A.
Probab=38.43  E-value=56  Score=20.16  Aligned_cols=20  Identities=15%  Similarity=0.317  Sum_probs=13.9

Q ss_pred             EEEEe-CCeEEEEcCCCceEE
Q 042843           97 VLRIS-DGNLVLFNESQLPIW  116 (484)
Q Consensus        97 ~l~l~-~G~LvL~~~~~~~vW  116 (484)
                      -+.++ +|++++.|.....||
T Consensus         6 gvav~~~g~i~VaD~~n~rV~   26 (28)
T PF01436_consen    6 GVAVDSDGNIYVADSGNHRVQ   26 (28)
T ss_dssp             EEEEETTSEEEEEECCCTEEE
T ss_pred             EEEEeCCCCEEEEECCCCEEE
Confidence            36777 888888887655554


No 58 
>cd05845 Ig2_L1-CAM_like Second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM) and similar proteins. Ig2_L1-CAM_like: domain similar to the second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM). L1 belongs to the L1 subfamily of cell adhesion molecules (CAMs) and is comprised of an extracellular region having six Ig-like domains, five fibronectin type III domains, a transmembrane region and an intracellular domain. L1 is primarily expressed in the nervous system and is involved in its development and function. L1 is associated with an X-linked recessive disorder, X-linked hydrocephalus, MASA syndrome, or spastic paraplegia type 1, that involves abnormalities of axonal growth.
Probab=38.21  E-value=46  Score=27.41  Aligned_cols=31  Identities=16%  Similarity=0.324  Sum_probs=21.2

Q ss_pred             CCCcEEEEcCCCCCCCCCCceEEEEe-CCeEEEE
Q 042843           76 SERTIVWVANREQPVSDRFSSVLRIS-DGNLVLF  108 (484)
Q Consensus        76 ~~~tvVW~ANr~~Pv~~~~~~~l~l~-~G~LvL~  108 (484)
                      |..++.|+-+....+.  ...++.++ +|+|.+.
T Consensus        32 P~P~i~W~~~~~~~i~--~~~Ri~~~~~GnL~fs   63 (95)
T cd05845          32 VPLRIYWMNSDLLHIT--QDERVSMGQNGNLYFA   63 (95)
T ss_pred             CCCEEEEECCCCcccc--ccccEEECCCceEEEE
Confidence            5667888844434454  35678888 8999874


No 59 
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=37.59  E-value=4.7e+02  Score=26.86  Aligned_cols=75  Identities=20%  Similarity=0.344  Sum_probs=43.1

Q ss_pred             CCcEEEEcCCCCCCCCC---CceEEEEe--CCeEEEEcC-CCceEEEeccCCC-------CCC----ceEEEEecCCCEE
Q 042843           77 ERTIVWVANREQPVSDR---FSSVLRIS--DGNLVLFNE-SQLPIWSTNLTAT-------SRR----SVEAVLLDEGNLV  139 (484)
Q Consensus        77 ~~tvVW~ANr~~Pv~~~---~~~~l~l~--~G~LvL~~~-~~~~vWst~~~~~-------~~~----~~~a~LlDsGNLV  139 (484)
                      ...++|..+...++..+   ....+.+.  +|.|+-+|. +|+++|+.+....       ...    .....-..+|.++
T Consensus       139 tG~~~W~~~~~~~~~ssP~v~~~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~~~~~~~sP~v~~~~v~~~~~~g~v~  218 (394)
T PRK11138        139 DGEVAWQTKVAGEALSRPVVSDGLVLVHTSNGMLQALNESDGAVKWTVNLDVPSLTLRGESAPATAFGGAIVGGDNGRVS  218 (394)
T ss_pred             CCCCcccccCCCceecCCEEECCEEEEECCCCEEEEEEccCCCEeeeecCCCCcccccCCCCCEEECCEEEEEcCCCEEE
Confidence            45678988765433210   13344444  788888885 6889998764311       000    0111224566777


Q ss_pred             EEeccCCCCcceee
Q 042843          140 LRDLSNNLSKPLWQ  153 (484)
Q Consensus       140 L~~~~~~~~~~lWQ  153 (484)
                      -.+..+  ++.+|+
T Consensus       219 a~d~~~--G~~~W~  230 (394)
T PRK11138        219 AVLMEQ--GQLIWQ  230 (394)
T ss_pred             EEEccC--Chhhhe
Confidence            666543  578997


No 60 
>PF06697 DUF1191:  Protein of unknown function (DUF1191);  InterPro: IPR010605 This family contains hypothetical plant proteins of unknown function.
Probab=36.70  E-value=15  Score=36.47  Aligned_cols=8  Identities=13%  Similarity=0.127  Sum_probs=3.9

Q ss_pred             EEEEEecc
Q 042843          425 IIYIKLAA  432 (484)
Q Consensus       425 ~~yikv~~  432 (484)
                      .+-++...
T Consensus       187 slVV~~~~  194 (278)
T PF06697_consen  187 SLVVPSPA  194 (278)
T ss_pred             EEEEcCCC
Confidence            44555443


No 61 
>PF00954 S_locus_glycop:  S-locus glycoprotein family;  InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=36.08  E-value=84  Score=26.21  Aligned_cols=58  Identities=12%  Similarity=0.220  Sum_probs=36.0

Q ss_pred             CCeeEEEEeecCCceeEEEEEccCCcEEEEeeCCCCCCCeEEEeecCCCCCcccccCCCC--cccc
Q 042843          249 ENESYFTYNVKDSTYTSRAFMDVSGQDKQMNWLPLPTNSWFLFWSQPRQQCEVYALCGQF--STCN  312 (484)
Q Consensus       249 ~~~~~~~~~~~~~~~~~rl~Ld~dG~l~~y~w~~~~~~~W~~~w~~p~d~C~~~~~CG~~--giC~  312 (484)
                      +...+.++.+.....+++++.+.+.+-....|.... .. +..+    ..|..+++|-.+  ..|.
T Consensus        42 ~~s~~~r~~ld~~G~l~~~~w~~~~~~W~~~~~~p~-d~-Cd~y----~~CG~~g~C~~~~~~~C~  101 (110)
T PF00954_consen   42 NSSVLSRLVLDSDGQLQRYIWNESTQSWSVFWSAPK-DQ-CDVY----GFCGPNGICNSNNSPKCS  101 (110)
T ss_pred             CCceEEEEEEeeeeEEEEEEEecCCCcEEEEEEecc-cC-CCCc----cccCCccEeCCCCCCceE
Confidence            334444455555557888888777777776775543 22 2222    689999999654  3575


No 62 
>PF14991 MLANA:  Protein melan-A; PDB: 2GTZ_F 2GT9_F 3MRO_P 2GUO_C 3MRQ_P 2GTW_C 3L6F_C 3MRP_P.
Probab=34.63  E-value=12  Score=31.63  Aligned_cols=13  Identities=8%  Similarity=0.225  Sum_probs=0.0

Q ss_pred             HheeeeeeeccCc
Q 042843          463 MLVYLGRRKTATV  475 (484)
Q Consensus       463 ~~~~~~~r~~~~~  475 (484)
                      +..|.++||...+
T Consensus        42 iGCWYckRRSGYk   54 (118)
T PF14991_consen   42 IGCWYCKRRSGYK   54 (118)
T ss_dssp             -------------
T ss_pred             Hhheeeeecchhh
Confidence            3444444444433


No 63 
>PF12191 stn_TNFRSF12A:  Tumour necrosis factor receptor stn_TNFRSF12A_TNFR domain;  InterPro: IPR022316 The tumour necrosis factor (TNF) receptor (TNFR) superfamily comprises more than 20 type-I transmembrane proteins. Family members are defined based on similarity in their extracellular domain - a region that contains many cysteine residues arranged in a specific repetitive pattern []. The cysteines allow formation of an extended rod-like structure, responsible for ligand binding []. Upon receptor activation, different intracellular signalling complexes are assembled for different members of the TNFR superfamily, depending on their intracellular domains and sequences []. Activation of TNFRs can therefore induce a range of disparate effects, including cell proliferation, differentiation, survival, or apoptotic cell death, depending upon the receptor involved []. TNFRs are widely distributed and play important roles in many crucial biological processes, such as lymphoid and neuronal development, innate and adaptive immunity, and maintenance of cellular homeostasis []. Drugs that manipulate their signalling have potential roles in the prevention and treatment of many diseases, such as viral infections, coronary heart disease, transplant rejection, and immune disease []. TNF receptor 12 (also known as TWEAK receptor, and fibroblast growth factor-inducible-14 (Fn14)) has been implicated in endothelial cell growth and migration []. The receptor may also play a role in cell-matrix interactions [].; PDB: 2KN0_A 2RPJ_A 2KMZ_A 2EQP_A.
Probab=32.94  E-value=12  Score=32.45  Aligned_cols=10  Identities=20%  Similarity=-0.077  Sum_probs=0.0

Q ss_pred             eeeeccCccc
Q 042843          468 GRRKTATVTT  477 (484)
Q Consensus       468 ~~r~~~~~~~  477 (484)
                      +||-|++++.
T Consensus       101 ~rrcrrr~~~  110 (129)
T PF12191_consen  101 WRRCRRREKF  110 (129)
T ss_dssp             ----------
T ss_pred             HhhhhccccC
Confidence            3544555554


No 64 
>PF13360 PQQ_2:  PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=32.07  E-value=97  Score=28.89  Aligned_cols=50  Identities=28%  Similarity=0.445  Sum_probs=28.9

Q ss_pred             CCeEEEEcC-CCceEEEeccCCCCCCceEE-EE---------ecCCCEEEEeccCCCCcceeee
Q 042843          102 DGNLVLFNE-SQLPIWSTNLTATSRRSVEA-VL---------LDEGNLVLRDLSNNLSKPLWQS  154 (484)
Q Consensus       102 ~G~LvL~~~-~~~~vWst~~~~~~~~~~~a-~L---------lDsGNLVL~~~~~~~~~~lWQS  154 (484)
                      +|.|...|. +|+.+|+.+..... ....+ .+         ..+|+|+..|..  +++++|+-
T Consensus         2 ~g~l~~~d~~tG~~~W~~~~~~~~-~~~~~~~~~~~~~v~~~~~~~~l~~~d~~--tG~~~W~~   62 (238)
T PF13360_consen    2 DGTLSALDPRTGKELWSYDLGPGI-GGPVATAVPDGGRVYVASGDGNLYALDAK--TGKVLWRF   62 (238)
T ss_dssp             TSEEEEEETTTTEEEEEEECSSSC-SSEEETEEEETTEEEEEETTSEEEEEETT--TSEEEEEE
T ss_pred             CCEEEEEECCCCCEEEEEECCCCC-CCccceEEEeCCEEEEEcCCCEEEEEECC--CCCEEEEe
Confidence            477777886 78889988642111 11111 12         366666666653  25778875


No 65 
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=31.14  E-value=3.9e+02  Score=27.16  Aligned_cols=75  Identities=13%  Similarity=0.353  Sum_probs=44.6

Q ss_pred             CCCcEEEEcCCCC-----CCCCCCceEEEEe--CCeEEEEcC-CCceEEEeccCCCCC------CceEEEEecCCCEEEE
Q 042843           76 SERTIVWVANREQ-----PVSDRFSSVLRIS--DGNLVLFNE-SQLPIWSTNLTATSR------RSVEAVLLDEGNLVLR  141 (484)
Q Consensus        76 ~~~tvVW~ANr~~-----Pv~~~~~~~l~l~--~G~LvL~~~-~~~~vWst~~~~~~~------~~~~a~LlDsGNLVL~  141 (484)
                      ....++|.-+-..     |+.  ....+.+.  +|.|+-+|. +|+++|+....+...      ......-..+|.|+..
T Consensus        83 ~tG~~~W~~~~~~~~~~~p~v--~~~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~a~  160 (377)
T TIGR03300        83 ETGKRLWRVDLDERLSGGVGA--DGGLVFVGTEKGEVIALDAEDGKELWRAKLSSEVLSPPLVANGLVVVRTNDGRLTAL  160 (377)
T ss_pred             cCCcEeeeecCCCCcccceEE--cCCEEEEEcCCCEEEEEECCCCcEeeeeccCceeecCCEEECCEEEEECCCCeEEEE
Confidence            3567888755433     333  23455554  899998886 689999876533210      1111222356777777


Q ss_pred             eccCCCCcceeee
Q 042843          142 DLSNNLSKPLWQS  154 (484)
Q Consensus       142 ~~~~~~~~~lWQS  154 (484)
                      |..  +++++|+-
T Consensus       161 d~~--tG~~~W~~  171 (377)
T TIGR03300       161 DAA--TGERLWTY  171 (377)
T ss_pred             EcC--CCceeeEE
Confidence            764  35789984


No 66 
>PF15330 SIT:  SHP2-interacting transmembrane adaptor protein, SIT
Probab=30.91  E-value=13  Score=31.44  Aligned_cols=9  Identities=22%  Similarity=0.154  Sum_probs=4.1

Q ss_pred             heeeeeeec
Q 042843          464 LVYLGRRKT  472 (484)
Q Consensus       464 ~~~~~~r~~  472 (484)
                      +.++.||.+
T Consensus        17 asl~~wr~~   25 (107)
T PF15330_consen   17 ASLLAWRMK   25 (107)
T ss_pred             HHHHHHHHH
Confidence            344445544


No 67 
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=28.62  E-value=18  Score=29.88  Aligned_cols=32  Identities=9%  Similarity=-0.111  Sum_probs=17.9

Q ss_pred             eEEEEEehHHHHHHHHHHH-HheeeeeeeccCc
Q 042843          444 VVIGGVVGSVAVVALIGLI-MLVYLGRRKTATV  475 (484)
Q Consensus       444 ~~i~~~v~~~~~~~~~~~~-~~~~~~~r~~~~~  475 (484)
                      .-.+++++++|.+++++.+ ++++++|..+|+|
T Consensus        63 ls~gaiagi~vg~~~~v~~lv~~l~w~f~~r~k   95 (96)
T PTZ00382         63 LSTGAIAGISVAVVAVVGGLVGFLCWWFVCRGK   95 (96)
T ss_pred             cccccEEEEEeehhhHHHHHHHHHhheeEEeec
Confidence            4466777776665544433 4455556555543


No 68 
>PF05393 Hum_adeno_E3A:  Human adenovirus early E3A glycoprotein;  InterPro: IPR008652 This family consists of several early glycoproteins (E3A), from human adenovirus type 2.; GO: 0016021 integral to membrane
Probab=28.53  E-value=14  Score=29.73  Aligned_cols=26  Identities=12%  Similarity=0.245  Sum_probs=11.1

Q ss_pred             ehHHHHHHHHHHHHheeeeeeeccCcc
Q 042843          450 VGSVAVVALIGLIMLVYLGRRKTATVT  476 (484)
Q Consensus       450 v~~~~~~~~~~~~~~~~~~~r~~~~~~  476 (484)
                      .+++.+.|++ ++++.+..|++|+|.+
T Consensus        37 ~lvI~~iFil-~VilwfvCC~kRkrsR   62 (94)
T PF05393_consen   37 FLVICGIFIL-LVILWFVCCKKRKRSR   62 (94)
T ss_pred             HHHHHHHHHH-HHHHHHHHHHHhhhcc
Confidence            3344343433 3334444455554443


No 69 
>PF14670 FXa_inhibition:  Coagulation Factor Xa inhibitory site; PDB: 3Q3K_B 1NFY_B 1LQD_A 1G2L_B 1IQF_L 2UWP_B 2VH6_B 3KQC_L 2P93_L 2BQW_A ....
Probab=28.02  E-value=19  Score=24.05  Aligned_cols=14  Identities=29%  Similarity=0.657  Sum_probs=9.9

Q ss_pred             CCccccccCCCccC
Q 042843          315 TERFCSCLKGFQQK  328 (484)
Q Consensus       315 ~~~~C~C~~GF~p~  328 (484)
                      ....|+|++||...
T Consensus        17 g~~~C~C~~Gy~L~   30 (36)
T PF14670_consen   17 GSYRCSCPPGYKLA   30 (36)
T ss_dssp             TSEEEE-STTEEE-
T ss_pred             CceEeECCCCCEEC
Confidence            45689999999875


No 70 
>PF02009 Rifin_STEVOR:  Rifin/stevor family;  InterPro: IPR002858 Malaria is still a major cause of mortality in many areas of the world. Plasmodium falciparum causes the most severe human form of the disease and is responsible for most fatalities. Severe cases of malaria can occur when the parasite invades and then proliferates within red blood cell erythrocytes. The parasite produces many variant antigenic proteins, encoded by multigene families, which are present on the surface of the infected erythrocyte and play important roles in virulence. A crucial survival mechanism for the malaria parasite is its ability to evade the immune response by switching these variant surface antigens. The high virulence of P. falciparum relative to other malarial parasites is in large part due to the fact that in this organism many of these surface antigens mediate the binding of infected erythrocytes to the vascular endothelium (cytoadherence) and non-infected erythrocytes (rosetting). This can lead to the accumulation of infected cells in the vasculature of a variety of organs, blocking the blood flow and reducing the oxygen supply. Clinical symptoms of severe infection can include fever, progressive anaemia, multi-organ dysfunction and coma. For more information see []. Several multicopy gene families have been described in Plasmodium falciparum, including the stevor family of subtelomeric open reading frames and the rif interspersed repetitive elements. Both families contain three predicted transmembrane segments. It has been proposed that stevor and rif are members of a larger superfamily that code for variant surface antigens [].
Probab=27.92  E-value=6.5  Score=39.49  Aligned_cols=28  Identities=21%  Similarity=0.368  Sum_probs=15.9

Q ss_pred             EehHHHHHHHHHHHHheeeeeeeccCcc
Q 042843          449 VVGSVAVVALIGLIMLVYLGRRKTATVT  476 (484)
Q Consensus       449 ~v~~~~~~~~~~~~~~~~~~~r~~~~~~  476 (484)
                      +++++|.+++++++.+.+.+||+|+-++
T Consensus       262 iiaIliIVLIMvIIYLILRYRRKKKmkK  289 (299)
T PF02009_consen  262 IIAILIIVLIMVIIYLILRYRRKKKMKK  289 (299)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHHhhhhH
Confidence            3333333444666777777888665443


No 71 
>PHA02887 EGF-like protein; Provisional
Probab=27.85  E-value=62  Score=27.73  Aligned_cols=30  Identities=30%  Similarity=0.775  Sum_probs=22.3

Q ss_pred             CCCCcc--cccCCCCccccC---CCCccccccCCCc
Q 042843          296 RQQCEV--YALCGQFSTCNQ---QTERFCSCLKGFQ  326 (484)
Q Consensus       296 ~d~C~~--~~~CG~~giC~~---~~~~~C~C~~GF~  326 (484)
                      -++|.-  .++|= +|.|..   -+.+.|.|++||.
T Consensus        83 f~pC~~eyk~YCi-HG~C~yI~dL~epsCrC~~GYt  117 (126)
T PHA02887         83 FEKCKNDFNDFCI-NGECMNIIDLDEKFCICNKGYT  117 (126)
T ss_pred             ccccChHhhCEee-CCEEEccccCCCceeECCCCcc
Confidence            367853  57787 789974   2458999999985


No 72 
>PF12946 EGF_MSP1_1:  MSP1 EGF domain 1;  InterPro: IPR024730 This EGF-like domain is found at the C terminus of the malaria parasite MSP1 protein. MSP1 is the merozoite surface protein 1. This domain is part of the C-terminal fragment that is proteolytically processed from the the rest of the protein and is left attached to the surface of the invading parasite [].; PDB: 1N1I_C 2FLG_A 1CEJ_A 2NPR_A 1B9W_A 1OB1_F.
Probab=27.84  E-value=18  Score=24.38  Aligned_cols=25  Identities=24%  Similarity=0.679  Sum_probs=16.3

Q ss_pred             ccCCCCccccC--CCCccccccCCCcc
Q 042843          303 ALCGQFSTCNQ--QTERFCSCLKGFQQ  327 (484)
Q Consensus       303 ~~CG~~giC~~--~~~~~C~C~~GF~p  327 (484)
                      ..|-.++-|-.  +.+..|.|++||..
T Consensus         5 ~~cP~NA~C~~~~dG~eecrCllgyk~   31 (37)
T PF12946_consen    5 TKCPANAGCFRYDDGSEECRCLLGYKK   31 (37)
T ss_dssp             S---TTEEEEEETTSEEEEEE-TTEEE
T ss_pred             ccCCCCcccEEcCCCCEEEEeeCCccc
Confidence            45778888963  35678999999975


No 73 
>PF15102 TMEM154:  TMEM154 protein family
Probab=27.67  E-value=48  Score=29.57  Aligned_cols=15  Identities=13%  Similarity=-0.248  Sum_probs=6.6

Q ss_pred             HheeeeeeeccCccc
Q 042843          463 MLVYLGRRKTATVTT  477 (484)
Q Consensus       463 ~~~~~~~r~~~~~~~  477 (484)
                      +..+++-.+.|++|.
T Consensus        73 l~vV~lv~~~kRkr~   87 (146)
T PF15102_consen   73 LSVVCLVIYYKRKRT   87 (146)
T ss_pred             HHHHHheeEEeeccc
Confidence            334444444444444


No 74 
>KOG1219 consensus Uncharacterized conserved protein, contains laminin, cadherin and EGF domains [Signal transduction mechanisms]
Probab=27.20  E-value=51  Score=41.80  Aligned_cols=24  Identities=25%  Similarity=0.509  Sum_probs=17.9

Q ss_pred             ccCCCCccccCC--CCccccccCCCc
Q 042843          303 ALCGQFSTCNQQ--TERFCSCLKGFQ  326 (484)
Q Consensus       303 ~~CG~~giC~~~--~~~~C~C~~GF~  326 (484)
                      ..|---|.|+..  +.-.|.||+-|.
T Consensus      3870 npCqhgG~C~~~~~ggy~CkCpsqys 3895 (4289)
T KOG1219|consen 3870 NPCQHGGTCISQPKGGYKCKCPSQYS 3895 (4289)
T ss_pred             CcccCCCEecCCCCCceEEeCccccc
Confidence            567778899853  335799999875


No 75 
>PF14870 PSII_BNR:  Photosynthesis system II assembly factor YCF48; PDB: 2XBG_A.
Probab=26.96  E-value=6.6e+02  Score=25.31  Aligned_cols=98  Identities=21%  Similarity=0.289  Sum_probs=0.0

Q ss_pred             CceEEEEe-CCeEEEEcCCCceEEEeccCCCCCCceEEEEecCCCEEEEeccCCCCcceeeecccCceeccCCceeeeec
Q 042843           94 FSSVLRIS-DGNLVLFNESQLPIWSTNLTATSRRSVEAVLLDEGNLVLRDLSNNLSKPLWQSFDHPAHTWIPGMKLTFNK  172 (484)
Q Consensus        94 ~~~~l~l~-~G~LvL~~~~~~~vWst~~~~~~~~~~~a~LlDsGNLVL~~~~~~~~~~lWQSFd~PTDTlLpgq~l~~n~  172 (484)
                      ....+..+ .|++......|.. |.............+.-+++|..|+....+    .+.+|+|.--+++.|-++..   
T Consensus       114 ~~~~~l~~~~G~iy~T~DgG~t-W~~~~~~~~gs~~~~~r~~dG~~vavs~~G----~~~~s~~~G~~~w~~~~r~~---  185 (302)
T PF14870_consen  114 DGSAELAGDRGAIYRTTDGGKT-WQAVVSETSGSINDITRSSDGRYVAVSSRG----NFYSSWDPGQTTWQPHNRNS---  185 (302)
T ss_dssp             TTEEEEEETT--EEEESSTTSS-EEEEE-S----EEEEEE-TTS-EEEEETTS----SEEEEE-TT-SS-EEEE--S---
T ss_pred             CCcEEEEcCCCcEEEeCCCCCC-eeEcccCCcceeEeEEECCCCcEEEEECcc----cEEEEecCCCccceEEccCc---


Q ss_pred             CCCCceEEEecCCCCCCCCceEEEEEcCCCCcEEEEEeeCCeeEEe
Q 042843          173 RNNVSQLITSWKNKENPAPGLFSLERAPDGSNQYVMLWNRSEQYWS  218 (484)
Q Consensus       173 ~~g~~~~L~Sw~s~~dps~G~y~l~~~~~g~~~~~l~~~~~~~Yw~  218 (484)
                          .++|.             .+.+.+++.  +.+.-+|.+.+..
T Consensus       186 ----~~riq-------------~~gf~~~~~--lw~~~~Gg~~~~s  212 (302)
T PF14870_consen  186 ----SRRIQ-------------SMGFSPDGN--LWMLARGGQIQFS  212 (302)
T ss_dssp             ----SS-EE-------------EEEE-TTS---EEEEETTTEEEEE
T ss_pred             ----cceeh-------------hceecCCCC--EEEEeCCcEEEEc


No 76 
>KOG1214 consensus Nidogen and related basement membrane protein proteins [Cell wall/membrane/envelope biogenesis; Extracellular structures]
Probab=25.57  E-value=51  Score=37.40  Aligned_cols=30  Identities=23%  Similarity=0.609  Sum_probs=23.6

Q ss_pred             CCCcccccCCCCccccCC-CCccccccCCCcc
Q 042843          297 QQCEVYALCGQFSTCNQQ-TERFCSCLKGFQQ  327 (484)
Q Consensus       297 d~C~~~~~CG~~giC~~~-~~~~C~C~~GF~p  327 (484)
                      |+|. +..|-+...|-+. .+..|.|.|||.-
T Consensus       828 DeC~-psrChp~A~CyntpgsfsC~C~pGy~G  858 (1289)
T KOG1214|consen  828 DECS-PSRCHPAATCYNTPGSFSCRCQPGYYG  858 (1289)
T ss_pred             cccC-ccccCCCceEecCCCcceeecccCccC
Confidence            6777 7889999999754 3567999999963


No 77 
>PF14575 EphA2_TM:  Ephrin type-A receptor 2 transmembrane domain; PDB: 3KUL_A 2XVD_A 2VX1_A 2VWV_A 2VX0_A 2VWY_A 2VWZ_A 2VWW_A 2VWU_A 2VWX_A ....
Probab=24.13  E-value=22  Score=27.93  Aligned_cols=9  Identities=22%  Similarity=0.217  Sum_probs=3.0

Q ss_pred             eeeeeeecc
Q 042843          465 VYLGRRKTA  473 (484)
Q Consensus       465 ~~~~~r~~~  473 (484)
                      +++++|+++
T Consensus        20 ~~~~~rr~~   28 (75)
T PF14575_consen   20 VIVCFRRCK   28 (75)
T ss_dssp             HHCCCTT--
T ss_pred             EEEEEeeEc
Confidence            334444443


No 78 
>PF06247 Plasmod_Pvs28:  Plasmodium ookinete surface protein Pvs28;  InterPro: IPR010423 This family consists of several ookinete surface protein (Pvs28) from several species of Plasmodium. Pvs25 and Pvs28 are expressed on the surface of ookinetes. These proteins are potential candidates for vaccine and induce antibodies that block the infectivity of Plasmodium vivax in immunised animals [].; GO: 0009986 cell surface, 0016020 membrane; PDB: 1Z3G_B 1Z1Y_B 1Z27_A.
Probab=21.21  E-value=33  Score=31.89  Aligned_cols=27  Identities=30%  Similarity=0.830  Sum_probs=19.6

Q ss_pred             cccCCCCccccCCC------CccccccCCCccC
Q 042843          302 YALCGQFSTCNQQT------ERFCSCLKGFQQK  328 (484)
Q Consensus       302 ~~~CG~~giC~~~~------~~~C~C~~GF~p~  328 (484)
                      .-.||.|+.|...+      .-.|.|.+||...
T Consensus        49 ~K~Cgdya~C~~~~~~~~~~~~~C~C~~gY~~~   81 (197)
T PF06247_consen   49 NKPCGDYAKCINQANKGEERAYKCDCINGYILK   81 (197)
T ss_dssp             TSEEETTEEEEE-SSTTSSTSEEEEE-TTEEES
T ss_pred             CccccchhhhhcCCCcccceeEEEecccCceee
Confidence            45799999998422      2469999999874


No 79 
>smart00564 PQQ beta-propeller repeat. Beta-propeller repeat occurring in enzymes with pyrrolo-quinoline quinone (PQQ) as cofactor, in Ire1p-like Ser/Thr kinases, and in prokaryotic dehydrogenases.
Probab=20.54  E-value=1.9e+02  Score=17.80  Aligned_cols=16  Identities=25%  Similarity=0.652  Sum_probs=9.4

Q ss_pred             CCeEEEEcC-CCceEEE
Q 042843          102 DGNLVLFNE-SQLPIWS  117 (484)
Q Consensus       102 ~G~LvL~~~-~~~~vWs  117 (484)
                      +|.|+-.|. +|..+|.
T Consensus        15 ~g~l~a~d~~~G~~~W~   31 (33)
T smart00564       15 DGTLYALDAKTGEILWT   31 (33)
T ss_pred             CCEEEEEEcccCcEEEE
Confidence            566665554 4566665


Done!