Query         043869
Match_columns 450
No_of_seqs    240 out of 1540
Neff          7.6 
Searched_HMMs 46136
Date          Fri Mar 29 09:04:23 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/043869.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/043869hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF01453 B_lectin:  D-mannose b  99.9 1.5E-28 3.2E-33  209.4   5.5  102   70-172     1-114 (114)
  2 cd00028 B_lectin Bulb-type man  99.9 1.1E-24 2.5E-29  186.2  14.5  107   36-146     7-116 (116)
  3 smart00108 B_lectin Bulb-type   99.9 2.1E-23 4.5E-28  177.9  14.2  106   36-145     7-114 (114)
  4 PF00954 S_locus_glycop:  S-loc  99.8 3.6E-19 7.9E-24  150.7  10.7   70  226-298    41-110 (110)
  5 PF08276 PAN_2:  PAN-like domai  99.5 1.8E-14   4E-19  110.4   5.3   58  317-374     4-66  (66)
  6 cd01098 PAN_AP_plant Plant PAN  99.5 6.2E-14 1.3E-18  112.2   7.3   71  318-389    12-84  (84)
  7 cd00129 PAN_APPLE PAN/APPLE-li  99.3 6.1E-12 1.3E-16   99.8   5.9   66  318-388     9-80  (80)
  8 smart00108 B_lectin Bulb-type   98.7 5.4E-08 1.2E-12   82.7   8.8   73   90-189    24-101 (114)
  9 smart00473 PAN_AP divergent su  98.7 7.4E-08 1.6E-12   75.3   7.5   70  318-387     4-77  (78)
 10 cd00028 B_lectin Bulb-type man  98.6 8.8E-08 1.9E-12   81.7   7.9   72   91-189    25-102 (116)
 11 PF01453 B_lectin:  D-mannose b  98.5 1.8E-06   4E-11   73.4  10.9  100   37-147    12-114 (114)
 12 cd01100 APPLE_Factor_XI_like S  97.9 1.1E-05 2.4E-10   62.9   4.1   48  322-369     8-57  (73)
 13 smart00223 APPLE APPLE domain.  95.2   0.026 5.7E-07   44.6   4.0   47  323-369     6-57  (79)
 14 PF00024 PAN_1:  PAN domain Thi  94.5    0.07 1.5E-06   41.2   4.7   51  319-369     3-56  (79)
 15 PF14295 PAN_4:  PAN domain; PD  94.3   0.041 8.9E-07   38.9   2.8   41  327-367     2-51  (51)
 16 PF08277 PAN_3:  PAN-like domai  90.9     1.1 2.5E-05   34.0   6.9   32  337-368    18-49  (71)
 17 smart00605 CW CW domain.        89.7     2.1 4.6E-05   34.7   8.0   55  337-391    20-77  (94)
 18 PF04478 Mid2:  Mid2 like cell   89.0    0.45 9.8E-06   42.1   3.6   38  412-450    50-87  (154)
 19 PF08693 SKG6:  Transmembrane a  88.7    0.45 9.8E-06   32.3   2.6   31  409-439     8-39  (40)
 20 cd01099 PAN_AP_HGF Subfamily o  87.9     1.9 4.2E-05   33.9   6.3   33  337-369    23-59  (80)
 21 PF07645 EGF_CA:  Calcium-bindi  87.6    0.24 5.2E-06   34.0   0.8   32  265-296     3-35  (42)
 22 cd00053 EGF Epidermal growth f  87.3    0.45 9.8E-06   30.2   2.0   30  267-296     2-31  (36)
 23 PF15102 TMEM154:  TMEM154 prot  86.4    0.92   2E-05   39.9   3.9   27  414-440    61-87  (146)
 24 smart00179 EGF_CA Calcium-bind  82.1     1.1 2.4E-05   29.3   2.1   31  265-295     3-33  (39)
 25 PTZ00382 Variant-specific surf  80.6     1.9 4.1E-05   35.4   3.3   14  426-439    82-95  (96)
 26 cd00054 EGF_CA Calcium-binding  78.4     1.8 3.8E-05   27.9   2.1   32  265-296     3-34  (38)
 27 PF01683 EB:  EB module;  Inter  76.5     2.5 5.3E-05   30.2   2.5   33  262-298    17-49  (52)
 28 PF02009 Rifin_STEVOR:  Rifin/s  75.6    0.51 1.1E-05   46.8  -1.7   32  415-446   258-291 (299)
 29 PF12947 EGF_3:  EGF domain;  I  74.9    0.65 1.4E-05   30.9  -0.8   27  270-296     5-31  (36)
 30 PF01102 Glycophorin_A:  Glycop  74.1    0.57 1.2E-05   40.1  -1.5   28  413-442    68-95  (122)
 31 PF01299 Lamp:  Lysosome-associ  73.1     2.5 5.5E-05   42.1   2.6   19  428-446   287-306 (306)
 32 PF02439 Adeno_E3_CR2:  Adenovi  70.8    0.97 2.1E-05   30.2  -0.7   29  415-443     8-36  (38)
 33 PF06697 DUF1191:  Protein of u  69.1     3.6 7.9E-05   40.2   2.5   36  413-448   214-250 (278)
 34 PF00008 EGF:  EGF-like domain   69.0     1.1 2.5E-05   28.7  -0.7   24  272-295     5-29  (32)
 35 PF12661 hEGF:  Human growth fa  68.8     1.1 2.4E-05   22.9  -0.6    9  287-295     1-9   (13)
 36 smart00181 EGF Epidermal growt  67.6     4.4 9.5E-05   25.9   1.9   25  271-296     6-30  (35)
 37 PF14575 EphA2_TM:  Ephrin type  66.6     1.5 3.1E-05   34.4  -0.6   25  415-439     3-27  (75)
 38 PTZ00046 rifin; Provisional     63.8     1.7 3.8E-05   43.9  -0.8   24  415-438   317-340 (358)
 39 TIGR01477 RIFIN variant surfac  63.3     1.8 3.9E-05   43.7  -0.9   33  415-447   312-346 (353)
 40 PF12662 cEGF:  Complement Clr-  62.3       4 8.6E-05   24.6   0.8   11  287-297     3-13  (24)
 41 PF13908 Shisa:  Wnt and FGF in  60.5     7.2 0.00016   35.6   2.7   20  411-430    77-96  (179)
 42 PF06024 DUF912:  Nucleopolyhed  58.8     7.7 0.00017   32.1   2.3    9  382-390    15-23  (101)
 43 PF06365 CD34_antigen:  CD34/Po  58.4     9.4  0.0002   35.7   3.0   36  413-448   101-139 (202)
 44 PF07974 EGF_2:  EGF-like domai  57.8     7.7 0.00017   25.0   1.7   24  271-296     6-29  (32)
 45 PF05454 DAG1:  Dystroglycan (D  54.3     4.2 9.1E-05   40.2   0.0   29  410-438   145-173 (290)
 46 PRK11138 outer membrane biogen  53.9      49  0.0011   33.9   7.9   71   70-142    88-186 (394)
 47 KOG1219 Uncharacterized conser  53.2      15 0.00032   45.7   4.1   32  265-297  3865-3897(4289)
 48 PF13360 PQQ_2:  PQQ-like domai  50.3 1.5E+02  0.0031   27.3   9.9   75   68-142    53-148 (238)
 49 PF09064 Tme5_EGF_like:  Thromb  49.8     9.4  0.0002   25.0   1.1   18  278-296    11-28  (34)
 50 PF15102 TMEM154:  TMEM154 prot  48.1      16 0.00034   32.3   2.6   31  415-445    58-89  (146)
 51 PF12768 Rax2:  Cortical protei  46.5      12 0.00025   37.1   1.7   27  408-434   226-252 (281)
 52 PF14670 FXa_inhibition:  Coagu  45.3     5.6 0.00012   26.4  -0.5   21  278-298    11-31  (36)
 53 PRK11138 outer membrane biogen  45.2 1.2E+02  0.0027   30.9   9.2   48   94-142   302-361 (394)
 54 cd05845 Ig2_L1-CAM_like Second  44.8      35 0.00075   27.9   4.0   35   68-103    31-65  (95)
 55 PF12690 BsuPI:  Intracellular   44.7      28 0.00061   27.5   3.4   16   97-113    27-42  (82)
 56 KOG0291 WD40-repeat-containing  43.4 4.7E+02    0.01   29.5  13.1   56   86-141   352-420 (893)
 57 TIGR03300 assembly_YfgL outer   43.3 1.7E+02  0.0038   29.4   9.9   50   93-142   248-305 (377)
 58 PF01436 NHL:  NHL repeat;  Int  41.3      49  0.0011   20.2   3.4   21   89-110     6-26  (28)
 59 PF13360 PQQ_2:  PQQ-like domai  41.2 1.3E+02  0.0028   27.7   8.0   73   70-142    12-102 (238)
 60 PF03302 VSP:  Giardia variant-  41.2      22 0.00048   36.9   2.9   15  425-439   382-396 (397)
 61 PF12877 DUF3827:  Domain of un  40.3      21 0.00045   38.9   2.5   30  410-439   267-297 (684)
 62 PF12191 stn_TNFRSF12A:  Tumour  40.1       8 0.00017   33.1  -0.4   37  413-450    79-118 (129)
 63 PHA02887 EGF-like protein; Pro  39.8      32 0.00069   29.2   3.0   35  260-295    79-117 (126)
 64 PF05337 CSF-1:  Macrophage col  39.4     9.8 0.00021   37.0   0.0   30  419-448   231-260 (285)
 65 PHA03099 epidermal growth fact  38.8      16 0.00034   31.5   1.1   31  417-447   104-134 (139)
 66 TIGR03300 assembly_YfgL outer   37.8 2.1E+02  0.0046   28.7   9.6   18  125-142   249-267 (377)
 67 PF15065 NCU-G1:  Lysosomal tra  37.6      12 0.00026   38.1   0.2   32  411-442   318-349 (350)
 68 COG3763 Uncharacterized protei  37.3     4.3 9.4E-05   31.0  -2.2   28  416-443     5-32  (71)
 69 PF01034 Syndecan:  Syndecan do  35.7      11 0.00024   28.4  -0.2    8  432-439    30-37  (64)
 70 PF14991 MLANA:  Protein melan-  35.1      12 0.00026   31.5  -0.1   17  424-440    36-52  (118)
 71 PF07354 Sp38:  Zona-pellucida-  34.3      52  0.0011   32.0   4.0   35   68-103    10-44  (271)
 72 KOG0640 mRNA cleavage stimulat  32.5 1.6E+02  0.0034   29.6   6.9   65   76-141   253-333 (430)
 73 KOG3637 Vitronectin receptor,   32.3      31 0.00067   40.3   2.5   22  412-433   978-1000(1030)
 74 TIGR03066 Gem_osc_para_1 Gemma  31.6 1.8E+02  0.0039   24.6   6.3   52   86-138    34-104 (111)
 75 PF06006 DUF905:  Bacterial pro  31.3      53  0.0011   25.1   2.8   20  128-147    34-53  (70)
 76 KOG4649 PQQ (pyrrolo-quinoline  29.3   2E+02  0.0043   28.3   6.9   40   71-111   169-213 (354)
 77 PF12946 EGF_MSP1_1:  MSP1 EGF   29.2      11 0.00025   25.2  -1.0   27  271-297     5-32  (37)
 78 PHA03290 envelope glycoprotein  29.1      52  0.0011   33.0   3.1   35   12-55     11-45  (357)
 79 TIGR03075 PQQ_enz_alc_DH PQQ-d  28.4 7.7E+02   0.017   26.5  13.2  121   50-178    73-250 (527)
 80 PF08114 PMP1_2:  ATPase proteo  28.1     6.4 0.00014   26.8  -2.3   19  429-447    23-41  (43)
 81 PF02480 Herpes_gE:  Alphaherpe  28.1      20 0.00042   37.8   0.0   30  302-331   237-267 (439)
 82 cd00216 PQQ_DH Dehydrogenases   27.6 1.9E+02  0.0042   30.7   7.5   71   69-141    37-135 (488)
 83 PF05545 FixQ:  Cbb3-type cytoc  27.6      21 0.00045   25.2   0.0   14  432-445    27-40  (49)
 84 PF11403 Yeast_MT:  Yeast metal  26.3      47   0.001   21.6   1.5   19  271-294    12-30  (40)
 85 PF12458 DUF3686:  ATPase invol  25.9 1.7E+02  0.0037   30.6   6.2   55   38-103   312-367 (448)
 86 PF00558 Vpu:  Vpu protein;  In  24.7      36 0.00078   27.0   0.9    6  441-446    34-39  (81)
 87 PTZ00382 Variant-specific surf  24.6      27 0.00059   28.6   0.3   28  415-442    68-95  (96)
 88 KOG1214 Nidogen and related ba  24.0      55  0.0012   36.8   2.5   31  265-296   828-858 (1289)
 89 PRK12785 fliL flagellar basal   23.1      95  0.0021   28.0   3.5   18  421-438    32-49  (166)
 90 PF02237 BPL_C:  Biotin protein  22.1 1.1E+02  0.0024   21.3   3.0   14   90-103    20-33  (48)
 91 PF06247 Plasmod_Pvs28:  Plasmo  22.1      23 0.00049   32.7  -0.7   41  264-309    39-88  (197)
 92 PF05568 ASFV_J13L:  African sw  21.8      17 0.00036   31.9  -1.6   16  432-447    49-64  (189)
 93 PRK01844 hypothetical protein;  20.5      11 0.00024   29.1  -2.6   25  418-442     7-31  (72)
 94 TIGR03503 conserved hypothetic  20.4      52  0.0011   33.8   1.3   21  419-440   353-373 (374)
 95 PF14316 DUF4381:  Domain of un  20.4      23  0.0005   31.2  -1.1   11  435-445    42-52  (146)
 96 KOG4289 Cadherin EGF LAG seven  20.2      50  0.0011   39.5   1.3   40  270-309  1244-1285(2531)
 97 cd05852 Ig5_Contactin-1 Fifth   20.0 1.1E+02  0.0023   23.1   2.8   34   68-103    13-46  (73)
 98 PTZ00208 65 kDa invariant surf  20.0      48   0.001   34.1   1.0   30  410-439   384-414 (436)

No 1  
>PF01453 B_lectin:  D-mannose binding lectin;  InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]:  Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein   This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity.  Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=99.95  E-value=1.5e-28  Score=209.35  Aligned_cols=102  Identities=46%  Similarity=0.740  Sum_probs=74.1

Q ss_pred             CCcEEEEecCCCCCCC--CccEEEEecCCcEEEEeCCCCceEEeecCC--C--CccEEEEecCCCeEEEecCCeeEEeec
Q 043869           70 EKTVVWTANRDNPPVS--SNATLMFNSEGRIVLRSGEQGQNSIIADNS--Q--SASSASMLDSGSFVLHNSDGKVIWQTF  143 (450)
Q Consensus        70 ~~tvVW~ANr~~Pv~~--~~~~L~l~~~G~LvL~d~~~g~~vW~st~~--~--~~~~a~LldsGNlVL~~~~~~~lWQSF  143 (450)
                      ++||||+|||+.|+.+  ...+|.|+.||+|+|.+. .++.+|.+..+  .  .+..|.|+|+|||||+|..+.+|||||
T Consensus         1 ~~tvvW~an~~~p~~~~s~~~~L~l~~dGnLvl~~~-~~~~iWss~~t~~~~~~~~~~~L~~~GNlvl~d~~~~~lW~Sf   79 (114)
T PF01453_consen    1 PRTVVWVANRNSPLTSSSGNYTLILQSDGNLVLYDS-NGSVIWSSNNTSGRGNSGCYLVLQDDGNLVLYDSSGNVLWQSF   79 (114)
T ss_dssp             ---------TTEEEEECETTEEEEEETTSEEEEEET-TTEEEEE--S-TTSS-SSEEEEEETTSEEEEEETTSEEEEEST
T ss_pred             CcccccccccccccccccccccceECCCCeEEEEcC-CCCEEEEecccCCccccCeEEEEeCCCCEEEEeecceEEEeec
Confidence            3699999999999953  348999999999999998 88899944244  2  478999999999999999999999999


Q ss_pred             CCCCCccCCCcccCC----C--CeEEeccCCCCCC
Q 043869          144 DHPTDTLLPTQRLSA----G--TELCSGISETDPS  172 (450)
Q Consensus       144 d~PTDTlLpgq~L~~----~--~~L~S~~s~~dps  172 (450)
                      ||||||+||+|+|+.    +  ..|+||++.+|||
T Consensus        80 ~~ptdt~L~~q~l~~~~~~~~~~~~~sw~s~~dps  114 (114)
T PF01453_consen   80 DYPTDTLLPGQKLGDGNVTGKNDSLTSWSSNTDPS  114 (114)
T ss_dssp             TSSS-EEEEEET--TSEEEEESTSSEEEESS----
T ss_pred             CCCccEEEeccCcccCCCccccceEEeECCCCCCC
Confidence            999999999999987    3  3599999999996


No 2  
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=99.92  E-value=1.1e-24  Score=186.23  Aligned_cols=107  Identities=37%  Similarity=0.669  Sum_probs=93.2

Q ss_pred             CCeEEeCCCeEEEEEEeCCCCCeeEEEEEEeecCCCcEEEEecCCCCCCCCccEEEEecCCcEEEEeCCCCceEEeecCC
Q 043869           36 NSSWRSPSGLYAFGFYPQRNGSRYYVGVFLAGIPEKTVVWTANRDNPPVSSNATLMFNSEGRIVLRSGEQGQNSIIADNS  115 (450)
Q Consensus        36 ~~~l~S~~g~F~lGF~~~~~~~~~~lgIw~~~~~~~tvVW~ANr~~Pv~~~~~~L~l~~~G~LvL~d~~~g~~vW~st~~  115 (450)
                      +++|+|+++.|++|||.......++++|||.+.+ +++||.||++.|.. ..++|.|+.||+|+|.|. +|.++| ++++
T Consensus         7 ~~~l~s~~~~f~~G~~~~~~q~~dgnlv~~~~~~-~~~vW~snt~~~~~-~~~~l~l~~dGnLvl~~~-~g~~vW-~S~~   82 (116)
T cd00028           7 GQTLVSSGSLFELGFFKLIMQSRDYNLILYKGSS-RTVVWVANRDNPSG-SSCTLTLQSDGNLVIYDG-SGTVVW-SSNT   82 (116)
T ss_pred             CCEEEeCCCcEEEecccCCCCCCeEEEEEEeCCC-CeEEEECCCCCCCC-CCEEEEEecCCCeEEEcC-CCcEEE-Eecc
Confidence            3899999999999999854332389999998876 78999999999844 678999999999999998 899999 5554


Q ss_pred             ---CCccEEEEecCCCeEEEecCCeeEEeecCCC
Q 043869          116 ---QSASSASMLDSGSFVLHNSDGKVIWQTFDHP  146 (450)
Q Consensus       116 ---~~~~~a~LldsGNlVL~~~~~~~lWQSFd~P  146 (450)
                         ....+|.|+|+|||||++.++.+||||||||
T Consensus        83 ~~~~~~~~~~L~ddGnlvl~~~~~~~~W~Sf~~P  116 (116)
T cd00028          83 TRVNGNYVLVLLDDGNLVLYDSDGNFLWQSFDYP  116 (116)
T ss_pred             cCCCCceEEEEeCCCCEEEECCCCCEEEcCCCCC
Confidence               3467899999999999999999999999999


No 3  
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=99.90  E-value=2.1e-23  Score=177.86  Aligned_cols=106  Identities=32%  Similarity=0.600  Sum_probs=91.6

Q ss_pred             CCeEEeCCCeEEEEEEeCCCCCeeEEEEEEeecCCCcEEEEecCCCCCCCCccEEEEecCCcEEEEeCCCCceEEeecCC
Q 043869           36 NSSWRSPSGLYAFGFYPQRNGSRYYVGVFLAGIPEKTVVWTANRDNPPVSSNATLMFNSEGRIVLRSGEQGQNSIIADNS  115 (450)
Q Consensus        36 ~~~l~S~~g~F~lGF~~~~~~~~~~lgIw~~~~~~~tvVW~ANr~~Pv~~~~~~L~l~~~G~LvL~d~~~g~~vW~st~~  115 (450)
                      ++.|+|+++.|++|||..... .++++|||...+ +++||+|||+.|+. .+++|.|++||+|+|.+. +|.++|++...
T Consensus         7 ~~~l~s~~~~f~~G~~~~~~q-~dgnlV~~~~~~-~~~vW~snt~~~~~-~~~~l~l~~dGnLvl~~~-~g~~vW~S~t~   82 (114)
T smart00108        7 GQTLVSGNSLFELGFFTLIMQ-NDYNLILYKSSS-RTVVWVANRDNPVS-DSCTLTLQSDGNLVLYDG-DGRVVWSSNTT   82 (114)
T ss_pred             CCEEecCCCcEeeeccccCCC-CCEEEEEEECCC-CcEEEECCCCCCCC-CCEEEEEeCCCCEEEEeC-CCCEEEEeccc
Confidence            389999999999999985433 488999999876 78999999999976 458999999999999998 89999944332


Q ss_pred             --CCccEEEEecCCCeEEEecCCeeEEeecCC
Q 043869          116 --QSASSASMLDSGSFVLHNSDGKVIWQTFDH  145 (450)
Q Consensus       116 --~~~~~a~LldsGNlVL~~~~~~~lWQSFd~  145 (450)
                        ....+|+|+|+|||||++.++.++||||||
T Consensus        83 ~~~~~~~~~L~ddGnlvl~~~~~~~~W~Sf~~  114 (114)
T smart00108       83 GANGNYVLVLLDDGNLVIYDSDGNFLWQSFDY  114 (114)
T ss_pred             CCCCceEEEEeCCCCEEEECCCCCEEeCCCCC
Confidence              345689999999999999999999999997


No 4  
>PF00954 S_locus_glycop:  S-locus glycoprotein family;  InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=99.80  E-value=3.6e-19  Score=150.74  Aligned_cols=70  Identities=33%  Similarity=0.835  Sum_probs=63.7

Q ss_pred             CCCceEEEEEEccCCcEEEEEecccCCCCceEEEccccCCCCCCcCCCCCCcccccCCCCCCCcCCCCCeecc
Q 043869          226 PIQGMMYLMKIDSDGIFRLYSYNLRWQNSTWSEVWPSTSEKCDPIGLCGFNSFCVLNDQTPNCTCLPGFVAIS  298 (450)
Q Consensus       226 ~~~~~~~r~~Ld~dG~lr~y~~~~~~~~~~W~~~w~~p~~~C~~~g~CG~~g~C~~~~~~~~C~C~~GF~~~~  298 (450)
                      .+...+.|++||.||++++|.|.  +..+.|..+|++|.++||+|+.||+||+|+. .+.+.|+||+||+|++
T Consensus        41 ~~~s~~~r~~ld~~G~l~~~~w~--~~~~~W~~~~~~p~d~Cd~y~~CG~~g~C~~-~~~~~C~Cl~GF~P~n  110 (110)
T PF00954_consen   41 SNSSVLSRLVLDSDGQLQRYIWN--ESTQSWSVFWSAPKDQCDVYGFCGPNGICNS-NNSPKCSCLPGFEPKN  110 (110)
T ss_pred             CCCceEEEEEEeeeeEEEEEEEe--cCCCcEEEEEEecccCCCCccccCCccEeCC-CCCCceECCCCcCCCc
Confidence            35567899999999999999998  7788999999999999999999999999987 4578999999999974


No 5  
>PF08276 PAN_2:  PAN-like domain;  InterPro: IPR013227 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs
Probab=99.51  E-value=1.8e-14  Score=110.41  Aligned_cols=58  Identities=26%  Similarity=0.496  Sum_probs=50.6

Q ss_pred             CcceEEeccccccCccccee-cccCHHHHHHHHhcCCCeEEEEec----CCceEeeeccccce
Q 043869          317 NKAIQELENTNWEDVSYNVL-SEITKEKCKQACLEDCNCEAALYK----NEECKMQRLPLRFG  374 (450)
Q Consensus       317 ~~~f~~l~~v~~p~~~~~~~-~~~s~~~C~~~CL~nCsC~A~~y~----~g~C~~~~~~L~~~  374 (450)
                      +++|++|++|++|+++...+ .++++++|+++||+||||+||+|.    +++|++|.++|+|.
T Consensus         4 ~d~F~~l~~~~~p~~~~~~~~~~~s~~~C~~~Cl~nCsC~Ayay~~~~~~~~C~lW~~~L~d~   66 (66)
T PF08276_consen    4 GDGFLKLPNMKLPDFDNAIVDSSVSLEECEKACLSNCSCTAYAYSNLSGGGGCLLWYGDLVDL   66 (66)
T ss_pred             CCEEEEECCeeCCCCcceeeecCCCHHHHHhhcCCCCCEeeEEeeccCCCCEEEEEcCEeecC
Confidence            38999999999999876654 568999999999999999999995    26799999998873


No 6  
>cd01098 PAN_AP_plant Plant PAN/APPLE-like domain; present in plant S-receptor protein kinases and secreted glycoproteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions. S-receptor protein kinases and S-locus glycoproteins are involved in sporophytic self-incompatibility response in Brassica, one of probably many molecular mechanisms, by which hermaphrodite flowering plants avoid self-fertilization.
Probab=99.49  E-value=6.2e-14  Score=112.24  Aligned_cols=71  Identities=20%  Similarity=0.447  Sum_probs=61.0

Q ss_pred             cceEEeccccccCcccceecccCHHHHHHHHhcCCCeEEEEec--CCceEeeeccccceeeeCCCCceEEEEec
Q 043869          318 KAIQELENTNWEDVSYNVLSEITKEKCKQACLEDCNCEAALYK--NEECKMQRLPLRFGKRNLRDSDITFVKVD  389 (450)
Q Consensus       318 ~~f~~l~~v~~p~~~~~~~~~~s~~~C~~~CL~nCsC~A~~y~--~g~C~~~~~~L~~~~~~~~~~~~~yiKv~  389 (450)
                      ++|++++++++|+..+.. ...++++|+++||+||+|+||+|.  +++|++|..++.+.+.....+.++||||+
T Consensus        12 ~~f~~~~~~~~~~~~~~~-~~~s~~~C~~~Cl~nCsC~a~~~~~~~~~C~~~~~~~~~~~~~~~~~~~~yiKv~   84 (84)
T cd01098          12 DGFLKLPDVKLPDNASAI-TAISLEECREACLSNCSCTAYAYNNGSGGCLLWNGLLNNLRSLSSGGGTLYLRLA   84 (84)
T ss_pred             CEEEEeCCeeCCCchhhh-ccCCHHHHHHHHhcCCCcceeeecCCCCeEEEEeceecceEeecCCCcEEEEEeC
Confidence            689999999999876654 667999999999999999999994  47899999999987765444589999985


No 7  
>cd00129 PAN_APPLE PAN/APPLE-like domain; present in N-terminal (N) domains of plasminogen/ hepatocyte growth factor proteins,  plasma prekallikrein/coagulation factor XI and microneme antigen proteins, plant receptor-like protein kinases, and various nematode and leech anti-platelet proteins. Common structural features include two disulfide bonds that link the alpha-helix to the central region of the protein. PAN domains have significant functional versatility, fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=99.27  E-value=6.1e-12  Score=99.83  Aligned_cols=66  Identities=9%  Similarity=0.150  Sum_probs=56.9

Q ss_pred             cceEEeccccccCcccceecccCHHHHHHHHhc---CCCeEEEEec--CCceEeeeccc-cceeeeCCCCceEEEEe
Q 043869          318 KAIQELENTNWEDVSYNVLSEITKEKCKQACLE---DCNCEAALYK--NEECKMQRLPL-RFGKRNLRDSDITFVKV  388 (450)
Q Consensus       318 ~~f~~l~~v~~p~~~~~~~~~~s~~~C~~~CL~---nCsC~A~~y~--~g~C~~~~~~L-~~~~~~~~~~~~~yiKv  388 (450)
                      ..|+++.++++|+...     .+++||++.|++   ||||.||+|.  +++|++|.++| .+.+....++.++|+|.
T Consensus         9 g~fl~~~~~klpd~~~-----~s~~eC~~~Cl~~~~nCsC~Aya~~~~~~gC~~W~~~l~~d~~~~~~~g~~Ly~r~   80 (80)
T cd00129           9 GTTLIKIALKIKTTKA-----NTADECANRCEKNGLPFSCKAFVFAKARKQCLWFPFNSMSGVRKEFSHGFDLYENK   80 (80)
T ss_pred             CeEEEeecccCCcccc-----cCHHHHHHHHhcCCCCCCceeeeccCCCCCeEEecCcchhhHHhccCCCceeEeEC
Confidence            6799999999998644     689999999999   9999999994  35899999999 88887766569999983


No 8  
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=98.73  E-value=5.4e-08  Score=82.75  Aligned_cols=73  Identities=26%  Similarity=0.536  Sum_probs=56.4

Q ss_pred             EEEecCCcEEEEeCCC-CceEEeecCC----CCccEEEEecCCCeEEEecCCeeEEeecCCCCCccCCCcccCCCCeEEe
Q 043869           90 LMFNSEGRIVLRSGEQ-GQNSIIADNS----QSASSASMLDSGSFVLHNSDGKVIWQTFDHPTDTLLPTQRLSAGTELCS  164 (450)
Q Consensus        90 L~l~~~G~LvL~d~~~-g~~vW~st~~----~~~~~a~LldsGNlVL~~~~~~~lWQSFd~PTDTlLpgq~L~~~~~L~S  164 (450)
                      +.+..||+||+.+. . +.++| ++++    .....+.|+++|||||++.++.++|+|=     |-              
T Consensus        24 ~~~q~dgnlV~~~~-~~~~~vW-~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S~-----t~--------------   82 (114)
T smart00108       24 LIMQNDYNLILYKS-SSRTVVW-VANRDNPVSDSCTLTLQSDGNLVLYDGDGRVVWSSN-----TT--------------   82 (114)
T ss_pred             cCCCCCEEEEEEEC-CCCcEEE-ECCCCCCCCCCEEEEEeCCCCEEEEeCCCCEEEEec-----cc--------------
Confidence            44567999999986 4 47999 6654    1236789999999999999899999971     10              


Q ss_pred             ccCCCCCCCCceEEEecCCCceeEc
Q 043869          165 GISETDPSTGKFRLKMQNDGNLVQY  189 (450)
Q Consensus       165 ~~s~~dps~G~f~l~~~~~g~~~l~  189 (450)
                            ...|.+.+.|+++|++++|
T Consensus        83 ------~~~~~~~~~L~ddGnlvl~  101 (114)
T smart00108       83 ------GANGNYVLVLLDDGNLVIY  101 (114)
T ss_pred             ------CCCCceEEEEeCCCCEEEE
Confidence                  1356688999999999998


No 9  
>smart00473 PAN_AP divergent subfamily of APPLE domains. Apple-like domains present in Plasminogen, C. elegans hypothetical ORFs and the extracellular portion of plant receptor-like protein kinases. Predicted to possess protein- and/or carbohydrate-binding functions.
Probab=98.67  E-value=7.4e-08  Score=75.28  Aligned_cols=70  Identities=23%  Similarity=0.378  Sum_probs=55.3

Q ss_pred             cceEEeccccccCcccceecccCHHHHHHHHhc-CCCeEEEEec--CCceEeee-ccccceeeeCCCCceEEEE
Q 043869          318 KAIQELENTNWEDVSYNVLSEITKEKCKQACLE-DCNCEAALYK--NEECKMQR-LPLRFGKRNLRDSDITFVK  387 (450)
Q Consensus       318 ~~f~~l~~v~~p~~~~~~~~~~s~~~C~~~CL~-nCsC~A~~y~--~g~C~~~~-~~L~~~~~~~~~~~~~yiK  387 (450)
                      ..|++++++.+++.........++++|++.|++ +|+|.|+.|.  ++.|++|. .++.+.+.....+.++|.|
T Consensus         4 ~~f~~~~~~~l~~~~~~~~~~~s~~~C~~~C~~~~~~C~s~~y~~~~~~C~l~~~~~~~~~~~~~~~~~~~y~~   77 (78)
T smart00473        4 DCFVRLPNTKLPGFSRIVISVASLEECASKCLNSNCSCRSFTYNNGTKGCLLWSESSLGDARLFPSGGVDLYEK   77 (78)
T ss_pred             ceeEEecCccCCCCcceeEcCCCHHHHHHHhCCCCCceEEEEEcCCCCEEEEeeCCccccceecccCCceeEEe
Confidence            568999999998654433456799999999999 9999999994  57899998 7777776444444677776


No 10 
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=98.65  E-value=8.8e-08  Score=81.71  Aligned_cols=72  Identities=25%  Similarity=0.513  Sum_probs=55.9

Q ss_pred             EEec-CCcEEEEeCCC-CceEEeecCC----CCccEEEEecCCCeEEEecCCeeEEeecCCCCCccCCCcccCCCCeEEe
Q 043869           91 MFNS-EGRIVLRSGEQ-GQNSIIADNS----QSASSASMLDSGSFVLHNSDGKVIWQTFDHPTDTLLPTQRLSAGTELCS  164 (450)
Q Consensus        91 ~l~~-~G~LvL~d~~~-g~~vW~st~~----~~~~~a~LldsGNlVL~~~~~~~lWQSFd~PTDTlLpgq~L~~~~~L~S  164 (450)
                      .... ||+|++.+. . +.++| ++++    .....+.|+++|||||+|.++.++|+|=..                   
T Consensus        25 ~~q~~dgnlv~~~~-~~~~~vW-~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S~~~-------------------   83 (116)
T cd00028          25 IMQSRDYNLILYKG-SSRTVVW-VANRDNPSGSSCTLTLQSDGNLVIYDGSGTVVWSSNTT-------------------   83 (116)
T ss_pred             CCCCCeEEEEEEeC-CCCeEEE-ECCCCCCCCCCEEEEEecCCCeEEEcCCCcEEEEeccc-------------------
Confidence            3454 899999976 4 47899 6654    245678999999999999999999996211                   


Q ss_pred             ccCCCCCCCCceEEEecCCCceeEc
Q 043869          165 GISETDPSTGKFRLKMQNDGNLVQY  189 (450)
Q Consensus       165 ~~s~~dps~G~f~l~~~~~g~~~l~  189 (450)
                            ...+.+.+.|++||++++|
T Consensus        84 ------~~~~~~~~~L~ddGnlvl~  102 (116)
T cd00028          84 ------RVNGNYVLVLLDDGNLVLY  102 (116)
T ss_pred             ------CCCCceEEEEeCCCCEEEE
Confidence                  0256788999999999998


No 11 
>PF01453 B_lectin:  D-mannose binding lectin;  InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]:  Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein   This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity.  Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=98.46  E-value=1.8e-06  Score=73.44  Aligned_cols=100  Identities=20%  Similarity=0.328  Sum_probs=69.1

Q ss_pred             CeEEeCCCeEEEEEEeCCCCCeeEEEEEEeecCCCcEEEEe-cCCCCCCCCccEEEEecCCcEEEEeCCCCceEEeecCC
Q 043869           37 SSWRSPSGLYAFGFYPQRNGSRYYVGVFLAGIPEKTVVWTA-NRDNPPVSSNATLMFNSEGRIVLRSGEQGQNSIIADNS  115 (450)
Q Consensus        37 ~~l~S~~g~F~lGF~~~~~~~~~~lgIw~~~~~~~tvVW~A-Nr~~Pv~~~~~~L~l~~~G~LvL~d~~~g~~vW~st~~  115 (450)
                      +.+.+.+|.+.|-|..+|+-      +.|..  ..+++|.. +...... ..+.+.|.++|||||+|. .+.++|+|...
T Consensus        12 ~p~~~~s~~~~L~l~~dGnL------vl~~~--~~~~iWss~~t~~~~~-~~~~~~L~~~GNlvl~d~-~~~~lW~Sf~~   81 (114)
T PF01453_consen   12 SPLTSSSGNYTLILQSDGNL------VLYDS--NGSVIWSSNNTSGRGN-SGCYLVLQDDGNLVLYDS-SGNVLWQSFDY   81 (114)
T ss_dssp             EEEEECETTEEEEEETTSEE------EEEET--TTEEEEE--S-TTSS--SSEEEEEETTSEEEEEET-TSEEEEESTTS
T ss_pred             cccccccccccceECCCCeE------EEEcC--CCCEEEEecccCCccc-cCeEEEEeCCCCEEEEee-cceEEEeecCC
Confidence            45656558999999987764      44443  45789999 4443332 478899999999999998 89999955433


Q ss_pred             CCccEEEEec--CCCeEEEecCCeeEEeecCCCC
Q 043869          116 QSASSASMLD--SGSFVLHNSDGKVIWQTFDHPT  147 (450)
Q Consensus       116 ~~~~~a~Lld--sGNlVL~~~~~~~lWQSFd~PT  147 (450)
                      ...+.+..++  .||++ +.....+.|.|=+.|+
T Consensus        82 ptdt~L~~q~l~~~~~~-~~~~~~~sw~s~~dps  114 (114)
T PF01453_consen   82 PTDTLLPGQKLGDGNVT-GKNDSLTSWSSNTDPS  114 (114)
T ss_dssp             SS-EEEEEET--TSEEE-EESTSSEEEESS----
T ss_pred             CccEEEeccCcccCCCc-cccceEEeECCCCCCC
Confidence            5566677777  88888 7666778999876663


No 12 
>cd01100 APPLE_Factor_XI_like Subfamily of PAN/APPLE-like domains; present in plasma prekallikrein/coagulation factor XI, microneme antigen proteins, and a few prokaryotic proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=97.92  E-value=1.1e-05  Score=62.88  Aligned_cols=48  Identities=21%  Similarity=0.460  Sum_probs=39.3

Q ss_pred             EeccccccCcccceecccCHHHHHHHHhcCCCeEEEEec--CCceEeeec
Q 043869          322 ELENTNWEDVSYNVLSEITKEKCKQACLEDCNCEAALYK--NEECKMQRL  369 (450)
Q Consensus       322 ~l~~v~~p~~~~~~~~~~s~~~C~~~CL~nCsC~A~~y~--~g~C~~~~~  369 (450)
                      .++++++++.+.......+.++|++.|+.+|+|.||.|.  .+.|+++..
T Consensus         8 ~~~~~~~~g~d~~~~~~~s~~~Cq~~C~~~~~C~afT~~~~~~~C~lk~~   57 (73)
T cd01100           8 QGSNVDFRGGDLSTVFASSAEQCQAACTADPGCLAFTYNTKSKKCFLKSS   57 (73)
T ss_pred             ccCCCccccCCcceeecCCHHHHHHHcCCCCCceEEEEECCCCeEEcccC
Confidence            346888888777655566899999999999999999993  478998865


No 13 
>smart00223 APPLE APPLE domain. Four-fold repeat in plasma kallikrein and coagulation factor XI. Factor XI apple 3 mediates binding to platelets. Factor XI apple 1 binds high-molecular-mass kininogen. Apple 4 in factor XI mediates dimer formation and binds to factor XIIa. Mutations in apple 4 cause factor XI deficiency, an inherited bleeding disorder.
Probab=95.22  E-value=0.026  Score=44.65  Aligned_cols=47  Identities=13%  Similarity=0.383  Sum_probs=40.1

Q ss_pred             eccccccCcccceecccCHHHHHHHHhcCCCeEEEEe--cCC---ceEeeec
Q 043869          323 LENTNWEDVSYNVLSEITKEKCKQACLEDCNCEAALY--KNE---ECKMQRL  369 (450)
Q Consensus       323 l~~v~~p~~~~~~~~~~s~~~C~~~CL~nCsC~A~~y--~~g---~C~~~~~  369 (450)
                      .+|++|++.+...+...+.++|++.|..+=.|.+|.|  .+.   .|+++..
T Consensus         6 ~~~~df~G~Dl~~~~~~~~~~Cq~~Ct~~~~C~~FTf~~~~~~~~~C~LK~s   57 (79)
T smart00223        6 YKNVDFRGSDINTVYVPSAQVCQKRCTSHPRCLFFTFSTNEPPEEKCLLKDS   57 (79)
T ss_pred             ccCccccCceeeeeecCCHHHHHHhhcCCCCccEEEeeCCCCCCCEeEeCcC
Confidence            4689999988887777899999999999999999999  345   7998754


No 14 
>PF00024 PAN_1:  PAN domain This Prosite entry concerns apple domains, a subset of PAN domains;  InterPro: IPR003014 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs It has been shown that, the N-terminal N domains of members of the plasminogen/hepatocyte growth factor family, the apple domains of the plasma prekallikrein/coagulation factor XI family, and domains of various nematode proteins belong to the same module superfamily, the PAN module []. PAN contains a conserved core of three disulphide bridges. In some members of the family there is an additional fourth disulphide bridge that links the N and C termini of the domain.; PDB: 1GP9_C 2QJ2_B 1GMO_H 1NK1_B 3MKP_B 1BHT_B 3HN4_A 1GMN_A 3HMS_A 3HMT_B ....
Probab=94.46  E-value=0.07  Score=41.20  Aligned_cols=51  Identities=20%  Similarity=0.429  Sum_probs=40.2

Q ss_pred             ceEEeccccccCcccceecccCHHHHHHHHhcCCC-eEEEEe--cCCceEeeec
Q 043869          319 AIQELENTNWEDVSYNVLSEITKEKCKQACLEDCN-CEAALY--KNEECKMQRL  369 (450)
Q Consensus       319 ~f~~l~~v~~p~~~~~~~~~~s~~~C~~~CL~nCs-C~A~~y--~~g~C~~~~~  369 (450)
                      .|..+++..+...........++++|.+.|+++=. |.++.|  ..+.|.+...
T Consensus         3 ~f~~~~~~~l~~~~~~~~~v~s~~~C~~~C~~~~~~C~s~~y~~~~~~C~L~~~   56 (79)
T PF00024_consen    3 AFERIPGYRLSGHSIKEINVPSLEECAQLCLNEPRRCKSFNYDPSSKTCYLSSS   56 (79)
T ss_dssp             TEEEEEEEEEESCEEEEEEESSHHHHHHHHHHSTT-ESEEEEETTTTEEEEECS
T ss_pred             CeEEECCEEEeCCcceEEcCCCHHHHHhhcCcCcccCCeEEEECCCCEEEEcCC
Confidence            47778888877755554444599999999999999 999999  3468998754


No 15 
>PF14295 PAN_4:  PAN domain; PDB: 2YIL_E 2YIP_C 2YIO_A.
Probab=94.29  E-value=0.041  Score=38.95  Aligned_cols=41  Identities=20%  Similarity=0.549  Sum_probs=17.2

Q ss_pred             cccCcccce--ecccCHHHHHHHHhcCCCeEEEEecC-------CceEee
Q 043869          327 NWEDVSYNV--LSEITKEKCKQACLEDCNCEAALYKN-------EECKMQ  367 (450)
Q Consensus       327 ~~p~~~~~~--~~~~s~~~C~~~CL~nCsC~A~~y~~-------g~C~~~  367 (450)
                      ++++.++..  ....+.++|.++|.++=.|.++.|..       +.|++|
T Consensus         2 d~~G~dl~~~~~~~~s~~~C~~~C~~~~~C~~~~~~~~~~~~~~~~C~LK   51 (51)
T PF14295_consen    2 DYPGGDLRSFPVTASSPEECQAACAADPGCQAFTFNPPGCPSSSGRCYLK   51 (51)
T ss_dssp             ----------------HHHHHHHHHTSTT--EEEEETTEE----------
T ss_pred             cccccccccccccCCCHHHHHHHccCCCCCCEEEEECCCcccccccccCC
Confidence            455544443  24568999999999999999999932       457764


No 16 
>PF08277 PAN_3:  PAN-like domain;  InterPro: IPR006583 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs The PAN-3 or CW is a domain associated with a number of Caenorhabditis elegans hypothetical proteins.
Probab=90.91  E-value=1.1  Score=33.99  Aligned_cols=32  Identities=25%  Similarity=0.641  Sum_probs=27.9

Q ss_pred             cccCHHHHHHHHhcCCCeEEEEecCCceEeee
Q 043869          337 SEITKEKCKQACLEDCNCEAALYKNEECKMQR  368 (450)
Q Consensus       337 ~~~s~~~C~~~CL~nCsC~A~~y~~g~C~~~~  368 (450)
                      ...+.++|-+.|.++=+|.++.+.++.|.++.
T Consensus        18 ~~~sw~~Cv~~C~~~~~C~la~~~~~~C~~y~   49 (71)
T PF08277_consen   18 TNTSWDDCVQKCYNDENCVLAYFDSGKCYLYN   49 (71)
T ss_pred             cCCCHHHHhHHhCCCCEEEEEEeCCCCEEEEE
Confidence            45678999999999999999998777899865


No 17 
>smart00605 CW CW domain.
Probab=89.74  E-value=2.1  Score=34.67  Aligned_cols=55  Identities=24%  Similarity=0.372  Sum_probs=39.5

Q ss_pred             cccCHHHHHHHHhcCCCeEEEEecC-CceEeeec-cccceeee-CCCCceEEEEecCC
Q 043869          337 SEITKEKCKQACLEDCNCEAALYKN-EECKMQRL-PLRFGKRN-LRDSDITFVKVDDA  391 (450)
Q Consensus       337 ~~~s~~~C~~~CL~nCsC~A~~y~~-g~C~~~~~-~L~~~~~~-~~~~~~~yiKv~~~  391 (450)
                      ...+-++|.+.|.++..|..+...+ ..|.+... .+...++. ...+..+=||+..+
T Consensus        20 ~~~sw~~Ci~~C~~~~~Cvlay~~~~~~C~~f~~~~~~~v~~~~~~~~~~VAfK~~~~   77 (94)
T smart00605       20 ATLSWDECIQKCYEDSNCVLAYGNSSETCYLFSYGTVLTVKKLSSSSGKKVAFKVSTD   77 (94)
T ss_pred             cCCCHHHHHHHHhCCCceEEEecCCCCceEEEEcCCeEEEEEccCCCCcEEEEEEeCC
Confidence            3467899999999999999887643 78987643 34555554 33446788888754


No 18 
>PF04478 Mid2:  Mid2 like cell wall stress sensor;  InterPro: IPR007567 This family represents a region near the C terminus of Mid2, which contains a transmembrane region. The remainder of the protein sequence is serine-rich and of low complexity, and is therefore impossible to align accurately. Mid2 is thought to act as a mechanosensor of cell wall stress. The C-terminal cytoplasmic region of Mid2 is known to interact with Rom2, a guanine nucleotide exchange factor (GEF) for Rho1, which is part of the cell wall integrity signalling pathway [].
Probab=88.98  E-value=0.45  Score=42.08  Aligned_cols=38  Identities=16%  Similarity=0.194  Sum_probs=15.2

Q ss_pred             EEEEehhhHHHHHHHHHHhheeeeEeeccceeecccCCC
Q 043869          412 NIVIICLFVTVVILISVVTFGIFIYRYRVGSYRRIQGNG  450 (450)
Q Consensus       412 ~i~i~~~~~~~~~l~~~~~~~~~~~r~~~~~~~~~~~~~  450 (450)
                      ++|.+.+-+|+.+|+.+++.+|++++|+ +|..=|.++|
T Consensus        50 IVIGvVVGVGg~ill~il~lvf~~c~r~-kktdfidSdG   87 (154)
T PF04478_consen   50 IVIGVVVGVGGPILLGILALVFIFCIRR-KKTDFIDSDG   87 (154)
T ss_pred             EEEEEEecccHHHHHHHHHhheeEEEec-ccCccccCCC
Confidence            4444433334444444433334333333 2334454443


No 19 
>PF08693 SKG6:  Transmembrane alpha-helix domain;  InterPro: IPR014805 SKG6 and AXL2 are membrane proteins that show polarised intracellular localisation [, ]. This entry represents the highly conserved transmembrane alpha-helical domain found in these proteins [, ]. The full-length AXL2 protein has a negative regulatory function in cytokinesis [].
Probab=88.67  E-value=0.45  Score=32.33  Aligned_cols=31  Identities=23%  Similarity=0.285  Sum_probs=16.0

Q ss_pred             cceEEEEehhhHHHHHHHHHHhheee-eEeec
Q 043869          409 LWKNIVIICLFVTVVILISVVTFGIF-IYRYR  439 (450)
Q Consensus       409 ~~~~i~i~~~~~~~~~l~~~~~~~~~-~~r~~  439 (450)
                      ++..-+.+++++.+..++++++.++| +|||+
T Consensus         8 ~~~vaIa~~VvVPV~vI~~vl~~~l~~~~rR~   39 (40)
T PF08693_consen    8 SNTVAIAVGVVVPVGVIIIVLGAFLFFWYRRK   39 (40)
T ss_pred             CceEEEEEEEEechHHHHHHHHHHhheEEecc
Confidence            34455666666665555444433344 45544


No 20 
>cd01099 PAN_AP_HGF Subfamily of PAN/APPLE-like domains; present in N-terminal (N) domains of plasminogen/hepatocyte growth factor proteins, and various proteins found in Bilateria, such as leech anti-platelet proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=87.89  E-value=1.9  Score=33.86  Aligned_cols=33  Identities=30%  Similarity=0.694  Sum_probs=27.3

Q ss_pred             cccCHHHHHHHHhc--CCCeEEEEe--cCCceEeeec
Q 043869          337 SEITKEKCKQACLE--DCNCEAALY--KNEECKMQRL  369 (450)
Q Consensus       337 ~~~s~~~C~~~CL~--nCsC~A~~y--~~g~C~~~~~  369 (450)
                      ...++++|.+.|++  +=.|.++.|  .++.|.+-..
T Consensus        23 ~~~s~~~C~~~C~~~~~f~CrSf~y~~~~~~C~L~~~   59 (80)
T cd01099          23 TVASLEECLRKCLEETEFTCRSFNYNYKSKECILSDE   59 (80)
T ss_pred             ecCCHHHHHHHhCCCCCceEeEEEEEcCCCEEEEeCC
Confidence            34789999999999  889999998  4678987543


No 21 
>PF07645 EGF_CA:  Calcium-binding EGF domain;  InterPro: IPR001881 A sequence of about forty amino-acid residues found in epidermal growth factor (EGF) has been shown [, , , , , ] to be present in a large number of membrane-bound and extracellular, mostly animal, proteins. Many of these proteins require calcium for their biological function and a calcium-binding site has been found at the N terminus of some EGF-like domains []. Calcium-binding may be crucial for numerous protein-protein interactions. For human coagulation factor IX it has been shown [] that the calcium-ligands form a pentagonal bipyramid. The first, third and fourth conserved negatively charged or polar residues are side chain ligands. The latter is possibly hydroxylated (see aspartic acid and asparagine hydroxylation site) []. A conserved aromatic residue, as well as the second conserved negative residue, are thought to be involved in stabilising the calcium-binding site. As in non-calcium binding EGF-like domains, there are six conserved cysteines and the structure of both types is very similar as calcium-binding induces only strictly local structural changes [].  +------------------+ +---------+ | | | | nxnnC-x(3,14)-C-x(3,7)-CxxbxxxxaxC-x(1,6)-C-x(8,13)-Cx | | +------------------+ 'n': negatively charged or polar residue [DEQN] 'b': possibly beta-hydroxylated residue [DN] 'a': aromatic amino acid 'C': cysteine, involved in disulphide bond 'x': any amino acid. ; GO: 0005509 calcium ion binding; PDB: 2VJ3_A 1TOZ_A 1LMJ_A 1UZQ_A 1UZK_A 1UZJ_B 1UZP_A 1EMO_A 1EMN_A 2RR0_A ....
Probab=87.59  E-value=0.24  Score=33.99  Aligned_cols=32  Identities=28%  Similarity=0.682  Sum_probs=25.3

Q ss_pred             CCCCCc-CCCCCCcccccCCCCCCCcCCCCCee
Q 043869          265 EKCDPI-GLCGFNSFCVLNDQTPNCTCLPGFVA  296 (450)
Q Consensus       265 ~~C~~~-g~CG~~g~C~~~~~~~~C~C~~GF~~  296 (450)
                      |+|... ..|..++.|......-.|.|++||+.
T Consensus         3 dEC~~~~~~C~~~~~C~N~~Gsy~C~C~~Gy~~   35 (42)
T PF07645_consen    3 DECAEGPHNCPENGTCVNTEGSYSCSCPPGYEL   35 (42)
T ss_dssp             STTTTTSSSSSTTSEEEEETTEEEEEESTTEEE
T ss_pred             cccCCCCCcCCCCCEEEcCCCCEEeeCCCCcEE
Confidence            567764 47999999986444678999999994


No 22 
>cd00053 EGF Epidermal growth factor domain, found in epidermal growth factor (EGF) presents in a large number of proteins, mostly animal; the list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied; the functional significance of EGF-like domains in what appear to be unrelated proteins is not yet clear; a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase); the domain includes six cysteine residues which have been shown to be involved in disulfide bonds; the main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet; Subdomains between the conserved cysteines vary in length; the region between the 5th and 6th cysteine contains two conserved glycines of which at  least  one  is  present  in  most EGF-like domains; a subset of these bind calcium.
Probab=87.34  E-value=0.45  Score=30.25  Aligned_cols=30  Identities=27%  Similarity=0.723  Sum_probs=22.7

Q ss_pred             CCCcCCCCCCcccccCCCCCCCcCCCCCee
Q 043869          267 CDPIGLCGFNSFCVLNDQTPNCTCLPGFVA  296 (450)
Q Consensus       267 C~~~g~CG~~g~C~~~~~~~~C~C~~GF~~  296 (450)
                      |.....|..++.|........|.|++||..
T Consensus         2 C~~~~~C~~~~~C~~~~~~~~C~C~~g~~g   31 (36)
T cd00053           2 CAASNPCSNGGTCVNTPGSYRCVCPPGYTG   31 (36)
T ss_pred             CCCCCCCCCCCEEecCCCCeEeECCCCCcc
Confidence            443467888999986444788999999964


No 23 
>PF15102 TMEM154:  TMEM154 protein family
Probab=86.35  E-value=0.92  Score=39.91  Aligned_cols=27  Identities=37%  Similarity=0.581  Sum_probs=12.1

Q ss_pred             EEehhhHHHHHHHHHHhheeeeEeecc
Q 043869          414 VIICLFVTVVILISVVTFGIFIYRYRV  440 (450)
Q Consensus       414 ~i~~~~~~~~~l~~~~~~~~~~~r~~~  440 (450)
                      +++.+++.+++|+++++.+++.+|||.
T Consensus        61 IlIP~VLLvlLLl~vV~lv~~~kRkr~   87 (146)
T PF15102_consen   61 ILIPLVLLVLLLLSVVCLVIYYKRKRT   87 (146)
T ss_pred             EeHHHHHHHHHHHHHHHheeEEeeccc
Confidence            333335555555554444443444443


No 24 
>smart00179 EGF_CA Calcium-binding EGF-like domain.
Probab=82.08  E-value=1.1  Score=29.28  Aligned_cols=31  Identities=26%  Similarity=0.669  Sum_probs=22.9

Q ss_pred             CCCCCcCCCCCCcccccCCCCCCCcCCCCCe
Q 043869          265 EKCDPIGLCGFNSFCVLNDQTPNCTCLPGFV  295 (450)
Q Consensus       265 ~~C~~~g~CG~~g~C~~~~~~~~C~C~~GF~  295 (450)
                      +.|.....|...+.|........|.|++||.
T Consensus         3 ~~C~~~~~C~~~~~C~~~~g~~~C~C~~g~~   33 (39)
T smart00179        3 DECASGNPCQNGGTCVNTVGSYRCECPPGYT   33 (39)
T ss_pred             ccCcCCCCcCCCCEeECCCCCeEeECCCCCc
Confidence            4565545788888998643356799999997


No 25 
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=80.63  E-value=1.9  Score=35.40  Aligned_cols=14  Identities=14%  Similarity=0.090  Sum_probs=8.2

Q ss_pred             HHHHhheeeeEeec
Q 043869          426 ISVVTFGIFIYRYR  439 (450)
Q Consensus       426 ~~~~~~~~~~~r~~  439 (450)
                      |+.++.|||++|||
T Consensus        82 lv~~l~w~f~~r~k   95 (96)
T PTZ00382         82 LVGFLCWWFVCRGK   95 (96)
T ss_pred             HHHHHhheeEEeec
Confidence            33455567777765


No 26 
>cd00054 EGF_CA Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.
Probab=78.41  E-value=1.8  Score=27.85  Aligned_cols=32  Identities=25%  Similarity=0.670  Sum_probs=22.7

Q ss_pred             CCCCCcCCCCCCcccccCCCCCCCcCCCCCee
Q 043869          265 EKCDPIGLCGFNSFCVLNDQTPNCTCLPGFVA  296 (450)
Q Consensus       265 ~~C~~~g~CG~~g~C~~~~~~~~C~C~~GF~~  296 (450)
                      +.|.....|...+.|........|.|++||.-
T Consensus         3 ~~C~~~~~C~~~~~C~~~~~~~~C~C~~g~~g   34 (38)
T cd00054           3 DECASGNPCQNGGTCVNTVGSYRCSCPPGYTG   34 (38)
T ss_pred             ccCCCCCCcCCCCEeECCCCCeEeECCCCCcC
Confidence            45654356877888976444567999999863


No 27 
>PF01683 EB:  EB module;  InterPro: IPR006149  The EB domain has no known function. It is found in several Caenorhabditis sp. and Drosophila sp. proteins. The domain contains 8 conserved cysteines that probably form four disulphide bridges and is found associated with kunitz domains IPR002223 from INTERPRO 
Probab=76.48  E-value=2.5  Score=30.23  Aligned_cols=33  Identities=33%  Similarity=0.755  Sum_probs=27.3

Q ss_pred             ccCCCCCCcCCCCCCcccccCCCCCCCcCCCCCeecc
Q 043869          262 STSEKCDPIGLCGFNSFCVLNDQTPNCTCLPGFVAIS  298 (450)
Q Consensus       262 ~p~~~C~~~g~CG~~g~C~~~~~~~~C~C~~GF~~~~  298 (450)
                      .|-+.|+...-|-.++.|..    ..|.|++||.+..
T Consensus        17 ~~g~~C~~~~qC~~~s~C~~----g~C~C~~g~~~~~   49 (52)
T PF01683_consen   17 QPGESCESDEQCIGGSVCVN----GRCQCPPGYVEVG   49 (52)
T ss_pred             CCCCCCCCcCCCCCcCEEcC----CEeECCCCCEecC
Confidence            35678999999999999953    6899999998753


No 28 
>PF02009 Rifin_STEVOR:  Rifin/stevor family;  InterPro: IPR002858 Malaria is still a major cause of mortality in many areas of the world. Plasmodium falciparum causes the most severe human form of the disease and is responsible for most fatalities. Severe cases of malaria can occur when the parasite invades and then proliferates within red blood cell erythrocytes. The parasite produces many variant antigenic proteins, encoded by multigene families, which are present on the surface of the infected erythrocyte and play important roles in virulence. A crucial survival mechanism for the malaria parasite is its ability to evade the immune response by switching these variant surface antigens. The high virulence of P. falciparum relative to other malarial parasites is in large part due to the fact that in this organism many of these surface antigens mediate the binding of infected erythrocytes to the vascular endothelium (cytoadherence) and non-infected erythrocytes (rosetting). This can lead to the accumulation of infected cells in the vasculature of a variety of organs, blocking the blood flow and reducing the oxygen supply. Clinical symptoms of severe infection can include fever, progressive anaemia, multi-organ dysfunction and coma. For more information see []. Several multicopy gene families have been described in Plasmodium falciparum, including the stevor family of subtelomeric open reading frames and the rif interspersed repetitive elements. Both families contain three predicted transmembrane segments. It has been proposed that stevor and rif are members of a larger superfamily that code for variant surface antigens [].
Probab=75.64  E-value=0.51  Score=46.84  Aligned_cols=32  Identities=22%  Similarity=0.515  Sum_probs=18.9

Q ss_pred             EehhhHHHHHHHHHHhheeeeEeec--cceeecc
Q 043869          415 IICLFVTVVILISVVTFGIFIYRYR--VGSYRRI  446 (450)
Q Consensus       415 i~~~~~~~~~l~~~~~~~~~~~r~~--~~~~~~~  446 (450)
                      |+++++.+++++.+++++|+|+|.|  ++..+|+
T Consensus       258 I~aSiiaIliIVLIMvIIYLILRYRRKKKmkKKl  291 (299)
T PF02009_consen  258 IIASIIAILIIVLIMVIIYLILRYRRKKKMKKKL  291 (299)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHHHHHHhhhhHHH
Confidence            4456666666666777778876533  3444444


No 29 
>PF12947 EGF_3:  EGF domain;  InterPro: IPR024731 This entry represents an EGF domain found in the the C terminus of malarial parasite merozoite surface protein 1 [], as well as other proteins.; PDB: 2NPR_A 1N1I_C 1B9W_A 1YO8_A 2RHP_A.
Probab=74.95  E-value=0.65  Score=30.86  Aligned_cols=27  Identities=33%  Similarity=0.700  Sum_probs=19.3

Q ss_pred             cCCCCCCcccccCCCCCCCcCCCCCee
Q 043869          270 IGLCGFNSFCVLNDQTPNCTCLPGFVA  296 (450)
Q Consensus       270 ~g~CG~~g~C~~~~~~~~C~C~~GF~~  296 (450)
                      .+-|.++..|+.......|.|.+||+-
T Consensus         5 ~~~C~~nA~C~~~~~~~~C~C~~Gy~G   31 (36)
T PF12947_consen    5 NGGCHPNATCTNTGGSYTCTCKPGYEG   31 (36)
T ss_dssp             GGGS-TTCEEEE-TTSEEEEE-CEEEC
T ss_pred             CCCCCCCcEeecCCCCEEeECCCCCcc
Confidence            356889999987555778999999963


No 30 
>PF01102 Glycophorin_A:  Glycophorin A;  InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=74.06  E-value=0.57  Score=40.15  Aligned_cols=28  Identities=21%  Similarity=0.279  Sum_probs=11.9

Q ss_pred             EEEehhhHHHHHHHHHHhheeeeEeeccce
Q 043869          413 IVIICLFVTVVILISVVTFGIFIYRYRVGS  442 (450)
Q Consensus       413 i~i~~~~~~~~~l~~~~~~~~~~~r~~~~~  442 (450)
                      .||++++.|++.+  |+++.|+++|+|++.
T Consensus        68 ~Ii~gv~aGvIg~--Illi~y~irR~~Kk~   95 (122)
T PF01102_consen   68 GIIFGVMAGVIGI--ILLISYCIRRLRKKS   95 (122)
T ss_dssp             HHHHHHHHHHHHH--HHHHHHHHHHHS---
T ss_pred             ehhHHHHHHHHHH--HHHHHHHHHHHhccC
Confidence            3444444444333  234456666655443


No 31 
>PF01299 Lamp:  Lysosome-associated membrane glycoprotein (Lamp);  InterPro: IPR002000 Lysosome-associated membrane glycoproteins (lamp) [] are integral membrane proteins, specific to lysosomes, and whose exact biological function is not yet clear. Structurally, the lamp proteins consist of two internally homologous lysosome-luminal domains separated by a proline-rich hinge region; at the C-terminal extremity there is a transmembrane region (TM) followed by a very short cytoplasmic tail (C). In each of the duplicated domains, there are two conserved disulphide bonds. This structure is schematically represented in the figure below.   +-----+ +-----+ +-----+ +-----+ | | | | | | | | xCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxxxCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxx +--------------------------++Hinge++--------------------------++TM++C+  In mammals, there are two closely related types of lamp: lamp-1 and lamp-2, which form major components of the lysosome membrane. In chicken lamp-1 is known as LEP100.  Also included in this entry is the macrophage protein CD68 (or macrosialin) [] is a heavily glycosylated integral membrane protein whose structure consists of a mucin-like domain followed by a proline-rich hinge; a single lamp-like domain; a transmembrane region and a short cytoplasmic tail.   Similar to CD68, mammalian lamp-3, which is expressed in lymphoid organs, dendritic cells and in lung, contains all the C-terminal regions but lacks the N-terminal lamp-like region []. In a lamp-family protein from nematodes [] only the part C-terminal to the hinge is conserved. ; GO: 0016020 membrane
Probab=73.13  E-value=2.5  Score=42.14  Aligned_cols=19  Identities=32%  Similarity=0.488  Sum_probs=13.5

Q ss_pred             HHhheeeeEeeccce-eecc
Q 043869          428 VVTFGIFIYRYRVGS-YRRI  446 (450)
Q Consensus       428 ~~~~~~~~~r~~~~~-~~~~  446 (450)
                      |++++|+|.|||.++ |+.|
T Consensus       287 ivLiaYli~Rrr~~~gYq~~  306 (306)
T PF01299_consen  287 IVLIAYLIGRRRSRAGYQSI  306 (306)
T ss_pred             HHHHhheeEecccccccccC
Confidence            455578888888766 7764


No 32 
>PF02439 Adeno_E3_CR2:  Adenovirus E3 region protein CR2;  InterPro: IPR003470 Early region 3 (E3) of human adenoviruses (Ads) codes for proteins that appear to control viral interactions with the host []. This region called CR1 (conserved region 1) [] is found three times in Human adenovirus 19 (a subgroup D adenovirus) 49 kDa protein in the E3 region. CR1 is also found in the 20.1 Kd protein of subgroup B adenoviruses. The function of this 80 amino acid region is unknown. This region is probably a divergent immunoglobulin domain.
Probab=70.81  E-value=0.97  Score=30.25  Aligned_cols=29  Identities=17%  Similarity=0.089  Sum_probs=13.0

Q ss_pred             EehhhHHHHHHHHHHhheeeeEeecccee
Q 043869          415 IICLFVTVVILISVVTFGIFIYRYRVGSY  443 (450)
Q Consensus       415 i~~~~~~~~~l~~~~~~~~~~~r~~~~~~  443 (450)
                      |+++|+..++++++++..|-..+||.+++
T Consensus         8 IIv~V~vg~~iiii~~~~YaCcykk~~~~   36 (38)
T PF02439_consen    8 IIVAVVVGMAIIIICMFYYACCYKKHRRQ   36 (38)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHHcccccc
Confidence            33344444455555555554434443333


No 33 
>PF06697 DUF1191:  Protein of unknown function (DUF1191);  InterPro: IPR010605 This family contains hypothetical plant proteins of unknown function.
Probab=69.10  E-value=3.6  Score=40.23  Aligned_cols=36  Identities=11%  Similarity=0.108  Sum_probs=19.4

Q ss_pred             EEEehhhHHHHHHHHHHhheee-eEeeccceeecccC
Q 043869          413 IVIICLFVTVVILISVVTFGIF-IYRYRVGSYRRIQG  448 (450)
Q Consensus       413 i~i~~~~~~~~~l~~~~~~~~~-~~r~~~~~~~~~~~  448 (450)
                      .+|.++++|+++|.++.+.+.. .+-+|++|-++|+.
T Consensus       214 ~iv~g~~~G~~~L~ll~~lv~~~vr~krk~k~~eMEr  250 (278)
T PF06697_consen  214 KIVVGVVGGVVLLGLLSLLVAMLVRYKRKKKIEEMER  250 (278)
T ss_pred             EEEEEehHHHHHHHHHHHHHHhhhhhhHHHHHHHHHH
Confidence            3455667777776655443333 33345555555543


No 34 
>PF00008 EGF:  EGF-like domain This is a sub-family of the Pfam entry This is a sub-family of the Pfam entry;  InterPro: IPR006209 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length.; GO: 0005515 protein binding; PDB: 1WHE_A 1CCF_A 1APO_A 1WHF_A 2VJ3_A 1TOZ_A 4D90_B 3CFW_A 1EDM_B 1IXA_A ....
Probab=68.96  E-value=1.1  Score=28.73  Aligned_cols=24  Identities=25%  Similarity=0.636  Sum_probs=18.9

Q ss_pred             CCCCCcccccCC-CCCCCcCCCCCe
Q 043869          272 LCGFNSFCVLND-QTPNCTCLPGFV  295 (450)
Q Consensus       272 ~CG~~g~C~~~~-~~~~C~C~~GF~  295 (450)
                      .|...|.|.... ....|.|++||.
T Consensus         5 ~C~n~g~C~~~~~~~y~C~C~~G~~   29 (32)
T PF00008_consen    5 PCQNGGTCIDLPGGGYTCECPPGYT   29 (32)
T ss_dssp             SSTTTEEEEEESTSEEEEEEBTTEE
T ss_pred             cCCCCeEEEeCCCCCEEeECCCCCc
Confidence            677788887644 467899999986


No 35 
>PF12661 hEGF:  Human growth factor-like EGF; PDB: 2YGQ_A 2E26_A 3A7Q_A 2YGP_A 2YGO_A 1HRE_A 1HAE_A 1HAF_A 1HRF_A.
Probab=68.76  E-value=1.1  Score=22.89  Aligned_cols=9  Identities=44%  Similarity=1.442  Sum_probs=6.6

Q ss_pred             CCcCCCCCe
Q 043869          287 NCTCLPGFV  295 (450)
Q Consensus       287 ~C~C~~GF~  295 (450)
                      .|.|++||.
T Consensus         1 ~C~C~~G~~    9 (13)
T PF12661_consen    1 TCQCPPGWT    9 (13)
T ss_dssp             EEEE-TTEE
T ss_pred             CccCcCCCc
Confidence            489999986


No 36 
>smart00181 EGF Epidermal growth factor-like domain.
Probab=67.61  E-value=4.4  Score=25.89  Aligned_cols=25  Identities=28%  Similarity=0.834  Sum_probs=18.8

Q ss_pred             CCCCCCcccccCCCCCCCcCCCCCee
Q 043869          271 GLCGFNSFCVLNDQTPNCTCLPGFVA  296 (450)
Q Consensus       271 g~CG~~g~C~~~~~~~~C~C~~GF~~  296 (450)
                      ..|... .|........|.|++||..
T Consensus         6 ~~C~~~-~C~~~~~~~~C~C~~g~~g   30 (35)
T smart00181        6 GPCSNG-TCINTPGSYTCSCPPGYTG   30 (35)
T ss_pred             CCCCCC-EEECCCCCeEeECCCCCcc
Confidence            456666 8876445788999999975


No 37 
>PF14575 EphA2_TM:  Ephrin type-A receptor 2 transmembrane domain; PDB: 3KUL_A 2XVD_A 2VX1_A 2VWV_A 2VX0_A 2VWY_A 2VWZ_A 2VWW_A 2VWU_A 2VWX_A ....
Probab=66.58  E-value=1.5  Score=34.36  Aligned_cols=25  Identities=28%  Similarity=0.458  Sum_probs=12.9

Q ss_pred             EehhhHHHHHHHHHHhheeeeEeec
Q 043869          415 IICLFVTVVILISVVTFGIFIYRYR  439 (450)
Q Consensus       415 i~~~~~~~~~l~~~~~~~~~~~r~~  439 (450)
                      +.++++|+++++++++..++++||+
T Consensus         3 i~~~~~g~~~ll~~v~~~~~~~rr~   27 (75)
T PF14575_consen    3 IASIIVGVLLLLVLVIIVIVCFRRC   27 (75)
T ss_dssp             HHHHHHHHHHHHHHHHHHHCCCTT-
T ss_pred             EehHHHHHHHHHHhheeEEEEEeeE
Confidence            4455666666655555444444443


No 38 
>PTZ00046 rifin; Provisional
Probab=63.78  E-value=1.7  Score=43.89  Aligned_cols=24  Identities=29%  Similarity=0.609  Sum_probs=15.6

Q ss_pred             EehhhHHHHHHHHHHhheeeeEee
Q 043869          415 IICLFVTVVILISVVTFGIFIYRY  438 (450)
Q Consensus       415 i~~~~~~~~~l~~~~~~~~~~~r~  438 (450)
                      |++++++.++++.+++++|++.|.
T Consensus       317 IiaSiiAIvVIVLIMvIIYLILRY  340 (358)
T PTZ00046        317 IIASIVAIVVIVLIMVIIYLILRY  340 (358)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHHh
Confidence            455666666666667777886553


No 39 
>TIGR01477 RIFIN variant surface antigen, rifin family. This model represents the rifin branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of rifin sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 20 bits.
Probab=63.28  E-value=1.8  Score=43.67  Aligned_cols=33  Identities=18%  Similarity=0.410  Sum_probs=18.9

Q ss_pred             EehhhHHHHHHHHHHhheeeeEe--eccceeeccc
Q 043869          415 IICLFVTVVILISVVTFGIFIYR--YRVGSYRRIQ  447 (450)
Q Consensus       415 i~~~~~~~~~l~~~~~~~~~~~r--~~~~~~~~~~  447 (450)
                      |++++++.++++.+++++|++.|  ||++-.+|||
T Consensus       312 IiaSiIAIvvIVLIMvIIYLILRYRRKKKMkKKLQ  346 (353)
T TIGR01477       312 IIASIIAILIIVLIMVIIYLILRYRRKKKMKKKLQ  346 (353)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHHhhhcchhHHHHH
Confidence            45566666666666777788655  3333334443


No 40 
>PF12662 cEGF:  Complement Clr-like EGF-like
Probab=62.25  E-value=4  Score=24.60  Aligned_cols=11  Identities=36%  Similarity=1.135  Sum_probs=9.5

Q ss_pred             CCcCCCCCeec
Q 043869          287 NCTCLPGFVAI  297 (450)
Q Consensus       287 ~C~C~~GF~~~  297 (450)
                      .|+|++||+..
T Consensus         3 ~C~C~~Gy~l~   13 (24)
T PF12662_consen    3 TCSCPPGYQLS   13 (24)
T ss_pred             EeeCCCCCcCC
Confidence            69999999864


No 41 
>PF13908 Shisa:  Wnt and FGF inhibitory regulator
Probab=60.50  E-value=7.2  Score=35.62  Aligned_cols=20  Identities=10%  Similarity=0.323  Sum_probs=10.7

Q ss_pred             eEEEEehhhHHHHHHHHHHh
Q 043869          411 KNIVIICLFVTVVILISVVT  430 (450)
Q Consensus       411 ~~i~i~~~~~~~~~l~~~~~  430 (450)
                      ...||+++++++++++++++
T Consensus        77 ~~~iivgvi~~Vi~Iv~~Iv   96 (179)
T PF13908_consen   77 ITGIIVGVICGVIAIVVLIV   96 (179)
T ss_pred             eeeeeeehhhHHHHHHHhHh
Confidence            34556666666655544333


No 42 
>PF06024 DUF912:  Nucleopolyhedrovirus protein of unknown function (DUF912);  InterPro: IPR009261 This entry is represented by Autographa californica nuclear polyhedrosis virus (AcMNPV), Orf78; it is a family of uncharacterised viral proteins.
Probab=58.82  E-value=7.7  Score=32.09  Aligned_cols=9  Identities=22%  Similarity=0.080  Sum_probs=5.5

Q ss_pred             ceEEEEecC
Q 043869          382 DITFVKVDD  390 (450)
Q Consensus       382 ~~~yiKv~~  390 (450)
                      .-+-+||+-
T Consensus        15 ~yIPLKLal   23 (101)
T PF06024_consen   15 DYIPLKLAL   23 (101)
T ss_pred             cceeeeeec
Confidence            346677774


No 43 
>PF06365 CD34_antigen:  CD34/Podocalyxin family;  InterPro: IPR013836 This family consists of several mammalian CD34 antigen proteins. The CD34 antigen is a human leukocyte membrane protein expressed specifically by lymphohematopoietic progenitor cells. CD34 is a phosphoprotein. Activation of protein kinase C (PKC) has been found to enhance CD34 phosphorylation [, ]. This family contains several eukaryotic podocalyxin proteins. Podocalyxin is a major membrane protein of the glomerular epithelium and is thought to be involved in maintenance of the architecture of the foot processes and filtration slits characteristic of this unique epithelium by virtue of its high negative charge. Podocalyxin functions as an anti-adhesin that maintains an open filtration pathway between neighbouring foot processes in the glomerular epithelium by charge repulsion [].
Probab=58.44  E-value=9.4  Score=35.68  Aligned_cols=36  Identities=14%  Similarity=0.140  Sum_probs=19.8

Q ss_pred             EEEehhhHHHHHHHH-HHhheeeeEeeccc--eeecccC
Q 043869          413 IVIICLFVTVVILIS-VVTFGIFIYRYRVG--SYRRIQG  448 (450)
Q Consensus       413 i~i~~~~~~~~~l~~-~~~~~~~~~r~~~~--~~~~~~~  448 (450)
                      ++|..++.|.++|++ +.+++||++.||.+  +-.|+.|
T Consensus       101 ~lI~lv~~g~~lLla~~~~~~Y~~~~Rrs~~~~~~rl~E  139 (202)
T PF06365_consen  101 TLIALVTSGSFLLLAILLGAGYCCHQRRSWSKKGQRLGE  139 (202)
T ss_pred             EEEehHHhhHHHHHHHHHHHHHHhhhhccCCcchhhhcc
Confidence            555566666444444 45556777666543  3444443


No 44 
>PF07974 EGF_2:  EGF-like domain;  InterPro: IPR013111 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length. This entry contains EGF domains found in a variety of extracellular and membrane proteins
Probab=57.84  E-value=7.7  Score=25.01  Aligned_cols=24  Identities=25%  Similarity=0.708  Sum_probs=19.0

Q ss_pred             CCCCCCcccccCCCCCCCcCCCCCee
Q 043869          271 GLCGFNSFCVLNDQTPNCTCLPGFVA  296 (450)
Q Consensus       271 g~CG~~g~C~~~~~~~~C~C~~GF~~  296 (450)
                      ..|...|.|...  ...|.|.+||.-
T Consensus         6 ~~C~~~G~C~~~--~g~C~C~~g~~G   29 (32)
T PF07974_consen    6 NICSGHGTCVSP--CGRCVCDSGYTG   29 (32)
T ss_pred             CccCCCCEEeCC--CCEEECCCCCcC
Confidence            469999999852  478999999863


No 45 
>PF05454 DAG1:  Dystroglycan (Dystrophin-associated glycoprotein 1);  InterPro: IPR008465 Dystroglycan is one of the dystrophin-associated glycoproteins, which is encoded by a 5.5 kb transcript in Homo sapiens. The protein product is cleaved into two non-covalently associated subunits, [alpha] (N-terminal) and [beta] (C-terminal). In skeletal muscle the dystroglycan complex works as a transmembrane linkage between the extracellular matrix and the cytoskeleton [alpha]-dystroglycan is extracellular and binds to merosin ([alpha]-2 laminin) in the basement membrane, while [beta]-dystroglycan is a transmembrane protein and binds to dystrophin, which is a large rod-like cytoskeletal protein, absent in Duchenne muscular dystrophy patients. Dystrophin binds to intracellular actin cables. In this way, the dystroglycan complex, which links the extracellular matrix to the intracellular actin cables, is thought to provide structural integrity in muscle tissues. The dystroglycan complex is also known to serve as an agrin receptor in muscle, where it may regulate agrin-induced acetylcholine receptor clustering at the neuromuscular junction. There is also evidence which suggests the function of dystroglycan as a part of the signal transduction pathway because it is shown that Grb2, a mediator of the Ras-related signal pathway, can interact with the cytoplasmic domain of dystroglycan. In general, aberrant expression of dystrophin-associated protein complex underlies the pathogenesis of Duchenne muscular dystrophy, Becker muscular dystrophy and severe childhood autosomal recessive muscular dystrophy. Interestingly, no genetic disease has been described for either [alpha]- or [beta]-dystroglycan. Dystroglycan is widely distributed in non-muscle tissues as well as in muscle tissues. During epithelial morphogenesis of kidney, the dystroglycan complex is shown to act as a receptor for the basement membrane. Dystroglycan expression in Mus musculus brain and neural retina has also been reported. However, the physiological role of dystroglycan in non-muscle tissues has remained unclear [].; PDB: 1EG4_P.
Probab=54.29  E-value=4.2  Score=40.22  Aligned_cols=29  Identities=17%  Similarity=0.272  Sum_probs=0.0

Q ss_pred             ceEEEEehhhHHHHHHHHHHhheeeeEee
Q 043869          410 WKNIVIICLFVTVVILISVVTFGIFIYRY  438 (450)
Q Consensus       410 ~~~i~i~~~~~~~~~l~~~~~~~~~~~r~  438 (450)
                      ....+|.++|+.+++|+++++++++.+||
T Consensus       145 yL~T~IpaVVI~~iLLIA~iIa~icyrrk  173 (290)
T PF05454_consen  145 YLHTFIPAVVIAAILLIAGIIACICYRRK  173 (290)
T ss_dssp             -----------------------------
T ss_pred             hHHHHHHHHHHHHHHHHHHHHHHHhhhhh
Confidence            33444566676666666555554443333


No 46 
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=53.91  E-value=49  Score=33.88  Aligned_cols=71  Identities=15%  Similarity=0.330  Sum_probs=42.7

Q ss_pred             CCcEEEEecCCC----------------CCCCCccEEEE-ecCCcEEEEeCCCCceEEeecCCC-----Cc-----cEEE
Q 043869           70 EKTVVWTANRDN----------------PPVSSNATLMF-NSEGRIVLRSGEQGQNSIIADNSQ-----SA-----SSAS  122 (450)
Q Consensus        70 ~~tvVW~ANr~~----------------Pv~~~~~~L~l-~~~G~LvL~d~~~g~~vW~st~~~-----~~-----~~a~  122 (450)
                      ...++|...-..                |+. .+..+-+ +.+|.|+-.|..+|.++| +....     ..     ....
T Consensus        88 tG~~~W~~~~~~~~~~~~~~~~~~~~~~~~v-~~~~v~v~~~~g~l~ald~~tG~~~W-~~~~~~~~~ssP~v~~~~v~v  165 (394)
T PRK11138         88 TGKEIWSVDLSEKDGWFSKNKSALLSGGVTV-AGGKVYIGSEKGQVYALNAEDGEVAW-QTKVAGEALSRPVVSDGLVLV  165 (394)
T ss_pred             CCcEeeEEcCCCcccccccccccccccccEE-ECCEEEEEcCCCEEEEEECCCCCCcc-cccCCCceecCCEEECCEEEE
Confidence            467899865432                222 2344545 467888877754799999 44321     11     1112


Q ss_pred             EecCCCeEEEec-CCeeEEee
Q 043869          123 MLDSGSFVLHNS-DGKVIWQT  142 (450)
Q Consensus       123 LldsGNlVL~~~-~~~~lWQS  142 (450)
                      ...+|.|+-.|. +|+++|+-
T Consensus       166 ~~~~g~l~ald~~tG~~~W~~  186 (394)
T PRK11138        166 HTSNGMLQALNESDGAVKWTV  186 (394)
T ss_pred             ECCCCEEEEEEccCCCEeeee
Confidence            245677888885 78899984


No 47 
>KOG1219 consensus Uncharacterized conserved protein, contains laminin, cadherin and EGF domains [Signal transduction mechanisms]
Probab=53.18  E-value=15  Score=45.68  Aligned_cols=32  Identities=16%  Similarity=0.539  Sum_probs=21.1

Q ss_pred             CCCCCcCCCCCCcccccCCC-CCCCcCCCCCeec
Q 043869          265 EKCDPIGLCGFNSFCVLNDQ-TPNCTCLPGFVAI  297 (450)
Q Consensus       265 ~~C~~~g~CG~~g~C~~~~~-~~~C~C~~GF~~~  297 (450)
                      +.|. ...|---|.|+.... .-.|.||+-|.-.
T Consensus      3865 d~C~-~npCqhgG~C~~~~~ggy~CkCpsqysG~ 3897 (4289)
T KOG1219|consen 3865 DPCN-DNPCQHGGTCISQPKGGYKCKCPSQYSGN 3897 (4289)
T ss_pred             cccc-cCcccCCCEecCCCCCceEEeCcccccCc
Confidence            4454 356666677775433 5689999988754


No 48 
>PF13360 PQQ_2:  PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=50.35  E-value=1.5e+02  Score=27.32  Aligned_cols=75  Identities=17%  Similarity=0.289  Sum_probs=46.5

Q ss_pred             cCCCcEEEEecCCCCCCC----CccE-EEEecCCcEEEEeCCCCceEEeecC-C---CC----------ccEEEE-ecCC
Q 043869           68 IPEKTVVWTANRDNPPVS----SNAT-LMFNSEGRIVLRSGEQGQNSIIADN-S---QS----------ASSASM-LDSG  127 (450)
Q Consensus        68 ~~~~tvVW~ANr~~Pv~~----~~~~-L~l~~~G~LvL~d~~~g~~vW~st~-~---~~----------~~~a~L-ldsG  127 (450)
                      +...+++|....+.++..    .... +..+.+|.|...|..+|.++|.... .   ..          ...+.+ ..+|
T Consensus        53 ~~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~g  132 (238)
T PF13360_consen   53 AKTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYALDAKTGKVLWSIYLTSSPPAGVRSSSSPAVDGDRLYVGTSSG  132 (238)
T ss_dssp             TTTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEEEETTTSCEEEEEEE-SSCTCSTB--SEEEEETTEEEEEETCS
T ss_pred             CCCCCEEEEeeccccccceeeecccccccccceeeeEecccCCcceeeeeccccccccccccccCceEecCEEEEEeccC
Confidence            346789999886655331    2333 4445678888888438999995211 1   10          111222 3488


Q ss_pred             CeEEEe-cCCeeEEee
Q 043869          128 SFVLHN-SDGKVIWQT  142 (450)
Q Consensus       128 NlVL~~-~~~~~lWQS  142 (450)
                      .++..| .+|+.+|+-
T Consensus       133 ~l~~~d~~tG~~~w~~  148 (238)
T PF13360_consen  133 KLVALDPKTGKLLWKY  148 (238)
T ss_dssp             EEEEEETTTTEEEEEE
T ss_pred             cEEEEecCCCcEEEEe
Confidence            999998 478999984


No 49 
>PF09064 Tme5_EGF_like:  Thrombomodulin like fifth domain, EGF-like;  InterPro: IPR015149 This domain adopts a fold similar to other EGF domains, with a flat major and a twisted minor beta sheet. Disulphide pairing, however, is not of the usual 1-3, 2-4, 5-6 type; rather 1-2, 3-4, 5-6 pairing is found. Its extended major sheet (strands beta-2 and beta-3 and the connecting loop) projects into thrombin's active site groove. This domain is required for interaction of thrombomodulin with thrombin, and subsequent activation of protein-C []. ; GO: 0004888 transmembrane signaling receptor activity, 0016021 integral to membrane
Probab=49.77  E-value=9.4  Score=24.96  Aligned_cols=18  Identities=22%  Similarity=0.648  Sum_probs=12.7

Q ss_pred             ccccCCCCCCCcCCCCCee
Q 043869          278 FCVLNDQTPNCTCLPGFVA  296 (450)
Q Consensus       278 ~C~~~~~~~~C~C~~GF~~  296 (450)
                      .|+. ++...|.||.||..
T Consensus        11 ~CDp-n~~~~C~CPeGyIl   28 (34)
T PF09064_consen   11 DCDP-NSPGQCFCPEGYIL   28 (34)
T ss_pred             ccCC-CCCCceeCCCceEe
Confidence            4544 23568999999976


No 50 
>PF15102 TMEM154:  TMEM154 protein family
Probab=48.12  E-value=16  Score=32.31  Aligned_cols=31  Identities=13%  Similarity=0.426  Sum_probs=20.0

Q ss_pred             EehhhHHHHHH-HHHHhheeeeEeeccceeec
Q 043869          415 IICLFVTVVIL-ISVVTFGIFIYRYRVGSYRR  445 (450)
Q Consensus       415 i~~~~~~~~~l-~~~~~~~~~~~r~~~~~~~~  445 (450)
                      |+-+++..++| +.++++++++.+.|+||.++
T Consensus        58 iLmIlIP~VLLvlLLl~vV~lv~~~kRkr~K~   89 (146)
T PF15102_consen   58 ILMILIPLVLLVLLLLSVVCLVIYYKRKRTKQ   89 (146)
T ss_pred             EEEEeHHHHHHHHHHHHHHHheeEEeecccCC
Confidence            44555564444 55566666688888888765


No 51 
>PF12768 Rax2:  Cortical protein marker for cell polarity
Probab=46.47  E-value=12  Score=37.06  Aligned_cols=27  Identities=7%  Similarity=0.173  Sum_probs=17.3

Q ss_pred             CcceEEEEehhhHHHHHHHHHHhheee
Q 043869          408 GLWKNIVIICLFVTVVILISVVTFGIF  434 (450)
Q Consensus       408 ~~~~~i~i~~~~~~~~~l~~~~~~~~~  434 (450)
                      +-.+++|.+++.+|.+.|+.++.+++.
T Consensus       226 ~G~VVlIslAiALG~v~ll~l~Gii~~  252 (281)
T PF12768_consen  226 RGFVVLISLAIALGTVFLLVLIGIILA  252 (281)
T ss_pred             ceEEEEEehHHHHHHHHHHHHHHHHHH
Confidence            334567777888888777766544443


No 52 
>PF14670 FXa_inhibition:  Coagulation Factor Xa inhibitory site; PDB: 3Q3K_B 1NFY_B 1LQD_A 1G2L_B 1IQF_L 2UWP_B 2VH6_B 3KQC_L 2P93_L 2BQW_A ....
Probab=45.26  E-value=5.6  Score=26.41  Aligned_cols=21  Identities=29%  Similarity=0.768  Sum_probs=13.3

Q ss_pred             ccccCCCCCCCcCCCCCeecc
Q 043869          278 FCVLNDQTPNCTCLPGFVAIS  298 (450)
Q Consensus       278 ~C~~~~~~~~C~C~~GF~~~~  298 (450)
                      +|........|+|++||....
T Consensus        11 ~C~~~~g~~~C~C~~Gy~L~~   31 (36)
T PF14670_consen   11 ICVNTPGSYRCSCPPGYKLAE   31 (36)
T ss_dssp             EEEEETTSEEEE-STTEEE-T
T ss_pred             CCccCCCceEeECCCCCEECc
Confidence            455433367899999998753


No 53 
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=45.17  E-value=1.2e+02  Score=30.86  Aligned_cols=48  Identities=15%  Similarity=0.144  Sum_probs=27.7

Q ss_pred             cCCcEEEEeCCCCceEEeecCC------CC-----ccEEEEecCCCeEEEec-CCeeEEee
Q 043869           94 SEGRIVLRSGEQGQNSIIADNS------QS-----ASSASMLDSGSFVLHNS-DGKVIWQT  142 (450)
Q Consensus        94 ~~G~LvL~d~~~g~~vW~st~~------~~-----~~~a~LldsGNlVL~~~-~~~~lWQS  142 (450)
                      .+|.|+..|..+|..+| +.+.      .+     .......++|.|...|. +|+++|+-
T Consensus       302 ~~g~l~ald~~tG~~~W-~~~~~~~~~~~sp~v~~g~l~v~~~~G~l~~ld~~tG~~~~~~  361 (394)
T PRK11138        302 QNDRVYALDTRGGVELW-SQSDLLHRLLTAPVLYNGYLVVGDSEGYLHWINREDGRFVAQQ  361 (394)
T ss_pred             CCCeEEEEECCCCcEEE-cccccCCCcccCCEEECCEEEEEeCCCEEEEEECCCCCEEEEE
Confidence            46777766654677788 3321      11     11122356777777775 67788875


No 54 
>cd05845 Ig2_L1-CAM_like Second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM) and similar proteins. Ig2_L1-CAM_like: domain similar to the second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM). L1 belongs to the L1 subfamily of cell adhesion molecules (CAMs) and is comprised of an extracellular region having six Ig-like domains, five fibronectin type III domains, a transmembrane region and an intracellular domain. L1 is primarily expressed in the nervous system and is involved in its development and function. L1 is associated with an X-linked recessive disorder, X-linked hydrocephalus, MASA syndrome, or spastic paraplegia type 1, that involves abnormalities of axonal growth.
Probab=44.77  E-value=35  Score=27.86  Aligned_cols=35  Identities=6%  Similarity=0.127  Sum_probs=23.1

Q ss_pred             cCCCcEEEEecCCCCCCCCccEEEEecCCcEEEEeC
Q 043869           68 IPEKTVVWTANRDNPPVSSNATLMFNSEGRIVLRSG  103 (450)
Q Consensus        68 ~~~~tvVW~ANr~~Pv~~~~~~L~l~~~G~LvL~d~  103 (450)
                      .|..++.|+-+....+. ......++.+|+|.+.+-
T Consensus        31 ~P~P~i~W~~~~~~~i~-~~~Ri~~~~~GnL~fs~v   65 (95)
T cd05845          31 AVPLRIYWMNSDLLHIT-QDERVSMGQNGNLYFANV   65 (95)
T ss_pred             CCCCEEEEECCCCcccc-ccccEEECCCceEEEEEE
Confidence            45678889844433344 466777877888887653


No 55 
>PF12690 BsuPI:  Intracellular proteinase inhibitor;  InterPro: IPR020481 BsuPI is a intracellular proteinase inhibitor that directly regulates the major intracellular proteinase (ISP-1) activity in vivo. It inhibits ISP-1 in the early stages of sporulation and then may be inactivated by a membrane-bound proteinase [].; PDB: 3ISY_A.
Probab=44.68  E-value=28  Score=27.55  Aligned_cols=16  Identities=13%  Similarity=0.219  Sum_probs=9.2

Q ss_pred             cEEEEeCCCCceEEeec
Q 043869           97 RIVLRSGEQGQNSIIAD  113 (450)
Q Consensus        97 ~LvL~d~~~g~~vW~st  113 (450)
                      +++|.|. +|..||.-+
T Consensus        27 D~~v~d~-~g~~vwrwS   42 (82)
T PF12690_consen   27 DFVVKDK-EGKEVWRWS   42 (82)
T ss_dssp             EEEEE-T-T--EEEETT
T ss_pred             EEEEECC-CCCEEEEec
Confidence            4778887 888888543


No 56 
>KOG0291 consensus WD40-repeat-containing subunit of the 18S rRNA processing complex [RNA processing and modification]
Probab=43.37  E-value=4.7e+02  Score=29.52  Aligned_cols=56  Identities=20%  Similarity=0.492  Sum_probs=39.4

Q ss_pred             CccEEEEecCCcEEEEeCCCCce-EEeecC----------CCCccEEEEecCCCeEEEec-CCee-EEe
Q 043869           86 SNATLMFNSEGRIVLRSGEQGQN-SIIADN----------SQSASSASMLDSGSFVLHNS-DGKV-IWQ  141 (450)
Q Consensus        86 ~~~~L~l~~~G~LvL~d~~~g~~-vW~st~----------~~~~~~a~LldsGNlVL~~~-~~~~-lWQ  141 (450)
                      .-..+.-+.||.++.+.+.+|.+ ||....          +++++.....-+||.+|-.. +|+| .|.
T Consensus       352 ~i~~l~YSpDgq~iaTG~eDgKVKvWn~~SgfC~vTFteHts~Vt~v~f~~~g~~llssSLDGtVRAwD  420 (893)
T KOG0291|consen  352 RITSLAYSPDGQLIATGAEDGKVKVWNTQSGFCFVTFTEHTSGVTAVQFTARGNVLLSSSLDGTVRAWD  420 (893)
T ss_pred             ceeeEEECCCCcEEEeccCCCcEEEEeccCceEEEEeccCCCceEEEEEEecCCEEEEeecCCeEEeee
Confidence            34668888999998887645554 894322          13566777889999999876 6765 665


No 57 
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=43.29  E-value=1.7e+02  Score=29.38  Aligned_cols=50  Identities=20%  Similarity=0.266  Sum_probs=24.9

Q ss_pred             ecCCcEEEEeCCCCceEEeecCC--C-----CccEEEEecCCCeEEEec-CCeeEEee
Q 043869           93 NSEGRIVLRSGEQGQNSIIADNS--Q-----SASSASMLDSGSFVLHNS-DGKVIWQT  142 (450)
Q Consensus        93 ~~~G~LvL~d~~~g~~vW~st~~--~-----~~~~a~LldsGNlVL~~~-~~~~lWQS  142 (450)
                      +.+|.|+..|..+|.++|.....  .     ....-...++|.++..|. +++.+|+.
T Consensus       248 ~~~g~l~a~d~~tG~~~W~~~~~~~~~p~~~~~~vyv~~~~G~l~~~d~~tG~~~W~~  305 (377)
T TIGR03300       248 SYQGRVAALDLRSGRVLWKRDASSYQGPAVDDNRLYVTDADGVVVALDRRSGSELWKN  305 (377)
T ss_pred             EcCCEEEEEECCCCcEEEeeccCCccCceEeCCEEEEECCCCeEEEEECCCCcEEEcc
Confidence            45666666665356677732211  0     111112234566666664 45667753


No 58 
>PF01436 NHL:  NHL repeat;  InterPro: IPR001258 The NHL repeat, named after NCL-1, HT2A and Lin-41, is found largely in a large number of eukaryotic and prokaryotic proteins. For example, the repeat is found in a variety of enzymes of the copper type II, ascorbate-dependent monooxygenase family which catalyse the C terminus alpha-amidation of biological peptides []. In many it occurs in tandem arrays, for example in the ringfinger beta-box, coiled-coil (RBCC) eukaryotic growth regulators []. The 'Brain Tumor' protein (Brat) is one such growth regulator that contains a 6-bladed NHL-repeat beta-propeller [, ].  The NHL repeats are also found in serine/threonine protein kinase (STPK) in diverse range of pathogenic bacteria. These STPK are transmembrane receptors with a intracellular N-terminal kinase domain and extracellular C-terminal sensor domain. In the STPK, PknD, from Mycobacterium tuberculosis, the sensor domain forms a rigid, six-bladed b-propeller composed of NHL repeats with a flexible tether to the transmembrane domain.; GO: 0005515 protein binding; PDB: 3FVZ_A 3FW0_A 1RWL_A 1RWI_A 1Q7F_A.
Probab=41.33  E-value=49  Score=20.23  Aligned_cols=21  Identities=14%  Similarity=0.273  Sum_probs=14.6

Q ss_pred             EEEEecCCcEEEEeCCCCceEE
Q 043869           89 TLMFNSEGRIVLRSGEQGQNSI  110 (450)
Q Consensus        89 ~L~l~~~G~LvL~d~~~g~~vW  110 (450)
                      -+.++.+|++++.|. ....||
T Consensus         6 gvav~~~g~i~VaD~-~n~rV~   26 (28)
T PF01436_consen    6 GVAVDSDGNIYVADS-GNHRVQ   26 (28)
T ss_dssp             EEEEETTSEEEEEEC-CCTEEE
T ss_pred             EEEEeCCCCEEEEEC-CCCEEE
Confidence            366777888888886 555555


No 59 
>PF13360 PQQ_2:  PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=41.19  E-value=1.3e+02  Score=27.70  Aligned_cols=73  Identities=16%  Similarity=0.335  Sum_probs=42.0

Q ss_pred             CCcEEEEecC----CCCC--C-CCccEEEE-ecCCcEEEEeCCCCceEEeecCCCC---c-----cEEEE-ecCCCeEEE
Q 043869           70 EKTVVWTANR----DNPP--V-SSNATLMF-NSEGRIVLRSGEQGQNSIIADNSQS---A-----SSASM-LDSGSFVLH  132 (450)
Q Consensus        70 ~~tvVW~ANr----~~Pv--~-~~~~~L~l-~~~G~LvL~d~~~g~~vW~st~~~~---~-----~~a~L-ldsGNlVL~  132 (450)
                      ....+|..+-    ..++  . ..+..+.+ +.+|.|+..|..+|..+|.......   .     ....+ ..+|.|+..
T Consensus        12 tG~~~W~~~~~~~~~~~~~~~~~~~~~v~~~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~~   91 (238)
T PF13360_consen   12 TGKELWSYDLGPGIGGPVATAVPDGGRVYVASGDGNLYALDAKTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYAL   91 (238)
T ss_dssp             TTEEEEEEECSSSCSSEEETEEEETTEEEEEETTSEEEEEETTTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEEE
T ss_pred             CCCEEEEEECCCCCCCccceEEEeCCEEEEEcCCCEEEEEECCCCCEEEEeeccccccceeeecccccccccceeeeEec
Confidence            5677888753    2222  1 02333444 4888999888547899995432211   1     11222 234556667


Q ss_pred             e-cCCeeEEee
Q 043869          133 N-SDGKVIWQT  142 (450)
Q Consensus       133 ~-~~~~~lWQS  142 (450)
                      | .+|+++|+.
T Consensus        92 d~~tG~~~W~~  102 (238)
T PF13360_consen   92 DAKTGKVLWSI  102 (238)
T ss_dssp             ETTTSCEEEEE
T ss_pred             ccCCcceeeee
Confidence            7 678999995


No 60 
>PF03302 VSP:  Giardia variant-specific surface protein;  InterPro: IPR005127 During infection, the intestinal protozoan parasite Giardia lamblia virus undergoes continuous antigenic variation which is determined by diversification of the parasite's major surface antigen, named VSP (variant surface protein).
Probab=41.19  E-value=22  Score=36.88  Aligned_cols=15  Identities=20%  Similarity=0.007  Sum_probs=11.2

Q ss_pred             HHHHHhheeeeEeec
Q 043869          425 LISVVTFGIFIYRYR  439 (450)
Q Consensus       425 l~~~~~~~~~~~r~~  439 (450)
                      -|+-++.||||.|.|
T Consensus       382 glvGfLcWwf~crgk  396 (397)
T PF03302_consen  382 GLVGFLCWWFICRGK  396 (397)
T ss_pred             HHHHHHhhheeeccc
Confidence            345577799999876


No 61 
>PF12877 DUF3827:  Domain of unknown function (DUF3827);  InterPro: IPR024606 The function of the proteins in this entry is not currently known, but one of the human proteins (Q9HCM3 from SWISSPROT) has been implicated in pilocytic astrocytomas [, , ]. In the majority of cases of pilocytic astrocytomas a tandem duplication produces an in-frame fusion of the gene encoding this protein and the BRAF oncogene. The resulting fusion protein has constitutive BRAF kinase activity and is capable of transforming cells. 
Probab=40.27  E-value=21  Score=38.88  Aligned_cols=30  Identities=13%  Similarity=0.280  Sum_probs=17.0

Q ss_pred             ceEEEEehhhHHHHHHHHHHhheee-eEeec
Q 043869          410 WKNIVIICLFVTVVILISVVTFGIF-IYRYR  439 (450)
Q Consensus       410 ~~~i~i~~~~~~~~~l~~~~~~~~~-~~r~~  439 (450)
                      ..++||+++++.++++++|++++++ ++|++
T Consensus       267 ~NlWII~gVlvPv~vV~~Iiiil~~~LCRk~  297 (684)
T PF12877_consen  267 NNLWIIAGVLVPVLVVLLIIIILYWKLCRKN  297 (684)
T ss_pred             CCeEEEehHhHHHHHHHHHHHHHHHHHhccc
Confidence            4456677777666665555444444 55543


No 62 
>PF12191 stn_TNFRSF12A:  Tumour necrosis factor receptor stn_TNFRSF12A_TNFR domain;  InterPro: IPR022316 The tumour necrosis factor (TNF) receptor (TNFR) superfamily comprises more than 20 type-I transmembrane proteins. Family members are defined based on similarity in their extracellular domain - a region that contains many cysteine residues arranged in a specific repetitive pattern []. The cysteines allow formation of an extended rod-like structure, responsible for ligand binding []. Upon receptor activation, different intracellular signalling complexes are assembled for different members of the TNFR superfamily, depending on their intracellular domains and sequences []. Activation of TNFRs can therefore induce a range of disparate effects, including cell proliferation, differentiation, survival, or apoptotic cell death, depending upon the receptor involved []. TNFRs are widely distributed and play important roles in many crucial biological processes, such as lymphoid and neuronal development, innate and adaptive immunity, and maintenance of cellular homeostasis []. Drugs that manipulate their signalling have potential roles in the prevention and treatment of many diseases, such as viral infections, coronary heart disease, transplant rejection, and immune disease []. TNF receptor 12 (also known as TWEAK receptor, and fibroblast growth factor-inducible-14 (Fn14)) has been implicated in endothelial cell growth and migration []. The receptor may also play a role in cell-matrix interactions [].; PDB: 2KN0_A 2RPJ_A 2KMZ_A 2EQP_A.
Probab=40.14  E-value=8  Score=33.14  Aligned_cols=37  Identities=22%  Similarity=0.434  Sum_probs=0.0

Q ss_pred             EEEehhhHHHHHHHHHHhheeeeEe---eccceeecccCCC
Q 043869          413 IVIICLFVTVVILISVVTFGIFIYR---YRVGSYRRIQGNG  450 (450)
Q Consensus       413 i~i~~~~~~~~~l~~~~~~~~~~~r---~~~~~~~~~~~~~  450 (450)
                      ..|..+++++++++.++. .++++|   ||.+-..-|+|+|
T Consensus        79 ~pi~~sal~v~lVl~lls-g~lv~rrcrrr~~~ttPIeeTg  118 (129)
T PF12191_consen   79 WPILGSALSVVLVLALLS-GFLVWRRCRRREKFTTPIEETG  118 (129)
T ss_dssp             -----------------------------------------
T ss_pred             hhhhhhHHHHHHHHHHHH-HHHHHhhhhccccCCCcccccC
Confidence            334445555554433322 233322   3334444677765


No 63 
>PHA02887 EGF-like protein; Provisional
Probab=39.80  E-value=32  Score=29.17  Aligned_cols=35  Identities=26%  Similarity=0.412  Sum_probs=25.1

Q ss_pred             ccccCCCCC--CcCCCCCCcccccCC--CCCCCcCCCCCe
Q 043869          260 WPSTSEKCD--PIGLCGFNSFCVLND--QTPNCTCLPGFV  295 (450)
Q Consensus       260 w~~p~~~C~--~~g~CG~~g~C~~~~--~~~~C~C~~GF~  295 (450)
                      ++..-++|.  ..++|= +|.|.+-.  +.|.|.|++||.
T Consensus        79 ~~~hf~pC~~eyk~YCi-HG~C~yI~dL~epsCrC~~GYt  117 (126)
T PHA02887         79 NSMFFEKCKNDFNDFCI-NGECMNIIDLDEKFCICNKGYT  117 (126)
T ss_pred             cccCccccChHhhCEee-CCEEEccccCCCceeECCCCcc
Confidence            344456785  367787 78997643  378999999985


No 64 
>PF05337 CSF-1:  Macrophage colony stimulating factor-1 (CSF-1);  InterPro: IPR008001 Colony stimulating factor 1 (CSF-1) is a homodimeric polypeptide growth factor whose primary function is to regulate the survival, proliferation, differentiation, and function of cells of the mononuclear phagocytic lineage. This lineage includes mononuclear phagocytic precursors, blood monocytes, tissue macrophages, osteoclasts, and microglia of the brain, all of which possess cell surface receptors for CSF-1. The protein has also been linked with male fertility [] and mutations in the Csf-1 gene have been found to cause osteopetrosis and failure of tooth eruption [].; GO: 0005125 cytokine activity, 0008083 growth factor activity, 0016021 integral to membrane; PDB: 3EJJ_A.
Probab=39.42  E-value=9.8  Score=37.04  Aligned_cols=30  Identities=33%  Similarity=0.470  Sum_probs=0.0

Q ss_pred             hHHHHHHHHHHhheeeeEeeccceeecccC
Q 043869          419 FVTVVILISVVTFGIFIYRYRVGSYRRIQG  448 (450)
Q Consensus       419 ~~~~~~l~~~~~~~~~~~r~~~~~~~~~~~  448 (450)
                      .+..|+++.++++.+++||+|++..++-|.
T Consensus       231 LVPSiILVLLaVGGLLfYr~rrRs~~e~q~  260 (285)
T PF05337_consen  231 LVPSIILVLLAVGGLLFYRRRRRSHREPQT  260 (285)
T ss_dssp             ------------------------------
T ss_pred             cccchhhhhhhccceeeecccccccccccc
Confidence            344555666777778888877666665543


No 65 
>PHA03099 epidermal growth factor-like protein (EGF-like protein); Provisional
Probab=38.84  E-value=16  Score=31.47  Aligned_cols=31  Identities=19%  Similarity=0.200  Sum_probs=13.5

Q ss_pred             hhhHHHHHHHHHHhheeeeEeeccceeeccc
Q 043869          417 CLFVTVVILISVVTFGIFIYRYRVGSYRRIQ  447 (450)
Q Consensus       417 ~~~~~~~~l~~~~~~~~~~~r~~~~~~~~~~  447 (450)
                      .+++++++++++..+.++++|.-++|+..+|
T Consensus       104 ~~il~il~~i~is~~~~~~yr~~r~~~~~~~  134 (139)
T PHA03099        104 PGIVLVLVGIIITCCLLSVYRFTRRTKLPLQ  134 (139)
T ss_pred             hHHHHHHHHHHHHHHHHhhheeeecccCchh
Confidence            3444444444444444444444333343343


No 66 
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=37.83  E-value=2.1e+02  Score=28.74  Aligned_cols=18  Identities=22%  Similarity=0.602  Sum_probs=12.1

Q ss_pred             cCCCeEEEec-CCeeEEee
Q 043869          125 DSGSFVLHNS-DGKVIWQT  142 (450)
Q Consensus       125 dsGNlVL~~~-~~~~lWQS  142 (450)
                      .+|+++.+|. +++.+|+.
T Consensus       249 ~~g~l~a~d~~tG~~~W~~  267 (377)
T TIGR03300       249 YQGRVAALDLRSGRVLWKR  267 (377)
T ss_pred             cCCEEEEEECCCCcEEEee
Confidence            4667777775 56778864


No 67 
>PF15065 NCU-G1:  Lysosomal transcription factor, NCU-G1
Probab=37.60  E-value=12  Score=38.13  Aligned_cols=32  Identities=19%  Similarity=0.122  Sum_probs=21.6

Q ss_pred             eEEEEehhhHHHHHHHHHHhheeeeEeeccce
Q 043869          411 KNIVIICLFVTVVILISVVTFGIFIYRYRVGS  442 (450)
Q Consensus       411 ~~i~i~~~~~~~~~l~~~~~~~~~~~r~~~~~  442 (450)
                      .+|+|+++.+|+-++++++.++|.+.||+++|
T Consensus       318 lvi~i~~vgLG~P~l~li~Ggl~v~~~r~r~~  349 (350)
T PF15065_consen  318 LVIMIMAVGLGVPLLLLILGGLYVCLRRRRKR  349 (350)
T ss_pred             HHHHHHHHHhhHHHHHHHHhhheEEEeccccC
Confidence            35667777788877777777777666555444


No 68 
>COG3763 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=37.28  E-value=4.3  Score=31.02  Aligned_cols=28  Identities=21%  Similarity=0.402  Sum_probs=17.3

Q ss_pred             ehhhHHHHHHHHHHhheeeeEeecccee
Q 043869          416 ICLFVTVVILISVVTFGIFIYRYRVGSY  443 (450)
Q Consensus       416 ~~~~~~~~~l~~~~~~~~~~~r~~~~~~  443 (450)
                      +++++.++.+++.++++||+.||.-+++
T Consensus         5 lail~ivl~ll~G~~~G~fiark~~~k~   32 (71)
T COG3763           5 LAILLIVLALLAGLIGGFFIARKQMKKQ   32 (71)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence            3445556666666777788877654443


No 69 
>PF01034 Syndecan:  Syndecan domain;  InterPro: IPR001050 The syndecans are transmembrane proteoglycans which are involved in the organisation of cytoskeleton and/or actin microfilaments, and have important roles as cell surface receptors during cell-cell and/or cell-matrix interactions [, ]. Structurally, these proteins consist of four separate domains:   A signal sequence; An extracellular domain (ectodomain) of variable length whose sequence is not evolutionary conserved in the various forms of syndecans. The ectodomain contains the sites of attachment of the heparan sulphate glycosaminoglycan side chains;  A transmembrane region;  A highly conserved cytoplasmic domain of about 30 to 35 residues, which could interact with cytoskeletal proteins.    The proteins known to belong to this family are:    Syndecan 1.  Syndecan 2 or fibroglycan.  Syndecan 3 or neuroglycan or N-syndecan.  Syndecan 4 or amphiglycan or ryudocan.  Drosophila syndecan.   Caenorhabditis elegans probable syndecan (F57C7.3).    Syndecan-4, a transmembrane heparan sulphate proteoglycan, is a coreceptor with integrins in cell adhesion. It has been suggested to form a ternary signalling complex with protein kinase Calpha and phosphatidylinositol 4,5-bisphosphate (PIP2). Structural studies have demonstrated that the cytoplasmic domain undergoes a conformational transition and forms a symmetric dimer in the presence of phospholipid activator PIP2, and whose overall structure in solution exhibits a twisted clamp shape having a cavity in the centre of dimeric interface. In addition, it has been observed that the syndecan-4 variable domain interacts, strongly, not only with fatty acyl groups but also the anionic head group of PIP2. These findings indicate that PIP2 promotes oligomerisation of the syndecan-4 cytoplasmic domain for transmembrane signalling and cell-matrix adhesion [, ].; GO: 0008092 cytoskeletal protein binding, 0016020 membrane; PDB: 1EJQ_B 1EJP_B 1YBO_C 1OBY_Q.
Probab=35.69  E-value=11  Score=28.36  Aligned_cols=8  Identities=50%  Similarity=0.816  Sum_probs=0.4

Q ss_pred             eeeeEeec
Q 043869          432 GIFIYRYR  439 (450)
Q Consensus       432 ~~~~~r~~  439 (450)
                      .++++|.|
T Consensus        30 lf~iyR~r   37 (64)
T PF01034_consen   30 LFLIYRMR   37 (64)
T ss_dssp             -------S
T ss_pred             HHHHHHHH
Confidence            34466644


No 70 
>PF14991 MLANA:  Protein melan-A; PDB: 2GTZ_F 2GT9_F 3MRO_P 2GUO_C 3MRQ_P 2GTW_C 3L6F_C 3MRP_P.
Probab=35.08  E-value=12  Score=31.46  Aligned_cols=17  Identities=12%  Similarity=-0.083  Sum_probs=0.0

Q ss_pred             HHHHHHhheeeeEeecc
Q 043869          424 ILISVVTFGIFIYRYRV  440 (450)
Q Consensus       424 ~l~~~~~~~~~~~r~~~  440 (450)
                      +.+.+++++|+.+||..
T Consensus        36 LgiLLliGCWYckRRSG   52 (118)
T PF14991_consen   36 LGILLLIGCWYCKRRSG   52 (118)
T ss_dssp             -----------------
T ss_pred             HHHHHHHhheeeeecch
Confidence            33344556777766643


No 71 
>PF07354 Sp38:  Zona-pellucida-binding protein (Sp38);  InterPro: IPR010857 This family contains a number of zona-pellucida-binding proteins that seem to be restricted to mammals. These are sperm proteins that bind to the 90 kDa family of zona pellucida glycoproteins in a calcium-dependent manner []. These represent some of the specific molecules that mediate the first steps of gamete interaction, allowing fertilisation to occur [].; GO: 0007339 binding of sperm to zona pellucida, 0005576 extracellular region
Probab=34.25  E-value=52  Score=32.03  Aligned_cols=35  Identities=17%  Similarity=0.439  Sum_probs=26.7

Q ss_pred             cCCCcEEEEecCCCCCCCCccEEEEecCCcEEEEeC
Q 043869           68 IPEKTVVWTANRDNPPVSSNATLMFNSEGRIVLRSG  103 (450)
Q Consensus        68 ~~~~tvVW~ANr~~Pv~~~~~~L~l~~~G~LvL~d~  103 (450)
                      +.+.+..|+--.+.++. .++.++||+.|.|++.|=
T Consensus        10 ~iDP~y~W~GP~g~~l~-gn~~~nIT~TG~L~~~~F   44 (271)
T PF07354_consen   10 LIDPTYLWTGPNGKPLS-GNSYVNITETGKLMFKNF   44 (271)
T ss_pred             cCCCceEEECCCCcccC-CCCeEEEccCceEEeecc
Confidence            34567788887777777 677888888888888653


No 72 
>KOG0640 consensus mRNA cleavage stimulating factor complex; subunit 1 [RNA processing and modification]
Probab=32.45  E-value=1.6e+02  Score=29.61  Aligned_cols=65  Identities=18%  Similarity=0.411  Sum_probs=44.9

Q ss_pred             EecCCCCCCCCccEEEEecCCcEEEEeCCCCc-eEEeecCC-------------CCccEEEEecCCCeEEEecCCe--eE
Q 043869           76 TANRDNPPVSSNATLMFNSEGRIVLRSGEQGQ-NSIIADNS-------------QSASSASMLDSGSFVLHNSDGK--VI  139 (450)
Q Consensus        76 ~ANr~~Pv~~~~~~L~l~~~G~LvL~d~~~g~-~vW~st~~-------------~~~~~a~LldsGNlVL~~~~~~--~l  139 (450)
                      .||.+....+.-..+.-++.|+|+++...+|. -+| -.-+             +.+.+|+...+|..+|.....+  -+
T Consensus       253 sanPd~qht~ai~~V~Ys~t~~lYvTaSkDG~Iklw-DGVS~rCv~t~~~AH~gsevcSa~Ftkn~kyiLsSG~DS~vkL  331 (430)
T KOG0640|consen  253 SANPDDQHTGAITQVRYSSTGSLYVTASKDGAIKLW-DGVSNRCVRTIGNAHGGSEVCSAVFTKNGKYILSSGKDSTVKL  331 (430)
T ss_pred             ecCcccccccceeEEEecCCccEEEEeccCCcEEee-ccccHHHHHHHHhhcCCceeeeEEEccCCeEEeecCCcceeee
Confidence            56766655544456777889999998663554 488 3321             2467899999999999876433  38


Q ss_pred             Ee
Q 043869          140 WQ  141 (450)
Q Consensus       140 WQ  141 (450)
                      |+
T Consensus       332 WE  333 (430)
T KOG0640|consen  332 WE  333 (430)
T ss_pred             ee
Confidence            98


No 73 
>KOG3637 consensus Vitronectin receptor, alpha subunit [Extracellular structures]
Probab=32.29  E-value=31  Score=40.35  Aligned_cols=22  Identities=18%  Similarity=0.190  Sum_probs=10.7

Q ss_pred             EEEEehhhHHHHHHHHH-Hhhee
Q 043869          412 NIVIICLFVTVVILISV-VTFGI  433 (450)
Q Consensus       412 ~i~i~~~~~~~~~l~~~-~~~~~  433 (450)
                      +.||+.++++.++||++ +++.|
T Consensus       978 ~wiIi~svl~GLLlL~llv~~Lw 1000 (1030)
T KOG3637|consen  978 LWIIILSVLGGLLLLALLVLLLW 1000 (1030)
T ss_pred             eeeehHHHHHHHHHHHHHHHHHH
Confidence            44455555555555444 44433


No 74 
>TIGR03066 Gem_osc_para_1 Gemmata obscuriglobus paralogous family TIGR03066. This model represents an uncharacterized paralogous family in Gemmata obscuriglobus UQM 2246, a member of the Planctomycetes. This family shows sequence similarity to TIGR03067, which is also found in Gemmata obscuriglobus as well as in a few other species.
Probab=31.61  E-value=1.8e+02  Score=24.58  Aligned_cols=52  Identities=19%  Similarity=0.366  Sum_probs=31.2

Q ss_pred             CccEEEEecCCcEEEEeCCCCce------EEe-ecCC-------C----CccE-EEEecCCCeEEEecCCee
Q 043869           86 SNATLMFNSEGRIVLRSGEQGQN------SII-ADNS-------Q----SASS-ASMLDSGSFVLHNSDGKV  138 (450)
Q Consensus        86 ~~~~L~l~~~G~LvL~d~~~g~~------vW~-st~~-------~----~~~~-a~LldsGNlVL~~~~~~~  138 (450)
                      +...|+|..||.|+|+.+ ++.-      -|. +.+.       .    .... ..=+++|-|||.|+++.+
T Consensus        34 ~~~~leF~~dGKL~v~~g-nng~~~~~~Gty~L~G~kLtL~~~p~g~t~k~~Vtv~~l~~~~Lvl~d~dg~~  104 (111)
T TIGR03066        34 DDVVIEFAKDGKLVVTIG-EKGKEVKADGTYKLDGNKLTLTLKAGGKEKKETLTVKKLTDDELVGKDPDGKK  104 (111)
T ss_pred             CceEEEEcCCCeEEEecC-CCCcEeccCceEEEECCEEEEEEcCCCccccceEEEEEecCCeEEEEcCCCCE
Confidence            456799999999998776 4321      131 1110       0    0011 123688999999988764


No 75 
>PF06006 DUF905:  Bacterial protein of unknown function (DUF905);  InterPro: IPR009253 This family consists of several short hypothetical proteobacterial proteins of unknown function.; PDB: 2HJJ_A.
Probab=31.35  E-value=53  Score=25.14  Aligned_cols=20  Identities=15%  Similarity=0.806  Sum_probs=11.7

Q ss_pred             CeEEEecCCeeEEeecCCCC
Q 043869          128 SFVLHNSDGKVIWQTFDHPT  147 (450)
Q Consensus       128 NlVL~~~~~~~lWQSFd~PT  147 (450)
                      .||++|.+|..+|..|.+-.
T Consensus        34 RlvvRd~~g~mvWRaWNFEp   53 (70)
T PF06006_consen   34 RLVVRDTEGQMVWRAWNFEP   53 (70)
T ss_dssp             EEEEE-SS--EEEEEESSST
T ss_pred             EEEEEcCCCcEEEEeeccCC
Confidence            46777777788888776543


No 76 
>KOG4649 consensus PQQ (pyrrolo-quinoline quinone) repeat protein [Secondary metabolites biosynthesis, transport and catabolism]
Probab=29.33  E-value=2e+02  Score=28.32  Aligned_cols=40  Identities=20%  Similarity=0.365  Sum_probs=28.3

Q ss_pred             CcEEEEecCCCCCCCC----ccEEEE-ecCCcEEEEeCCCCceEEe
Q 043869           71 KTVVWTANRDNPPVSS----NATLMF-NSEGRIVLRSGEQGQNSII  111 (450)
Q Consensus        71 ~tvVW~ANr~~Pv~~~----~~~L~l-~~~G~LvL~d~~~g~~vW~  111 (450)
                      .+..|.|.|..|+-.+    +....+ +-||+|.-.++ .|+.||+
T Consensus       169 ~~~~w~~~~~~PiF~splcv~~sv~i~~VdG~l~~f~~-sG~qvwr  213 (354)
T KOG4649|consen  169 STEFWAATRFGPIFASPLCVGSSVIITTVDGVLTSFDE-SGRQVWR  213 (354)
T ss_pred             cceehhhhcCCccccCceeccceEEEEEeccEEEEEcC-CCcEEEe
Confidence            4788999998887633    233333 57888877777 8888884


No 77 
>PF12946 EGF_MSP1_1:  MSP1 EGF domain 1;  InterPro: IPR024730 This EGF-like domain is found at the C terminus of the malaria parasite MSP1 protein. MSP1 is the merozoite surface protein 1. This domain is part of the C-terminal fragment that is proteolytically processed from the the rest of the protein and is left attached to the surface of the invading parasite [].; PDB: 1N1I_C 2FLG_A 1CEJ_A 2NPR_A 1B9W_A 1OB1_F.
Probab=29.18  E-value=11  Score=25.16  Aligned_cols=27  Identities=30%  Similarity=0.701  Sum_probs=17.3

Q ss_pred             CCCCCCcccccCCC-CCCCcCCCCCeec
Q 043869          271 GLCGFNSFCVLNDQ-TPNCTCLPGFVAI  297 (450)
Q Consensus       271 g~CG~~g~C~~~~~-~~~C~C~~GF~~~  297 (450)
                      ..|-.|+-|....+ ...|.|++||+..
T Consensus         5 ~~cP~NA~C~~~~dG~eecrCllgyk~~   32 (37)
T PF12946_consen    5 TKCPANAGCFRYDDGSEECRCLLGYKKV   32 (37)
T ss_dssp             S---TTEEEEEETTSEEEEEE-TTEEEE
T ss_pred             ccCCCCcccEEcCCCCEEEEeeCCcccc
Confidence            45777888865443 6789999999864


No 78 
>PHA03290 envelope glycoprotein I; Provisional
Probab=29.13  E-value=52  Score=33.03  Aligned_cols=35  Identities=17%  Similarity=0.070  Sum_probs=17.9

Q ss_pred             hhhhccccccccccCCCCcccCCCCCeEEeCCCeEEEEEEeCCC
Q 043869           12 GCFTAAAQKKHSNISIGSSLSPTGNSSWRSPSGLYAFGFYPQRN   55 (450)
Q Consensus        12 ~~~~~~~~~~~~~i~~g~~l~~~~~~~l~S~~g~F~lGF~~~~~   55 (450)
                      ++|....++  .-+-.|..++      +.+ ++.-+.||...++
T Consensus        11 ~~~~i~~~~--gIVyRG~~VS------L~v-DsSa~v~f~~~g~   45 (357)
T PHA03290         11 MIFGIQCAA--AIIFKGDHIS------LQV-NSSATSIFIKMGN   45 (357)
T ss_pred             HHHHhhhee--EEEEECCeEE------EEE-CCccceeeecCCC
Confidence            555433322  2466677653      333 4455677876433


No 79 
>TIGR03075 PQQ_enz_alc_DH PQQ-dependent dehydrogenase, methanol/ethanol family. This protein family has a phylogenetic distribution very similar to that coenzyme PQQ biosynthesis enzymes, as shown by partial phylogenetic profiling. Genes in this family often are found adjacent to the PQQ biosynthesis genes themselves. An unusual, strained disulfide bond between adjacent Cys residues contributes to PQQ-binding, as does a Trp residue that is part of a PQQ enzyme repeat (see pfam01011). Characterized members include the dehydrogenase subunit of a membrane-anchored, three subunit alcohol (ethanol) dehydrogenase of Gluconobacter suboxydans, a homodimeric ethanol dehydrogenase in Pseudomonas aeruginosa, and the large subunit of an alpha2/beta2 heterotetrameric methanol dehydrogenase in Methylobacterium extorquens.
Probab=28.43  E-value=7.7e+02  Score=26.53  Aligned_cols=121  Identities=15%  Similarity=0.313  Sum_probs=0.0

Q ss_pred             EEeCCCCCeeEEEEEEeecCCCcEEEEecCCCCCCCCc---------------cEEEE-ecCCcEEEEeCCCCceEEeec
Q 043869           50 FYPQRNGSRYYVGVFLAGIPEKTVVWTANRDNPPVSSN---------------ATLMF-NSEGRIVLRSGEQGQNSIIAD  113 (450)
Q Consensus        50 F~~~~~~~~~~lgIw~~~~~~~tvVW~ANr~~Pv~~~~---------------~~L~l-~~~G~LvL~d~~~g~~vW~st  113 (450)
                      |+.....  ...+|   ......++|..+...|.....               .++.+ +.+|.|+-.|..+|.++| +.
T Consensus        73 yv~s~~g--~v~Al---Da~TGk~lW~~~~~~~~~~~~~~~~~~~~rg~av~~~~v~v~t~dg~l~ALDa~TGk~~W-~~  146 (527)
T TIGR03075        73 YVTTSYS--RVYAL---DAKTGKELWKYDPKLPDDVIPVMCCDVVNRGVALYDGKVFFGTLDARLVALDAKTGKVVW-SK  146 (527)
T ss_pred             EEECCCC--cEEEE---ECCCCceeeEecCCCCcccccccccccccccceEECCEEEEEcCCCEEEEEECCCCCEEe-ec


Q ss_pred             CCC-------CccEEEEec--------------CCCeEEEec-CCeeEEeecCCCCC------------------ccCCC
Q 043869          114 NSQ-------SASSASMLD--------------SGSFVLHNS-DGKVIWQTFDHPTD------------------TLLPT  153 (450)
Q Consensus       114 ~~~-------~~~~a~Lld--------------sGNlVL~~~-~~~~lWQSFd~PTD------------------TlLpg  153 (450)
                      ...       ..+...+.+              +|.++-+|. +|+.+|+--.-|.+                  |+ +|
T Consensus       147 ~~~~~~~~~~~tssP~v~~g~Vivg~~~~~~~~~G~v~AlD~~TG~~lW~~~~~p~~~~~~~~~~~~~~~~~~~~tw-~~  225 (527)
T TIGR03075       147 KNGDYKAGYTITAAPLVVKGKVITGISGGEFGVRGYVTAYDAKTGKLVWRRYTVPGDMGYLDKADKPVGGEPGAKTW-PG  225 (527)
T ss_pred             ccccccccccccCCcEEECCEEEEeecccccCCCcEEEEEECCCCceeEeccCcCCCcccccccccccccccccCCC-CC


Q ss_pred             cccCCCCeEEeccCCC-CCCCCceEE
Q 043869          154 QRLSAGTELCSGISET-DPSTGKFRL  178 (450)
Q Consensus       154 q~L~~~~~L~S~~s~~-dps~G~f~l  178 (450)
                      +....+.- ..|-..+ |+..|...+
T Consensus       226 ~~~~~gg~-~~W~~~s~D~~~~lvy~  250 (527)
T TIGR03075       226 DAWKTGGG-ATWGTGSYDPETNLIYF  250 (527)
T ss_pred             CccccCCC-CccCceeEcCCCCeEEE


No 80 
>PF08114 PMP1_2:  ATPase proteolipid family;  InterPro: IPR012589 This family consists of small proteolipids associated with the plasma membrane H+ ATPase. Two proteolipids (PMP1 and PMP2) are associated with the ATPase and both genes are similarly expressed in the wild-type strain of yeast. No modification of the level of transcription of one PMP gene is detected in a strain deleted of the other. Though both proteolipids show similarity with other small proteolipids associated with other cation -transporting ATPases, their functions remain unclear [].
Probab=28.12  E-value=6.4  Score=26.76  Aligned_cols=19  Identities=32%  Similarity=0.450  Sum_probs=9.4

Q ss_pred             HhheeeeEeeccceeeccc
Q 043869          429 VTFGIFIYRYRVGSYRRIQ  447 (450)
Q Consensus       429 ~~~~~~~~r~~~~~~~~~~  447 (450)
                      .+...|+|||-..|.+-+|
T Consensus        23 ~iva~~iYRKw~aRkr~l~   41 (43)
T PF08114_consen   23 GIVALFIYRKWQARKRALQ   41 (43)
T ss_pred             HHHHHHHHHHHHHHHHHHh
Confidence            3334456666555544443


No 81 
>PF02480 Herpes_gE:  Alphaherpesvirus glycoprotein E;  InterPro: IPR003404 Glycoprotein E (gE) of Alphaherpesvirus forms a complex with glycoprotein I (gI), functioning as an immunoglobulin G (IgG) Fc binding protein. gE is involved in virus spread but is not essential for propagation [].; GO: 0016020 membrane; PDB: 2GJ7_F 2GIY_B.
Probab=28.10  E-value=20  Score=37.80  Aligned_cols=30  Identities=17%  Similarity=0.363  Sum_probs=14.6

Q ss_pred             CCCCCccCCcccccCC-cceEEeccccccCc
Q 043869          302 WTAGCERNYTAESCGN-KAIQELENTNWEDV  331 (450)
Q Consensus       302 ~s~GC~r~~~l~~C~~-~~f~~l~~v~~p~~  331 (450)
                      .+.+|.+......|.+ ..+.+..++.+.++
T Consensus       237 ~y~~C~~~~~~~~C~~~~~~~~~~~~~~~~~  267 (439)
T PF02480_consen  237 RYANCSPSGWPRRCPSTSHIEPVPGLRWASN  267 (439)
T ss_dssp             EEEEEBTTC-TTTTEEEEEE---TTEEE-TT
T ss_pred             hhcCCCCCCCcCCCCchhccCcCccccccCC
Confidence            3678888644445754 34444556666543


No 82 
>cd00216 PQQ_DH Dehydrogenases with pyrrolo-quinoline quinone (PQQ) as cofactor, like ethanol, methanol, and membrane bound glucose dehydrogenases. The alignment model contains an 8-bladed beta-propeller.
Probab=27.61  E-value=1.9e+02  Score=30.66  Aligned_cols=71  Identities=18%  Similarity=0.348  Sum_probs=42.6

Q ss_pred             CCCcEEEEecCC-------CCCCCCccEEEE-ecCCcEEEEeCCCCceEEeecCC-CC--------c---------cEEE
Q 043869           69 PEKTVVWTANRD-------NPPVSSNATLMF-NSEGRIVLRSGEQGQNSIIADNS-QS--------A---------SSAS  122 (450)
Q Consensus        69 ~~~tvVW~ANr~-------~Pv~~~~~~L~l-~~~G~LvL~d~~~g~~vW~st~~-~~--------~---------~~a~  122 (450)
                      ...+++|..+-.       .|+. .+.++.+ +.+|.|+-.|..+|.++| +... ..        .         ....
T Consensus        37 ~~~~~~W~~~~~~~~~~~~sPvv-~~g~vy~~~~~g~l~AlD~~tG~~~W-~~~~~~~~~~~~~~~~~~g~~~~~~~~V~  114 (488)
T cd00216          37 KKLKVAWTFSTGDERGQEGTPLV-VDGDMYFTTSHSALFALDAATGKVLW-RYDPKLPADRGCCDVVNRGVAYWDPRKVF  114 (488)
T ss_pred             hcceeeEEEECCCCCCcccCCEE-ECCEEEEeCCCCcEEEEECCCChhhc-eeCCCCCccccccccccCCcEEccCCeEE
Confidence            345678887643       3555 3455555 457988877764789999 4322 10        0         0111


Q ss_pred             E-ecCCCeEEEec-CCeeEEe
Q 043869          123 M-LDSGSFVLHNS-DGKVIWQ  141 (450)
Q Consensus       123 L-ldsGNlVL~~~-~~~~lWQ  141 (450)
                      + ..+|.++-+|. +++.+|+
T Consensus       115 v~~~~g~v~AlD~~TG~~~W~  135 (488)
T cd00216         115 FGTFDGRLVALDAETGKQVWK  135 (488)
T ss_pred             EecCCCeEEEEECCCCCEeee
Confidence            1 23677777776 6889999


No 83 
>PF05545 FixQ:  Cbb3-type cytochrome oxidase component FixQ;  InterPro: IPR008621 This family consists of several Cbb3-type cytochrome oxidase components (FixQ/CcoQ). FixQ is found in nitrogen fixing bacteria. Since nitrogen fixation is an energy-consuming process, effective symbioses depend on operation of a respiratory chain with a high affinity for O2, closely coupled to ATP production. This requirement is fulfilled by a special three-subunit terminal oxidase (cytochrome terminal oxidase cbb3), which was first identified in Bradyrhizobium japonicum as the product of the fixNOQP operon [].
Probab=27.59  E-value=21  Score=25.24  Aligned_cols=14  Identities=0%  Similarity=-0.249  Sum_probs=6.7

Q ss_pred             eeeeEeeccceeec
Q 043869          432 GIFIYRYRVGSYRR  445 (450)
Q Consensus       432 ~~~~~r~~~~~~~~  445 (450)
                      +|..++++++++.+
T Consensus        27 ~w~~~~~~k~~~e~   40 (49)
T PF05545_consen   27 IWAYRPRNKKRFEE   40 (49)
T ss_pred             HHHHcccchhhHHH
Confidence            33344455555544


No 84 
>PF11403 Yeast_MT:  Yeast metallothionein;  InterPro: IPR022710  Metallothioneins are characterised by an abundance of cysteine residues and a lack of generic secondary structure motifs. This protein functions in primary metal storage, transport and detoxification []. For the first 40 residues in the protein the polypeptide wraps around the metal by forming two large parallel loops separated by a deep cleft containing the metal cluster []. ; PDB: 1AQS_A 1AQR_A 1RJU_V 1FMY_A 1AOO_A 1AQQ_A.
Probab=26.30  E-value=47  Score=21.58  Aligned_cols=19  Identities=37%  Similarity=0.931  Sum_probs=9.2

Q ss_pred             CCCCCCcccccCCCCCCCcCCCCC
Q 043869          271 GLCGFNSFCVLNDQTPNCTCLPGF  294 (450)
Q Consensus       271 g~CG~~g~C~~~~~~~~C~C~~GF  294 (450)
                      |.|-.|.-|     ...|+||.|-
T Consensus        12 gscknneqc-----qkscscptgc   30 (40)
T PF11403_consen   12 GSCKNNEQC-----QKSCSCPTGC   30 (40)
T ss_dssp             STTTT-TTS-----TTS-SS-TTT
T ss_pred             CCccChHHH-----hhcCCCCCCC
Confidence            444444444     4579998764


No 85 
>PF12458 DUF3686:  ATPase involved in DNA repair ;  InterPro: IPR020958  This entry represents an N-terminal domain associated with ATPases and some uncharacterised proteins; it is approximately 450 amino acids in length and contains two conserved sequence motifs: DVF and SPNGED. 
Probab=25.91  E-value=1.7e+02  Score=30.60  Aligned_cols=55  Identities=24%  Similarity=0.392  Sum_probs=30.4

Q ss_pred             eEEeCCC-eEEEEEEeCCCCCeeEEEEEEeecCCCcEEEEecCCCCCCCCccEEEEecCCcEEEEeC
Q 043869           38 SWRSPSG-LYAFGFYPQRNGSRYYVGVFLAGIPEKTVVWTANRDNPPVSSNATLMFNSEGRIVLRSG  103 (450)
Q Consensus        38 ~l~S~~g-~F~lGF~~~~~~~~~~lgIw~~~~~~~tvVW~ANr~~Pv~~~~~~L~l~~~G~LvL~d~  103 (450)
                      .+.|||| ++-.-||.+..+  .|+=+-|+-|...    +   .+|+...  -..+.+||.|++..+
T Consensus       312 ~vrSPNGEDvLYvF~~~~~g--~~~Ll~YN~I~k~----v---~tPi~ch--G~alf~DG~l~~fra  367 (448)
T PF12458_consen  312 KVRSPNGEDVLYVFYAREEG--RYLLLPYNLIRKE----V---ATPIICH--GYALFEDGRLVYFRA  367 (448)
T ss_pred             EecCCCCceEEEEEEECCCC--cEEEEechhhhhh----h---cCCeecc--ceeEecCCEEEEEec
Confidence            4567777 455556665555  4555556554322    1   2466522  245667777777654


No 86 
>PF00558 Vpu:  Vpu protein;  InterPro: IPR008187 The Human immunodeficiency virus 1 (HIV-1) Vpu protein acts in the degradation of CD4 in the endoplasmic reticulum and in the enhancement of virion release from the plasma membrane of infected cells [].; GO: 0019076 release of virus from host; PDB: 2JPX_A 1PI8_A 2GOH_A 2GOF_A 1PI7_A 1PJE_A 1VPU_A 2K7Y_A.
Probab=24.71  E-value=36  Score=27.01  Aligned_cols=6  Identities=33%  Similarity=0.468  Sum_probs=0.4

Q ss_pred             ceeecc
Q 043869          441 GSYRRI  446 (450)
Q Consensus       441 ~~~~~~  446 (450)
                      +|.+||
T Consensus        34 ~rqrkI   39 (81)
T PF00558_consen   34 KRQRKI   39 (81)
T ss_dssp             -----C
T ss_pred             HHHHhH
Confidence            333443


No 87 
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=24.65  E-value=27  Score=28.58  Aligned_cols=28  Identities=18%  Similarity=0.087  Sum_probs=18.1

Q ss_pred             EehhhHHHHHHHHHHhheeeeEeeccce
Q 043869          415 IICLFVTVVILISVVTFGIFIYRYRVGS  442 (450)
Q Consensus       415 i~~~~~~~~~l~~~~~~~~~~~r~~~~~  442 (450)
                      |.++++++++++.+++++.+++..+++|
T Consensus        68 iagi~vg~~~~v~~lv~~l~w~f~~r~k   95 (96)
T PTZ00382         68 IAGISVAVVAVVGGLVGFLCWWFVCRGK   95 (96)
T ss_pred             EEEEEeehhhHHHHHHHHHhheeEEeec
Confidence            4556667777666677666677777543


No 88 
>KOG1214 consensus Nidogen and related basement membrane protein proteins [Cell wall/membrane/envelope biogenesis; Extracellular structures]
Probab=23.99  E-value=55  Score=36.83  Aligned_cols=31  Identities=23%  Similarity=0.657  Sum_probs=24.4

Q ss_pred             CCCCCcCCCCCCcccccCCCCCCCcCCCCCee
Q 043869          265 EKCDPIGLCGFNSFCVLNDQTPNCTCLPGFVA  296 (450)
Q Consensus       265 ~~C~~~g~CG~~g~C~~~~~~~~C~C~~GF~~  296 (450)
                      |+|. +..|-++..|.+..+.-.|.|-|||.-
T Consensus       828 DeC~-psrChp~A~CyntpgsfsC~C~pGy~G  858 (1289)
T KOG1214|consen  828 DECS-PSRCHPAATCYNTPGSFSCRCQPGYYG  858 (1289)
T ss_pred             cccC-ccccCCCceEecCCCcceeecccCccC
Confidence            5666 788999999987555778999998863


No 89 
>PRK12785 fliL flagellar basal body-associated protein FliL; Reviewed
Probab=23.13  E-value=95  Score=28.02  Aligned_cols=18  Identities=22%  Similarity=0.312  Sum_probs=9.0

Q ss_pred             HHHHHHHHHhheeeeEee
Q 043869          421 TVVILISVVTFGIFIYRY  438 (450)
Q Consensus       421 ~~~~l~~~~~~~~~~~r~  438 (450)
                      .+++++.+..+.||+...
T Consensus        32 ~~lll~~~g~g~~f~~~~   49 (166)
T PRK12785         32 AAVLLLGGGGGGFFFFFS   49 (166)
T ss_pred             HHHHHHhcchheEEEEEe
Confidence            344444444556665553


No 90 
>PF02237 BPL_C:  Biotin protein ligase C terminal domain;  InterPro: IPR003142 This C-terminal domain has an SH3-like barrel fold, the function of which is unknown. It is found associated with prokaryotic bifunctional transcriptional repressors [] and eukaryotic enzymes involved in biotin utilization [, ].   In Escherichia coli the biotin operon repressor (BirA) is a bifunctional protein. BirA acts both as the acetyl-coA carboxylase biotin holoenzyme synthetase (6.3.4.15 from EC) and as the biotin operon repressor. DNA sequence analysis of mutations indicates that the helix-turn-helix DNA binding region is located at the N terminus while mutations affecting enzyme function, although mapping over a large region, are found mainly in the central part of the protein's primary sequence [].; GO: 0006464 protein modification process; PDB: 3RUX_A 2CGH_A 3L1A_B 3L2Z_A 1HXD_A 1BIB_A 2EWN_B 1BIA_A 2EJ9_A 3FJP_A ....
Probab=22.13  E-value=1.1e+02  Score=21.31  Aligned_cols=14  Identities=14%  Similarity=0.427  Sum_probs=6.9

Q ss_pred             EEEecCCcEEEEeC
Q 043869           90 LMFNSEGRIVLRSG  103 (450)
Q Consensus        90 L~l~~~G~LvL~d~  103 (450)
                      .-++++|.|+|...
T Consensus        20 ~gId~~G~L~v~~~   33 (48)
T PF02237_consen   20 EGIDDDGALLVRTE   33 (48)
T ss_dssp             EEEETTSEEEEEET
T ss_pred             EEECCCCEEEEEEC
Confidence            33455555555444


No 91 
>PF06247 Plasmod_Pvs28:  Plasmodium ookinete surface protein Pvs28;  InterPro: IPR010423 This family consists of several ookinete surface protein (Pvs28) from several species of Plasmodium. Pvs25 and Pvs28 are expressed on the surface of ookinetes. These proteins are potential candidates for vaccine and induce antibodies that block the infectivity of Plasmodium vivax in immunised animals [].; GO: 0009986 cell surface, 0016020 membrane; PDB: 1Z3G_B 1Z1Y_B 1Z27_A.
Probab=22.09  E-value=23  Score=32.66  Aligned_cols=41  Identities=24%  Similarity=0.619  Sum_probs=25.5

Q ss_pred             CCCCCC----cCCCCCCcccccCCC-----CCCCcCCCCCeeccCCCCCCCCccC
Q 043869          264 SEKCDP----IGLCGFNSFCVLNDQ-----TPNCTCLPGFVAISKGNWTAGCERN  309 (450)
Q Consensus       264 ~~~C~~----~g~CG~~g~C~~~~~-----~~~C~C~~GF~~~~~~~~s~GC~r~  309 (450)
                      ...|+.    .-.||.|+.|....+     .-.|.|.+||....     .-|+|.
T Consensus        39 kv~C~~~e~~~K~Cgdya~C~~~~~~~~~~~~~C~C~~gY~~~~-----~vCvp~   88 (197)
T PF06247_consen   39 KVECDKLENVNKPCGDYAKCINQANKGEERAYKCDCINGYILKQ-----GVCVPN   88 (197)
T ss_dssp             ----SG-GGTTSEEETTEEEEE-SSTTSSTSEEEEE-TTEEESS-----SSEEEG
T ss_pred             ceecCcccccCccccchhhhhcCCCcccceeEEEecccCceeeC-----CeEchh
Confidence            345654    568999999976432     34699999999863     347764


No 92 
>PF05568 ASFV_J13L:  African swine fever virus J13L protein;  InterPro: IPR008385 This family consists of several African swine fever virus (ASFV) j13L proteins [, , ].
Probab=21.84  E-value=17  Score=31.95  Aligned_cols=16  Identities=6%  Similarity=-0.099  Sum_probs=7.1

Q ss_pred             eeeeEeeccceeeccc
Q 043869          432 GIFIYRYRVGSYRRIQ  447 (450)
Q Consensus       432 ~~~~~r~~~~~~~~~~  447 (450)
                      ++++.+||+|+-.-|.
T Consensus        49 i~lcssRKkKaaAAi~   64 (189)
T PF05568_consen   49 IYLCSSRKKKAAAAIE   64 (189)
T ss_pred             HHHHhhhhHHHHhhhh
Confidence            3444444445444443


No 93 
>PRK01844 hypothetical protein; Provisional
Probab=20.46  E-value=11  Score=29.07  Aligned_cols=25  Identities=36%  Similarity=0.537  Sum_probs=13.7

Q ss_pred             hhHHHHHHHHHHhheeeeEeeccce
Q 043869          418 LFVTVVILISVVTFGIFIYRYRVGS  442 (450)
Q Consensus       418 ~~~~~~~l~~~~~~~~~~~r~~~~~  442 (450)
                      +++.++.+++.++++||+.||.-++
T Consensus         7 I~l~I~~li~G~~~Gff~ark~~~k   31 (72)
T PRK01844          7 ILVGVVALVAGVALGFFIARKYMMN   31 (72)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHHHH
Confidence            3444445555556667776665433


No 94 
>TIGR03503 conserved hypothetical protein TIGR03503. This set of conserved hypothetical protein has a phylogenetic range that closely matches that of TIGR03501, a putative C-terminal protein targeting signal.
Probab=20.39  E-value=52  Score=33.82  Aligned_cols=21  Identities=29%  Similarity=0.518  Sum_probs=11.5

Q ss_pred             hHHHHHHHHHHhheeeeEeecc
Q 043869          419 FVTVVILISVVTFGIFIYRYRV  440 (450)
Q Consensus       419 ~~~~~~l~~~~~~~~~~~r~~~  440 (450)
                      +.++++ +++.+.+|+++|||+
T Consensus       353 ~~N~v~-lllg~~~~~~~rk~k  373 (374)
T TIGR03503       353 VGNVVI-LLLGGIGFFVWRKKK  373 (374)
T ss_pred             hhhhhh-hhhheeeEEEEEEee
Confidence            434444 444556677777664


No 95 
>PF14316 DUF4381:  Domain of unknown function (DUF4381)
Probab=20.36  E-value=23  Score=31.17  Aligned_cols=11  Identities=45%  Similarity=0.767  Sum_probs=5.3

Q ss_pred             eEeeccceeec
Q 043869          435 IYRYRVGSYRR  445 (450)
Q Consensus       435 ~~r~~~~~~~~  445 (450)
                      .+|+|+.+|+|
T Consensus        42 ~r~~~~~~yrr   52 (146)
T PF14316_consen   42 WRRWRRNRYRR   52 (146)
T ss_pred             HHHHHccHHHH
Confidence            33444455654


No 96 
>KOG4289 consensus Cadherin EGF LAG seven-pass G-type receptor [Signal transduction mechanisms]
Probab=20.21  E-value=50  Score=39.45  Aligned_cols=40  Identities=28%  Similarity=0.562  Sum_probs=28.4

Q ss_pred             cCCCCCCcccccCCCCCCCcCCCCCeeccCC--CCCCCCccC
Q 043869          270 IGLCGFNSFCVLNDQTPNCTCLPGFVAISKG--NWTAGCERN  309 (450)
Q Consensus       270 ~g~CG~~g~C~~~~~~~~C~C~~GF~~~~~~--~~s~GC~r~  309 (450)
                      .+.||++|-|..-....+|.|-|||.-..=+  ..+.-|++.
T Consensus      1244 s~pC~nng~C~srEggYtCeCrpg~tGehCEvs~~agrCvpG 1285 (2531)
T KOG4289|consen 1244 SGPCGNNGRCRSREGGYTCECRPGFTGEHCEVSARAGRCVPG 1285 (2531)
T ss_pred             cCCCCCCCceEEecCceeEEecCCccccceeeecccCccccc
Confidence            6899999999874557889999999643211  134556654


No 97 
>cd05852 Ig5_Contactin-1 Fifth Ig domain of contactin-1. Ig5_Contactin-1: fifth Ig domain of the neural cell adhesion molecule contactin-1. Contactins are comprised of six Ig domains followed by four fibronectin type III (FnIII) domains anchored to the membrane by glycosylphosphatidylinositol. Contactin-1 is differentially expressed in tumor tissues and may through a RhoA mechanism, facilitate invasion and metastasis of human lung adenocarcinoma.
Probab=20.05  E-value=1.1e+02  Score=23.14  Aligned_cols=34  Identities=15%  Similarity=0.332  Sum_probs=22.8

Q ss_pred             cCCCcEEEEecCCCCCCCCccEEEEecCCcEEEEeC
Q 043869           68 IPEKTVVWTANRDNPPVSSNATLMFNSEGRIVLRSG  103 (450)
Q Consensus        68 ~~~~tvVW~ANr~~Pv~~~~~~L~l~~~G~LvL~d~  103 (450)
                      .|..++.|.=+. .++. .+....+..+|.|.|.+.
T Consensus        13 ~P~p~v~W~k~~-~~l~-~~~r~~~~~~g~L~I~~v   46 (73)
T cd05852          13 APKPKFSWSKGT-ELLV-NNSRISIWDDGSLEILNI   46 (73)
T ss_pred             eCCCEEEEEeCC-Eecc-cCCCEEEcCCCEEEECcC
Confidence            466788898654 3444 345677777888888654


No 98 
>PTZ00208 65 kDa invariant surface glycoprotein; Provisional
Probab=20.02  E-value=48  Score=34.10  Aligned_cols=30  Identities=20%  Similarity=0.371  Sum_probs=16.8

Q ss_pred             ceEEEEehhhHHHHHHHHHHhheee-eEeec
Q 043869          410 WKNIVIICLFVTVVILISVVTFGIF-IYRYR  439 (450)
Q Consensus       410 ~~~i~i~~~~~~~~~l~~~~~~~~~-~~r~~  439 (450)
                      +..+||.++.+-+++|.+++...|. ++|||
T Consensus       384 ~~~~i~~avl~p~~il~~~~~~~~~~v~rrr  414 (436)
T PTZ00208        384 RTAMIILAVLVPAIILAIIAVAFFIMVKRRR  414 (436)
T ss_pred             hhHHHHHHHHHHHHHHHHHHHHhheeeeecc
Confidence            3456677777777776655443333 44444


Done!