Query         038524
Match_columns 175
No_of_seqs    153 out of 833
Neff          6.2 
Searched_HMMs 46136
Date          Fri Mar 29 11:57:44 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/038524.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/038524hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 cd06899 lectin_legume_LecRK_Ar 100.0 2.4E-28 5.3E-33  203.3  11.3  102    5-114   124-235 (236)
  2 PF00139 Lectin_legB:  Legume l  99.9 1.8E-27 3.9E-32  197.5   9.0  103    5-112   125-236 (236)
  3 cd01951 lectin_L-type legume l  99.9 3.2E-22 6.9E-27  164.1  11.5  102    5-112   111-222 (223)
  4 cd07308 lectin_leg-like legume  98.0 2.9E-05 6.2E-10   63.6   7.9   59   52-112   152-216 (218)
  5 cd06902 lectin_ERGIC-53_ERGL E  97.5 0.00042 9.1E-09   57.7   7.9   59   52-112   155-222 (225)
  6 cd06903 lectin_EMP46_EMP47 EMP  97.0  0.0047   1E-07   51.2   8.2   59   50-111   149-211 (215)
  7 PF03388 Lectin_leg-like:  Legu  96.5  0.0097 2.1E-07   49.5   6.7   55   51-109   158-223 (229)
  8 cd06901 lectin_VIP36_VIPL VIP3  96.3   0.021 4.5E-07   48.3   8.0   59   51-112   156-221 (248)
  9 KOG3839 Lectin VIP36, involved  95.6   0.022 4.8E-07   50.2   5.2   57   55-111   212-272 (351)
 10 PF15065 NCU-G1:  Lysosomal tra  95.3   0.011 2.3E-07   52.6   2.1   26   89-114   281-306 (350)
 11 KOG3838 Mannose lectin ERGIC-5  94.6    0.11 2.4E-06   47.0   6.6   60   53-114   188-256 (497)
 12 PF08693 SKG6:  Transmembrane a  93.4   0.069 1.5E-06   33.1   2.1    8  157-164    32-39  (40)
 13 PF04478 Mid2:  Mid2 like cell   89.7    0.22 4.7E-06   39.5   1.9   28  138-165    49-79  (154)
 14 PF15102 TMEM154:  TMEM154 prot  89.5    0.34 7.3E-06   38.1   2.8   17  138-154    58-74  (146)
 15 PF01102 Glycophorin_A:  Glycop  87.4    0.34 7.3E-06   37.0   1.6   11  156-166    84-94  (122)
 16 PF14610 DUF4448:  Protein of u  85.3    0.76 1.6E-05   36.9   2.7   26  138-163   159-184 (189)
 17 cd06900 lectin_VcfQ VcfQ bacte  85.0     4.2   9E-05   34.7   7.1   28   83-110   225-252 (255)
 18 PF05454 DAG1:  Dystroglycan (D  83.8    0.33 7.2E-06   42.2   0.0   24  141-164   151-174 (290)
 19 PF01299 Lamp:  Lysosome-associ  83.2    0.81 1.8E-05   39.5   2.1   30  138-167   270-301 (306)
 20 PTZ00382 Variant-specific surf  77.9    0.95 2.1E-05   33.0   0.7   14  136-149    64-77  (96)
 21 PF14575 EphA2_TM:  Ephrin type  76.6     1.4 2.9E-05   30.7   1.1   11  155-165    18-28  (75)
 22 PF06697 DUF1191:  Protein of u  75.9     4.8  0.0001   34.9   4.5   15  139-153   215-229 (278)
 23 PF14283 DUF4366:  Domain of un  75.7      11 0.00023   31.6   6.4   29   41-71     72-100 (218)
 24 PF14991 MLANA:  Protein melan-  75.3    0.73 1.6E-05   34.8  -0.5   15  155-169    42-56  (118)
 25 PHA03265 envelope glycoprotein  67.7     4.6 9.9E-05   36.3   2.6   28  136-164   349-376 (402)
 26 PF03302 VSP:  Giardia variant-  65.2     4.7  0.0001   36.3   2.2   19  136-154   365-383 (397)
 27 PF06024 DUF912:  Nucleopolyhed  61.8     5.6 0.00012   29.1   1.8   10  155-164    81-90  (101)
 28 PF12877 DUF3827:  Domain of un  59.6      14 0.00031   35.6   4.4   18  137-154   269-286 (684)
 29 PF02009 Rifin_STEVOR:  Rifin/s  56.3       5 0.00011   35.0   0.8    7  155-161   274-280 (299)
 30 PF11857 DUF3377:  Domain of un  54.2      20 0.00043   25.1   3.4   18  137-154    30-47  (74)
 31 PF12768 Rax2:  Cortical protei  49.8      14  0.0003   31.9   2.5    9  138-146   229-237 (281)
 32 PF15345 TMEM51:  Transmembrane  48.3      16 0.00035   30.9   2.5   10  157-166    77-86  (233)
 33 PTZ00208 65 kDa invariant surf  45.5     9.3  0.0002   34.9   0.8   30  139-168   388-418 (436)
 34 PF05337 CSF-1:  Macrophage col  44.1     7.5 0.00016   33.7   0.0   29  137-165   224-253 (285)
 35 PF12191 stn_TNFRSF12A:  Tumour  43.2     7.9 0.00017   29.8   0.0   24  139-162    79-103 (129)
 36 PLN03150 hypothetical protein;  42.6      25 0.00055   33.3   3.3    6  158-163   566-571 (623)
 37 PF03229 Alpha_GJ:  Alphavirus   42.2      40 0.00087   25.7   3.6   28  137-164    82-111 (126)
 38 PF06365 CD34_antigen:  CD34/Po  41.9      46 0.00099   27.6   4.3   26  138-163   101-128 (202)
 39 TIGR01167 LPXTG_anchor LPXTG-m  41.1      31 0.00067   19.5   2.4    6  159-164    27-32  (34)
 40 PF02480 Herpes_gE:  Alphaherpe  40.7     9.1  0.0002   35.1   0.0    9  105-113   293-301 (439)
 41 PF15050 SCIMP:  SCIMP protein   40.7     9.9 0.00021   29.2   0.2    8  155-162    26-33  (133)
 42 TIGR01478 STEVOR variant surfa  39.0      20 0.00043   31.3   1.8    7  158-164   281-287 (295)
 43 PTZ00370 STEVOR; Provisional    38.8      20 0.00044   31.3   1.8    7  158-164   277-283 (296)
 44 TIGR03370 PEPCTERM_Roseo varia  37.7      20 0.00043   20.1   1.1   11  155-165    15-25  (26)
 45 PF13268 DUF4059:  Protein of u  35.7      22 0.00048   24.7   1.3   24  143-166    13-36  (72)
 46 PF08374 Protocadherin:  Protoc  34.8      65  0.0014   27.0   4.1   17  138-154    38-54  (221)
 47 PF12248 Methyltransf_FA:  Farn  33.1 1.8E+02  0.0039   20.7   5.9   41   49-94     49-94  (102)
 48 PF13908 Shisa:  Wnt and FGF in  32.3      41 0.00088   26.5   2.5   11  138-148    79-89  (179)
 49 PF03597 CcoS:  Cytochrome oxid  31.7      36 0.00077   21.4   1.6   27  142-168     6-32  (45)
 50 PRK05886 yajC preprotein trans  31.6      35 0.00076   25.5   1.9   11  155-165    17-27  (109)
 51 PF11353 DUF3153:  Protein of u  31.5      35 0.00076   27.8   2.0   25   73-97    117-142 (209)
 52 PF14914 LRRC37AB_C:  LRRC37A/B  31.2      29 0.00062   27.5   1.4   18  137-154   119-136 (154)
 53 PTZ00046 rifin; Provisional     30.5      33 0.00071   30.8   1.8    6  155-160   333-338 (358)
 54 TIGR02595 PEP_exosort PEP-CTER  30.0      45 0.00097   18.4   1.7    7  158-164    17-23  (26)
 55 TIGR03141 cytochro_ccmD heme e  29.9      36 0.00077   21.1   1.4   11  142-152     9-19  (45)
 56 TIGR03521 GldG gliding-associa  29.2      36 0.00079   31.9   2.0   12  155-166   540-551 (552)
 57 TIGR01582 FDH-beta formate deh  29.0      71  0.0015   27.6   3.6   11  120-130   228-238 (283)
 58 COG4736 CcoQ Cbb3-type cytochr  28.9      56  0.0012   21.9   2.3   10  156-165    26-35  (60)
 59 PF04689 S1FA:  DNA binding pro  28.4      67  0.0015   22.0   2.6   19  136-154    11-29  (69)
 60 PF10661 EssA:  WXG100 protein   26.2      43 0.00092   26.1   1.6   21  141-161   121-141 (145)
 61 PHA03264 envelope glycoprotein  25.8      73  0.0016   29.0   3.2   18  136-153   359-376 (416)
 62 PRK07021 fliL flagellar basal   25.5      69  0.0015   25.0   2.7    8  155-162    35-42  (162)
 63 PHA03291 envelope glycoprotein  24.2      58  0.0013   29.4   2.2   25  138-162   288-313 (401)
 64 PRK15471 chain length determin  23.5      60  0.0013   28.5   2.2    7  136-142   293-299 (325)
 65 PF06809 NPDC1:  Neural prolife  23.1 1.4E+02  0.0029   26.7   4.2    7   89-95    144-150 (341)
 66 PRK10381 LPS O-antigen length   23.1      46   0.001   29.8   1.4    9  136-144   337-345 (377)
 67 PLN00113 leucine-rich repeat r  23.0      89  0.0019   30.6   3.5   12  139-150   630-641 (968)
 68 PRK00523 hypothetical protein;  22.5      53  0.0011   22.9   1.3    8  155-162    21-28  (72)
 69 PF03988 DUF347:  Repeat of Unk  22.0      62  0.0014   20.9   1.5   16  139-154    25-40  (55)
 70 KOG3637 Vitronectin receptor,   21.4 1.1E+02  0.0024   31.2   3.8   14  138-151   979-992 (1030)
 71 COG4282 SMI1 Protein involved   21.2      62  0.0013   26.4   1.6   15   10-24    123-137 (191)
 72 PRK11638 lipopolysaccharide bi  21.2      43 0.00094   29.6   0.8    8  157-164   334-341 (342)
 73 PF02699 YajC:  Preprotein tran  20.7      73  0.0016   22.2   1.8    7  156-162    16-22  (82)
 74 TIGR00847 ccoS cytochrome oxid  20.3 1.1E+02  0.0023   19.8   2.4   13  155-167    20-32  (51)
 75 PF07332 DUF1469:  Protein of u  20.3      56  0.0012   23.8   1.2    8  167-174   109-116 (121)
 76 TIGR01006 polys_exp_MPA1 polys  20.1      50  0.0011   26.7   0.9   14  158-171   197-210 (226)

No 1  
>cd06899 lectin_legume_LecRK_Arcelin_ConA legume lectins, lectin-like receptor kinases, arcelin, concanavalinA, and alpha-amylase inhibitor. This alignment model includes the legume lectins (also known as agglutinins), the arcelin (also known as phytohemagglutinin-L) family of lectin-like defense proteins, the LecRK family of lectin-like receptor kinases, concanavalinA (ConA), and an alpha-amylase inhibitor.  Arcelin is a major seed glycoprotein discovered in kidney beans (Phaseolus vulgaris) that has insecticidal properties and protects the seeds from predation by larvae of various bruchids.  Arcelin is devoid of monosaccharide binding properties and lacks a key metal-binding loop that is present in other members of this family.  Phytohaemagglutinin (PHA) is a lectin found in plants, especially beans, that affects cell metabolism by inducing mitosis and by altering the permeability of the cell membrane to various proteins.  PHA agglutinates most mammalian red blood cell types by bindin
Probab=99.95  E-value=2.4e-28  Score=203.25  Aligned_cols=102  Identities=47%  Similarity=0.685  Sum_probs=90.5

Q ss_pred             eeecCCCCCCCCCeeEEEcCCCccccceecceecCCCCCcccccccCCCcEEEEEEeeCCCceEEEee----------ee
Q 038524            5 RYQSIQFNDTSNNHVGIDVNRLTSVGSVPATYFSDEKGTNKSLKLISGDPMQTWIDYMGSEKLHEIQC----------LS   74 (175)
Q Consensus         5 T~~n~e~~D~~~nHVGIdiNs~~S~~s~~~~~~~~~~~~~~~~~l~sG~~~~vwIdYd~~~~~L~v~l----------Pl   74 (175)
                      |++|.+++||++||||||+|++.|..+..   |+.     ..+.|.+|+.++|||+||+.+++|+|.+          |+
T Consensus       124 T~~n~~~~D~~~nHigIdvn~~~S~~~~~---~~~-----~~~~l~~g~~~~v~I~Y~~~~~~L~V~l~~~~~~~~~~~~  195 (236)
T cd06899         124 TFQNPEFGDPDDNHVGIDVNSLVSVKAGY---WDD-----DGGKLKSGKPMQAWIDYDSSSKRLSVTLAYSGVAKPKKPL  195 (236)
T ss_pred             cccCcccCCCCCCeEEEEcCCcccceeec---ccc-----ccccccCCCeEEEEEEEcCCCCEEEEEEEeCCCCCCcCCE
Confidence            67888889999999999999987776544   322     2345789999999999999999999988          68


Q ss_pred             eeeeecCCccCCCceEEEEEeecCCceeeeEEeeeEEeeC
Q 038524           75 LSTSVDLSQLLLDTMCVGFSAATGSLATEHYILGCSLNKS  114 (175)
Q Consensus        75 ls~~idLs~~l~~~~yVGFSAsTG~~~~~h~IlsWsF~~~  114 (175)
                      |+.++||+.+|+++||||||||||...|.|+|++|+|++.
T Consensus       196 ls~~vdL~~~l~~~~~vGFSasTG~~~~~h~i~sWsF~s~  235 (236)
T cd06899         196 LSYPVDLSKVLPEEVYVGFSASTGLLTELHYILSWSFSSN  235 (236)
T ss_pred             EEEeccHHHhCCCceEEEEEeEcCCCcceEEEEEEEEEcC
Confidence            9999999999999999999999999999999999999875


No 2  
>PF00139 Lectin_legB:  Legume lectin domain;  InterPro: IPR001220 Legume lectins are one of the largest lectin families with more than 70 lectins reported. Leguminous plant lectins resemble each other in their physicochemical properties although they differ in their carbohydrate specificities. They consist of two or four subunits with relative molecular mass of 30 kDa and each subunit has one carbohydrate-binding site. The interaction with sugars requires tightly bound calcium and manganese ions. The structural similarities of these lectins are reported by the primary structural analyses and X-ray crystallographic studies. X-ray studies have shown that the folding of the polypeptide chains in the region of the carbohydrate-binding sites is also similar, despite differences in the primary sequences. The carbohydrate-binding sites of these lectins consist of two conserved amino acids on beta pleated sheets. One of these loops contains transition metals, calcium and manganese, which keep the amino acid residues of the sugar-binding site at the required positions. Amino acid sequences of this loop play an important role in the carbohydrate-binding specificities of these lectins. These lectins bind either glucose/mannose or galactose. The exact function of legume lectins is not known but they may be involved in the attachment of nitrogen-fixing bacteria to legumes and in the protection against pathogens. Some legume lectins are proteolytically processed to produce two chains, beta (which corresponds to the N-terminal) and alpha (C-terminal) (IPR000985 from INTERPRO). The lectin concanavalin A (conA) from jack bean is exceptional in that the two chains are transposed and ligated (by formation of a new peptide bond). The N terminus of mature conA thus corresponds to that of the alpha chain and the C terminus to the beta chain.; GO: 0005488 binding; PDB: 1VLN_B 2GDF_C 2JE9_C 2JEC_C 1DGL_B 2P37_B 2CWM_A 2P34_D 2OW4_A 3IPV_B ....
Probab=99.94  E-value=1.8e-27  Score=197.46  Aligned_cols=103  Identities=40%  Similarity=0.603  Sum_probs=89.9

Q ss_pred             eeecCCCCCCCCCeeEEEcCCCccccceecceecCCCCCcccccccCCCcEEEEEEeeCCCceEEEee---------eee
Q 038524            5 RYQSIQFNDTSNNHVGIDVNRLTSVGSVPATYFSDEKGTNKSLKLISGDPMQTWIDYMGSEKLHEIQC---------LSL   75 (175)
Q Consensus         5 T~~n~e~~D~~~nHVGIdiNs~~S~~s~~~~~~~~~~~~~~~~~l~sG~~~~vwIdYd~~~~~L~v~l---------Pll   75 (175)
                      |++|++++||++||||||+|++.|..+.+++++.     .....|.+|+.++|||+||+.+++|.|++         |+|
T Consensus       125 T~~N~~~~d~~~nHIgI~~n~~~s~~~~~~~~~~-----~~~~~l~~g~~~~v~I~Yd~~~~~L~V~l~~~~~~~~~~~l  199 (236)
T PF00139_consen  125 TYKNPEYNDPDDNHIGIDVNSVVSNKTASAGYYS-----SPSFSLSDGKWHTVWIDYDASTKRLSVYLDDNSSKPSSPVL  199 (236)
T ss_dssp             TSTCGGGTTTSSSEEEEEESSSSESEEEE----E-----EEEHHHGTTSEEEEEEEEETTTTEEEEEEEETTTTSEEEEE
T ss_pred             eeecccccccCCCEEEEECCCCcccccccccccc-----cccccccCCcEEEEEEEEcCCccEEEEEEecccCCCcceeE
Confidence            5678889999999999999999999988877652     35678999999999999999999999988         689


Q ss_pred             eeeecCCccCCCceEEEEEeecCCceeeeEEeeeEEe
Q 038524           76 STSVDLSQLLLDTMCVGFSAATGSLATEHYILGCSLN  112 (175)
Q Consensus        76 s~~idLs~~l~~~~yVGFSAsTG~~~~~h~IlsWsF~  112 (175)
                      +..+||+.+|+++||||||||||...|.|.|++|+|+
T Consensus       200 ~~~vdL~~~l~~~v~vGFsasTG~~~~~h~I~sW~F~  236 (236)
T PF00139_consen  200 SVNVDLSAVLPEQVYVGFSASTGGSYQTHDILSWSFS  236 (236)
T ss_dssp             EEE--HHHHSCSEEEEEEEEEESSSSEEEEEEEEEEE
T ss_pred             EEEEchHHhcCCCcEEEEEeecCCCcceEEEEEEEeC
Confidence            9999999999999999999999999999999999995


No 3  
>cd01951 lectin_L-type legume lectins. The L-type (legume-type) lectins are a highly diverse family of carbohydrate binding proteins that generally display no enzymatic activity toward the sugars they bind.  This family includes arcelin, concanavalinA, the lectin-like receptor kinases, the ERGIC-53/VIP36/EMP46 type1 transmembrane proteins, and an alpha-amylase inhibitor.  L-type lectins have a dome-shaped beta-barrel carbohydrate recognition domain with a curved seven-stranded beta-sheet referred to as the "front face" and a flat six-stranded beta-sheet referred to as the "back face".  This domain homodimerizes so that adjacent back sheets form a contiguous 12-stranded sheet and homotetramers occur by a back-to-back association of these homodimers.  Though L-type lectins exhibit both sequence and structural similarity to one another, their carbohydrate binding specificities differ widely.
Probab=99.88  E-value=3.2e-22  Score=164.05  Aligned_cols=102  Identities=28%  Similarity=0.309  Sum_probs=84.2

Q ss_pred             eeecCCCCCCCCCeeEEEcCCCcccc--ceecceecCCCCCcccccccCCCcEEEEEEeeCCCceEEEee--------ee
Q 038524            5 RYQSIQFNDTSNNHVGIDVNRLTSVG--SVPATYFSDEKGTNKSLKLISGDPMQTWIDYMGSEKLHEIQC--------LS   74 (175)
Q Consensus         5 T~~n~e~~D~~~nHVGIdiNs~~S~~--s~~~~~~~~~~~~~~~~~l~sG~~~~vwIdYd~~~~~L~v~l--------Pl   74 (175)
                      |++|.+++||+.||||||+|+..+..  ..+.+++.      .+....+|+.++|||+||+.+++|+|.+        |.
T Consensus       111 T~~N~~~~dp~~~higi~~n~~~~~~~~~~~~~~~~------~~~~~~~g~~~~v~I~Y~~~~~~L~v~l~~~~~~~~~~  184 (223)
T cd01951         111 TYKNDDNNDPNGNHISIDVNGNGNNTALATSLGSAS------LPNGTGLGNEHTVRITYDPTTNTLTVYLDNGSTLTSLD  184 (223)
T ss_pred             ccccCCCCCCCCCEEEEEcCCCCCCcccccccceee------CCCccCCCCEEEEEEEEeCCCCEEEEEECCCCcccccc
Confidence            78898888999999999999987641  11222221      1122223999999999999999999988        58


Q ss_pred             eeeeecCCccCCCceEEEEEeecCCceeeeEEeeeEEe
Q 038524           75 LSTSVDLSQLLLDTMCVGFSAATGSLATEHYILGCSLN  112 (175)
Q Consensus        75 ls~~idLs~~l~~~~yVGFSAsTG~~~~~h~IlsWsF~  112 (175)
                      ++.++||+..++++||||||||||...+.|+|++|+|+
T Consensus       185 l~~~~~l~~~~~~~~yvGFTAsTG~~~~~h~V~~wsf~  222 (223)
T cd01951         185 ITIPVDLIQLGPTKAYFGFTASTGGLTNLHDILNWSFT  222 (223)
T ss_pred             EEEeeeecccCCCcEEEEEEcccCCCcceeEEEEEEec
Confidence            89999999999999999999999999999999999996


No 4  
>cd07308 lectin_leg-like legume-like lectins: ERGIC-53, ERGL, VIP36, VIPL, EMP46, and EMP47. The legume-like (leg-like) lectins are eukaryotic intracellular sugar transport proteins with a carbohydrate recognition domain similar to that of the legume lectins.  This domain binds high-mannose-type oligosaccharides for transport from the endoplasmic reticulum to the Golgi complex.  These leg-like lectins include ERGIC-53, ERGL, VIP36, VIPL, EMP46, EMP47, and the UIP5 (ULP1-interacting protein 5) precursor protein.  Leg-like lectins have different intracellular distributions and dynamics in the endoplasmic reticulum-Golgi system of the secretory pathway and interact with N-glycans of glycoproteins in a calcium-dependent manner, suggesting a role in glycoprotein sorting and trafficking.  L-type lectins have a dome-shaped beta-barrel carbohydrate recognition domain with a curved seven-stranded beta-sheet referred to as the "front face" and a flat six-stranded beta-sheet referred to as the "ba
Probab=97.99  E-value=2.9e-05  Score=63.62  Aligned_cols=59  Identities=24%  Similarity=0.389  Sum_probs=44.1

Q ss_pred             CCcEEEEEEeeCCCceEEEee-e----eeeeeecCCc-cCCCceEEEEEeecCCceeeeEEeeeEEe
Q 038524           52 GDPMQTWIDYMGSEKLHEIQC-L----SLSTSVDLSQ-LLLDTMCVGFSAATGSLATEHYILGCSLN  112 (175)
Q Consensus        52 G~~~~vwIdYd~~~~~L~v~l-P----lls~~idLs~-~l~~~~yVGFSAsTG~~~~~h~IlsWsF~  112 (175)
                      +++++++|.|+  .+.|.|.+ +    --..-.++.. .+++.+|+||||+||...+.|.|++|.+.
T Consensus       152 ~~~~~~~I~y~--~~~l~v~i~~~~~~~~~~c~~~~~~~l~~~~y~G~sA~tg~~~d~~dIls~~~~  216 (218)
T cd07308         152 NAPTTLRISYL--NNTLKVDITYSEGNNWKECFTVEDVILPSQGYFGFSAQTGDLSDNHDILSVHTY  216 (218)
T ss_pred             CCCeEEEEEEE--CCEEEEEEeCCCCCCccEEEEcCCcccCCCCEEEEEeccCCCcCcEEEEEEEee
Confidence            68999999999  56777665 1    0111222333 46788999999999999999999999874


No 5  
>cd06902 lectin_ERGIC-53_ERGL ERGIC-53 and ERGL type 1 transmembrane proteins, N-terminal lectin domain. ERGIC-53 and ERGL, N-terminal carbohydrate recognition domain. ERGIC-53 and ERGL are eukaryotic mannose-binding type 1 transmembrane proteins of the early secretory pathway that transport newly synthesized glycoproteins from the endoplasmic reticulum (ER) to the ER-Golgi intermediate compartment (ERGIC).  ERGIC-53 and ERGL have an N-terminal lectin-like carbohydrate recognition domain (represented by this alignment model) as well as a C-terminal transmembrane domain.  ERGIC-53 functions as a 'cargo receptor' to facilitate the export of glycoproteins with different characteristics from the ER, while the ERGIC-53-like protein (ERGL) which may act as a regulator of ERGIC-53.  In mammals, ERGIC-53 forms a complex with MCFD2 (multi-coagulation factor deficiency 2) which then recruits blood coagulation factors V and VIII.  Mutations in either MCFD2 or ERGIC-53 cause a mild form of inherite
Probab=97.53  E-value=0.00042  Score=57.74  Aligned_cols=59  Identities=24%  Similarity=0.293  Sum_probs=42.8

Q ss_pred             CCcEEEEEEeeCCCceEEEee-----e---eeeeeecCCc-cCCCceEEEEEeecCCceeeeEEeeeEEe
Q 038524           52 GDPMQTWIDYMGSEKLHEIQC-----L---SLSTSVDLSQ-LLLDTMCVGFSAATGSLATEHYILGCSLN  112 (175)
Q Consensus        52 G~~~~vwIdYd~~~~~L~v~l-----P---lls~~idLs~-~l~~~~yVGFSAsTG~~~~~h~IlsWsF~  112 (175)
                      ..+.++.|.|...  .|.|.+     +   ....-.++.. .||+..|+||||+||...+.|.|++|+|.
T Consensus       155 ~~p~~~rI~Y~~~--~l~V~~d~~~~~~~~~~~~Cf~~~~v~LP~~~yfGiSA~Tg~l~d~hDIls~~~~  222 (225)
T cd06902         155 PYPVRAKITYYQN--VLTVSINNGFTPNKDDYELCTRVENMVLPPNGYFGVSAATGGLADDHDVLSFLTF  222 (225)
T ss_pred             CCCeEEEEEEECC--eEEEEEeCCcCCCCCcccEEEecCCeeCCCCCEEEEEecCCCCCCcEeEEEEEEe
Confidence            5689999999985  466543     1   0111122222 46778999999999999999999999985


No 6  
>cd06903 lectin_EMP46_EMP47 EMP46 and EMP47 type 1 transmembrane proteins, N-terminal lectin domain. EMP46 and EMP47, N-terminal carbohydrate recognition domain. EMP46 and EMP47 are fungal type-I transmembrane proteins that cycle between the endoplasmic reticulum and the golgi apparatus and are thought to function as cargo receptors that transport newly synthesized glycoproteins.  EMP47 is a receptor for EMP46 responsible for the selective transport of EMP46 by forming hetero-oligomerization between the two proteins. EMP46 and EMP47 have an N-terminal lectin-like carbohydrate recognition domain (represented by this alignment model) as well as a C-terminal transmembrane domain. EMP46 and EMP47 are 45% sequence-identical to one another and have sequence homology to a class of intracellular lectins defined by ERGIC-53 and VIP36.  L-type lectins have a dome-shaped beta-barrel carbohydrate recognition domain with a curved seven-stranded beta-sheet referred to as the "front face" and a flat s
Probab=96.95  E-value=0.0047  Score=51.21  Aligned_cols=59  Identities=27%  Similarity=0.334  Sum_probs=44.1

Q ss_pred             cCCCcEEEEEEeeCCCceEEEee---eeeee-eecCCccCCCceEEEEEeecCCceeeeEEeeeEE
Q 038524           50 ISGDPMQTWIDYMGSEKLHEIQC---LSLST-SVDLSQLLLDTMCVGFSAATGSLATEHYILGCSL  111 (175)
Q Consensus        50 ~sG~~~~vwIdYd~~~~~L~v~l---Plls~-~idLs~~l~~~~yVGFSAsTG~~~~~h~IlsWsF  111 (175)
                      .++.+.++.|.|....+.|.|.+   .++.. .+.|..   ...|+||||+||...+.|.|++-.+
T Consensus       149 n~~~p~~iri~Y~~~~~~l~v~vd~~~Cf~~~~v~lP~---~~y~fGiSAaTg~~~d~hdIl~~~~  211 (215)
T cd06903         149 DSGVPSTIRLSYDALNSLFKVQVDNRLCFQTDKVQLPQ---GGYRFGITAANADNPESFEILKLKV  211 (215)
T ss_pred             CCCCCEEEEEEEECCCCEEEEEECCCEEEecCCeecCC---CCCEEEEEEcCCCCCCcEEEEEEEE
Confidence            45678999999999777888776   33332 333332   4568999999999999999996544


No 7  
>PF03388 Lectin_leg-like:  Legume-like lectin family;  InterPro: IPR005052  Lectins are structurally diverse proteins that bind to specific carbohydrates. This family includes the VIP36 and ERGIC-53 lectins. These two proteins were the first members of the family of animal lectins similar to the leguminous plant lectins []. The alignment for this family is towards the N terminus, where the similarity of VIP36 and ERGIC-53 is greatest. Although they have been identified as a family of animal lectins, this alignment also includes yeast sequences[].  ERGIC-53 is a 53kDa protein, localised to the intermediate region between the endoplasmic reticulum and the Golgi apparatus (ER-Golgi-Intermediate Compartment, ERGIC). It was identified as a calcium-dependent, mannose-specific lectin []. Its dysfunction has been associated with combined factors V and VIII deficiency, suggesting an important and substrate-specific role for ERGIC-53 in the glycoprotein-secreting pathway [,]. The L-type lectin-like domain has an overall globular shape composed of a beta-sandwich of two major twisted antiparallel beta-sheets. The beta-sandwich comprises a major concave beta-sheet and a minor convex beta-sheet, in a variation of the jelly roll fold [, , , ]. ; GO: 0016020 membrane; PDB: 3A4U_A 3LCP_B 2A6Z_A 2A71_C 2A70_B 2A6Y_A 2A6X_A 2A6W_B 2A6V_B 2E6V_B ....
Probab=96.45  E-value=0.0097  Score=49.47  Aligned_cols=55  Identities=35%  Similarity=0.411  Sum_probs=40.6

Q ss_pred             CCCcEEEEEEeeCCCceEEEe---e-------eeeee-eecCCccCCCceEEEEEeecCCceeeeEEeee
Q 038524           51 SGDPMQTWIDYMGSEKLHEIQ---C-------LSLST-SVDLSQLLLDTMCVGFSAATGSLATEHYILGC  109 (175)
Q Consensus        51 sG~~~~vwIdYd~~~~~L~v~---l-------Plls~-~idLs~~l~~~~yVGFSAsTG~~~~~h~IlsW  109 (175)
                      .+.+.++.|.|......+.+.   +       .++.. .+    .||+..|+||||+||...+.|.|++=
T Consensus       158 ~~~p~~~ri~Y~~~~l~v~id~~~~~~~~~~~~Cf~~~~v----~LP~~~yfGvSA~Tg~~~d~hdi~s~  223 (229)
T PF03388_consen  158 SDVPTRIRISYSKNTLTVSIDSNYLKNQDDWELCFTTDGV----DLPEGYYFGVSAATGELSDNHDILSV  223 (229)
T ss_dssp             ESSEEEEEEEEETTEEEEEEETSCCSECCTTEEEEEESTE----EGGSSBEEEEEEEESSSGGEEEEEEE
T ss_pred             CCCCEEEEEEEECCeEEEEEecccccCCcCCcEEEEcCCe----ecCCCCEEEEEecCCCCCCcEEEEEE
Confidence            456789999999875555444   1       34443 23    35778899999999999999999964


No 8  
>cd06901 lectin_VIP36_VIPL VIP36 and VIPL type 1 transmembrane proteins, lectin domain. The vesicular integral protein of 36 kDa (VIP36) is a type 1 transmembrane protein of the mammalian early secretory pathway that acts as a cargo receptor transporting high mannose type glycoproteins between the Golgi and the endoplasmic reticulum (ER).  Lectins of the early secretory pathway are involved in the selective transport of newly synthesized glycoproteins from the ER to the ER-Golgi intermediate compartment (ERGIC). The most prominent cycling lectin is the mannose-binding type1 membrane protein ERGIC-53, which functions as a cargo receptor to facilitate export of glycoproteins from the ER. L-type lectins have a dome-shaped beta-barrel carbohydrate recognition domain with a curved seven-stranded beta-sheet referred to as the "front face" and a flat six-stranded beta-sheet referred to as the "back face".  This domain homodimerizes so that adjacent back sheets form a contiguous 12-stranded she
Probab=96.30  E-value=0.021  Score=48.32  Aligned_cols=59  Identities=22%  Similarity=0.166  Sum_probs=40.4

Q ss_pred             CCCcEEEEEEeeCCCceEEEee-------eeeeeeecCCccCCCceEEEEEeecCCceeeeEEeeeEEe
Q 038524           51 SGDPMQTWIDYMGSEKLHEIQC-------LSLSTSVDLSQLLLDTMCVGFSAATGSLATEHYILGCSLN  112 (175)
Q Consensus        51 sG~~~~vwIdYd~~~~~L~v~l-------Plls~~idLs~~l~~~~yVGFSAsTG~~~~~h~IlsWsF~  112 (175)
                      .+.+.+++|.|......|.+..       .++..   -.-.||...|+||||+||...+.|.|++-.+-
T Consensus       156 ~~~~t~~rI~Y~~~~l~v~vd~~~~~~w~~Cf~~---~~v~LP~~~yfGiSA~Tg~~sd~hdIlsv~~~  221 (248)
T cd06901         156 KDHDTFVAIRYSKGRLTVMTDIDGKNEWKECFDV---TGVRLPTGYYFGASAATGDLSDNHDIISMKLY  221 (248)
T ss_pred             CCCCeEEEEEEECCeEEEEEecCCCCceeeeEEe---CCeecCCCCEEEEEecCCCCCCcEEEEEEEEe
Confidence            4567889999997543333332       12221   12245677999999999999999999976664


No 9  
>KOG3839 consensus Lectin VIP36, involved in the transport of glycoproteins carrying high mannose-type glycans [Intracellular trafficking, secretion, and vesicular transport]
Probab=95.62  E-value=0.022  Score=50.16  Aligned_cols=57  Identities=25%  Similarity=0.205  Sum_probs=39.6

Q ss_pred             EEEEEEeeCCCceEEEee--e-eeeeeecCCc-cCCCceEEEEEeecCCceeeeEEeeeEE
Q 038524           55 MQTWIDYMGSEKLHEIQC--L-SLSTSVDLSQ-LLLDTMCVGFSAATGSLATEHYILGCSL  111 (175)
Q Consensus        55 ~~vwIdYd~~~~~L~v~l--P-lls~~idLs~-~l~~~~yVGFSAsTG~~~~~h~IlsWsF  111 (175)
                      ..+-|.|+..+.++...+  | -...-.+|.. .||.--|+|+||+||...+.|.|++=.+
T Consensus       212 t~~~iry~~~~l~~~~dl~~~~~~~~c~~~n~v~lp~g~~fg~SasTGdlSd~HdivS~kl  272 (351)
T KOG3839|consen  212 TLVVIRYEKKTLSISIDLEGPNEWIDCFSLNNVELPLGYFFGVSASTGDLSDSHDIVSLKL  272 (351)
T ss_pred             ceeEEEecCCceEEEEecCCCceeeeeeeecceecccceEEeeeeccCccchhhHHHHhhh
Confidence            456788888554444444  4 2344455555 4567789999999999999999997543


No 10 
>PF15065 NCU-G1:  Lysosomal transcription factor, NCU-G1
Probab=95.26  E-value=0.011  Score=52.60  Aligned_cols=26  Identities=12%  Similarity=0.123  Sum_probs=23.4

Q ss_pred             eEEEEEeecCCceeeeEEeeeEEeeC
Q 038524           89 MCVGFSAATGSLATEHYILGCSLNKS  114 (175)
Q Consensus        89 ~yVGFSAsTG~~~~~h~IlsWsF~~~  114 (175)
                      +.|=|.++++..+..+..++|+|...
T Consensus       281 ~nvSFG~~gDgfY~~t~ylsWt~~~G  306 (350)
T PF15065_consen  281 LNVSFGTSGDGFYWATNYLSWTFLIG  306 (350)
T ss_pred             EEEEeccCCCCcccccceEEEEEecc
Confidence            77888888888899999999999986


No 11 
>KOG3838 consensus Mannose lectin ERGIC-53, involved in glycoprotein traffic [Intracellular trafficking, secretion, and vesicular transport]
Probab=94.57  E-value=0.11  Score=47.03  Aligned_cols=60  Identities=32%  Similarity=0.419  Sum_probs=42.8

Q ss_pred             CcEEEEEEeeCCCceEEEee-----ee--eeeeecCCc-cCCCceEEEEEeecCCceeeeEEeeeE-EeeC
Q 038524           53 DPMQTWIDYMGSEKLHEIQC-----LS--LSTSVDLSQ-LLLDTMCVGFSAATGSLATEHYILGCS-LNKS  114 (175)
Q Consensus        53 ~~~~vwIdYd~~~~~L~v~l-----Pl--ls~~idLs~-~l~~~~yVGFSAsTG~~~~~h~IlsWs-F~~~  114 (175)
                      -++.|.|+|-+  ++|.|-+     |.  ...-++... +||..-|+|.|||||....-|+||+.. |+..
T Consensus       188 yPvRarItY~~--nvLtv~innGmtp~d~yE~C~rve~~~lp~nGyFGvSAATGgLADDHDVl~FltfsL~  256 (497)
T KOG3838|consen  188 YPVRARITYYG--NVLTVMINNGMTPSDDYEFCVRVENLLLPPNGYFGVSAATGGLADDHDVLSFLTFSLS  256 (497)
T ss_pred             CCceEEEEEec--cEEEEEEcCCCCCCCCcceeEeccceeccCCCeeeeeecccccccccceeeeEEeeec
Confidence            37899999986  4677655     43  111233334 357889999999999999999999874 4443


No 12 
>PF08693 SKG6:  Transmembrane alpha-helix domain;  InterPro: IPR014805 SKG6 and AXL2 are membrane proteins that show polarised intracellular localisation [, ]. This entry represents the highly conserved transmembrane alpha-helical domain found in these proteins [, ]. The full-length AXL2 protein has a negative regulatory function in cytokinesis [].
Probab=93.36  E-value=0.069  Score=33.13  Aligned_cols=8  Identities=25%  Similarity=0.488  Sum_probs=3.5

Q ss_pred             hhhheecc
Q 038524          157 AVYIVRKK  164 (175)
Q Consensus       157 ~~~~~rr~  164 (175)
                      +++++||+
T Consensus        32 l~~~~rR~   39 (40)
T PF08693_consen   32 LFFWYRRK   39 (40)
T ss_pred             hheEEecc
Confidence            33344544


No 13 
>PF04478 Mid2:  Mid2 like cell wall stress sensor;  InterPro: IPR007567 This family represents a region near the C terminus of Mid2, which contains a transmembrane region. The remainder of the protein sequence is serine-rich and of low complexity, and is therefore impossible to align accurately. Mid2 is thought to act as a mechanosensor of cell wall stress. The C-terminal cytoplasmic region of Mid2 is known to interact with Rom2, a guanine nucleotide exchange factor (GEF) for Rho1, which is part of the cell wall integrity signalling pathway [].
Probab=89.73  E-value=0.22  Score=39.51  Aligned_cols=28  Identities=18%  Similarity=0.164  Sum_probs=11.6

Q ss_pred             eEEeh--hHHHHHHHHHHH-Hhhhhheeccc
Q 038524          138 LVTCV--ILVALVSVLTTV-GAAVYIVRKKK  165 (175)
Q Consensus       138 ~l~i~--l~~~~~~~~~~~-~~~~~~~rr~~  165 (175)
                      .++||  ++++++++++++ ++|++++|++|
T Consensus        49 nIVIGvVVGVGg~ill~il~lvf~~c~r~kk   79 (154)
T PF04478_consen   49 NIVIGVVVGVGGPILLGILALVFIFCIRRKK   79 (154)
T ss_pred             cEEEEEEecccHHHHHHHHHhheeEEEeccc
Confidence            34444  444444444333 33334444443


No 14 
>PF15102 TMEM154:  TMEM154 protein family
Probab=89.46  E-value=0.34  Score=38.12  Aligned_cols=17  Identities=18%  Similarity=0.301  Sum_probs=8.4

Q ss_pred             eEEehhHHHHHHHHHHH
Q 038524          138 LVTCVILVALVSVLTTV  154 (175)
Q Consensus       138 ~l~i~l~~~~~~~~~~~  154 (175)
                      .+.|+++++++++++++
T Consensus        58 iLmIlIP~VLLvlLLl~   74 (146)
T PF15102_consen   58 ILMILIPLVLLVLLLLS   74 (146)
T ss_pred             EEEEeHHHHHHHHHHHH
Confidence            45666665444333333


No 15 
>PF01102 Glycophorin_A:  Glycophorin A;  InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=87.45  E-value=0.34  Score=37.02  Aligned_cols=11  Identities=27%  Similarity=0.217  Sum_probs=4.5

Q ss_pred             hhhhheecccc
Q 038524          156 AAVYIVRKKKY  166 (175)
Q Consensus       156 ~~~~~~rr~~~  166 (175)
                      ++|++||++|+
T Consensus        84 i~y~irR~~Kk   94 (122)
T PF01102_consen   84 ISYCIRRLRKK   94 (122)
T ss_dssp             HHHHHHHHS--
T ss_pred             HHHHHHHHhcc
Confidence            34445555544


No 16 
>PF14610 DUF4448:  Protein of unknown function (DUF4448)
Probab=85.33  E-value=0.76  Score=36.94  Aligned_cols=26  Identities=15%  Similarity=0.156  Sum_probs=15.8

Q ss_pred             eEEehhHHHHHHHHHHHHhhhhheec
Q 038524          138 LVTCVILVALVSVLTTVGAAVYIVRK  163 (175)
Q Consensus       138 ~l~i~l~~~~~~~~~~~~~~~~~~rr  163 (175)
                      .+.|+|+++.+++++++.+++++.||
T Consensus       159 ~laI~lPvvv~~~~~~~~~~~~~~R~  184 (189)
T PF14610_consen  159 ALAIALPVVVVVLALIMYGFFFWNRK  184 (189)
T ss_pred             eEEEEccHHHHHHHHHHHhhheeecc
Confidence            78888888766665555333333333


No 17 
>cd06900 lectin_VcfQ VcfQ bacterial pilus biogenesis protein, lectin domain. This family includes bacterial proteins homologous to the VcfQ (also known as MshQ) bacterial pilus biogenesis protein.  VcfQ is encoded by the vcfQ gene of the type IV pilus gene cluster of Vibrio cholerae and is essential for type IV pilus assembly.  VcfQ has a Laminin G-like domain as well as an L-type lectin domain.
Probab=85.04  E-value=4.2  Score=34.75  Aligned_cols=28  Identities=18%  Similarity=0.338  Sum_probs=24.5

Q ss_pred             ccCCCceEEEEEeecCCceeeeEEeeeE
Q 038524           83 QLLLDTMCVGFSAATGSLATEHYILGCS  110 (175)
Q Consensus        83 ~~l~~~~yVGFSAsTG~~~~~h~IlsWs  110 (175)
                      .-+|+..+++|++|||..+-.|.|-...
T Consensus       225 ~avP~~f~lS~TgSTGgstN~HEIdnf~  252 (255)
T cd06900         225 DAIPENFYLSFTGSTGGSTNTHEIDNFQ  252 (255)
T ss_pred             CCCCccEEEEEEecCCCcccceeecceE
Confidence            6789999999999999999999987543


No 18 
>PF05454 DAG1:  Dystroglycan (Dystrophin-associated glycoprotein 1);  InterPro: IPR008465 Dystroglycan is one of the dystrophin-associated glycoproteins, which is encoded by a 5.5 kb transcript in Homo sapiens. The protein product is cleaved into two non-covalently associated subunits, [alpha] (N-terminal) and [beta] (C-terminal). In skeletal muscle the dystroglycan complex works as a transmembrane linkage between the extracellular matrix and the cytoskeleton [alpha]-dystroglycan is extracellular and binds to merosin ([alpha]-2 laminin) in the basement membrane, while [beta]-dystroglycan is a transmembrane protein and binds to dystrophin, which is a large rod-like cytoskeletal protein, absent in Duchenne muscular dystrophy patients. Dystrophin binds to intracellular actin cables. In this way, the dystroglycan complex, which links the extracellular matrix to the intracellular actin cables, is thought to provide structural integrity in muscle tissues. The dystroglycan complex is also known to serve as an agrin receptor in muscle, where it may regulate agrin-induced acetylcholine receptor clustering at the neuromuscular junction. There is also evidence which suggests the function of dystroglycan as a part of the signal transduction pathway because it is shown that Grb2, a mediator of the Ras-related signal pathway, can interact with the cytoplasmic domain of dystroglycan. In general, aberrant expression of dystrophin-associated protein complex underlies the pathogenesis of Duchenne muscular dystrophy, Becker muscular dystrophy and severe childhood autosomal recessive muscular dystrophy. Interestingly, no genetic disease has been described for either [alpha]- or [beta]-dystroglycan. Dystroglycan is widely distributed in non-muscle tissues as well as in muscle tissues. During epithelial morphogenesis of kidney, the dystroglycan complex is shown to act as a receptor for the basement membrane. Dystroglycan expression in Mus musculus brain and neural retina has also been reported. However, the physiological role of dystroglycan in non-muscle tissues has remained unclear [].; PDB: 1EG4_P.
Probab=83.82  E-value=0.33  Score=42.16  Aligned_cols=24  Identities=17%  Similarity=0.325  Sum_probs=0.0

Q ss_pred             ehhHHHHHHHHHHHHhhhhheecc
Q 038524          141 CVILVALVSVLTTVGAAVYIVRKK  164 (175)
Q Consensus       141 i~l~~~~~~~~~~~~~~~~~~rr~  164 (175)
                      +++.++++++++++++++++||||
T Consensus       151 paVVI~~iLLIA~iIa~icyrrkR  174 (290)
T PF05454_consen  151 PAVVIAAILLIAGIIACICYRRKR  174 (290)
T ss_dssp             ------------------------
T ss_pred             HHHHHHHHHHHHHHHHHHhhhhhh
Confidence            333333333333333344444443


No 19 
>PF01299 Lamp:  Lysosome-associated membrane glycoprotein (Lamp);  InterPro: IPR002000 Lysosome-associated membrane glycoproteins (lamp) [] are integral membrane proteins, specific to lysosomes, and whose exact biological function is not yet clear. Structurally, the lamp proteins consist of two internally homologous lysosome-luminal domains separated by a proline-rich hinge region; at the C-terminal extremity there is a transmembrane region (TM) followed by a very short cytoplasmic tail (C). In each of the duplicated domains, there are two conserved disulphide bonds. This structure is schematically represented in the figure below.   +-----+ +-----+ +-----+ +-----+ | | | | | | | | xCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxxxCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxx +--------------------------++Hinge++--------------------------++TM++C+  In mammals, there are two closely related types of lamp: lamp-1 and lamp-2, which form major components of the lysosome membrane. In chicken lamp-1 is known as LEP100.  Also included in this entry is the macrophage protein CD68 (or macrosialin) [] is a heavily glycosylated integral membrane protein whose structure consists of a mucin-like domain followed by a proline-rich hinge; a single lamp-like domain; a transmembrane region and a short cytoplasmic tail.   Similar to CD68, mammalian lamp-3, which is expressed in lymphoid organs, dendritic cells and in lung, contains all the C-terminal regions but lacks the N-terminal lamp-like region []. In a lamp-family protein from nematodes [] only the part C-terminal to the hinge is conserved. ; GO: 0016020 membrane
Probab=83.19  E-value=0.81  Score=39.48  Aligned_cols=30  Identities=20%  Similarity=0.043  Sum_probs=13.3

Q ss_pred             eEEehhHHHHHHH--HHHHHhhhhheecccch
Q 038524          138 LVTCVILVALVSV--LTTVGAAVYIVRKKKYD  167 (175)
Q Consensus       138 ~l~i~l~~~~~~~--~~~~~~~~~~~rr~~~~  167 (175)
                      ..++.+.+|++++  +++++++|++.|||+.+
T Consensus       270 ~~~vPIaVG~~La~lvlivLiaYli~Rrr~~~  301 (306)
T PF01299_consen  270 SDLVPIAVGAALAGLVLIVLIAYLIGRRRSRA  301 (306)
T ss_pred             cchHHHHHHHHHHHHHHHHHHhheeEeccccc
Confidence            3444444443333  33333445555555444


No 20 
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=77.94  E-value=0.95  Score=32.97  Aligned_cols=14  Identities=29%  Similarity=0.159  Sum_probs=6.3

Q ss_pred             ceeEEehhHHHHHH
Q 038524          136 QKLVTCVILVALVS  149 (175)
Q Consensus       136 ~~~l~i~l~~~~~~  149 (175)
                      +...++|+++++++
T Consensus        64 s~gaiagi~vg~~~   77 (96)
T PTZ00382         64 STGAIAGISVAVVA   77 (96)
T ss_pred             ccccEEEEEeehhh
Confidence            34444554454333


No 21 
>PF14575 EphA2_TM:  Ephrin type-A receptor 2 transmembrane domain; PDB: 3KUL_A 2XVD_A 2VX1_A 2VWV_A 2VX0_A 2VWY_A 2VWZ_A 2VWW_A 2VWU_A 2VWX_A ....
Probab=76.64  E-value=1.4  Score=30.71  Aligned_cols=11  Identities=18%  Similarity=0.111  Sum_probs=4.0

Q ss_pred             Hhhhhheeccc
Q 038524          155 GAAVYIVRKKK  165 (175)
Q Consensus       155 ~~~~~~~rr~~  165 (175)
                      +++++++||.+
T Consensus        18 ~~~~~~~rr~~   28 (75)
T PF14575_consen   18 IIVIVCFRRCK   28 (75)
T ss_dssp             HHHHCCCTT--
T ss_pred             eeEEEEEeeEc
Confidence            34444444443


No 22 
>PF06697 DUF1191:  Protein of unknown function (DUF1191);  InterPro: IPR010605 This family contains hypothetical plant proteins of unknown function.
Probab=75.92  E-value=4.8  Score=34.89  Aligned_cols=15  Identities=7%  Similarity=-0.002  Sum_probs=6.9

Q ss_pred             EEehhHHHHHHHHHH
Q 038524          139 VTCVILVALVSVLTT  153 (175)
Q Consensus       139 l~i~l~~~~~~~~~~  153 (175)
                      +++|+.++.++++++
T Consensus       215 iv~g~~~G~~~L~ll  229 (278)
T PF06697_consen  215 IVVGVVGGVVLLGLL  229 (278)
T ss_pred             EEEEehHHHHHHHHH
Confidence            344544554444444


No 23 
>PF14283 DUF4366:  Domain of unknown function (DUF4366)
Probab=75.75  E-value=11  Score=31.55  Aligned_cols=29  Identities=14%  Similarity=0.077  Sum_probs=24.7

Q ss_pred             CCCcccccccCCCcEEEEEEeeCCCceEEEe
Q 038524           41 KGTNKSLKLISGDPMQTWIDYMGSEKLHEIQ   71 (175)
Q Consensus        41 ~~~~~~~~l~sG~~~~vwIdYd~~~~~L~v~   71 (175)
                      +.+|..+.-++|+.++.-||+|....  +|+
T Consensus        72 ~kQFiTv~Tk~gn~FyliIDr~~~~e--nV~  100 (218)
T PF14283_consen   72 GKQFITVTTKSGNTFYLIIDRDEEGE--NVY  100 (218)
T ss_pred             CcEEEEEEecCCCEEEEEEecCCCcc--eEE
Confidence            44688888999999999999999977  564


No 24 
>PF14991 MLANA:  Protein melan-A; PDB: 2GTZ_F 2GT9_F 3MRO_P 2GUO_C 3MRQ_P 2GTW_C 3L6F_C 3MRP_P.
Probab=75.31  E-value=0.73  Score=34.84  Aligned_cols=15  Identities=20%  Similarity=0.341  Sum_probs=0.0

Q ss_pred             Hhhhhheecccchhh
Q 038524          155 GAAVYIVRKKKYDEV  169 (175)
Q Consensus       155 ~~~~~~~rr~~~~e~  169 (175)
                      +.++++|||..|+-+
T Consensus        42 iGCWYckRRSGYk~L   56 (118)
T PF14991_consen   42 IGCWYCKRRSGYKTL   56 (118)
T ss_dssp             ---------------
T ss_pred             Hhheeeeecchhhhh
Confidence            566778877665443


No 25 
>PHA03265 envelope glycoprotein D; Provisional
Probab=67.73  E-value=4.6  Score=36.25  Aligned_cols=28  Identities=18%  Similarity=0.030  Sum_probs=15.8

Q ss_pred             ceeEEehhHHHHHHHHHHHHhhhhheecc
Q 038524          136 QKLVTCVILVALVSVLTTVGAAVYIVRKK  164 (175)
Q Consensus       136 ~~~l~i~l~~~~~~~~~~~~~~~~~~rr~  164 (175)
                      .++++||+.+++++++-+ ++++++||||
T Consensus       349 ~~g~~ig~~i~glv~vg~-il~~~~rr~k  376 (402)
T PHA03265        349 FVGISVGLGIAGLVLVGV-ILYVCLRRKK  376 (402)
T ss_pred             ccceEEccchhhhhhhhH-HHHHHhhhhh
Confidence            356777877766554432 3445555554


No 26 
>PF03302 VSP:  Giardia variant-specific surface protein;  InterPro: IPR005127 During infection, the intestinal protozoan parasite Giardia lamblia virus undergoes continuous antigenic variation which is determined by diversification of the parasite's major surface antigen, named VSP (variant surface protein).
Probab=65.17  E-value=4.7  Score=36.33  Aligned_cols=19  Identities=26%  Similarity=0.165  Sum_probs=13.8

Q ss_pred             ceeEEehhHHHHHHHHHHH
Q 038524          136 QKLVTCVILVALVSVLTTV  154 (175)
Q Consensus       136 ~~~l~i~l~~~~~~~~~~~  154 (175)
                      ++..++|++|+++++|-.|
T Consensus       365 stgaIaGIsvavvvvVggl  383 (397)
T PF03302_consen  365 STGAIAGISVAVVVVVGGL  383 (397)
T ss_pred             cccceeeeeehhHHHHHHH
Confidence            6778888888876665544


No 27 
>PF06024 DUF912:  Nucleopolyhedrovirus protein of unknown function (DUF912);  InterPro: IPR009261 This entry is represented by Autographa californica nuclear polyhedrosis virus (AcMNPV), Orf78; it is a family of uncharacterised viral proteins.
Probab=61.78  E-value=5.6  Score=29.05  Aligned_cols=10  Identities=20%  Similarity=0.152  Sum_probs=4.1

Q ss_pred             Hhhhhheecc
Q 038524          155 GAAVYIVRKK  164 (175)
Q Consensus       155 ~~~~~~~rr~  164 (175)
                      +.+|++.|.+
T Consensus        81 IyYFVILRer   90 (101)
T PF06024_consen   81 IYYFVILRER   90 (101)
T ss_pred             heEEEEEecc
Confidence            3344444433


No 28 
>PF12877 DUF3827:  Domain of unknown function (DUF3827);  InterPro: IPR024606 The function of the proteins in this entry is not currently known, but one of the human proteins (Q9HCM3 from SWISSPROT) has been implicated in pilocytic astrocytomas [, , ]. In the majority of cases of pilocytic astrocytomas a tandem duplication produces an in-frame fusion of the gene encoding this protein and the BRAF oncogene. The resulting fusion protein has constitutive BRAF kinase activity and is capable of transforming cells. 
Probab=59.64  E-value=14  Score=35.56  Aligned_cols=18  Identities=22%  Similarity=0.351  Sum_probs=8.1

Q ss_pred             eeEEehhHHHHHHHHHHH
Q 038524          137 KLVTCVILVALVSVLTTV  154 (175)
Q Consensus       137 ~~l~i~l~~~~~~~~~~~  154 (175)
                      ..|++|+.+.++++++++
T Consensus       269 lWII~gVlvPv~vV~~Ii  286 (684)
T PF12877_consen  269 LWIIAGVLVPVLVVLLII  286 (684)
T ss_pred             eEEEehHhHHHHHHHHHH
Confidence            445556544444333333


No 29 
>PF02009 Rifin_STEVOR:  Rifin/stevor family;  InterPro: IPR002858 Malaria is still a major cause of mortality in many areas of the world. Plasmodium falciparum causes the most severe human form of the disease and is responsible for most fatalities. Severe cases of malaria can occur when the parasite invades and then proliferates within red blood cell erythrocytes. The parasite produces many variant antigenic proteins, encoded by multigene families, which are present on the surface of the infected erythrocyte and play important roles in virulence. A crucial survival mechanism for the malaria parasite is its ability to evade the immune response by switching these variant surface antigens. The high virulence of P. falciparum relative to other malarial parasites is in large part due to the fact that in this organism many of these surface antigens mediate the binding of infected erythrocytes to the vascular endothelium (cytoadherence) and non-infected erythrocytes (rosetting). This can lead to the accumulation of infected cells in the vasculature of a variety of organs, blocking the blood flow and reducing the oxygen supply. Clinical symptoms of severe infection can include fever, progressive anaemia, multi-organ dysfunction and coma. For more information see []. Several multicopy gene families have been described in Plasmodium falciparum, including the stevor family of subtelomeric open reading frames and the rif interspersed repetitive elements. Both families contain three predicted transmembrane segments. It has been proposed that stevor and rif are members of a larger superfamily that code for variant surface antigens [].
Probab=56.25  E-value=5  Score=35.04  Aligned_cols=7  Identities=0%  Similarity=-0.263  Sum_probs=2.9

Q ss_pred             Hhhhhhe
Q 038524          155 GAAVYIV  161 (175)
Q Consensus       155 ~~~~~~~  161 (175)
                      ++++++|
T Consensus       274 IIYLILR  280 (299)
T PF02009_consen  274 IIYLILR  280 (299)
T ss_pred             HHHHHHH
Confidence            3444443


No 30 
>PF11857 DUF3377:  Domain of unknown function (DUF3377);  InterPro: IPR021805  This domain is functionally uncharacterised and found at the C terminus of peptidases belonging to MEROPS peptidase family M10A, membrane-type matrix metallopeptidases (clan MA). ; GO: 0004222 metalloendopeptidase activity
Probab=54.22  E-value=20  Score=25.09  Aligned_cols=18  Identities=22%  Similarity=0.327  Sum_probs=10.6

Q ss_pred             eeEEehhHHHHHHHHHHH
Q 038524          137 KLVTCVILVALVSVLTTV  154 (175)
Q Consensus       137 ~~l~i~l~~~~~~~~~~~  154 (175)
                      +.+.+.+++.+++.++++
T Consensus        30 ~avaVviPl~L~LCiLvl   47 (74)
T PF11857_consen   30 NAVAVVIPLVLLLCILVL   47 (74)
T ss_pred             eEEEEeHHHHHHHHHHHH
Confidence            455666666655555555


No 31 
>PF12768 Rax2:  Cortical protein marker for cell polarity
Probab=49.81  E-value=14  Score=31.89  Aligned_cols=9  Identities=22%  Similarity=0.331  Sum_probs=4.2

Q ss_pred             eEEehhHHH
Q 038524          138 LVTCVILVA  146 (175)
Q Consensus       138 ~l~i~l~~~  146 (175)
                      ++.|++.+|
T Consensus       229 VVlIslAiA  237 (281)
T PF12768_consen  229 VVLISLAIA  237 (281)
T ss_pred             EEEEehHHH
Confidence            344444444


No 32 
>PF15345 TMEM51:  Transmembrane protein 51
Probab=48.32  E-value=16  Score=30.89  Aligned_cols=10  Identities=20%  Similarity=0.298  Sum_probs=4.3

Q ss_pred             hhhheecccc
Q 038524          157 AVYIVRKKKY  166 (175)
Q Consensus       157 ~~~~~rr~~~  166 (175)
                      |+-+|.|||.
T Consensus        77 CL~IR~KRr~   86 (233)
T PF15345_consen   77 CLSIRDKRRR   86 (233)
T ss_pred             HHHHHHHHHH
Confidence            3445544443


No 33 
>PTZ00208 65 kDa invariant surface glycoprotein; Provisional
Probab=45.49  E-value=9.3  Score=34.87  Aligned_cols=30  Identities=13%  Similarity=0.235  Sum_probs=13.1

Q ss_pred             EEehhHHHHHHHHHHHH-hhhhheecccchh
Q 038524          139 VTCVILVALVSVLTTVG-AAVYIVRKKKYDE  168 (175)
Q Consensus       139 l~i~l~~~~~~~~~~~~-~~~~~~rr~~~~e  168 (175)
                      +++++.+..++|++..+ +++++||||..+|
T Consensus       388 i~~avl~p~~il~~~~~~~~~~v~rrr~~~~  418 (436)
T PTZ00208        388 IILAVLVPAIILAIIAVAFFIMVKRRRNSSE  418 (436)
T ss_pred             HHHHHHHHHHHHHHHHHHhheeeeeccCCch
Confidence            33444443333332223 4444555555444


No 34 
>PF05337 CSF-1:  Macrophage colony stimulating factor-1 (CSF-1);  InterPro: IPR008001 Colony stimulating factor 1 (CSF-1) is a homodimeric polypeptide growth factor whose primary function is to regulate the survival, proliferation, differentiation, and function of cells of the mononuclear phagocytic lineage. This lineage includes mononuclear phagocytic precursors, blood monocytes, tissue macrophages, osteoclasts, and microglia of the brain, all of which possess cell surface receptors for CSF-1. The protein has also been linked with male fertility [] and mutations in the Csf-1 gene have been found to cause osteopetrosis and failure of tooth eruption [].; GO: 0005125 cytokine activity, 0008083 growth factor activity, 0016021 integral to membrane; PDB: 3EJJ_A.
Probab=44.15  E-value=7.5  Score=33.71  Aligned_cols=29  Identities=14%  Similarity=0.311  Sum_probs=0.0

Q ss_pred             eeEEehhHHHHHHHHHHH-Hhhhhheeccc
Q 038524          137 KLVTCVILVALVSVLTTV-GAAVYIVRKKK  165 (175)
Q Consensus       137 ~~l~i~l~~~~~~~~~~~-~~~~~~~rr~~  165 (175)
                      ..++.-|.+.++++|++. +.+++||||+|
T Consensus       224 p~~vf~lLVPSiILVLLaVGGLLfYr~rrR  253 (285)
T PF05337_consen  224 PGFVFYLLVPSIILVLLAVGGLLFYRRRRR  253 (285)
T ss_dssp             ------------------------------
T ss_pred             Ccccccccccchhhhhhhccceeeeccccc
Confidence            345566667666666555 33444444433


No 35 
>PF12191 stn_TNFRSF12A:  Tumour necrosis factor receptor stn_TNFRSF12A_TNFR domain;  InterPro: IPR022316 The tumour necrosis factor (TNF) receptor (TNFR) superfamily comprises more than 20 type-I transmembrane proteins. Family members are defined based on similarity in their extracellular domain - a region that contains many cysteine residues arranged in a specific repetitive pattern []. The cysteines allow formation of an extended rod-like structure, responsible for ligand binding []. Upon receptor activation, different intracellular signalling complexes are assembled for different members of the TNFR superfamily, depending on their intracellular domains and sequences []. Activation of TNFRs can therefore induce a range of disparate effects, including cell proliferation, differentiation, survival, or apoptotic cell death, depending upon the receptor involved []. TNFRs are widely distributed and play important roles in many crucial biological processes, such as lymphoid and neuronal development, innate and adaptive immunity, and maintenance of cellular homeostasis []. Drugs that manipulate their signalling have potential roles in the prevention and treatment of many diseases, such as viral infections, coronary heart disease, transplant rejection, and immune disease []. TNF receptor 12 (also known as TWEAK receptor, and fibroblast growth factor-inducible-14 (Fn14)) has been implicated in endothelial cell growth and migration []. The receptor may also play a role in cell-matrix interactions [].; PDB: 2KN0_A 2RPJ_A 2KMZ_A 2EQP_A.
Probab=43.16  E-value=7.9  Score=29.81  Aligned_cols=24  Identities=8%  Similarity=-0.057  Sum_probs=0.0

Q ss_pred             EEehhHHHHHHHHHHH-Hhhhhhee
Q 038524          139 VTCVILVALVSVLTTV-GAAVYIVR  162 (175)
Q Consensus       139 l~i~l~~~~~~~~~~~-~~~~~~~r  162 (175)
                      ..|+.++.++++++.+ ..++++||
T Consensus        79 ~pi~~sal~v~lVl~llsg~lv~rr  103 (129)
T PF12191_consen   79 WPILGSALSVVLVLALLSGFLVWRR  103 (129)
T ss_dssp             -------------------------
T ss_pred             hhhhhhHHHHHHHHHHHHHHHHHhh
Confidence            3343344444444343 33344443


No 36 
>PLN03150 hypothetical protein; Provisional
Probab=42.61  E-value=25  Score=33.31  Aligned_cols=6  Identities=17%  Similarity=0.446  Sum_probs=2.3

Q ss_pred             hhheec
Q 038524          158 VYIVRK  163 (175)
Q Consensus       158 ~~~~rr  163 (175)
                      +++|||
T Consensus       566 ~~~~~r  571 (623)
T PLN03150        566 CWWKRR  571 (623)
T ss_pred             hheeeh
Confidence            334433


No 37 
>PF03229 Alpha_GJ:  Alphavirus glycoprotein J;  InterPro: IPR004913 The exact function of the herpesvirus glycoprotein J is unknown, but it appears to play a role in the inhibition of apotosis of the host cell [].; GO: 0019050 suppression by virus of host apoptosis
Probab=42.18  E-value=40  Score=25.72  Aligned_cols=28  Identities=18%  Similarity=0.148  Sum_probs=14.2

Q ss_pred             eeEEehhHHHHHHHHHHH--Hhhhhheecc
Q 038524          137 KLVTCVILVALVSVLTTV--GAAVYIVRKK  164 (175)
Q Consensus       137 ~~l~i~l~~~~~~~~~~~--~~~~~~~rr~  164 (175)
                      ..+++++.++++.++.+.  +...++||+.
T Consensus        82 ~d~aLp~VIGGLcaL~LaamGA~~LLrR~c  111 (126)
T PF03229_consen   82 VDFALPLVIGGLCALTLAAMGAGALLRRCC  111 (126)
T ss_pred             cccchhhhhhHHHHHHHHHHHHHHHHHHHH
Confidence            346666666655444333  4445554433


No 38 
>PF06365 CD34_antigen:  CD34/Podocalyxin family;  InterPro: IPR013836 This family consists of several mammalian CD34 antigen proteins. The CD34 antigen is a human leukocyte membrane protein expressed specifically by lymphohematopoietic progenitor cells. CD34 is a phosphoprotein. Activation of protein kinase C (PKC) has been found to enhance CD34 phosphorylation [, ]. This family contains several eukaryotic podocalyxin proteins. Podocalyxin is a major membrane protein of the glomerular epithelium and is thought to be involved in maintenance of the architecture of the foot processes and filtration slits characteristic of this unique epithelium by virtue of its high negative charge. Podocalyxin functions as an anti-adhesin that maintains an open filtration pathway between neighbouring foot processes in the glomerular epithelium by charge repulsion [].
Probab=41.94  E-value=46  Score=27.55  Aligned_cols=26  Identities=8%  Similarity=0.118  Sum_probs=11.2

Q ss_pred             eEEehhHHHHHHHHHHH--Hhhhhheec
Q 038524          138 LVTCVILVALVSVLTTV--GAAVYIVRK  163 (175)
Q Consensus       138 ~l~i~l~~~~~~~~~~~--~~~~~~~rr  163 (175)
                      .|+..+.++++++++++  ++++++.||
T Consensus       101 ~lI~lv~~g~~lLla~~~~~~Y~~~~Rr  128 (202)
T PF06365_consen  101 TLIALVTSGSFLLLAILLGAGYCCHQRR  128 (202)
T ss_pred             EEEehHHhhHHHHHHHHHHHHHHhhhhc
Confidence            34444444444444333  334444444


No 39 
>TIGR01167 LPXTG_anchor LPXTG-motif cell wall anchor domain. A common feature of this proteins containing this domain appears to be a high proportion of charged and zwitterionic residues immediatedly upstream of the LPXTG motif. This model differs from other descriptions of the LPXTG region by including a portion of that upstream charged region.
Probab=41.10  E-value=31  Score=19.47  Aligned_cols=6  Identities=17%  Similarity=0.523  Sum_probs=2.4

Q ss_pred             hheecc
Q 038524          159 YIVRKK  164 (175)
Q Consensus       159 ~~~rr~  164 (175)
                      +++||+
T Consensus        27 ~~~~rk   32 (34)
T TIGR01167        27 LLRKRK   32 (34)
T ss_pred             Hheecc
Confidence            333443


No 40 
>PF02480 Herpes_gE:  Alphaherpesvirus glycoprotein E;  InterPro: IPR003404 Glycoprotein E (gE) of Alphaherpesvirus forms a complex with glycoprotein I (gI), functioning as an immunoglobulin G (IgG) Fc binding protein. gE is involved in virus spread but is not essential for propagation [].; GO: 0016020 membrane; PDB: 2GJ7_F 2GIY_B.
Probab=40.75  E-value=9.1  Score=35.09  Aligned_cols=9  Identities=0%  Similarity=-0.090  Sum_probs=5.4

Q ss_pred             EEeeeEEee
Q 038524          105 YILGCSLNK  113 (175)
Q Consensus       105 ~IlsWsF~~  113 (175)
                      .+.+|....
T Consensus       293 hv~aW~yt~  301 (439)
T PF02480_consen  293 HVEAWTYTL  301 (439)
T ss_dssp             EEEEEEEEE
T ss_pred             eeeeeEEEE
Confidence            456776653


No 41 
>PF15050 SCIMP:  SCIMP protein
Probab=40.70  E-value=9.9  Score=29.17  Aligned_cols=8  Identities=0%  Similarity=-0.566  Sum_probs=3.7

Q ss_pred             Hhhhhhee
Q 038524          155 GAAVYIVR  162 (175)
Q Consensus       155 ~~~~~~~r  162 (175)
                      ++++++|+
T Consensus        26 IlyCvcR~   33 (133)
T PF15050_consen   26 ILYCVCRW   33 (133)
T ss_pred             HHHHHHHH
Confidence            34444553


No 42 
>TIGR01478 STEVOR variant surface antigen, stevor family. This model represents the stevor branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of stevor sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 8 bits.
Probab=38.96  E-value=20  Score=31.28  Aligned_cols=7  Identities=29%  Similarity=0.254  Sum_probs=2.9

Q ss_pred             hhheecc
Q 038524          158 VYIVRKK  164 (175)
Q Consensus       158 ~~~~rr~  164 (175)
                      +++|||+
T Consensus       281 WlyrrRK  287 (295)
T TIGR01478       281 WLYRRRK  287 (295)
T ss_pred             HHHHhhc
Confidence            3344444


No 43 
>PTZ00370 STEVOR; Provisional
Probab=38.77  E-value=20  Score=31.28  Aligned_cols=7  Identities=29%  Similarity=0.254  Sum_probs=3.0

Q ss_pred             hhheecc
Q 038524          158 VYIVRKK  164 (175)
Q Consensus       158 ~~~~rr~  164 (175)
                      +++|||+
T Consensus       277 wlyrrRK  283 (296)
T PTZ00370        277 WLYRRRK  283 (296)
T ss_pred             HHHHhhc
Confidence            3344444


No 44 
>TIGR03370 PEPCTERM_Roseo variant PEP-CTERM putative exosortase signal, Roseobacter type. A probable protein export sorting signal, PEP-CTERM, was described by Haft, et al. (PubMed:16930487). It is predicted to interact with a putative transpeptidase we designate exosortase. Most examples of this signal are recognized by model TIGR02595, but some unusual clades require different models. This model describes a variant with conserved motif VPLPA, rather than VPEP. This variant is found prominently in two members of the Rhodobacterales, namely Jannaschia sp. CCS1 and Roseobacter denitrificans OCh 114. One interesting member protein has a full-length duplication and therefore two copies of this putative sorting domain.
Probab=37.67  E-value=20  Score=20.14  Aligned_cols=11  Identities=18%  Similarity=0.416  Sum_probs=5.0

Q ss_pred             Hhhhhheeccc
Q 038524          155 GAAVYIVRKKK  165 (175)
Q Consensus       155 ~~~~~~~rr~~  165 (175)
                      +.+...|||+|
T Consensus        15 ggl~~~rRRrk   25 (26)
T TIGR03370        15 GGLGAMRRRRR   25 (26)
T ss_pred             HHHHHHHHhhc
Confidence            33444555543


No 45 
>PF13268 DUF4059:  Protein of unknown function (DUF4059)
Probab=35.72  E-value=22  Score=24.71  Aligned_cols=24  Identities=21%  Similarity=0.228  Sum_probs=10.2

Q ss_pred             hHHHHHHHHHHHHhhhhheecccc
Q 038524          143 ILVALVSVLTTVGAAVYIVRKKKY  166 (175)
Q Consensus       143 l~~~~~~~~~~~~~~~~~~rr~~~  166 (175)
                      +.+++++++++.+.+.++|.++|+
T Consensus        13 L~ls~i~V~~~~~~wi~~Ra~~~~   36 (72)
T PF13268_consen   13 LLLSSILVLLVSGIWILWRALRKK   36 (72)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHcC
Confidence            344443333333444445544444


No 46 
>PF08374 Protocadherin:  Protocadherin;  InterPro: IPR013585 The structure of protocadherins is similar to that of classic cadherins (IPR002126 from INTERPRO), but they also have some unique features associated with the cytoplasmic domains. They are expressed in a variety of organisms and are found in high concentrations in the brain where they seem to be localised mainly at cell-cell contact sites. Their expression seems to be developmentally regulated []. 
Probab=34.79  E-value=65  Score=27.04  Aligned_cols=17  Identities=12%  Similarity=0.387  Sum_probs=7.2

Q ss_pred             eEEehhHHHHHHHHHHH
Q 038524          138 LVTCVILVALVSVLTTV  154 (175)
Q Consensus       138 ~l~i~l~~~~~~~~~~~  154 (175)
                      .+++|+..++++++|++
T Consensus        38 ~I~iaiVAG~~tVILVI   54 (221)
T PF08374_consen   38 KIMIAIVAGIMTVILVI   54 (221)
T ss_pred             eeeeeeecchhhhHHHH
Confidence            44444444444444333


No 47 
>PF12248 Methyltransf_FA:  Farnesoic acid 0-methyl transferase;  InterPro: IPR022041  This domain, found in farnesoic acid O-methyl transferase, is approximately 110 amino acids in length. Farnesoic acid O-methyl transferase (FAMeT) is the enzyme that catalyses the formation of methyl farnesoate (MF) from farnesoic acid (FA) in the biosynthetic pathway of juvenile hormone (JH) []. 
Probab=33.06  E-value=1.8e+02  Score=20.71  Aligned_cols=41  Identities=20%  Similarity=0.239  Sum_probs=28.9

Q ss_pred             ccCCCcEEEEEEeeCCCceEEEee-----eeeeeeecCCccCCCceEEEEE
Q 038524           49 LISGDPMQTWIDYMGSEKLHEIQC-----LSLSTSVDLSQLLLDTMCVGFS   94 (175)
Q Consensus        49 l~sG~~~~vwIdYd~~~~~L~v~l-----Plls~~idLs~~l~~~~yVGFS   94 (175)
                      |...+...-||.+++  ..+.|+.     |+|+.. |-.  -..--|||||
T Consensus        49 ls~~e~~~fwI~~~~--G~I~vg~~g~~~pfl~~~-Dp~--~~~v~yvGft   94 (102)
T PF12248_consen   49 LSPSEFRMFWISWRD--GTIRVGRGGEDEPFLEWT-DPE--PIPVNYVGFT   94 (102)
T ss_pred             CCCCccEEEEEEECC--CEEEEEECCCccEEEEEE-CCC--CCcccEEEEe
Confidence            467888999999765  4666665     888876 322  3356799994


No 48 
>PF13908 Shisa:  Wnt and FGF inhibitory regulator
Probab=32.30  E-value=41  Score=26.54  Aligned_cols=11  Identities=0%  Similarity=0.187  Sum_probs=4.7

Q ss_pred             eEEehhHHHHH
Q 038524          138 LVTCVILVALV  148 (175)
Q Consensus       138 ~l~i~l~~~~~  148 (175)
                      .+++++.++++
T Consensus        79 ~iivgvi~~Vi   89 (179)
T PF13908_consen   79 GIIVGVICGVI   89 (179)
T ss_pred             eeeeehhhHHH
Confidence            34444444333


No 49 
>PF03597 CcoS:  Cytochrome oxidase maturation protein cbb3-type;  InterPro: IPR004714 Cytochrome cbb3 oxidases are found almost exclusively in Proteobacteria, and represent a distinctive class of proton-pumping respiratory haem-copper oxidases (HCO) that lack many of the key structural features that contribute to the reaction cycle of the intensely studied mitochondrial cytochrome c oxidase (CcO). Expression of cytochrome cbb3 oxidase allows human pathogens to colonise anoxic tissues and agronomically important diazotrophs to sustain nitrogen fixation []. Genes encoding a cytochrome cbb3 oxidase were initially designated fixNOQP (ccoNOQP), the ccoNOQP operon is always found close to a second gene cluster, known as fixGHIS (ccoGHIS) whose expression is necessary for the assembly of a functional cbb3 oxidase. On the basis of their derived amino acid sequences each of the four proteins encoded by the ccoGHIS operon are thought to be membrane-bound. It has been suggested that they may function in concert as a multi-subunit complex, possibly playing a role in the uptake and metabolism of copper required for the assembly of the binuclear centre of cytochrome cbb3 oxidase. 
Probab=31.72  E-value=36  Score=21.43  Aligned_cols=27  Identities=26%  Similarity=0.498  Sum_probs=11.6

Q ss_pred             hhHHHHHHHHHHHHhhhhheecccchh
Q 038524          142 VILVALVSVLTTVGAAVYIVRKKKYDE  168 (175)
Q Consensus       142 ~l~~~~~~~~~~~~~~~~~~rr~~~~e  168 (175)
                      -++++.++.+++++++++.-|+.++.+
T Consensus         6 lip~sl~l~~~~l~~f~Wavk~GQfdD   32 (45)
T PF03597_consen    6 LIPVSLILGLIALAAFLWAVKSGQFDD   32 (45)
T ss_pred             HHHHHHHHHHHHHHHHHHHHccCCCCC
Confidence            344444333333334444445555533


No 50 
>PRK05886 yajC preprotein translocase subunit YajC; Validated
Probab=31.62  E-value=35  Score=25.51  Aligned_cols=11  Identities=18%  Similarity=-0.055  Sum_probs=5.2

Q ss_pred             Hhhhhheeccc
Q 038524          155 GAAVYIVRKKK  165 (175)
Q Consensus       155 ~~~~~~~rr~~  165 (175)
                      ++|+++|+.+|
T Consensus        17 ~yF~~iRPQkK   27 (109)
T PRK05886         17 FMYFASRRQRK   27 (109)
T ss_pred             HHHHHccHHHH
Confidence            34455554443


No 51 
>PF11353 DUF3153:  Protein of unknown function (DUF3153);  InterPro: IPR021499  This family of proteins with unknown function appear to be restricted to Cyanobacteria. Some members are annotated as membrane proteins however this cannot be confirmed. 
Probab=31.45  E-value=35  Score=27.76  Aligned_cols=25  Identities=28%  Similarity=0.252  Sum_probs=14.2

Q ss_pred             eeeeeeecCCccCCC-ceEEEEEeec
Q 038524           73 LSLSTSVDLSQLLLD-TMCVGFSAAT   97 (175)
Q Consensus        73 Plls~~idLs~~l~~-~~yVGFSAsT   97 (175)
                      .-|...+||+..-.. ..=+-|+=++
T Consensus       117 ~~L~~~lDL~~L~~~~~ldl~f~l~~  142 (209)
T PF11353_consen  117 YRLDLDLDLRSLPDLPGLDLEFSLST  142 (209)
T ss_pred             EEEEEEeehhhcCCCCcceEEEEEeC
Confidence            446777888765432 2345555544


No 52 
>PF14914 LRRC37AB_C:  LRRC37A/B like protein 1 C-terminal domain
Probab=31.22  E-value=29  Score=27.53  Aligned_cols=18  Identities=17%  Similarity=0.265  Sum_probs=9.6

Q ss_pred             eeEEehhHHHHHHHHHHH
Q 038524          137 KLVTCVILVALVSVLTTV  154 (175)
Q Consensus       137 ~~l~i~l~~~~~~~~~~~  154 (175)
                      ..+++++++.+++.++++
T Consensus       119 nklilaisvtvv~~ilii  136 (154)
T PF14914_consen  119 NKLILAISVTVVVMILII  136 (154)
T ss_pred             chhHHHHHHHHHHHHHHH
Confidence            356666666554444333


No 53 
>PTZ00046 rifin; Provisional
Probab=30.50  E-value=33  Score=30.82  Aligned_cols=6  Identities=0%  Similarity=-0.197  Sum_probs=2.2

Q ss_pred             Hhhhhh
Q 038524          155 GAAVYI  160 (175)
Q Consensus       155 ~~~~~~  160 (175)
                      ++++++
T Consensus       333 IIYLIL  338 (358)
T PTZ00046        333 IIYLIL  338 (358)
T ss_pred             HHHHHH
Confidence            333333


No 54 
>TIGR02595 PEP_exosort PEP-CTERM putative exosortase interaction domain. This model describes a 25-residue domain that includes a near-invariant Pro-Glu-Pro (PEP) motif, a thirteen residue strongly hydrophobic sequence likely to span the membrane, and a five-residue strongly basic motif that often contains four Arg residues. In nearly every case, this motif is found within nine residues, and usually within five residues, of the extreme C-terminus of the protein. Proteins with this motif typically have signal sequences at the N-terminus. This region appears many times per genome or not at all, and co-occurs in genomes with a proposed protein-sorting integral membrane protein we designate exosortase (see TIGR02602). PEP-CTERM proteins frequently are poorly conserved, Ser/Thr-rich proteins and may become extensively modified proteinaceous constituents of extracellular material in bacterial biofilms.
Probab=30.04  E-value=45  Score=18.35  Aligned_cols=7  Identities=14%  Similarity=0.605  Sum_probs=3.1

Q ss_pred             hhheecc
Q 038524          158 VYIVRKK  164 (175)
Q Consensus       158 ~~~~rr~  164 (175)
                      +..|||+
T Consensus        17 ~~~rrrk   23 (26)
T TIGR02595        17 LLLRRRR   23 (26)
T ss_pred             HHHhhcc
Confidence            3444444


No 55 
>TIGR03141 cytochro_ccmD heme exporter protein CcmD. The model for this protein family describes a small, hydrophobic, and only moderately well-conserved protein, tricky to identify accurately for all of these reasons. However, members are found as part of large operons involved in heme export across the inner membrane for assembly of c-type cytochromes in a large number of bacteria. The gray zone between the trusted cutoff (13.0) and noise cutoff (4.75) includes both low-scoring examples and false-positive matches to hydrophobic domains of longer proteins.
Probab=29.93  E-value=36  Score=21.12  Aligned_cols=11  Identities=0%  Similarity=0.141  Sum_probs=4.4

Q ss_pred             hhHHHHHHHHH
Q 038524          142 VILVALVSVLT  152 (175)
Q Consensus       142 ~l~~~~~~~~~  152 (175)
                      ..+-+..++++
T Consensus         9 W~sYg~t~l~l   19 (45)
T TIGR03141         9 WLAYGITALVL   19 (45)
T ss_pred             HHHHHHHHHHH
Confidence            33444433333


No 56 
>TIGR03521 GldG gliding-associated putative ABC transporter substrate-binding component GldG. Members of this protein family are exclusive to the Bacteroidetes phylum (previously Cytophaga-Flavobacteria-Bacteroides). GldG is a protein linked to a type of rapid surface gliding motility found in certain Bacteroidetes, such as Flavobacterium johnsoniae and Cytophaga hutchinsonii. Knockouts of GldG abolish the gliding phenotype. GldG, along with GldA and GldF are believed to compose an ABC transporter and are observed as an operon. Gliding motility appears closely linked to chitin utilization in the model species Flavobacterium johnsoniae. Bacteroidetes with members of this protein family appear to have all of the genes associated with gliding motility.
Probab=29.20  E-value=36  Score=31.86  Aligned_cols=12  Identities=42%  Similarity=0.797  Sum_probs=6.5

Q ss_pred             Hhhhhheecccc
Q 038524          155 GAAVYIVRKKKY  166 (175)
Q Consensus       155 ~~~~~~~rr~~~  166 (175)
                      ++++++|||+||
T Consensus       540 G~~~~~~Rrr~~  551 (552)
T TIGR03521       540 GLSFTYIRKRKY  551 (552)
T ss_pred             HHHHHHHHHhhc
Confidence            344455666665


No 57 
>TIGR01582 FDH-beta formate dehydrogenase, beta subunit, Fe-S containing. In addition to the gamma proteobacteria, a sequence from Aquifex aolicus falls within the scope of this model. This appears to be the case for the alpha, gamma and epsilon (accessory protein TIGR01562) chains as well.
Probab=28.98  E-value=71  Score=27.57  Aligned_cols=11  Identities=18%  Similarity=0.111  Sum_probs=6.0

Q ss_pred             CCCCCCCCCCC
Q 038524          120 LNISTLPSFHL  130 (175)
Q Consensus       120 l~~s~lp~~p~  130 (175)
                      ..+..||..|.
T Consensus       228 ~~~~~lp~~p~  238 (283)
T TIGR01582       228 KDYQDLPEDPR  238 (283)
T ss_pred             HHhcCCCCCCc
Confidence            34445676654


No 58 
>COG4736 CcoQ Cbb3-type cytochrome oxidase, subunit 3 [Posttranslational modification, protein turnover, chaperones]
Probab=28.90  E-value=56  Score=21.91  Aligned_cols=10  Identities=20%  Similarity=-0.110  Sum_probs=4.3

Q ss_pred             hhhhheeccc
Q 038524          156 AAVYIVRKKK  165 (175)
Q Consensus       156 ~~~~~~rr~~  165 (175)
                      +++.+|+++|
T Consensus        26 i~~ayr~~~K   35 (60)
T COG4736          26 IYFAYRPGKK   35 (60)
T ss_pred             HHHHhcccch
Confidence            3444444443


No 59 
>PF04689 S1FA:  DNA binding protein S1FA;  InterPro: IPR006779  S1FA is an unusual small plant peptide of only 70 amino acids with a basic domain which contains a nuclear localization signal and a putative DNA binding helix. S1FA is highly conserved between dicotyledonous and monocotyledonous plants and may be a DNA-binding protein that specifically recognises the negative promoter element S1F [].; GO: 0003677 DNA binding, 0006355 regulation of transcription, DNA-dependent, 0005634 nucleus
Probab=28.40  E-value=67  Score=21.99  Aligned_cols=19  Identities=16%  Similarity=0.141  Sum_probs=11.9

Q ss_pred             ceeEEehhHHHHHHHHHHH
Q 038524          136 QKLVTCVILVALVSVLTTV  154 (175)
Q Consensus       136 ~~~l~i~l~~~~~~~~~~~  154 (175)
                      +-++++.+.++.+++++++
T Consensus        11 nPGlIVLlvV~g~ll~flv   29 (69)
T PF04689_consen   11 NPGLIVLLVVAGLLLVFLV   29 (69)
T ss_pred             CCCeEEeehHHHHHHHHHH
Confidence            3457777777766665555


No 60 
>PF10661 EssA:  WXG100 protein secretion system (Wss), protein EssA;  InterPro: IPR018920  The Wss (WXG100 protein secretion system) in Staphylococcus aureus seems to be encoded by a locus of eight ORFs, called ess (eSAT-6 secretion system) []. This locus encodes, amongst several other proteins, EssA, a protein predicted to possess one transmembrane domain. Due to its predicted membrane location and its absolute requirement for WXG100 protein secretion, it has been speculated that EssA could form a secretion apparatus in conjunction with YukC and YukAB. Proteins homologous to EssA, YukC, EsaA and YukD were absent from mycobacteria [].   Members of this family are associated with type VII secretion of WXG100 family targets in the Firmicutes, but not in the Actinobacteria. This highly divergent protein family consists largely of a central region of highly polar low-complexity sequence containing occasional LF motifs in weak repeats about 17 residues in length, flanked by hydrophobic N- and C-terminal regions. 
Probab=26.24  E-value=43  Score=26.14  Aligned_cols=21  Identities=10%  Similarity=0.074  Sum_probs=9.7

Q ss_pred             ehhHHHHHHHHHHHHhhhhhe
Q 038524          141 CVILVALVSVLTTVGAAVYIV  161 (175)
Q Consensus       141 i~l~~~~~~~~~~~~~~~~~~  161 (175)
                      +++++++++++++.+++..+|
T Consensus       121 i~~~i~g~ll~i~~giy~~~r  141 (145)
T PF10661_consen  121 ILLSIGGILLAICGGIYVVLR  141 (145)
T ss_pred             HHHHHHHHHHHHHHHHHHHHH
Confidence            444455554444444444444


No 61 
>PHA03264 envelope glycoprotein D; Provisional
Probab=25.78  E-value=73  Score=28.99  Aligned_cols=18  Identities=11%  Similarity=0.032  Sum_probs=10.3

Q ss_pred             ceeEEehhHHHHHHHHHH
Q 038524          136 QKLVTCVILVALVSVLTT  153 (175)
Q Consensus       136 ~~~l~i~l~~~~~~~~~~  153 (175)
                      ...+.+|+.++..+++++
T Consensus       359 ~~~~~vg~~~a~~~i~~~  376 (416)
T PHA03264        359 ARPVIVGTGIAAAAIACV  376 (416)
T ss_pred             cceeeeehhhhHHHHHHH
Confidence            345667777766444443


No 62 
>PRK07021 fliL flagellar basal body-associated protein FliL; Reviewed
Probab=25.48  E-value=69  Score=25.01  Aligned_cols=8  Identities=13%  Similarity=0.480  Sum_probs=3.4

Q ss_pred             Hhhhhhee
Q 038524          155 GAAVYIVR  162 (175)
Q Consensus       155 ~~~~~~~r  162 (175)
                      +++|++.+
T Consensus        35 g~~~~~~~   42 (162)
T PRK07021         35 GYSWWLSK   42 (162)
T ss_pred             HHHHHhhc
Confidence            34444443


No 63 
>PHA03291 envelope glycoprotein I; Provisional
Probab=24.17  E-value=58  Score=29.43  Aligned_cols=25  Identities=12%  Similarity=0.151  Sum_probs=14.4

Q ss_pred             eEEehhHHHHHHHHHHH-Hhhhhhee
Q 038524          138 LVTCVILVALVSVLTTV-GAAVYIVR  162 (175)
Q Consensus       138 ~l~i~l~~~~~~~~~~~-~~~~~~~r  162 (175)
                      .+.|+++.+.++++++. .++++.|+
T Consensus       288 iiQiAIPasii~cV~lGSC~Ccl~R~  313 (401)
T PHA03291        288 IIQIAIPASIIACVFLGSCACCLHRR  313 (401)
T ss_pred             hheeccchHHHHHhhhhhhhhhhhhh
Confidence            45667777666666555 44555443


No 64 
>PRK15471 chain length determinant protein WzzB; Provisional
Probab=23.46  E-value=60  Score=28.53  Aligned_cols=7  Identities=43%  Similarity=0.524  Sum_probs=3.1

Q ss_pred             ceeEEeh
Q 038524          136 QKLVTCV  142 (175)
Q Consensus       136 ~~~l~i~  142 (175)
                      ++.+++.
T Consensus       293 kr~lIli  299 (325)
T PRK15471        293 KKAITLV  299 (325)
T ss_pred             cchhHHH
Confidence            4444443


No 65 
>PF06809 NPDC1:  Neural proliferation differentiation control-1 protein (NPDC1);  InterPro: IPR009635 This family consists of several neural proliferation differentiation control-1 (NPDC1) proteins. NPDC1 plays a role in the control of neural cell proliferation and differentiation. It has been suggested that NPDC1 may be involved in the development of several secretion glands. This family also contains the C-terminal region of the Caenorhabditis elegans protein CAB-1 (Q93249 from SWISSPROT) which is known to interact with AEX-3 [].; GO: 0016021 integral to membrane
Probab=23.12  E-value=1.4e+02  Score=26.71  Aligned_cols=7  Identities=43%  Similarity=0.781  Sum_probs=4.1

Q ss_pred             eEEEEEe
Q 038524           89 MCVGFSA   95 (175)
Q Consensus        89 ~yVGFSA   95 (175)
                      +..|||.
T Consensus       144 ~tl~~s~  150 (341)
T PF06809_consen  144 ATLGFSE  150 (341)
T ss_pred             ccccccc
Confidence            5566664


No 66 
>PRK10381 LPS O-antigen length regulator; Provisional
Probab=23.10  E-value=46  Score=29.80  Aligned_cols=9  Identities=11%  Similarity=0.257  Sum_probs=4.2

Q ss_pred             ceeEEehhH
Q 038524          136 QKLVTCVIL  144 (175)
Q Consensus       136 ~~~l~i~l~  144 (175)
                      ++.+++.++
T Consensus       337 kr~lIlvl~  345 (377)
T PRK10381        337 GKALIVILA  345 (377)
T ss_pred             chhHHHHHH
Confidence            455544433


No 67 
>PLN00113 leucine-rich repeat receptor-like protein kinase; Provisional
Probab=22.98  E-value=89  Score=30.56  Aligned_cols=12  Identities=8%  Similarity=-0.075  Sum_probs=4.7

Q ss_pred             EEehhHHHHHHH
Q 038524          139 VTCVILVALVSV  150 (175)
Q Consensus       139 l~i~l~~~~~~~  150 (175)
                      +++++.++++++
T Consensus       630 ~~~~~~~~~~~~  641 (968)
T PLN00113        630 FYITCTLGAFLV  641 (968)
T ss_pred             eehhHHHHHHHH
Confidence            344444443333


No 68 
>PRK00523 hypothetical protein; Provisional
Probab=22.46  E-value=53  Score=22.88  Aligned_cols=8  Identities=0%  Similarity=-0.330  Sum_probs=3.8

Q ss_pred             Hhhhhhee
Q 038524          155 GAAVYIVR  162 (175)
Q Consensus       155 ~~~~~~~r  162 (175)
                      +.+|+.||
T Consensus        21 ~Gffiark   28 (72)
T PRK00523         21 IGYFVSKK   28 (72)
T ss_pred             HHHHHHHH
Confidence            34455444


No 69 
>PF03988 DUF347:  Repeat of Unknown Function (DUF347) ;  InterPro: IPR007136 This repeat is found as four tandem repeats in a family of bacterial membrane proteins. Each repeat contains two transmembrane regions and a conserved tryptophan.
Probab=22.05  E-value=62  Score=20.87  Aligned_cols=16  Identities=6%  Similarity=0.046  Sum_probs=7.7

Q ss_pred             EEehhHHHHHHHHHHH
Q 038524          139 VTCVILVALVSVLTTV  154 (175)
Q Consensus       139 l~i~l~~~~~~~~~~~  154 (175)
                      +.++...+++++.+++
T Consensus        25 lglg~~~~~~~~~~~l   40 (55)
T PF03988_consen   25 LGLGYLISTLIFAALL   40 (55)
T ss_pred             cCccHHHHHHHHHHHH
Confidence            4455555544444444


No 70 
>KOG3637 consensus Vitronectin receptor, alpha subunit [Extracellular structures]
Probab=21.35  E-value=1.1e+02  Score=31.24  Aligned_cols=14  Identities=7%  Similarity=0.143  Sum_probs=5.7

Q ss_pred             eEEehhHHHHHHHH
Q 038524          138 LVTCVILVALVSVL  151 (175)
Q Consensus       138 ~l~i~l~~~~~~~~  151 (175)
                      .++|+..+++++++
T Consensus       979 wiIi~svl~GLLlL  992 (1030)
T KOG3637|consen  979 WIIILSVLGGLLLL  992 (1030)
T ss_pred             eeehHHHHHHHHHH
Confidence            34444444444333


No 71 
>COG4282 SMI1 Protein involved in beta-1,3-glucan synthesis [Carbohydrate transport and metabolism]
Probab=21.19  E-value=62  Score=26.35  Aligned_cols=15  Identities=40%  Similarity=0.709  Sum_probs=12.9

Q ss_pred             CCCCCCCCeeEEEcC
Q 038524           10 QFNDTSNNHVGIDVN   24 (175)
Q Consensus        10 e~~D~~~nHVGIdiN   24 (175)
                      -+.|+-+||++||+-
T Consensus       123 L~~d~~Gnhi~IDLa  137 (191)
T COG4282         123 LFGDPRGNHICIDLA  137 (191)
T ss_pred             ecccCCCCeEEEecC
Confidence            468999999999984


No 72 
>PRK11638 lipopolysaccharide biosynthesis protein WzzE; Provisional
Probab=21.16  E-value=43  Score=29.64  Aligned_cols=8  Identities=13%  Similarity=0.322  Sum_probs=3.7

Q ss_pred             hhhheecc
Q 038524          157 AVYIVRKK  164 (175)
Q Consensus       157 ~~~~~rr~  164 (175)
                      +.++||++
T Consensus       334 ~vL~r~~~  341 (342)
T PRK11638        334 VALTRRRR  341 (342)
T ss_pred             eeEeecCC
Confidence            34455543


No 73 
>PF02699 YajC:  Preprotein translocase subunit;  InterPro: IPR003849 Secretion across the inner membrane in some Gram-negative bacteria occurs via the preprotein translocase pathway. Proteins are produced in the cytoplasm as precursors, and require a chaperone subunit to direct them to the translocase component []. From there, the mature proteins are either targeted to the outer membrane, or remain as periplasmic proteins []. The translocase protein subunits are encoded on the bacterial chromosome.  The translocase itself comprises 7 proteins, including a chaperone (SecB), ATPase (SecA), an integral membrane complex (SecY, SecE and SecG), and two additional membrane proteins that promote the release of the mature peptide into the periplasm (SecD and SecF) []. Other cytoplasmic/periplasmic proteins play a part in preprotein translocase activity, namely YidC and YajC []. The latter is bound in a complex to SecD and SecF, and plays a part in stabilising and regulating secretion through the SecYEG integral membrane component via SecA [].  Homologues of the YajC gene have been found in a range of pathogenic and commensal microbes. Brucella abortis YajC- and SecD-like proteins were shown to stimulate a Th1 cell-mediated immune response in mice, and conferred protection when challenged with B.abortis []. Therefore, these proteins may have an antigenic role as well as a secretory one in virulent bacteria []. A number of previously uncharacterised "hypothetical" proteins also show similarity to E.coli YajC, suggesting that this family is wider than first thought [].  More recently, the precise interactions between the E.coli SecYEG complex, SecD, SecF, YajC and YidC have been studied []. Rather than acting individually, the four proteins form a heterotetrameric complex and associate with the SecYEG heterotrimeric complex []. The SecF and YajC subunits link the complex to the integral membrane translocase. ; PDB: 2RDD_B.
Probab=20.69  E-value=73  Score=22.21  Aligned_cols=7  Identities=14%  Similarity=-0.083  Sum_probs=3.0

Q ss_pred             hhhhhee
Q 038524          156 AAVYIVR  162 (175)
Q Consensus       156 ~~~~~~r  162 (175)
                      +++.+|.
T Consensus        16 yf~~~rp   22 (82)
T PF02699_consen   16 YFLMIRP   22 (82)
T ss_dssp             HHHTHHH
T ss_pred             hhheecH
Confidence            4444443


No 74 
>TIGR00847 ccoS cytochrome oxidase maturation protein, cbb3-type. CcoS from Rhodobacter capsulatus has been shown essential for incorporation of redox-active prosthetic groups (heme, Cu) into cytochrome cbb(3) oxidase. FixS of Bradyrhizobium japonicum appears to have the same function. Members of this family are found so far in organisms with a cbb3-type cytochrome oxidase, including Neisseria meningitidis, Helicobacter pylori, Campylobacter jejuni, Caulobacter crescentus, Bradyrhizobium japonicum, and Rhodobacter capsulatus.
Probab=20.35  E-value=1.1e+02  Score=19.82  Aligned_cols=13  Identities=23%  Similarity=0.473  Sum_probs=5.9

Q ss_pred             Hhhhhheecccch
Q 038524          155 GAAVYIVRKKKYD  167 (175)
Q Consensus       155 ~~~~~~~rr~~~~  167 (175)
                      +++++.-|+.+|.
T Consensus        20 ~~f~Wavk~GQfD   32 (51)
T TIGR00847        20 VAFLWSLKSGQYD   32 (51)
T ss_pred             HHHHHHHccCCCC
Confidence            3444444555543


No 75 
>PF07332 DUF1469:  Protein of unknown function (DUF1469);  InterPro: IPR009937 This entry represents proteins found in hypothetical bacterial proteins where is is annotated as ycf49 or ycf49-like. The function is not known.
Probab=20.31  E-value=56  Score=23.77  Aligned_cols=8  Identities=38%  Similarity=0.397  Sum_probs=4.7

Q ss_pred             hhhhcccC
Q 038524          167 DEVYEDWE  174 (175)
Q Consensus       167 ~e~~edwE  174 (175)
                      +|+.+|++
T Consensus       109 ~~l~~d~~  116 (121)
T PF07332_consen  109 AELKEDIA  116 (121)
T ss_pred             HHHHHHHH
Confidence            55666654


No 76 
>TIGR01006 polys_exp_MPA1 polysaccharide export protein, MPA1 family, Gram-positive type. This family contains members from Low GC Gram-positive bacteria; they are proposed to have a function in the export of complex polysaccharides.
Probab=20.07  E-value=50  Score=26.74  Aligned_cols=14  Identities=21%  Similarity=0.067  Sum_probs=7.1

Q ss_pred             hhheecccchhhhc
Q 038524          158 VYIVRKKKYDEVYE  171 (175)
Q Consensus       158 ~~~~rr~~~~e~~e  171 (175)
                      .++.++-|-.|..|
T Consensus       197 ~~~d~~i~~~~d~~  210 (226)
T TIGR01006       197 ELLDTRVKRPEDVE  210 (226)
T ss_pred             HHHhCCcCCHHHHH
Confidence            34555555555444


Done!