Query 038524
Match_columns 175
No_of_seqs 153 out of 833
Neff 6.2
Searched_HMMs 46136
Date Fri Mar 29 11:57:44 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/038524.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/038524hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 cd06899 lectin_legume_LecRK_Ar 100.0 2.4E-28 5.3E-33 203.3 11.3 102 5-114 124-235 (236)
2 PF00139 Lectin_legB: Legume l 99.9 1.8E-27 3.9E-32 197.5 9.0 103 5-112 125-236 (236)
3 cd01951 lectin_L-type legume l 99.9 3.2E-22 6.9E-27 164.1 11.5 102 5-112 111-222 (223)
4 cd07308 lectin_leg-like legume 98.0 2.9E-05 6.2E-10 63.6 7.9 59 52-112 152-216 (218)
5 cd06902 lectin_ERGIC-53_ERGL E 97.5 0.00042 9.1E-09 57.7 7.9 59 52-112 155-222 (225)
6 cd06903 lectin_EMP46_EMP47 EMP 97.0 0.0047 1E-07 51.2 8.2 59 50-111 149-211 (215)
7 PF03388 Lectin_leg-like: Legu 96.5 0.0097 2.1E-07 49.5 6.7 55 51-109 158-223 (229)
8 cd06901 lectin_VIP36_VIPL VIP3 96.3 0.021 4.5E-07 48.3 8.0 59 51-112 156-221 (248)
9 KOG3839 Lectin VIP36, involved 95.6 0.022 4.8E-07 50.2 5.2 57 55-111 212-272 (351)
10 PF15065 NCU-G1: Lysosomal tra 95.3 0.011 2.3E-07 52.6 2.1 26 89-114 281-306 (350)
11 KOG3838 Mannose lectin ERGIC-5 94.6 0.11 2.4E-06 47.0 6.6 60 53-114 188-256 (497)
12 PF08693 SKG6: Transmembrane a 93.4 0.069 1.5E-06 33.1 2.1 8 157-164 32-39 (40)
13 PF04478 Mid2: Mid2 like cell 89.7 0.22 4.7E-06 39.5 1.9 28 138-165 49-79 (154)
14 PF15102 TMEM154: TMEM154 prot 89.5 0.34 7.3E-06 38.1 2.8 17 138-154 58-74 (146)
15 PF01102 Glycophorin_A: Glycop 87.4 0.34 7.3E-06 37.0 1.6 11 156-166 84-94 (122)
16 PF14610 DUF4448: Protein of u 85.3 0.76 1.6E-05 36.9 2.7 26 138-163 159-184 (189)
17 cd06900 lectin_VcfQ VcfQ bacte 85.0 4.2 9E-05 34.7 7.1 28 83-110 225-252 (255)
18 PF05454 DAG1: Dystroglycan (D 83.8 0.33 7.2E-06 42.2 0.0 24 141-164 151-174 (290)
19 PF01299 Lamp: Lysosome-associ 83.2 0.81 1.8E-05 39.5 2.1 30 138-167 270-301 (306)
20 PTZ00382 Variant-specific surf 77.9 0.95 2.1E-05 33.0 0.7 14 136-149 64-77 (96)
21 PF14575 EphA2_TM: Ephrin type 76.6 1.4 2.9E-05 30.7 1.1 11 155-165 18-28 (75)
22 PF06697 DUF1191: Protein of u 75.9 4.8 0.0001 34.9 4.5 15 139-153 215-229 (278)
23 PF14283 DUF4366: Domain of un 75.7 11 0.00023 31.6 6.4 29 41-71 72-100 (218)
24 PF14991 MLANA: Protein melan- 75.3 0.73 1.6E-05 34.8 -0.5 15 155-169 42-56 (118)
25 PHA03265 envelope glycoprotein 67.7 4.6 9.9E-05 36.3 2.6 28 136-164 349-376 (402)
26 PF03302 VSP: Giardia variant- 65.2 4.7 0.0001 36.3 2.2 19 136-154 365-383 (397)
27 PF06024 DUF912: Nucleopolyhed 61.8 5.6 0.00012 29.1 1.8 10 155-164 81-90 (101)
28 PF12877 DUF3827: Domain of un 59.6 14 0.00031 35.6 4.4 18 137-154 269-286 (684)
29 PF02009 Rifin_STEVOR: Rifin/s 56.3 5 0.00011 35.0 0.8 7 155-161 274-280 (299)
30 PF11857 DUF3377: Domain of un 54.2 20 0.00043 25.1 3.4 18 137-154 30-47 (74)
31 PF12768 Rax2: Cortical protei 49.8 14 0.0003 31.9 2.5 9 138-146 229-237 (281)
32 PF15345 TMEM51: Transmembrane 48.3 16 0.00035 30.9 2.5 10 157-166 77-86 (233)
33 PTZ00208 65 kDa invariant surf 45.5 9.3 0.0002 34.9 0.8 30 139-168 388-418 (436)
34 PF05337 CSF-1: Macrophage col 44.1 7.5 0.00016 33.7 0.0 29 137-165 224-253 (285)
35 PF12191 stn_TNFRSF12A: Tumour 43.2 7.9 0.00017 29.8 0.0 24 139-162 79-103 (129)
36 PLN03150 hypothetical protein; 42.6 25 0.00055 33.3 3.3 6 158-163 566-571 (623)
37 PF03229 Alpha_GJ: Alphavirus 42.2 40 0.00087 25.7 3.6 28 137-164 82-111 (126)
38 PF06365 CD34_antigen: CD34/Po 41.9 46 0.00099 27.6 4.3 26 138-163 101-128 (202)
39 TIGR01167 LPXTG_anchor LPXTG-m 41.1 31 0.00067 19.5 2.4 6 159-164 27-32 (34)
40 PF02480 Herpes_gE: Alphaherpe 40.7 9.1 0.0002 35.1 0.0 9 105-113 293-301 (439)
41 PF15050 SCIMP: SCIMP protein 40.7 9.9 0.00021 29.2 0.2 8 155-162 26-33 (133)
42 TIGR01478 STEVOR variant surfa 39.0 20 0.00043 31.3 1.8 7 158-164 281-287 (295)
43 PTZ00370 STEVOR; Provisional 38.8 20 0.00044 31.3 1.8 7 158-164 277-283 (296)
44 TIGR03370 PEPCTERM_Roseo varia 37.7 20 0.00043 20.1 1.1 11 155-165 15-25 (26)
45 PF13268 DUF4059: Protein of u 35.7 22 0.00048 24.7 1.3 24 143-166 13-36 (72)
46 PF08374 Protocadherin: Protoc 34.8 65 0.0014 27.0 4.1 17 138-154 38-54 (221)
47 PF12248 Methyltransf_FA: Farn 33.1 1.8E+02 0.0039 20.7 5.9 41 49-94 49-94 (102)
48 PF13908 Shisa: Wnt and FGF in 32.3 41 0.00088 26.5 2.5 11 138-148 79-89 (179)
49 PF03597 CcoS: Cytochrome oxid 31.7 36 0.00077 21.4 1.6 27 142-168 6-32 (45)
50 PRK05886 yajC preprotein trans 31.6 35 0.00076 25.5 1.9 11 155-165 17-27 (109)
51 PF11353 DUF3153: Protein of u 31.5 35 0.00076 27.8 2.0 25 73-97 117-142 (209)
52 PF14914 LRRC37AB_C: LRRC37A/B 31.2 29 0.00062 27.5 1.4 18 137-154 119-136 (154)
53 PTZ00046 rifin; Provisional 30.5 33 0.00071 30.8 1.8 6 155-160 333-338 (358)
54 TIGR02595 PEP_exosort PEP-CTER 30.0 45 0.00097 18.4 1.7 7 158-164 17-23 (26)
55 TIGR03141 cytochro_ccmD heme e 29.9 36 0.00077 21.1 1.4 11 142-152 9-19 (45)
56 TIGR03521 GldG gliding-associa 29.2 36 0.00079 31.9 2.0 12 155-166 540-551 (552)
57 TIGR01582 FDH-beta formate deh 29.0 71 0.0015 27.6 3.6 11 120-130 228-238 (283)
58 COG4736 CcoQ Cbb3-type cytochr 28.9 56 0.0012 21.9 2.3 10 156-165 26-35 (60)
59 PF04689 S1FA: DNA binding pro 28.4 67 0.0015 22.0 2.6 19 136-154 11-29 (69)
60 PF10661 EssA: WXG100 protein 26.2 43 0.00092 26.1 1.6 21 141-161 121-141 (145)
61 PHA03264 envelope glycoprotein 25.8 73 0.0016 29.0 3.2 18 136-153 359-376 (416)
62 PRK07021 fliL flagellar basal 25.5 69 0.0015 25.0 2.7 8 155-162 35-42 (162)
63 PHA03291 envelope glycoprotein 24.2 58 0.0013 29.4 2.2 25 138-162 288-313 (401)
64 PRK15471 chain length determin 23.5 60 0.0013 28.5 2.2 7 136-142 293-299 (325)
65 PF06809 NPDC1: Neural prolife 23.1 1.4E+02 0.0029 26.7 4.2 7 89-95 144-150 (341)
66 PRK10381 LPS O-antigen length 23.1 46 0.001 29.8 1.4 9 136-144 337-345 (377)
67 PLN00113 leucine-rich repeat r 23.0 89 0.0019 30.6 3.5 12 139-150 630-641 (968)
68 PRK00523 hypothetical protein; 22.5 53 0.0011 22.9 1.3 8 155-162 21-28 (72)
69 PF03988 DUF347: Repeat of Unk 22.0 62 0.0014 20.9 1.5 16 139-154 25-40 (55)
70 KOG3637 Vitronectin receptor, 21.4 1.1E+02 0.0024 31.2 3.8 14 138-151 979-992 (1030)
71 COG4282 SMI1 Protein involved 21.2 62 0.0013 26.4 1.6 15 10-24 123-137 (191)
72 PRK11638 lipopolysaccharide bi 21.2 43 0.00094 29.6 0.8 8 157-164 334-341 (342)
73 PF02699 YajC: Preprotein tran 20.7 73 0.0016 22.2 1.8 7 156-162 16-22 (82)
74 TIGR00847 ccoS cytochrome oxid 20.3 1.1E+02 0.0023 19.8 2.4 13 155-167 20-32 (51)
75 PF07332 DUF1469: Protein of u 20.3 56 0.0012 23.8 1.2 8 167-174 109-116 (121)
76 TIGR01006 polys_exp_MPA1 polys 20.1 50 0.0011 26.7 0.9 14 158-171 197-210 (226)
No 1
>cd06899 lectin_legume_LecRK_Arcelin_ConA legume lectins, lectin-like receptor kinases, arcelin, concanavalinA, and alpha-amylase inhibitor. This alignment model includes the legume lectins (also known as agglutinins), the arcelin (also known as phytohemagglutinin-L) family of lectin-like defense proteins, the LecRK family of lectin-like receptor kinases, concanavalinA (ConA), and an alpha-amylase inhibitor. Arcelin is a major seed glycoprotein discovered in kidney beans (Phaseolus vulgaris) that has insecticidal properties and protects the seeds from predation by larvae of various bruchids. Arcelin is devoid of monosaccharide binding properties and lacks a key metal-binding loop that is present in other members of this family. Phytohaemagglutinin (PHA) is a lectin found in plants, especially beans, that affects cell metabolism by inducing mitosis and by altering the permeability of the cell membrane to various proteins. PHA agglutinates most mammalian red blood cell types by bindin
Probab=99.95 E-value=2.4e-28 Score=203.25 Aligned_cols=102 Identities=47% Similarity=0.685 Sum_probs=90.5
Q ss_pred eeecCCCCCCCCCeeEEEcCCCccccceecceecCCCCCcccccccCCCcEEEEEEeeCCCceEEEee----------ee
Q 038524 5 RYQSIQFNDTSNNHVGIDVNRLTSVGSVPATYFSDEKGTNKSLKLISGDPMQTWIDYMGSEKLHEIQC----------LS 74 (175)
Q Consensus 5 T~~n~e~~D~~~nHVGIdiNs~~S~~s~~~~~~~~~~~~~~~~~l~sG~~~~vwIdYd~~~~~L~v~l----------Pl 74 (175)
|++|.+++||++||||||+|++.|..+.. |+. ..+.|.+|+.++|||+||+.+++|+|.+ |+
T Consensus 124 T~~n~~~~D~~~nHigIdvn~~~S~~~~~---~~~-----~~~~l~~g~~~~v~I~Y~~~~~~L~V~l~~~~~~~~~~~~ 195 (236)
T cd06899 124 TFQNPEFGDPDDNHVGIDVNSLVSVKAGY---WDD-----DGGKLKSGKPMQAWIDYDSSSKRLSVTLAYSGVAKPKKPL 195 (236)
T ss_pred cccCcccCCCCCCeEEEEcCCcccceeec---ccc-----ccccccCCCeEEEEEEEcCCCCEEEEEEEeCCCCCCcCCE
Confidence 67888889999999999999987776544 322 2345789999999999999999999988 68
Q ss_pred eeeeecCCccCCCceEEEEEeecCCceeeeEEeeeEEeeC
Q 038524 75 LSTSVDLSQLLLDTMCVGFSAATGSLATEHYILGCSLNKS 114 (175)
Q Consensus 75 ls~~idLs~~l~~~~yVGFSAsTG~~~~~h~IlsWsF~~~ 114 (175)
|+.++||+.+|+++||||||||||...|.|+|++|+|++.
T Consensus 196 ls~~vdL~~~l~~~~~vGFSasTG~~~~~h~i~sWsF~s~ 235 (236)
T cd06899 196 LSYPVDLSKVLPEEVYVGFSASTGLLTELHYILSWSFSSN 235 (236)
T ss_pred EEEeccHHHhCCCceEEEEEeEcCCCcceEEEEEEEEEcC
Confidence 9999999999999999999999999999999999999875
No 2
>PF00139 Lectin_legB: Legume lectin domain; InterPro: IPR001220 Legume lectins are one of the largest lectin families with more than 70 lectins reported. Leguminous plant lectins resemble each other in their physicochemical properties although they differ in their carbohydrate specificities. They consist of two or four subunits with relative molecular mass of 30 kDa and each subunit has one carbohydrate-binding site. The interaction with sugars requires tightly bound calcium and manganese ions. The structural similarities of these lectins are reported by the primary structural analyses and X-ray crystallographic studies. X-ray studies have shown that the folding of the polypeptide chains in the region of the carbohydrate-binding sites is also similar, despite differences in the primary sequences. The carbohydrate-binding sites of these lectins consist of two conserved amino acids on beta pleated sheets. One of these loops contains transition metals, calcium and manganese, which keep the amino acid residues of the sugar-binding site at the required positions. Amino acid sequences of this loop play an important role in the carbohydrate-binding specificities of these lectins. These lectins bind either glucose/mannose or galactose. The exact function of legume lectins is not known but they may be involved in the attachment of nitrogen-fixing bacteria to legumes and in the protection against pathogens. Some legume lectins are proteolytically processed to produce two chains, beta (which corresponds to the N-terminal) and alpha (C-terminal) (IPR000985 from INTERPRO). The lectin concanavalin A (conA) from jack bean is exceptional in that the two chains are transposed and ligated (by formation of a new peptide bond). The N terminus of mature conA thus corresponds to that of the alpha chain and the C terminus to the beta chain.; GO: 0005488 binding; PDB: 1VLN_B 2GDF_C 2JE9_C 2JEC_C 1DGL_B 2P37_B 2CWM_A 2P34_D 2OW4_A 3IPV_B ....
Probab=99.94 E-value=1.8e-27 Score=197.46 Aligned_cols=103 Identities=40% Similarity=0.603 Sum_probs=89.9
Q ss_pred eeecCCCCCCCCCeeEEEcCCCccccceecceecCCCCCcccccccCCCcEEEEEEeeCCCceEEEee---------eee
Q 038524 5 RYQSIQFNDTSNNHVGIDVNRLTSVGSVPATYFSDEKGTNKSLKLISGDPMQTWIDYMGSEKLHEIQC---------LSL 75 (175)
Q Consensus 5 T~~n~e~~D~~~nHVGIdiNs~~S~~s~~~~~~~~~~~~~~~~~l~sG~~~~vwIdYd~~~~~L~v~l---------Pll 75 (175)
|++|++++||++||||||+|++.|..+.+++++. .....|.+|+.++|||+||+.+++|.|++ |+|
T Consensus 125 T~~N~~~~d~~~nHIgI~~n~~~s~~~~~~~~~~-----~~~~~l~~g~~~~v~I~Yd~~~~~L~V~l~~~~~~~~~~~l 199 (236)
T PF00139_consen 125 TYKNPEYNDPDDNHIGIDVNSVVSNKTASAGYYS-----SPSFSLSDGKWHTVWIDYDASTKRLSVYLDDNSSKPSSPVL 199 (236)
T ss_dssp TSTCGGGTTTSSSEEEEEESSSSESEEEE----E-----EEEHHHGTTSEEEEEEEEETTTTEEEEEEEETTTTSEEEEE
T ss_pred eeecccccccCCCEEEEECCCCcccccccccccc-----cccccccCCcEEEEEEEEcCCccEEEEEEecccCCCcceeE
Confidence 5678889999999999999999999988877652 35678999999999999999999999988 689
Q ss_pred eeeecCCccCCCceEEEEEeecCCceeeeEEeeeEEe
Q 038524 76 STSVDLSQLLLDTMCVGFSAATGSLATEHYILGCSLN 112 (175)
Q Consensus 76 s~~idLs~~l~~~~yVGFSAsTG~~~~~h~IlsWsF~ 112 (175)
+..+||+.+|+++||||||||||...|.|.|++|+|+
T Consensus 200 ~~~vdL~~~l~~~v~vGFsasTG~~~~~h~I~sW~F~ 236 (236)
T PF00139_consen 200 SVNVDLSAVLPEQVYVGFSASTGGSYQTHDILSWSFS 236 (236)
T ss_dssp EEE--HHHHSCSEEEEEEEEEESSSSEEEEEEEEEEE
T ss_pred EEEEchHHhcCCCcEEEEEeecCCCcceEEEEEEEeC
Confidence 9999999999999999999999999999999999995
No 3
>cd01951 lectin_L-type legume lectins. The L-type (legume-type) lectins are a highly diverse family of carbohydrate binding proteins that generally display no enzymatic activity toward the sugars they bind. This family includes arcelin, concanavalinA, the lectin-like receptor kinases, the ERGIC-53/VIP36/EMP46 type1 transmembrane proteins, and an alpha-amylase inhibitor. L-type lectins have a dome-shaped beta-barrel carbohydrate recognition domain with a curved seven-stranded beta-sheet referred to as the "front face" and a flat six-stranded beta-sheet referred to as the "back face". This domain homodimerizes so that adjacent back sheets form a contiguous 12-stranded sheet and homotetramers occur by a back-to-back association of these homodimers. Though L-type lectins exhibit both sequence and structural similarity to one another, their carbohydrate binding specificities differ widely.
Probab=99.88 E-value=3.2e-22 Score=164.05 Aligned_cols=102 Identities=28% Similarity=0.309 Sum_probs=84.2
Q ss_pred eeecCCCCCCCCCeeEEEcCCCcccc--ceecceecCCCCCcccccccCCCcEEEEEEeeCCCceEEEee--------ee
Q 038524 5 RYQSIQFNDTSNNHVGIDVNRLTSVG--SVPATYFSDEKGTNKSLKLISGDPMQTWIDYMGSEKLHEIQC--------LS 74 (175)
Q Consensus 5 T~~n~e~~D~~~nHVGIdiNs~~S~~--s~~~~~~~~~~~~~~~~~l~sG~~~~vwIdYd~~~~~L~v~l--------Pl 74 (175)
|++|.+++||+.||||||+|+..+.. ..+.+++. .+....+|+.++|||+||+.+++|+|.+ |.
T Consensus 111 T~~N~~~~dp~~~higi~~n~~~~~~~~~~~~~~~~------~~~~~~~g~~~~v~I~Y~~~~~~L~v~l~~~~~~~~~~ 184 (223)
T cd01951 111 TYKNDDNNDPNGNHISIDVNGNGNNTALATSLGSAS------LPNGTGLGNEHTVRITYDPTTNTLTVYLDNGSTLTSLD 184 (223)
T ss_pred ccccCCCCCCCCCEEEEEcCCCCCCcccccccceee------CCCccCCCCEEEEEEEEeCCCCEEEEEECCCCcccccc
Confidence 78898888999999999999987641 11222221 1122223999999999999999999988 58
Q ss_pred eeeeecCCccCCCceEEEEEeecCCceeeeEEeeeEEe
Q 038524 75 LSTSVDLSQLLLDTMCVGFSAATGSLATEHYILGCSLN 112 (175)
Q Consensus 75 ls~~idLs~~l~~~~yVGFSAsTG~~~~~h~IlsWsF~ 112 (175)
++.++||+..++++||||||||||...+.|+|++|+|+
T Consensus 185 l~~~~~l~~~~~~~~yvGFTAsTG~~~~~h~V~~wsf~ 222 (223)
T cd01951 185 ITIPVDLIQLGPTKAYFGFTASTGGLTNLHDILNWSFT 222 (223)
T ss_pred EEEeeeecccCCCcEEEEEEcccCCCcceeEEEEEEec
Confidence 89999999999999999999999999999999999996
No 4
>cd07308 lectin_leg-like legume-like lectins: ERGIC-53, ERGL, VIP36, VIPL, EMP46, and EMP47. The legume-like (leg-like) lectins are eukaryotic intracellular sugar transport proteins with a carbohydrate recognition domain similar to that of the legume lectins. This domain binds high-mannose-type oligosaccharides for transport from the endoplasmic reticulum to the Golgi complex. These leg-like lectins include ERGIC-53, ERGL, VIP36, VIPL, EMP46, EMP47, and the UIP5 (ULP1-interacting protein 5) precursor protein. Leg-like lectins have different intracellular distributions and dynamics in the endoplasmic reticulum-Golgi system of the secretory pathway and interact with N-glycans of glycoproteins in a calcium-dependent manner, suggesting a role in glycoprotein sorting and trafficking. L-type lectins have a dome-shaped beta-barrel carbohydrate recognition domain with a curved seven-stranded beta-sheet referred to as the "front face" and a flat six-stranded beta-sheet referred to as the "ba
Probab=97.99 E-value=2.9e-05 Score=63.62 Aligned_cols=59 Identities=24% Similarity=0.389 Sum_probs=44.1
Q ss_pred CCcEEEEEEeeCCCceEEEee-e----eeeeeecCCc-cCCCceEEEEEeecCCceeeeEEeeeEEe
Q 038524 52 GDPMQTWIDYMGSEKLHEIQC-L----SLSTSVDLSQ-LLLDTMCVGFSAATGSLATEHYILGCSLN 112 (175)
Q Consensus 52 G~~~~vwIdYd~~~~~L~v~l-P----lls~~idLs~-~l~~~~yVGFSAsTG~~~~~h~IlsWsF~ 112 (175)
+++++++|.|+ .+.|.|.+ + --..-.++.. .+++.+|+||||+||...+.|.|++|.+.
T Consensus 152 ~~~~~~~I~y~--~~~l~v~i~~~~~~~~~~c~~~~~~~l~~~~y~G~sA~tg~~~d~~dIls~~~~ 216 (218)
T cd07308 152 NAPTTLRISYL--NNTLKVDITYSEGNNWKECFTVEDVILPSQGYFGFSAQTGDLSDNHDILSVHTY 216 (218)
T ss_pred CCCeEEEEEEE--CCEEEEEEeCCCCCCccEEEEcCCcccCCCCEEEEEeccCCCcCcEEEEEEEee
Confidence 68999999999 56777665 1 0111222333 46788999999999999999999999874
No 5
>cd06902 lectin_ERGIC-53_ERGL ERGIC-53 and ERGL type 1 transmembrane proteins, N-terminal lectin domain. ERGIC-53 and ERGL, N-terminal carbohydrate recognition domain. ERGIC-53 and ERGL are eukaryotic mannose-binding type 1 transmembrane proteins of the early secretory pathway that transport newly synthesized glycoproteins from the endoplasmic reticulum (ER) to the ER-Golgi intermediate compartment (ERGIC). ERGIC-53 and ERGL have an N-terminal lectin-like carbohydrate recognition domain (represented by this alignment model) as well as a C-terminal transmembrane domain. ERGIC-53 functions as a 'cargo receptor' to facilitate the export of glycoproteins with different characteristics from the ER, while the ERGIC-53-like protein (ERGL) which may act as a regulator of ERGIC-53. In mammals, ERGIC-53 forms a complex with MCFD2 (multi-coagulation factor deficiency 2) which then recruits blood coagulation factors V and VIII. Mutations in either MCFD2 or ERGIC-53 cause a mild form of inherite
Probab=97.53 E-value=0.00042 Score=57.74 Aligned_cols=59 Identities=24% Similarity=0.293 Sum_probs=42.8
Q ss_pred CCcEEEEEEeeCCCceEEEee-----e---eeeeeecCCc-cCCCceEEEEEeecCCceeeeEEeeeEEe
Q 038524 52 GDPMQTWIDYMGSEKLHEIQC-----L---SLSTSVDLSQ-LLLDTMCVGFSAATGSLATEHYILGCSLN 112 (175)
Q Consensus 52 G~~~~vwIdYd~~~~~L~v~l-----P---lls~~idLs~-~l~~~~yVGFSAsTG~~~~~h~IlsWsF~ 112 (175)
..+.++.|.|... .|.|.+ + ....-.++.. .||+..|+||||+||...+.|.|++|+|.
T Consensus 155 ~~p~~~rI~Y~~~--~l~V~~d~~~~~~~~~~~~Cf~~~~v~LP~~~yfGiSA~Tg~l~d~hDIls~~~~ 222 (225)
T cd06902 155 PYPVRAKITYYQN--VLTVSINNGFTPNKDDYELCTRVENMVLPPNGYFGVSAATGGLADDHDVLSFLTF 222 (225)
T ss_pred CCCeEEEEEEECC--eEEEEEeCCcCCCCCcccEEEecCCeeCCCCCEEEEEecCCCCCCcEeEEEEEEe
Confidence 5689999999985 466543 1 0111122222 46778999999999999999999999985
No 6
>cd06903 lectin_EMP46_EMP47 EMP46 and EMP47 type 1 transmembrane proteins, N-terminal lectin domain. EMP46 and EMP47, N-terminal carbohydrate recognition domain. EMP46 and EMP47 are fungal type-I transmembrane proteins that cycle between the endoplasmic reticulum and the golgi apparatus and are thought to function as cargo receptors that transport newly synthesized glycoproteins. EMP47 is a receptor for EMP46 responsible for the selective transport of EMP46 by forming hetero-oligomerization between the two proteins. EMP46 and EMP47 have an N-terminal lectin-like carbohydrate recognition domain (represented by this alignment model) as well as a C-terminal transmembrane domain. EMP46 and EMP47 are 45% sequence-identical to one another and have sequence homology to a class of intracellular lectins defined by ERGIC-53 and VIP36. L-type lectins have a dome-shaped beta-barrel carbohydrate recognition domain with a curved seven-stranded beta-sheet referred to as the "front face" and a flat s
Probab=96.95 E-value=0.0047 Score=51.21 Aligned_cols=59 Identities=27% Similarity=0.334 Sum_probs=44.1
Q ss_pred cCCCcEEEEEEeeCCCceEEEee---eeeee-eecCCccCCCceEEEEEeecCCceeeeEEeeeEE
Q 038524 50 ISGDPMQTWIDYMGSEKLHEIQC---LSLST-SVDLSQLLLDTMCVGFSAATGSLATEHYILGCSL 111 (175)
Q Consensus 50 ~sG~~~~vwIdYd~~~~~L~v~l---Plls~-~idLs~~l~~~~yVGFSAsTG~~~~~h~IlsWsF 111 (175)
.++.+.++.|.|....+.|.|.+ .++.. .+.|.. ...|+||||+||...+.|.|++-.+
T Consensus 149 n~~~p~~iri~Y~~~~~~l~v~vd~~~Cf~~~~v~lP~---~~y~fGiSAaTg~~~d~hdIl~~~~ 211 (215)
T cd06903 149 DSGVPSTIRLSYDALNSLFKVQVDNRLCFQTDKVQLPQ---GGYRFGITAANADNPESFEILKLKV 211 (215)
T ss_pred CCCCCEEEEEEEECCCCEEEEEECCCEEEecCCeecCC---CCCEEEEEEcCCCCCCcEEEEEEEE
Confidence 45678999999999777888776 33332 333332 4568999999999999999996544
No 7
>PF03388 Lectin_leg-like: Legume-like lectin family; InterPro: IPR005052 Lectins are structurally diverse proteins that bind to specific carbohydrates. This family includes the VIP36 and ERGIC-53 lectins. These two proteins were the first members of the family of animal lectins similar to the leguminous plant lectins []. The alignment for this family is towards the N terminus, where the similarity of VIP36 and ERGIC-53 is greatest. Although they have been identified as a family of animal lectins, this alignment also includes yeast sequences[]. ERGIC-53 is a 53kDa protein, localised to the intermediate region between the endoplasmic reticulum and the Golgi apparatus (ER-Golgi-Intermediate Compartment, ERGIC). It was identified as a calcium-dependent, mannose-specific lectin []. Its dysfunction has been associated with combined factors V and VIII deficiency, suggesting an important and substrate-specific role for ERGIC-53 in the glycoprotein-secreting pathway [,]. The L-type lectin-like domain has an overall globular shape composed of a beta-sandwich of two major twisted antiparallel beta-sheets. The beta-sandwich comprises a major concave beta-sheet and a minor convex beta-sheet, in a variation of the jelly roll fold [, , , ]. ; GO: 0016020 membrane; PDB: 3A4U_A 3LCP_B 2A6Z_A 2A71_C 2A70_B 2A6Y_A 2A6X_A 2A6W_B 2A6V_B 2E6V_B ....
Probab=96.45 E-value=0.0097 Score=49.47 Aligned_cols=55 Identities=35% Similarity=0.411 Sum_probs=40.6
Q ss_pred CCCcEEEEEEeeCCCceEEEe---e-------eeeee-eecCCccCCCceEEEEEeecCCceeeeEEeee
Q 038524 51 SGDPMQTWIDYMGSEKLHEIQ---C-------LSLST-SVDLSQLLLDTMCVGFSAATGSLATEHYILGC 109 (175)
Q Consensus 51 sG~~~~vwIdYd~~~~~L~v~---l-------Plls~-~idLs~~l~~~~yVGFSAsTG~~~~~h~IlsW 109 (175)
.+.+.++.|.|......+.+. + .++.. .+ .||+..|+||||+||...+.|.|++=
T Consensus 158 ~~~p~~~ri~Y~~~~l~v~id~~~~~~~~~~~~Cf~~~~v----~LP~~~yfGvSA~Tg~~~d~hdi~s~ 223 (229)
T PF03388_consen 158 SDVPTRIRISYSKNTLTVSIDSNYLKNQDDWELCFTTDGV----DLPEGYYFGVSAATGELSDNHDILSV 223 (229)
T ss_dssp ESSEEEEEEEEETTEEEEEEETSCCSECCTTEEEEEESTE----EGGSSBEEEEEEEESSSGGEEEEEEE
T ss_pred CCCCEEEEEEEECCeEEEEEecccccCCcCCcEEEEcCCe----ecCCCCEEEEEecCCCCCCcEEEEEE
Confidence 456789999999875555444 1 34443 23 35778899999999999999999964
No 8
>cd06901 lectin_VIP36_VIPL VIP36 and VIPL type 1 transmembrane proteins, lectin domain. The vesicular integral protein of 36 kDa (VIP36) is a type 1 transmembrane protein of the mammalian early secretory pathway that acts as a cargo receptor transporting high mannose type glycoproteins between the Golgi and the endoplasmic reticulum (ER). Lectins of the early secretory pathway are involved in the selective transport of newly synthesized glycoproteins from the ER to the ER-Golgi intermediate compartment (ERGIC). The most prominent cycling lectin is the mannose-binding type1 membrane protein ERGIC-53, which functions as a cargo receptor to facilitate export of glycoproteins from the ER. L-type lectins have a dome-shaped beta-barrel carbohydrate recognition domain with a curved seven-stranded beta-sheet referred to as the "front face" and a flat six-stranded beta-sheet referred to as the "back face". This domain homodimerizes so that adjacent back sheets form a contiguous 12-stranded she
Probab=96.30 E-value=0.021 Score=48.32 Aligned_cols=59 Identities=22% Similarity=0.166 Sum_probs=40.4
Q ss_pred CCCcEEEEEEeeCCCceEEEee-------eeeeeeecCCccCCCceEEEEEeecCCceeeeEEeeeEEe
Q 038524 51 SGDPMQTWIDYMGSEKLHEIQC-------LSLSTSVDLSQLLLDTMCVGFSAATGSLATEHYILGCSLN 112 (175)
Q Consensus 51 sG~~~~vwIdYd~~~~~L~v~l-------Plls~~idLs~~l~~~~yVGFSAsTG~~~~~h~IlsWsF~ 112 (175)
.+.+.+++|.|......|.+.. .++.. -.-.||...|+||||+||...+.|.|++-.+-
T Consensus 156 ~~~~t~~rI~Y~~~~l~v~vd~~~~~~w~~Cf~~---~~v~LP~~~yfGiSA~Tg~~sd~hdIlsv~~~ 221 (248)
T cd06901 156 KDHDTFVAIRYSKGRLTVMTDIDGKNEWKECFDV---TGVRLPTGYYFGASAATGDLSDNHDIISMKLY 221 (248)
T ss_pred CCCCeEEEEEEECCeEEEEEecCCCCceeeeEEe---CCeecCCCCEEEEEecCCCCCCcEEEEEEEEe
Confidence 4567889999997543333332 12221 12245677999999999999999999976664
No 9
>KOG3839 consensus Lectin VIP36, involved in the transport of glycoproteins carrying high mannose-type glycans [Intracellular trafficking, secretion, and vesicular transport]
Probab=95.62 E-value=0.022 Score=50.16 Aligned_cols=57 Identities=25% Similarity=0.205 Sum_probs=39.6
Q ss_pred EEEEEEeeCCCceEEEee--e-eeeeeecCCc-cCCCceEEEEEeecCCceeeeEEeeeEE
Q 038524 55 MQTWIDYMGSEKLHEIQC--L-SLSTSVDLSQ-LLLDTMCVGFSAATGSLATEHYILGCSL 111 (175)
Q Consensus 55 ~~vwIdYd~~~~~L~v~l--P-lls~~idLs~-~l~~~~yVGFSAsTG~~~~~h~IlsWsF 111 (175)
..+-|.|+..+.++...+ | -...-.+|.. .||.--|+|+||+||...+.|.|++=.+
T Consensus 212 t~~~iry~~~~l~~~~dl~~~~~~~~c~~~n~v~lp~g~~fg~SasTGdlSd~HdivS~kl 272 (351)
T KOG3839|consen 212 TLVVIRYEKKTLSISIDLEGPNEWIDCFSLNNVELPLGYFFGVSASTGDLSDSHDIVSLKL 272 (351)
T ss_pred ceeEEEecCCceEEEEecCCCceeeeeeeecceecccceEEeeeeccCccchhhHHHHhhh
Confidence 456788888554444444 4 2344455555 4567789999999999999999997543
No 10
>PF15065 NCU-G1: Lysosomal transcription factor, NCU-G1
Probab=95.26 E-value=0.011 Score=52.60 Aligned_cols=26 Identities=12% Similarity=0.123 Sum_probs=23.4
Q ss_pred eEEEEEeecCCceeeeEEeeeEEeeC
Q 038524 89 MCVGFSAATGSLATEHYILGCSLNKS 114 (175)
Q Consensus 89 ~yVGFSAsTG~~~~~h~IlsWsF~~~ 114 (175)
+.|=|.++++..+..+..++|+|...
T Consensus 281 ~nvSFG~~gDgfY~~t~ylsWt~~~G 306 (350)
T PF15065_consen 281 LNVSFGTSGDGFYWATNYLSWTFLIG 306 (350)
T ss_pred EEEEeccCCCCcccccceEEEEEecc
Confidence 77888888888899999999999986
No 11
>KOG3838 consensus Mannose lectin ERGIC-53, involved in glycoprotein traffic [Intracellular trafficking, secretion, and vesicular transport]
Probab=94.57 E-value=0.11 Score=47.03 Aligned_cols=60 Identities=32% Similarity=0.419 Sum_probs=42.8
Q ss_pred CcEEEEEEeeCCCceEEEee-----ee--eeeeecCCc-cCCCceEEEEEeecCCceeeeEEeeeE-EeeC
Q 038524 53 DPMQTWIDYMGSEKLHEIQC-----LS--LSTSVDLSQ-LLLDTMCVGFSAATGSLATEHYILGCS-LNKS 114 (175)
Q Consensus 53 ~~~~vwIdYd~~~~~L~v~l-----Pl--ls~~idLs~-~l~~~~yVGFSAsTG~~~~~h~IlsWs-F~~~ 114 (175)
-++.|.|+|-+ ++|.|-+ |. ...-++... +||..-|+|.|||||....-|+||+.. |+..
T Consensus 188 yPvRarItY~~--nvLtv~innGmtp~d~yE~C~rve~~~lp~nGyFGvSAATGgLADDHDVl~FltfsL~ 256 (497)
T KOG3838|consen 188 YPVRARITYYG--NVLTVMINNGMTPSDDYEFCVRVENLLLPPNGYFGVSAATGGLADDHDVLSFLTFSLS 256 (497)
T ss_pred CCceEEEEEec--cEEEEEEcCCCCCCCCcceeEeccceeccCCCeeeeeecccccccccceeeeEEeeec
Confidence 37899999986 4677655 43 111233334 357889999999999999999999874 4443
No 12
>PF08693 SKG6: Transmembrane alpha-helix domain; InterPro: IPR014805 SKG6 and AXL2 are membrane proteins that show polarised intracellular localisation [, ]. This entry represents the highly conserved transmembrane alpha-helical domain found in these proteins [, ]. The full-length AXL2 protein has a negative regulatory function in cytokinesis [].
Probab=93.36 E-value=0.069 Score=33.13 Aligned_cols=8 Identities=25% Similarity=0.488 Sum_probs=3.5
Q ss_pred hhhheecc
Q 038524 157 AVYIVRKK 164 (175)
Q Consensus 157 ~~~~~rr~ 164 (175)
+++++||+
T Consensus 32 l~~~~rR~ 39 (40)
T PF08693_consen 32 LFFWYRRK 39 (40)
T ss_pred hheEEecc
Confidence 33344544
No 13
>PF04478 Mid2: Mid2 like cell wall stress sensor; InterPro: IPR007567 This family represents a region near the C terminus of Mid2, which contains a transmembrane region. The remainder of the protein sequence is serine-rich and of low complexity, and is therefore impossible to align accurately. Mid2 is thought to act as a mechanosensor of cell wall stress. The C-terminal cytoplasmic region of Mid2 is known to interact with Rom2, a guanine nucleotide exchange factor (GEF) for Rho1, which is part of the cell wall integrity signalling pathway [].
Probab=89.73 E-value=0.22 Score=39.51 Aligned_cols=28 Identities=18% Similarity=0.164 Sum_probs=11.6
Q ss_pred eEEeh--hHHHHHHHHHHH-Hhhhhheeccc
Q 038524 138 LVTCV--ILVALVSVLTTV-GAAVYIVRKKK 165 (175)
Q Consensus 138 ~l~i~--l~~~~~~~~~~~-~~~~~~~rr~~ 165 (175)
.++|| ++++++++++++ ++|++++|++|
T Consensus 49 nIVIGvVVGVGg~ill~il~lvf~~c~r~kk 79 (154)
T PF04478_consen 49 NIVIGVVVGVGGPILLGILALVFIFCIRRKK 79 (154)
T ss_pred cEEEEEEecccHHHHHHHHHhheeEEEeccc
Confidence 34444 444444444333 33334444443
No 14
>PF15102 TMEM154: TMEM154 protein family
Probab=89.46 E-value=0.34 Score=38.12 Aligned_cols=17 Identities=18% Similarity=0.301 Sum_probs=8.4
Q ss_pred eEEehhHHHHHHHHHHH
Q 038524 138 LVTCVILVALVSVLTTV 154 (175)
Q Consensus 138 ~l~i~l~~~~~~~~~~~ 154 (175)
.+.|+++++++++++++
T Consensus 58 iLmIlIP~VLLvlLLl~ 74 (146)
T PF15102_consen 58 ILMILIPLVLLVLLLLS 74 (146)
T ss_pred EEEEeHHHHHHHHHHHH
Confidence 45666665444333333
No 15
>PF01102 Glycophorin_A: Glycophorin A; InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=87.45 E-value=0.34 Score=37.02 Aligned_cols=11 Identities=27% Similarity=0.217 Sum_probs=4.5
Q ss_pred hhhhheecccc
Q 038524 156 AAVYIVRKKKY 166 (175)
Q Consensus 156 ~~~~~~rr~~~ 166 (175)
++|++||++|+
T Consensus 84 i~y~irR~~Kk 94 (122)
T PF01102_consen 84 ISYCIRRLRKK 94 (122)
T ss_dssp HHHHHHHHS--
T ss_pred HHHHHHHHhcc
Confidence 34445555544
No 16
>PF14610 DUF4448: Protein of unknown function (DUF4448)
Probab=85.33 E-value=0.76 Score=36.94 Aligned_cols=26 Identities=15% Similarity=0.156 Sum_probs=15.8
Q ss_pred eEEehhHHHHHHHHHHHHhhhhheec
Q 038524 138 LVTCVILVALVSVLTTVGAAVYIVRK 163 (175)
Q Consensus 138 ~l~i~l~~~~~~~~~~~~~~~~~~rr 163 (175)
.+.|+|+++.+++++++.+++++.||
T Consensus 159 ~laI~lPvvv~~~~~~~~~~~~~~R~ 184 (189)
T PF14610_consen 159 ALAIALPVVVVVLALIMYGFFFWNRK 184 (189)
T ss_pred eEEEEccHHHHHHHHHHHhhheeecc
Confidence 78888888766665555333333333
No 17
>cd06900 lectin_VcfQ VcfQ bacterial pilus biogenesis protein, lectin domain. This family includes bacterial proteins homologous to the VcfQ (also known as MshQ) bacterial pilus biogenesis protein. VcfQ is encoded by the vcfQ gene of the type IV pilus gene cluster of Vibrio cholerae and is essential for type IV pilus assembly. VcfQ has a Laminin G-like domain as well as an L-type lectin domain.
Probab=85.04 E-value=4.2 Score=34.75 Aligned_cols=28 Identities=18% Similarity=0.338 Sum_probs=24.5
Q ss_pred ccCCCceEEEEEeecCCceeeeEEeeeE
Q 038524 83 QLLLDTMCVGFSAATGSLATEHYILGCS 110 (175)
Q Consensus 83 ~~l~~~~yVGFSAsTG~~~~~h~IlsWs 110 (175)
.-+|+..+++|++|||..+-.|.|-...
T Consensus 225 ~avP~~f~lS~TgSTGgstN~HEIdnf~ 252 (255)
T cd06900 225 DAIPENFYLSFTGSTGGSTNTHEIDNFQ 252 (255)
T ss_pred CCCCccEEEEEEecCCCcccceeecceE
Confidence 6789999999999999999999987543
No 18
>PF05454 DAG1: Dystroglycan (Dystrophin-associated glycoprotein 1); InterPro: IPR008465 Dystroglycan is one of the dystrophin-associated glycoproteins, which is encoded by a 5.5 kb transcript in Homo sapiens. The protein product is cleaved into two non-covalently associated subunits, [alpha] (N-terminal) and [beta] (C-terminal). In skeletal muscle the dystroglycan complex works as a transmembrane linkage between the extracellular matrix and the cytoskeleton [alpha]-dystroglycan is extracellular and binds to merosin ([alpha]-2 laminin) in the basement membrane, while [beta]-dystroglycan is a transmembrane protein and binds to dystrophin, which is a large rod-like cytoskeletal protein, absent in Duchenne muscular dystrophy patients. Dystrophin binds to intracellular actin cables. In this way, the dystroglycan complex, which links the extracellular matrix to the intracellular actin cables, is thought to provide structural integrity in muscle tissues. The dystroglycan complex is also known to serve as an agrin receptor in muscle, where it may regulate agrin-induced acetylcholine receptor clustering at the neuromuscular junction. There is also evidence which suggests the function of dystroglycan as a part of the signal transduction pathway because it is shown that Grb2, a mediator of the Ras-related signal pathway, can interact with the cytoplasmic domain of dystroglycan. In general, aberrant expression of dystrophin-associated protein complex underlies the pathogenesis of Duchenne muscular dystrophy, Becker muscular dystrophy and severe childhood autosomal recessive muscular dystrophy. Interestingly, no genetic disease has been described for either [alpha]- or [beta]-dystroglycan. Dystroglycan is widely distributed in non-muscle tissues as well as in muscle tissues. During epithelial morphogenesis of kidney, the dystroglycan complex is shown to act as a receptor for the basement membrane. Dystroglycan expression in Mus musculus brain and neural retina has also been reported. However, the physiological role of dystroglycan in non-muscle tissues has remained unclear [].; PDB: 1EG4_P.
Probab=83.82 E-value=0.33 Score=42.16 Aligned_cols=24 Identities=17% Similarity=0.325 Sum_probs=0.0
Q ss_pred ehhHHHHHHHHHHHHhhhhheecc
Q 038524 141 CVILVALVSVLTTVGAAVYIVRKK 164 (175)
Q Consensus 141 i~l~~~~~~~~~~~~~~~~~~rr~ 164 (175)
+++.++++++++++++++++||||
T Consensus 151 paVVI~~iLLIA~iIa~icyrrkR 174 (290)
T PF05454_consen 151 PAVVIAAILLIAGIIACICYRRKR 174 (290)
T ss_dssp ------------------------
T ss_pred HHHHHHHHHHHHHHHHHHhhhhhh
Confidence 333333333333333344444443
No 19
>PF01299 Lamp: Lysosome-associated membrane glycoprotein (Lamp); InterPro: IPR002000 Lysosome-associated membrane glycoproteins (lamp) [] are integral membrane proteins, specific to lysosomes, and whose exact biological function is not yet clear. Structurally, the lamp proteins consist of two internally homologous lysosome-luminal domains separated by a proline-rich hinge region; at the C-terminal extremity there is a transmembrane region (TM) followed by a very short cytoplasmic tail (C). In each of the duplicated domains, there are two conserved disulphide bonds. This structure is schematically represented in the figure below. +-----+ +-----+ +-----+ +-----+ | | | | | | | | xCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxxxCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxx +--------------------------++Hinge++--------------------------++TM++C+ In mammals, there are two closely related types of lamp: lamp-1 and lamp-2, which form major components of the lysosome membrane. In chicken lamp-1 is known as LEP100. Also included in this entry is the macrophage protein CD68 (or macrosialin) [] is a heavily glycosylated integral membrane protein whose structure consists of a mucin-like domain followed by a proline-rich hinge; a single lamp-like domain; a transmembrane region and a short cytoplasmic tail. Similar to CD68, mammalian lamp-3, which is expressed in lymphoid organs, dendritic cells and in lung, contains all the C-terminal regions but lacks the N-terminal lamp-like region []. In a lamp-family protein from nematodes [] only the part C-terminal to the hinge is conserved. ; GO: 0016020 membrane
Probab=83.19 E-value=0.81 Score=39.48 Aligned_cols=30 Identities=20% Similarity=0.043 Sum_probs=13.3
Q ss_pred eEEehhHHHHHHH--HHHHHhhhhheecccch
Q 038524 138 LVTCVILVALVSV--LTTVGAAVYIVRKKKYD 167 (175)
Q Consensus 138 ~l~i~l~~~~~~~--~~~~~~~~~~~rr~~~~ 167 (175)
..++.+.+|++++ +++++++|++.|||+.+
T Consensus 270 ~~~vPIaVG~~La~lvlivLiaYli~Rrr~~~ 301 (306)
T PF01299_consen 270 SDLVPIAVGAALAGLVLIVLIAYLIGRRRSRA 301 (306)
T ss_pred cchHHHHHHHHHHHHHHHHHHhheeEeccccc
Confidence 3444444443333 33333445555555444
No 20
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=77.94 E-value=0.95 Score=32.97 Aligned_cols=14 Identities=29% Similarity=0.159 Sum_probs=6.3
Q ss_pred ceeEEehhHHHHHH
Q 038524 136 QKLVTCVILVALVS 149 (175)
Q Consensus 136 ~~~l~i~l~~~~~~ 149 (175)
+...++|+++++++
T Consensus 64 s~gaiagi~vg~~~ 77 (96)
T PTZ00382 64 STGAIAGISVAVVA 77 (96)
T ss_pred ccccEEEEEeehhh
Confidence 34444554454333
No 21
>PF14575 EphA2_TM: Ephrin type-A receptor 2 transmembrane domain; PDB: 3KUL_A 2XVD_A 2VX1_A 2VWV_A 2VX0_A 2VWY_A 2VWZ_A 2VWW_A 2VWU_A 2VWX_A ....
Probab=76.64 E-value=1.4 Score=30.71 Aligned_cols=11 Identities=18% Similarity=0.111 Sum_probs=4.0
Q ss_pred Hhhhhheeccc
Q 038524 155 GAAVYIVRKKK 165 (175)
Q Consensus 155 ~~~~~~~rr~~ 165 (175)
+++++++||.+
T Consensus 18 ~~~~~~~rr~~ 28 (75)
T PF14575_consen 18 IIVIVCFRRCK 28 (75)
T ss_dssp HHHHCCCTT--
T ss_pred eeEEEEEeeEc
Confidence 34444444443
No 22
>PF06697 DUF1191: Protein of unknown function (DUF1191); InterPro: IPR010605 This family contains hypothetical plant proteins of unknown function.
Probab=75.92 E-value=4.8 Score=34.89 Aligned_cols=15 Identities=7% Similarity=-0.002 Sum_probs=6.9
Q ss_pred EEehhHHHHHHHHHH
Q 038524 139 VTCVILVALVSVLTT 153 (175)
Q Consensus 139 l~i~l~~~~~~~~~~ 153 (175)
+++|+.++.++++++
T Consensus 215 iv~g~~~G~~~L~ll 229 (278)
T PF06697_consen 215 IVVGVVGGVVLLGLL 229 (278)
T ss_pred EEEEehHHHHHHHHH
Confidence 344544554444444
No 23
>PF14283 DUF4366: Domain of unknown function (DUF4366)
Probab=75.75 E-value=11 Score=31.55 Aligned_cols=29 Identities=14% Similarity=0.077 Sum_probs=24.7
Q ss_pred CCCcccccccCCCcEEEEEEeeCCCceEEEe
Q 038524 41 KGTNKSLKLISGDPMQTWIDYMGSEKLHEIQ 71 (175)
Q Consensus 41 ~~~~~~~~l~sG~~~~vwIdYd~~~~~L~v~ 71 (175)
+.+|..+.-++|+.++.-||+|.... +|+
T Consensus 72 ~kQFiTv~Tk~gn~FyliIDr~~~~e--nV~ 100 (218)
T PF14283_consen 72 GKQFITVTTKSGNTFYLIIDRDEEGE--NVY 100 (218)
T ss_pred CcEEEEEEecCCCEEEEEEecCCCcc--eEE
Confidence 44688888999999999999999977 564
No 24
>PF14991 MLANA: Protein melan-A; PDB: 2GTZ_F 2GT9_F 3MRO_P 2GUO_C 3MRQ_P 2GTW_C 3L6F_C 3MRP_P.
Probab=75.31 E-value=0.73 Score=34.84 Aligned_cols=15 Identities=20% Similarity=0.341 Sum_probs=0.0
Q ss_pred Hhhhhheecccchhh
Q 038524 155 GAAVYIVRKKKYDEV 169 (175)
Q Consensus 155 ~~~~~~~rr~~~~e~ 169 (175)
+.++++|||..|+-+
T Consensus 42 iGCWYckRRSGYk~L 56 (118)
T PF14991_consen 42 IGCWYCKRRSGYKTL 56 (118)
T ss_dssp ---------------
T ss_pred Hhheeeeecchhhhh
Confidence 566778877665443
No 25
>PHA03265 envelope glycoprotein D; Provisional
Probab=67.73 E-value=4.6 Score=36.25 Aligned_cols=28 Identities=18% Similarity=0.030 Sum_probs=15.8
Q ss_pred ceeEEehhHHHHHHHHHHHHhhhhheecc
Q 038524 136 QKLVTCVILVALVSVLTTVGAAVYIVRKK 164 (175)
Q Consensus 136 ~~~l~i~l~~~~~~~~~~~~~~~~~~rr~ 164 (175)
.++++||+.+++++++-+ ++++++||||
T Consensus 349 ~~g~~ig~~i~glv~vg~-il~~~~rr~k 376 (402)
T PHA03265 349 FVGISVGLGIAGLVLVGV-ILYVCLRRKK 376 (402)
T ss_pred ccceEEccchhhhhhhhH-HHHHHhhhhh
Confidence 356777877766554432 3445555554
No 26
>PF03302 VSP: Giardia variant-specific surface protein; InterPro: IPR005127 During infection, the intestinal protozoan parasite Giardia lamblia virus undergoes continuous antigenic variation which is determined by diversification of the parasite's major surface antigen, named VSP (variant surface protein).
Probab=65.17 E-value=4.7 Score=36.33 Aligned_cols=19 Identities=26% Similarity=0.165 Sum_probs=13.8
Q ss_pred ceeEEehhHHHHHHHHHHH
Q 038524 136 QKLVTCVILVALVSVLTTV 154 (175)
Q Consensus 136 ~~~l~i~l~~~~~~~~~~~ 154 (175)
++..++|++|+++++|-.|
T Consensus 365 stgaIaGIsvavvvvVggl 383 (397)
T PF03302_consen 365 STGAIAGISVAVVVVVGGL 383 (397)
T ss_pred cccceeeeeehhHHHHHHH
Confidence 6778888888876665544
No 27
>PF06024 DUF912: Nucleopolyhedrovirus protein of unknown function (DUF912); InterPro: IPR009261 This entry is represented by Autographa californica nuclear polyhedrosis virus (AcMNPV), Orf78; it is a family of uncharacterised viral proteins.
Probab=61.78 E-value=5.6 Score=29.05 Aligned_cols=10 Identities=20% Similarity=0.152 Sum_probs=4.1
Q ss_pred Hhhhhheecc
Q 038524 155 GAAVYIVRKK 164 (175)
Q Consensus 155 ~~~~~~~rr~ 164 (175)
+.+|++.|.+
T Consensus 81 IyYFVILRer 90 (101)
T PF06024_consen 81 IYYFVILRER 90 (101)
T ss_pred heEEEEEecc
Confidence 3344444433
No 28
>PF12877 DUF3827: Domain of unknown function (DUF3827); InterPro: IPR024606 The function of the proteins in this entry is not currently known, but one of the human proteins (Q9HCM3 from SWISSPROT) has been implicated in pilocytic astrocytomas [, , ]. In the majority of cases of pilocytic astrocytomas a tandem duplication produces an in-frame fusion of the gene encoding this protein and the BRAF oncogene. The resulting fusion protein has constitutive BRAF kinase activity and is capable of transforming cells.
Probab=59.64 E-value=14 Score=35.56 Aligned_cols=18 Identities=22% Similarity=0.351 Sum_probs=8.1
Q ss_pred eeEEehhHHHHHHHHHHH
Q 038524 137 KLVTCVILVALVSVLTTV 154 (175)
Q Consensus 137 ~~l~i~l~~~~~~~~~~~ 154 (175)
..|++|+.+.++++++++
T Consensus 269 lWII~gVlvPv~vV~~Ii 286 (684)
T PF12877_consen 269 LWIIAGVLVPVLVVLLII 286 (684)
T ss_pred eEEEehHhHHHHHHHHHH
Confidence 445556544444333333
No 29
>PF02009 Rifin_STEVOR: Rifin/stevor family; InterPro: IPR002858 Malaria is still a major cause of mortality in many areas of the world. Plasmodium falciparum causes the most severe human form of the disease and is responsible for most fatalities. Severe cases of malaria can occur when the parasite invades and then proliferates within red blood cell erythrocytes. The parasite produces many variant antigenic proteins, encoded by multigene families, which are present on the surface of the infected erythrocyte and play important roles in virulence. A crucial survival mechanism for the malaria parasite is its ability to evade the immune response by switching these variant surface antigens. The high virulence of P. falciparum relative to other malarial parasites is in large part due to the fact that in this organism many of these surface antigens mediate the binding of infected erythrocytes to the vascular endothelium (cytoadherence) and non-infected erythrocytes (rosetting). This can lead to the accumulation of infected cells in the vasculature of a variety of organs, blocking the blood flow and reducing the oxygen supply. Clinical symptoms of severe infection can include fever, progressive anaemia, multi-organ dysfunction and coma. For more information see []. Several multicopy gene families have been described in Plasmodium falciparum, including the stevor family of subtelomeric open reading frames and the rif interspersed repetitive elements. Both families contain three predicted transmembrane segments. It has been proposed that stevor and rif are members of a larger superfamily that code for variant surface antigens [].
Probab=56.25 E-value=5 Score=35.04 Aligned_cols=7 Identities=0% Similarity=-0.263 Sum_probs=2.9
Q ss_pred Hhhhhhe
Q 038524 155 GAAVYIV 161 (175)
Q Consensus 155 ~~~~~~~ 161 (175)
++++++|
T Consensus 274 IIYLILR 280 (299)
T PF02009_consen 274 IIYLILR 280 (299)
T ss_pred HHHHHHH
Confidence 3444443
No 30
>PF11857 DUF3377: Domain of unknown function (DUF3377); InterPro: IPR021805 This domain is functionally uncharacterised and found at the C terminus of peptidases belonging to MEROPS peptidase family M10A, membrane-type matrix metallopeptidases (clan MA). ; GO: 0004222 metalloendopeptidase activity
Probab=54.22 E-value=20 Score=25.09 Aligned_cols=18 Identities=22% Similarity=0.327 Sum_probs=10.6
Q ss_pred eeEEehhHHHHHHHHHHH
Q 038524 137 KLVTCVILVALVSVLTTV 154 (175)
Q Consensus 137 ~~l~i~l~~~~~~~~~~~ 154 (175)
+.+.+.+++.+++.++++
T Consensus 30 ~avaVviPl~L~LCiLvl 47 (74)
T PF11857_consen 30 NAVAVVIPLVLLLCILVL 47 (74)
T ss_pred eEEEEeHHHHHHHHHHHH
Confidence 455666666655555555
No 31
>PF12768 Rax2: Cortical protein marker for cell polarity
Probab=49.81 E-value=14 Score=31.89 Aligned_cols=9 Identities=22% Similarity=0.331 Sum_probs=4.2
Q ss_pred eEEehhHHH
Q 038524 138 LVTCVILVA 146 (175)
Q Consensus 138 ~l~i~l~~~ 146 (175)
++.|++.+|
T Consensus 229 VVlIslAiA 237 (281)
T PF12768_consen 229 VVLISLAIA 237 (281)
T ss_pred EEEEehHHH
Confidence 344444444
No 32
>PF15345 TMEM51: Transmembrane protein 51
Probab=48.32 E-value=16 Score=30.89 Aligned_cols=10 Identities=20% Similarity=0.298 Sum_probs=4.3
Q ss_pred hhhheecccc
Q 038524 157 AVYIVRKKKY 166 (175)
Q Consensus 157 ~~~~~rr~~~ 166 (175)
|+-+|.|||.
T Consensus 77 CL~IR~KRr~ 86 (233)
T PF15345_consen 77 CLSIRDKRRR 86 (233)
T ss_pred HHHHHHHHHH
Confidence 3445544443
No 33
>PTZ00208 65 kDa invariant surface glycoprotein; Provisional
Probab=45.49 E-value=9.3 Score=34.87 Aligned_cols=30 Identities=13% Similarity=0.235 Sum_probs=13.1
Q ss_pred EEehhHHHHHHHHHHHH-hhhhheecccchh
Q 038524 139 VTCVILVALVSVLTTVG-AAVYIVRKKKYDE 168 (175)
Q Consensus 139 l~i~l~~~~~~~~~~~~-~~~~~~rr~~~~e 168 (175)
+++++.+..++|++..+ +++++||||..+|
T Consensus 388 i~~avl~p~~il~~~~~~~~~~v~rrr~~~~ 418 (436)
T PTZ00208 388 IILAVLVPAIILAIIAVAFFIMVKRRRNSSE 418 (436)
T ss_pred HHHHHHHHHHHHHHHHHHhheeeeeccCCch
Confidence 33444443333332223 4444555555444
No 34
>PF05337 CSF-1: Macrophage colony stimulating factor-1 (CSF-1); InterPro: IPR008001 Colony stimulating factor 1 (CSF-1) is a homodimeric polypeptide growth factor whose primary function is to regulate the survival, proliferation, differentiation, and function of cells of the mononuclear phagocytic lineage. This lineage includes mononuclear phagocytic precursors, blood monocytes, tissue macrophages, osteoclasts, and microglia of the brain, all of which possess cell surface receptors for CSF-1. The protein has also been linked with male fertility [] and mutations in the Csf-1 gene have been found to cause osteopetrosis and failure of tooth eruption [].; GO: 0005125 cytokine activity, 0008083 growth factor activity, 0016021 integral to membrane; PDB: 3EJJ_A.
Probab=44.15 E-value=7.5 Score=33.71 Aligned_cols=29 Identities=14% Similarity=0.311 Sum_probs=0.0
Q ss_pred eeEEehhHHHHHHHHHHH-Hhhhhheeccc
Q 038524 137 KLVTCVILVALVSVLTTV-GAAVYIVRKKK 165 (175)
Q Consensus 137 ~~l~i~l~~~~~~~~~~~-~~~~~~~rr~~ 165 (175)
..++.-|.+.++++|++. +.+++||||+|
T Consensus 224 p~~vf~lLVPSiILVLLaVGGLLfYr~rrR 253 (285)
T PF05337_consen 224 PGFVFYLLVPSIILVLLAVGGLLFYRRRRR 253 (285)
T ss_dssp ------------------------------
T ss_pred Ccccccccccchhhhhhhccceeeeccccc
Confidence 345566667666666555 33444444433
No 35
>PF12191 stn_TNFRSF12A: Tumour necrosis factor receptor stn_TNFRSF12A_TNFR domain; InterPro: IPR022316 The tumour necrosis factor (TNF) receptor (TNFR) superfamily comprises more than 20 type-I transmembrane proteins. Family members are defined based on similarity in their extracellular domain - a region that contains many cysteine residues arranged in a specific repetitive pattern []. The cysteines allow formation of an extended rod-like structure, responsible for ligand binding []. Upon receptor activation, different intracellular signalling complexes are assembled for different members of the TNFR superfamily, depending on their intracellular domains and sequences []. Activation of TNFRs can therefore induce a range of disparate effects, including cell proliferation, differentiation, survival, or apoptotic cell death, depending upon the receptor involved []. TNFRs are widely distributed and play important roles in many crucial biological processes, such as lymphoid and neuronal development, innate and adaptive immunity, and maintenance of cellular homeostasis []. Drugs that manipulate their signalling have potential roles in the prevention and treatment of many diseases, such as viral infections, coronary heart disease, transplant rejection, and immune disease []. TNF receptor 12 (also known as TWEAK receptor, and fibroblast growth factor-inducible-14 (Fn14)) has been implicated in endothelial cell growth and migration []. The receptor may also play a role in cell-matrix interactions [].; PDB: 2KN0_A 2RPJ_A 2KMZ_A 2EQP_A.
Probab=43.16 E-value=7.9 Score=29.81 Aligned_cols=24 Identities=8% Similarity=-0.057 Sum_probs=0.0
Q ss_pred EEehhHHHHHHHHHHH-Hhhhhhee
Q 038524 139 VTCVILVALVSVLTTV-GAAVYIVR 162 (175)
Q Consensus 139 l~i~l~~~~~~~~~~~-~~~~~~~r 162 (175)
..|+.++.++++++.+ ..++++||
T Consensus 79 ~pi~~sal~v~lVl~llsg~lv~rr 103 (129)
T PF12191_consen 79 WPILGSALSVVLVLALLSGFLVWRR 103 (129)
T ss_dssp -------------------------
T ss_pred hhhhhhHHHHHHHHHHHHHHHHHhh
Confidence 3343344444444343 33344443
No 36
>PLN03150 hypothetical protein; Provisional
Probab=42.61 E-value=25 Score=33.31 Aligned_cols=6 Identities=17% Similarity=0.446 Sum_probs=2.3
Q ss_pred hhheec
Q 038524 158 VYIVRK 163 (175)
Q Consensus 158 ~~~~rr 163 (175)
+++|||
T Consensus 566 ~~~~~r 571 (623)
T PLN03150 566 CWWKRR 571 (623)
T ss_pred hheeeh
Confidence 334433
No 37
>PF03229 Alpha_GJ: Alphavirus glycoprotein J; InterPro: IPR004913 The exact function of the herpesvirus glycoprotein J is unknown, but it appears to play a role in the inhibition of apotosis of the host cell [].; GO: 0019050 suppression by virus of host apoptosis
Probab=42.18 E-value=40 Score=25.72 Aligned_cols=28 Identities=18% Similarity=0.148 Sum_probs=14.2
Q ss_pred eeEEehhHHHHHHHHHHH--Hhhhhheecc
Q 038524 137 KLVTCVILVALVSVLTTV--GAAVYIVRKK 164 (175)
Q Consensus 137 ~~l~i~l~~~~~~~~~~~--~~~~~~~rr~ 164 (175)
..+++++.++++.++.+. +...++||+.
T Consensus 82 ~d~aLp~VIGGLcaL~LaamGA~~LLrR~c 111 (126)
T PF03229_consen 82 VDFALPLVIGGLCALTLAAMGAGALLRRCC 111 (126)
T ss_pred cccchhhhhhHHHHHHHHHHHHHHHHHHHH
Confidence 346666666655444333 4445554433
No 38
>PF06365 CD34_antigen: CD34/Podocalyxin family; InterPro: IPR013836 This family consists of several mammalian CD34 antigen proteins. The CD34 antigen is a human leukocyte membrane protein expressed specifically by lymphohematopoietic progenitor cells. CD34 is a phosphoprotein. Activation of protein kinase C (PKC) has been found to enhance CD34 phosphorylation [, ]. This family contains several eukaryotic podocalyxin proteins. Podocalyxin is a major membrane protein of the glomerular epithelium and is thought to be involved in maintenance of the architecture of the foot processes and filtration slits characteristic of this unique epithelium by virtue of its high negative charge. Podocalyxin functions as an anti-adhesin that maintains an open filtration pathway between neighbouring foot processes in the glomerular epithelium by charge repulsion [].
Probab=41.94 E-value=46 Score=27.55 Aligned_cols=26 Identities=8% Similarity=0.118 Sum_probs=11.2
Q ss_pred eEEehhHHHHHHHHHHH--Hhhhhheec
Q 038524 138 LVTCVILVALVSVLTTV--GAAVYIVRK 163 (175)
Q Consensus 138 ~l~i~l~~~~~~~~~~~--~~~~~~~rr 163 (175)
.|+..+.++++++++++ ++++++.||
T Consensus 101 ~lI~lv~~g~~lLla~~~~~~Y~~~~Rr 128 (202)
T PF06365_consen 101 TLIALVTSGSFLLLAILLGAGYCCHQRR 128 (202)
T ss_pred EEEehHHhhHHHHHHHHHHHHHHhhhhc
Confidence 34444444444444333 334444444
No 39
>TIGR01167 LPXTG_anchor LPXTG-motif cell wall anchor domain. A common feature of this proteins containing this domain appears to be a high proportion of charged and zwitterionic residues immediatedly upstream of the LPXTG motif. This model differs from other descriptions of the LPXTG region by including a portion of that upstream charged region.
Probab=41.10 E-value=31 Score=19.47 Aligned_cols=6 Identities=17% Similarity=0.523 Sum_probs=2.4
Q ss_pred hheecc
Q 038524 159 YIVRKK 164 (175)
Q Consensus 159 ~~~rr~ 164 (175)
+++||+
T Consensus 27 ~~~~rk 32 (34)
T TIGR01167 27 LLRKRK 32 (34)
T ss_pred Hheecc
Confidence 333443
No 40
>PF02480 Herpes_gE: Alphaherpesvirus glycoprotein E; InterPro: IPR003404 Glycoprotein E (gE) of Alphaherpesvirus forms a complex with glycoprotein I (gI), functioning as an immunoglobulin G (IgG) Fc binding protein. gE is involved in virus spread but is not essential for propagation [].; GO: 0016020 membrane; PDB: 2GJ7_F 2GIY_B.
Probab=40.75 E-value=9.1 Score=35.09 Aligned_cols=9 Identities=0% Similarity=-0.090 Sum_probs=5.4
Q ss_pred EEeeeEEee
Q 038524 105 YILGCSLNK 113 (175)
Q Consensus 105 ~IlsWsF~~ 113 (175)
.+.+|....
T Consensus 293 hv~aW~yt~ 301 (439)
T PF02480_consen 293 HVEAWTYTL 301 (439)
T ss_dssp EEEEEEEEE
T ss_pred eeeeeEEEE
Confidence 456776653
No 41
>PF15050 SCIMP: SCIMP protein
Probab=40.70 E-value=9.9 Score=29.17 Aligned_cols=8 Identities=0% Similarity=-0.566 Sum_probs=3.7
Q ss_pred Hhhhhhee
Q 038524 155 GAAVYIVR 162 (175)
Q Consensus 155 ~~~~~~~r 162 (175)
++++++|+
T Consensus 26 IlyCvcR~ 33 (133)
T PF15050_consen 26 ILYCVCRW 33 (133)
T ss_pred HHHHHHHH
Confidence 34444553
No 42
>TIGR01478 STEVOR variant surface antigen, stevor family. This model represents the stevor branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of stevor sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 8 bits.
Probab=38.96 E-value=20 Score=31.28 Aligned_cols=7 Identities=29% Similarity=0.254 Sum_probs=2.9
Q ss_pred hhheecc
Q 038524 158 VYIVRKK 164 (175)
Q Consensus 158 ~~~~rr~ 164 (175)
+++|||+
T Consensus 281 WlyrrRK 287 (295)
T TIGR01478 281 WLYRRRK 287 (295)
T ss_pred HHHHhhc
Confidence 3344444
No 43
>PTZ00370 STEVOR; Provisional
Probab=38.77 E-value=20 Score=31.28 Aligned_cols=7 Identities=29% Similarity=0.254 Sum_probs=3.0
Q ss_pred hhheecc
Q 038524 158 VYIVRKK 164 (175)
Q Consensus 158 ~~~~rr~ 164 (175)
+++|||+
T Consensus 277 wlyrrRK 283 (296)
T PTZ00370 277 WLYRRRK 283 (296)
T ss_pred HHHHhhc
Confidence 3344444
No 44
>TIGR03370 PEPCTERM_Roseo variant PEP-CTERM putative exosortase signal, Roseobacter type. A probable protein export sorting signal, PEP-CTERM, was described by Haft, et al. (PubMed:16930487). It is predicted to interact with a putative transpeptidase we designate exosortase. Most examples of this signal are recognized by model TIGR02595, but some unusual clades require different models. This model describes a variant with conserved motif VPLPA, rather than VPEP. This variant is found prominently in two members of the Rhodobacterales, namely Jannaschia sp. CCS1 and Roseobacter denitrificans OCh 114. One interesting member protein has a full-length duplication and therefore two copies of this putative sorting domain.
Probab=37.67 E-value=20 Score=20.14 Aligned_cols=11 Identities=18% Similarity=0.416 Sum_probs=5.0
Q ss_pred Hhhhhheeccc
Q 038524 155 GAAVYIVRKKK 165 (175)
Q Consensus 155 ~~~~~~~rr~~ 165 (175)
+.+...|||+|
T Consensus 15 ggl~~~rRRrk 25 (26)
T TIGR03370 15 GGLGAMRRRRR 25 (26)
T ss_pred HHHHHHHHhhc
Confidence 33444555543
No 45
>PF13268 DUF4059: Protein of unknown function (DUF4059)
Probab=35.72 E-value=22 Score=24.71 Aligned_cols=24 Identities=21% Similarity=0.228 Sum_probs=10.2
Q ss_pred hHHHHHHHHHHHHhhhhheecccc
Q 038524 143 ILVALVSVLTTVGAAVYIVRKKKY 166 (175)
Q Consensus 143 l~~~~~~~~~~~~~~~~~~rr~~~ 166 (175)
+.+++++++++.+.+.++|.++|+
T Consensus 13 L~ls~i~V~~~~~~wi~~Ra~~~~ 36 (72)
T PF13268_consen 13 LLLSSILVLLVSGIWILWRALRKK 36 (72)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHcC
Confidence 344443333333444445544444
No 46
>PF08374 Protocadherin: Protocadherin; InterPro: IPR013585 The structure of protocadherins is similar to that of classic cadherins (IPR002126 from INTERPRO), but they also have some unique features associated with the cytoplasmic domains. They are expressed in a variety of organisms and are found in high concentrations in the brain where they seem to be localised mainly at cell-cell contact sites. Their expression seems to be developmentally regulated [].
Probab=34.79 E-value=65 Score=27.04 Aligned_cols=17 Identities=12% Similarity=0.387 Sum_probs=7.2
Q ss_pred eEEehhHHHHHHHHHHH
Q 038524 138 LVTCVILVALVSVLTTV 154 (175)
Q Consensus 138 ~l~i~l~~~~~~~~~~~ 154 (175)
.+++|+..++++++|++
T Consensus 38 ~I~iaiVAG~~tVILVI 54 (221)
T PF08374_consen 38 KIMIAIVAGIMTVILVI 54 (221)
T ss_pred eeeeeeecchhhhHHHH
Confidence 44444444444444333
No 47
>PF12248 Methyltransf_FA: Farnesoic acid 0-methyl transferase; InterPro: IPR022041 This domain, found in farnesoic acid O-methyl transferase, is approximately 110 amino acids in length. Farnesoic acid O-methyl transferase (FAMeT) is the enzyme that catalyses the formation of methyl farnesoate (MF) from farnesoic acid (FA) in the biosynthetic pathway of juvenile hormone (JH) [].
Probab=33.06 E-value=1.8e+02 Score=20.71 Aligned_cols=41 Identities=20% Similarity=0.239 Sum_probs=28.9
Q ss_pred ccCCCcEEEEEEeeCCCceEEEee-----eeeeeeecCCccCCCceEEEEE
Q 038524 49 LISGDPMQTWIDYMGSEKLHEIQC-----LSLSTSVDLSQLLLDTMCVGFS 94 (175)
Q Consensus 49 l~sG~~~~vwIdYd~~~~~L~v~l-----Plls~~idLs~~l~~~~yVGFS 94 (175)
|...+...-||.+++ ..+.|+. |+|+.. |-. -..--|||||
T Consensus 49 ls~~e~~~fwI~~~~--G~I~vg~~g~~~pfl~~~-Dp~--~~~v~yvGft 94 (102)
T PF12248_consen 49 LSPSEFRMFWISWRD--GTIRVGRGGEDEPFLEWT-DPE--PIPVNYVGFT 94 (102)
T ss_pred CCCCccEEEEEEECC--CEEEEEECCCccEEEEEE-CCC--CCcccEEEEe
Confidence 467888999999765 4666665 888876 322 3356799994
No 48
>PF13908 Shisa: Wnt and FGF inhibitory regulator
Probab=32.30 E-value=41 Score=26.54 Aligned_cols=11 Identities=0% Similarity=0.187 Sum_probs=4.7
Q ss_pred eEEehhHHHHH
Q 038524 138 LVTCVILVALV 148 (175)
Q Consensus 138 ~l~i~l~~~~~ 148 (175)
.+++++.++++
T Consensus 79 ~iivgvi~~Vi 89 (179)
T PF13908_consen 79 GIIVGVICGVI 89 (179)
T ss_pred eeeeehhhHHH
Confidence 34444444333
No 49
>PF03597 CcoS: Cytochrome oxidase maturation protein cbb3-type; InterPro: IPR004714 Cytochrome cbb3 oxidases are found almost exclusively in Proteobacteria, and represent a distinctive class of proton-pumping respiratory haem-copper oxidases (HCO) that lack many of the key structural features that contribute to the reaction cycle of the intensely studied mitochondrial cytochrome c oxidase (CcO). Expression of cytochrome cbb3 oxidase allows human pathogens to colonise anoxic tissues and agronomically important diazotrophs to sustain nitrogen fixation []. Genes encoding a cytochrome cbb3 oxidase were initially designated fixNOQP (ccoNOQP), the ccoNOQP operon is always found close to a second gene cluster, known as fixGHIS (ccoGHIS) whose expression is necessary for the assembly of a functional cbb3 oxidase. On the basis of their derived amino acid sequences each of the four proteins encoded by the ccoGHIS operon are thought to be membrane-bound. It has been suggested that they may function in concert as a multi-subunit complex, possibly playing a role in the uptake and metabolism of copper required for the assembly of the binuclear centre of cytochrome cbb3 oxidase.
Probab=31.72 E-value=36 Score=21.43 Aligned_cols=27 Identities=26% Similarity=0.498 Sum_probs=11.6
Q ss_pred hhHHHHHHHHHHHHhhhhheecccchh
Q 038524 142 VILVALVSVLTTVGAAVYIVRKKKYDE 168 (175)
Q Consensus 142 ~l~~~~~~~~~~~~~~~~~~rr~~~~e 168 (175)
-++++.++.+++++++++.-|+.++.+
T Consensus 6 lip~sl~l~~~~l~~f~Wavk~GQfdD 32 (45)
T PF03597_consen 6 LIPVSLILGLIALAAFLWAVKSGQFDD 32 (45)
T ss_pred HHHHHHHHHHHHHHHHHHHHccCCCCC
Confidence 344444333333334444445555533
No 50
>PRK05886 yajC preprotein translocase subunit YajC; Validated
Probab=31.62 E-value=35 Score=25.51 Aligned_cols=11 Identities=18% Similarity=-0.055 Sum_probs=5.2
Q ss_pred Hhhhhheeccc
Q 038524 155 GAAVYIVRKKK 165 (175)
Q Consensus 155 ~~~~~~~rr~~ 165 (175)
++|+++|+.+|
T Consensus 17 ~yF~~iRPQkK 27 (109)
T PRK05886 17 FMYFASRRQRK 27 (109)
T ss_pred HHHHHccHHHH
Confidence 34455554443
No 51
>PF11353 DUF3153: Protein of unknown function (DUF3153); InterPro: IPR021499 This family of proteins with unknown function appear to be restricted to Cyanobacteria. Some members are annotated as membrane proteins however this cannot be confirmed.
Probab=31.45 E-value=35 Score=27.76 Aligned_cols=25 Identities=28% Similarity=0.252 Sum_probs=14.2
Q ss_pred eeeeeeecCCccCCC-ceEEEEEeec
Q 038524 73 LSLSTSVDLSQLLLD-TMCVGFSAAT 97 (175)
Q Consensus 73 Plls~~idLs~~l~~-~~yVGFSAsT 97 (175)
.-|...+||+..-.. ..=+-|+=++
T Consensus 117 ~~L~~~lDL~~L~~~~~ldl~f~l~~ 142 (209)
T PF11353_consen 117 YRLDLDLDLRSLPDLPGLDLEFSLST 142 (209)
T ss_pred EEEEEEeehhhcCCCCcceEEEEEeC
Confidence 446777888765432 2345555544
No 52
>PF14914 LRRC37AB_C: LRRC37A/B like protein 1 C-terminal domain
Probab=31.22 E-value=29 Score=27.53 Aligned_cols=18 Identities=17% Similarity=0.265 Sum_probs=9.6
Q ss_pred eeEEehhHHHHHHHHHHH
Q 038524 137 KLVTCVILVALVSVLTTV 154 (175)
Q Consensus 137 ~~l~i~l~~~~~~~~~~~ 154 (175)
..+++++++.+++.++++
T Consensus 119 nklilaisvtvv~~ilii 136 (154)
T PF14914_consen 119 NKLILAISVTVVVMILII 136 (154)
T ss_pred chhHHHHHHHHHHHHHHH
Confidence 356666666554444333
No 53
>PTZ00046 rifin; Provisional
Probab=30.50 E-value=33 Score=30.82 Aligned_cols=6 Identities=0% Similarity=-0.197 Sum_probs=2.2
Q ss_pred Hhhhhh
Q 038524 155 GAAVYI 160 (175)
Q Consensus 155 ~~~~~~ 160 (175)
++++++
T Consensus 333 IIYLIL 338 (358)
T PTZ00046 333 IIYLIL 338 (358)
T ss_pred HHHHHH
Confidence 333333
No 54
>TIGR02595 PEP_exosort PEP-CTERM putative exosortase interaction domain. This model describes a 25-residue domain that includes a near-invariant Pro-Glu-Pro (PEP) motif, a thirteen residue strongly hydrophobic sequence likely to span the membrane, and a five-residue strongly basic motif that often contains four Arg residues. In nearly every case, this motif is found within nine residues, and usually within five residues, of the extreme C-terminus of the protein. Proteins with this motif typically have signal sequences at the N-terminus. This region appears many times per genome or not at all, and co-occurs in genomes with a proposed protein-sorting integral membrane protein we designate exosortase (see TIGR02602). PEP-CTERM proteins frequently are poorly conserved, Ser/Thr-rich proteins and may become extensively modified proteinaceous constituents of extracellular material in bacterial biofilms.
Probab=30.04 E-value=45 Score=18.35 Aligned_cols=7 Identities=14% Similarity=0.605 Sum_probs=3.1
Q ss_pred hhheecc
Q 038524 158 VYIVRKK 164 (175)
Q Consensus 158 ~~~~rr~ 164 (175)
+..|||+
T Consensus 17 ~~~rrrk 23 (26)
T TIGR02595 17 LLLRRRR 23 (26)
T ss_pred HHHhhcc
Confidence 3444444
No 55
>TIGR03141 cytochro_ccmD heme exporter protein CcmD. The model for this protein family describes a small, hydrophobic, and only moderately well-conserved protein, tricky to identify accurately for all of these reasons. However, members are found as part of large operons involved in heme export across the inner membrane for assembly of c-type cytochromes in a large number of bacteria. The gray zone between the trusted cutoff (13.0) and noise cutoff (4.75) includes both low-scoring examples and false-positive matches to hydrophobic domains of longer proteins.
Probab=29.93 E-value=36 Score=21.12 Aligned_cols=11 Identities=0% Similarity=0.141 Sum_probs=4.4
Q ss_pred hhHHHHHHHHH
Q 038524 142 VILVALVSVLT 152 (175)
Q Consensus 142 ~l~~~~~~~~~ 152 (175)
..+-+..++++
T Consensus 9 W~sYg~t~l~l 19 (45)
T TIGR03141 9 WLAYGITALVL 19 (45)
T ss_pred HHHHHHHHHHH
Confidence 33444433333
No 56
>TIGR03521 GldG gliding-associated putative ABC transporter substrate-binding component GldG. Members of this protein family are exclusive to the Bacteroidetes phylum (previously Cytophaga-Flavobacteria-Bacteroides). GldG is a protein linked to a type of rapid surface gliding motility found in certain Bacteroidetes, such as Flavobacterium johnsoniae and Cytophaga hutchinsonii. Knockouts of GldG abolish the gliding phenotype. GldG, along with GldA and GldF are believed to compose an ABC transporter and are observed as an operon. Gliding motility appears closely linked to chitin utilization in the model species Flavobacterium johnsoniae. Bacteroidetes with members of this protein family appear to have all of the genes associated with gliding motility.
Probab=29.20 E-value=36 Score=31.86 Aligned_cols=12 Identities=42% Similarity=0.797 Sum_probs=6.5
Q ss_pred Hhhhhheecccc
Q 038524 155 GAAVYIVRKKKY 166 (175)
Q Consensus 155 ~~~~~~~rr~~~ 166 (175)
++++++|||+||
T Consensus 540 G~~~~~~Rrr~~ 551 (552)
T TIGR03521 540 GLSFTYIRKRKY 551 (552)
T ss_pred HHHHHHHHHhhc
Confidence 344455666665
No 57
>TIGR01582 FDH-beta formate dehydrogenase, beta subunit, Fe-S containing. In addition to the gamma proteobacteria, a sequence from Aquifex aolicus falls within the scope of this model. This appears to be the case for the alpha, gamma and epsilon (accessory protein TIGR01562) chains as well.
Probab=28.98 E-value=71 Score=27.57 Aligned_cols=11 Identities=18% Similarity=0.111 Sum_probs=6.0
Q ss_pred CCCCCCCCCCC
Q 038524 120 LNISTLPSFHL 130 (175)
Q Consensus 120 l~~s~lp~~p~ 130 (175)
..+..||..|.
T Consensus 228 ~~~~~lp~~p~ 238 (283)
T TIGR01582 228 KDYQDLPEDPR 238 (283)
T ss_pred HHhcCCCCCCc
Confidence 34445676654
No 58
>COG4736 CcoQ Cbb3-type cytochrome oxidase, subunit 3 [Posttranslational modification, protein turnover, chaperones]
Probab=28.90 E-value=56 Score=21.91 Aligned_cols=10 Identities=20% Similarity=-0.110 Sum_probs=4.3
Q ss_pred hhhhheeccc
Q 038524 156 AAVYIVRKKK 165 (175)
Q Consensus 156 ~~~~~~rr~~ 165 (175)
+++.+|+++|
T Consensus 26 i~~ayr~~~K 35 (60)
T COG4736 26 IYFAYRPGKK 35 (60)
T ss_pred HHHHhcccch
Confidence 3444444443
No 59
>PF04689 S1FA: DNA binding protein S1FA; InterPro: IPR006779 S1FA is an unusual small plant peptide of only 70 amino acids with a basic domain which contains a nuclear localization signal and a putative DNA binding helix. S1FA is highly conserved between dicotyledonous and monocotyledonous plants and may be a DNA-binding protein that specifically recognises the negative promoter element S1F [].; GO: 0003677 DNA binding, 0006355 regulation of transcription, DNA-dependent, 0005634 nucleus
Probab=28.40 E-value=67 Score=21.99 Aligned_cols=19 Identities=16% Similarity=0.141 Sum_probs=11.9
Q ss_pred ceeEEehhHHHHHHHHHHH
Q 038524 136 QKLVTCVILVALVSVLTTV 154 (175)
Q Consensus 136 ~~~l~i~l~~~~~~~~~~~ 154 (175)
+-++++.+.++.+++++++
T Consensus 11 nPGlIVLlvV~g~ll~flv 29 (69)
T PF04689_consen 11 NPGLIVLLVVAGLLLVFLV 29 (69)
T ss_pred CCCeEEeehHHHHHHHHHH
Confidence 3457777777766665555
No 60
>PF10661 EssA: WXG100 protein secretion system (Wss), protein EssA; InterPro: IPR018920 The Wss (WXG100 protein secretion system) in Staphylococcus aureus seems to be encoded by a locus of eight ORFs, called ess (eSAT-6 secretion system) []. This locus encodes, amongst several other proteins, EssA, a protein predicted to possess one transmembrane domain. Due to its predicted membrane location and its absolute requirement for WXG100 protein secretion, it has been speculated that EssA could form a secretion apparatus in conjunction with YukC and YukAB. Proteins homologous to EssA, YukC, EsaA and YukD were absent from mycobacteria []. Members of this family are associated with type VII secretion of WXG100 family targets in the Firmicutes, but not in the Actinobacteria. This highly divergent protein family consists largely of a central region of highly polar low-complexity sequence containing occasional LF motifs in weak repeats about 17 residues in length, flanked by hydrophobic N- and C-terminal regions.
Probab=26.24 E-value=43 Score=26.14 Aligned_cols=21 Identities=10% Similarity=0.074 Sum_probs=9.7
Q ss_pred ehhHHHHHHHHHHHHhhhhhe
Q 038524 141 CVILVALVSVLTTVGAAVYIV 161 (175)
Q Consensus 141 i~l~~~~~~~~~~~~~~~~~~ 161 (175)
+++++++++++++.+++..+|
T Consensus 121 i~~~i~g~ll~i~~giy~~~r 141 (145)
T PF10661_consen 121 ILLSIGGILLAICGGIYVVLR 141 (145)
T ss_pred HHHHHHHHHHHHHHHHHHHHH
Confidence 444455554444444444444
No 61
>PHA03264 envelope glycoprotein D; Provisional
Probab=25.78 E-value=73 Score=28.99 Aligned_cols=18 Identities=11% Similarity=0.032 Sum_probs=10.3
Q ss_pred ceeEEehhHHHHHHHHHH
Q 038524 136 QKLVTCVILVALVSVLTT 153 (175)
Q Consensus 136 ~~~l~i~l~~~~~~~~~~ 153 (175)
...+.+|+.++..+++++
T Consensus 359 ~~~~~vg~~~a~~~i~~~ 376 (416)
T PHA03264 359 ARPVIVGTGIAAAAIACV 376 (416)
T ss_pred cceeeeehhhhHHHHHHH
Confidence 345667777766444443
No 62
>PRK07021 fliL flagellar basal body-associated protein FliL; Reviewed
Probab=25.48 E-value=69 Score=25.01 Aligned_cols=8 Identities=13% Similarity=0.480 Sum_probs=3.4
Q ss_pred Hhhhhhee
Q 038524 155 GAAVYIVR 162 (175)
Q Consensus 155 ~~~~~~~r 162 (175)
+++|++.+
T Consensus 35 g~~~~~~~ 42 (162)
T PRK07021 35 GYSWWLSK 42 (162)
T ss_pred HHHHHhhc
Confidence 34444443
No 63
>PHA03291 envelope glycoprotein I; Provisional
Probab=24.17 E-value=58 Score=29.43 Aligned_cols=25 Identities=12% Similarity=0.151 Sum_probs=14.4
Q ss_pred eEEehhHHHHHHHHHHH-Hhhhhhee
Q 038524 138 LVTCVILVALVSVLTTV-GAAVYIVR 162 (175)
Q Consensus 138 ~l~i~l~~~~~~~~~~~-~~~~~~~r 162 (175)
.+.|+++.+.++++++. .++++.|+
T Consensus 288 iiQiAIPasii~cV~lGSC~Ccl~R~ 313 (401)
T PHA03291 288 IIQIAIPASIIACVFLGSCACCLHRR 313 (401)
T ss_pred hheeccchHHHHHhhhhhhhhhhhhh
Confidence 45667777666666555 44555443
No 64
>PRK15471 chain length determinant protein WzzB; Provisional
Probab=23.46 E-value=60 Score=28.53 Aligned_cols=7 Identities=43% Similarity=0.524 Sum_probs=3.1
Q ss_pred ceeEEeh
Q 038524 136 QKLVTCV 142 (175)
Q Consensus 136 ~~~l~i~ 142 (175)
++.+++.
T Consensus 293 kr~lIli 299 (325)
T PRK15471 293 KKAITLV 299 (325)
T ss_pred cchhHHH
Confidence 4444443
No 65
>PF06809 NPDC1: Neural proliferation differentiation control-1 protein (NPDC1); InterPro: IPR009635 This family consists of several neural proliferation differentiation control-1 (NPDC1) proteins. NPDC1 plays a role in the control of neural cell proliferation and differentiation. It has been suggested that NPDC1 may be involved in the development of several secretion glands. This family also contains the C-terminal region of the Caenorhabditis elegans protein CAB-1 (Q93249 from SWISSPROT) which is known to interact with AEX-3 [].; GO: 0016021 integral to membrane
Probab=23.12 E-value=1.4e+02 Score=26.71 Aligned_cols=7 Identities=43% Similarity=0.781 Sum_probs=4.1
Q ss_pred eEEEEEe
Q 038524 89 MCVGFSA 95 (175)
Q Consensus 89 ~yVGFSA 95 (175)
+..|||.
T Consensus 144 ~tl~~s~ 150 (341)
T PF06809_consen 144 ATLGFSE 150 (341)
T ss_pred ccccccc
Confidence 5566664
No 66
>PRK10381 LPS O-antigen length regulator; Provisional
Probab=23.10 E-value=46 Score=29.80 Aligned_cols=9 Identities=11% Similarity=0.257 Sum_probs=4.2
Q ss_pred ceeEEehhH
Q 038524 136 QKLVTCVIL 144 (175)
Q Consensus 136 ~~~l~i~l~ 144 (175)
++.+++.++
T Consensus 337 kr~lIlvl~ 345 (377)
T PRK10381 337 GKALIVILA 345 (377)
T ss_pred chhHHHHHH
Confidence 455544433
No 67
>PLN00113 leucine-rich repeat receptor-like protein kinase; Provisional
Probab=22.98 E-value=89 Score=30.56 Aligned_cols=12 Identities=8% Similarity=-0.075 Sum_probs=4.7
Q ss_pred EEehhHHHHHHH
Q 038524 139 VTCVILVALVSV 150 (175)
Q Consensus 139 l~i~l~~~~~~~ 150 (175)
+++++.++++++
T Consensus 630 ~~~~~~~~~~~~ 641 (968)
T PLN00113 630 FYITCTLGAFLV 641 (968)
T ss_pred eehhHHHHHHHH
Confidence 344444443333
No 68
>PRK00523 hypothetical protein; Provisional
Probab=22.46 E-value=53 Score=22.88 Aligned_cols=8 Identities=0% Similarity=-0.330 Sum_probs=3.8
Q ss_pred Hhhhhhee
Q 038524 155 GAAVYIVR 162 (175)
Q Consensus 155 ~~~~~~~r 162 (175)
+.+|+.||
T Consensus 21 ~Gffiark 28 (72)
T PRK00523 21 IGYFVSKK 28 (72)
T ss_pred HHHHHHHH
Confidence 34455444
No 69
>PF03988 DUF347: Repeat of Unknown Function (DUF347) ; InterPro: IPR007136 This repeat is found as four tandem repeats in a family of bacterial membrane proteins. Each repeat contains two transmembrane regions and a conserved tryptophan.
Probab=22.05 E-value=62 Score=20.87 Aligned_cols=16 Identities=6% Similarity=0.046 Sum_probs=7.7
Q ss_pred EEehhHHHHHHHHHHH
Q 038524 139 VTCVILVALVSVLTTV 154 (175)
Q Consensus 139 l~i~l~~~~~~~~~~~ 154 (175)
+.++...+++++.+++
T Consensus 25 lglg~~~~~~~~~~~l 40 (55)
T PF03988_consen 25 LGLGYLISTLIFAALL 40 (55)
T ss_pred cCccHHHHHHHHHHHH
Confidence 4455555544444444
No 70
>KOG3637 consensus Vitronectin receptor, alpha subunit [Extracellular structures]
Probab=21.35 E-value=1.1e+02 Score=31.24 Aligned_cols=14 Identities=7% Similarity=0.143 Sum_probs=5.7
Q ss_pred eEEehhHHHHHHHH
Q 038524 138 LVTCVILVALVSVL 151 (175)
Q Consensus 138 ~l~i~l~~~~~~~~ 151 (175)
.++|+..+++++++
T Consensus 979 wiIi~svl~GLLlL 992 (1030)
T KOG3637|consen 979 WIIILSVLGGLLLL 992 (1030)
T ss_pred eeehHHHHHHHHHH
Confidence 34444444444333
No 71
>COG4282 SMI1 Protein involved in beta-1,3-glucan synthesis [Carbohydrate transport and metabolism]
Probab=21.19 E-value=62 Score=26.35 Aligned_cols=15 Identities=40% Similarity=0.709 Sum_probs=12.9
Q ss_pred CCCCCCCCeeEEEcC
Q 038524 10 QFNDTSNNHVGIDVN 24 (175)
Q Consensus 10 e~~D~~~nHVGIdiN 24 (175)
-+.|+-+||++||+-
T Consensus 123 L~~d~~Gnhi~IDLa 137 (191)
T COG4282 123 LFGDPRGNHICIDLA 137 (191)
T ss_pred ecccCCCCeEEEecC
Confidence 468999999999984
No 72
>PRK11638 lipopolysaccharide biosynthesis protein WzzE; Provisional
Probab=21.16 E-value=43 Score=29.64 Aligned_cols=8 Identities=13% Similarity=0.322 Sum_probs=3.7
Q ss_pred hhhheecc
Q 038524 157 AVYIVRKK 164 (175)
Q Consensus 157 ~~~~~rr~ 164 (175)
+.++||++
T Consensus 334 ~vL~r~~~ 341 (342)
T PRK11638 334 VALTRRRR 341 (342)
T ss_pred eeEeecCC
Confidence 34455543
No 73
>PF02699 YajC: Preprotein translocase subunit; InterPro: IPR003849 Secretion across the inner membrane in some Gram-negative bacteria occurs via the preprotein translocase pathway. Proteins are produced in the cytoplasm as precursors, and require a chaperone subunit to direct them to the translocase component []. From there, the mature proteins are either targeted to the outer membrane, or remain as periplasmic proteins []. The translocase protein subunits are encoded on the bacterial chromosome. The translocase itself comprises 7 proteins, including a chaperone (SecB), ATPase (SecA), an integral membrane complex (SecY, SecE and SecG), and two additional membrane proteins that promote the release of the mature peptide into the periplasm (SecD and SecF) []. Other cytoplasmic/periplasmic proteins play a part in preprotein translocase activity, namely YidC and YajC []. The latter is bound in a complex to SecD and SecF, and plays a part in stabilising and regulating secretion through the SecYEG integral membrane component via SecA []. Homologues of the YajC gene have been found in a range of pathogenic and commensal microbes. Brucella abortis YajC- and SecD-like proteins were shown to stimulate a Th1 cell-mediated immune response in mice, and conferred protection when challenged with B.abortis []. Therefore, these proteins may have an antigenic role as well as a secretory one in virulent bacteria []. A number of previously uncharacterised "hypothetical" proteins also show similarity to E.coli YajC, suggesting that this family is wider than first thought []. More recently, the precise interactions between the E.coli SecYEG complex, SecD, SecF, YajC and YidC have been studied []. Rather than acting individually, the four proteins form a heterotetrameric complex and associate with the SecYEG heterotrimeric complex []. The SecF and YajC subunits link the complex to the integral membrane translocase. ; PDB: 2RDD_B.
Probab=20.69 E-value=73 Score=22.21 Aligned_cols=7 Identities=14% Similarity=-0.083 Sum_probs=3.0
Q ss_pred hhhhhee
Q 038524 156 AAVYIVR 162 (175)
Q Consensus 156 ~~~~~~r 162 (175)
+++.+|.
T Consensus 16 yf~~~rp 22 (82)
T PF02699_consen 16 YFLMIRP 22 (82)
T ss_dssp HHHTHHH
T ss_pred hhheecH
Confidence 4444443
No 74
>TIGR00847 ccoS cytochrome oxidase maturation protein, cbb3-type. CcoS from Rhodobacter capsulatus has been shown essential for incorporation of redox-active prosthetic groups (heme, Cu) into cytochrome cbb(3) oxidase. FixS of Bradyrhizobium japonicum appears to have the same function. Members of this family are found so far in organisms with a cbb3-type cytochrome oxidase, including Neisseria meningitidis, Helicobacter pylori, Campylobacter jejuni, Caulobacter crescentus, Bradyrhizobium japonicum, and Rhodobacter capsulatus.
Probab=20.35 E-value=1.1e+02 Score=19.82 Aligned_cols=13 Identities=23% Similarity=0.473 Sum_probs=5.9
Q ss_pred Hhhhhheecccch
Q 038524 155 GAAVYIVRKKKYD 167 (175)
Q Consensus 155 ~~~~~~~rr~~~~ 167 (175)
+++++.-|+.+|.
T Consensus 20 ~~f~Wavk~GQfD 32 (51)
T TIGR00847 20 VAFLWSLKSGQYD 32 (51)
T ss_pred HHHHHHHccCCCC
Confidence 3444444555543
No 75
>PF07332 DUF1469: Protein of unknown function (DUF1469); InterPro: IPR009937 This entry represents proteins found in hypothetical bacterial proteins where is is annotated as ycf49 or ycf49-like. The function is not known.
Probab=20.31 E-value=56 Score=23.77 Aligned_cols=8 Identities=38% Similarity=0.397 Sum_probs=4.7
Q ss_pred hhhhcccC
Q 038524 167 DEVYEDWE 174 (175)
Q Consensus 167 ~e~~edwE 174 (175)
+|+.+|++
T Consensus 109 ~~l~~d~~ 116 (121)
T PF07332_consen 109 AELKEDIA 116 (121)
T ss_pred HHHHHHHH
Confidence 55666654
No 76
>TIGR01006 polys_exp_MPA1 polysaccharide export protein, MPA1 family, Gram-positive type. This family contains members from Low GC Gram-positive bacteria; they are proposed to have a function in the export of complex polysaccharides.
Probab=20.07 E-value=50 Score=26.74 Aligned_cols=14 Identities=21% Similarity=0.067 Sum_probs=7.1
Q ss_pred hhheecccchhhhc
Q 038524 158 VYIVRKKKYDEVYE 171 (175)
Q Consensus 158 ~~~~rr~~~~e~~e 171 (175)
.++.++-|-.|..|
T Consensus 197 ~~~d~~i~~~~d~~ 210 (226)
T TIGR01006 197 ELLDTRVKRPEDVE 210 (226)
T ss_pred HHHhCCcCCHHHHH
Confidence 34555555555444
Done!