Query 036189
Match_columns 241
No_of_seqs 151 out of 675
Neff 7.2
Searched_HMMs 46136
Date Fri Mar 29 10:17:54 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/036189.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/036189hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PLN03160 uncharacterized prote 100.0 2.1E-37 4.6E-42 265.9 25.6 199 1-223 1-204 (219)
2 PF03168 LEA_2: Late embryogen 99.4 2.1E-12 4.5E-17 96.6 10.3 99 105-215 1-101 (101)
3 smart00769 WHy Water Stress an 98.3 1.2E-05 2.6E-10 60.6 11.4 61 96-160 11-72 (100)
4 PF07092 DUF1356: Protein of u 97.5 0.012 2.6E-07 51.1 16.6 83 74-160 96-181 (238)
5 PF12751 Vac7: Vacuolar segreg 97.4 0.001 2.2E-08 61.3 8.9 73 53-132 307-379 (387)
6 COG5608 LEA14-like dessication 95.4 1.2 2.6E-05 36.2 17.8 107 77-199 31-138 (161)
7 PLN03160 uncharacterized prote 91.2 1.5 3.2E-05 37.8 8.5 102 38-152 32-146 (219)
8 TIGR02588 conserved hypothetic 87.8 1.4 3.1E-05 34.5 5.2 48 60-113 13-62 (122)
9 PF09307 MHC2-interact: CLIP, 80.2 0.54 1.2E-05 36.4 0.0 36 41-77 24-59 (114)
10 PRK10893 lipopolysaccharide ex 65.2 37 0.0008 28.5 7.6 30 75-105 37-66 (192)
11 KOG3950 Gamma/delta sarcoglyca 64.6 9.1 0.0002 33.6 3.8 22 97-118 105-126 (292)
12 PF09624 DUF2393: Protein of u 64.4 34 0.00074 27.1 7.0 64 61-132 28-93 (149)
13 PF06072 Herpes_US9: Alphaherp 60.5 4.1 8.8E-05 27.8 0.7 8 65-72 52-59 (60)
14 PF14155 DUF4307: Domain of un 60.0 76 0.0016 24.3 8.6 28 127-160 71-100 (112)
15 COG1580 FliL Flagellar basal b 59.5 27 0.00059 28.6 5.6 22 53-74 21-42 (159)
16 PF07787 DUF1625: Protein of u 58.2 10 0.00023 33.0 3.2 17 60-76 232-248 (248)
17 PRK13183 psbN photosystem II r 56.9 17 0.00036 23.5 3.1 21 57-77 13-33 (46)
18 COG3671 Predicted membrane pro 55.4 3.9 8.4E-05 31.9 -0.0 33 43-75 68-103 (125)
19 PRK07021 fliL flagellar basal 53.8 59 0.0013 26.4 6.8 17 116-132 77-93 (162)
20 PRK05529 cell division protein 53.3 30 0.00064 30.4 5.2 43 78-121 58-128 (255)
21 PF12505 DUF3712: Protein of u 53.2 1E+02 0.0022 23.7 9.5 66 136-211 3-69 (125)
22 CHL00020 psbN photosystem II p 52.8 18 0.0004 23.0 2.7 21 56-76 9-29 (43)
23 PF08113 CoxIIa: Cytochrome c 52.6 6 0.00013 23.8 0.5 13 60-72 12-24 (34)
24 PF01102 Glycophorin_A: Glycop 50.5 6.3 0.00014 30.9 0.4 25 61-85 76-101 (122)
25 PRK06531 yajC preprotein trans 50.2 7.3 0.00016 30.2 0.8 13 66-78 12-24 (113)
26 PF04478 Mid2: Mid2 like cell 48.2 3.6 7.9E-05 33.5 -1.2 42 61-119 62-103 (154)
27 PF02468 PsbN: Photosystem II 47.6 16 0.00034 23.4 1.8 18 59-76 12-29 (43)
28 PF14927 Neurensin: Neurensin 47.4 48 0.0011 26.6 5.1 12 60-71 54-65 (140)
29 KOG0810 SNARE protein Syntaxin 44.6 7.6 0.00016 35.1 0.1 14 37-50 263-276 (297)
30 COG4698 Uncharacterized protei 42.8 23 0.0005 29.7 2.7 30 66-95 26-58 (197)
31 PF02009 Rifin_STEVOR: Rifin/s 42.3 9.6 0.00021 34.5 0.4 17 59-75 264-280 (299)
32 PF11322 DUF3124: Protein of u 41.4 1.7E+02 0.0037 23.1 7.2 53 96-154 19-73 (125)
33 PF04790 Sarcoglycan_1: Sarcog 40.9 2.6E+02 0.0056 24.8 11.6 17 98-114 84-100 (264)
34 PF06637 PV-1: PV-1 protein (P 40.7 36 0.00079 31.8 3.8 12 60-71 38-49 (442)
35 PHA02844 putative transmembran 39.9 30 0.00065 24.7 2.5 11 62-72 59-69 (75)
36 PF14283 DUF4366: Domain of un 38.3 32 0.0007 29.7 3.0 21 61-81 170-190 (218)
37 PF13131 DUF3951: Protein of u 36.7 25 0.00054 23.2 1.6 31 50-80 3-34 (53)
38 PF11906 DUF3426: Protein of u 36.5 2E+02 0.0044 22.4 7.3 57 81-141 48-106 (149)
39 PRK08455 fliL flagellar basal 36.1 52 0.0011 27.5 3.8 15 118-132 103-117 (182)
40 PF04573 SPC22: Signal peptida 35.9 2.4E+02 0.0052 23.4 7.8 32 97-132 65-97 (175)
41 PF06024 DUF912: Nucleopolyhed 35.6 48 0.001 24.9 3.2 14 63-76 75-89 (101)
42 PF10907 DUF2749: Protein of u 35.3 51 0.0011 22.9 3.0 16 62-77 13-28 (66)
43 PF03302 VSP: Giardia variant- 33.8 14 0.00031 34.6 0.1 34 43-77 362-396 (397)
44 PF09911 DUF2140: Uncharacteri 33.3 69 0.0015 26.8 4.2 20 60-79 12-31 (187)
45 PF15012 DUF4519: Domain of un 33.3 36 0.00078 23.0 1.9 15 63-77 42-56 (56)
46 PF09865 DUF2092: Predicted pe 33.0 3E+02 0.0065 23.6 8.1 36 96-132 35-72 (214)
47 PF06092 DUF943: Enterobacteri 32.9 28 0.0006 28.5 1.6 16 61-76 13-28 (157)
48 PTZ00116 signal peptidase; Pro 32.5 2.8E+02 0.0061 23.4 7.6 52 76-131 36-94 (185)
49 PF12505 DUF3712: Protein of u 32.0 1.1E+02 0.0023 23.5 4.8 27 98-125 98-124 (125)
50 PHA02650 hypothetical protein; 31.9 30 0.00065 25.0 1.5 11 62-72 60-70 (81)
51 PF05478 Prominin: Prominin; 31.9 44 0.00096 34.3 3.3 27 41-67 133-159 (806)
52 PF11395 DUF2873: Protein of u 31.2 22 0.00048 21.9 0.6 9 66-74 24-32 (43)
53 PF15145 DUF4577: Domain of un 30.4 46 0.001 25.7 2.4 26 49-76 63-88 (128)
54 PF02038 ATP1G1_PLM_MAT8: ATP1 30.1 71 0.0015 21.0 2.9 16 53-68 18-33 (50)
55 TIGR01477 RIFIN variant surfac 29.9 20 0.00044 33.1 0.4 24 53-76 312-335 (353)
56 PTZ00046 rifin; Provisional 28.9 22 0.00047 33.0 0.4 24 53-76 317-340 (358)
57 PRK05696 fliL flagellar basal 27.8 1.5E+02 0.0032 24.2 5.2 17 116-132 85-101 (170)
58 PF04505 Dispanin: Interferon- 27.7 1.3E+02 0.0028 21.6 4.3 8 63-70 33-40 (82)
59 PRK12785 fliL flagellar basal 27.7 1.5E+02 0.0033 24.2 5.2 16 117-132 86-101 (166)
60 COG5009 MrcA Membrane carboxyp 27.6 29 0.00063 35.3 1.1 31 52-82 8-38 (797)
61 PHA03093 EEV glycoprotein; Pro 27.0 40 0.00087 28.3 1.6 21 109-130 97-117 (185)
62 PF10614 CsgF: Type VIII secre 26.9 44 0.00096 26.9 1.8 11 71-81 23-33 (142)
63 PF13396 PLDc_N: Phospholipase 24.8 84 0.0018 19.5 2.6 16 62-77 31-46 (46)
64 PHA02673 ORF109 EEV glycoprote 24.4 28 0.00061 28.5 0.3 8 115-122 78-85 (161)
65 PF15050 SCIMP: SCIMP protein 23.7 31 0.00068 27.0 0.4 10 66-75 23-32 (133)
66 PF04790 Sarcoglycan_1: Sarcog 23.6 95 0.0021 27.6 3.5 11 211-222 199-209 (264)
67 PF06835 LptC: Lipopolysacchar 23.5 78 0.0017 25.0 2.8 51 78-130 32-82 (176)
68 PF05545 FixQ: Cbb3-type cytoc 23.2 48 0.001 21.2 1.2 13 66-78 22-34 (49)
69 COG4736 CcoQ Cbb3-type cytochr 23.0 49 0.0011 22.7 1.2 13 66-78 22-34 (60)
70 PF01034 Syndecan: Syndecan do 22.9 28 0.0006 24.2 0.0 15 62-76 22-36 (64)
71 COG1589 FtsQ Cell division sep 22.9 1.2E+02 0.0026 26.7 4.0 30 61-90 40-69 (269)
72 PF12321 DUF3634: Protein of u 22.8 32 0.00068 26.4 0.3 17 67-83 10-28 (108)
73 PTZ00382 Variant-specific surf 22.8 55 0.0012 24.4 1.6 15 61-75 78-93 (96)
74 PRK14759 potassium-transportin 22.7 41 0.00089 19.6 0.7 19 60-78 10-28 (29)
75 PF08693 SKG6: Transmembrane a 22.3 74 0.0016 20.0 1.8 11 66-76 28-38 (40)
76 PF13473 Cupredoxin_1: Cupredo 22.1 66 0.0014 23.7 1.9 35 80-114 19-55 (104)
77 PHA03049 IMV membrane protein; 21.9 36 0.00077 23.8 0.4 17 60-76 9-25 (68)
78 PF13800 Sigma_reg_N: Sigma fa 21.7 33 0.00072 25.2 0.2 12 104-115 56-67 (96)
79 PF05961 Chordopox_A13L: Chord 21.4 40 0.00086 23.6 0.5 17 60-76 9-25 (68)
80 PF00927 Transglut_C: Transglu 21.2 3.3E+02 0.0071 19.9 5.6 61 97-160 12-76 (107)
81 PF15018 InaF-motif: TRP-inter 21.1 1.4E+02 0.003 18.5 2.8 20 60-79 17-37 (38)
82 COG5294 Uncharacterized protei 20.9 2.5E+02 0.0054 21.7 4.8 13 101-113 53-65 (113)
83 COG2332 CcmE Cytochrome c-type 20.9 3.3E+02 0.0072 22.2 5.7 32 101-132 71-102 (153)
84 PHA03265 envelope glycoprotein 20.5 58 0.0013 30.2 1.5 23 47-71 350-372 (402)
85 PF01299 Lamp: Lysosome-associ 20.5 81 0.0017 28.3 2.5 18 61-78 282-299 (306)
86 PF02158 Neuregulin: Neureguli 20.5 35 0.00077 31.9 0.1 23 48-72 9-32 (404)
87 PF14828 Amnionless: Amnionles 20.4 61 0.0013 30.9 1.8 21 61-81 349-369 (437)
88 PF09049 SNN_transmemb: Stanni 20.3 1.2E+02 0.0025 17.9 2.2 16 52-67 14-29 (33)
No 1
>PLN03160 uncharacterized protein; Provisional
Probab=100.00 E-value=2.1e-37 Score=265.86 Aligned_cols=199 Identities=13% Similarity=0.139 Sum_probs=159.6
Q ss_pred CccccccccccccCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCchhhhHHHHHHHHHHHHHHHHHhheeeEEecCCCCE
Q 036189 1 MEENEIIQNSAHRCPSKVYPLTTGDISQLPPSRPPHYQHFQIKKLPKLIIITLLVVAASISLTALICILMYFTLGPKLPS 80 (241)
Q Consensus 1 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~p~~~~p~~~~~~~~~~~rc~~~~~~~~~~~i~llgl~~lil~lv~rPk~P~ 80 (241)
|-|+||+||+|-.+++. .+|.++. .+++++.+|++|++||+|++.++ ++++++++.++|++||||+|+
T Consensus 1 ~~~~~~~~p~a~~~~~~-----~~d~~~~----~~~~~~~~r~~~~~c~~~~~a~~---l~l~~v~~~l~~~vfrPk~P~ 68 (219)
T PLN03160 1 MAETEQVRPLAPAAFRL-----RSDEEEA----TNHLKKTRRRNCIKCCGCITATL---LILATTILVLVFTVFRVKDPV 68 (219)
T ss_pred CCccccCCCCCCCcccc-----cCchhhc----CcchhccccccceEEHHHHHHHH---HHHHHHHHheeeEEEEccCCe
Confidence 89999999999988872 1222221 12222234555666655444333 355677788889999999999
Q ss_pred EEEeeEEEeeeecCC-----CceeEEEEEEEEEeCCCCeeEEEEccEEEEEEeCCcccceeeeccCCCceecCCCeEEEE
Q 036189 81 LHLDTFSVSNFTIGS-----TNLIAKWDFNLTFKNPDHLWQIYLDYIECIALNHDHFPIAINHSVSPPFKVKPMKKSTIH 155 (241)
Q Consensus 81 f~V~s~~l~~f~~~~-----~~l~~~~~~~l~v~NPN~k~~i~Y~~~~v~v~Y~g~~~~~lg~~~vp~F~q~~~~tt~v~ 155 (241)
|+|++++|++|+++. ..+|++++++++++|||+ ++|+|+++++.++|+|+. +|++.+|+|+|++++++.++
T Consensus 69 ~~v~~v~l~~~~~~~~~~~~~~~n~tl~~~v~v~NPN~-~~~~Y~~~~~~v~Y~g~~---vG~a~~p~g~~~ar~T~~l~ 144 (219)
T PLN03160 69 IKMNGVTVTKLELINNTTLRPGTNITLIADVSVKNPNV-ASFKYSNTTTTIYYGGTV---VGEARTPPGKAKARRTMRMN 144 (219)
T ss_pred EEEEEEEEeeeeeccCCCCceeEEEEEEEEEEEECCCc-eeEEEcCeEEEEEECCEE---EEEEEcCCcccCCCCeEEEE
Confidence 999999999999864 357888889999999999 899999999999999999 99999999999999999999
Q ss_pred EEEEeCCceeecCHHHHHHHHHHhhCCceEEEEEEEEEEEEEEEEeceEEEeeeeEEEEecceEEeee
Q 036189 156 VQLATGDSLIFLNHQLLQKINSQRRNGRMVVFGLAVRAKTRFTGVSWLWWTEFANLMYTCLDLKVGFK 223 (241)
Q Consensus 156 v~l~~~~~~~~l~~~~~~~l~~d~~~G~~v~~~v~v~~~vr~kv~~g~~~~~~~~~~v~C~~l~V~~~ 223 (241)
+++.. .....+.. ..|..|..+|. ++|++++++++++++ |+++++++.++++| +++|++.
T Consensus 145 ~tv~~-~~~~~~~~---~~L~~D~~~G~-v~l~~~~~v~gkVkv--~~i~k~~v~~~v~C-~v~V~~~ 204 (219)
T PLN03160 145 VTVDI-IPDKILSV---PGLLTDISSGL-LNMNSYTRIGGKVKI--LKIIKKHVVVKMNC-TMTVNIT 204 (219)
T ss_pred EEEEE-Eeceeccc---hhHHHHhhCCe-EEEEEEEEEEEEEEE--EEEEEEEEEEEEEe-EEEEECC
Confidence 99765 21122221 46888999999 999999999999999 99999999999999 9999883
No 2
>PF03168 LEA_2: Late embryogenesis abundant protein; InterPro: IPR004864 Different types of LEA proteins are expressed at different stages of late embryogenesis in higher plant seed embryos and under conditions of dehydration stress [, ]. The function of these proteins is unknown. ; PDB: 3BUT_A 1XO8_A 1YYC_A.
Probab=99.41 E-value=2.1e-12 Score=96.59 Aligned_cols=99 Identities=20% Similarity=0.332 Sum_probs=72.8
Q ss_pred EEEEeCCCCeeEEEEccEEEEEEeCCcccceee-eccCCCceecCCCeEEEEEEEEeCCceeecCHHHHHHHHHHhhCCc
Q 036189 105 NLTFKNPDHLWQIYLDYIECIALNHDHFPIAIN-HSVSPPFKVKPMKKSTIHVQLATGDSLIFLNHQLLQKINSQRRNGR 183 (241)
Q Consensus 105 ~l~v~NPN~k~~i~Y~~~~v~v~Y~g~~~~~lg-~~~vp~F~q~~~~tt~v~v~l~~~~~~~~l~~~~~~~l~~d~~~G~ 183 (241)
+|+++|||. ++++|+++++.++|+|+. +| ....++|+|++++++.+.+.+.... ..+.+.+.++. .|.
T Consensus 1 ~l~v~NPN~-~~i~~~~~~~~v~~~g~~---v~~~~~~~~~~i~~~~~~~v~~~v~~~~------~~l~~~l~~~~-~~~ 69 (101)
T PF03168_consen 1 TLSVRNPNS-FGIRYDSIEYDVYYNGQR---VGTGGSLPPFTIPARSSTTVPVPVSVDY------SDLPRLLKDLL-AGR 69 (101)
T ss_dssp EEEEEESSS-S-EEEEEEEEEEEESSSE---EEEEEECE-EEESSSCEEEEEEEEEEEH------HHHHHHHHHHH-HTT
T ss_pred CEEEECCCc-eeEEEeCEEEEEEECCEE---EECccccCCeEECCCCcEEEEEEEEEcH------HHHHHHHHhhh-ccc
Confidence 589999999 999999999999999998 99 7789999999999999988877721 22245666666 556
Q ss_pred eEEEEEEEEEEEEEEE-EeceEEEeeeeEEEEe
Q 036189 184 MVVFGLAVRAKTRFTG-VSWLWWTEFANLMYTC 215 (241)
Q Consensus 184 ~v~~~v~v~~~vr~kv-~~g~~~~~~~~~~v~C 215 (241)
..+++.+++++++++ ..+.+.+.++.++.+|
T Consensus 70 -~~~~v~~~~~g~~~v~~~~~~~~~~v~~~~~~ 101 (101)
T PF03168_consen 70 -VPFDVTYRIRGTFKVLGTPIFGSVRVPVSCEC 101 (101)
T ss_dssp -SCEEEEEEEEEEEE-EE-TTTSCEEEEEEEEE
T ss_pred -cceEEEEEEEEEEEEcccceeeeEEEeEEeEC
Confidence 677777888888883 2244454555555554
No 3
>smart00769 WHy Water Stress and Hypersensitive response.
Probab=98.33 E-value=1.2e-05 Score=60.63 Aligned_cols=61 Identities=10% Similarity=0.132 Sum_probs=56.3
Q ss_pred CceeEEEEEEEEEeCCCCeeEEEEccEEEEEEeCCcccceeeeccCC-CceecCCCeEEEEEEEEe
Q 036189 96 TNLIAKWDFNLTFKNPDHLWQIYLDYIECIALNHDHFPIAINHSVSP-PFKVKPMKKSTIHVQLAT 160 (241)
Q Consensus 96 ~~l~~~~~~~l~v~NPN~k~~i~Y~~~~v~v~Y~g~~~~~lg~~~vp-~F~q~~~~tt~v~v~l~~ 160 (241)
+.++.++.+++.+.|||. ..+.|+.++..++|+|.. +|++..+ ++..++++++.+.+.+..
T Consensus 11 ~~~~~~~~l~l~v~NPN~-~~l~~~~~~y~l~~~g~~---v~~g~~~~~~~ipa~~~~~v~v~~~~ 72 (100)
T smart00769 11 SGLEIEIVLKVKVQNPNP-FPIPVNGLSYDLYLNGVE---LGSGEIPDSGTLPGNGRTVLDVPVTV 72 (100)
T ss_pred cceEEEEEEEEEEECCCC-CccccccEEEEEEECCEE---EEEEEcCCCcEECCCCcEEEEEEEEe
Confidence 367789999999999998 999999999999999999 9999986 799999999999888877
No 4
>PF07092 DUF1356: Protein of unknown function (DUF1356); InterPro: IPR009790 This family consists of several hypothetical mammalian proteins of around 250 residues in length. The function of this family is unknown.
Probab=97.49 E-value=0.012 Score=51.12 Aligned_cols=83 Identities=11% Similarity=0.178 Sum_probs=58.5
Q ss_pred ecCCCCEEEEeeEEEee--eecCCCceeEEEEEEEEEeCCCCeeEEEEccEEEEEEeCCcccceeeeccCCC-ceecCCC
Q 036189 74 LGPKLPSLHLDTFSVSN--FTIGSTNLIAKWDFNLTFKNPDHLWQIYLDYIECIALNHDHFPIAINHSVSPP-FKVKPMK 150 (241)
Q Consensus 74 ~rPk~P~f~V~s~~l~~--f~~~~~~l~~~~~~~l~v~NPN~k~~i~Y~~~~v~v~Y~g~~~~~lg~~~vp~-F~q~~~~ 150 (241)
+-||.-.++-.++.... |+-+.+.+..+++-.+.++|||- ..+.-.++.+.+.|...- +|.+.... ...++++
T Consensus 96 LfPRsV~v~~~gv~s~~V~f~~~~~~v~l~itn~lNIsN~NF-y~V~Vt~~s~qv~~~~~V---VG~~~~~~~~~I~Prs 171 (238)
T PF07092_consen 96 LFPRSVTVSPVGVKSVTVSFNPDKSTVQLNITNTLNISNPNF-YPVTVTNLSIQVLYMKTV---VGKGKNSNITVIGPRS 171 (238)
T ss_pred EeCcEEEEecCcEEEEEEEEeCCCCEEEEEEEEEEEccCCCE-EEEEEEeEEEEEEEEEeE---EeeeEecceEEecccC
Confidence 34664444333322222 33333568889999999999996 999999999999998877 99887654 4677777
Q ss_pred eEEEEEEEEe
Q 036189 151 KSTIHVQLAT 160 (241)
Q Consensus 151 tt~v~v~l~~ 160 (241)
.+.+..++..
T Consensus 172 ~~q~~~tV~t 181 (238)
T PF07092_consen 172 SKQVNYTVKT 181 (238)
T ss_pred CceEEEEeeE
Confidence 7777666555
No 5
>PF12751 Vac7: Vacuolar segregation subunit 7; InterPro: IPR024260 Vac7 is localised at the vacuole membrane, a location which is consistent with its involvement in vacuole morphology and inheritance []. Vac7 has been shown to function as an upstream regulator of the Fab1 lipid kinase pathway []. The Fab1 lipid pathway is important for correct regulation of membrane trafficking events.
Probab=97.36 E-value=0.001 Score=61.26 Aligned_cols=73 Identities=14% Similarity=0.238 Sum_probs=44.8
Q ss_pred HHHHHHHHHHHHHHhheeeEEecCCCCEEEEeeEEEeeeecCCCceeEEEEEEEEEeCCCCeeEEEEccEEEEEEeCCcc
Q 036189 53 LLVVAASISLTALICILMYFTLGPKLPSLHLDTFSVSNFTIGSTNLIAKWDFNLTFKNPDHLWQIYLDYIECIALNHDHF 132 (241)
Q Consensus 53 ~~~~~~~i~llgl~~lil~lv~rPk~P~f~V~s~~l~~f~~~~~~l~~~~~~~l~v~NPN~k~~i~Y~~~~v~v~Y~g~~ 132 (241)
++.+++++++.|++.++|. .-+| --.|+=..|.+.-.+ .--.-|+++|.+.|||. +.|.-++.++.++-+-..
T Consensus 307 ~~~i~~lL~ig~~~gFv~A-ttKp---L~~v~v~~I~NVlaS--~qELmfdl~V~A~NPn~-~~V~I~d~dldIFAKS~y 379 (387)
T PF12751_consen 307 YLSILLLLVIGFAIGFVFA-TTKP---LTDVQVVSIQNVLAS--EQELMFDLTVEAFNPNW-FTVTIDDMDLDIFAKSRY 379 (387)
T ss_pred HHHHHHHHHHHHHHHhhhh-cCcc---cccceEEEeeeeeec--cceEEEeeEEEEECCCe-EEEEeccceeeeEecCCc
Confidence 3343333444444554444 3333 333333444443333 34466889999999998 999999999999876554
No 6
>COG5608 LEA14-like dessication related protein [Defense mechanisms]
Probab=95.37 E-value=1.2 Score=36.21 Aligned_cols=107 Identities=19% Similarity=0.159 Sum_probs=73.6
Q ss_pred CCCEEEEeeEEEeeeecCCCceeEEEEEEEEEeCCCCeeEEEEccEEEEEEeCCcccceeeeccC-CCceecCCCeEEEE
Q 036189 77 KLPSLHLDTFSVSNFTIGSTNLIAKWDFNLTFKNPDHLWQIYLDYIECIALNHDHFPIAINHSVS-PPFKVKPMKKSTIH 155 (241)
Q Consensus 77 k~P~f~V~s~~l~~f~~~~~~l~~~~~~~l~v~NPN~k~~i~Y~~~~v~v~Y~g~~~~~lg~~~v-p~F~q~~~~tt~v~ 155 (241)
+.|...--.+..-... ...-.+-.++.++|||. ..+--..++..+|-+|-. +|.+.. .++..++++...+.
T Consensus 31 ~~p~ve~~ka~wGkvt----~s~~EiV~t~KiyNPN~-fPipVtgl~y~vymN~Ik---i~eG~~~k~~~v~p~S~~tvd 102 (161)
T COG5608 31 KKPGVESMKAKWGKVT----NSETEIVGTLKIYNPNP-FPIPVTGLQYAVYMNDIK---IGEGEILKGTTVPPNSRETVD 102 (161)
T ss_pred CCCCceEEEEEEEEEe----ccceEEEEEEEecCCCC-cceeeeceEEEEEEcceE---eeccccccceEECCCCeEEEE
Confidence 4455555555554432 24457888999999998 999999999999999988 999875 56999999999998
Q ss_pred EEEEeCCceeecCHHHHHHHHHHhhCCceEEEEEEEEEEEEEEE
Q 036189 156 VQLATGDSLIFLNHQLLQKINSQRRNGRMVVFGLAVRAKTRFTG 199 (241)
Q Consensus 156 v~l~~~~~~~~l~~~~~~~l~~d~~~G~~v~~~v~v~~~vr~kv 199 (241)
+.+.. + .+..-+.......+|.+-.+++++ +..+++
T Consensus 103 v~l~~-d-----~~~~ke~w~~hi~ngErs~Ir~~i--~~~v~v 138 (161)
T COG5608 103 VPLRL-D-----NSKIKEWWVTHIENGERSTIRVRI--KGVVKV 138 (161)
T ss_pred EEEEE-e-----hHHHHHHHHHHhhccCcccEEEEE--EEEEEE
Confidence 88877 2 222334455556677632333333 334455
No 7
>PLN03160 uncharacterized protein; Provisional
Probab=91.17 E-value=1.5 Score=37.80 Aligned_cols=102 Identities=11% Similarity=0.073 Sum_probs=52.4
Q ss_pred CCCCCCCchhhhHHHHHHHHHHHHHHHHHhheeeEEecCC--CCEEEEeeEEEe-------eeecCC----CceeEEEEE
Q 036189 38 QHFQIKKLPKLIIITLLVVAASISLTALICILMYFTLGPK--LPSLHLDTFSVS-------NFTIGS----TNLIAKWDF 104 (241)
Q Consensus 38 ~~~~~~~~~rc~~~~~~~~~~~i~llgl~~lil~lv~rPk--~P~f~V~s~~l~-------~f~~~~----~~l~~~~~~ 104 (241)
+|+++.+||.|++..+++++ +++++++++++=.=+|+ .-.++|+++.+. .+|++- ..-|.|. +
T Consensus 32 ~r~~~~~c~~~~~a~~l~l~---~v~~~l~~~vfrPk~P~~~v~~v~l~~~~~~~~~~~~~~~n~tl~~~v~v~NPN~-~ 107 (219)
T PLN03160 32 RRRNCIKCCGCITATLLILA---TTILVLVFTVFRVKDPVIKMNGVTVTKLELINNTTLRPGTNITLIADVSVKNPNV-A 107 (219)
T ss_pred ccccceEEHHHHHHHHHHHH---HHHHheeeEEEEccCCeEEEEEEEEeeeeeccCCCCceeEEEEEEEEEEEECCCc-e
Confidence 45667778888887776664 23333344444345553 345666666553 233321 1123444 3
Q ss_pred EEEEeCCCCeeEEEEccEEEEEEeCCcccceeeeccCCCceecCCCeE
Q 036189 105 NLTFKNPDHLWQIYLDYIECIALNHDHFPIAINHSVSPPFKVKPMKKS 152 (241)
Q Consensus 105 ~l~v~NPN~k~~i~Y~~~~v~v~Y~g~~~~~lg~~~vp~F~q~~~~tt 152 (241)
.+.-. |..+.++|+...+.-. . +..+.++++.+..-+.+
T Consensus 108 ~~~Y~--~~~~~v~Y~g~~vG~a----~---~p~g~~~ar~T~~l~~t 146 (219)
T PLN03160 108 SFKYS--NTTTTIYYGGTVVGEA----R---TPPGKAKARRTMRMNVT 146 (219)
T ss_pred eEEEc--CeEEEEEECCEEEEEE----E---cCCcccCCCCeEEEEEE
Confidence 44443 4458889977654321 2 33344455555555544
No 8
>TIGR02588 conserved hypothetical protein TIGR02588. The function of this protein is unknown. It is always found as part of a two-gene operon with TIGR02587, a protein that appears to span the membrane seven times. It is found in Nostoc sp. PCC 7120, Agrobacterium tumefaciens, Sinorhizobium meliloti, and Gloeobacter violaceus, so far, all of which are bacterial.
Probab=87.79 E-value=1.4 Score=34.48 Aligned_cols=48 Identities=15% Similarity=0.241 Sum_probs=32.2
Q ss_pred HHHHHHHhheee--EEecCCCCEEEEeeEEEeeeecCCCceeEEEEEEEEEeCCCC
Q 036189 60 ISLTALICILMY--FTLGPKLPSLHLDTFSVSNFTIGSTNLIAKWDFNLTFKNPDH 113 (241)
Q Consensus 60 i~llgl~~lil~--lv~rPk~P~f~V~s~~l~~f~~~~~~l~~~~~~~l~v~NPN~ 113 (241)
+++++++.+++| +.-+++.|.+.+......+ .....+-+-++++|--.
T Consensus 13 ~ill~viglv~y~~l~~~~~pp~l~v~~~~~~r------~~~gqyyVpF~V~N~gg 62 (122)
T TIGR02588 13 LILAAMFGLVAYDWLRYSNKAAVLEVAPAEVER------MQTGQYYVPFAIHNLGG 62 (122)
T ss_pred HHHHHHHHHHHHHhhccCCCCCeEEEeehheeE------EeCCEEEEEEEEEeCCC
Confidence 456666667775 5566788999888877655 23345667777777654
No 9
>PF09307 MHC2-interact: CLIP, MHC2 interacting; InterPro: IPR015386 This domain is found in MHC class II-associated invariant chain (Ii), and in class II invariant chain-associated peptide (CLIP), and is required for association with class II major histocompatibility complex (MHC II) in the MHC II processing pathway []. Ii plays a critical role in the assembly of the MHC, as well as in MHC II antigen processing by stabilising peptide-free class II alpha/beta heterodimers in a complex soon after their synthesis and directing transport of the complex from the endoplasmic reticulum to compartments where peptide loading of class II takes place []. In antigen-presenting cells (APCs), loading of MHC II molecules with peptides is regulated by Ii, which blocks MHC II antigen-binding sites in pre-endosomal compartments []. Several factors modulate the surface expression of MHC II molecules via post-Golgi mechanisms, including CLIP. The Invariant chain contains a single transmembrane domain. Ii first assembles into a trimer and then associates with three class II alpha/beta MHC heterodimers. Although the membrane-proximal region of the Ii luminal domain is structurally disordered, the C-terminal segment of the luminal domain is largely alpha-helical and contains a major interaction site for the Ii trimer []. More information about these proteins can be found at Protein of the Month: MHC [].; GO: 0042289 MHC class II protein binding, 0006886 intracellular protein transport, 0006955 immune response, 0019882 antigen processing and presentation, 0016020 membrane; PDB: 1A6A_C 3QXD_F 3QXA_F 3PDO_C 1MUJ_C 3PGD_F 3PGC_F.
Probab=80.18 E-value=0.54 Score=36.39 Aligned_cols=36 Identities=19% Similarity=0.296 Sum_probs=0.0
Q ss_pred CCCCchhhhHHHHHHHHHHHHHHHHHhheeeEEecCC
Q 036189 41 QIKKLPKLIIITLLVVAASISLTALICILMYFTLGPK 77 (241)
Q Consensus 41 ~~~~~~rc~~~~~~~~~~~i~llgl~~lil~lv~rPk 77 (241)
+|.+|.|++.++.+.+++.++|+|- ++..|++|.=+
T Consensus 24 ~~~s~sra~~vagltvLa~LLiAGQ-a~TaYfv~~Qk 59 (114)
T PF09307_consen 24 QRGSCSRALKVAGLTVLACLLIAGQ-AVTAYFVFQQK 59 (114)
T ss_dssp -------------------------------------
T ss_pred CCCCccchhHHHHHHHHHHHHHHhH-HHHHHHHHHhH
Confidence 4567889999888777766777775 45556666653
No 10
>PRK10893 lipopolysaccharide exporter periplasmic protein; Provisional
Probab=65.16 E-value=37 Score=28.54 Aligned_cols=30 Identities=10% Similarity=0.008 Sum_probs=22.7
Q ss_pred cCCCCEEEEeeEEEeeeecCCCceeEEEEEE
Q 036189 75 GPKLPSLHLDTFSVSNFTIGSTNLIAKWDFN 105 (241)
Q Consensus 75 rPk~P~f~V~s~~l~~f~~~~~~l~~~~~~~ 105 (241)
.++.|.|.+++++...|+.++ .+++.++..
T Consensus 37 ~~~~Pdy~~~~~~~~~yd~~G-~l~y~l~a~ 66 (192)
T PRK10893 37 NNNDPTYQSQHTDTVVYNPEG-ALSYKLVAQ 66 (192)
T ss_pred CCCCCCEEEeccEEEEECCCC-CEEEEEEec
Confidence 356799999999999999875 555555443
No 11
>KOG3950 consensus Gamma/delta sarcoglycan [Cytoskeleton]
Probab=64.57 E-value=9.1 Score=33.61 Aligned_cols=22 Identities=14% Similarity=0.102 Sum_probs=16.4
Q ss_pred ceeEEEEEEEEEeCCCCeeEEE
Q 036189 97 NLIAKWDFNLTFKNPDHLWQIY 118 (241)
Q Consensus 97 ~l~~~~~~~l~v~NPN~k~~i~ 118 (241)
.+...=++++.++|||.++.=+
T Consensus 105 ~~~S~rnvtvnarn~~g~v~~~ 126 (292)
T KOG3950|consen 105 YLQSARNVTVNARNPNGKVTGQ 126 (292)
T ss_pred EEEeccCeeEEccCCCCceeee
Confidence 3455667899999999887533
No 12
>PF09624 DUF2393: Protein of unknown function (DUF2393); InterPro: IPR013417 The function of this protein is unknown. It is always found as part of a two-gene operon with IPR013416 from INTERPRO, a protein that appears to span the membrane seven times. It has so far been found in the bacteria Anabaena sp. (strain PCC 7120), Agrobacterium tumefaciens, Rhizobium meliloti, and Gloeobacter violaceus.
Probab=64.44 E-value=34 Score=27.14 Aligned_cols=64 Identities=20% Similarity=0.177 Sum_probs=39.4
Q ss_pred HHHHHHhheeeEEecC--CCCEEEEeeEEEeeeecCCCceeEEEEEEEEEeCCCCeeEEEEccEEEEEEeCCcc
Q 036189 61 SLTALICILMYFTLGP--KLPSLHLDTFSVSNFTIGSTNLIAKWDFNLTFKNPDHLWQIYLDYIECIALNHDHF 132 (241)
Q Consensus 61 ~llgl~~lil~lv~rP--k~P~f~V~s~~l~~f~~~~~~l~~~~~~~l~v~NPN~k~~i~Y~~~~v~v~Y~g~~ 132 (241)
+++.++.+++|.++.. +.+..++.+.+- ++.+ -.+.+..+++|-.+ ..+..=.+++.+...+..
T Consensus 28 i~~~~~~~~~~~~l~~~~~~~~~~~~~~~~--l~~~-----~~~~v~g~V~N~g~-~~i~~c~i~~~l~~~~~~ 93 (149)
T PF09624_consen 28 ILAFLIPFFGYYWLDKYLKKIELTLTSQKR--LQYS-----ESFYVDGTVTNTGK-FTIKKCKITVKLYNDKQV 93 (149)
T ss_pred HHHHHHHHHHHHHHhhhcCCceEEEeeeee--eeec-----cEEEEEEEEEECCC-CEeeEEEEEEEEEeCCCc
Confidence 3333445555554444 445665555443 3333 46777899999887 777777788888885543
No 13
>PF06072 Herpes_US9: Alphaherpesvirus tegument protein US9; InterPro: IPR009278 This family consists of several US9 and related proteins from the Alphaherpesviruses. The function of the US9 protein is unknown although in Bovine herpesvirus 5 Us9 is essential for the anterograde spread of the virus from the olfactory mucosa to the bulb [].; GO: 0019033 viral tegument
Probab=60.47 E-value=4.1 Score=27.80 Aligned_cols=8 Identities=13% Similarity=0.330 Sum_probs=3.5
Q ss_pred HHhheeeE
Q 036189 65 LICILMYF 72 (241)
Q Consensus 65 l~~lil~l 72 (241)
+-+++.|+
T Consensus 52 lG~~~~~~ 59 (60)
T PF06072_consen 52 LGALVAWH 59 (60)
T ss_pred HHHHhhcc
Confidence 34444443
No 14
>PF14155 DUF4307: Domain of unknown function (DUF4307)
Probab=60.00 E-value=76 Score=24.25 Aligned_cols=28 Identities=14% Similarity=0.210 Sum_probs=16.7
Q ss_pred EeCCcccceeee--ccCCCceecCCCeEEEEEEEEe
Q 036189 127 LNHDHFPIAINH--SVSPPFKVKPMKKSTIHVQLAT 160 (241)
Q Consensus 127 ~Y~g~~~~~lg~--~~vp~F~q~~~~tt~v~v~l~~ 160 (241)
.|.+.. +|. ..+|+ +...+..+.+++..
T Consensus 71 ~~d~ae---VGrreV~vp~---~~~~~~~~~v~v~T 100 (112)
T PF14155_consen 71 DYDGAE---VGRREVLVPP---SGERTVRVTVTVRT 100 (112)
T ss_pred eCCCCE---EEEEEEEECC---CCCcEEEEEEEEEe
Confidence 345554 773 45677 55556666666665
No 15
>COG1580 FliL Flagellar basal body-associated protein [Cell motility and secretion]
Probab=59.50 E-value=27 Score=28.60 Aligned_cols=22 Identities=23% Similarity=0.240 Sum_probs=13.1
Q ss_pred HHHHHHHHHHHHHHhheeeEEe
Q 036189 53 LLVVAASISLTALICILMYFTL 74 (241)
Q Consensus 53 ~~~~~~~i~llgl~~lil~lv~ 74 (241)
++++++.++++++.+..+|+..
T Consensus 21 ~liv~ivl~~~a~~~~~~~~~~ 42 (159)
T COG1580 21 LLIVLIVLLALAGAGYFFWFGS 42 (159)
T ss_pred HHHHHHHHHHHHHHHHHHhhhc
Confidence 3344434566666777777765
No 16
>PF07787 DUF1625: Protein of unknown function (DUF1625); InterPro: IPR012430 Sequences making up this family are derived from hypothetical proteins expressed by both prokaryotic and eukaryotic species. The region in question is approximately 250 residues long.
Probab=58.16 E-value=10 Score=32.97 Aligned_cols=17 Identities=29% Similarity=0.448 Sum_probs=12.3
Q ss_pred HHHHHHHhheeeEEecC
Q 036189 60 ISLTALICILMYFTLGP 76 (241)
Q Consensus 60 i~llgl~~lil~lv~rP 76 (241)
+.+..+++.+.|+.|||
T Consensus 232 ~~lsl~~Ia~aW~~yRP 248 (248)
T PF07787_consen 232 FSLSLLTIALAWLFYRP 248 (248)
T ss_pred HHHHHHHHHHhheeeCc
Confidence 34445577788999998
No 17
>PRK13183 psbN photosystem II reaction center protein N; Provisional
Probab=56.89 E-value=17 Score=23.52 Aligned_cols=21 Identities=29% Similarity=0.358 Sum_probs=16.9
Q ss_pred HHHHHHHHHHhheeeEEecCC
Q 036189 57 AASISLTALICILMYFTLGPK 77 (241)
Q Consensus 57 ~~~i~llgl~~lil~lv~rPk 77 (241)
++..++++++...+|..|-|.
T Consensus 13 ~i~~lL~~~TgyaiYtaFGpp 33 (46)
T PRK13183 13 TILAILLALTGFGIYTAFGPP 33 (46)
T ss_pred HHHHHHHHHhhheeeeccCCc
Confidence 334578999999999999983
No 18
>COG3671 Predicted membrane protein [Function unknown]
Probab=55.43 E-value=3.9 Score=31.88 Aligned_cols=33 Identities=0% Similarity=0.064 Sum_probs=19.3
Q ss_pred CCchhhhHHHHHHHHHHHHHHHHHh---heeeEEec
Q 036189 43 KKLPKLIIITLLVVAASISLTALIC---ILMYFTLG 75 (241)
Q Consensus 43 ~~~~rc~~~~~~~~~~~i~llgl~~---lil~lv~r 75 (241)
+.+++|+++.++++++.++.+|+++ +-+|.++|
T Consensus 68 RTFw~~vl~~iIg~Llt~lgiGv~i~~AlgvW~i~R 103 (125)
T COG3671 68 RTFWLAVLWWIIGLLLTFLGIGVVILVALGVWYIYR 103 (125)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 4567777777777765555556533 33455554
No 19
>PRK07021 fliL flagellar basal body-associated protein FliL; Reviewed
Probab=53.82 E-value=59 Score=26.36 Aligned_cols=17 Identities=12% Similarity=-0.089 Sum_probs=10.6
Q ss_pred EEEEccEEEEEEeCCcc
Q 036189 116 QIYLDYIECIALNHDHF 132 (241)
Q Consensus 116 ~i~Y~~~~v~v~Y~g~~ 132 (241)
+-+|=..++++.+.+..
T Consensus 77 ~~rylkv~i~L~~~~~~ 93 (162)
T PRK07021 77 ADRVLYVGLTLRLPDEA 93 (162)
T ss_pred CceEEEEEEEEEECCHH
Confidence 35676677777666543
No 20
>PRK05529 cell division protein FtsQ; Provisional
Probab=53.32 E-value=30 Score=30.38 Aligned_cols=43 Identities=14% Similarity=0.094 Sum_probs=27.8
Q ss_pred CCEEEEeeEEEeeeecCC--------------CceeE--------------EEEEEEEEeCCCCeeEEEEcc
Q 036189 78 LPSLHLDTFSVSNFTIGS--------------TNLIA--------------KWDFNLTFKNPDHLWQIYLDY 121 (241)
Q Consensus 78 ~P~f~V~s~~l~~f~~~~--------------~~l~~--------------~~~~~l~v~NPN~k~~i~Y~~ 121 (241)
.|.|.|.++.|++-..-+ +.+.. -=++.++-+.||. +.|.-.+
T Consensus 58 Sp~~~v~~I~V~Gn~~vs~~eI~~~~~~~~g~~l~~vd~~~~~~~l~~~P~V~sa~V~r~~P~t-l~I~V~E 128 (255)
T PRK05529 58 SPLLALRSIEVAGNMRVKPQDIVAALRDQFGKPLPLVDPETVRKKLAAFPLIRSYSVESKPPGT-IVVRVVE 128 (255)
T ss_pred CCceEEEEEEEECCccCCHHHHHHHhcccCCCcceeECHHHHHHHHhcCCCEeEEEEEEeCCCE-EEEEEEE
Confidence 589999999998643221 11111 1257788899997 7777654
No 21
>PF12505 DUF3712: Protein of unknown function (DUF3712); InterPro: IPR022185 This domain family is found in eukaryotes, and is approximately 130 amino acids in length.
Probab=53.20 E-value=1e+02 Score=23.65 Aligned_cols=66 Identities=12% Similarity=-0.053 Sum_probs=38.9
Q ss_pred eeeccCCCceecCCCeE-EEEEEEEeCCceeecCHHHHHHHHHHhhCCceEEEEEEEEEEEEEEEEeceEEEeeeeE
Q 036189 136 INHSVSPPFKVKPMKKS-TIHVQLATGDSLIFLNHQLLQKINSQRRNGRMVVFGLAVRAKTRFTGVSWLWWTEFANL 211 (241)
Q Consensus 136 lg~~~vp~F~q~~~~tt-~v~v~l~~~~~~~~l~~~~~~~l~~d~~~G~~v~~~v~v~~~vr~kv~~g~~~~~~~~~ 211 (241)
+|...+|+..-.+..+. .++..+.. .+.+...++.++.-....+.+.++.+ ...++ |.++.....+
T Consensus 3 f~~~~lP~~~~~~~~~~~~~~~~l~i------~d~~~f~~f~~~~~~~~~~~l~l~g~--~~~~~--g~l~~~~i~~ 69 (125)
T PF12505_consen 3 FATLDLPQIKIKGNGTISIIDQTLTI------TDQDAFTQFVTALLFNEEVTLTLRGK--TDTHL--GGLPFSGIPF 69 (125)
T ss_pred eEEEECCCEEecCCceEEEeeeeEEe------cCHHHHHHHHHHHHhCCcEEEEEEEe--eeEEE--ccEEEEEEee
Confidence 88889999888332222 23233333 45566778887774433255555544 57777 8886544443
No 22
>CHL00020 psbN photosystem II protein N
Probab=52.85 E-value=18 Score=23.02 Aligned_cols=21 Identities=19% Similarity=0.249 Sum_probs=16.8
Q ss_pred HHHHHHHHHHHhheeeEEecC
Q 036189 56 VAASISLTALICILMYFTLGP 76 (241)
Q Consensus 56 ~~~~i~llgl~~lil~lv~rP 76 (241)
+++..++++++...+|..|-|
T Consensus 9 i~i~~ll~~~Tgy~iYtaFGp 29 (43)
T CHL00020 9 IFISGLLVSFTGYALYTAFGQ 29 (43)
T ss_pred HHHHHHHHHhhheeeeeccCC
Confidence 333457889999999999998
No 23
>PF08113 CoxIIa: Cytochrome c oxidase subunit IIa family; InterPro: IPR012538 This family consists of the cytochrome c oxidase subunit IIa family. The bax-type cytochrome c oxidase from Thermus thermophilus is known as a two subunit enzyme. From its crystal structure, it was discovered that an additional transmembrane helix, subunit IIa, spans the membrane. This subunit consists of 34 residues forming one helix across the membrane. The presence of this subunit seems to be important for the function of cytochrome c oxidases [].; PDB: 2QPD_C 3QJR_C 3EH5_C 3BVD_C 3S39_C 3QJU_C 3QJS_C 4EV3_C 3QJT_C 4FA7_C ....
Probab=52.59 E-value=6 Score=23.78 Aligned_cols=13 Identities=8% Similarity=0.508 Sum_probs=8.6
Q ss_pred HHHHHHHhheeeE
Q 036189 60 ISLTALICILMYF 72 (241)
Q Consensus 60 i~llgl~~lil~l 72 (241)
+.++++++|++|+
T Consensus 12 v~iLt~~ILvFWf 24 (34)
T PF08113_consen 12 VMILTAFILVFWF 24 (34)
T ss_dssp HHHHHHHHHHHHH
T ss_pred HHHHHHHHHHHHH
Confidence 3466677777774
No 24
>PF01102 Glycophorin_A: Glycophorin A; InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=50.51 E-value=6.3 Score=30.93 Aligned_cols=25 Identities=16% Similarity=0.251 Sum_probs=9.9
Q ss_pred HHHHHHhheeeEEec-CCCCEEEEee
Q 036189 61 SLTALICILMYFTLG-PKLPSLHLDT 85 (241)
Q Consensus 61 ~llgl~~lil~lv~r-Pk~P~f~V~s 85 (241)
-++|+++||+|++-| =|++...++.
T Consensus 76 GvIg~Illi~y~irR~~Kk~~~~~~p 101 (122)
T PF01102_consen 76 GVIGIILLISYCIRRLRKKSSSDVQP 101 (122)
T ss_dssp HHHHHHHHHHHHHHHHS---------
T ss_pred HHHHHHHHHHHHHHHHhccCCCCCCC
Confidence 445666777777654 3445555544
No 25
>PRK06531 yajC preprotein translocase subunit YajC; Validated
Probab=50.16 E-value=7.3 Score=30.16 Aligned_cols=13 Identities=15% Similarity=0.286 Sum_probs=8.4
Q ss_pred HhheeeEEecCCC
Q 036189 66 ICILMYFTLGPKL 78 (241)
Q Consensus 66 ~~lil~lv~rPk~ 78 (241)
++.++||.+||+.
T Consensus 12 ~~~i~yf~iRPQk 24 (113)
T PRK06531 12 MLGLIFFMQRQQK 24 (113)
T ss_pred HHHHHHheechHH
Confidence 3444567799964
No 26
>PF04478 Mid2: Mid2 like cell wall stress sensor; InterPro: IPR007567 This family represents a region near the C terminus of Mid2, which contains a transmembrane region. The remainder of the protein sequence is serine-rich and of low complexity, and is therefore impossible to align accurately. Mid2 is thought to act as a mechanosensor of cell wall stress. The C-terminal cytoplasmic region of Mid2 is known to interact with Rom2, a guanine nucleotide exchange factor (GEF) for Rho1, which is part of the cell wall integrity signalling pathway [].
Probab=48.23 E-value=3.6 Score=33.47 Aligned_cols=42 Identities=12% Similarity=0.241 Sum_probs=26.1
Q ss_pred HHHHHHhheeeEEecCCCCEEEEeeEEEeeeecCCCceeEEEEEEEEEeCCCCeeEEEE
Q 036189 61 SLTALICILMYFTLGPKLPSLHLDTFSVSNFTIGSTNLIAKWDFNLTFKNPDHLWQIYL 119 (241)
Q Consensus 61 ~llgl~~lil~lv~rPk~P~f~V~s~~l~~f~~~~~~l~~~~~~~l~v~NPN~k~~i~Y 119 (241)
+|+++++++||+..|+|.=.| ++..+ . .+++.++|+--.++|
T Consensus 62 ill~il~lvf~~c~r~kktdf---------idSdG-k-------vvtay~~n~~~~~w~ 103 (154)
T PF04478_consen 62 ILLGILALVFIFCIRRKKTDF---------IDSDG-K-------VVTAYRSNKLTKWWY 103 (154)
T ss_pred HHHHHHHhheeEEEecccCcc---------ccCCC-c-------EEEEEcCchHHHHHH
Confidence 456778888888999986443 22221 1 356777776555555
No 27
>PF02468 PsbN: Photosystem II reaction centre N protein (psbN); InterPro: IPR003398 Oxygenic photosynthesis uses two multi-subunit photosystems (I and II) located in the cell membranes of cyanobacteria and in the thylakoid membranes of chloroplasts in plants and algae. Photosystem II (PSII) has a P680 reaction centre containing chlorophyll 'a' that uses light energy to carry out the oxidation (splitting) of water molecules, and to produce ATP via a proton pump. Photosystem I (PSI) has a P700 reaction centre containing chlorophyll that takes the electron and associated hydrogen donated from PSII to reduce NADP+ to NADPH. Both ATP and NADPH are subsequently used in the light-independent reactions to convert carbon dioxide to glucose using the hydrogen atom extracted from water by PSII, releasing oxygen as a by-product. PSII is a multisubunit protein-pigment complex containing polypeptides both intrinsic and extrinsic to the photosynthetic membrane [, ]. Within the core of the complex, the chlorophyll and beta-carotene pigments are mainly bound to the antenna proteins CP43 (PsbC) and CP47 (PsbB), which pass the excitation energy on to the reaction centre proteins D1 (Qb, PsbA) and D2 (Qa, PsbD) that bind all the redox-active cofactors involved in the energy conversion process. The PSII oxygen-evolving complex (OEC) oxidises water to provide protons for use by PSI, and consists of OEE1 (PsbO), OEE2 (PsbP) and OEE3 (PsbQ). The remaining subunits in PSII are of low molecular weight (less than 10 kDa), and are involved in PSII assembly, stabilisation, dimerisation, and photo-protection []. This family represents the low molecular weight transmembrane protein PsbN found in PSII. PsbN may have a role in PSII stability, however its actual function unknown. PsbN does not appear to be essential for photoautotrophic growth or normal PSII function.; GO: 0015979 photosynthesis, 0009523 photosystem II, 0009539 photosystem II reaction center, 0016020 membrane
Probab=47.57 E-value=16 Score=23.37 Aligned_cols=18 Identities=28% Similarity=0.506 Sum_probs=15.2
Q ss_pred HHHHHHHHhheeeEEecC
Q 036189 59 SISLTALICILMYFTLGP 76 (241)
Q Consensus 59 ~i~llgl~~lil~lv~rP 76 (241)
..++++++...+|..|.|
T Consensus 12 ~~~lv~~Tgy~iYtaFGp 29 (43)
T PF02468_consen 12 SCLLVSITGYAIYTAFGP 29 (43)
T ss_pred HHHHHHHHhhhhhheeCC
Confidence 457888899999999987
No 28
>PF14927 Neurensin: Neurensin
Probab=47.42 E-value=48 Score=26.61 Aligned_cols=12 Identities=8% Similarity=0.263 Sum_probs=7.0
Q ss_pred HHHHHHHhheee
Q 036189 60 ISLTALICILMY 71 (241)
Q Consensus 60 i~llgl~~lil~ 71 (241)
++++|++++++-
T Consensus 54 ~Ll~Gi~~l~vg 65 (140)
T PF14927_consen 54 LLLLGIVALTVG 65 (140)
T ss_pred HHHHHHHHHHhh
Confidence 357777555553
No 29
>KOG0810 consensus SNARE protein Syntaxin 1 and related proteins [Intracellular trafficking, secretion, and vesicular transport]
Probab=44.63 E-value=7.6 Score=35.10 Aligned_cols=14 Identities=36% Similarity=0.214 Sum_probs=5.5
Q ss_pred CCCCCCCCchhhhH
Q 036189 37 YQHFQIKKLPKLII 50 (241)
Q Consensus 37 ~~~~~~~~~~rc~~ 50 (241)
||+..|++-|.|++
T Consensus 263 ~qkkaRK~k~i~ii 276 (297)
T KOG0810|consen 263 YQKKARKWKIIIII 276 (297)
T ss_pred HHHHhhhceeeeeh
Confidence 44443333333333
No 30
>COG4698 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=42.76 E-value=23 Score=29.68 Aligned_cols=30 Identities=27% Similarity=0.497 Sum_probs=20.0
Q ss_pred HhheeeEEecCCCCEEEEeeEEE---eeeecCC
Q 036189 66 ICILMYFTLGPKLPSLHLDTFSV---SNFTIGS 95 (241)
Q Consensus 66 ~~lil~lv~rPk~P~f~V~s~~l---~~f~~~~ 95 (241)
+++++.+++.|+.|..++.+++= ..|.+++
T Consensus 26 ~~~i~~~vlsp~ee~t~~~~a~~~~~~~fqitt 58 (197)
T COG4698 26 AVLIALFVLSPREEPTHLEDASEKSEKSFQITT 58 (197)
T ss_pred HHHhheeeccCCCCCchhhccCcccceeEEEEc
Confidence 36666778999997777766544 3355543
No 31
>PF02009 Rifin_STEVOR: Rifin/stevor family; InterPro: IPR002858 Malaria is still a major cause of mortality in many areas of the world. Plasmodium falciparum causes the most severe human form of the disease and is responsible for most fatalities. Severe cases of malaria can occur when the parasite invades and then proliferates within red blood cell erythrocytes. The parasite produces many variant antigenic proteins, encoded by multigene families, which are present on the surface of the infected erythrocyte and play important roles in virulence. A crucial survival mechanism for the malaria parasite is its ability to evade the immune response by switching these variant surface antigens. The high virulence of P. falciparum relative to other malarial parasites is in large part due to the fact that in this organism many of these surface antigens mediate the binding of infected erythrocytes to the vascular endothelium (cytoadherence) and non-infected erythrocytes (rosetting). This can lead to the accumulation of infected cells in the vasculature of a variety of organs, blocking the blood flow and reducing the oxygen supply. Clinical symptoms of severe infection can include fever, progressive anaemia, multi-organ dysfunction and coma. For more information see []. Several multicopy gene families have been described in Plasmodium falciparum, including the stevor family of subtelomeric open reading frames and the rif interspersed repetitive elements. Both families contain three predicted transmembrane segments. It has been proposed that stevor and rif are members of a larger superfamily that code for variant surface antigens [].
Probab=42.28 E-value=9.6 Score=34.48 Aligned_cols=17 Identities=29% Similarity=0.647 Sum_probs=11.6
Q ss_pred HHHHHHHHhheeeEEec
Q 036189 59 SISLTALICILMYFTLG 75 (241)
Q Consensus 59 ~i~llgl~~lil~lv~r 75 (241)
+|+++.++.+|+||+||
T Consensus 264 aIliIVLIMvIIYLILR 280 (299)
T PF02009_consen 264 AILIIVLIMVIIYLILR 280 (299)
T ss_pred HHHHHHHHHHHHHHHHH
Confidence 34555667788888766
No 32
>PF11322 DUF3124: Protein of unknown function (DUF3124); InterPro: IPR021471 This bacterial family of proteins has no known function.
Probab=41.35 E-value=1.7e+02 Score=23.05 Aligned_cols=53 Identities=15% Similarity=0.295 Sum_probs=34.8
Q ss_pred CceeEEEEEEEEEeCCCCeeEEEEccEEEEEEe--CCcccceeeeccCCCceecCCCeEEE
Q 036189 96 TNLIAKWDFNLTFKNPDHLWQIYLDYIECIALN--HDHFPIAINHSVSPPFKVKPMKKSTI 154 (241)
Q Consensus 96 ~~l~~~~~~~l~v~NPN~k~~i~Y~~~~v~v~Y--~g~~~~~lg~~~vp~F~q~~~~tt~v 154 (241)
.....+|+++|++||.+.+-.|+-.+. -|| +|.. +-..--.|.+.++-.+..+
T Consensus 19 ~~~~~~Lt~tLSiRNtd~~~~i~i~~v---~Yydt~G~l---vr~yl~~Pi~L~Pl~t~~~ 73 (125)
T PF11322_consen 19 KHRPFNLTATLSIRNTDPTDPIYITSV---DYYDTDGKL---VRSYLDKPIYLKPLATTEF 73 (125)
T ss_pred CCceEeEEEEEEEEcCCCCCCEEEEEE---EEECCCCeE---hHHhcCCCeEcCCCceEEE
Confidence 466789999999999888777765433 234 3443 4444445677777777655
No 33
>PF04790 Sarcoglycan_1: Sarcoglycan complex subunit protein; InterPro: IPR006875 The dystrophin glycoprotein complex (DGC) is a membrane-spanning complex that links the interior cytoskeleton to the extracellular matrix in muscle. The sarcoglycan complex is a subcomplex within the DGC and is composed of several muscle-specific, transmembrane proteins (alpha-, beta-, gamma-, delta- and zeta-sarcoglycan). The sarcoglycans are asparagine-linked glycosylated proteins with single transmembrane domains. This family contains beta, gamma and delta members [, ].; GO: 0007010 cytoskeleton organization, 0016012 sarcoglycan complex, 0016021 integral to membrane
Probab=40.95 E-value=2.6e+02 Score=24.81 Aligned_cols=17 Identities=12% Similarity=0.202 Sum_probs=11.7
Q ss_pred eeEEEEEEEEEeCCCCe
Q 036189 98 LIAKWDFNLTFKNPDHL 114 (241)
Q Consensus 98 l~~~~~~~l~v~NPN~k 114 (241)
+..+=++.+.++|.|..
T Consensus 84 i~s~~~v~~~~r~~~g~ 100 (264)
T PF04790_consen 84 IQSSRNVTLNARNENGS 100 (264)
T ss_pred EEecCceEEEEecCCCc
Confidence 44444577778888876
No 34
>PF06637 PV-1: PV-1 protein (PLVAP); InterPro: IPR009538 This family consists of several PV-1 (PLVAP) proteins, which seem to be specific to mammals. PV-1 is a novel protein component of the endothelial fenestral and stomatal diaphragms []. The function of this family is unknown.
Probab=40.72 E-value=36 Score=31.83 Aligned_cols=12 Identities=17% Similarity=0.606 Sum_probs=5.9
Q ss_pred HHHHHHHhheee
Q 036189 60 ISLTALICILMY 71 (241)
Q Consensus 60 i~llgl~~lil~ 71 (241)
+||+|++.+.+|
T Consensus 38 LIIlgLVLFmVY 49 (442)
T PF06637_consen 38 LIILGLVLFMVY 49 (442)
T ss_pred HHHHHHHHHHhh
Confidence 455555444444
No 35
>PHA02844 putative transmembrane protein; Provisional
Probab=39.85 E-value=30 Score=24.68 Aligned_cols=11 Identities=18% Similarity=0.431 Sum_probs=5.3
Q ss_pred HHHHHhheeeE
Q 036189 62 LTALICILMYF 72 (241)
Q Consensus 62 llgl~~lil~l 72 (241)
++.++.+.+||
T Consensus 59 ~~~~~~~flYL 69 (75)
T PHA02844 59 VFATFLTFLYL 69 (75)
T ss_pred HHHHHHHHHHH
Confidence 33344455565
No 36
>PF14283 DUF4366: Domain of unknown function (DUF4366)
Probab=38.28 E-value=32 Score=29.66 Aligned_cols=21 Identities=14% Similarity=0.018 Sum_probs=11.2
Q ss_pred HHHHHHhheeeEEecCCCCEE
Q 036189 61 SLTALICILMYFTLGPKLPSL 81 (241)
Q Consensus 61 ~llgl~~lil~lv~rPk~P~f 81 (241)
+++|..+..+|-++|||....
T Consensus 170 ~l~gGGa~yYfK~~K~K~~~~ 190 (218)
T PF14283_consen 170 ALIGGGAYYYFKFYKPKQEEK 190 (218)
T ss_pred HHhhcceEEEEEEeccccccc
Confidence 334443343334888886543
No 37
>PF13131 DUF3951: Protein of unknown function (DUF3951)
Probab=36.75 E-value=25 Score=23.24 Aligned_cols=31 Identities=19% Similarity=0.310 Sum_probs=17.5
Q ss_pred HHHHHHHHHHHHHHHHHhheeeEE-ecCCCCE
Q 036189 50 IITLLVVAASISLTALICILMYFT-LGPKLPS 80 (241)
Q Consensus 50 ~~~~~~~~~~i~llgl~~lil~lv-~rPk~P~ 80 (241)
+.++.++++.++++.++.++.|-. .+-+.|.
T Consensus 3 L~tiG~~~~~~~I~~lIgfity~mfV~K~s~q 34 (53)
T PF13131_consen 3 LLTIGIILFTIFIFFLIGFITYKMFVKKASPQ 34 (53)
T ss_pred hHHHHHHHHHHHHHHHHHHHHHHhheecCCCc
Confidence 345555555566777777777744 3333343
No 38
>PF11906 DUF3426: Protein of unknown function (DUF3426); InterPro: IPR021834 This family of proteins are functionally uncharacterised. This protein is found in bacteria. Proteins in this family are typically between 262 to 463 amino acids in length.
Probab=36.46 E-value=2e+02 Score=22.44 Aligned_cols=57 Identities=12% Similarity=0.099 Sum_probs=38.4
Q ss_pred EEEeeEEEeeeecCC-CceeEEEEEEEEEeCCCCeeEEEEccEEEEEE-eCCcccceeeeccC
Q 036189 81 LHLDTFSVSNFTIGS-TNLIAKWDFNLTFKNPDHLWQIYLDYIECIAL-NHDHFPIAINHSVS 141 (241)
Q Consensus 81 f~V~s~~l~~f~~~~-~~l~~~~~~~l~v~NPN~k~~i~Y~~~~v~v~-Y~g~~~~~lg~~~v 141 (241)
-.++.+++....+.. +.-.-.+.++.+++|... ....|-.++++++ -+|+. +++-.+
T Consensus 48 ~~~~~l~i~~~~~~~~~~~~~~l~v~g~i~N~~~-~~~~~P~l~l~L~D~~g~~---l~~r~~ 106 (149)
T PF11906_consen 48 RDIDALKIESSDLRPVPDGPGVLVVSGTIRNRAD-FPQALPALELSLLDAQGQP---LARRVF 106 (149)
T ss_pred cCcceEEEeeeeEEeecCCCCEEEEEEEEEeCCC-CcccCceEEEEEECCCCCE---EEEEEE
Confidence 355555555444432 234567888999999987 7888888888887 46665 665444
No 39
>PRK08455 fliL flagellar basal body-associated protein FliL; Reviewed
Probab=36.07 E-value=52 Score=27.47 Aligned_cols=15 Identities=0% Similarity=-0.420 Sum_probs=10.8
Q ss_pred EEccEEEEEEeCCcc
Q 036189 118 YLDYIECIALNHDHF 132 (241)
Q Consensus 118 ~Y~~~~v~v~Y~g~~ 132 (241)
+|=...+.+.+.+..
T Consensus 103 ryLkv~i~Le~~~~~ 117 (182)
T PRK08455 103 RYLKTSISLELSNEK 117 (182)
T ss_pred eEEEEEEEEEECCHh
Confidence 787777777776653
No 40
>PF04573 SPC22: Signal peptidase subunit; InterPro: IPR007653 Translocation of polypeptide chains across the endoplasmic reticulum membrane is triggered by signal sequences. During translocation of the nascent chain through the membrane, the signal sequence of most secretory and membrane proteins is cleaved off. Cleavage occurs by the signal peptidase complex (SPC), which consists of four subunits in yeast and five in mammals. This family is is described as similar to microsomal signal peptidase 23 kDa subunit. Found in eukaryotes [, ].; GO: 0008233 peptidase activity, 0006465 signal peptide processing, 0005787 signal peptidase complex, 0016021 integral to membrane
Probab=35.91 E-value=2.4e+02 Score=23.38 Aligned_cols=32 Identities=13% Similarity=-0.028 Sum_probs=21.1
Q ss_pred ceeEEEEEEEE-EeCCCCeeEEEEccEEEEEEeCCcc
Q 036189 97 NLIAKWDFNLT-FKNPDHLWQIYLDYIECIALNHDHF 132 (241)
Q Consensus 97 ~l~~~~~~~l~-v~NPN~k~~i~Y~~~~v~v~Y~g~~ 132 (241)
.++.+++++++ .-|=|.|.-+-| +.+.|.+..
T Consensus 65 ~i~fdl~aDls~lfnWNtKq~Fvy----v~A~Y~t~~ 97 (175)
T PF04573_consen 65 KITFDLDADLSPLFNWNTKQLFVY----VTAEYETPK 97 (175)
T ss_pred EEEEEeccCcccceeeeeeEEEEE----EEEEECCCC
Confidence 45555555555 478888888777 667776653
No 41
>PF06024 DUF912: Nucleopolyhedrovirus protein of unknown function (DUF912); InterPro: IPR009261 This entry is represented by Autographa californica nuclear polyhedrosis virus (AcMNPV), Orf78; it is a family of uncharacterised viral proteins.
Probab=35.59 E-value=48 Score=24.88 Aligned_cols=14 Identities=21% Similarity=0.534 Sum_probs=6.6
Q ss_pred HHHHhheeeE-EecC
Q 036189 63 TALICILMYF-TLGP 76 (241)
Q Consensus 63 lgl~~lil~l-v~rP 76 (241)
+.++.+|.|+ ++|=
T Consensus 75 lVily~IyYFVILRe 89 (101)
T PF06024_consen 75 LVILYAIYYFVILRE 89 (101)
T ss_pred HHHHhhheEEEEEec
Confidence 3334445555 4553
No 42
>PF10907 DUF2749: Protein of unknown function (DUF2749); InterPro: IPR024475 This bacterial family of proteins represent the TrbJ and TrbK genes of the Ti plasmid conjugative transfer operon [].
Probab=35.35 E-value=51 Score=22.95 Aligned_cols=16 Identities=13% Similarity=0.270 Sum_probs=12.4
Q ss_pred HHHHHhheeeEEecCC
Q 036189 62 LTALICILMYFTLGPK 77 (241)
Q Consensus 62 llgl~~lil~lv~rPk 77 (241)
+.+.+..+.|++.+|+
T Consensus 13 vaa~a~~atwviVq~~ 28 (66)
T PF10907_consen 13 VAAAAGAATWVIVQPR 28 (66)
T ss_pred HHhhhceeEEEEECCC
Confidence 4445778889999998
No 43
>PF03302 VSP: Giardia variant-specific surface protein; InterPro: IPR005127 During infection, the intestinal protozoan parasite Giardia lamblia virus undergoes continuous antigenic variation which is determined by diversification of the parasite's major surface antigen, named VSP (variant surface protein).
Probab=33.75 E-value=14 Score=34.64 Aligned_cols=34 Identities=24% Similarity=0.256 Sum_probs=19.6
Q ss_pred CCchhhhHHHHHHHHHHHHHHHHHhheee-EEecCC
Q 036189 43 KKLPKLIIITLLVVAASISLTALICILMY-FTLGPK 77 (241)
Q Consensus 43 ~~~~rc~~~~~~~~~~~i~llgl~~lil~-lv~rPk 77 (241)
++++-..+. .|.+.++|++.||+.|+.| |+.|=|
T Consensus 362 s~LstgaIa-GIsvavvvvVgglvGfLcWwf~crgk 396 (397)
T PF03302_consen 362 SGLSTGAIA-GISVAVVVVVGGLVGFLCWWFICRGK 396 (397)
T ss_pred cccccccee-eeeehhHHHHHHHHHHHhhheeeccc
Confidence 455556663 3334334567777777766 577644
No 44
>PF09911 DUF2140: Uncharacterized protein conserved in bacteria (DUF2140); InterPro: IPR018672 This family of conserved hypothetical proteins has no known function.
Probab=33.27 E-value=69 Score=26.79 Aligned_cols=20 Identities=15% Similarity=0.331 Sum_probs=14.2
Q ss_pred HHHHHHHhheeeEEecCCCC
Q 036189 60 ISLTALICILMYFTLGPKLP 79 (241)
Q Consensus 60 i~llgl~~lil~lv~rPk~P 79 (241)
.+++++++.+++.+++|..|
T Consensus 12 a~~l~~~~~~~~~~~~~~~~ 31 (187)
T PF09911_consen 12 ALNLAFVIVVFFRLFQPSEP 31 (187)
T ss_pred HHHHHHHhheeeEEEccCCC
Confidence 34555667777788999866
No 45
>PF15012 DUF4519: Domain of unknown function (DUF4519)
Probab=33.25 E-value=36 Score=22.96 Aligned_cols=15 Identities=20% Similarity=0.605 Sum_probs=10.4
Q ss_pred HHHHhheeeEEecCC
Q 036189 63 TALICILMYFTLGPK 77 (241)
Q Consensus 63 lgl~~lil~lv~rPk 77 (241)
+-++++++|+.-||+
T Consensus 42 ~~~Ivv~vy~kTRP~ 56 (56)
T PF15012_consen 42 FLFIVVFVYLKTRPR 56 (56)
T ss_pred HHHHhheeEEeccCC
Confidence 334677888888884
No 46
>PF09865 DUF2092: Predicted periplasmic protein (DUF2092); InterPro: IPR019207 This entry represents various hypothetical prokaryotic proteins of unknown function.
Probab=32.98 E-value=3e+02 Score=23.57 Aligned_cols=36 Identities=14% Similarity=0.002 Sum_probs=29.8
Q ss_pred CceeEEEEEEEEEeCCCCeeEEEEc--cEEEEEEeCCcc
Q 036189 96 TNLIAKWDFNLTFKNPDHLWQIYLD--YIECIALNHDHF 132 (241)
Q Consensus 96 ~~l~~~~~~~l~v~NPN~k~~i~Y~--~~~v~v~Y~g~~ 132 (241)
..+.+.-+.+|.++=||+ +.+.+. ..+..++|.|..
T Consensus 35 qklq~~~~~~v~v~RPdk-lr~~~~gd~~~~~~~yDGkt 72 (214)
T PF09865_consen 35 QKLQFSSSGTVTVQRPDK-LRIDRRGDGADREFYYDGKT 72 (214)
T ss_pred ceEEEEEEEEEEEeCCCe-EEEEEEcCCcceEEEECCCE
Confidence 578888899999999997 988883 356789998886
No 47
>PF06092 DUF943: Enterobacterial putative membrane protein (DUF943); InterPro: IPR010351 This family consists of several hypothetical proteins from Escherichia coli, Yersinia pestis and Salmonella typhi.
Probab=32.94 E-value=28 Score=28.54 Aligned_cols=16 Identities=38% Similarity=0.555 Sum_probs=11.1
Q ss_pred HHHHHHhheeeEEecC
Q 036189 61 SLTALICILMYFTLGP 76 (241)
Q Consensus 61 ~llgl~~lil~lv~rP 76 (241)
+++++++.++|+.+||
T Consensus 13 ~l~~~~~y~~W~~~rp 28 (157)
T PF06092_consen 13 FLLACILYFLWLTLRP 28 (157)
T ss_pred HHHHHHHHhhhhccCC
Confidence 3444444888999999
No 48
>PTZ00116 signal peptidase; Provisional
Probab=32.49 E-value=2.8e+02 Score=23.36 Aligned_cols=52 Identities=12% Similarity=0.041 Sum_probs=31.8
Q ss_pred CCCCEEEEeeEEEeeeecCC------CceeEEEEEEEE-EeCCCCeeEEEEccEEEEEEeCCc
Q 036189 76 PKLPSLHLDTFSVSNFTIGS------TNLIAKWDFNLT-FKNPDHLWQIYLDYIECIALNHDH 131 (241)
Q Consensus 76 Pk~P~f~V~s~~l~~f~~~~------~~l~~~~~~~l~-v~NPN~k~~i~Y~~~~v~v~Y~g~ 131 (241)
...|..+|+=..|.+|...+ ..++.+++++++ .-|=|.|.-|-| +.+.|.+.
T Consensus 36 ~~~~~~~i~v~~V~~~~~~~~~~~D~a~i~fdl~~DL~~lfnWNtKqlFvy----v~a~Y~t~ 94 (185)
T PTZ00116 36 EKEMSTNIKVKSVKRLVYNRHIKGDEAVLSLDLSYDMSKAFNWNLKQLFLY----VLVTYETP 94 (185)
T ss_pred CCCceeeEEEeecccccccCCCCceeEEEEEeeccCchhcCCccccEEEEE----EEEEEcCC
Confidence 34455666555556676432 245566666665 468888888877 66677554
No 49
>PF12505 DUF3712: Protein of unknown function (DUF3712); InterPro: IPR022185 This domain family is found in eukaryotes, and is approximately 130 amino acids in length.
Probab=32.04 E-value=1.1e+02 Score=23.54 Aligned_cols=27 Identities=19% Similarity=0.121 Sum_probs=19.7
Q ss_pred eeEEEEEEEEEeCCCCeeEEEEccEEEE
Q 036189 98 LIAKWDFNLTFKNPDHLWQIYLDYIECI 125 (241)
Q Consensus 98 l~~~~~~~l~v~NPN~k~~i~Y~~~~v~ 125 (241)
-.+++.+++++.||.. +++..+++...
T Consensus 98 ~g~~~~~~~~l~NPS~-~ti~lG~v~~~ 124 (125)
T PF12505_consen 98 DGINLNATVTLPNPSP-LTIDLGNVTLN 124 (125)
T ss_pred CcEEEEEEEEEcCCCe-EEEEeccEEEe
Confidence 3567888888999987 77766665543
No 50
>PHA02650 hypothetical protein; Provisional
Probab=31.90 E-value=30 Score=25.00 Aligned_cols=11 Identities=9% Similarity=0.329 Sum_probs=5.2
Q ss_pred HHHHHhheeeE
Q 036189 62 LTALICILMYF 72 (241)
Q Consensus 62 llgl~~lil~l 72 (241)
++.++.+.+||
T Consensus 60 ~i~~l~~flYL 70 (81)
T PHA02650 60 IIVALFSFFVF 70 (81)
T ss_pred HHHHHHHHHHH
Confidence 33344455555
No 51
>PF05478 Prominin: Prominin; InterPro: IPR008795 The prominins are an emerging family of proteins that, among the multispan membrane proteins, display a novel topology. Mouse and Homo sapiens prominin and (Mus musculus) prominin-like 1 (PROML1) are predicted to contain five membrane spanning domains, with an N-terminal domain exposed to the extracellular space followed by four, alternating small cytoplasmic and large extracellular, loops and a cytoplasmic C-terminal domain []. The exact function of prominin is unknown although in humans defects in PROM1, the gene coding for prominin, cause retinal degeneration [].; GO: 0016021 integral to membrane
Probab=31.86 E-value=44 Score=34.29 Aligned_cols=27 Identities=22% Similarity=0.273 Sum_probs=15.6
Q ss_pred CCCCchhhhHHHHHHHHHHHHHHHHHh
Q 036189 41 QIKKLPKLIIITLLVVAASISLTALIC 67 (241)
Q Consensus 41 ~~~~~~rc~~~~~~~~~~~i~llgl~~ 67 (241)
++..|.|+++..+++++++++++|+++
T Consensus 133 ~~~~c~R~~l~~~L~~~~~~il~g~i~ 159 (806)
T PF05478_consen 133 KNDACRRGCLGILLLLLTLIILFGVIC 159 (806)
T ss_pred cccccchHHHHHHHHHHHHHHHHHHHH
Confidence 445566666655555555566666654
No 52
>PF11395 DUF2873: Protein of unknown function (DUF2873); InterPro: IPR021532 This entry is represented by the human SARS coronavirus, Orf7b; it is a family of uncharacterised viral proteins.
Probab=31.22 E-value=22 Score=21.94 Aligned_cols=9 Identities=33% Similarity=0.962 Sum_probs=4.6
Q ss_pred HhheeeEEe
Q 036189 66 ICILMYFTL 74 (241)
Q Consensus 66 ~~lil~lv~ 74 (241)
...|||+++
T Consensus 24 mliif~f~l 32 (43)
T PF11395_consen 24 MLIIFWFSL 32 (43)
T ss_pred HHHHHHHHH
Confidence 444556554
No 53
>PF15145 DUF4577: Domain of unknown function (DUF4577)
Probab=30.45 E-value=46 Score=25.72 Aligned_cols=26 Identities=19% Similarity=0.501 Sum_probs=15.2
Q ss_pred hHHHHHHHHHHHHHHHHHhheeeEEecC
Q 036189 49 IIITLLVVAASISLTALICILMYFTLGP 76 (241)
Q Consensus 49 ~~~~~~~~~~~i~llgl~~lil~lv~rP 76 (241)
+++.+++++ ++-++++.+++||+++-
T Consensus 63 ffvglii~L--ivSLaLVsFvIFLiiQT 88 (128)
T PF15145_consen 63 FFVGLIIVL--IVSLALVSFVIFLIIQT 88 (128)
T ss_pred hHHHHHHHH--HHHHHHHHHHHHheeec
Confidence 333444443 56666677777777764
No 54
>PF02038 ATP1G1_PLM_MAT8: ATP1G1/PLM/MAT8 family; InterPro: IPR000272 The FXYD protein family contains at least seven members in mammals []. Two other family members that are not obvious orthologs of any identified mammalian FXYD protein exist in zebrafish. All these proteins share a signature sequence of six conserved amino acids comprising the FXYD motif in the NH2-terminus, and two glycines and one serine residue in the transmembrane domain. FXYD proteins are widely distributed in mammalian tissues with prominent expression in tissues that perform fluid and solute transport or that are electrically excitable. Initial functional characterisation suggested that FXYD proteins act as channels or as modulators of ion channels however studies have revealed that most FXYD proteins have another specific function and act as tissue-specific regulatory subunits of the Na,K-ATPase. Each of these auxiliary subunits produces a distinct functional effect on the transport characteristics of the Na,K-ATPase that is adjusted to the specific functional demands of the tissue in which the FXYD protein is expressed. FXYD proteins appear to preferentially associate with Na,K-ATPase alpha1-beta isozymes, and affect their function in a way that render them operationally complementary or supplementary to coexisting isozymes.; GO: 0005216 ion channel activity, 0006811 ion transport, 0016020 membrane; PDB: 2JO1_A 2JP3_A 2ZXE_G 3A3Y_G 3N23_E 3B8E_H 3KDP_G 3N2F_E.
Probab=30.15 E-value=71 Score=21.05 Aligned_cols=16 Identities=19% Similarity=0.318 Sum_probs=8.5
Q ss_pred HHHHHHHHHHHHHHhh
Q 036189 53 LLVVAASISLTALICI 68 (241)
Q Consensus 53 ~~~~~~~i~llgl~~l 68 (241)
.+++.++++++||+++
T Consensus 18 GLi~A~vlfi~Gi~ii 33 (50)
T PF02038_consen 18 GLIFAGVLFILGILII 33 (50)
T ss_dssp HHHHHHHHHHHHHHHH
T ss_pred chHHHHHHHHHHHHHH
Confidence 3444444566666544
No 55
>TIGR01477 RIFIN variant surface antigen, rifin family. This model represents the rifin branch of the rifin/stevor family (pfam02009) of predicted variant surface antigens as found in Plasmodium falciparum. This model is based on a set of rifin sequences kindly provided by Matt Berriman from the Sanger Center. This is a global model and assesses a penalty for incomplete sequence. Additional fragmentary sequences may be found with the fragment model and a cutoff of 20 bits.
Probab=29.90 E-value=20 Score=33.10 Aligned_cols=24 Identities=21% Similarity=0.459 Sum_probs=15.1
Q ss_pred HHHHHHHHHHHHHHhheeeEEecC
Q 036189 53 LLVVAASISLTALICILMYFTLGP 76 (241)
Q Consensus 53 ~~~~~~~i~llgl~~lil~lv~rP 76 (241)
++.-+++|+++.++.+|+||+||=
T Consensus 312 IiaSiIAIvvIVLIMvIIYLILRY 335 (353)
T TIGR01477 312 IIASIIAILIIVLIMVIIYLILRY 335 (353)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHh
Confidence 333333455666678888988764
No 56
>PTZ00046 rifin; Provisional
Probab=28.86 E-value=22 Score=33.00 Aligned_cols=24 Identities=21% Similarity=0.473 Sum_probs=15.1
Q ss_pred HHHHHHHHHHHHHHhheeeEEecC
Q 036189 53 LLVVAASISLTALICILMYFTLGP 76 (241)
Q Consensus 53 ~~~~~~~i~llgl~~lil~lv~rP 76 (241)
++.-++.|+++.++.+|+||+||=
T Consensus 317 IiaSiiAIvVIVLIMvIIYLILRY 340 (358)
T PTZ00046 317 IIASIVAIVVIVLIMVIIYLILRY 340 (358)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHh
Confidence 333333455666678888998774
No 57
>PRK05696 fliL flagellar basal body-associated protein FliL; Reviewed
Probab=27.80 E-value=1.5e+02 Score=24.25 Aligned_cols=17 Identities=12% Similarity=-0.115 Sum_probs=12.2
Q ss_pred EEEEccEEEEEEeCCcc
Q 036189 116 QIYLDYIECIALNHDHF 132 (241)
Q Consensus 116 ~i~Y~~~~v~v~Y~g~~ 132 (241)
+-+|=...+++.+++..
T Consensus 85 ~~ryLkv~i~l~~~d~~ 101 (170)
T PRK05696 85 RDRLVQIKVQLMVRGSD 101 (170)
T ss_pred CceEEEEEEEEEECCHH
Confidence 36787788888777654
No 58
>PF04505 Dispanin: Interferon-induced transmembrane protein; InterPro: IPR007593 This family includes the human leukocyte antigen CD225, which is an interferon inducible transmembrane protein, and is associated with interferon induced cell growth suppression [].; GO: 0009607 response to biotic stimulus, 0016021 integral to membrane
Probab=27.67 E-value=1.3e+02 Score=21.59 Aligned_cols=8 Identities=13% Similarity=0.467 Sum_probs=4.0
Q ss_pred HHHHhhee
Q 036189 63 TALICILM 70 (241)
Q Consensus 63 lgl~~lil 70 (241)
+|++++++
T Consensus 33 lGi~Ai~~ 40 (82)
T PF04505_consen 33 LGIVAIVY 40 (82)
T ss_pred HHHHHhee
Confidence 55555543
No 59
>PRK12785 fliL flagellar basal body-associated protein FliL; Reviewed
Probab=27.66 E-value=1.5e+02 Score=24.19 Aligned_cols=16 Identities=6% Similarity=-0.027 Sum_probs=10.6
Q ss_pred EEEccEEEEEEeCCcc
Q 036189 117 IYLDYIECIALNHDHF 132 (241)
Q Consensus 117 i~Y~~~~v~v~Y~g~~ 132 (241)
.+|=.+.+.+.+.+..
T Consensus 86 ~ryLkv~i~L~~~~~~ 101 (166)
T PRK12785 86 VQYLKLKVVLEVKDEK 101 (166)
T ss_pred ceEEEEEEEEEECCHH
Confidence 4677777777776653
No 60
>COG5009 MrcA Membrane carboxypeptidase/penicillin-binding protein [Cell envelope biogenesis, outer membrane]
Probab=27.64 E-value=29 Score=35.29 Aligned_cols=31 Identities=26% Similarity=0.360 Sum_probs=18.6
Q ss_pred HHHHHHHHHHHHHHHhheeeEEecCCCCEEE
Q 036189 52 TLLVVAASISLTALICILMYFTLGPKLPSLH 82 (241)
Q Consensus 52 ~~~~~~~~i~llgl~~lil~lv~rPk~P~f~ 82 (241)
+++++++++++.+.+++++++.+.|+.|.+.
T Consensus 8 ~l~i~~~~~l~g~~~~~~~~~~~~~dLPd~~ 38 (797)
T COG5009 8 LLGILVTLILLGAGALAGLYLYISPDLPDVE 38 (797)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHhccCCChH
Confidence 3333333334444466777778889988765
No 61
>PHA03093 EEV glycoprotein; Provisional
Probab=27.02 E-value=40 Score=28.27 Aligned_cols=21 Identities=14% Similarity=-0.088 Sum_probs=10.8
Q ss_pred eCCCCeeEEEEccEEEEEEeCC
Q 036189 109 KNPDHLWQIYLDYIECIALNHD 130 (241)
Q Consensus 109 ~NPN~k~~i~Y~~~~v~v~Y~g 130 (241)
.|-+= -||.|+.--..+.++.
T Consensus 97 ~~~~C-~GI~~~~~C~~~~~ep 117 (185)
T PHA03093 97 HKESC-KGIVYDGSCYIFHSEP 117 (185)
T ss_pred ccCcC-CCeecCCEeEEecCCC
Confidence 34443 4677775544444433
No 62
>PF10614 CsgF: Type VIII secretion system (T8SS), CsgF protein; InterPro: IPR018893 Fimbriae are cell-surface protein polymers, of e.g. Escherichia coli and Salmonella spp, that mediate interactions important for host and environmental persistence, development of biofilms, motility, colonisation and invasion of cells, and conjugation. Four general assembly pathways for different fimbriae have been proposed, one of which is extracellular nucleation-precipitation (ENP), that differs from the others in that fibre-growth occurs extracellularly. Thin aggregative fimbriae (Tafi) are the only fimbriae dependent on the ENP pathway. Tafi were first identified in Salmonella spp. and the controlling operon termed agf; however subsequent isolation of the homologous operon in E. coli led to its being called csg. Tafi are known as curli because, in the absence of extracellular polysaccharides, their morphology appears curled; however, when expressed with such polysaccharides their morphology appears as a tangled amorphous matrix []. CsgF is one of three putative curli assembly factors appearing to act as a nucleator protein. Unlike eukaryotic amyloid formation, curli biogenesis is a productive pathway requiring a specific assembly machinery [].
Probab=26.93 E-value=44 Score=26.91 Aligned_cols=11 Identities=27% Similarity=0.392 Sum_probs=9.4
Q ss_pred eEEecCCCCEE
Q 036189 71 YFTLGPKLPSL 81 (241)
Q Consensus 71 ~lv~rPk~P~f 81 (241)
=|||+|..|.|
T Consensus 23 eLVY~PvNPsF 33 (142)
T PF10614_consen 23 ELVYTPVNPSF 33 (142)
T ss_pred heEeeccCCCC
Confidence 38999999976
No 63
>PF13396 PLDc_N: Phospholipase_D-nuclease N-terminal
Probab=24.82 E-value=84 Score=19.53 Aligned_cols=16 Identities=25% Similarity=0.584 Sum_probs=11.8
Q ss_pred HHHHHhheeeEEecCC
Q 036189 62 LTALICILMYFTLGPK 77 (241)
Q Consensus 62 llgl~~lil~lv~rPk 77 (241)
++-++..++|++++.|
T Consensus 31 ~~P~iG~i~Yl~~gr~ 46 (46)
T PF13396_consen 31 FFPIIGPILYLIFGRK 46 (46)
T ss_pred HHHHHHHhheEEEeCC
Confidence 4566788889888764
No 64
>PHA02673 ORF109 EEV glycoprotein; Provisional
Probab=24.37 E-value=28 Score=28.53 Aligned_cols=8 Identities=13% Similarity=-0.338 Sum_probs=4.7
Q ss_pred eEEEEccE
Q 036189 115 WQIYLDYI 122 (241)
Q Consensus 115 ~~i~Y~~~ 122 (241)
-||+|+.-
T Consensus 78 ~GI~~~~~ 85 (161)
T PHA02673 78 DGINAGNK 85 (161)
T ss_pred CCcccCCe
Confidence 45666654
No 65
>PF15050 SCIMP: SCIMP protein
Probab=23.67 E-value=31 Score=26.99 Aligned_cols=10 Identities=10% Similarity=0.458 Sum_probs=6.1
Q ss_pred HhheeeEEec
Q 036189 66 ICILMYFTLG 75 (241)
Q Consensus 66 ~~lil~lv~r 75 (241)
+.||+|.++|
T Consensus 23 lglIlyCvcR 32 (133)
T PF15050_consen 23 LGLILYCVCR 32 (133)
T ss_pred HHHHHHHHHH
Confidence 5666665555
No 66
>PF04790 Sarcoglycan_1: Sarcoglycan complex subunit protein; InterPro: IPR006875 The dystrophin glycoprotein complex (DGC) is a membrane-spanning complex that links the interior cytoskeleton to the extracellular matrix in muscle. The sarcoglycan complex is a subcomplex within the DGC and is composed of several muscle-specific, transmembrane proteins (alpha-, beta-, gamma-, delta- and zeta-sarcoglycan). The sarcoglycans are asparagine-linked glycosylated proteins with single transmembrane domains. This family contains beta, gamma and delta members [, ].; GO: 0007010 cytoskeleton organization, 0016012 sarcoglycan complex, 0016021 integral to membrane
Probab=23.64 E-value=95 Score=27.56 Aligned_cols=11 Identities=0% Similarity=0.145 Sum_probs=5.8
Q ss_pred EEEEecceEEee
Q 036189 211 LMYTCLDLKVGF 222 (241)
Q Consensus 211 ~~v~C~~l~V~~ 222 (241)
+.+.| .--+.+
T Consensus 199 I~~~a-~~di~L 209 (264)
T PF04790_consen 199 IEASA-RQDISL 209 (264)
T ss_pred EEEEe-cCCEEE
Confidence 56666 444444
No 67
>PF06835 LptC: Lipopolysaccharide-assembly, LptC-related; InterPro: IPR010664 This family consists of several related groups of proteins one of which is the LptC family. LptC is involved in lipopolysaccharide-assembly on the outer membrane of Gram-negative organisms. The cell envelope of Gram-negative bacteria consists of an inner (IM) and an outer membrane (OM) separated by an aqueous compartment, the periplasm, which contains the peptidoglycan layer. The OM is an asymmetric bilayer, with phospholipids in the inner leaflet and lipopolysaccharides (LPS) facing outward [, ]. The OM is an effective permeability barrier that protects the cells from toxic compounds, such as antibiotics and detergents, thus allowing bacteria to inhabit several different and often hostile environments. LPS is responsible for the permeability properties of the OM. LPS consists of the lipid A moiety (a glucosamine-based phospholipid) linked to the short core oligosaccharide and the distal O-antigen polysaccharide chain. The core oligosaccharide can be further divided into an inner core, composed of 3-deoxy-D-mannooctulosanate (KDO) and heptose, and an outer core, which has a somewhat variable structure. LPS is essential in most Gram-negative bacteria, with the notable exception of Neisseria meningitidis. The biogenesis of the OM implies that the individual components are transported from the site of synthesis to their final destination outside the IM by crossing both hydrophilic and hydrophobic compartments. The machinery and the energy source that drive this process are not yet fully understood. The lipid A-core moiety and the O-antigen repeat units are synthesized at the cytoplasmic face of the IM and are separately exported via two independent transport systems, namely, the O-antigen transporter Wzx (RfbX) [, ] and the ATP binding cassette (ABC) transporter MsbA that flips the lipid A-core moiety from the inner leaflet to the outer leaflet of the IM [, , ]. O-antigen repeat units are then polymerised in the periplasm by the Wzy polymerase and ligated to the lipid A-core moiety by the WaaL ligase [see, , ]. The LPS transport machinery is composed of LptA, LptB, LptC, LptD, LptE. This supported by the fact, that depletion of any of one of these proteins blocks the LPS assembly pathway and results in very similar OM biogenesis defects. Moreover, the location of at least one of these five proteins in every cellular compartment suggests a model for how the LPS assembly pathway is organised and ordered in space []. Required for the translocation of lipopolysaccharide (LPS) from the inner membrane to the outer membrane [].; PDB: 3MY2_A.
Probab=23.45 E-value=78 Score=24.96 Aligned_cols=51 Identities=16% Similarity=0.214 Sum_probs=12.7
Q ss_pred CCEEEEeeEEEeeeecCCCceeEEEEEEEEEeCCCCeeEEEEccEEEEEEeCC
Q 036189 78 LPSLHLDTFSVSNFTIGSTNLIAKWDFNLTFKNPDHLWQIYLDYIECIALNHD 130 (241)
Q Consensus 78 ~P~f~V~s~~l~~f~~~~~~l~~~~~~~l~v~NPN~k~~i~Y~~~~v~v~Y~g 130 (241)
.|.+.++++++..++-.+ .+...++..=..+++|.+. ++.+...+..+-.+
T Consensus 32 ~~~~~~~~~~~~~~~~~G-~~~~~l~A~~~~~~~~~~~-~~l~~p~~~~~~~~ 82 (176)
T PF06835_consen 32 DPDYSIENFTLTQYDEDG-KLQWKLTAERAEHYPNSDT-VELEDPSLIIYDDD 82 (176)
T ss_dssp ----------------------EEEE-SSEEEETTTTE-EEEES-EEEEE-TT
T ss_pred CCcEEEEeeEEEEECCCC-CEEEEEEEeEEEEecCCCc-EEEeccEEEEEeCC
Confidence 345555555555443321 3444444443346665532 44445555444443
No 68
>PF05545 FixQ: Cbb3-type cytochrome oxidase component FixQ; InterPro: IPR008621 This family consists of several Cbb3-type cytochrome oxidase components (FixQ/CcoQ). FixQ is found in nitrogen fixing bacteria. Since nitrogen fixation is an energy-consuming process, effective symbioses depend on operation of a respiratory chain with a high affinity for O2, closely coupled to ATP production. This requirement is fulfilled by a special three-subunit terminal oxidase (cytochrome terminal oxidase cbb3), which was first identified in Bradyrhizobium japonicum as the product of the fixNOQP operon [].
Probab=23.20 E-value=48 Score=21.24 Aligned_cols=13 Identities=8% Similarity=0.248 Sum_probs=7.1
Q ss_pred HhheeeEEecCCC
Q 036189 66 ICILMYFTLGPKL 78 (241)
Q Consensus 66 ~~lil~lv~rPk~ 78 (241)
.+.++|.+|+|+.
T Consensus 22 F~gi~~w~~~~~~ 34 (49)
T PF05545_consen 22 FIGIVIWAYRPRN 34 (49)
T ss_pred HHHHHHHHHcccc
Confidence 3444444678863
No 69
>COG4736 CcoQ Cbb3-type cytochrome oxidase, subunit 3 [Posttranslational modification, protein turnover, chaperones]
Probab=23.00 E-value=49 Score=22.68 Aligned_cols=13 Identities=23% Similarity=0.570 Sum_probs=8.0
Q ss_pred HhheeeEEecCCC
Q 036189 66 ICILMYFTLGPKL 78 (241)
Q Consensus 66 ~~lil~lv~rPk~ 78 (241)
.+.++|.+|||+.
T Consensus 22 fiavi~~ayr~~~ 34 (60)
T COG4736 22 FIAVIYFAYRPGK 34 (60)
T ss_pred HHHHHHHHhcccc
Confidence 3445566788863
No 70
>PF01034 Syndecan: Syndecan domain; InterPro: IPR001050 The syndecans are transmembrane proteoglycans which are involved in the organisation of cytoskeleton and/or actin microfilaments, and have important roles as cell surface receptors during cell-cell and/or cell-matrix interactions [, ]. Structurally, these proteins consist of four separate domains: A signal sequence; An extracellular domain (ectodomain) of variable length whose sequence is not evolutionary conserved in the various forms of syndecans. The ectodomain contains the sites of attachment of the heparan sulphate glycosaminoglycan side chains; A transmembrane region; A highly conserved cytoplasmic domain of about 30 to 35 residues, which could interact with cytoskeletal proteins. The proteins known to belong to this family are: Syndecan 1. Syndecan 2 or fibroglycan. Syndecan 3 or neuroglycan or N-syndecan. Syndecan 4 or amphiglycan or ryudocan. Drosophila syndecan. Caenorhabditis elegans probable syndecan (F57C7.3). Syndecan-4, a transmembrane heparan sulphate proteoglycan, is a coreceptor with integrins in cell adhesion. It has been suggested to form a ternary signalling complex with protein kinase Calpha and phosphatidylinositol 4,5-bisphosphate (PIP2). Structural studies have demonstrated that the cytoplasmic domain undergoes a conformational transition and forms a symmetric dimer in the presence of phospholipid activator PIP2, and whose overall structure in solution exhibits a twisted clamp shape having a cavity in the centre of dimeric interface. In addition, it has been observed that the syndecan-4 variable domain interacts, strongly, not only with fatty acyl groups but also the anionic head group of PIP2. These findings indicate that PIP2 promotes oligomerisation of the syndecan-4 cytoplasmic domain for transmembrane signalling and cell-matrix adhesion [, ].; GO: 0008092 cytoskeletal protein binding, 0016020 membrane; PDB: 1EJQ_B 1EJP_B 1YBO_C 1OBY_Q.
Probab=22.94 E-value=28 Score=24.17 Aligned_cols=15 Identities=13% Similarity=0.295 Sum_probs=0.0
Q ss_pred HHHHHhheeeEEecC
Q 036189 62 LTALICILMYFTLGP 76 (241)
Q Consensus 62 llgl~~lil~lv~rP 76 (241)
+++++++|++++||=
T Consensus 22 ll~ailLIlf~iyR~ 36 (64)
T PF01034_consen 22 LLFAILLILFLIYRM 36 (64)
T ss_dssp ---------------
T ss_pred HHHHHHHHHHHHHHH
Confidence 444456666777764
No 71
>COG1589 FtsQ Cell division septal protein [Cell envelope biogenesis, outer membrane]
Probab=22.90 E-value=1.2e+02 Score=26.68 Aligned_cols=30 Identities=23% Similarity=0.383 Sum_probs=24.6
Q ss_pred HHHHHHhheeeEEecCCCCEEEEeeEEEee
Q 036189 61 SLTALICILMYFTLGPKLPSLHLDTFSVSN 90 (241)
Q Consensus 61 ~llgl~~lil~lv~rPk~P~f~V~s~~l~~ 90 (241)
+++++.++++|...-+..|.|.+..+.+++
T Consensus 40 ~~~~~~~~~~~~~~~~~~~~~~i~~v~v~G 69 (269)
T COG1589 40 VLLLLVLVVLWVLILLSLPYFPIRKVSVSG 69 (269)
T ss_pred HHHHHHHHHHheehhhhcCCccceEEEEec
Confidence 455567778888888999999999999986
No 72
>PF12321 DUF3634: Protein of unknown function (DUF3634); InterPro: IPR022090 This family of proteins is found in bacteria. Proteins in this family are typically between 103 and 114 amino acids in length.
Probab=22.77 E-value=32 Score=26.44 Aligned_cols=17 Identities=12% Similarity=0.444 Sum_probs=9.4
Q ss_pred hheeeEE--ecCCCCEEEE
Q 036189 67 CILMYFT--LGPKLPSLHL 83 (241)
Q Consensus 67 ~lil~lv--~rPk~P~f~V 83 (241)
++|+||+ .|-..|.|.|
T Consensus 10 ~li~~Lv~~~r~~~~vf~i 28 (108)
T PF12321_consen 10 ALIFWLVFVDRRGLPVFEI 28 (108)
T ss_pred HHHHHHHHccccCceEEEE
Confidence 3777765 3433466654
No 73
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=22.76 E-value=55 Score=24.38 Aligned_cols=15 Identities=20% Similarity=0.401 Sum_probs=7.3
Q ss_pred HHHHHHhheee-EEec
Q 036189 61 SLTALICILMY-FTLG 75 (241)
Q Consensus 61 ~llgl~~lil~-lv~r 75 (241)
++.+|+.+++| +++|
T Consensus 78 ~v~~lv~~l~w~f~~r 93 (96)
T PTZ00382 78 VVGGLVGFLCWWFVCR 93 (96)
T ss_pred HHHHHHHHHhheeEEe
Confidence 34444444555 4555
No 74
>PRK14759 potassium-transporting ATPase subunit F; Provisional
Probab=22.75 E-value=41 Score=19.59 Aligned_cols=19 Identities=26% Similarity=0.308 Sum_probs=10.7
Q ss_pred HHHHHHHhheeeEEecCCC
Q 036189 60 ISLTALICILMYFTLGPKL 78 (241)
Q Consensus 60 i~llgl~~lil~lv~rPk~ 78 (241)
++.+|+.+-.+|-++||.+
T Consensus 10 ~va~~L~vYL~~ALlrPEr 28 (29)
T PRK14759 10 AVSLGLLIYLTYALLRPER 28 (29)
T ss_pred HHHHHHHHHHHHHHhCccc
Confidence 3444555555555688853
No 75
>PF08693 SKG6: Transmembrane alpha-helix domain; InterPro: IPR014805 SKG6 and AXL2 are membrane proteins that show polarised intracellular localisation [, ]. This entry represents the highly conserved transmembrane alpha-helical domain found in these proteins [, ]. The full-length AXL2 protein has a negative regulatory function in cytokinesis [].
Probab=22.26 E-value=74 Score=19.97 Aligned_cols=11 Identities=9% Similarity=0.350 Sum_probs=5.7
Q ss_pred HhheeeEEecC
Q 036189 66 ICILMYFTLGP 76 (241)
Q Consensus 66 ~~lil~lv~rP 76 (241)
++++||+++|-
T Consensus 28 l~~~l~~~~rR 38 (40)
T PF08693_consen 28 LGAFLFFWYRR 38 (40)
T ss_pred HHHHhheEEec
Confidence 34455555664
No 76
>PF13473 Cupredoxin_1: Cupredoxin-like domain; PDB: 1IBZ_D 1IC0_E 1IBY_D.
Probab=22.14 E-value=66 Score=23.66 Aligned_cols=35 Identities=26% Similarity=0.384 Sum_probs=11.6
Q ss_pred EEEEeeEEEeeeecCCCceeEE--EEEEEEEeCCCCe
Q 036189 80 SLHLDTFSVSNFTIGSTNLIAK--WDFNLTFKNPDHL 114 (241)
Q Consensus 80 ~f~V~s~~l~~f~~~~~~l~~~--~~~~l~v~NPN~k 114 (241)
.-....++++++.++++.+... =.++++++|.+..
T Consensus 19 ~~~~v~I~~~~~~f~P~~i~v~~G~~v~l~~~N~~~~ 55 (104)
T PF13473_consen 19 AAQTVTITVTDFGFSPSTITVKAGQPVTLTFTNNDSR 55 (104)
T ss_dssp -----------EEEES-EEEEETTCEEEEEEEE-SSS
T ss_pred ccccccccccCCeEecCEEEEcCCCeEEEEEEECCCC
Confidence 3334455555666554433332 2456777776653
No 77
>PHA03049 IMV membrane protein; Provisional
Probab=21.87 E-value=36 Score=23.77 Aligned_cols=17 Identities=18% Similarity=0.288 Sum_probs=10.8
Q ss_pred HHHHHHHhheeeEEecC
Q 036189 60 ISLTALICILMYFTLGP 76 (241)
Q Consensus 60 i~llgl~~lil~lv~rP 76 (241)
++.+.+++||+|-+|+-
T Consensus 9 iICVaIi~lIvYgiYnk 25 (68)
T PHA03049 9 IICVVIIGLIVYGIYNK 25 (68)
T ss_pred HHHHHHHHHHHHHHHhc
Confidence 34455577777777764
No 78
>PF13800 Sigma_reg_N: Sigma factor regulator N-terminal
Probab=21.67 E-value=33 Score=25.18 Aligned_cols=12 Identities=8% Similarity=0.277 Sum_probs=6.0
Q ss_pred EEEEEeCCCCee
Q 036189 104 FNLTFKNPDHLW 115 (241)
Q Consensus 104 ~~l~v~NPN~k~ 115 (241)
..+.+.-||-.+
T Consensus 56 ~~~~it~PN~~~ 67 (96)
T PF13800_consen 56 LAIEITYPNIYI 67 (96)
T ss_pred HHHHhcCCCEeE
Confidence 344455566433
No 79
>PF05961 Chordopox_A13L: Chordopoxvirus A13L protein; InterPro: IPR009236 This family consists of A13L proteins from the Chordopoxviruses. A13L or p8 is one of the three most abundant membrane proteins of the intracellular mature Vaccinia virus [].
Probab=21.40 E-value=40 Score=23.62 Aligned_cols=17 Identities=24% Similarity=0.350 Sum_probs=10.3
Q ss_pred HHHHHHHhheeeEEecC
Q 036189 60 ISLTALICILMYFTLGP 76 (241)
Q Consensus 60 i~llgl~~lil~lv~rP 76 (241)
++.+.++++|+|-+|+-
T Consensus 9 ~ICVaii~lIlY~iYnr 25 (68)
T PF05961_consen 9 IICVAIIGLILYGIYNR 25 (68)
T ss_pred HHHHHHHHHHHHHHHhc
Confidence 34445567777766654
No 80
>PF00927 Transglut_C: Transglutaminase family, C-terminal ig like domain; InterPro: IPR008958 Synonym(s): Protein-glutamine gamma-glutamyltransferase, Fibrinoligase, TGase Transglutaminases catalyse the post-translational modification of proteins at glutamine residues, with formation of isopeptide bonds. Members of the transglutaminase family usually have three domains: N-terminal (IPR001102 from INTERPRO), middle (IPR013808 from INTERPRO) and C-terminal. The middle domain is usually well conserved, but family members can display major differences in their N- and C-terminal domains, although their overall structure is conserved []. This entry represents the C-terminal domain found in transglutaminases, which consists of an immunoglobulin-like beta-sandwich consisting of seven strands in two sheets with a Greek key topology. The best known transglutaminase is blood coagulation factor XIII, a plasma tetrameric protein composed of two catalytic A subunits and two non-catalytic B subunits. Factor XIII is responsible for cross-linking fibrin chains, thus stabilising the fibrin clot. Protein-glutamine gamma-glutamyltransferases (2.3.2.13 from EC) are calcium-dependent enzymes that catalyse the cross-linking of proteins by promoting the formation of isopeptide bonds between the gamma-carboxyl group of a glutamine in one polypeptide chain and the epsilon-amino group of a lysine in a second polypeptide chain. TGases also catalyse the conjugation of polyamines to proteins [, ].; GO: 0003810 protein-glutamine gamma-glutamyltransferase activity, 0018149 peptide cross-linking; PDB: 2XZZ_A 1GGY_B 1FIE_B 1GGU_B 1GGT_B 1F13_A 1QRK_B 1EVU_A 1EX0_B 1L9N_B ....
Probab=21.22 E-value=3.3e+02 Score=19.90 Aligned_cols=61 Identities=11% Similarity=0.158 Sum_probs=32.9
Q ss_pred ceeEEEEEEEEEeCCCCee--EEEEccEEEEEEeCCcccceee--eccCCCceecCCCeEEEEEEEEe
Q 036189 97 NLIAKWDFNLTFKNPDHLW--QIYLDYIECIALNHDHFPIAIN--HSVSPPFKVKPMKKSTIHVQLAT 160 (241)
Q Consensus 97 ~l~~~~~~~l~v~NPN~k~--~i~Y~~~~v~v~Y~g~~~~~lg--~~~vp~F~q~~~~tt~v~v~l~~ 160 (241)
.+.-++++.+++.||...- .+.-.=....++|.|.. .. .........+++++..+...+.-
T Consensus 12 ~vG~d~~v~v~~~N~~~~~l~~v~~~l~~~~v~ytG~~---~~~~~~~~~~~~l~p~~~~~~~~~i~p 76 (107)
T PF00927_consen 12 VVGQDFTVSVSFTNPSSEPLRNVSLNLCAFTVEYTGLT---RDQFKKEKFEVTLKPGETKSVEVTITP 76 (107)
T ss_dssp BTTSEEEEEEEEEE-SSS-EECEEEEEEEEEEECTTTE---EEEEEEEEEEEEE-TTEEEEEEEEE-H
T ss_pred cCCCCEEEEEEEEeCCcCccccceeEEEEEEEEECCcc---cccEeEEEcceeeCCCCEEEEEEEEEc
Confidence 4556889999999996521 12222234567888874 32 22334455566666666555543
No 81
>PF15018 InaF-motif: TRP-interacting helix
Probab=21.10 E-value=1.4e+02 Score=18.53 Aligned_cols=20 Identities=30% Similarity=0.700 Sum_probs=9.8
Q ss_pred HHHHHHHhheeeE-EecCCCC
Q 036189 60 ISLTALICILMYF-TLGPKLP 79 (241)
Q Consensus 60 i~llgl~~lil~l-v~rPk~P 79 (241)
+-+.++...+.|+ +..|+.|
T Consensus 17 VSl~Ai~LsiYY~f~W~p~~~ 37 (38)
T PF15018_consen 17 VSLAAIVLSIYYIFFWDPDMP 37 (38)
T ss_pred HHHHHHHHHHHHheeeCCCCC
Confidence 3445554555553 4456543
No 82
>COG5294 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=20.87 E-value=2.5e+02 Score=21.71 Aligned_cols=13 Identities=15% Similarity=0.416 Sum_probs=9.4
Q ss_pred EEEEEEEEeCCCC
Q 036189 101 KWDFNLTFKNPDH 113 (241)
Q Consensus 101 ~~~~~l~v~NPN~ 113 (241)
-.+.++++.|-|.
T Consensus 53 ~y~y~i~ayn~~G 65 (113)
T COG5294 53 GYEYTITAYNKNG 65 (113)
T ss_pred cceeeehhhccCC
Confidence 4567788888776
No 83
>COG2332 CcmE Cytochrome c-type biogenesis protein CcmE [Posttranslational modification, protein turnover, chaperones]
Probab=20.85 E-value=3.3e+02 Score=22.19 Aligned_cols=32 Identities=3% Similarity=-0.064 Sum_probs=18.8
Q ss_pred EEEEEEEEeCCCCeeEEEEccEEEEEEeCCcc
Q 036189 101 KWDFNLTFKNPDHLWQIYLDYIECIALNHDHF 132 (241)
Q Consensus 101 ~~~~~l~v~NPN~k~~i~Y~~~~v~v~Y~g~~ 132 (241)
++.+.+.+.--|+++.+.|..+-=+++=+|+.
T Consensus 71 ~~~v~F~vtD~~~~v~V~Y~GiLPDLFREGQg 102 (153)
T COG2332 71 SLKVSFVVTDGNKSVTVSYEGILPDLFREGQG 102 (153)
T ss_pred CcEEEEEEecCCceEEEEEeccCchhhhcCCe
Confidence 34445555566666777776665555555554
No 84
>PHA03265 envelope glycoprotein D; Provisional
Probab=20.52 E-value=58 Score=30.22 Aligned_cols=23 Identities=17% Similarity=0.177 Sum_probs=13.0
Q ss_pred hhhHHHHHHHHHHHHHHHHHhheee
Q 036189 47 KLIIITLLVVAASISLTALICILMY 71 (241)
Q Consensus 47 rc~~~~~~~~~~~i~llgl~~lil~ 71 (241)
.++.+.+++. .++++|+|+.++|
T Consensus 350 ~g~~ig~~i~--glv~vg~il~~~~ 372 (402)
T PHA03265 350 VGISVGLGIA--GLVLVGVILYVCL 372 (402)
T ss_pred cceEEccchh--hhhhhhHHHHHHh
Confidence 3444444333 3577777766666
No 85
>PF01299 Lamp: Lysosome-associated membrane glycoprotein (Lamp); InterPro: IPR002000 Lysosome-associated membrane glycoproteins (lamp) [] are integral membrane proteins, specific to lysosomes, and whose exact biological function is not yet clear. Structurally, the lamp proteins consist of two internally homologous lysosome-luminal domains separated by a proline-rich hinge region; at the C-terminal extremity there is a transmembrane region (TM) followed by a very short cytoplasmic tail (C). In each of the duplicated domains, there are two conserved disulphide bonds. This structure is schematically represented in the figure below. +-----+ +-----+ +-----+ +-----+ | | | | | | | | xCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxxxCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxx +--------------------------++Hinge++--------------------------++TM++C+ In mammals, there are two closely related types of lamp: lamp-1 and lamp-2, which form major components of the lysosome membrane. In chicken lamp-1 is known as LEP100. Also included in this entry is the macrophage protein CD68 (or macrosialin) [] is a heavily glycosylated integral membrane protein whose structure consists of a mucin-like domain followed by a proline-rich hinge; a single lamp-like domain; a transmembrane region and a short cytoplasmic tail. Similar to CD68, mammalian lamp-3, which is expressed in lymphoid organs, dendritic cells and in lung, contains all the C-terminal regions but lacks the N-terminal lamp-like region []. In a lamp-family protein from nematodes [] only the part C-terminal to the hinge is conserved. ; GO: 0016020 membrane
Probab=20.50 E-value=81 Score=28.28 Aligned_cols=18 Identities=17% Similarity=0.222 Sum_probs=11.6
Q ss_pred HHHHHHhheeeEEecCCC
Q 036189 61 SLTALICILMYFTLGPKL 78 (241)
Q Consensus 61 ~llgl~~lil~lv~rPk~ 78 (241)
+.+.|++||.||+.|=|.
T Consensus 282 a~lvlivLiaYli~Rrr~ 299 (306)
T PF01299_consen 282 AGLVLIVLIAYLIGRRRS 299 (306)
T ss_pred HHHHHHHHHhheeEeccc
Confidence 344456777888877553
No 86
>PF02158 Neuregulin: Neuregulin family; InterPro: IPR002154 Neuregulins are a sub-family of EGF-like molecules that have been shown to play multiple essential roles in vertebrate embryogenesis including: cardiac development, Schwann cell and oligodendrocyte differentiation, some aspects of neuronal development, as well as the formation of neuromuscular synapses [, ]. Included in the family are heregulin; neu differentiation factor; acetylcholine receptor synthesis stimulator; glial growth factor; and sensory and motor-neuron derived factor []. Multiple family members are generated by alternate splicing or by use of several cell type-specific transcription initiation sites. In general, they bind to and activate the erbB family of receptor tyrosine kinases (erbB2 (HER2), erbB3 (HER3), and erbB4 (HER4)), functioning both as heterodimers and homodimers. The transmembrane forms of neuregulin 1 (NRG1) are present within synaptic vesicles, including those containing glutamate []. After exocytosis, NRG1 is in the presynaptic membrane, where the ectodomain of NRG1 may be cleaved off. The ectodomain then migrates across the synaptic cleft and binds to and activates a member of the EGF-receptor family on the postsynaptic membrane. This has been shown to increase the expression of certain glutamate-receptor subunits. NRG1 appears to signal for glutamate-receptor subunit expression, localisation, and /or phosphorylation facilitating subsequent glutamate transmission. The NRG1 gene has been identified as a potential gene determining susceptibility to schizophrenia by a combination of genetic linkage and association approaches []. ; GO: 0005102 receptor binding, 0009790 embryo development; PDB: 1HRE_A 1HAE_A 1HAF_A 1HRF_A.
Probab=20.47 E-value=35 Score=31.89 Aligned_cols=23 Identities=26% Similarity=0.671 Sum_probs=0.0
Q ss_pred hhHHHHHHHHHHHHHHHHHhhe-eeE
Q 036189 48 LIIITLLVVAASISLTALICIL-MYF 72 (241)
Q Consensus 48 c~~~~~~~~~~~i~llgl~~li-l~l 72 (241)
-+-|++|||. |+++|+++.+ +|.
T Consensus 9 VLTITgIcva--LlVVGi~Cvv~aYC 32 (404)
T PF02158_consen 9 VLTITGICVA--LLVVGIVCVVDAYC 32 (404)
T ss_dssp --------------------------
T ss_pred hhhhhhhhHH--HHHHHHHHHHHHHH
Confidence 3445555554 7889999998 884
No 87
>PF14828 Amnionless: Amnionless
Probab=20.43 E-value=61 Score=30.89 Aligned_cols=21 Identities=33% Similarity=0.393 Sum_probs=12.0
Q ss_pred HHHHHHhheeeEEecCCCCEE
Q 036189 61 SLTALICILMYFTLGPKLPSL 81 (241)
Q Consensus 61 ~llgl~~lil~lv~rPk~P~f 81 (241)
++++++++++|+.+.|+.|.+
T Consensus 349 llv~ll~~~~ll~~~~~~~~l 369 (437)
T PF14828_consen 349 LLVALLFGVILLYRLPRNPSL 369 (437)
T ss_pred HHHHHHHHhheEEeccccccc
Confidence 444555555565565666655
No 88
>PF09049 SNN_transmemb: Stannin transmembrane; InterPro: IPR015135 This region consists of a single highly hydrophobic transmembrane helix that transverses the lipid bilayer at a 20 degree angle with respect to the membrane normal. It contains a conserved cysteine residue (Cys32) that, together with Cys34 found in the stannin unstructured linker domain, constitutes the putative trimethyltin-binding site that resides at the end of the transmembrane domain close to the lipid/solvent interface []. ; PDB: 1ZZA_A.
Probab=20.31 E-value=1.2e+02 Score=17.85 Aligned_cols=16 Identities=13% Similarity=0.343 Sum_probs=6.4
Q ss_pred HHHHHHHHHHHHHHHh
Q 036189 52 TLLVVAASISLTALIC 67 (241)
Q Consensus 52 ~~~~~~~~i~llgl~~ 67 (241)
+++++++.+..+|+.+
T Consensus 14 ti~viliavaalg~li 29 (33)
T PF09049_consen 14 TIIVILIAVAALGALI 29 (33)
T ss_dssp HHHHHHHHHHHHHHHH
T ss_pred EehhHHHHHHHHhhhh
Confidence 4444443333444433
Done!