Query 019981
Match_columns 333
No_of_seqs 77 out of 79
Neff 3.3
Searched_HMMs 46136
Date Fri Mar 29 06:04:41 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/019981.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/019981hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PRK08097 ligB NAD-dependent DN 95.0 0.026 5.7E-07 59.2 4.4 49 125-173 8-69 (562)
2 cd00114 LIGANc NAD+ dependent 94.5 0.038 8.3E-07 53.6 3.8 27 148-174 12-39 (307)
3 PRK00398 rpoP DNA-directed RNA 94.4 0.039 8.5E-07 39.3 2.9 37 274-326 5-41 (46)
4 PF11023 DUF2614: Protein of u 93.9 0.058 1.3E-06 46.6 3.3 64 242-323 37-102 (114)
5 PF01653 DNA_ligase_aden: NAD- 93.7 0.072 1.6E-06 51.9 4.1 25 149-173 17-42 (315)
6 COG2051 RPS27A Ribosomal prote 93.7 0.071 1.5E-06 42.3 3.3 45 270-329 17-61 (67)
7 smart00532 LIGANc Ligase N fam 92.8 0.11 2.4E-06 53.0 3.9 25 149-173 15-40 (441)
8 TIGR00575 dnlj DNA ligase, NAD 92.7 0.12 2.6E-06 54.9 4.2 29 146-174 5-34 (652)
9 PRK07956 ligA NAD-dependent DN 92.6 0.12 2.5E-06 55.2 3.9 26 148-173 18-44 (665)
10 TIGR01206 lysW lysine biosynth 92.5 0.21 4.5E-06 37.9 4.1 44 271-328 1-46 (54)
11 PRK00420 hypothetical protein; 91.6 0.17 3.6E-06 43.4 3.0 46 257-319 8-53 (112)
12 TIGR02098 MJ0042_CXXC MJ0042 f 91.5 0.16 3.6E-06 34.4 2.3 36 272-318 2-37 (38)
13 PF13240 zinc_ribbon_2: zinc-r 90.0 0.14 3.1E-06 32.5 0.9 9 275-283 2-10 (23)
14 cd00114 LIGANc NAD+ dependent 89.8 0.39 8.5E-06 46.7 4.1 35 96-130 4-38 (307)
15 PF06044 DRP: Dam-replacing fa 89.8 0.14 3E-06 49.4 1.0 51 256-322 19-69 (254)
16 PRK14351 ligA NAD-dependent DN 89.6 0.35 7.6E-06 52.0 4.0 25 149-173 46-71 (689)
17 PRK00415 rps27e 30S ribosomal 89.6 0.27 5.9E-06 38.2 2.3 41 269-324 8-48 (59)
18 PF13248 zf-ribbon_3: zinc-rib 89.6 0.16 3.4E-06 32.7 0.9 10 273-282 3-12 (26)
19 COG2888 Predicted Zn-ribbon RN 89.3 0.17 3.7E-06 39.6 1.0 34 267-305 22-55 (61)
20 PF01653 DNA_ligase_aden: NAD- 87.7 0.7 1.5E-05 45.1 4.3 36 95-130 7-42 (315)
21 smart00532 LIGANc Ligase N fam 87.6 0.59 1.3E-05 47.8 3.8 36 95-130 5-40 (441)
22 PRK14350 ligA NAD-dependent DN 87.2 0.67 1.5E-05 49.8 4.2 24 150-173 20-44 (669)
23 PRK14890 putative Zn-ribbon RN 87.0 0.32 7E-06 37.8 1.3 30 270-305 23-53 (59)
24 PF01667 Ribosomal_S27e: Ribos 86.3 1.1 2.4E-05 34.3 3.8 42 270-326 5-46 (55)
25 smart00531 TFIIE Transcription 86.1 0.43 9.2E-06 41.4 1.7 55 260-324 87-141 (147)
26 PRK08097 ligB NAD-dependent DN 85.8 0.76 1.7E-05 48.5 3.6 36 95-130 34-69 (562)
27 PF08271 TF_Zn_Ribbon: TFIIB z 85.0 0.58 1.3E-05 32.9 1.7 32 274-320 2-33 (43)
28 PRK05978 hypothetical protein; 85.0 0.44 9.6E-06 42.6 1.3 34 271-319 32-65 (148)
29 PRK14351 ligA NAD-dependent DN 84.5 1.1 2.3E-05 48.4 4.2 36 95-130 36-71 (689)
30 PRK07956 ligA NAD-dependent DN 84.5 1.1 2.4E-05 48.0 4.2 36 95-130 9-44 (665)
31 PF14255 Cys_rich_CPXG: Cystei 83.6 0.85 1.8E-05 34.4 2.1 37 274-320 2-38 (52)
32 PHA00626 hypothetical protein 81.8 1.1 2.4E-05 34.9 2.2 35 274-319 2-36 (59)
33 smart00659 RPOLCX RNA polymera 81.5 1.9 4.1E-05 31.3 3.2 34 274-324 4-38 (44)
34 PRK02935 hypothetical protein; 81.2 1.6 3.4E-05 37.8 3.1 36 270-323 68-103 (110)
35 TIGR00575 dnlj DNA ligase, NAD 80.6 1.5 3.3E-05 46.8 3.5 27 104-130 7-33 (652)
36 PLN00209 ribosomal protein S27 80.2 2 4.3E-05 35.8 3.2 45 266-325 30-74 (86)
37 TIGR00373 conserved hypothetic 79.8 0.46 9.9E-06 42.1 -0.6 54 257-325 94-147 (158)
38 PRK14350 ligA NAD-dependent DN 79.2 1.7 3.7E-05 46.8 3.3 36 95-130 9-44 (669)
39 COG0272 Lig NAD-dependent DNA 78.9 2.1 4.6E-05 46.3 3.8 24 151-174 23-47 (667)
40 PF07754 DUF1610: Domain of un 78.8 0.88 1.9E-05 29.7 0.6 11 270-280 14-24 (24)
41 PTZ00083 40S ribosomal protein 78.6 2.4 5.2E-05 35.2 3.3 45 266-325 29-73 (85)
42 PF05876 Terminase_GpA: Phage 78.3 1.3 2.9E-05 46.1 2.1 54 264-323 192-246 (557)
43 PF10571 UPF0547: Uncharacteri 78.1 1.2 2.7E-05 29.2 1.2 9 274-282 2-10 (26)
44 PRK09710 lar restriction allev 77.0 1.7 3.7E-05 34.4 1.9 34 273-319 7-40 (64)
45 PF14803 Nudix_N_2: Nudix N-te 76.4 1.8 3.8E-05 30.1 1.6 29 275-315 3-31 (34)
46 PRK06266 transcription initiat 76.0 0.62 1.4E-05 42.1 -0.9 55 256-325 101-155 (178)
47 PF14353 CpXC: CpXC protein 75.1 2.9 6.2E-05 35.0 2.9 40 274-319 3-51 (128)
48 PF09851 SHOCT: Short C-termin 74.6 4.8 0.0001 27.0 3.3 25 147-173 5-30 (31)
49 PF14354 Lar_restr_allev: Rest 74.1 2.6 5.7E-05 30.9 2.2 33 274-314 5-37 (61)
50 TIGR03831 YgiT_finger YgiT-typ 72.7 2.1 4.6E-05 29.2 1.3 19 266-284 22-44 (46)
51 smart00661 RPOL9 RNA polymeras 71.8 5.7 0.00012 28.1 3.4 34 274-321 2-35 (52)
52 PF06906 DUF1272: Protein of u 70.7 1.8 3.9E-05 33.6 0.7 13 270-282 39-51 (57)
53 smart00834 CxxC_CXXC_SSSS Puta 70.1 3.8 8.3E-05 27.6 2.1 29 274-315 7-35 (41)
54 PF01096 TFIIS_C: Transcriptio 68.8 3.5 7.6E-05 28.9 1.7 31 274-305 2-33 (39)
55 smart00778 Prim_Zn_Ribbon Zinc 68.6 2.2 4.8E-05 30.2 0.7 13 272-284 3-16 (37)
56 PF13719 zinc_ribbon_5: zinc-r 68.0 3.7 7.9E-05 28.4 1.7 35 272-317 2-36 (37)
57 COG5349 Uncharacterized protei 68.0 2 4.4E-05 37.9 0.5 33 271-320 20-54 (126)
58 PF09851 SHOCT: Short C-termin 68.0 7.1 0.00015 26.2 3.0 26 103-130 5-30 (31)
59 PF09862 DUF2089: Protein of u 67.3 4.4 9.5E-05 35.0 2.4 23 275-317 1-23 (113)
60 smart00440 ZnF_C2C2 C2C2 Zinc 66.9 4.9 0.00011 28.3 2.2 32 274-305 2-33 (40)
61 PF07282 OrfB_Zn_ribbon: Putat 65.2 4.9 0.00011 30.1 2.0 28 272-315 28-55 (69)
62 TIGR03655 anti_R_Lar restricti 65.2 5.5 0.00012 29.2 2.2 36 274-318 3-38 (53)
63 COG1645 Uncharacterized Zn-fin 64.1 5.3 0.00011 35.4 2.3 39 259-315 15-53 (131)
64 COG3813 Uncharacterized protei 63.4 3.1 6.7E-05 34.2 0.7 15 270-284 39-53 (84)
65 TIGR00340 zpr1_rel ZPR1-relate 63.1 4.7 0.0001 36.4 1.8 20 264-283 20-39 (163)
66 PF03367 zf-ZPR1: ZPR1 zinc-fi 62.7 4 8.6E-05 36.5 1.3 20 265-284 23-42 (161)
67 smart00709 Zpr1 Duplicated dom 62.0 4.9 0.00011 36.1 1.7 22 264-285 21-42 (160)
68 PRK00432 30S ribosomal protein 61.2 5.9 0.00013 29.4 1.8 29 270-315 18-46 (50)
69 PF12773 DZR: Double zinc ribb 61.1 3.9 8.6E-05 28.9 0.8 9 275-283 15-23 (50)
70 PF08273 Prim_Zn_Ribbon: Zinc- 60.9 3.6 7.8E-05 29.5 0.6 15 272-286 3-18 (40)
71 PF05191 ADK_lid: Adenylate ki 60.6 4.8 0.0001 28.1 1.1 33 274-320 3-35 (36)
72 PF09538 FYDLN_acid: Protein o 59.9 6.8 0.00015 33.3 2.2 32 271-319 8-39 (108)
73 PF04216 FdhE: Protein involve 59.7 9.1 0.0002 36.4 3.2 43 246-288 146-189 (290)
74 TIGR03830 CxxCG_CxxCG_HTH puta 58.7 6 0.00013 32.2 1.6 39 275-316 1-41 (127)
75 COG0272 Lig NAD-dependent DNA 56.5 13 0.00027 40.6 3.9 38 93-130 9-46 (667)
76 PF09723 Zn-ribbon_8: Zinc rib 55.3 10 0.00022 26.7 2.1 28 274-314 7-34 (42)
77 PRK01103 formamidopyrimidine/5 54.8 10 0.00022 36.0 2.7 32 266-305 235-270 (274)
78 PF03604 DNA_RNApol_7kD: DNA d 54.4 7.6 0.00017 26.7 1.3 27 275-318 3-29 (32)
79 PRK08665 ribonucleotide-diphos 54.2 13 0.00027 40.8 3.5 13 273-286 725-737 (752)
80 PF11746 DUF3303: Protein of u 53.9 9.9 0.00021 31.0 2.1 70 96-165 11-89 (91)
81 PF14319 Zn_Tnp_IS91: Transpos 52.9 8.4 0.00018 32.4 1.6 29 271-315 41-69 (111)
82 PF13717 zinc_ribbon_4: zinc-r 52.6 11 0.00023 26.1 1.8 33 272-315 2-34 (36)
83 PF07508 Recombinase: Recombin 50.6 13 0.00029 28.9 2.3 20 155-174 82-101 (102)
84 PRK14714 DNA polymerase II lar 49.8 13 0.00028 43.3 2.9 22 267-288 662-683 (1337)
85 TIGR02300 FYDLN_acid conserved 49.7 12 0.00025 33.3 2.0 32 271-319 8-39 (129)
86 PRK13130 H/ACA RNA-protein com 49.3 7.8 0.00017 29.8 0.8 15 270-284 15-29 (56)
87 COG4306 Uncharacterized protei 49.0 10 0.00023 34.1 1.6 39 274-319 41-81 (160)
88 PRK14810 formamidopyrimidine-D 48.4 11 0.00024 35.9 1.9 25 273-305 245-269 (272)
89 COG5525 Bacteriophage tail ass 48.3 13 0.00028 40.2 2.4 66 261-328 216-281 (611)
90 PRK10445 endonuclease VIII; Pr 48.1 12 0.00026 35.5 2.0 25 273-305 236-260 (263)
91 TIGR00097 HMP-P_kinase phospho 48.1 68 0.0015 29.3 6.8 57 120-184 115-172 (254)
92 COG3677 Transposase and inacti 47.7 17 0.00036 31.6 2.7 49 265-324 23-71 (129)
93 TIGR00577 fpg formamidopyrimid 47.3 12 0.00026 35.7 1.8 24 274-305 247-270 (272)
94 COG1998 RPS31 Ribosomal protei 47.2 14 0.0003 28.3 1.8 19 268-286 15-34 (51)
95 COG1096 Predicted RNA-binding 46.0 16 0.00034 34.3 2.3 41 265-328 142-182 (188)
96 PF01807 zf-CHC2: CHC2 zinc fi 45.8 11 0.00025 30.6 1.3 31 271-314 32-62 (97)
97 TIGR02605 CxxC_CxxC_SSSS putat 45.4 22 0.00047 25.4 2.6 28 274-314 7-34 (52)
98 PF10083 DUF2321: Uncharacteri 45.1 12 0.00026 34.2 1.5 37 274-317 41-79 (158)
99 PF03119 DNA_ligase_ZBD: NAD-d 44.3 13 0.00029 24.5 1.2 11 275-285 2-12 (28)
100 PRK13945 formamidopyrimidine-D 43.6 15 0.00031 35.3 1.8 25 273-305 255-279 (282)
101 PF09889 DUF2116: Uncharacteri 43.3 11 0.00023 29.3 0.7 11 273-283 4-14 (59)
102 PF12677 DUF3797: Domain of un 43.1 14 0.0003 28.1 1.2 12 272-283 13-24 (49)
103 PF08996 zf-DNA_Pol: DNA Polym 42.2 10 0.00022 34.2 0.6 45 264-315 10-54 (188)
104 PRK14811 formamidopyrimidine-D 41.8 16 0.00035 34.9 1.8 28 274-315 237-264 (269)
105 PF10263 SprT-like: SprT-like 41.5 22 0.00048 30.0 2.5 36 269-318 120-155 (157)
106 PRK14892 putative transcriptio 40.8 20 0.00043 30.3 2.0 32 274-318 23-54 (99)
107 PF12760 Zn_Tnp_IS1595: Transp 40.7 23 0.00051 25.2 2.1 22 275-305 21-42 (46)
108 PF05502 Dynactin_p62: Dynacti 40.4 17 0.00036 37.9 1.8 45 269-322 23-68 (483)
109 TIGR01054 rgy reverse gyrase. 39.7 12 0.00026 42.8 0.7 17 269-285 4-20 (1171)
110 PRK00398 rpoP DNA-directed RNA 39.7 29 0.00064 24.5 2.5 24 300-329 3-26 (46)
111 PF06677 Auto_anti-p27: Sjogre 38.8 27 0.00059 25.2 2.2 25 260-284 5-29 (41)
112 PRK00464 nrdR transcriptional 37.8 26 0.00055 31.5 2.4 37 274-316 2-38 (154)
113 PF11793 FANCL_C: FANCL C-term 37.4 11 0.00024 29.2 0.0 19 265-283 48-66 (70)
114 PLN02919 haloacid dehalogenase 37.3 2.5E+02 0.0055 32.0 10.4 30 94-123 85-114 (1057)
115 PF04380 BMFP: Membrane fusoge 36.8 55 0.0012 26.1 3.8 36 94-130 25-60 (79)
116 PF05129 Elf1: Transcription e 36.7 28 0.00061 28.1 2.2 37 273-319 23-59 (81)
117 PRK04023 DNA polymerase II lar 36.5 23 0.0005 40.6 2.2 17 270-286 624-640 (1121)
118 PLN00049 carboxyl-terminal pro 36.4 70 0.0015 31.9 5.4 71 95-172 2-80 (389)
119 PRK12495 hypothetical protein; 35.7 32 0.0007 33.1 2.8 28 258-285 28-55 (226)
120 COG3877 Uncharacterized protei 35.1 29 0.00062 30.5 2.1 26 272-317 6-31 (122)
121 COG0675 Transposase and inacti 34.9 25 0.00053 31.9 1.8 22 273-315 310-331 (364)
122 PF08209 Sgf11: Sgf11 (transcr 34.7 16 0.00036 25.3 0.5 10 274-283 6-15 (33)
123 PF04280 Tim44: Tim44-like dom 34.1 21 0.00045 29.8 1.1 37 139-175 21-62 (147)
124 PRK12412 pyridoxal kinase; Rev 33.9 1.5E+02 0.0032 27.6 6.7 57 119-183 119-176 (268)
125 cd07110 ALDH_F10_BADH Arabidop 33.1 1.2E+02 0.0026 30.3 6.4 70 116-185 238-334 (456)
126 TIGR01562 FdhE formate dehydro 32.9 27 0.00059 34.6 1.9 11 272-282 224-234 (305)
127 PRK08351 DNA-directed RNA poly 32.8 20 0.00044 28.0 0.8 16 274-289 17-34 (61)
128 COG4443 Uncharacterized protei 32.5 28 0.0006 28.2 1.5 18 113-130 52-70 (72)
129 PRK00133 metG methionyl-tRNA s 31.9 40 0.00087 36.0 3.1 16 308-323 171-186 (673)
130 PF06170 DUF983: Protein of un 31.3 21 0.00045 29.2 0.7 17 267-283 3-19 (86)
131 PF09567 RE_MamI: MamI restric 31.1 21 0.00046 35.4 0.8 13 273-285 83-95 (314)
132 cd01169 HMPP_kinase 4-amino-5- 31.1 2.1E+02 0.0046 25.4 7.1 53 123-183 119-172 (242)
133 PF06221 zf-C2HC5: Putative zi 30.6 25 0.00053 27.2 0.9 13 272-284 35-47 (57)
134 PF05416 Peptidase_C37: Southa 30.3 17 0.00037 38.3 0.0 43 116-166 252-297 (535)
135 TIGR01031 rpmF_bact ribosomal 29.9 32 0.00069 26.0 1.4 13 273-285 27-39 (55)
136 PRK12286 rpmF 50S ribosomal pr 29.7 34 0.00073 26.1 1.5 12 273-284 28-39 (57)
137 COG1996 RPC10 DNA-directed RNA 29.2 41 0.00089 25.4 1.9 30 274-319 8-37 (49)
138 PF09863 DUF2090: Uncharacteri 29.0 1.4E+02 0.0031 30.1 6.1 49 94-142 188-246 (311)
139 PF14206 Cys_rich_CPCC: Cystei 28.5 25 0.00054 28.6 0.7 12 272-283 1-12 (78)
140 PF02829 3H: 3H domain; Inter 28.2 94 0.002 26.1 4.0 32 144-176 50-95 (98)
141 PRK14714 DNA polymerase II lar 28.1 33 0.00071 40.2 1.7 12 271-282 678-689 (1337)
142 PF12767 SAGA-Tad1: Transcript 28.1 59 0.0013 30.5 3.2 36 92-130 19-55 (252)
143 COG1675 TFA1 Transcription ini 28.0 15 0.00033 33.9 -0.8 39 273-326 114-152 (176)
144 PRK08176 pdxK pyridoxal-pyrido 27.8 1.5E+02 0.0033 27.9 5.8 53 123-183 143-196 (281)
145 PF15616 TerY-C: TerY-C metal 27.7 51 0.0011 29.2 2.5 45 268-321 73-120 (131)
146 PF03965 Penicillinase_R: Peni 27.0 1.9E+02 0.0042 23.6 5.7 33 95-131 1-33 (115)
147 PRK04011 peptide chain release 26.7 28 0.00061 35.4 0.8 36 272-319 328-363 (411)
148 PF03317 ELF: ELF protein; In 26.5 1.5E+02 0.0034 28.8 5.6 106 50-177 144-255 (284)
149 smart00400 ZnF_CHCC zinc finge 26.3 39 0.00085 24.6 1.3 27 272-305 2-28 (55)
150 cd07114 ALDH_DhaS Uncharacteri 26.1 1.8E+02 0.004 29.1 6.4 69 116-184 237-332 (457)
151 PRK09401 reverse gyrase; Revie 25.7 28 0.00061 40.0 0.7 16 269-284 4-19 (1176)
152 COG5257 GCD11 Translation init 25.7 54 0.0012 33.9 2.6 44 265-329 52-95 (415)
153 COG1326 Uncharacterized archae 25.2 30 0.00065 32.8 0.7 41 271-319 5-47 (201)
154 TIGR00398 metG methionyl-tRNA 25.0 57 0.0012 33.3 2.7 48 272-323 136-183 (530)
155 cd07078 ALDH NAD(P)+ dependent 24.9 2.2E+02 0.0048 27.9 6.6 68 117-184 215-309 (432)
156 PRK03564 formate dehydrogenase 24.9 53 0.0011 32.7 2.3 24 144-169 104-127 (309)
157 PF05193 Peptidase_M16_C: Pept 24.9 1.2E+02 0.0025 24.4 3.9 33 94-129 152-184 (184)
158 PF03884 DUF329: Domain of unk 24.8 26 0.00057 27.0 0.2 13 272-284 2-14 (57)
159 COG1779 C4-type Zn-finger prot 24.7 43 0.00092 31.8 1.6 30 269-305 11-48 (201)
160 PTZ00381 aldehyde dehydrogenas 24.7 1.7E+02 0.0037 30.2 6.0 67 116-183 224-314 (493)
161 cd02661 Peptidase_C19E A subfa 24.6 79 0.0017 28.6 3.3 24 300-329 182-205 (304)
162 PF06750 DiS_P_DiS: Bacterial 24.5 40 0.00086 27.6 1.2 26 245-283 44-69 (92)
163 TIGR00100 hypA hydrogenase nic 24.1 44 0.00095 28.2 1.4 14 267-280 65-78 (115)
164 PF12207 DUF3600: Domain of un 24.0 49 0.0011 30.5 1.7 66 91-156 72-159 (162)
165 TIGR01391 dnaG DNA primase, ca 23.9 48 0.001 33.5 1.9 28 271-305 33-60 (415)
166 PRK00448 polC DNA polymerase I 23.8 46 0.00099 39.4 1.9 36 274-318 910-945 (1437)
167 PF14485 DUF4431: Domain of un 23.7 67 0.0014 23.8 2.1 22 159-183 5-26 (48)
168 cd07120 ALDH_PsfA-ACA09737 Pse 23.7 2.3E+02 0.0049 28.8 6.6 70 116-185 236-332 (455)
169 PF13058 DUF3920: Protein of u 23.4 36 0.00078 30.1 0.7 47 84-130 30-81 (126)
170 KOG2589 Histone tail methylase 23.3 37 0.0008 35.3 0.9 30 279-317 229-258 (453)
171 PRK15398 aldehyde dehydrogenas 23.3 1.7E+02 0.0036 30.1 5.6 59 116-176 247-316 (465)
172 PRK06427 bifunctional hydroxy- 23.0 3.1E+02 0.0066 25.0 6.7 54 123-184 124-179 (266)
173 PF00412 LIM: LIM domain; Int 23.0 34 0.00074 24.1 0.5 37 275-317 1-37 (58)
174 TIGR01405 polC_Gram_pos DNA po 23.0 48 0.001 38.5 1.8 36 274-318 685-720 (1213)
175 PRK03824 hypA hydrogenase nick 22.9 76 0.0017 27.6 2.7 12 269-280 67-78 (135)
176 COG3464 Transposase and inacti 22.9 62 0.0013 32.8 2.4 47 270-316 36-87 (402)
177 PF08976 DUF1880: Domain of un 22.5 39 0.00085 29.7 0.8 19 115-133 2-22 (118)
178 PF13913 zf-C2HC_2: zinc-finge 22.3 40 0.00087 21.5 0.6 8 274-281 4-11 (25)
179 KOG3457 Sec61 protein transloc 22.2 47 0.001 27.9 1.2 30 218-247 46-75 (88)
180 PF05907 DUF866: Eukaryotic pr 22.2 67 0.0014 28.9 2.2 45 269-319 27-77 (161)
181 KOG1372 GDP-mannose 4,6 dehydr 22.0 47 0.001 33.4 1.4 39 99-137 257-309 (376)
182 PF10751 DUF2535: Protein of u 21.9 88 0.0019 26.1 2.7 40 136-175 21-68 (83)
183 PF14952 zf-tcix: Putative tre 21.8 41 0.00088 25.1 0.7 9 274-282 13-21 (44)
184 TIGR00777 ahpD alkylhydroperox 21.8 47 0.001 30.8 1.2 26 150-175 79-105 (177)
185 PRK12380 hydrogenase nickel in 21.6 51 0.0011 27.8 1.3 14 267-280 65-78 (113)
186 COG2960 Uncharacterized protei 21.5 1.1E+02 0.0024 26.4 3.3 51 92-151 32-82 (103)
187 PF09845 DUF2072: Zn-ribbon co 21.5 44 0.00096 29.8 0.9 21 268-289 16-36 (131)
188 PRK03564 formate dehydrogenase 21.4 1.2E+02 0.0026 30.3 4.0 13 271-283 186-198 (309)
189 cd07092 ALDH_ABALDH-YdcW Esche 21.4 2.8E+02 0.006 27.7 6.5 68 116-183 235-328 (450)
190 PF04475 DUF555: Protein of un 21.4 44 0.00095 28.8 0.8 15 270-284 45-59 (102)
191 COG3024 Uncharacterized protei 21.3 44 0.00096 26.7 0.8 14 270-283 5-18 (65)
192 smart00709 Zpr1 Duplicated dom 21.3 46 0.001 29.9 1.0 6 274-279 2-7 (160)
193 PF04328 DUF466: Protein of un 21.2 1.8E+02 0.0038 22.8 4.1 34 144-177 26-59 (65)
194 PF09334 tRNA-synt_1g: tRNA sy 21.1 62 0.0013 32.5 2.0 9 271-279 135-143 (391)
195 TIGR00310 ZPR1_znf ZPR1 zinc f 21.1 47 0.001 30.8 1.1 20 264-283 22-41 (192)
196 cd02674 Peptidase_C19R A subfa 20.8 99 0.0021 27.0 3.0 25 299-329 103-127 (230)
197 cd07105 ALDH_SaliADH Salicylal 20.8 2.7E+02 0.0059 27.7 6.4 70 116-185 219-310 (432)
198 PF11682 DUF3279: Protein of u 20.7 53 0.0012 29.0 1.3 15 307-321 29-43 (128)
199 PHA02942 putative transposase; 20.7 69 0.0015 32.3 2.2 27 273-316 326-352 (383)
200 PF04423 Rad50_zn_hook: Rad50 20.7 43 0.00093 24.4 0.6 6 312-317 26-31 (54)
201 PF08278 DnaG_DnaB_bind: DNA p 20.6 1.8E+02 0.0039 23.8 4.3 48 94-150 79-126 (127)
202 PF09297 zf-NADH-PPase: NADH p 20.5 34 0.00074 22.6 0.0 10 274-283 23-32 (32)
203 PF12207 DUF3600: Domain of un 20.4 83 0.0018 29.0 2.5 65 112-176 35-122 (162)
204 PF09930 DUF2162: Predicted tr 20.4 79 0.0017 30.1 2.4 35 247-282 73-115 (224)
205 PF04216 FdhE: Protein involve 20.4 58 0.0013 31.1 1.6 15 270-284 209-223 (290)
206 PLN02278 succinic semialdehyde 20.2 2.9E+02 0.0063 28.4 6.6 68 116-183 278-372 (498)
207 PF09855 DUF2082: Nucleic-acid 20.2 1.1E+02 0.0023 24.1 2.7 42 274-322 2-52 (64)
208 TIGR02159 PA_CoA_Oxy4 phenylac 20.1 34 0.00073 30.4 -0.1 38 272-318 105-142 (146)
No 1
>PRK08097 ligB NAD-dependent DNA ligase LigB; Reviewed
Probab=94.99 E-value=0.026 Score=59.16 Aligned_cols=49 Identities=35% Similarity=0.531 Sum_probs=34.5
Q ss_pred hHhhhhhcCCe---eEEeChhh--HH--HHHH-----HHhhhc-CCCccChHHHHHHHHHHh
Q 019981 125 LKEELMWEGSS---VVMLSSAE--QK--FLEA-----SMAYVA-GKPIMSDEEYDKLKQKLK 173 (333)
Q Consensus 125 LkEeL~weGSs---vv~L~~~E--q~--fLEA-----~~AY~s-GkPimsDeeFD~LK~kLk 173 (333)
|---|.|..|- |.+++..| ++ .|.+ -.+||. |+|+|||+|||+|..+|+
T Consensus 8 ~~~~~~~~~~~~~~~~~~~~~~~~~~i~~L~~~l~~~~~~YY~~~~p~IsD~eYD~L~~eL~ 69 (562)
T PRK08097 8 LISLLLWSSSAWAVCPDWSPARAQEEIAALQQQLAQWDDAYWRQGKSEVDDEVYDQLRARLT 69 (562)
T ss_pred HHHHHHhcccccccCCCCCHHHHHHHHHHHHHHHHHHHHHHHhCCCCCCChHHHHHHHHHHH
Confidence 34457898887 55666644 11 2222 246775 999999999999999997
No 2
>cd00114 LIGANc NAD+ dependent DNA ligase adenylation domain. DNA ligases catalyze the crucial step of joining the breaks in duplex DNA during DNA replication, repair and recombination, utilizing either ATP or NAD(+) as a cofactor, but using the same basic reaction mechanism. The enzyme reacts with the cofactor to form a phosphoamide-linked AMP with the amino group of a conserved Lysine in the KXDG motif, and subsequently transfers it to the DNA substrate to yield adenylated DNA. This alignment contains members of the NAD+ dependent subfamily only.
Probab=94.46 E-value=0.038 Score=53.62 Aligned_cols=27 Identities=33% Similarity=0.641 Sum_probs=22.9
Q ss_pred HHHHhhhc-CCCccChHHHHHHHHHHhh
Q 019981 148 EASMAYVA-GKPIMSDEEYDKLKQKLKM 174 (333)
Q Consensus 148 EA~~AY~s-GkPimsDeeFD~LK~kLk~ 174 (333)
++-.+||. |+|+|||+|||.|.++|+.
T Consensus 12 ~~~~~YY~~~~p~IsD~eYD~L~~~L~~ 39 (307)
T cd00114 12 KHDYRYYVLDEPSVSDAEYDRLYRELRA 39 (307)
T ss_pred HHHHHHHhCCCCCCChHHHHHHHHHHHH
Confidence 34456887 9999999999999999973
No 3
>PRK00398 rpoP DNA-directed RNA polymerase subunit P; Provisional
Probab=94.44 E-value=0.039 Score=39.28 Aligned_cols=37 Identities=22% Similarity=0.613 Sum_probs=26.1
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccceeee
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRLIT 326 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~it 326 (333)
-|||||.++-. -. .. ..++|+. ||..+.|....++|.
T Consensus 5 ~C~~CG~~~~~-~~-------~~--~~~~Cp~------CG~~~~~~~~~~~v~ 41 (46)
T PRK00398 5 KCARCGREVEL-DE-------YG--TGVRCPY------CGYRILFKERPPVVK 41 (46)
T ss_pred ECCCCCCEEEE-CC-------CC--CceECCC------CCCeEEEccCCCcce
Confidence 59999997543 11 11 1689999 999999986666553
No 4
>PF11023 DUF2614: Protein of unknown function (DUF2614); InterPro: IPR020912 This entry describes proteins of unknown function, which are thought to be membrane proteins.; GO: 0005887 integral to plasma membrane
Probab=93.88 E-value=0.058 Score=46.59 Aligned_cols=64 Identities=19% Similarity=0.447 Sum_probs=37.1
Q ss_pred hHHHHHhhhhHHHHHHHHHHhhhhc--ceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEe
Q 019981 242 FIFTWFAAVPLIVYLSQSLTKLIVR--ESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 242 fi~t~~~~~P~i~~~a~~Lt~l~~~--D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~ 319 (333)
++.+.++.+.++..+++...=.|.+ -.-+.--.|||||.+....= +.. -|.+ |+++|+.|
T Consensus 37 ~im~ifmllG~L~~l~S~~VYfwIGmlStkav~V~CP~C~K~TKmLG---------r~D---~CM~------C~~pLTLd 98 (114)
T PF11023_consen 37 IIMVIFMLLGLLAILASTAVYFWIGMLSTKAVQVECPNCGKQTKMLG---------RVD---ACMH------CKEPLTLD 98 (114)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHhhhhcccceeeECCCCCChHhhhc---------hhh---ccCc------CCCcCccC
Confidence 4444455555555444443333321 11223334999999987652 222 6877 99999999
Q ss_pred ccce
Q 019981 320 SNTR 323 (333)
Q Consensus 320 tk~R 323 (333)
+..+
T Consensus 99 ~~le 102 (114)
T PF11023_consen 99 PSLE 102 (114)
T ss_pred chhh
Confidence 7654
No 5
>PF01653 DNA_ligase_aden: NAD-dependent DNA ligase adenylation domain; InterPro: IPR013839 DNA ligase (polydeoxyribonucleotide synthase) is the enzyme that joins two DNA fragments by catalyzing the formation of an internucleotide ester bond between phosphate and deoxyribose. It is active during DNA replication, DNA repair and DNA recombination. There are two forms of DNA ligase: one requires ATP (6.5.1.1 from EC), the other NAD (6.5.1.2 from EC). This entry represents the N-terminal adenylation domain of NAD-dependent DNA ligases. These are proteins of about 75 to 85 Kd whose sequence is well conserved [, ]. They also show similarity to yicF, an Escherichia coli hypothetical protein of 63 Kd. Despite a complete lack of detectable sequence similarity, the fold of the central core of this adenyaltion domain shares homology with the equivalent region of ATP-dependent DNA ligases [, ].; GO: 0003911 DNA ligase (NAD+) activity; PDB: 1ZAU_A 3SGI_A 1B04_A 3JSL_A 3JSN_A 1DGS_A 1V9P_A 3PN1_A 3BAC_A 3UQ8_A ....
Probab=93.74 E-value=0.072 Score=51.86 Aligned_cols=25 Identities=52% Similarity=0.875 Sum_probs=21.6
Q ss_pred HHHhhhc-CCCccChHHHHHHHHHHh
Q 019981 149 ASMAYVA-GKPIMSDEEYDKLKQKLK 173 (333)
Q Consensus 149 A~~AY~s-GkPimsDeeFD~LK~kLk 173 (333)
+-.+||. |+|+|||+|||.|..+|+
T Consensus 17 ~~~~YY~~~~p~isD~eYD~l~~~L~ 42 (315)
T PF01653_consen 17 HNYAYYNLGEPIISDAEYDQLFRELK 42 (315)
T ss_dssp HHHHHHTTSSSSSSHHHHHHHHHHHH
T ss_pred HHHHHhcCCCCCCCHHHHHHHHHHHH
Confidence 3457777 899999999999999986
No 6
>COG2051 RPS27A Ribosomal protein S27E [Translation, ribosomal structure and biogenesis]
Probab=93.70 E-value=0.071 Score=42.33 Aligned_cols=45 Identities=29% Similarity=0.659 Sum_probs=36.4
Q ss_pred eeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccceeeecCC
Q 019981 270 ILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRLITLPE 329 (333)
Q Consensus 270 iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~itl~~ 329 (333)
-|+--||.||.|--.| .+++..|+|.. ||+.|..-+.-++...++
T Consensus 17 Fl~VkCpdC~N~q~vF---------shast~V~C~~------CG~~l~~PTGGka~i~~~ 61 (67)
T COG2051 17 FLRVKCPDCGNEQVVF---------SHASTVVTCLI------CGTTLAEPTGGKAKISGK 61 (67)
T ss_pred EEEEECCCCCCEEEEe---------ccCceEEEecc------cccEEEecCCCeEEeeee
Confidence 4667799999998888 45777889987 999999888877766654
No 7
>smart00532 LIGANc Ligase N family.
Probab=92.81 E-value=0.11 Score=52.99 Aligned_cols=25 Identities=44% Similarity=0.769 Sum_probs=22.0
Q ss_pred HHHhhhc-CCCccChHHHHHHHHHHh
Q 019981 149 ASMAYVA-GKPIMSDEEYDKLKQKLK 173 (333)
Q Consensus 149 A~~AY~s-GkPimsDeeFD~LK~kLk 173 (333)
+-.+||. |+|+|||+|||+|..+|+
T Consensus 15 ~~~~YY~~~~p~IsD~eYD~L~~eL~ 40 (441)
T smart00532 15 HDYRYYVLDAPIISDAEYDRLMRELK 40 (441)
T ss_pred HHHHHHhcCCCCCChHHHHHHHHHHH
Confidence 3456886 999999999999999997
No 8
>TIGR00575 dnlj DNA ligase, NAD-dependent. The member of this family from Treponema pallidum differs in having three rather than just one copy of the BRCT (BRCA1 C Terminus) domain (pfam00533) at the C-terminus. It is included in the seed.
Probab=92.69 E-value=0.12 Score=54.87 Aligned_cols=29 Identities=31% Similarity=0.556 Sum_probs=24.4
Q ss_pred HHHHHHhhhc-CCCccChHHHHHHHHHHhh
Q 019981 146 FLEASMAYVA-GKPIMSDEEYDKLKQKLKM 174 (333)
Q Consensus 146 fLEA~~AY~s-GkPimsDeeFD~LK~kLk~ 174 (333)
.-++-.+||. |+|+|||+|||.|.++|+.
T Consensus 5 l~~~~~~YY~~~~p~IsD~eYD~L~~~L~~ 34 (652)
T TIGR00575 5 IRHHDYRYYVLDEPSISDAEYDRLYRELQE 34 (652)
T ss_pred HHHHHHHHHhCCCCCCChHHHHHHHHHHHH
Confidence 3445667886 9999999999999999974
No 9
>PRK07956 ligA NAD-dependent DNA ligase LigA; Validated
Probab=92.59 E-value=0.12 Score=55.20 Aligned_cols=26 Identities=38% Similarity=0.629 Sum_probs=22.9
Q ss_pred HHHHhhh-cCCCccChHHHHHHHHHHh
Q 019981 148 EASMAYV-AGKPIMSDEEYDKLKQKLK 173 (333)
Q Consensus 148 EA~~AY~-sGkPimsDeeFD~LK~kLk 173 (333)
++-.+|| .|+|+|||+|||.|.++|+
T Consensus 18 ~~~~~YY~~~~p~IsD~eYD~L~~~L~ 44 (665)
T PRK07956 18 HHAYAYYVLDAPSISDAEYDRLYRELV 44 (665)
T ss_pred HHHHHHHhCCCCCCChHHHHHHHHHHH
Confidence 3445788 9999999999999999997
No 10
>TIGR01206 lysW lysine biosynthesis protein LysW. This very small, poorly characterized protein has been shown essential in Thermus thermophilus for an unusual pathway of Lys biosynthesis from aspartate by way of alpha-aminoadipate (AAA) rather than diaminopimelate. It is found also in Deinococcus radiodurans and Pyrococcus horikoshii, which appear to share the AAA pathway.
Probab=92.54 E-value=0.21 Score=37.88 Aligned_cols=44 Identities=27% Similarity=0.538 Sum_probs=28.4
Q ss_pred eecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEecc--ceeeecC
Q 019981 271 LKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSN--TRLITLP 328 (333)
Q Consensus 271 LKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk--~R~itl~ 328 (333)
|+..||.||+++-- +..-.--.+.|+. ||..|+.-++ .||-..|
T Consensus 1 ~~~~CP~CG~~iev--------~~~~~GeiV~Cp~------CGaeleVv~~~p~~L~~ap 46 (54)
T TIGR01206 1 MQFECPDCGAEIEL--------ENPELGELVICDE------CGAELEVVSLDPLRLEAAP 46 (54)
T ss_pred CccCCCCCCCEEec--------CCCccCCEEeCCC------CCCEEEEEeCCCCEEEeCc
Confidence 36789999998732 2222233678887 9999998755 4443333
No 11
>PRK00420 hypothetical protein; Validated
Probab=91.60 E-value=0.17 Score=43.44 Aligned_cols=46 Identities=17% Similarity=0.457 Sum_probs=32.9
Q ss_pred HHHHHhhhhcceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEe
Q 019981 257 SQSLTKLIVRESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 257 a~~Lt~l~~~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~ 319 (333)
++.+..++++=...|-..||.||.+.|.+-. ..+-|++ ||..+...
T Consensus 8 ~k~~a~~Ll~Ga~ml~~~CP~Cg~pLf~lk~-----------g~~~Cp~------Cg~~~~v~ 53 (112)
T PRK00420 8 VKKAAELLLKGAKMLSKHCPVCGLPLFELKD-----------GEVVCPV------HGKVYIVK 53 (112)
T ss_pred HHHHHHHHHhHHHHccCCCCCCCCcceecCC-----------CceECCC------CCCeeeec
Confidence 3445556666566688999999999887632 2567877 99977764
No 12
>TIGR02098 MJ0042_CXXC MJ0042 family finger-like domain. This domain contains a CXXCX(19)CXXC motif suggestive of both zinc fingers and thioredoxin, usually found at the N-terminus of prokaryotic proteins. One partially characterized gene, agmX, is among a large set in Myxococcus whose interruption affects adventurous gliding motility.
Probab=91.46 E-value=0.16 Score=34.41 Aligned_cols=36 Identities=25% Similarity=0.587 Sum_probs=22.9
Q ss_pred ecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEE
Q 019981 272 KGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVY 318 (333)
Q Consensus 272 KG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f 318 (333)
+-.||+||+.+..=-.. + .....+++|++ ||..+..
T Consensus 2 ~~~CP~C~~~~~v~~~~---~--~~~~~~v~C~~------C~~~~~~ 37 (38)
T TIGR02098 2 RIQCPNCKTSFRVVDSQ---L--GANGGKVRCGK------CGHVWYA 37 (38)
T ss_pred EEECCCCCCEEEeCHHH---c--CCCCCEEECCC------CCCEEEe
Confidence 45799999985422111 1 12223799999 9998764
No 13
>PF13240 zinc_ribbon_2: zinc-ribbon domain
Probab=89.97 E-value=0.14 Score=32.51 Aligned_cols=9 Identities=67% Similarity=1.497 Sum_probs=8.0
Q ss_pred CCCCccccc
Q 019981 275 CPNCGTENV 283 (333)
Q Consensus 275 CPNCGeEv~ 283 (333)
||+||.+|-
T Consensus 2 Cp~CG~~~~ 10 (23)
T PF13240_consen 2 CPNCGAEIE 10 (23)
T ss_pred CcccCCCCC
Confidence 999999874
No 14
>cd00114 LIGANc NAD+ dependent DNA ligase adenylation domain. DNA ligases catalyze the crucial step of joining the breaks in duplex DNA during DNA replication, repair and recombination, utilizing either ATP or NAD(+) as a cofactor, but using the same basic reaction mechanism. The enzyme reacts with the cofactor to form a phosphoamide-linked AMP with the amino group of a conserved Lysine in the KXDG motif, and subsequently transfers it to the DNA substrate to yield adenylated DNA. This alignment contains members of the NAD+ dependent subfamily only.
Probab=89.77 E-value=0.39 Score=46.73 Aligned_cols=35 Identities=26% Similarity=0.366 Sum_probs=27.8
Q ss_pred hHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhh
Q 019981 96 GELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELM 130 (333)
Q Consensus 96 ge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~ 130 (333)
-++...--.+=..||..+.+++||+|||.|.++|.
T Consensus 4 ~~L~~~i~~~~~~YY~~~~p~IsD~eYD~L~~~L~ 38 (307)
T cd00114 4 AELRELLNKHDYRYYVLDEPSVSDAEYDRLYRELR 38 (307)
T ss_pred HHHHHHHHHHHHHHHhCCCCCCChHHHHHHHHHHH
Confidence 34555555556678888999999999999999984
No 15
>PF06044 DRP: Dam-replacing family; InterPro: IPR010324 Dam-replacing protein (DRP) is a restriction endonuclease that is flanked by pseudo-transposable small repeat elements. The replacement of Dam-methylase by DRP allows phase variation through slippage-like mechanisms in several pathogenic isolates of Neisseria meningitidis [].; PDB: 4ESJ_A.
Probab=89.77 E-value=0.14 Score=49.40 Aligned_cols=51 Identities=25% Similarity=0.511 Sum_probs=24.7
Q ss_pred HHHHHHhhhhcceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccc
Q 019981 256 LSQSLTKLIVRESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNT 322 (333)
Q Consensus 256 ~a~~Lt~l~~~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~ 322 (333)
.|..||..| |+=.+.|||||.+..+=| +.|+....+.|++ |+...+..++.
T Consensus 19 ~aRVltE~W----v~~n~yCP~Cg~~~L~~f------~NN~PVaDF~C~~------C~eeyELKSk~ 69 (254)
T PF06044_consen 19 IARVLTEDW----VAENMYCPNCGSKPLSKF------ENNRPVADFYCPN------CNEEYELKSKK 69 (254)
T ss_dssp HHHHHHHHH----HHHH---TTT--SS-EE--------------EEE-TT------T--EEEEEEEE
T ss_pred hhHHHHHHH----HHHCCcCCCCCChhHhhc------cCCCccceeECCC------CchHHhhhhhc
Confidence 344455544 455678999999977666 5699999999999 99998888775
No 16
>PRK14351 ligA NAD-dependent DNA ligase LigA; Provisional
Probab=89.62 E-value=0.35 Score=51.96 Aligned_cols=25 Identities=28% Similarity=0.610 Sum_probs=21.7
Q ss_pred HHHhhh-cCCCccChHHHHHHHHHHh
Q 019981 149 ASMAYV-AGKPIMSDEEYDKLKQKLK 173 (333)
Q Consensus 149 A~~AY~-sGkPimsDeeFD~LK~kLk 173 (333)
+-.+|| .|+|+|||+|||.|.++|+
T Consensus 46 ~~~~YY~~~~p~IsD~eYD~L~~eL~ 71 (689)
T PRK14351 46 HDHRYYVEADPVIADRAYDALFARLQ 71 (689)
T ss_pred HHHHHHhCCCCCCChHHHHHHHHHHH
Confidence 345687 5799999999999999997
No 17
>PRK00415 rps27e 30S ribosomal protein S27e; Reviewed
Probab=89.56 E-value=0.27 Score=38.18 Aligned_cols=41 Identities=32% Similarity=0.735 Sum_probs=31.8
Q ss_pred eeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEecccee
Q 019981 269 LILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRL 324 (333)
Q Consensus 269 ~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~ 324 (333)
--|+--||.|+.|...| ..++..|+|.+ ||+.|.--+..++
T Consensus 8 ~F~~VkCp~C~n~q~vF---------sha~t~V~C~~------Cg~~L~~PtGGKa 48 (59)
T PRK00415 8 RFLKVKCPDCGNEQVVF---------SHASTVVRCLV------CGKTLAEPTGGKA 48 (59)
T ss_pred eEEEEECCCCCCeEEEE---------ecCCcEEECcc------cCCCcccCCCcce
Confidence 34677899999999888 45777899988 9999876554443
No 18
>PF13248 zf-ribbon_3: zinc-ribbon domain
Probab=89.56 E-value=0.16 Score=32.73 Aligned_cols=10 Identities=60% Similarity=1.168 Sum_probs=8.2
Q ss_pred cCCCCCcccc
Q 019981 273 GPCPNCGTEN 282 (333)
Q Consensus 273 G~CPNCGeEv 282 (333)
-.|||||.+|
T Consensus 3 ~~Cp~Cg~~~ 12 (26)
T PF13248_consen 3 MFCPNCGAEI 12 (26)
T ss_pred CCCcccCCcC
Confidence 3699999976
No 19
>COG2888 Predicted Zn-ribbon RNA-binding protein with a function in translation [Translation, ribosomal structure and biogenesis]
Probab=89.27 E-value=0.17 Score=39.58 Aligned_cols=34 Identities=26% Similarity=0.550 Sum_probs=21.6
Q ss_pred ceeeeecCCCCCccccceeccccccccCCCCccceeCCC
Q 019981 267 ESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSN 305 (333)
Q Consensus 267 D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~ 305 (333)
+--+.+=+||||||++-- + .-....-.|..+|++
T Consensus 22 ~e~~v~F~CPnCGe~~I~--R---c~~CRk~g~~Y~Cp~ 55 (61)
T COG2888 22 GETAVKFPCPNCGEVEIY--R---CAKCRKLGNPYRCPK 55 (61)
T ss_pred CCceeEeeCCCCCceeee--h---hhhHHHcCCceECCC
Confidence 334677899999965431 1 113445567889988
No 20
>PF01653 DNA_ligase_aden: NAD-dependent DNA ligase adenylation domain; InterPro: IPR013839 DNA ligase (polydeoxyribonucleotide synthase) is the enzyme that joins two DNA fragments by catalyzing the formation of an internucleotide ester bond between phosphate and deoxyribose. It is active during DNA replication, DNA repair and DNA recombination. There are two forms of DNA ligase: one requires ATP (6.5.1.1 from EC), the other NAD (6.5.1.2 from EC). This entry represents the N-terminal adenylation domain of NAD-dependent DNA ligases. These are proteins of about 75 to 85 Kd whose sequence is well conserved [, ]. They also show similarity to yicF, an Escherichia coli hypothetical protein of 63 Kd. Despite a complete lack of detectable sequence similarity, the fold of the central core of this adenyaltion domain shares homology with the equivalent region of ATP-dependent DNA ligases [, ].; GO: 0003911 DNA ligase (NAD+) activity; PDB: 1ZAU_A 3SGI_A 1B04_A 3JSL_A 3JSN_A 1DGS_A 1V9P_A 3PN1_A 3BAC_A 3UQ8_A ....
Probab=87.71 E-value=0.7 Score=45.12 Aligned_cols=36 Identities=33% Similarity=0.553 Sum_probs=28.3
Q ss_pred hhHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhh
Q 019981 95 LGELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELM 130 (333)
Q Consensus 95 lge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~ 130 (333)
+-++...-..+=..||..+.++|||+|||.|.++|.
T Consensus 7 i~~L~~~i~~~~~~YY~~~~p~isD~eYD~l~~~L~ 42 (315)
T PF01653_consen 7 IEELRKEINRHNYAYYNLGEPIISDAEYDQLFRELK 42 (315)
T ss_dssp HHHHHHHHHHHHHHHHTTSSSSSSHHHHHHHHHHHH
T ss_pred HHHHHHHHHHHHHHHhcCCCCCCCHHHHHHHHHHHH
Confidence 344555555666789999999999999999998863
No 21
>smart00532 LIGANc Ligase N family.
Probab=87.56 E-value=0.59 Score=47.82 Aligned_cols=36 Identities=25% Similarity=0.420 Sum_probs=28.7
Q ss_pred hhHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhh
Q 019981 95 LGELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELM 130 (333)
Q Consensus 95 lge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~ 130 (333)
+-++...-..+-..||..+.+++||+|||.|.++|.
T Consensus 5 i~~L~~~i~~~~~~YY~~~~p~IsD~eYD~L~~eL~ 40 (441)
T smart00532 5 ISELRKLLNKHDYRYYVLDAPIISDAEYDRLMRELK 40 (441)
T ss_pred HHHHHHHHHHHHHHHHhcCCCCCChHHHHHHHHHHH
Confidence 445555555566678889999999999999999995
No 22
>PRK14350 ligA NAD-dependent DNA ligase LigA; Provisional
Probab=87.17 E-value=0.67 Score=49.76 Aligned_cols=24 Identities=29% Similarity=0.463 Sum_probs=20.8
Q ss_pred HHhhh-cCCCccChHHHHHHHHHHh
Q 019981 150 SMAYV-AGKPIMSDEEYDKLKQKLK 173 (333)
Q Consensus 150 ~~AY~-sGkPimsDeeFD~LK~kLk 173 (333)
-.+|| .|+|+|||++||.|..+|+
T Consensus 20 ~~~YY~~~~p~IsD~~YD~L~~eL~ 44 (669)
T PRK14350 20 DKEYYVDSSPSVEDFTYDKALLRLQ 44 (669)
T ss_pred HHHHHhCCCCCCChHHHHHHHHHHH
Confidence 35677 5799999999999999996
No 23
>PRK14890 putative Zn-ribbon RNA-binding protein; Provisional
Probab=86.98 E-value=0.32 Score=37.81 Aligned_cols=30 Identities=27% Similarity=0.477 Sum_probs=19.2
Q ss_pred eeecCCCCCccc-cceeccccccccCCCCccceeCCC
Q 019981 270 ILKGPCPNCGTE-NVSFFGTILSISSGGTTNTINCSN 305 (333)
Q Consensus 270 iLKG~CPNCGeE-v~aFfgtilsv~s~~~~n~vkC~~ 305 (333)
+.+=.|||||++ +.-=- .-..-.|.++|++
T Consensus 23 ~~~F~CPnCG~~~I~RC~------~CRk~~~~Y~CP~ 53 (59)
T PRK14890 23 AVKFLCPNCGEVIIYRCE------KCRKQSNPYTCPK 53 (59)
T ss_pred cCEeeCCCCCCeeEeech------hHHhcCCceECCC
Confidence 455689999998 44321 2344557888888
No 24
>PF01667 Ribosomal_S27e: Ribosomal protein S27; InterPro: IPR000592 Ribosomes are the particles that catalyse mRNA-directed protein synthesis in all organisms. The codons of the mRNA are exposed on the ribosome to allow tRNA binding. This leads to the incorporation of amino acids into the growing polypeptide chain in accordance with the genetic information. Incoming amino acid monomers enter the ribosomal A site in the form of aminoacyl-tRNAs complexed with elongation factor Tu (EF-Tu) and GTP. The growing polypeptide chain, situated in the P site as peptidyl-tRNA, is then transferred to aminoacyl-tRNA and the new peptidyl-tRNA, extended by one residue, is translocated to the P site with the aid the elongation factor G (EF-G) and GTP as the deacylated tRNA is released from the ribosome through one or more exit sites [, ]. About 2/3 of the mass of the ribosome consists of RNA and 1/3 of protein. The proteins are named in accordance with the subunit of the ribosome which they belong to - the small (S1 to S31) and the large (L1 to L44). Usually they decorate the rRNA cores of the subunits. Many ribosomal proteins, particularly those of the large subunit, are composed of a globular, surfaced-exposed domain with long finger-like projections that extend into the rRNA core to stabilise its structure. Most of the proteins interact with multiple RNA elements, often from different domains. In the large subunit, about 1/3 of the 23S rRNA nucleotides are at least in van der Waal's contact with protein, and L22 interacts with all six domains of the 23S rRNA. Proteins S4 and S7, which initiate assembly of the 16S rRNA, are located at junctions of five and four RNA helices, respectively. In this way proteins serve to organise and stabilise the rRNA tertiary structure. While the crucial activities of decoding and peptide transfer are RNA based, proteins play an active role in functions that may have evolved to streamline the process of protein synthesis. In addition to their function in the ribosome, many ribosomal proteins have some function 'outside' the ribosome [, ]. A number of eukaryotic and archaeal ribosomal proteins can be grouped on the basis of sequence similarities. One of these families include mammalian, yeast, Chlamydomonas reinhardtii and Entamoeba histolytica S27, and Methanocaldococcus jannaschii (Methanococcus jannaschii) MJ0250 []. These proteins have from 62 to 87 amino acids. They contain, in their central section, a putative zinc-finger region of the type C-x(2)-C-x(14)-C-x(2)-C.; GO: 0003735 structural constituent of ribosome, 0006412 translation, 0005622 intracellular, 0005840 ribosome; PDB: 1QXF_A 3IZ6_X 2XZN_6 2XZM_6 3U5G_b 3IZB_X 3U5C_b.
Probab=86.34 E-value=1.1 Score=34.32 Aligned_cols=42 Identities=19% Similarity=0.508 Sum_probs=27.4
Q ss_pred eeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccceeee
Q 019981 270 ILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRLIT 326 (333)
Q Consensus 270 iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~it 326 (333)
-|+--||.|+++...| ..++..|+|.+ |++.|.--+..++..
T Consensus 5 Fm~VkCp~C~~~q~vF---------Sha~t~V~C~~------Cg~~L~~PtGGKa~l 46 (55)
T PF01667_consen 5 FMDVKCPGCYNIQTVF---------SHAQTVVKCVV------CGTVLAQPTGGKARL 46 (55)
T ss_dssp EEEEE-TTT-SEEEEE---------TT-SS-EE-SS------STSEEEEE-SSSEEE
T ss_pred EEEEECCCCCCeeEEE---------ecCCeEEEccc------CCCEecCCCCcCeEE
Confidence 3566799999999887 55777899988 999998776654443
No 25
>smart00531 TFIIE Transcription initiation factor IIE.
Probab=86.08 E-value=0.43 Score=41.41 Aligned_cols=55 Identities=24% Similarity=0.453 Sum_probs=32.0
Q ss_pred HHhhhhcceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEecccee
Q 019981 260 LTKLIVRESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRL 324 (333)
Q Consensus 260 Lt~l~~~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~ 324 (333)
|..-+..+.--..=-|||||.. |+|-- .... +..+.++.|++ ||..|++.-....
T Consensus 87 L~~~l~~e~~~~~Y~Cp~C~~~-y~~~e-a~~~--~d~~~~f~Cp~------Cg~~l~~~dn~~~ 141 (147)
T smart00531 87 LEDKLEDETNNAYYKCPNCQSK-YTFLE-ANQL--LDMDGTFTCPR------CGEELEEDDNSEP 141 (147)
T ss_pred HHHHHhcccCCcEEECcCCCCE-eeHHH-HHHh--cCCCCcEECCC------CCCEEEEcCchhh
Confidence 4444433333334459999954 44532 2211 12355699999 9999999865444
No 26
>PRK08097 ligB NAD-dependent DNA ligase LigB; Reviewed
Probab=85.76 E-value=0.76 Score=48.55 Aligned_cols=36 Identities=28% Similarity=0.535 Sum_probs=28.6
Q ss_pred hhHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhh
Q 019981 95 LGELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELM 130 (333)
Q Consensus 95 lge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~ 130 (333)
+-++...-..+=..||..+.|++||+|||.|.+||.
T Consensus 34 i~~L~~~l~~~~~~YY~~~~p~IsD~eYD~L~~eL~ 69 (562)
T PRK08097 34 IAALQQQLAQWDDAYWRQGKSEVDDEVYDQLRARLT 69 (562)
T ss_pred HHHHHHHHHHHHHHHHhCCCCCCChHHHHHHHHHHH
Confidence 445555555556679999999999999999999985
No 27
>PF08271 TF_Zn_Ribbon: TFIIB zinc-binding; InterPro: IPR013137 Zinc finger (Znf) domains are relatively small protein motifs which contain multiple finger-like protrusions that make tandem contacts with their target molecule. Some of these domains bind zinc, but many do not; instead binding other metals such as iron, or no metal at all. For example, some family members form salt bridges to stabilise the finger-like folds. They were first identified as a DNA-binding motif in transcription factor TFIIIA from Xenopus laevis (African clawed frog), however they are now recognised to bind DNA, RNA, protein and/or lipid substrates [, , , , ]. Their binding properties depend on the amino acid sequence of the finger domains and of the linker between fingers, as well as on the higher-order structures and the number of fingers. Znf domains are often found in clusters, where fingers can have different binding specificities. There are many superfamilies of Znf motifs, varying in both sequence and structure. They display considerable versatility in binding modes, even between members of the same class (e.g. some bind DNA, others protein), suggesting that Znf motifs are stable scaffolds that have evolved specialised functions. For example, Znf-containing proteins function in gene transcription, translation, mRNA trafficking, cytoskeleton organisation, epithelial development, cell adhesion, protein folding, chromatin remodelling and zinc sensing, to name but a few []. Zinc-binding motifs are stable structures, and they rarely undergo conformational changes upon binding their target. This entry represents a zinc finger motif found in transcription factor IIB (TFIIB). In eukaryotes the initiation of transcription of protein encoding genes by the polymerase II complexe (Pol II) is modulated by general and specific transcription factors. The general transcription factors operate through common promoters elements (such as the TATA box). At least seven different proteins associate to form the general transcription factors: TFIIA, -IIB, -IID, -IIE, -IIF, -IIG, and -IIH []. TFIIB and TFIID are responsible for promoter recognition and interaction with pol II; together with Pol II, they form a minimal initiation complex capable of transcription under certain conditions. The TATA box of a Pol II promoter is bound in the initiation complex by the TBP subunit of TFIID, which bends the DNA around the C-terminal domain of TFIIB whereas the N-terminal zinc finger of TFIIB interacts with Pol II [, ]. The TFIIB zinc finger adopts a zinc ribbon fold characterised by two beta-hairpins forming two structurally similar zinc-binding sub-sites []. The zinc finger contacts the rbp1 subunit of Pol II through its dock domain, a conserved region of about 70 amino acids located close to the polymerase active site []. In the Pol II complex this surface is located near the RNA exit groove. Interestingly this sequence is best conserved in the three polymerases that utilise a TFIIB-like general transcription factor (Pol II, Pol III, and archaeal RNA polymerase) but not in Pol I []. More information about these proteins can be found at Protein of the Month: Zinc Fingers [].; GO: 0008270 zinc ion binding, 0006355 regulation of transcription, DNA-dependent; PDB: 1VD4_A 1PFT_A 3K1F_M 3K7A_M 1RO4_A 1RLY_A 1DL6_A.
Probab=85.03 E-value=0.58 Score=32.90 Aligned_cols=32 Identities=31% Similarity=0.828 Sum_probs=22.1
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEec
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDS 320 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~t 320 (333)
-||+||... ..+ .....++-|.+ ||..++.+.
T Consensus 2 ~Cp~Cg~~~-~~~--------D~~~g~~vC~~------CG~Vl~e~~ 33 (43)
T PF08271_consen 2 KCPNCGSKE-IVF--------DPERGELVCPN------CGLVLEENI 33 (43)
T ss_dssp SBTTTSSSE-EEE--------ETTTTEEEETT------T-BBEE-TT
T ss_pred CCcCCcCCc-eEE--------cCCCCeEECCC------CCCEeeccc
Confidence 499999977 233 34566778998 999988663
No 28
>PRK05978 hypothetical protein; Provisional
Probab=84.97 E-value=0.44 Score=42.58 Aligned_cols=34 Identities=29% Similarity=0.839 Sum_probs=23.3
Q ss_pred eecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEe
Q 019981 271 LKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 271 LKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~ 319 (333)
|+|-||+||+.-. |++-+.+ +-+|++ ||..+++.
T Consensus 32 l~grCP~CG~G~L--F~g~Lkv-------~~~C~~------CG~~~~~~ 65 (148)
T PRK05978 32 FRGRCPACGEGKL--FRAFLKP-------VDHCAA------CGEDFTHH 65 (148)
T ss_pred HcCcCCCCCCCcc--ccccccc-------CCCccc------cCCccccC
Confidence 6899999998643 3333322 446776 99988775
No 29
>PRK14351 ligA NAD-dependent DNA ligase LigA; Provisional
Probab=84.53 E-value=1.1 Score=48.35 Aligned_cols=36 Identities=19% Similarity=0.390 Sum_probs=30.3
Q ss_pred hhHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhh
Q 019981 95 LGELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELM 130 (333)
Q Consensus 95 lge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~ 130 (333)
+-++...--.+=..||..+.+++||+|||.|.++|.
T Consensus 36 i~~L~~~i~~~~~~YY~~~~p~IsD~eYD~L~~eL~ 71 (689)
T PRK14351 36 AEQLREAIREHDHRYYVEADPVIADRAYDALFARLQ 71 (689)
T ss_pred HHHHHHHHHHHHHHHHhCCCCCCChHHHHHHHHHHH
Confidence 556666666666789999999999999999999996
No 30
>PRK07956 ligA NAD-dependent DNA ligase LigA; Validated
Probab=84.50 E-value=1.1 Score=48.01 Aligned_cols=36 Identities=31% Similarity=0.504 Sum_probs=28.2
Q ss_pred hhHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhh
Q 019981 95 LGELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELM 130 (333)
Q Consensus 95 lge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~ 130 (333)
+-++...--.+=..||..+.+++||+|||.|.++|.
T Consensus 9 i~~L~~~i~~~~~~YY~~~~p~IsD~eYD~L~~~L~ 44 (665)
T PRK07956 9 IEELREELNHHAYAYYVLDAPSISDAEYDRLYRELV 44 (665)
T ss_pred HHHHHHHHHHHHHHHHhCCCCCCChHHHHHHHHHHH
Confidence 344444555555678889999999999999999986
No 31
>PF14255 Cys_rich_CPXG: Cysteine-rich CPXCG
Probab=83.63 E-value=0.85 Score=34.38 Aligned_cols=37 Identities=24% Similarity=0.506 Sum_probs=25.9
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEec
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDS 320 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~t 320 (333)
.||.||+.+-...-. |.+++.-+ -||.||-+++.|.-
T Consensus 2 ~CPyCge~~~~~iD~-----s~~~Q~yi-----EDC~vCC~PI~~~v 38 (52)
T PF14255_consen 2 QCPYCGEPIEILIDP-----SAGDQEYI-----EDCQVCCRPIEVQV 38 (52)
T ss_pred CCCCCCCeeEEEEec-----CCCCeeEE-----eehhhcCCccEEEE
Confidence 599999999988743 33333222 35667999999873
No 32
>PHA00626 hypothetical protein
Probab=81.82 E-value=1.1 Score=34.92 Aligned_cols=35 Identities=31% Similarity=0.722 Sum_probs=22.9
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEe
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~ 319 (333)
.||+||..+..==|.+ +.-.++.+|+. ||-..+-|
T Consensus 2 ~CP~CGS~~Ivrcg~c-----r~~snrYkCkd------CGY~ft~~ 36 (59)
T PHA00626 2 SCPKCGSGNIAKEKTM-----RGWSDDYVCCD------CGYNDSKD 36 (59)
T ss_pred CCCCCCCceeeeecee-----cccCcceEcCC------CCCeechh
Confidence 5999999655433332 34467888877 98765544
No 33
>smart00659 RPOLCX RNA polymerase subunit CX. present in RNA polymerase I, II and III
Probab=81.55 E-value=1.9 Score=31.26 Aligned_cols=34 Identities=32% Similarity=0.819 Sum_probs=25.1
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEe-cccee
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYD-SNTRL 324 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~-tk~R~ 324 (333)
-|.+||.||..- ....++|++ ||..+.|- ++.|.
T Consensus 4 ~C~~Cg~~~~~~-----------~~~~irC~~------CG~rIlyK~R~~~~ 38 (44)
T smart00659 4 ICGECGRENEIK-----------SKDVVRCRE------CGYRILYKKRTKRL 38 (44)
T ss_pred ECCCCCCEeecC-----------CCCceECCC------CCceEEEEeCCCce
Confidence 499999997632 235689999 99999987 34443
No 34
>PRK02935 hypothetical protein; Provisional
Probab=81.24 E-value=1.6 Score=37.77 Aligned_cols=36 Identities=17% Similarity=0.468 Sum_probs=25.7
Q ss_pred eeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccce
Q 019981 270 ILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTR 323 (333)
Q Consensus 270 iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R 323 (333)
+..-.||||+.+..--=+. . -|-+ |+++|+.|...+
T Consensus 68 avqV~CP~C~K~TKmLGrv---------D---~CM~------C~~PLTLd~~le 103 (110)
T PRK02935 68 AVQVICPSCEKPTKMLGRV---------D---ACMH------CNQPLTLDRSLE 103 (110)
T ss_pred ceeeECCCCCchhhhccce---------e---ecCc------CCCcCCcCcccc
Confidence 3444899999988754222 1 5777 999999987654
No 35
>TIGR00575 dnlj DNA ligase, NAD-dependent. The member of this family from Treponema pallidum differs in having three rather than just one copy of the BRCT (BRCA1 C Terminus) domain (pfam00533) at the C-terminus. It is included in the seed.
Probab=80.61 E-value=1.5 Score=46.79 Aligned_cols=27 Identities=26% Similarity=0.411 Sum_probs=23.0
Q ss_pred HHHHHhhhcCccccChHHHhhhHhhhh
Q 019981 104 QALQAFYYEGKAVMSNEEFDNLKEELM 130 (333)
Q Consensus 104 ~Al~~fY~~gk~~~sdeefd~LkEeL~ 130 (333)
.+-..||..+.|++||+|||.|.++|.
T Consensus 7 ~~~~~YY~~~~p~IsD~eYD~L~~~L~ 33 (652)
T TIGR00575 7 HHDYRYYVLDEPSISDAEYDRLYRELQ 33 (652)
T ss_pred HHHHHHHhCCCCCCChHHHHHHHHHHH
Confidence 344568888999999999999999985
No 36
>PLN00209 ribosomal protein S27; Provisional
Probab=80.25 E-value=2 Score=35.77 Aligned_cols=45 Identities=16% Similarity=0.408 Sum_probs=34.5
Q ss_pred cceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccceee
Q 019981 266 RESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRLI 325 (333)
Q Consensus 266 ~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~i 325 (333)
.|---|+--||.|+.+...| ..++..|.|.+ ||+.|.--+..+..
T Consensus 30 PnS~Fm~VkCp~C~n~q~VF---------ShA~t~V~C~~------Cg~~L~~PTGGKa~ 74 (86)
T PLN00209 30 PNSFFMDVKCQGCFNITTVF---------SHSQTVVVCGS------CQTVLCQPTGGKAR 74 (86)
T ss_pred CCCEEEEEECCCCCCeeEEE---------ecCceEEEccc------cCCEeeccCCCCeE
Confidence 34456788899999999888 45777889988 99999766654443
No 37
>TIGR00373 conserved hypothetical protein TIGR00373. This family of proteins is, so far, restricted to archaeal genomes. The family appears to be distantly related to the N-terminal region of the eukaryotic transcription initiation factor IIE alpha chain.
Probab=79.79 E-value=0.46 Score=42.06 Aligned_cols=54 Identities=22% Similarity=0.398 Sum_probs=33.3
Q ss_pred HHHHHhhhhcceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccceee
Q 019981 257 SQSLTKLIVRESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRLI 325 (333)
Q Consensus 257 a~~Lt~l~~~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~i 325 (333)
...|...+..+.--.-=-||||+. -++|--.+ .+...|++ ||..|++.-++.+|
T Consensus 94 ~~~lk~~l~~e~~~~~Y~Cp~c~~-r~tf~eA~--------~~~F~Cp~------Cg~~L~~~dn~~~i 147 (158)
T TIGR00373 94 AKKLREKLEFETNNMFFICPNMCV-RFTFNEAM--------ELNFTCPR------CGAMLDYLDNSEAI 147 (158)
T ss_pred HHHHHHHHhhccCCCeEECCCCCc-EeeHHHHH--------HcCCcCCC------CCCEeeeccCHHHH
Confidence 334444444433333345999994 35665331 24678887 99999998776655
No 38
>PRK14350 ligA NAD-dependent DNA ligase LigA; Provisional
Probab=79.18 E-value=1.7 Score=46.77 Aligned_cols=36 Identities=11% Similarity=0.200 Sum_probs=27.1
Q ss_pred hhHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhh
Q 019981 95 LGELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELM 130 (333)
Q Consensus 95 lge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~ 130 (333)
+-++...--..=..||..+.|++||+|||.|.+||.
T Consensus 9 i~~L~~~i~~~~~~YY~~~~p~IsD~~YD~L~~eL~ 44 (669)
T PRK14350 9 ILDLKKLIRKWDKEYYVDSSPSVEDFTYDKALLRLQ 44 (669)
T ss_pred HHHHHHHHHHHHHHHHhCCCCCCChHHHHHHHHHHH
Confidence 444444444445568889999999999999999984
No 39
>COG0272 Lig NAD-dependent DNA ligase (contains BRCT domain type II) [DNA replication, recombination, and repair]
Probab=78.88 E-value=2.1 Score=46.34 Aligned_cols=24 Identities=38% Similarity=0.635 Sum_probs=21.1
Q ss_pred Hhhhc-CCCccChHHHHHHHHHHhh
Q 019981 151 MAYVA-GKPIMSDEEYDKLKQKLKM 174 (333)
Q Consensus 151 ~AY~s-GkPimsDeeFD~LK~kLk~ 174 (333)
.+||- ++|+|+|+|||+|.++|..
T Consensus 23 ~~Yyv~d~P~VsD~eYD~L~reL~~ 47 (667)
T COG0272 23 YRYYVLDAPSVSDAEYDQLYRELQE 47 (667)
T ss_pred HHHhccCCCCCChHHHHHHHHHHHH
Confidence 46666 9999999999999999975
No 40
>PF07754 DUF1610: Domain of unknown function (DUF1610); InterPro: IPR011668 This domain is found in archaeal species. It is likely to bind zinc via its four well-conserved cysteine residues.
Probab=78.76 E-value=0.88 Score=29.71 Aligned_cols=11 Identities=55% Similarity=1.259 Sum_probs=8.3
Q ss_pred eeecCCCCCcc
Q 019981 270 ILKGPCPNCGT 280 (333)
Q Consensus 270 iLKG~CPNCGe 280 (333)
+..=+|||||+
T Consensus 14 ~v~f~CPnCG~ 24 (24)
T PF07754_consen 14 AVPFPCPNCGF 24 (24)
T ss_pred CceEeCCCCCC
Confidence 44568999996
No 41
>PTZ00083 40S ribosomal protein S27; Provisional
Probab=78.64 E-value=2.4 Score=35.19 Aligned_cols=45 Identities=16% Similarity=0.477 Sum_probs=34.7
Q ss_pred cceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccceee
Q 019981 266 RESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRLI 325 (333)
Q Consensus 266 ~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~i 325 (333)
.|---|+--||.|+.+...| ..++..|.|.+ |++.|.--+..++.
T Consensus 29 PnS~Fm~VkCp~C~n~q~VF---------ShA~t~V~C~~------Cg~~L~~PTGGKa~ 73 (85)
T PTZ00083 29 PNSYFMDVKCPGCSQITTVF---------SHAQTVVLCGG------CSSQLCQPTGGKAK 73 (85)
T ss_pred CCCeEEEEECCCCCCeeEEE---------ecCceEEEccc------cCCEeeccCCCCeE
Confidence 34456788899999999888 45777889988 99999766655443
No 42
>PF05876 Terminase_GpA: Phage terminase large subunit (GpA); InterPro: IPR008866 This entry is represented by Bacteriophage lambda, GpA. The characteristics of the protein distribution suggest prophage matches in addition to the phage matches. This entry consists of several phage terminase large subunit proteins as well as related sequences from several bacterial species. The DNA packaging enzyme of bacteriophage lambda, terminase, is a heteromultimer composed of a small subunit, gpNu1, and a large subunit, gpA, products of the Nu1 and A genes, respectively. Terminase is involved in the site-specific binding and cutting of the DNA in the initial stages of packaging. It is now known that gpA is actively involved in late stages of packaging, including DNA translocation, and that this enzyme contains separate functional domains for its early and late packaging activities [].
Probab=78.30 E-value=1.3 Score=46.13 Aligned_cols=54 Identities=22% Similarity=0.407 Sum_probs=38.1
Q ss_pred hhcceeeeecCCCCCccccceeccccccccC-CCCccceeCCCCccccccCceeEEeccce
Q 019981 264 IVRESLILKGPCPNCGTENVSFFGTILSISS-GGTTNTINCSNLTFCFSCGTTMVYDSNTR 323 (333)
Q Consensus 264 ~~~D~~iLKG~CPNCGeEv~aFfgtilsv~s-~~~~n~vkC~~~aeCHVC~t~L~f~tk~R 323 (333)
...|---..-|||+||++..-=|..+.--++ ...+..+.|++ ||+.++=.-+.+
T Consensus 192 ~~sdqr~~~vpCPhCg~~~~l~~~~l~w~~~~~~~~a~y~C~~------Cg~~i~e~~k~~ 246 (557)
T PF05876_consen 192 EESDQRRYYVPCPHCGEEQVLEWENLKWDKGEAPETARYVCPH------CGCEIEEHDKRR 246 (557)
T ss_pred HhCCceEEEccCCCCCCCccccccceeecCCCCccceEEECCC------CcCCCCHHHHhh
Confidence 4566667788999999987755666543322 45677888888 999887654444
No 43
>PF10571 UPF0547: Uncharacterised protein family UPF0547; InterPro: IPR018886 This domain may well be a type of zinc-finger as it carries two pairs of highly conserved cysteine residues though with no accompanying histidines. Several members are annotated as putative helicases.
Probab=78.10 E-value=1.2 Score=29.17 Aligned_cols=9 Identities=56% Similarity=1.486 Sum_probs=8.2
Q ss_pred CCCCCcccc
Q 019981 274 PCPNCGTEN 282 (333)
Q Consensus 274 ~CPNCGeEv 282 (333)
.||+|+++|
T Consensus 2 ~CP~C~~~V 10 (26)
T PF10571_consen 2 TCPECGAEV 10 (26)
T ss_pred cCCCCcCCc
Confidence 499999998
No 44
>PRK09710 lar restriction alleviation and modification protein; Reviewed
Probab=77.01 E-value=1.7 Score=34.39 Aligned_cols=34 Identities=26% Similarity=0.617 Sum_probs=22.1
Q ss_pred cCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEe
Q 019981 273 GPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 273 G~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~ 319 (333)
-|||.||.++...= . .+..-.++|+. |+..-.|.
T Consensus 7 KPCPFCG~~~~~v~-~------~~g~~~v~C~~------CgA~~~~~ 40 (64)
T PRK09710 7 KPCPFCGCPSVTVK-A------ISGYYRAKCNG------CESRTGYG 40 (64)
T ss_pred cCCCCCCCceeEEE-e------cCceEEEEcCC------CCcCcccc
Confidence 39999999987542 1 12233456655 99876665
No 45
>PF14803 Nudix_N_2: Nudix N-terminal; PDB: 3CNG_C.
Probab=76.44 E-value=1.8 Score=30.09 Aligned_cols=29 Identities=31% Similarity=0.785 Sum_probs=15.4
Q ss_pred CCCCccccceeccccccccCCCCccceeCCCCccccccCce
Q 019981 275 CPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTT 315 (333)
Q Consensus 275 CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~ 315 (333)
||+||.++- +.|+.+.+..+.-|++ ||..
T Consensus 3 C~~CG~~l~------~~ip~gd~r~R~vC~~------Cg~I 31 (34)
T PF14803_consen 3 CPQCGGPLE------RRIPEGDDRERLVCPA------CGFI 31 (34)
T ss_dssp -TTT--B-E------EE--TT-SS-EEEETT------TTEE
T ss_pred cccccChhh------hhcCCCCCccceECCC------CCCE
Confidence 999999963 2345667778888887 8853
No 46
>PRK06266 transcription initiation factor E subunit alpha; Validated
Probab=75.98 E-value=0.62 Score=42.13 Aligned_cols=55 Identities=24% Similarity=0.393 Sum_probs=34.2
Q ss_pred HHHHHHhhhhcceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccceee
Q 019981 256 LSQSLTKLIVRESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRLI 325 (333)
Q Consensus 256 ~a~~Lt~l~~~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~i 325 (333)
+...|...+..+.--.-=-||||+.. |+|.-. -.+...|++ ||..|++.-++..|
T Consensus 101 ~~~klk~~l~~e~~~~~Y~Cp~C~~r-ytf~eA--------~~~~F~Cp~------Cg~~L~~~dn~~~~ 155 (178)
T PRK06266 101 ELKKLKEQLEEEENNMFFFCPNCHIR-FTFDEA--------MEYGFRCPQ------CGEMLEEYDNSELI 155 (178)
T ss_pred HHHHHHHHhhhccCCCEEECCCCCcE-EeHHHH--------hhcCCcCCC------CCCCCeecccHHHH
Confidence 34445555444433344459999944 566532 235678887 99999998665544
No 47
>PF14353 CpXC: CpXC protein
Probab=75.08 E-value=2.9 Score=34.98 Aligned_cols=40 Identities=25% Similarity=0.586 Sum_probs=25.5
Q ss_pred CCCCCccccceeccccccccC---------CCCccceeCCCCccccccCceeEEe
Q 019981 274 PCPNCGTENVSFFGTILSISS---------GGTTNTINCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s---------~~~~n~vkC~~~aeCHVC~t~L~f~ 319 (333)
.||+||++...=+=++..... +++-+.+.|++ ||.....+
T Consensus 3 tCP~C~~~~~~~v~~~I~~~~~p~l~e~il~g~l~~~~CP~------Cg~~~~~~ 51 (128)
T PF14353_consen 3 TCPHCGHEFEFEVWTSINADEDPELKEKILDGSLFSFTCPS------CGHKFRLE 51 (128)
T ss_pred CCCCCCCeeEEEEEeEEcCcCCHHHHHHHHcCCcCEEECCC------CCCceecC
Confidence 699999986644433222111 45567899999 99865443
No 48
>PF09851 SHOCT: Short C-terminal domain; InterPro: IPR018649 This family of hypothetical prokaryotic proteins has no known function.
Probab=74.63 E-value=4.8 Score=26.98 Aligned_cols=25 Identities=40% Similarity=0.590 Sum_probs=20.0
Q ss_pred HHHHHh-hhcCCCccChHHHHHHHHHHh
Q 019981 147 LEASMA-YVAGKPIMSDEEYDKLKQKLK 173 (333)
Q Consensus 147 LEA~~A-Y~sGkPimsDeeFD~LK~kLk 173 (333)
|+.+.. |-+| +||++||++.|.+|.
T Consensus 5 L~~L~~l~~~G--~IseeEy~~~k~~ll 30 (31)
T PF09851_consen 5 LEKLKELYDKG--EISEEEYEQKKARLL 30 (31)
T ss_pred HHHHHHHHHcC--CCCHHHHHHHHHHHh
Confidence 555555 7777 799999999999884
No 49
>PF14354 Lar_restr_allev: Restriction alleviation protein Lar
Probab=74.09 E-value=2.6 Score=30.92 Aligned_cols=33 Identities=30% Similarity=0.652 Sum_probs=20.4
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCc
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGT 314 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t 314 (333)
|||=||.+........... .+....|.|++ ||.
T Consensus 5 PCPFCG~~~~~~~~~~~~~--~~~~~~V~C~~------Cga 37 (61)
T PF14354_consen 5 PCPFCGSADVLIRQDEGFD--YGMYYYVECTD------CGA 37 (61)
T ss_pred CCCCCCCcceEeecccCCC--CCCEEEEEcCC------CCC
Confidence 8999998887765431100 00005677777 988
No 50
>TIGR03831 YgiT_finger YgiT-type zinc finger domain. This domain model describes a small domain with two copies of a putative zinc-binding motif CXXC (usually CXXCG). Most member proteins consist largely of this domain or else carry an additional C-terminal helix-turn-helix domain, resembling that of the phage protein Cro and modeled by pfam01381.
Probab=72.68 E-value=2.1 Score=29.22 Aligned_cols=19 Identities=37% Similarity=1.064 Sum_probs=13.0
Q ss_pred cceeeeec-C---CCCCccccce
Q 019981 266 RESLILKG-P---CPNCGTENVS 284 (333)
Q Consensus 266 ~D~~iLKG-~---CPNCGeEv~a 284 (333)
...+++++ | ||+|||+.++
T Consensus 22 ~~~~~i~~vp~~~C~~CGE~~~~ 44 (46)
T TIGR03831 22 GELIVIENVPALVCPQCGEEYLD 44 (46)
T ss_pred CEEEEEeCCCccccccCCCEeeC
Confidence 34455544 5 9999998764
No 51
>smart00661 RPOL9 RNA polymerase subunit 9.
Probab=71.78 E-value=5.7 Score=28.08 Aligned_cols=34 Identities=24% Similarity=0.579 Sum_probs=20.5
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEecc
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSN 321 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk 321 (333)
-||+||+-+ +. +.....+...|+. ||-..--+++
T Consensus 2 FCp~Cg~~l--~~------~~~~~~~~~vC~~------Cg~~~~~~~~ 35 (52)
T smart00661 2 FCPKCGNML--IP------KEGKEKRRFVCRK------CGYEEPIEQK 35 (52)
T ss_pred CCCCCCCcc--cc------ccCCCCCEEECCc------CCCeEECCCc
Confidence 399999843 22 2222335778888 9976544433
No 52
>PF06906 DUF1272: Protein of unknown function (DUF1272); InterPro: IPR010696 This family consists of several hypothetical bacterial proteins of around 80 residues in length. This family contains a number of conserved cysteine residues and its function is unknown.
Probab=70.69 E-value=1.8 Score=33.59 Aligned_cols=13 Identities=62% Similarity=1.347 Sum_probs=10.7
Q ss_pred eeecCCCCCcccc
Q 019981 270 ILKGPCPNCGTEN 282 (333)
Q Consensus 270 iLKG~CPNCGeEv 282 (333)
.|.|-|||||-|.
T Consensus 39 ~l~~~CPNCgGel 51 (57)
T PF06906_consen 39 MLNGVCPNCGGEL 51 (57)
T ss_pred HhcCcCcCCCCcc
Confidence 3589999999875
No 53
>smart00834 CxxC_CXXC_SSSS Putative regulatory protein. CxxC_CXXC_SSSS represents a region of about 41 amino acids found in a number of small proteins in a wide range of bacteria. The region usually begins with the initiator Met and contains two CxxC motifs separated by 17 amino acids. One protein in this entry has been noted as a putative regulatory protein, designated FmdB. Most proteins in this entry have a C-terminal region containing highly degenerate sequence.
Probab=70.12 E-value=3.8 Score=27.58 Aligned_cols=29 Identities=21% Similarity=0.568 Sum_probs=19.4
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCce
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTT 315 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~ 315 (333)
.||+||.+.-...+. .+...+.|++ ||..
T Consensus 7 ~C~~Cg~~fe~~~~~-------~~~~~~~CP~------Cg~~ 35 (41)
T smart00834 7 RCEDCGHTFEVLQKI-------SDDPLATCPE------CGGD 35 (41)
T ss_pred EcCCCCCEEEEEEec-------CCCCCCCCCC------CCCc
Confidence 699999966555421 2245677888 9983
No 54
>PF01096 TFIIS_C: Transcription factor S-II (TFIIS); InterPro: IPR001222 Zinc finger (Znf) domains are relatively small protein motifs which contain multiple finger-like protrusions that make tandem contacts with their target molecule. Some of these domains bind zinc, but many do not; instead binding other metals such as iron, or no metal at all. For example, some family members form salt bridges to stabilise the finger-like folds. They were first identified as a DNA-binding motif in transcription factor TFIIIA from Xenopus laevis (African clawed frog), however they are now recognised to bind DNA, RNA, protein and/or lipid substrates [, , , , ]. Their binding properties depend on the amino acid sequence of the finger domains and of the linker between fingers, as well as on the higher-order structures and the number of fingers. Znf domains are often found in clusters, where fingers can have different binding specificities. There are many superfamilies of Znf motifs, varying in both sequence and structure. They display considerable versatility in binding modes, even between members of the same class (e.g. some bind DNA, others protein), suggesting that Znf motifs are stable scaffolds that have evolved specialised functions. For example, Znf-containing proteins function in gene transcription, translation, mRNA trafficking, cytoskeleton organisation, epithelial development, cell adhesion, protein folding, chromatin remodelling and zinc sensing, to name but a few []. Zinc-binding motifs are stable structures, and they rarely undergo conformational changes upon binding their target. This entry represents a zinc finger motif found in transcription factor IIs (TFIIS). In eukaryotes the initiation of transcription of protein encoding genes by polymerase II (Pol II) is modulated by general and specific transcription factors. The general transcription factors operate through common promoters elements (such as the TATA box). At least eight different proteins associate to form the general transcription factors: TFIIA, -IIB, -IID, -IIE, -IIF, -IIG, -IIH and -IIS []. During mRNA elongation, Pol II can encounter DNA sequences that cause reverse movement of the enzyme. Such backtracking involves extrusion of the RNA 3'-end into the pore, and can lead to transcriptional arrest. Escape from arrest requires cleavage of the extruded RNA with the help of TFIIS, which induces mRNA cleavage by enhancing the intrinsic nuclease activity of RNA polymerase (Pol) II, past template-encoded pause sites []. TFIIS extends from the polymerase surface via a pore to the internal active site. Two essential and invariant acidic residues in a TFIIS loop complement the Pol II active site and could position a metal ion and a water molecule for hydrolytic RNA cleavage. TFIIS also induces extensive structural changes in Pol II that would realign nucleic acids in the active centre. TFIIS is a protein of about 300 amino acids. It contains three regions: a variable N-terminal domain not required for TFIIS activity; a conserved central domain required for Pol II binding; and a conserved C-terminal C4-type zinc finger essential for RNA cleavage. The zinc finger folds in a conformation termed a zinc ribbon [] characterised by a three-stranded antiparallel beta-sheet and two beta-hairpins. A backbone model for Pol II-TFIIS complex was obtained from X-ray analysis. It shows that a beta hairpin protrudes from the zinc finger and complements the pol II active site []. Some viral proteins also contain the TFIIS zinc ribbon C-terminal domain. The Vaccinia virus protein, unlike its eukaryotic homologue, is an integral RNA polymerase subunit rather than a readily separable transcription factor []. More information about these proteins can be found at Protein of the Month: Zinc Fingers [].; GO: 0003676 nucleic acid binding, 0008270 zinc ion binding, 0006351 transcription, DNA-dependent; PDB: 3M4O_I 3S14_I 2E2J_I 4A3J_I 3HOZ_I 1TWA_I 3S1Q_I 3S1N_I 1TWG_I 3I4M_I ....
Probab=68.82 E-value=3.5 Score=28.86 Aligned_cols=31 Identities=32% Similarity=0.666 Sum_probs=16.1
Q ss_pred CCCCCccccceeccccccc-cCCCCccceeCCC
Q 019981 274 PCPNCGTENVSFFGTILSI-SSGGTTNTINCSN 305 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv-~s~~~~n~vkC~~ 305 (333)
.||+||...-.|| .+.+= .-...+--+.|.+
T Consensus 2 ~Cp~Cg~~~a~~~-~~Q~rsaDE~~T~fy~C~~ 33 (39)
T PF01096_consen 2 KCPKCGHNEAVFF-QIQTRSADEPMTLFYVCCN 33 (39)
T ss_dssp --SSS-SSEEEEE-EESSSSSSSSSEEEEEESS
T ss_pred CCcCCCCCeEEEE-EeeccCCCCCCeEEEEeCC
Confidence 6999999998888 22111 1122344566766
No 55
>smart00778 Prim_Zn_Ribbon Zinc-binding domain of primase-helicase. This region represents the zinc binding domain. It is found in the N-terminal region of the bacteriophage P4 alpha protein, which is a multifunctional protein with origin recognition, helicase and primase activities.
Probab=68.57 E-value=2.2 Score=30.17 Aligned_cols=13 Identities=54% Similarity=1.411 Sum_probs=9.6
Q ss_pred ecCCCCCccc-cce
Q 019981 272 KGPCPNCGTE-NVS 284 (333)
Q Consensus 272 KG~CPNCGeE-v~a 284 (333)
++|||+||-. -|.
T Consensus 3 ~~pCP~CGG~DrFr 16 (37)
T smart00778 3 HGPCPNCGGSDRFR 16 (37)
T ss_pred ccCCCCCCCccccc
Confidence 5899999763 445
No 56
>PF13719 zinc_ribbon_5: zinc-ribbon domain
Probab=68.04 E-value=3.7 Score=28.41 Aligned_cols=35 Identities=23% Similarity=0.451 Sum_probs=20.6
Q ss_pred ecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeE
Q 019981 272 KGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMV 317 (333)
Q Consensus 272 KG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~ 317 (333)
+-.||||++.-. =++--+ ....-.|+|++ |+....
T Consensus 2 ~i~CP~C~~~f~---v~~~~l--~~~~~~vrC~~------C~~~f~ 36 (37)
T PF13719_consen 2 IITCPNCQTRFR---VPDDKL--PAGGRKVRCPK------CGHVFR 36 (37)
T ss_pred EEECCCCCceEE---cCHHHc--ccCCcEEECCC------CCcEee
Confidence 346999997532 111112 23444899999 987653
No 57
>COG5349 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=68.00 E-value=2 Score=37.88 Aligned_cols=33 Identities=30% Similarity=0.873 Sum_probs=22.4
Q ss_pred eecCCCCCccccc--eeccccccccCCCCccceeCCCCccccccCceeEEec
Q 019981 271 LKGPCPNCGTENV--SFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDS 320 (333)
Q Consensus 271 LKG~CPNCGeEv~--aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~t 320 (333)
|+|-||||||--. .|.+. .=.|.+ ||-.+-|..
T Consensus 20 l~grCP~CGeGrLF~gFLK~-----------~p~C~a------CG~dyg~~~ 54 (126)
T COG5349 20 LRGRCPRCGEGRLFRGFLKV-----------VPACEA------CGLDYGFAD 54 (126)
T ss_pred hcCCCCCCCCchhhhhhccc-----------Cchhhh------ccccccCCc
Confidence 7899999999643 44422 124665 998887753
No 58
>PF09851 SHOCT: Short C-terminal domain; InterPro: IPR018649 This family of hypothetical prokaryotic proteins has no known function.
Probab=68.00 E-value=7.1 Score=26.16 Aligned_cols=26 Identities=35% Similarity=0.684 Sum_probs=21.4
Q ss_pred HHHHHHhhhcCccccChHHHhhhHhhhh
Q 019981 103 LQALQAFYYEGKAVMSNEEFDNLKEELM 130 (333)
Q Consensus 103 l~Al~~fY~~gk~~~sdeefd~LkEeL~ 130 (333)
+..|...|..| .+|+|||...|..|+
T Consensus 5 L~~L~~l~~~G--~IseeEy~~~k~~ll 30 (31)
T PF09851_consen 5 LEKLKELYDKG--EISEEEYEQKKARLL 30 (31)
T ss_pred HHHHHHHHHcC--CCCHHHHHHHHHHHh
Confidence 46677888777 699999999999874
No 59
>PF09862 DUF2089: Protein of unknown function (DUF2089); InterPro: IPR018658 This family consists of various hypothetical prokaryotic proteins.
Probab=67.29 E-value=4.4 Score=34.98 Aligned_cols=23 Identities=43% Similarity=1.081 Sum_probs=17.4
Q ss_pred CCCCccccceeccccccccCCCCccceeCCCCccccccCceeE
Q 019981 275 CPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMV 317 (333)
Q Consensus 275 CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~ 317 (333)
||.||.+... .+++|++ |++.++
T Consensus 1 CPvCg~~l~v--------------t~l~C~~------C~t~i~ 23 (113)
T PF09862_consen 1 CPVCGGELVV--------------TRLKCPS------CGTEIE 23 (113)
T ss_pred CCCCCCceEE--------------EEEEcCC------CCCEEE
Confidence 8999877532 4678888 998875
No 60
>smart00440 ZnF_C2C2 C2C2 Zinc finger. Nucleic-acid-binding motif in transcriptional elongation factor TFIIS and RNA polymerases.
Probab=66.91 E-value=4.9 Score=28.32 Aligned_cols=32 Identities=28% Similarity=0.657 Sum_probs=19.3
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCC
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSN 305 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~ 305 (333)
+||+||...-.||-.-..-.....+--++|.+
T Consensus 2 ~Cp~C~~~~a~~~q~Q~RsaDE~mT~fy~C~~ 33 (40)
T smart00440 2 PCPKCGNREATFFQLQTRSADEPMTVFYVCTK 33 (40)
T ss_pred cCCCCCCCeEEEEEEcccCCCCCCeEEEEeCC
Confidence 69999988888873211111223345567776
No 61
>PF07282 OrfB_Zn_ribbon: Putative transposase DNA-binding domain; InterPro: IPR010095 This entry represents a region of a sequence similarity between a family of putative transposases of Thermoanaerobacter tengcongensis, smaller related proteins from Bacillus anthracis, putative transposes described by IPR001959 from INTERPRO, and other proteins. More information about these proteins can be found at Protein of the Month: Transposase [].
Probab=65.18 E-value=4.9 Score=30.10 Aligned_cols=28 Identities=32% Similarity=0.841 Sum_probs=20.7
Q ss_pred ecCCCCCccccceeccccccccCCCCccceeCCCCccccccCce
Q 019981 272 KGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTT 315 (333)
Q Consensus 272 KG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~ 315 (333)
.-.||.||..+.. ........|++ ||..
T Consensus 28 Sq~C~~CG~~~~~----------~~~~r~~~C~~------Cg~~ 55 (69)
T PF07282_consen 28 SQTCPRCGHRNKK----------RRSGRVFTCPN------CGFE 55 (69)
T ss_pred ccCccCccccccc----------ccccceEEcCC------CCCE
Confidence 3459999998877 23456778888 8876
No 62
>TIGR03655 anti_R_Lar restriction alleviation protein, Lar family. Restriction alleviation proteins provide a countermeasure to host cell restriction enzyme defense against foreign DNA such as phage or plasmids. This family consists of homologs to the phage antirestriction protein Lar, and most members belong to phage genomes or prophage regions of bacterial genomes.
Probab=65.16 E-value=5.5 Score=29.19 Aligned_cols=36 Identities=28% Similarity=0.630 Sum_probs=22.0
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEE
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVY 318 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f 318 (333)
|||-||.+...|.... .......-++|.+ ||.....
T Consensus 3 PCPfCGg~~~~~~~~~---~~~~~~~~~~C~~------Cga~~~~ 38 (53)
T TIGR03655 3 PCPFCGGADVYLRRGF---DPLDLSHYFECST------CGASGPV 38 (53)
T ss_pred CCCCCCCcceeeEecc---CCCCCEEEEECCC------CCCCccc
Confidence 8999999888664110 0122233346776 9887665
No 63
>COG1645 Uncharacterized Zn-finger containing protein [General function prediction only]
Probab=64.09 E-value=5.3 Score=35.44 Aligned_cols=39 Identities=28% Similarity=0.669 Sum_probs=28.7
Q ss_pred HHHhhhhcceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCce
Q 019981 259 SLTKLIVRESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTT 315 (333)
Q Consensus 259 ~Lt~l~~~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~ 315 (333)
.+.+++++=+..|--.||-||.+.|..=| +|-|++ ||..
T Consensus 15 ~iA~lLl~GAkML~~hCp~Cg~PLF~KdG------------~v~CPv------C~~~ 53 (131)
T COG1645 15 KIAELLLQGAKMLAKHCPKCGTPLFRKDG------------EVFCPV------CGYR 53 (131)
T ss_pred HHHHHHHhhhHHHHhhCcccCCcceeeCC------------eEECCC------CCce
Confidence 34466677777777899999999998433 456776 9963
No 64
>COG3813 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=63.44 E-value=3.1 Score=34.20 Aligned_cols=15 Identities=60% Similarity=1.143 Sum_probs=12.3
Q ss_pred eeecCCCCCccccce
Q 019981 270 ILKGPCPNCGTENVS 284 (333)
Q Consensus 270 iLKG~CPNCGeEv~a 284 (333)
.|.|.|||||-|..+
T Consensus 39 ~l~g~CPnCGGelv~ 53 (84)
T COG3813 39 RLHGLCPNCGGELVA 53 (84)
T ss_pred hhcCcCCCCCchhhc
Confidence 478999999988654
No 65
>TIGR00340 zpr1_rel ZPR1-related zinc finger protein. A model ZPR1_znf (TIGR00310) has been created to describe the domain shared by this protein and ZPR1.
Probab=63.05 E-value=4.7 Score=36.44 Aligned_cols=20 Identities=15% Similarity=0.403 Sum_probs=15.5
Q ss_pred hhcceeeeecCCCCCccccc
Q 019981 264 IVRESLILKGPCPNCGTENV 283 (333)
Q Consensus 264 ~~~D~~iLKG~CPNCGeEv~ 283 (333)
.+++.+|+...||+||-.+.
T Consensus 20 ~F~evii~sf~C~~CGyr~~ 39 (163)
T TIGR00340 20 YFGKIMLSTYICEKCGYRST 39 (163)
T ss_pred CcceEEEEEEECCCCCCchh
Confidence 47888888888888886665
No 66
>PF03367 zf-ZPR1: ZPR1 zinc-finger domain; InterPro: IPR004457 Zinc finger (Znf) domains are relatively small protein motifs which contain multiple finger-like protrusions that make tandem contacts with their target molecule. Some of these domains bind zinc, but many do not; instead binding other metals such as iron, or no metal at all. For example, some family members form salt bridges to stabilise the finger-like folds. They were first identified as a DNA-binding motif in transcription factor TFIIIA from Xenopus laevis (African clawed frog), however they are now recognised to bind DNA, RNA, protein and/or lipid substrates [, , , , ]. Their binding properties depend on the amino acid sequence of the finger domains and of the linker between fingers, as well as on the higher-order structures and the number of fingers. Znf domains are often found in clusters, where fingers can have different binding specificities. There are many superfamilies of Znf motifs, varying in both sequence and structure. They display considerable versatility in binding modes, even between members of the same class (e.g. some bind DNA, others protein), suggesting that Znf motifs are stable scaffolds that have evolved specialised functions. For example, Znf-containing proteins function in gene transcription, translation, mRNA trafficking, cytoskeleton organisation, epithelial development, cell adhesion, protein folding, chromatin remodelling and zinc sensing, to name but a few []. Zinc-binding motifs are stable structures, and they rarely undergo conformational changes upon binding their target. This entry represents ZPR1-type zinc finger domains. An orthologous protein found once in each of the completed archaeal genomes corresponds to a zinc finger-containing domain repeated as the N-terminal and C-terminal halves of the mouse protein ZPR1. ZPR1 is an experimentally proven zinc-binding protein that binds the tyrosine kinase domain of the epidermal growth factor receptor (EGFR); binding is inhibited by EGF stimulation and tyrosine phosphorylation, and activation by EGF is followed by some redistribution of ZPR1 to the nucleus. By analogy, other proteins with the ZPR1 zinc finger domain may be regulatory proteins that sense protein phosphorylation state and/or participate in signal transduction (see also IPR004470 from INTERPRO). Deficiencies in ZPR1 may contribute to neurodegenerative disorders. ZPR1 appears to be down-regulated in patients with spinal muscular atrophy (SMA), a disease characterised by degeneration of the alpha-motor neurons in the spinal cord that can arise from mutations affecting the expression of Survival Motor Neurons (SMN) []. ZPR1 interacts with complexes formed by SMN [], and may act as a modifier that effects the severity of SMA. More information about these proteins can be found at Protein of the Month: Zinc Fingers [].; GO: 0008270 zinc ion binding; PDB: 2QKD_A.
Probab=62.71 E-value=4 Score=36.53 Aligned_cols=20 Identities=30% Similarity=0.680 Sum_probs=11.5
Q ss_pred hcceeeeecCCCCCccccce
Q 019981 265 VRESLILKGPCPNCGTENVS 284 (333)
Q Consensus 265 ~~D~~iLKG~CPNCGeEv~a 284 (333)
+++.+|+...||+||-.+..
T Consensus 23 F~evii~sf~C~~CGyk~~e 42 (161)
T PF03367_consen 23 FKEVIIMSFECEHCGYKNNE 42 (161)
T ss_dssp TEEEEEEEEE-TTT--EEEE
T ss_pred CceEEEEEeECCCCCCEeee
Confidence 66777777777777766653
No 67
>smart00709 Zpr1 Duplicated domain in the epidermal growth factor- and elongation factor-1alpha-binding protein Zpr1. Also present in archaeal proteins.
Probab=62.01 E-value=4.9 Score=36.09 Aligned_cols=22 Identities=32% Similarity=0.628 Sum_probs=16.3
Q ss_pred hhcceeeeecCCCCCcccccee
Q 019981 264 IVRESLILKGPCPNCGTENVSF 285 (333)
Q Consensus 264 ~~~D~~iLKG~CPNCGeEv~aF 285 (333)
-+++.+|+...||+||-.+...
T Consensus 21 ~F~evii~sf~C~~CGyk~~ev 42 (160)
T smart00709 21 YFREVIIMSFECEHCGYRNNEV 42 (160)
T ss_pred CcceEEEEEEECCCCCCccceE
Confidence 3778888888888888776643
No 68
>PRK00432 30S ribosomal protein S27ae; Validated
Probab=61.22 E-value=5.9 Score=29.40 Aligned_cols=29 Identities=31% Similarity=0.802 Sum_probs=19.3
Q ss_pred eeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCce
Q 019981 270 ILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTT 315 (333)
Q Consensus 270 iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~ 315 (333)
.+.--||+||.+ |... ...+..|.. ||-.
T Consensus 18 ~~~~fCP~Cg~~---~m~~--------~~~r~~C~~------Cgyt 46 (50)
T PRK00432 18 RKNKFCPRCGSG---FMAE--------HLDRWHCGK------CGYT 46 (50)
T ss_pred EccCcCcCCCcc---hhec--------cCCcEECCC------cCCE
Confidence 456699999987 3321 125778877 8864
No 69
>PF12773 DZR: Double zinc ribbon
Probab=61.08 E-value=3.9 Score=28.91 Aligned_cols=9 Identities=56% Similarity=1.316 Sum_probs=4.2
Q ss_pred CCCCccccc
Q 019981 275 CPNCGTENV 283 (333)
Q Consensus 275 CPNCGeEv~ 283 (333)
||+||+.+.
T Consensus 15 C~~CG~~l~ 23 (50)
T PF12773_consen 15 CPHCGTPLP 23 (50)
T ss_pred ChhhcCChh
Confidence 444444444
No 70
>PF08273 Prim_Zn_Ribbon: Zinc-binding domain of primase-helicase; InterPro: IPR013237 This entry is represented by bacteriophage T7 Gp4. The characteristics of the protein distribution suggest prophage matches in addition to the phage matches. This entry represents a zinc binding domain found in the N-terminal region of the bacteriophage T7 Gp4 and P4 alpha protein. P4 is a multifunctional protein with origin recognition, helicase and primase activities [, , ].; GO: 0003896 DNA primase activity, 0004386 helicase activity, 0008270 zinc ion binding; PDB: 1NUI_B.
Probab=60.86 E-value=3.6 Score=29.54 Aligned_cols=15 Identities=47% Similarity=1.271 Sum_probs=8.0
Q ss_pred ecCCCCCccccc-eec
Q 019981 272 KGPCPNCGTENV-SFF 286 (333)
Q Consensus 272 KG~CPNCGeEv~-aFf 286 (333)
.+|||+||-.-. ..|
T Consensus 3 h~pCP~CGG~DrFri~ 18 (40)
T PF08273_consen 3 HGPCPICGGKDRFRIF 18 (40)
T ss_dssp EE--TTTT-TTTEEEE
T ss_pred CCCCCCCcCccccccC
Confidence 589999987544 534
No 71
>PF05191 ADK_lid: Adenylate kinase, active site lid; InterPro: IPR007862 Adenylate kinases (ADK; 2.7.4.3 from EC) are phosphotransferases that catalyse the Mg-dependent reversible conversion of ATP and AMP to two molecules of ADP, an essential reaction for many processes in living cells. In large variants of adenylate kinase, the AMP and ATP substrates are buried in a domain that undergoes conformational changes from an open to a closed state when bound to substrate; the ligand is then contained within a highly specific environment required for catalysis. Adenylate kinase is a 3-domain protein consisting of a large central CORE domain flanked by a LID domain on one side and the AMP-binding NMPbind domain on the other []. The LID domain binds ATP and covers the phosphates at the active site. The substrates first bind the CORE domain, followed by closure of the active site by the LID and NMPbind domains. Comparisons of adenylate kinases have revealed a particular divergence in the active site lid. In some organisms, particularly the Gram-positive bacteria, residues in the lid domain have been mutated to cysteines and these cysteine residues (two CX(n)C motifs) are responsible for the binding of a zinc ion. The bound zinc ion in the lid domain is clearly structurally homologous to Zinc-finger domains. However, it is unclear whether the adenylate kinase lid is a novel zinc-finger DNA/RNA binding domain, or that the lid bound zinc serves a purely structural function [].; GO: 0004017 adenylate kinase activity; PDB: 3BE4_A 2OSB_B 2ORI_A 2EU8_A 3DL0_A 1P3J_A 2QAJ_A 2OO7_A 2P3S_A 3DKV_A ....
Probab=60.62 E-value=4.8 Score=28.09 Aligned_cols=33 Identities=30% Similarity=0.597 Sum_probs=22.8
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEec
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDS 320 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~t 320 (333)
-||+||.---..|.. ....-+|.+ ||..|+-|.
T Consensus 3 ~C~~Cg~~Yh~~~~p--------P~~~~~Cd~------cg~~L~qR~ 35 (36)
T PF05191_consen 3 ICPKCGRIYHIEFNP--------PKVEGVCDN------CGGELVQRK 35 (36)
T ss_dssp EETTTTEEEETTTB----------SSTTBCTT------TTEBEBEEG
T ss_pred CcCCCCCccccccCC--------CCCCCccCC------CCCeeEeCC
Confidence 499999876666633 455668887 999887654
No 72
>PF09538 FYDLN_acid: Protein of unknown function (FYDLN_acid); InterPro: IPR012644 Members of this family are bacterial proteins with a conserved motif [KR]FYDLN, sometimes flanked by a pair of CXXC motifs, followed by a long region of low complexity sequence in which roughly half the residues are Asp and Glu, including multiple runs of five or more acidic residues. The function of members of this family is unknown.
Probab=59.92 E-value=6.8 Score=33.31 Aligned_cols=32 Identities=31% Similarity=0.783 Sum_probs=23.9
Q ss_pred eecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEe
Q 019981 271 LKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 271 LKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~ 319 (333)
.|--||+||+.-|-. |+ .-+-|+. ||+...-.
T Consensus 8 tKR~Cp~CG~kFYDL---------nk--~PivCP~------CG~~~~~~ 39 (108)
T PF09538_consen 8 TKRTCPSCGAKFYDL---------NK--DPIVCPK------CGTEFPPE 39 (108)
T ss_pred CcccCCCCcchhccC---------CC--CCccCCC------CCCccCcc
Confidence 477899999976654 44 3467888 99977666
No 73
>PF04216 FdhE: Protein involved in formate dehydrogenase formation; InterPro: IPR006452 This family of sequences describe an accessory protein required for the assembly of formate dehydrogenase of certain proteobacteria although not present in the final complex []. The exact nature of the function of FdhE in the assembly of the complex is unknown, but considering the presence of selenocysteine, molybdopterin, iron-sulphur clusters and cytochrome b556, it is likely to be involved in the insertion of cofactors. ; GO: 0005737 cytoplasm; PDB: 2FIY_B.
Probab=59.66 E-value=9.1 Score=36.43 Aligned_cols=43 Identities=26% Similarity=0.542 Sum_probs=19.8
Q ss_pred HHhhhhHHHHHHHHHHhhhhcceeeeecCCCCCccc-cceeccc
Q 019981 246 WFAAVPLIVYLSQSLTKLIVRESLILKGPCPNCGTE-NVSFFGT 288 (333)
Q Consensus 246 ~~~~~P~i~~~a~~Lt~l~~~D~~iLKG~CPNCGeE-v~aFfgt 288 (333)
|.+.-|+.-..+..+..-+.....-.+|-||.||.. +.+.+..
T Consensus 146 ~aaL~~~~~~~a~~l~~~~~~~~~w~~g~CPvCGs~P~~s~l~~ 189 (290)
T PF04216_consen 146 WAALQPFLAALAAALDAALLPPEGWQRGYCPVCGSPPVLSVLRG 189 (290)
T ss_dssp HHHHHHHHHHHHHT--TTSSS---TT-SS-TTT---EEEEEEE-
T ss_pred HHHHHHHHHHHHHhccccccccCCccCCcCCCCCCcCceEEEec
Confidence 444446666666555544444445567999999987 5566643
No 74
>TIGR03830 CxxCG_CxxCG_HTH putative zinc finger/helix-turn-helix protein, YgiT family. This model describes a family of predicted regulatory proteins with a conserved zinc finger/HTH architecture. The amino-terminal region contains a novel domain, featuring two CXXC motifs and occuring in a number of small bacterial proteins as well as in the present family. The carboxyl-terminal region consists of a helix-turn-helix domain, modeled by pfam01381. The predicted function is DNA binding and transcriptional regulation.
Probab=58.72 E-value=6 Score=32.18 Aligned_cols=39 Identities=31% Similarity=0.562 Sum_probs=18.2
Q ss_pred CCCCcc-ccceecccc-ccccCCCCccceeCCCCccccccCcee
Q 019981 275 CPNCGT-ENVSFFGTI-LSISSGGTTNTINCSNLTFCFSCGTTM 316 (333)
Q Consensus 275 CPNCGe-Ev~aFfgti-lsv~s~~~~n~vkC~~~aeCHVC~t~L 316 (333)
||.||. +...-+.+. .++ .+....+ .-+-..|..||..+
T Consensus 1 C~~C~~~~~~~~~~~~~~~~--~G~~~~v-~~~~~~C~~CGe~~ 41 (127)
T TIGR03830 1 CPICGSGELVRDVKDEPYTY--KGESITI-GVPGWYCPACGEEL 41 (127)
T ss_pred CCCCCCccceeeeecceEEE--cCEEEEE-eeeeeECCCCCCEE
Confidence 999995 343333221 122 2222233 11124566698863
No 75
>COG0272 Lig NAD-dependent DNA ligase (contains BRCT domain type II) [DNA replication, recombination, and repair]
Probab=56.48 E-value=13 Score=40.65 Aligned_cols=38 Identities=24% Similarity=0.358 Sum_probs=30.8
Q ss_pred cchhHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhh
Q 019981 93 KSLGELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELM 130 (333)
Q Consensus 93 ~slge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~ 130 (333)
+-+.++.+.-.+.-.+||..+.|.++|.|||.|..||.
T Consensus 9 ~~i~~L~~~L~~~~~~Yyv~d~P~VsD~eYD~L~reL~ 46 (667)
T COG0272 9 EEIEELRELLNKHDYRYYVLDAPSVSDAEYDQLYRELQ 46 (667)
T ss_pred HHHHHHHHHHHHHHHHHhccCCCCCChHHHHHHHHHHH
Confidence 34566666666667789999999999999999988874
No 76
>PF09723 Zn-ribbon_8: Zinc ribbon domain; InterPro: IPR013429 This entry represents a region of about 41 amino acids found in a number of small proteins in a wide range of bacteria. The region usually begins with the initiator Met and contains two CxxC motifs separated by 17 amino acids. One protein in this entry has been noted as a putative regulatory protein, designated FmdB []. Most proteins in this entry have a C-terminal region containing highly degenerate sequence.
Probab=55.25 E-value=10 Score=26.75 Aligned_cols=28 Identities=25% Similarity=0.701 Sum_probs=19.4
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCc
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGT 314 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t 314 (333)
-|++||.+.-.+.. ..+...+.|++ ||.
T Consensus 7 ~C~~Cg~~fe~~~~-------~~~~~~~~CP~------Cg~ 34 (42)
T PF09723_consen 7 RCEECGHEFEVLQS-------ISEDDPVPCPE------CGS 34 (42)
T ss_pred EeCCCCCEEEEEEE-------cCCCCCCcCCC------CCC
Confidence 49999987666652 22256778887 998
No 77
>PRK01103 formamidopyrimidine/5-formyluracil/ 5-hydroxymethyluracil DNA glycosylase; Validated
Probab=54.76 E-value=10 Score=36.00 Aligned_cols=32 Identities=34% Similarity=0.686 Sum_probs=20.3
Q ss_pred cceeeeec----CCCCCccccceeccccccccCCCCccceeCCC
Q 019981 266 RESLILKG----PCPNCGTENVSFFGTILSISSGGTTNTINCSN 305 (333)
Q Consensus 266 ~D~~iLKG----~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~ 305 (333)
++.+.+.| |||.||+++.... -++++ ++-|++
T Consensus 235 ~~~l~Vy~R~g~pC~~Cg~~I~~~~------~~gR~--t~~CP~ 270 (274)
T PRK01103 235 QQSLQVYGREGEPCRRCGTPIEKIK------QGGRS--TFFCPR 270 (274)
T ss_pred cceeEEcCCCCCCCCCCCCeeEEEE------ECCCC--cEECcC
Confidence 34445554 8999999987433 23444 457877
No 78
>PF03604 DNA_RNApol_7kD: DNA directed RNA polymerase, 7 kDa subunit; InterPro: IPR006591 DNA-dependent RNA polymerase catalyzes the transcription of DNA into RNA using the four ribonucleoside triphosphates as substrates. Each class of RNA polymerase is assembled from 9 to 15 different polypeptides. Rbp10 (RNA polymerase CX) is a domain found in RNA polymerase subunit 10; present in RNA polymerase I, II and III.; GO: 0003677 DNA binding, 0003899 DNA-directed RNA polymerase activity, 0006351 transcription, DNA-dependent; PDB: 2PMZ_Z 3HKZ_X 2NVX_L 3S1Q_L 2JA6_L 3S17_L 3HOW_L 3HOV_L 3PO2_L 3HOZ_L ....
Probab=54.44 E-value=7.6 Score=26.66 Aligned_cols=27 Identities=33% Similarity=0.977 Sum_probs=18.5
Q ss_pred CCCCccccceeccccccccCCCCccceeCCCCccccccCceeEE
Q 019981 275 CPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVY 318 (333)
Q Consensus 275 CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f 318 (333)
|..||.+|. + + ....++|++ ||..+.|
T Consensus 3 C~~Cg~~~~--~------~---~~~~irC~~------CG~RIly 29 (32)
T PF03604_consen 3 CGECGAEVE--L------K---PGDPIRCPE------CGHRILY 29 (32)
T ss_dssp ESSSSSSE---B------S---TSSTSSBSS------SS-SEEB
T ss_pred CCcCCCeeE--c------C---CCCcEECCc------CCCeEEE
Confidence 889999997 2 1 123579999 9988776
No 79
>PRK08665 ribonucleotide-diphosphate reductase subunit alpha; Validated
Probab=54.16 E-value=13 Score=40.78 Aligned_cols=13 Identities=38% Similarity=1.068 Sum_probs=9.0
Q ss_pred cCCCCCccccceec
Q 019981 273 GPCPNCGTENVSFF 286 (333)
Q Consensus 273 G~CPNCGeEv~aFf 286 (333)
+.||.||+ ...|-
T Consensus 725 ~~Cp~Cg~-~l~~~ 737 (752)
T PRK08665 725 GACPECGS-ILEHE 737 (752)
T ss_pred CCCCCCCc-ccEEC
Confidence 56999994 45553
No 80
>PF11746 DUF3303: Protein of unknown function (DUF3303); InterPro: IPR021734 Several members are annotated as being LysM domain-like proteins, but these did not match any LysM domains reported in the literature.
Probab=53.91 E-value=9.9 Score=31.02 Aligned_cols=70 Identities=27% Similarity=0.307 Sum_probs=49.3
Q ss_pred hHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhhh-cCCeeEEeChhhHHHHHHHHh-hhcC-------CCccChHHH
Q 019981 96 GELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELMW-EGSSVVMLSSAEQKFLEASMA-YVAG-------KPIMSDEEY 165 (333)
Q Consensus 96 ge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~w-eGSsvv~L~~~Eq~fLEA~~A-Y~sG-------kPimsDeeF 165 (333)
++..+.--+++.+|+..|++...-|.|..|..=.+= .|..++.+..+..+-|-+-.+ ..+. .|+|+|+|+
T Consensus 11 ~~~~~~~~~~~~~~~~~G~~~~~peG~~~l~rw~~~~~g~g~~i~eadd~~~l~~~~~~W~~~fg~~~ei~Pv~~d~e~ 89 (91)
T PF11746_consen 11 GESQQEAYKAFERFMESGAPGDPPEGFKVLGRWHDPGGGRGFAIVEADDAKALFKHFAPWRDLFGMEFEITPVMTDEEA 89 (91)
T ss_pred cccchhHHHHHHHHHhcCCCCCCCCCEEEEEEEEecCCCcEEEEEEeCCHHHHHHHHhhhhhccCceEEEEecccHHHh
Confidence 455566778999999999887777777666443222 666777777777666666544 4444 699999997
No 81
>PF14319 Zn_Tnp_IS91: Transposase zinc-binding domain
Probab=52.93 E-value=8.4 Score=32.44 Aligned_cols=29 Identities=31% Similarity=0.810 Sum_probs=22.2
Q ss_pred eecCCCCCccccceeccccccccCCCCccceeCCCCccccccCce
Q 019981 271 LKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTT 315 (333)
Q Consensus 271 LKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~ 315 (333)
..-.|++||.+-+.++ .|.+| .|..|+..
T Consensus 41 ~~~~C~~Cg~~~~~~~---------------SCk~R-~CP~C~~~ 69 (111)
T PF14319_consen 41 HRYRCEDCGHEKIVYN---------------SCKNR-HCPSCQAK 69 (111)
T ss_pred ceeecCCCCceEEecC---------------cccCc-CCCCCCCh
Confidence 3447999999998887 46665 77779874
No 82
>PF13717 zinc_ribbon_4: zinc-ribbon domain
Probab=52.60 E-value=11 Score=26.09 Aligned_cols=33 Identities=27% Similarity=0.529 Sum_probs=20.4
Q ss_pred ecCCCCCccccceeccccccccCCCCccceeCCCCccccccCce
Q 019981 272 KGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTT 315 (333)
Q Consensus 272 KG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~ 315 (333)
+-.||||++.-. =++--|+ ....+++|++ |+..
T Consensus 2 ~i~Cp~C~~~y~---i~d~~ip--~~g~~v~C~~------C~~~ 34 (36)
T PF13717_consen 2 IITCPNCQAKYE---IDDEKIP--PKGRKVRCSK------CGHV 34 (36)
T ss_pred EEECCCCCCEEe---CCHHHCC--CCCcEEECCC------CCCE
Confidence 346999997532 1222233 3445899999 9865
No 83
>PF07508 Recombinase: Recombinase; InterPro: IPR011109 This domain is usually found associated with IPR006119 from INTERPRO in putative integrases/recombinases of mobile genetic elements of diverse bacteria and phages.
Probab=50.57 E-value=13 Score=28.87 Aligned_cols=20 Identities=35% Similarity=0.717 Sum_probs=17.1
Q ss_pred cCCCccChHHHHHHHHHHhh
Q 019981 155 AGKPIMSDEEYDKLKQKLKM 174 (333)
Q Consensus 155 sGkPimsDeeFD~LK~kLk~ 174 (333)
.-.|||++++|+++...|+.
T Consensus 82 ~~~~IIs~~~f~~vq~~l~~ 101 (102)
T PF07508_consen 82 YHPPIISEEEFERVQKKLDE 101 (102)
T ss_pred CCCCccCHHHHHHHHHHHhc
Confidence 34699999999999999863
No 84
>PRK14714 DNA polymerase II large subunit; Provisional
Probab=49.84 E-value=13 Score=43.29 Aligned_cols=22 Identities=32% Similarity=0.530 Sum_probs=17.5
Q ss_pred ceeeeecCCCCCccccceeccc
Q 019981 267 ESLILKGPCPNCGTENVSFFGT 288 (333)
Q Consensus 267 D~~iLKG~CPNCGeEv~aFfgt 288 (333)
++.+-.--||+||++++.+|=.
T Consensus 662 eVEV~~rkCPkCG~~t~~~fCP 683 (1337)
T PRK14714 662 EVEVGRRRCPSCGTETYENRCP 683 (1337)
T ss_pred EEEEEEEECCCCCCccccccCc
Confidence 4667788999999999887643
No 85
>TIGR02300 FYDLN_acid conserved hypothetical protein TIGR02300. Members of this family are bacterial proteins with a conserved motif [KR]FYDLN, sometimes flanked by a pair of CXXC motifs, followed by a long region of low complexity sequence in which roughly half the residues are Asp and Glu, including multiple runs of five or more acidic residues. The function of members of this family is unknown.
Probab=49.66 E-value=12 Score=33.32 Aligned_cols=32 Identities=19% Similarity=0.310 Sum_probs=22.5
Q ss_pred eecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEe
Q 019981 271 LKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 271 LKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~ 319 (333)
.|--||+||+.-|-. |+ .-+.|+. ||+...-.
T Consensus 8 tKr~Cp~cg~kFYDL---------nk--~p~vcP~------cg~~~~~~ 39 (129)
T TIGR02300 8 TKRICPNTGSKFYDL---------NR--RPAVSPY------TGEQFPPE 39 (129)
T ss_pred ccccCCCcCcccccc---------CC--CCccCCC------cCCccCcc
Confidence 477899999875543 33 4568888 99975444
No 86
>PRK13130 H/ACA RNA-protein complex component Nop10p; Reviewed
Probab=49.33 E-value=7.8 Score=29.76 Aligned_cols=15 Identities=40% Similarity=0.844 Sum_probs=10.3
Q ss_pred eeecCCCCCccccce
Q 019981 270 ILKGPCPNCGTENVS 284 (333)
Q Consensus 270 iLKG~CPNCGeEv~a 284 (333)
-||..||+||++..+
T Consensus 15 TLk~~CP~CG~~t~~ 29 (56)
T PRK13130 15 TLKEICPVCGGKTKN 29 (56)
T ss_pred EccccCcCCCCCCCC
Confidence 357778888877553
No 87
>COG4306 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=49.04 E-value=10 Score=34.15 Aligned_cols=39 Identities=21% Similarity=0.646 Sum_probs=25.8
Q ss_pred CCCCCccccc--eeccccccccCCCCccceeCCCCccccccCceeEEe
Q 019981 274 PCPNCGTENV--SFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 274 ~CPNCGeEv~--aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~ 319 (333)
.||.|-+.++ -|+-+.|+.-+. .+= -++||-||+..-+-
T Consensus 41 qcp~csasirgd~~vegvlglg~d-----ye~--psfchncgs~fpwt 81 (160)
T COG4306 41 QCPICSASIRGDYYVEGVLGLGGD-----YEP--PSFCHNCGSRFPWT 81 (160)
T ss_pred cCCccCCcccccceeeeeeccCCC-----CCC--cchhhcCCCCCCcH
Confidence 5999999998 455555655332 233 25888899976543
No 88
>PRK14810 formamidopyrimidine-DNA glycosylase; Provisional
Probab=48.39 E-value=11 Score=35.87 Aligned_cols=25 Identities=28% Similarity=0.608 Sum_probs=17.3
Q ss_pred cCCCCCccccceeccccccccCCCCccceeCCC
Q 019981 273 GPCPNCGTENVSFFGTILSISSGGTTNTINCSN 305 (333)
Q Consensus 273 G~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~ 305 (333)
-|||.||.++..-. .++++ ++-|++
T Consensus 245 ~pCprCG~~I~~~~------~~gR~--t~~CP~ 269 (272)
T PRK14810 245 EPCLNCKTPIRRVV------VAGRS--SHYCPH 269 (272)
T ss_pred CcCCCCCCeeEEEE------ECCCc--cEECcC
Confidence 39999999996443 24444 567877
No 89
>COG5525 Bacteriophage tail assembly protein [General function prediction only]
Probab=48.34 E-value=13 Score=40.18 Aligned_cols=66 Identities=23% Similarity=0.323 Sum_probs=36.0
Q ss_pred HhhhhcceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccceeeecC
Q 019981 261 TKLIVRESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRLITLP 328 (333)
Q Consensus 261 t~l~~~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~itl~ 328 (333)
..+...|..=-.-+||.||++-.==|+...+--+......--| +.-||.|.+.+......|=|-+-
T Consensus 216 ~~y~~gd~rr~yvpCPHCGe~q~l~~~e~~~~~g~~~~~~~~~--~~~c~h~~~~i~~~~~~~gv~~~ 281 (611)
T COG5525 216 RAYNAGDQRRFYVPCPHCGEEQQLKFGEKSGPRGLKDTPAEAA--FIQCEHCGCVIRPKLNGRGVCLR 281 (611)
T ss_pred HHhhhccceeEEeeCCCCCchhhccccccCCCcCcccchhhhh--hhhccccCceeeeeccCccchhc
Confidence 3445668888888999999976533333221111222211111 23455599999884444444333
No 90
>PRK10445 endonuclease VIII; Provisional
Probab=48.15 E-value=12 Score=35.54 Aligned_cols=25 Identities=20% Similarity=0.288 Sum_probs=17.2
Q ss_pred cCCCCCccccceeccccccccCCCCccceeCCC
Q 019981 273 GPCPNCGTENVSFFGTILSISSGGTTNTINCSN 305 (333)
Q Consensus 273 G~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~ 305 (333)
.+||.||.++..-. -++++ ++-|++
T Consensus 236 ~~Cp~Cg~~I~~~~------~~gR~--t~~CP~ 260 (263)
T PRK10445 236 EACERCGGIIEKTT------LSSRP--FYWCPG 260 (263)
T ss_pred CCCCCCCCEeEEEE------ECCCC--cEECCC
Confidence 58999999987443 23444 567877
No 91
>TIGR00097 HMP-P_kinase phosphomethylpyrimidine kinase. This model represents phosphomethylpyrimidine kinase, the ThiD protein of thiamine biosynthesis. The protein is commonly observed within operons containing other thiamine biosynthesis genes. Numerous examples are fusion proteins with other thiamine-biosynthetic domains. Saccaromyces has three recent paralogs, two of which are isofunctional and score above the trusted cutoff. The third shows a longer branch length in a phylogenetic tree and scores below the trusted cutoff, as do putative second copies in a number of species.
Probab=48.10 E-value=68 Score=29.34 Aligned_cols=57 Identities=21% Similarity=0.388 Sum_probs=40.2
Q ss_pred HHHhhhHhhhhhcCCeeEEeChhhHHHHHHHHhhhcCCCccChHHHHHHHHHHhhcCCc-eeeecC
Q 019981 120 EEFDNLKEELMWEGSSVVMLSSAEQKFLEASMAYVAGKPIMSDEEYDKLKQKLKMEGSE-IVVEGP 184 (333)
Q Consensus 120 eefd~LkEeL~weGSsvv~L~~~Eq~fLEA~~AY~sGkPimsDeeFD~LK~kLk~~GS~-VVvk~P 184 (333)
+..+.++++| .....+++.+..|-+.| .|.++-+.++..+.-.+|...|-+ |++++.
T Consensus 115 ~~~~~~~~~l-l~~~dvitpN~~Ea~~L-------~g~~~~~~~~~~~~a~~l~~~g~~~Vvvt~G 172 (254)
T TIGR00097 115 EAIEALRKRL-LPLATLITPNLPEAEAL-------LGTKIRTEQDMIKAAKKLRELGPKAVLIKGG 172 (254)
T ss_pred HHHHHHHHhc-cccccEecCCHHHHHHH-------hCCCCCCHHHHHHHHHHHHhcCCCEEEEeCC
Confidence 3345566654 35677999999998876 366666767777777888877865 777764
No 92
>COG3677 Transposase and inactivated derivatives [DNA replication, recombination, and repair]
Probab=47.66 E-value=17 Score=31.58 Aligned_cols=49 Identities=22% Similarity=0.426 Sum_probs=31.6
Q ss_pred hcceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEecccee
Q 019981 265 VRESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRL 324 (333)
Q Consensus 265 ~~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~ 324 (333)
.....+.+--||-|+.++ .+.-. ...+...+.+|+. |+........+++
T Consensus 23 ~~~~~~~~~~cP~C~s~~-~~k~g----~~~~~~qRyrC~~------C~~tf~~~~~~~~ 71 (129)
T COG3677 23 AIRMQITKVNCPRCKSSN-VVKIG----GIRRGHQRYKCKS------CGSTFTVETGSPL 71 (129)
T ss_pred HHhhhcccCcCCCCCccc-eeeEC----CccccccccccCC------cCcceeeeccCcc
Confidence 344556677899999999 33211 1222245666666 9999998876554
No 93
>TIGR00577 fpg formamidopyrimidine-DNA glycosylase (fpg). All proteins in the FPG family with known functions are FAPY-DNA glycosylases that function in base excision repair. Homologous to endonuclease VIII (nei). This family is based on the phylogenomic analysis of JA Eisen (1999, Ph.D. Thesis, Stanford University).
Probab=47.26 E-value=12 Score=35.67 Aligned_cols=24 Identities=33% Similarity=0.631 Sum_probs=16.9
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCC
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSN 305 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~ 305 (333)
|||.||+++.... -++++ .+-|++
T Consensus 247 pC~~Cg~~I~~~~------~~gR~--t~~CP~ 270 (272)
T TIGR00577 247 PCRRCGTPIEKIK------VGGRG--THFCPQ 270 (272)
T ss_pred CCCCCCCeeEEEE------ECCCC--CEECCC
Confidence 8999999987543 23444 557877
No 94
>COG1998 RPS31 Ribosomal protein S27AE [Translation, ribosomal structure and biogenesis]
Probab=47.21 E-value=14 Score=28.28 Aligned_cols=19 Identities=21% Similarity=0.345 Sum_probs=13.9
Q ss_pred eeeeecCCCCCccccc-eec
Q 019981 268 SLILKGPCPNCGTENV-SFF 286 (333)
Q Consensus 268 ~~iLKG~CPNCGeEv~-aFf 286 (333)
++-++--||+||.-+| |.-
T Consensus 15 v~rk~~~CPrCG~gvfmA~H 34 (51)
T COG1998 15 VKRKNRFCPRCGPGVFMADH 34 (51)
T ss_pred EEEccccCCCCCCcchhhhc
Confidence 4556778999999876 443
No 95
>COG1096 Predicted RNA-binding protein (consists of S1 domain and a Zn-ribbon domain) [Translation, ribosomal structure and biogenesis]
Probab=46.01 E-value=16 Score=34.32 Aligned_cols=41 Identities=24% Similarity=0.535 Sum_probs=30.6
Q ss_pred hcceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccceeeecC
Q 019981 265 VRESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRLITLP 328 (333)
Q Consensus 265 ~~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~itl~ 328 (333)
.||+=.++.-|+||+++..- +. +.++|+| ||. +.+|.|-.+
T Consensus 142 ~~dlGVI~A~CsrC~~~L~~--~~----------~~l~Cp~------Cg~-----tEkRKia~~ 182 (188)
T COG1096 142 GNDLGVIYARCSRCRAPLVK--KG----------NMLKCPN------CGN-----TEKRKIAKD 182 (188)
T ss_pred CCcceEEEEEccCCCcceEE--cC----------cEEECCC------CCC-----EEeeeeccc
Confidence 78898899999999998765 22 5789999 995 345555433
No 96
>PF01807 zf-CHC2: CHC2 zinc finger; InterPro: IPR002694 Zinc finger (Znf) domains are relatively small protein motifs which contain multiple finger-like protrusions that make tandem contacts with their target molecule. Some of these domains bind zinc, but many do not; instead binding other metals such as iron, or no metal at all. For example, some family members form salt bridges to stabilise the finger-like folds. They were first identified as a DNA-binding motif in transcription factor TFIIIA from Xenopus laevis (African clawed frog), however they are now recognised to bind DNA, RNA, protein and/or lipid substrates [, , , , ]. Their binding properties depend on the amino acid sequence of the finger domains and of the linker between fingers, as well as on the higher-order structures and the number of fingers. Znf domains are often found in clusters, where fingers can have different binding specificities. There are many superfamilies of Znf motifs, varying in both sequence and structure. They display considerable versatility in binding modes, even between members of the same class (e.g. some bind DNA, others protein), suggesting that Znf motifs are stable scaffolds that have evolved specialised functions. For example, Znf-containing proteins function in gene transcription, translation, mRNA trafficking, cytoskeleton organisation, epithelial development, cell adhesion, protein folding, chromatin remodelling and zinc sensing, to name but a few []. Zinc-binding motifs are stable structures, and they rarely undergo conformational changes upon binding their target. This entry represents CycHisCysCys (CHC2) type zinc finger domains, which are found in bacteria and viruses. More information about these proteins can be found at Protein of the Month: Zinc Fingers [].; GO: 0003677 DNA binding, 0003896 DNA primase activity, 0008270 zinc ion binding, 0006260 DNA replication; PDB: 1D0Q_B 2AU3_A.
Probab=45.77 E-value=11 Score=30.57 Aligned_cols=31 Identities=29% Similarity=0.593 Sum_probs=15.7
Q ss_pred eecCCCCCccccceeccccccccCCCCccceeCCCCccccccCc
Q 019981 271 LKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGT 314 (333)
Q Consensus 271 LKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t 314 (333)
..+.||-|++.+-+|. | +.+.+..+|.. ||.
T Consensus 32 ~~~~CPfH~d~~pS~~-----i--~~~k~~~~Cf~------Cg~ 62 (97)
T PF01807_consen 32 YRCLCPFHDDKTPSFS-----I--NPDKNRFKCFG------CGK 62 (97)
T ss_dssp EEE--SSS--SS--EE-----E--ETTTTEEEETT------T--
T ss_pred EEEECcCCCCCCCceE-----E--ECCCCeEEECC------CCC
Confidence 5688999999888775 2 33556677755 885
No 97
>TIGR02605 CxxC_CxxC_SSSS putative regulatory protein, FmdB family. This model represents a region of about 50 amino acids found in a number of small proteins in a wide range of bacteria. The region begins usually with the initiator Met and contains two CxxC motifs separated by 17 amino acids. One member of this family is has been noted as a putative regulatory protein, designated FmdB (PubMed:8841393). Most members of this family have a C-terminal region containing highly degenerate sequence, such as SSTSESTKSSGSSGSSGSSESKASGSTEKSTSSTTAAAAV in Mycobacterium tuberculosis and VAVGGSAPAPSPAPRAGGGGGGCCGGGCCG in Streptomyces avermitilis. These low complexity regions, which are not included in the model, resemble low-complexity C-terminal regions of some heterocycle-containing bacteriocin precursors.
Probab=45.44 E-value=22 Score=25.37 Aligned_cols=28 Identities=21% Similarity=0.538 Sum_probs=17.2
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCc
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGT 314 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t 314 (333)
-|++||.+.-.+.. ..+...+.|++ ||.
T Consensus 7 ~C~~Cg~~fe~~~~-------~~~~~~~~CP~------Cg~ 34 (52)
T TIGR02605 7 RCTACGHRFEVLQK-------MSDDPLATCPE------CGG 34 (52)
T ss_pred EeCCCCCEeEEEEe-------cCCCCCCCCCC------CCC
Confidence 49999975544431 11234566777 998
No 98
>PF10083 DUF2321: Uncharacterized protein conserved in bacteria (DUF2321); InterPro: IPR016891 This entry is represented by Bacteriophage 'Lactobacillus prophage Lj928', Orf-Ljo1454. The characteristics of the protein distribution suggest prophage matches in addition to the phage matches. There is currently no experimental data for members of this group or their homologues, nor do they exhibit features indicative of any function.
Probab=45.09 E-value=12 Score=34.21 Aligned_cols=37 Identities=24% Similarity=0.641 Sum_probs=21.8
Q ss_pred CCCCCcccccee--ccccccccCCCCccceeCCCCccccccCceeE
Q 019981 274 PCPNCGTENVSF--FGTILSISSGGTTNTINCSNLTFCFSCGTTMV 317 (333)
Q Consensus 274 ~CPNCGeEv~aF--fgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~ 317 (333)
.||||++.+.-. +-+.+++ +..+. .-+.||-||.+.-
T Consensus 41 ~Cp~C~~~IrG~y~v~gv~~~---g~~~~----~PsYC~~CGkpyP 79 (158)
T PF10083_consen 41 SCPNCSTPIRGDYHVEGVFGL---GGHYE----APSYCHNCGKPYP 79 (158)
T ss_pred HCcCCCCCCCCceecCCeeee---CCCCC----CChhHHhCCCCCc
Confidence 599999999832 2233333 22221 2357777998753
No 99
>PF03119 DNA_ligase_ZBD: NAD-dependent DNA ligase C4 zinc finger domain; InterPro: IPR004149 Zinc finger (Znf) domains are relatively small protein motifs which contain multiple finger-like protrusions that make tandem contacts with their target molecule. Some of these domains bind zinc, but many do not; instead binding other metals such as iron, or no metal at all. For example, some family members form salt bridges to stabilise the finger-like folds. They were first identified as a DNA-binding motif in transcription factor TFIIIA from Xenopus laevis (African clawed frog), however they are now recognised to bind DNA, RNA, protein and/or lipid substrates [, , , , ]. Their binding properties depend on the amino acid sequence of the finger domains and of the linker between fingers, as well as on the higher-order structures and the number of fingers. Znf domains are often found in clusters, where fingers can have different binding specificities. There are many superfamilies of Znf motifs, varying in both sequence and structure. They display considerable versatility in binding modes, even between members of the same class (e.g. some bind DNA, others protein), suggesting that Znf motifs are stable scaffolds that have evolved specialised functions. For example, Znf-containing proteins function in gene transcription, translation, mRNA trafficking, cytoskeleton organisation, epithelial development, cell adhesion, protein folding, chromatin remodelling and zinc sensing, to name but a few []. Zinc-binding motifs are stable structures, and they rarely undergo conformational changes upon binding their target. This entry represents the zinc finger domain found in NAD-dependent DNA ligases. DNA ligases catalyse the crucial step of joining the breaks in duplex DNA during DNA replication, repair and recombination, utilizing either ATP or NAD(+) as a cofactor []. This domain is a small zinc binding motif that is presumably DNA binding. It is found only in NAD-dependent DNA ligases. More information about these proteins can be found at Protein of the Month: Zinc Fingers [].; GO: 0003911 DNA ligase (NAD+) activity, 0006260 DNA replication, 0006281 DNA repair; PDB: 1DGS_A 1V9P_B 2OWO_A.
Probab=44.31 E-value=13 Score=24.52 Aligned_cols=11 Identities=45% Similarity=1.029 Sum_probs=5.5
Q ss_pred CCCCcccccee
Q 019981 275 CPNCGTENVSF 285 (333)
Q Consensus 275 CPNCGeEv~aF 285 (333)
||.||+++...
T Consensus 2 CP~C~s~l~~~ 12 (28)
T PF03119_consen 2 CPVCGSKLVRE 12 (28)
T ss_dssp -TTT--BEEE-
T ss_pred cCCCCCEeEcC
Confidence 88888888743
No 100
>PRK13945 formamidopyrimidine-DNA glycosylase; Provisional
Probab=43.59 E-value=15 Score=35.26 Aligned_cols=25 Identities=36% Similarity=0.658 Sum_probs=17.3
Q ss_pred cCCCCCccccceeccccccccCCCCccceeCCC
Q 019981 273 GPCPNCGTENVSFFGTILSISSGGTTNTINCSN 305 (333)
Q Consensus 273 G~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~ 305 (333)
-|||.||+++..-. -++++ .+-|++
T Consensus 255 ~pC~~Cg~~I~~~~------~~gR~--t~~CP~ 279 (282)
T PRK13945 255 KPCRKCGTPIERIK------LAGRS--THWCPN 279 (282)
T ss_pred CCCCcCCCeeEEEE------ECCCc--cEECCC
Confidence 39999999987543 23444 567887
No 101
>PF09889 DUF2116: Uncharacterized protein containing a Zn-ribbon (DUF2116); InterPro: IPR019216 This entry contains various hypothetical prokaryotic proteins whose functions are unknown. They contain a conserved zinc ribbon motif in the N-terminal part and a predicted transmembrane segment in the C-terminal part.
Probab=43.34 E-value=11 Score=29.30 Aligned_cols=11 Identities=36% Similarity=0.842 Sum_probs=8.8
Q ss_pred cCCCCCccccc
Q 019981 273 GPCPNCGTENV 283 (333)
Q Consensus 273 G~CPNCGeEv~ 283 (333)
.-|||||+++-
T Consensus 4 kHC~~CG~~Ip 14 (59)
T PF09889_consen 4 KHCPVCGKPIP 14 (59)
T ss_pred CcCCcCCCcCC
Confidence 36999998874
No 102
>PF12677 DUF3797: Domain of unknown function (DUF3797); InterPro: IPR024256 This presumed domain is functionally uncharacterised. This domain family is found in bacteria and viruses, and is approximately 50 amino acids in length. There is a conserved CGN sequence motif.
Probab=43.13 E-value=14 Score=28.08 Aligned_cols=12 Identities=42% Similarity=1.129 Sum_probs=10.1
Q ss_pred ecCCCCCccccc
Q 019981 272 KGPCPNCGTENV 283 (333)
Q Consensus 272 KG~CPNCGeEv~ 283 (333)
.+.||+||.+..
T Consensus 13 Y~~Cp~CGN~~v 24 (49)
T PF12677_consen 13 YCKCPKCGNDKV 24 (49)
T ss_pred hccCcccCCcEe
Confidence 789999998753
No 103
>PF08996 zf-DNA_Pol: DNA Polymerase alpha zinc finger; InterPro: IPR015088 Zinc finger (Znf) domains are relatively small protein motifs which contain multiple finger-like protrusions that make tandem contacts with their target molecule. Some of these domains bind zinc, but many do not; instead binding other metals such as iron, or no metal at all. For example, some family members form salt bridges to stabilise the finger-like folds. They were first identified as a DNA-binding motif in transcription factor TFIIIA from Xenopus laevis (African clawed frog), however they are now recognised to bind DNA, RNA, protein and/or lipid substrates [, , , , ]. Their binding properties depend on the amino acid sequence of the finger domains and of the linker between fingers, as well as on the higher-order structures and the number of fingers. Znf domains are often found in clusters, where fingers can have different binding specificities. There are many superfamilies of Znf motifs, varying in both sequence and structure. They display considerable versatility in binding modes, even between members of the same class (e.g. some bind DNA, others protein), suggesting that Znf motifs are stable scaffolds that have evolved specialised functions. For example, Znf-containing proteins function in gene transcription, translation, mRNA trafficking, cytoskeleton organisation, epithelial development, cell adhesion, protein folding, chromatin remodelling and zinc sensing, to name but a few []. Zinc-binding motifs are stable structures, and they rarely undergo conformational changes upon binding their target. The DNA Polymerase alpha zinc finger domain adopts an alpha-helix-like structure, followed by three turns, all of which involve proline. The resulting motif is a helix-turn-helix motif, in contrast to other zinc finger domains, which show anti-parallel sheet and helix conformation. Zinc binding occurs due to the presence of four cysteine residues positioned to bind the metal centre in a tetrahedral coordination geometry. The function of this domain is uncertain: it has been proposed that the zinc finger motif may be an essential part of the DNA binding domain []. More information about these proteins can be found at Protein of the Month: Zinc Fingers [].; GO: 0001882 nucleoside binding, 0003887 DNA-directed DNA polymerase activity, 0006260 DNA replication; PDB: 3FLO_D 1N5G_A 1K0P_A 1K18_A.
Probab=42.16 E-value=10 Score=34.23 Aligned_cols=45 Identities=31% Similarity=0.572 Sum_probs=21.8
Q ss_pred hhcceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCce
Q 019981 264 IVRESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTT 315 (333)
Q Consensus 264 ~~~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~ 315 (333)
-++|..-|+=.||.|++++. |=|...+.......+...|++ |+..
T Consensus 10 rf~~c~~l~~~C~~C~~~~~-f~g~~~~~~~~~~~~~~~C~~------C~~~ 54 (188)
T PF08996_consen 10 RFKDCEPLKLTCPSCGTEFE-FPGVFEEDGDDVSPSGLQCPN------CSTP 54 (188)
T ss_dssp TTTT---EEEE-TTT--EEE-E-SSS--SSEEEETTEEEETT------T--B
T ss_pred HhcCCCceEeECCCCCCCcc-ccccccCCccccccCcCcCCC------CCCc
Confidence 57888899999999999863 332222122233456778887 8873
No 104
>PRK14811 formamidopyrimidine-DNA glycosylase; Provisional
Probab=41.79 E-value=16 Score=34.86 Aligned_cols=28 Identities=36% Similarity=0.793 Sum_probs=19.6
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCce
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTT 315 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~ 315 (333)
|||.||+++..-. -++++ ++-|++ |...
T Consensus 237 pC~~Cg~~I~~~~------~~gR~--ty~Cp~------CQ~~ 264 (269)
T PRK14811 237 PCPRCGTPIEKIV------VGGRG--THFCPQ------CQPL 264 (269)
T ss_pred CCCcCCCeeEEEE------ECCCC--cEECCC------CcCC
Confidence 8999999987543 23444 568887 8754
No 105
>PF10263 SprT-like: SprT-like family; InterPro: IPR006640 This is a family of uncharacterised bacterial proteins which includes Escherichia coli SprT (P39902 from SWISSPROT). SprT is described as a regulator of bolA gene in stationary phase []. The majority of members contain the metallopeptidase zinc binding signature which has a HExxH motif, however there is no evidence for them being metallopeptidases.
Probab=41.54 E-value=22 Score=30.02 Aligned_cols=36 Identities=25% Similarity=0.532 Sum_probs=25.0
Q ss_pred eeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEE
Q 019981 269 LILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVY 318 (333)
Q Consensus 269 ~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f 318 (333)
-.-.-.|++||.++...-. ....+..|.. |+..|++
T Consensus 120 ~~~~~~C~~C~~~~~r~~~--------~~~~~~~C~~------C~~~l~~ 155 (157)
T PF10263_consen 120 KKYVYRCPSCGREYKRHRR--------SKRKRYRCGR------CGGPLVQ 155 (157)
T ss_pred cceEEEcCCCCCEeeeecc--------cchhhEECCC------CCCEEEE
Confidence 3446679999999865542 1334577887 9988875
No 106
>PRK14892 putative transcription elongation factor Elf1; Provisional
Probab=40.78 E-value=20 Score=30.27 Aligned_cols=32 Identities=31% Similarity=0.802 Sum_probs=18.3
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEE
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVY 318 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f 318 (333)
.|||||+.. ++|+=.+....+.|++ ||.--..
T Consensus 23 ~CP~Cge~~-------v~v~~~k~~~h~~C~~------CG~y~~~ 54 (99)
T PRK14892 23 ECPRCGKVS-------ISVKIKKNIAIITCGN------CGLYTEF 54 (99)
T ss_pred ECCCCCCeE-------eeeecCCCcceEECCC------CCCccCE
Confidence 599999532 2222233445566666 9976443
No 107
>PF12760 Zn_Tnp_IS1595: Transposase zinc-ribbon domain; InterPro: IPR024442 This zinc binding domain is found in a range of transposase proteins such as ISSPO8, ISSOD11, ISRSSP2 etc. It may be a zinc-binding beta ribbon domain that could bind DNA.
Probab=40.74 E-value=23 Score=25.17 Aligned_cols=22 Identities=27% Similarity=0.724 Sum_probs=15.4
Q ss_pred CCCCccccceeccccccccCCCCccceeCCC
Q 019981 275 CPNCGTENVSFFGTILSISSGGTTNTINCSN 305 (333)
Q Consensus 275 CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~ 305 (333)
||.||......+. +....+|..
T Consensus 21 CP~Cg~~~~~~~~---------~~~~~~C~~ 42 (46)
T PF12760_consen 21 CPHCGSTKHYRLK---------TRGRYRCKA 42 (46)
T ss_pred CCCCCCeeeEEeC---------CCCeEECCC
Confidence 9999998333332 267888877
No 108
>PF05502 Dynactin_p62: Dynactin p62 family; InterPro: IPR008603 Dynactin is a multi-subunit complex and a required cofactor for most, or all, o f the cellular processes powered by the microtubule-based motor cytoplasmic dyn ein. p62 binds directly to the Arp1 subunit of dynactin [, ].
Probab=40.41 E-value=17 Score=37.86 Aligned_cols=45 Identities=27% Similarity=0.312 Sum_probs=27.8
Q ss_pred eeeecCCCCCccccceeccccccccCCCCccceeCC-CCccccccCceeEEeccc
Q 019981 269 LILKGPCPNCGTENVSFFGTILSISSGGTTNTINCS-NLTFCFSCGTTMVYDSNT 322 (333)
Q Consensus 269 ~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~-~~aeCHVC~t~L~f~tk~ 322 (333)
.|..--||||-+++-+= .......+|. |=-+|.+|...|...+..
T Consensus 23 Ei~~~yCp~CL~~~p~~---------e~~~~~nrC~r~Cf~CP~C~~~L~~~~~~ 68 (483)
T PF05502_consen 23 EIDSYYCPNCLFEVPSS---------EARSEKNRCSRNCFDCPICFSPLSVRASD 68 (483)
T ss_pred ccceeECccccccCChh---------hheeccceeccccccCCCCCCcceeEecc
Confidence 34455699999887521 1112233564 445677799999988543
No 109
>TIGR01054 rgy reverse gyrase. Generally, these gyrases are encoded as a single polypeptide. An exception was found in Methanopyrus kandleri, where enzyme is split within the topoisomerase domain, yielding a heterodimer of gene products designated RgyB and RgyA.
Probab=39.74 E-value=12 Score=42.81 Aligned_cols=17 Identities=41% Similarity=0.747 Sum_probs=14.3
Q ss_pred eeeecCCCCCcccccee
Q 019981 269 LILKGPCPNCGTENVSF 285 (333)
Q Consensus 269 ~iLKG~CPNCGeEv~aF 285 (333)
.+-++.|||||.++.+.
T Consensus 4 ~~y~~~CPnCgg~i~~~ 20 (1171)
T TIGR01054 4 AVYSNLCPNCGGEISSE 20 (1171)
T ss_pred chhcCCCCCCCCccchh
Confidence 35689999999999875
No 110
>PRK00398 rpoP DNA-directed RNA polymerase subunit P; Provisional
Probab=39.68 E-value=29 Score=24.53 Aligned_cols=24 Identities=21% Similarity=0.480 Sum_probs=20.2
Q ss_pred ceeCCCCccccccCceeEEeccceeeecCC
Q 019981 300 TINCSNLTFCFSCGTTMVYDSNTRLITLPE 329 (333)
Q Consensus 300 ~vkC~~~aeCHVC~t~L~f~tk~R~itl~~ 329 (333)
.++|++ ||..++|+.....++.|.
T Consensus 3 ~y~C~~------CG~~~~~~~~~~~~~Cp~ 26 (46)
T PRK00398 3 EYKCAR------CGREVELDEYGTGVRCPY 26 (46)
T ss_pred EEECCC------CCCEEEECCCCCceECCC
Confidence 578999 999999998776777775
No 111
>PF06677 Auto_anti-p27: Sjogren's syndrome/scleroderma autoantigen 1 (Autoantigen p27); InterPro: IPR009563 The proteins in this entry are functionally uncharacterised and include several proteins that characterise Sjogren's syndrome/scleroderma autoantigen 1 (Autoantigen p27). It is thought that the potential association of anti-p27 with anti-centromere antibodies suggests that autoantigen p27 might play a role in mitosis [].
Probab=38.76 E-value=27 Score=25.16 Aligned_cols=25 Identities=24% Similarity=0.647 Sum_probs=18.5
Q ss_pred HHhhhhcceeeeecCCCCCccccce
Q 019981 260 LTKLIVRESLILKGPCPNCGTENVS 284 (333)
Q Consensus 260 Lt~l~~~D~~iLKG~CPNCGeEv~a 284 (333)
+..++.+=...|--.||.||.+.+.
T Consensus 5 m~~~LL~G~~ML~~~Cp~C~~PL~~ 29 (41)
T PF06677_consen 5 MGEYLLQGWTMLDEHCPDCGTPLMR 29 (41)
T ss_pred HHHHHHHhHhHhcCccCCCCCeeEE
Confidence 4455555667788899999988775
No 112
>PRK00464 nrdR transcriptional regulator NrdR; Validated
Probab=37.81 E-value=26 Score=31.54 Aligned_cols=37 Identities=19% Similarity=0.533 Sum_probs=20.2
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCcee
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTM 316 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L 316 (333)
-||-||.+-..-.-+..-=+||.-.-...|++ ||+..
T Consensus 2 ~cp~c~~~~~~~~~s~~~~~~~~~~~~~~c~~------c~~~f 38 (154)
T PRK00464 2 RCPFCGHPDTRVIDSRPAEDGNAIRRRRECLA------CGKRF 38 (154)
T ss_pred cCCCCCCCCCEeEeccccCCCCceeeeeeccc------cCCcc
Confidence 49999986543332221112223333367887 99864
No 113
>PF11793 FANCL_C: FANCL C-terminal domain; PDB: 3K1L_A.
Probab=37.38 E-value=11 Score=29.20 Aligned_cols=19 Identities=21% Similarity=0.518 Sum_probs=11.4
Q ss_pred hcceeeeecCCCCCccccc
Q 019981 265 VRESLILKGPCPNCGTENV 283 (333)
Q Consensus 265 ~~D~~iLKG~CPNCGeEv~ 283 (333)
++.+..+.|.||.|.+++.
T Consensus 48 ~~~~~~~~G~CP~C~~~i~ 66 (70)
T PF11793_consen 48 RQSFIPIFGECPYCSSPIS 66 (70)
T ss_dssp S-TTT--EEE-TTT-SEEE
T ss_pred CeeecccccCCcCCCCeee
Confidence 4557788999999999875
No 114
>PLN02919 haloacid dehalogenase-like hydrolase family protein
Probab=37.32 E-value=2.5e+02 Score=32.03 Aligned_cols=30 Identities=20% Similarity=0.135 Sum_probs=14.7
Q ss_pred chhHHHHHHHHHHHHhhhcCccccChHHHh
Q 019981 94 SLGELEQEFLQALQAFYYEGKAVMSNEEFD 123 (333)
Q Consensus 94 slge~E~~fl~Al~~fY~~gk~~~sdeefd 123 (333)
||-+-|..+.+|++..+.+---.++.+++.
T Consensus 85 TLiDS~~~~~~a~~~~~~~~G~~it~e~~~ 114 (1057)
T PLN02919 85 VLCNSEEPSRRAAVDVFAEMGVEVTVEDFV 114 (1057)
T ss_pred CeEeChHHHHHHHHHHHHHcCCCCCHHHHH
Confidence 444445555566555554322234555553
No 115
>PF04380 BMFP: Membrane fusogenic activity; InterPro: IPR007475 BMFP consists of two structural domains, a coiled-coil C-terminal domain via which the protein self-associates as a trimer, and an N-terminal domain disordered at neutral pH but adopting an amphipathic alpha-helical structure in the presence of phospholipid vesicles, high ionic strength, acidic pH or SDS. BMFP interacts with phospholipid vesicles though the predicted amphipathic alpha-helix induced in the N-terminal half of the protein and promotes aggregation and fusion of vesicles in vitro.
Probab=36.75 E-value=55 Score=26.10 Aligned_cols=36 Identities=28% Similarity=0.344 Sum_probs=30.5
Q ss_pred chhHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhh
Q 019981 94 SLGELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELM 130 (333)
Q Consensus 94 slge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~ 130 (333)
...|.|..+-..+++.+.. ....|.||||.+++.|.
T Consensus 25 ~~~e~e~~~r~~l~~~l~k-ldlVtREEFd~q~~~L~ 60 (79)
T PF04380_consen 25 PREEIEKNIRARLQSALSK-LDLVTREEFDAQKAVLA 60 (79)
T ss_pred hHHHHHHHHHHHHHHHHHH-CCCCcHHHHHHHHHHHH
Confidence 4467899999999999875 77899999999999853
No 116
>PF05129 Elf1: Transcription elongation factor Elf1 like; InterPro: IPR007808 This family of uncharacterised, mostly short, proteins contain a putative zinc binding domain with four conserved cysteines.; PDB: 1WII_A.
Probab=36.68 E-value=28 Score=28.10 Aligned_cols=37 Identities=19% Similarity=0.390 Sum_probs=18.5
Q ss_pred cCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEe
Q 019981 273 GPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 273 G~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~ 319 (333)
=.||.|+.+.-.=+.-+ .......+.|.+ ||..-++.
T Consensus 23 F~CPfC~~~~sV~v~id----kk~~~~~~~C~~------Cg~~~~~~ 59 (81)
T PF05129_consen 23 FDCPFCNHEKSVSVKID----KKEGIGILSCRV------CGESFQTK 59 (81)
T ss_dssp ---TTT--SS-EEEEEE----TTTTEEEEEESS------S--EEEEE
T ss_pred EcCCcCCCCCeEEEEEE----ccCCEEEEEecC------CCCeEEEc
Confidence 36999997766555332 235667788887 87655544
No 117
>PRK04023 DNA polymerase II large subunit; Validated
Probab=36.46 E-value=23 Score=40.59 Aligned_cols=17 Identities=35% Similarity=0.682 Sum_probs=10.6
Q ss_pred eeecCCCCCccccceec
Q 019981 270 ILKGPCPNCGTENVSFF 286 (333)
Q Consensus 270 iLKG~CPNCGeEv~aFf 286 (333)
+-.--||.||++.+.|+
T Consensus 624 Vg~RfCpsCG~~t~~fr 640 (1121)
T PRK04023 624 IGRRKCPSCGKETFYRR 640 (1121)
T ss_pred ccCccCCCCCCcCCccc
Confidence 33446888888764443
No 118
>PLN00049 carboxyl-terminal processing protease; Provisional
Probab=36.45 E-value=70 Score=31.92 Aligned_cols=71 Identities=17% Similarity=0.247 Sum_probs=45.2
Q ss_pred hhHHHHHHHHH---HHHhhhcCccccChHHHhhhHhhhhhcCCeeEEeChhhHHHHHHHH---h-hhcCCC-ccChHHHH
Q 019981 95 LGELEQEFLQA---LQAFYYEGKAVMSNEEFDNLKEELMWEGSSVVMLSSAEQKFLEASM---A-YVAGKP-IMSDEEYD 166 (333)
Q Consensus 95 lge~E~~fl~A---l~~fY~~gk~~~sdeefd~LkEeL~weGSsvv~L~~~Eq~fLEA~~---A-Y~sGkP-imsDeeFD 166 (333)
+-|..|+|.+| ++.+||+.+ |...+|+.+||...|.- . +...+ ++..|+. + .-+.-- .++-++|.
T Consensus 2 ~~~~~~~f~e~w~~v~~~~~d~~--~~g~dW~~~~e~y~~~~-~---~~~~~-~~~~~i~~ml~~L~D~hs~y~~~~~~~ 74 (389)
T PLN00049 2 LTEENLLFLEAWRTVDRAYVDKT--FNGQSWFRYRENALKNE-P---MNTRE-ETYAAIRKMLATLDDPFTRFLEPEKFK 74 (389)
T ss_pred CccHHHHHHHHHHHHHHHHcCcc--ccccCHHHHHHHHhhcc-C---CCcHH-HHHHHHHHHHhhCCCCcccCcCHHHHH
Confidence 34678999998 567888765 89999999999999964 2 22222 3333322 1 101111 66788888
Q ss_pred HHHHHH
Q 019981 167 KLKQKL 172 (333)
Q Consensus 167 ~LK~kL 172 (333)
.+....
T Consensus 75 ~~~~~~ 80 (389)
T PLN00049 75 SLRSGT 80 (389)
T ss_pred HHHHhc
Confidence 776543
No 119
>PRK12495 hypothetical protein; Provisional
Probab=35.70 E-value=32 Score=33.12 Aligned_cols=28 Identities=14% Similarity=0.522 Sum_probs=22.9
Q ss_pred HHHHhhhhcceeeeecCCCCCcccccee
Q 019981 258 QSLTKLIVRESLILKGPCPNCGTENVSF 285 (333)
Q Consensus 258 ~~Lt~l~~~D~~iLKG~CPNCGeEv~aF 285 (333)
+.+..|+++=...+---||.||.++|.+
T Consensus 28 ~~ma~lL~~gatmsa~hC~~CG~PIpa~ 55 (226)
T PRK12495 28 ERMSELLLQGATMTNAHCDECGDPIFRH 55 (226)
T ss_pred HHHHHHHHhhcccchhhcccccCcccCC
Confidence 4466777777888888999999999955
No 120
>COG3877 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=35.12 E-value=29 Score=30.49 Aligned_cols=26 Identities=38% Similarity=0.888 Sum_probs=19.5
Q ss_pred ecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeE
Q 019981 272 KGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMV 317 (333)
Q Consensus 272 KG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~ 317 (333)
--+||-||++... -+.+|++ ||++..
T Consensus 6 ~~~cPvcg~~~iV--------------TeL~c~~------~etTVr 31 (122)
T COG3877 6 INRCPVCGRKLIV--------------TELKCSN------CETTVR 31 (122)
T ss_pred CCCCCccccccee--------------EEEecCC------CCceEe
Confidence 3579999987542 3679999 998764
No 121
>COG0675 Transposase and inactivated derivatives [DNA replication, recombination, and repair]
Probab=34.89 E-value=25 Score=31.90 Aligned_cols=22 Identities=32% Similarity=0.990 Sum_probs=15.5
Q ss_pred cCCCCCccccceeccccccccCCCCccceeCCCCccccccCce
Q 019981 273 GPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTT 315 (333)
Q Consensus 273 G~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~ 315 (333)
-.||+||. + ..-..+|++ ||..
T Consensus 310 ~~C~~cg~----~-----------~~r~~~C~~------cg~~ 331 (364)
T COG0675 310 KTCPCCGH----L-----------SGRLFKCPR------CGFV 331 (364)
T ss_pred ccccccCC----c-----------cceeEECCC------CCCe
Confidence 46999999 1 223568888 8864
No 122
>PF08209 Sgf11: Sgf11 (transcriptional regulation protein); InterPro: IPR013246 The Sgf11 family is a SAGA complex subunit in Saccharomyces cerevisiae (Baker's yeast). The SAGA complex is a multisubunit protein complex involved in transcriptional regulation. SAGA combines proteins involved in interactions with DNA-bound activators and TATA-binding protein (TBP), as well as enzymes for histone acetylation and deubiquitylation [].; PDB: 3M99_B 2LO2_A 3MHH_C 3MHS_C.
Probab=34.69 E-value=16 Score=25.32 Aligned_cols=10 Identities=50% Similarity=1.255 Sum_probs=7.5
Q ss_pred CCCCCccccc
Q 019981 274 PCPNCGTENV 283 (333)
Q Consensus 274 ~CPNCGeEv~ 283 (333)
.||||+-+|-
T Consensus 6 ~C~nC~R~v~ 15 (33)
T PF08209_consen 6 ECPNCGRPVA 15 (33)
T ss_dssp E-TTTSSEEE
T ss_pred ECCCCcCCcc
Confidence 5999998875
No 123
>PF04280 Tim44: Tim44-like domain; InterPro: IPR007379 Tim44 is an essential component of the machinery that mediates the translocation of nuclear-encoded proteins across the mitochondrial inner membrane []. Tim44 is thought to bind phospholipids of the mitochondrial inner membrane both by electrostatic interactions and by penetrating the polar head group region [].; GO: 0015450 P-P-bond-hydrolysis-driven protein transmembrane transporter activity, 0006886 intracellular protein transport, 0005744 mitochondrial inner membrane presequence translocase complex; PDB: 2CW9_A 2FXT_A 3QK9_A.
Probab=34.14 E-value=21 Score=29.77 Aligned_cols=37 Identities=30% Similarity=0.668 Sum_probs=29.9
Q ss_pred eChhhHHHHHHHHhhhcCC-----CccChHHHHHHHHHHhhc
Q 019981 139 LSSAEQKFLEASMAYVAGK-----PIMSDEEYDKLKQKLKME 175 (333)
Q Consensus 139 L~~~Eq~fLEA~~AY~sGk-----PimsDeeFD~LK~kLk~~ 175 (333)
+...++.|+....||.+|+ +++++++|..++.++++.
T Consensus 21 ~~~ak~~f~~i~~A~~~~D~~~l~~~~t~~~~~~~~~~i~~~ 62 (147)
T PF04280_consen 21 LEEAKEAFLPIQEAWAKGDLEALRPLLTEELYERLQAEIKAR 62 (147)
T ss_dssp HHHHHHTHHHHHHHHHHT-HHHHHHHB-HHHHHHHHHHHHHH
T ss_pred HHHHHHHHHHHHHHHHcCCHHHHHHHhCHHHHHHHHHHHHHH
Confidence 4456777888777899985 899999999999999988
No 124
>PRK12412 pyridoxal kinase; Reviewed
Probab=33.92 E-value=1.5e+02 Score=27.62 Aligned_cols=57 Identities=26% Similarity=0.334 Sum_probs=42.1
Q ss_pred hHHHhhhHhhhhhcCCeeEEeChhhHHHHHHHHhhhcCCCccChHHHHHHHHHHhhcCCc-eeeec
Q 019981 119 NEEFDNLKEELMWEGSSVVMLSSAEQKFLEASMAYVAGKPIMSDEEYDKLKQKLKMEGSE-IVVEG 183 (333)
Q Consensus 119 deefd~LkEeL~weGSsvv~L~~~Eq~fLEA~~AY~sGkPimsDeeFD~LK~kLk~~GS~-VVvk~ 183 (333)
++..+.++++|. ....+++.+..|-+.| .|.++-+.++..+.-.+|...|-+ |++++
T Consensus 119 ~~~~~~~~~~ll-~~advitpN~~Ea~~L-------~g~~~~~~~~~~~aa~~l~~~g~~~ViIt~ 176 (268)
T PRK12412 119 PETNDCLRDVLV-PKALVVTPNLFEAYQL-------SGVKINSLEDMKEAAKKIHALGAKYVLIKG 176 (268)
T ss_pred hHHHHHHHHhhh-ccceEEcCCHHHHHHH-------hCcCCCCHHHHHHHHHHHHhcCCCEEEEec
Confidence 345567787765 5678999999998877 477777777777777888877864 66664
No 125
>cd07110 ALDH_F10_BADH Arabidopsis betaine aldehyde dehydrogenase 1 and 2, ALDH family 10A8 and 10A9-like. Present in this CD are the Arabidopsis betaine aldehyde dehydrogenase (BADH) 1 (chloroplast) and 2 (mitochondria), also known as, aldehyde dehydrogenase family 10 member A8 and aldehyde dehydrogenase family 10 member A9, respectively, and are putative dehydration- and salt-inducible BADHs (EC 1.2.1.8) that catalyze the oxidation of betaine aldehyde to the compatible solute glycine betaine.
Probab=33.12 E-value=1.2e+02 Score=30.35 Aligned_cols=70 Identities=23% Similarity=0.464 Sum_probs=47.5
Q ss_pred ccChHHHhhhHhhhhhc-----CCe-----eEEeCh-hhHHHHHHHHh----hhcCC---------CccChHHHHHHHHH
Q 019981 116 VMSNEEFDNLKEELMWE-----GSS-----VVMLSS-AEQKFLEASMA----YVAGK---------PIMSDEEYDKLKQK 171 (333)
Q Consensus 116 ~~sdeefd~LkEeL~we-----GSs-----vv~L~~-~Eq~fLEA~~A----Y~sGk---------PimsDeeFD~LK~k 171 (333)
++.|.+.|..-+.+.|. |.. .+.+-+ .-.+|++++.+ +.-|. |+++.+.+++++.-
T Consensus 238 V~~dadl~~aa~~i~~~~~~~~GQ~C~a~~rv~V~~~i~d~f~~~l~~~~~~~~~g~p~~~~~~~Gpli~~~~~~~~~~~ 317 (456)
T cd07110 238 VFDDADLEKAVEWAMFGCFWNNGQICSATSRLLVHESIADAFLERLATAAEAIRVGDPLEEGVRLGPLVSQAQYEKVLSF 317 (456)
T ss_pred ECCCCCHHHHHHHHHHHHHhcCCCCCCCCceEEEcHHHHHHHHHHHHHHHHhcCCCCCCCCCCCcCCCCCHHHHHHHHHH
Confidence 45678888888888773 433 344443 34578888654 33443 68899999999988
Q ss_pred Hhh---cCCceeeecCe
Q 019981 172 LKM---EGSEIVVEGPR 185 (333)
Q Consensus 172 Lk~---~GS~VVvk~Pr 185 (333)
+.+ .|.+++.-|.+
T Consensus 318 v~~a~~~Ga~~~~gg~~ 334 (456)
T cd07110 318 IARGKEEGARLLCGGRR 334 (456)
T ss_pred HHHHHhCCCEEEeCCCc
Confidence 865 67787775543
No 126
>TIGR01562 FdhE formate dehydrogenase accessory protein FdhE. The only sequence scoring between trusted and noise is that from Aquifex aeolicus, which shows certain structural differences from the proteobacterial forms in the alignment. However it is notable that A. aeolicus also has a sequence scoring above trusted to the alpha subunit of formate dehydrogenase (TIGR01553).
Probab=32.94 E-value=27 Score=34.59 Aligned_cols=11 Identities=18% Similarity=0.712 Sum_probs=7.7
Q ss_pred ecCCCCCcccc
Q 019981 272 KGPCPNCGTEN 282 (333)
Q Consensus 272 KG~CPNCGeEv 282 (333)
-.-||+||+.-
T Consensus 224 R~~C~~Cg~~~ 234 (305)
T TIGR01562 224 RVKCSHCEESK 234 (305)
T ss_pred CccCCCCCCCC
Confidence 55688888764
No 127
>PRK08351 DNA-directed RNA polymerase subunit E''; Validated
Probab=32.82 E-value=20 Score=28.03 Aligned_cols=16 Identities=31% Similarity=1.097 Sum_probs=12.0
Q ss_pred CCCCCccccc--eecccc
Q 019981 274 PCPNCGTENV--SFFGTI 289 (333)
Q Consensus 274 ~CPNCGeEv~--aFfgti 289 (333)
-|||||.+.+ .+||-+
T Consensus 17 ~CP~Cgs~~~T~~W~G~v 34 (61)
T PRK08351 17 RCPVCGSRDLSDEWFDLV 34 (61)
T ss_pred cCCCCcCCccccccccEE
Confidence 5999999875 566643
No 128
>COG4443 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=32.45 E-value=28 Score=28.20 Aligned_cols=18 Identities=50% Similarity=0.772 Sum_probs=15.0
Q ss_pred Ccc-ccChHHHhhhHhhhh
Q 019981 113 GKA-VMSNEEFDNLKEELM 130 (333)
Q Consensus 113 gk~-~~sdeefd~LkEeL~ 130 (333)
||- +||||||..||+.|+
T Consensus 52 GKGiTLt~eE~~~l~d~l~ 70 (72)
T COG4443 52 GKGITLTNEEFKALKDLLN 70 (72)
T ss_pred cCceeecHHHHHHHHHHHh
Confidence 444 899999999999874
No 129
>PRK00133 metG methionyl-tRNA synthetase; Reviewed
Probab=31.92 E-value=40 Score=35.97 Aligned_cols=16 Identities=19% Similarity=0.119 Sum_probs=11.7
Q ss_pred cccccCceeEEeccce
Q 019981 308 FCFSCGTTMVYDSNTR 323 (333)
Q Consensus 308 eCHVC~t~L~f~tk~R 323 (333)
.|..||.+++++....
T Consensus 171 ~~~~~g~~~e~~~~~~ 186 (673)
T PRK00133 171 KSAISGATPVLKESEH 186 (673)
T ss_pred ccccCCCcceEEecce
Confidence 3666999999886543
No 130
>PF06170 DUF983: Protein of unknown function (DUF983); InterPro: IPR009325 This family consists of several bacterial proteins of unknown function.
Probab=31.31 E-value=21 Score=29.22 Aligned_cols=17 Identities=29% Similarity=0.696 Sum_probs=12.2
Q ss_pred ceeeeecCCCCCccccc
Q 019981 267 ESLILKGPCPNCGTENV 283 (333)
Q Consensus 267 D~~iLKG~CPNCGeEv~ 283 (333)
..+-+.-.||+||++-.
T Consensus 3 g~Lk~~~~C~~CG~d~~ 19 (86)
T PF06170_consen 3 GYLKVAPRCPHCGLDYS 19 (86)
T ss_pred ccccCCCcccccCCccc
Confidence 45567778999988754
No 131
>PF09567 RE_MamI: MamI restriction endonuclease; InterPro: IPR019067 There are four classes of restriction endonucleases: types I, II,III and IV. All types of enzymes recognise specific short DNA sequences and carry out the endonucleolytic cleavage of DNA to give specific double-stranded fragments with terminal 5'-phosphates. They differ in their recognition sequence, subunit composition, cleavage position, and cofactor requirements [, ], as summarised below: Type I enzymes (3.1.21.3 from EC) cleave at sites remote from recognition site; require both ATP and S-adenosyl-L-methionine to function; multifunctional protein with both restriction and methylase (2.1.1.72 from EC) activities. Type II enzymes (3.1.21.4 from EC) cleave within or at short specific distances from recognition site; most require magnesium; single function (restriction) enzymes independent of methylase. Type III enzymes (3.1.21.5 from EC) cleave at sites a short distance from recognition site; require ATP (but doesn't hydrolyse it); S-adenosyl-L-methionine stimulates reaction but is not required; exists as part of a complex with a modification methylase methylase (2.1.1.72 from EC). Type IV enzymes target methylated DNA. Type II restriction endonucleases (3.1.21.4 from EC) are components of prokaryotic DNA restriction-modification mechanisms that protect the organism against invading foreign DNA. These site-specific deoxyribonucleases catalyse the endonucleolytic cleavage of DNA to give specific double-stranded fragments with terminal 5'-phosphates. Of the 3000 restriction endonucleases that have been characterised, most are homodimeric or tetrameric enzymes that cleave target DNA at sequence-specific sites close to the recognition site. For homodimeric enzymes, the recognition site is usually a palindromic sequence 4-8 bp in length. Most enzymes require magnesium ions as a cofactor for catalysis. Although they can vary in their mode of recognition, many restriction endonucleases share a similar structural core comprising four beta-strands and one alpha-helix, as well as a similar mechanism of cleavage, suggesting a common ancestral origin []. However, there is still considerable diversity amongst restriction endonucleases [, ]. The target site recognition process triggers large conformational changes of the enzyme and the target DNA, leading to the activation of the catalytic centres. Like other DNA binding proteins, restriction enzymes are capable of non-specific DNA binding as well, which is the prerequisite for efficient target site location by facilitated diffusion. Non-specific binding usually does not involve interactions with the bases but only with the DNA backbone []. This entry includes the MamI restriction endonuclease which recognises and cleaves GATNN^NNATC. ; GO: 0003677 DNA binding, 0009036 Type II site-specific deoxyribonuclease activity, 0009307 DNA restriction-modification system
Probab=31.12 E-value=21 Score=35.36 Aligned_cols=13 Identities=38% Similarity=0.989 Sum_probs=11.6
Q ss_pred cCCCCCcccccee
Q 019981 273 GPCPNCGTENVSF 285 (333)
Q Consensus 273 G~CPNCGeEv~aF 285 (333)
|.|-|||.+|.++
T Consensus 83 ~~C~~CGa~V~~~ 95 (314)
T PF09567_consen 83 GKCNNCGANVSRL 95 (314)
T ss_pred hhhccccceeeeh
Confidence 7899999999887
No 132
>cd01169 HMPP_kinase 4-amino-5-hydroxymethyl-2-methyl-pyrimidine phosphate kinase (HMPP-kinase) catalyzes two consecutive phosphorylation steps in the thiamine phosphate biosynthesis pathway, leading to the synthesis of vitamin B1. The first step is the phosphorylation of the hydroxyl group of HMP to form 4-amino-5-hydroxymethyl-2-methyl-pyrimidine phosphate (HMP-P) and then the phophorylation of HMP-P to form 4-amino-5-hydroxymethyl-2-methyl-pyrimidine pyrophosphate (HMP-PP), which is the substrate for the thiamine synthase coupling reaction.
Probab=31.11 E-value=2.1e+02 Score=25.36 Aligned_cols=53 Identities=23% Similarity=0.394 Sum_probs=37.0
Q ss_pred hhhHhhhhhcCCeeEEeChhhHHHHHHHHhhhcCCCccChHHHHHHHHHHhhcCC-ceeeec
Q 019981 123 DNLKEELMWEGSSVVMLSSAEQKFLEASMAYVAGKPIMSDEEYDKLKQKLKMEGS-EIVVEG 183 (333)
Q Consensus 123 d~LkEeL~weGSsvv~L~~~Eq~fLEA~~AY~sGkPimsDeeFD~LK~kLk~~GS-~VVvk~ 183 (333)
+.++++| +....+++.+..|-+.| .|.++-++++-.+...+|...|- .|++++
T Consensus 119 ~~~~~~l-l~~~dvitpN~~Ea~~L-------~g~~~~~~~~~~~~~~~l~~~g~~~Vvit~ 172 (242)
T cd01169 119 EALRELL-LPLATLITPNLPEAELL-------TGLEIATEEDMMKAAKALLALGAKAVLIKG 172 (242)
T ss_pred HHHHHHh-hccCeEEeCCHHHHHHH-------hCCCCCCHHHHHHHHHHHHhcCCCEEEEec
Confidence 4566654 67788999999998877 36666666555556677777775 466664
No 133
>PF06221 zf-C2HC5: Putative zinc finger motif, C2HC5-type; InterPro: IPR009349 Zinc finger (Znf) domains are relatively small protein motifs which contain multiple finger-like protrusions that make tandem contacts with their target molecule. Some of these domains bind zinc, but many do not; instead binding other metals such as iron, or no metal at all. For example, some family members form salt bridges to stabilise the finger-like folds. They were first identified as a DNA-binding motif in transcription factor TFIIIA from Xenopus laevis (African clawed frog), however they are now recognised to bind DNA, RNA, protein and/or lipid substrates [, , , , ]. Their binding properties depend on the amino acid sequence of the finger domains and of the linker between fingers, as well as on the higher-order structures and the number of fingers. Znf domains are often found in clusters, where fingers can have different binding specificities. There are many superfamilies of Znf motifs, varying in both sequence and structure. They display considerable versatility in binding modes, even between members of the same class (e.g. some bind DNA, others protein), suggesting that Znf motifs are stable scaffolds that have evolved specialised functions. For example, Znf-containing proteins function in gene transcription, translation, mRNA trafficking, cytoskeleton organisation, epithelial development, cell adhesion, protein folding, chromatin remodelling and zinc sensing, to name but a few []. Zinc-binding motifs are stable structures, and they rarely undergo conformational changes upon binding their target. This zinc finger appears to be common in activating signal cointegrator 1/thyroid receptor interacting protein 4. More information about these proteins can be found at Protein of the Month: Zinc Fingers [].; GO: 0008270 zinc ion binding, 0006355 regulation of transcription, DNA-dependent, 0005634 nucleus
Probab=30.56 E-value=25 Score=27.17 Aligned_cols=13 Identities=62% Similarity=1.260 Sum_probs=11.4
Q ss_pred ecCCCCCccccce
Q 019981 272 KGPCPNCGTENVS 284 (333)
Q Consensus 272 KG~CPNCGeEv~a 284 (333)
.||||-||+++.+
T Consensus 35 ~~pC~fCg~~l~~ 47 (57)
T PF06221_consen 35 LGPCPFCGTPLLS 47 (57)
T ss_pred cCcCCCCCCcccC
Confidence 7999999998875
No 134
>PF05416 Peptidase_C37: Southampton virus-type processing peptidase; InterPro: IPR001665 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This group of cysteine peptidases belong to the MEROPS peptidase family C37, (clan PA(C)). The type example is calicivirin from Southampton virus, an endopeptidase that cleaves the polyprotein at sites N-terminal to itself, liberating the polyprotein helicase. Southampton virus is a positive-stranded ssRNA virus belonging to the Caliciviruses, which are viruses that cause gastroenteritis. The calicivirus genome contains two open reading frames, ORF1 and ORF2. ORF1 encodes a non-structural polypeptide, which has RNA helicase, cysteine protease and RNA polymerase activity []. The regions of the polyprotein in which these activities lie are similar to proteins produced by the picornaviruses []. ORF2 encodes a structural, capsid protein. Two different families of caliciviruses can be distinguished on the basis of sequence similarity, namely the Norwalk-like viruses or small round structured viruses (SRSVs), and those classed as non-SRSVs.; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis; PDB: 2FYQ_A 2FYR_A 1WQS_D 4ASH_A 2IPH_B.
Probab=30.25 E-value=17 Score=38.33 Aligned_cols=43 Identities=26% Similarity=0.382 Sum_probs=0.0
Q ss_pred ccChHHHhh---hHhhhhhcCCeeEEeChhhHHHHHHHHhhhcCCCccChHHHH
Q 019981 116 VMSNEEFDN---LKEELMWEGSSVVMLSSAEQKFLEASMAYVAGKPIMSDEEYD 166 (333)
Q Consensus 116 ~~sdeefd~---LkEeL~weGSsvv~L~~~Eq~fLEA~~AY~sGkPimsDeeFD 166 (333)
-|||||||. |||| |.|.-- =|+|||..+-||+.-.+..-.++|
T Consensus 252 GLSDEEYDEyKkiREe--r~g~YS------IeEYLqdReRy~Eela~~~a~~~~ 297 (535)
T PF05416_consen 252 GLSDEEYDEYKKIREE--RGGKYS------IEEYLQDRERYEEELAEAQATEED 297 (535)
T ss_dssp ------------------------------------------------------
T ss_pred CCChhHHHHHHHHHHH--hcCCcc------HHHHHHHHHHHHHHhhhhhhhhcc
Confidence 399999885 5666 665421 178999999999887766544444
No 135
>TIGR01031 rpmF_bact ribosomal protein L32. This protein describes bacterial ribosomal protein L32. The noise cutoff is set low enough to include the equivalent protein from mitochondria and chloroplasts. No related proteins from the Archaea nor from the eukaryotic cytosol are detected by this model. This model is a fragment model; the putative L32 of some species shows similarity only toward the N-terminus.
Probab=29.94 E-value=32 Score=26.02 Aligned_cols=13 Identities=38% Similarity=0.892 Sum_probs=9.4
Q ss_pred cCCCCCcccccee
Q 019981 273 GPCPNCGTENVSF 285 (333)
Q Consensus 273 G~CPNCGeEv~aF 285 (333)
..||+||+....-
T Consensus 27 ~~C~~cG~~~~~H 39 (55)
T TIGR01031 27 VVCPNCGEFKLPH 39 (55)
T ss_pred eECCCCCCcccCe
Confidence 3499999966544
No 136
>PRK12286 rpmF 50S ribosomal protein L32; Reviewed
Probab=29.66 E-value=34 Score=26.14 Aligned_cols=12 Identities=42% Similarity=1.099 Sum_probs=9.3
Q ss_pred cCCCCCccccce
Q 019981 273 GPCPNCGTENVS 284 (333)
Q Consensus 273 G~CPNCGeEv~a 284 (333)
-.||+||+-...
T Consensus 28 ~~C~~CG~~~~~ 39 (57)
T PRK12286 28 VECPNCGEPKLP 39 (57)
T ss_pred eECCCCCCccCC
Confidence 359999987665
No 137
>COG1996 RPC10 DNA-directed RNA polymerase, subunit RPC10 (contains C4-type Zn-finger) [Transcription]
Probab=29.18 E-value=41 Score=25.40 Aligned_cols=30 Identities=30% Similarity=0.747 Sum_probs=22.4
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEe
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~ 319 (333)
-|-.||.++ . .....-.++|+. ||..+-|-
T Consensus 8 ~C~~Cg~~~-----~-----~~~~~~~irCp~------Cg~rIl~K 37 (49)
T COG1996 8 KCARCGREV-----E-----LDQETRGIRCPY------CGSRILVK 37 (49)
T ss_pred EhhhcCCee-----e-----hhhccCceeCCC------CCcEEEEe
Confidence 488999998 1 123455789999 99988886
No 138
>PF09863 DUF2090: Uncharacterized protein conserved in bacteria (DUF2090); InterPro: IPR018659 This domain, found in various prokaryotic carbohydrate kinases, has no known function.
Probab=29.03 E-value=1.4e+02 Score=30.06 Aligned_cols=49 Identities=16% Similarity=0.276 Sum_probs=40.1
Q ss_pred chhHHHHHHHHHHHHhhhcCc-------cccChHHHhhhHhhhhhcCCe---eEEeChh
Q 019981 94 SLGELEQEFLQALQAFYYEGK-------AVMSNEEFDNLKEELMWEGSS---VVMLSSA 142 (333)
Q Consensus 94 slge~E~~fl~Al~~fY~~gk-------~~~sdeefd~LkEeL~weGSs---vv~L~~~ 142 (333)
+....+.-|..+|++||+-|- +-||.+.|.++-+...=.++- ||+||.+
T Consensus 188 ~~~~~~~~~~~ai~r~Y~lGI~PDWWKLep~s~~~W~~i~~~I~~~Dp~crGvVvLGLd 246 (311)
T PF09863_consen 188 DMPVDDDTYARAIERFYNLGIKPDWWKLEPLSAAAWQAIEALIEERDPYCRGVVVLGLD 246 (311)
T ss_pred CCCCChHHHHHHHHHHHHcCCCCCeeccCCCCHHHHHHHHHHHHHhCCCceeEEEecCC
Confidence 345568899999999999874 467999999999888777774 7899874
No 139
>PF14206 Cys_rich_CPCC: Cysteine-rich CPCC
Probab=28.54 E-value=25 Score=28.63 Aligned_cols=12 Identities=50% Similarity=1.107 Sum_probs=9.7
Q ss_pred ecCCCCCccccc
Q 019981 272 KGPCPNCGTENV 283 (333)
Q Consensus 272 KG~CPNCGeEv~ 283 (333)
|-+||+||..++
T Consensus 1 K~~CPCCg~~Tl 12 (78)
T PF14206_consen 1 KYPCPCCGYYTL 12 (78)
T ss_pred CccCCCCCcEEe
Confidence 568999998765
No 140
>PF02829 3H: 3H domain; InterPro: IPR004173 The 3H domain is named after its three highly conserved histidine residues. The 3H domain appears to be a small molecule-binding domain, based on its occurrence with other domains []. Several proteins carrying this domain are transcriptional regulators from the biotin repressor family. The transcription regulator TM1602 from Thermotoga maritima is a DNA-binding protein thought to belong to a family of de novo NAD synthesis pathway regulators. TM1602 has an N-terminal DNA-binding domain and a C-terminal 3H regulatory domain. The N-terminal domain appears to bind to the NAD promoter region and repress the de novo NAD biosynthesis operon, while the C-terminal 3H domain may bind to nicotinamide, nicotinic acid, or other substrate/products []. The 3H domain has a 2-layer alpha/beta sandwich fold.; GO: 0005488 binding; PDB: 1J5Y_A.
Probab=28.23 E-value=94 Score=26.08 Aligned_cols=32 Identities=34% Similarity=0.618 Sum_probs=24.7
Q ss_pred HHHHHHHHhhhcCCCcc--------------ChHHHHHHHHHHhhcC
Q 019981 144 QKFLEASMAYVAGKPIM--------------SDEEYDKLKQKLKMEG 176 (333)
Q Consensus 144 q~fLEA~~AY~sGkPim--------------sDeeFD~LK~kLk~~G 176 (333)
++|++.+..+ .++|+. +.+.+|+++++|+.+|
T Consensus 50 ~~Fi~~l~~~-~~~~Ls~LT~GvH~HtI~a~~~e~l~~I~~~L~~~G 95 (98)
T PF02829_consen 50 DKFIEKLEKS-KAKPLSSLTGGVHYHTIEAPDEEDLDKIEEALKKKG 95 (98)
T ss_dssp HHHHHHHHH---S--STTGGGGEEEEEEEESSHHHHHHHHHHHHHTT
T ss_pred HHHHHHHhcc-CCcchHHhcCCEeeEEEEECCHHHHHHHHHHHHHCC
Confidence 8999999888 788875 4689999999999988
No 141
>PRK14714 DNA polymerase II large subunit; Provisional
Probab=28.15 E-value=33 Score=40.19 Aligned_cols=12 Identities=42% Similarity=1.207 Sum_probs=7.7
Q ss_pred eecCCCCCcccc
Q 019981 271 LKGPCPNCGTEN 282 (333)
Q Consensus 271 LKG~CPNCGeEv 282 (333)
-.+-||+||+..
T Consensus 678 ~~~fCP~CGs~t 689 (1337)
T PRK14714 678 YENRCPDCGTHT 689 (1337)
T ss_pred ccccCcccCCcC
Confidence 345777777764
No 142
>PF12767 SAGA-Tad1: Transcriptional regulator of RNA polII, SAGA, subunit; InterPro: IPR024738 The yeast Spt-Ada-Gcn5-Acetyl (SAGA) transferase complex is a multifunctional coactivator involved in multiple cellular processes [], including regulation of transcription by RNA polymerase II [, ]. It is formed of five major modular subunits and shows a high degree of structural conservation to human TFTC and STAGA []. This entry represents Ada1 (known as Tada1 in higher eukaryotes), one of the subunits that constitute the SAGA core. It also functions as a component of the SALSA and SLIK complexes. ; GO: 0070461 SAGA-type complex
Probab=28.09 E-value=59 Score=30.48 Aligned_cols=36 Identities=31% Similarity=0.544 Sum_probs=30.8
Q ss_pred ccchh-HHHHHHHHHHHHhhhcCccccChHHHhhhHhhhh
Q 019981 92 KKSLG-ELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELM 130 (333)
Q Consensus 92 k~slg-e~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~ 130 (333)
.+.|| |....|.+.|..|+-. -+|.+|||.+=..+.
T Consensus 19 ~~~LG~~~~~~Y~~~l~~fl~~---klsk~Efd~~~~~~L 55 (252)
T PF12767_consen 19 QKRLGPDRWKKYFQSLKRFLSG---KLSKEEFDKECRRIL 55 (252)
T ss_pred HHHHChHHHHHHHHHHHHHHHh---ccCHHHHHHHHHHHh
Confidence 45789 9999999999999985 489999999877754
No 143
>COG1675 TFA1 Transcription initiation factor IIE, alpha subunit [Transcription]
Probab=27.97 E-value=15 Score=33.85 Aligned_cols=39 Identities=26% Similarity=0.502 Sum_probs=26.0
Q ss_pred cCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccceeee
Q 019981 273 GPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRLIT 326 (333)
Q Consensus 273 G~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~it 326 (333)
=-||||.... +|=-. -.+...||. ||..|++....+.|+
T Consensus 114 y~C~~~~~r~-sfdeA--------~~~~F~Cp~------Cg~~L~~~d~s~~i~ 152 (176)
T COG1675 114 YVCPNCHVKY-SFDEA--------MELGFTCPK------CGEDLEEYDSSEEIE 152 (176)
T ss_pred eeCCCCCCcc-cHHHH--------HHhCCCCCC------CCchhhhccchHHHH
Confidence 3599998764 33211 123468888 999999998877664
No 144
>PRK08176 pdxK pyridoxal-pyridoxamine kinase/hydroxymethylpyrimidine kinase; Reviewed
Probab=27.84 E-value=1.5e+02 Score=27.86 Aligned_cols=53 Identities=15% Similarity=0.142 Sum_probs=39.6
Q ss_pred hhhHhhhhhcCCeeEEeChhhHHHHHHHHhhhcCCCccChHHHHHHHHHHhhcCC-ceeeec
Q 019981 123 DNLKEELMWEGSSVVMLSSAEQKFLEASMAYVAGKPIMSDEEYDKLKQKLKMEGS-EIVVEG 183 (333)
Q Consensus 123 d~LkEeL~weGSsvv~L~~~Eq~fLEA~~AY~sGkPimsDeeFD~LK~kLk~~GS-~VVvk~ 183 (333)
..+|++| .....+++.+..|-++| .|.++.++++..+.-.+|...|. .|++++
T Consensus 143 ~~~~~~L-l~~advitPN~~Ea~~L-------~g~~~~~~~~~~~~~~~l~~~g~~~VvIT~ 196 (281)
T PRK08176 143 EAYRQHL-LPLAQGLTPNIFELEIL-------TGKPCRTLDSAIAAAKSLLSDTLKWVVITS 196 (281)
T ss_pred HHHHHHh-HhhcCEeCCCHHHHHHH-------hCCCCCCHHHHHHHHHHHHhcCCCEEEEee
Confidence 4566655 57788999999998887 47787788777777777877785 466664
No 145
>PF15616 TerY-C: TerY-C metal binding domain
Probab=27.73 E-value=51 Score=29.22 Aligned_cols=45 Identities=24% Similarity=0.658 Sum_probs=29.4
Q ss_pred eeeeecCCCCCccc-ccee--ccccccccCCCCccceeCCCCccccccCceeEEecc
Q 019981 268 SLILKGPCPNCGTE-NVSF--FGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSN 321 (333)
Q Consensus 268 ~~iLKG~CPNCGeE-v~aF--fgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk 321 (333)
-+|=..-||.||++ .|+- =|-+.=|. ....+.|+. ||....|...
T Consensus 73 eL~g~PgCP~CGn~~~fa~C~CGkl~Ci~---g~~~~~CPw------Cg~~g~~~~~ 120 (131)
T PF15616_consen 73 ELIGAPGCPHCGNQYAFAVCGCGKLFCID---GEGEVTCPW------CGNEGSFGAG 120 (131)
T ss_pred HhcCCCCCCCCcChhcEEEecCCCEEEeC---CCCCEECCC------CCCeeeeccc
Confidence 34445889999998 3322 13333332 244788888 9999999865
No 146
>PF03965 Penicillinase_R: Penicillinase repressor; InterPro: IPR005650 Proteins in this entry are transcriptional regulators found in a variety of bacteria and a small number of archaea. Many are BlaI/MecI proteins which regulate resistance to penicillins (beta-lactams), though at least one protein (Q47839 from SWISSPROT) appears to be involved in the regulation of copper homeostasis []. BlaI regulators repress the expression of penicillin-degrading enzymes (penicillinases) until the cell encounters the antiobiotic, at which point repression ceases and penicillinase expression occurs, allowing cell growth []. MecI regulators repress the expression of MecA, a cell-wall biosynthetic enzyme not inhibited by penicillins at clinically achievable concentrations, until the presence of the antibiotic is detected []. At this point repression ends and MecA expression occurs which, together with the switching off of the penicillin-sensitive enzymes, allows the cell to grow despite the presence of antibiotic.; GO: 0003677 DNA binding, 0045892 negative regulation of transcription, DNA-dependent; PDB: 2G9W_A 2K4B_A 1XSD_A 1SD4_A 1SD7_A 1SD6_A 2P7C_B 1P6R_A 1OKR_B 2D45_B ....
Probab=27.03 E-value=1.9e+02 Score=23.62 Aligned_cols=33 Identities=33% Similarity=0.515 Sum_probs=24.3
Q ss_pred hhHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhhh
Q 019981 95 LGELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELMW 131 (333)
Q Consensus 95 lge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~w 131 (333)
|++.|...++.+|.. |. .-..|=.+.|+++..|
T Consensus 1 Ls~~E~~IM~~lW~~---~~-~t~~eI~~~l~~~~~~ 33 (115)
T PF03965_consen 1 LSDLELEIMEILWES---GE-ATVREIHEALPEERSW 33 (115)
T ss_dssp --HHHHHHHHHHHHH---SS-EEHHHHHHHHCTTSS-
T ss_pred CCHHHHHHHHHHHhC---CC-CCHHHHHHHHHhcccc
Confidence 688999999999973 33 5558888899998778
No 147
>PRK04011 peptide chain release factor 1; Provisional
Probab=26.74 E-value=28 Score=35.37 Aligned_cols=36 Identities=25% Similarity=0.534 Sum_probs=25.5
Q ss_pred ecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEe
Q 019981 272 KGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 272 KG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~ 319 (333)
.--||+||.+..-++.. ......-.|++ ||..++..
T Consensus 328 ~~~c~~c~~~~~~~~~~------~~~~~~~~c~~------~~~~~~~~ 363 (411)
T PRK04011 328 TYKCPNCGYEEEKTVKR------REELPEKTCPK------CGSELEIV 363 (411)
T ss_pred EEEcCCCCcceeeeccc------ccccccccCcc------cCcccccc
Confidence 45699999988777743 33455567776 99887764
No 148
>PF03317 ELF: ELF protein; InterPro: IPR004990 This is a family of hypothetical proteins from cereal crops.
Probab=26.45 E-value=1.5e+02 Score=28.80 Aligned_cols=106 Identities=20% Similarity=0.245 Sum_probs=67.9
Q ss_pred ccccccccccCCcccccccCcccCccccccccccccccccccccchhHHHHHHHHHHHHhhhcCccccC----hHHHhhh
Q 019981 50 FTVRRRSFVLPSKATTDQQGQVEGDEVVDSKILQYCSIDKKEKKSLGELEQEFLQALQAFYYEGKAVMS----NEEFDNL 125 (333)
Q Consensus 50 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~d~~i~~yCsiD~~~k~slge~E~~fl~Al~~fY~~gk~~~s----deefd~L 125 (333)
++-|-|+|++..+-+.-.+.-++|=-...+++++|=-|..- .-.+|+.+|-..=.++.+ ..-|..|
T Consensus 144 IPGRLrLFLMEE~~S~~R~DliQefvalY~r~g~~LPiEPY----------lleealrSYlD~i~atD~fsiLqAaYQdL 213 (284)
T PF03317_consen 144 IPGRLRLFLMEEKLSSMRQDLIQEFVALYQRSGPVLPIEPY----------LLEEALRSYLDHIHATDSFSILQAAYQDL 213 (284)
T ss_pred CcchhhhhhhHhHHHHHHHHHHHHHHHHHHccCCcccccHH----------HHHHHHHHHHHhhcccccHHHHHHHHHHH
Confidence 35566778877766655455555533456666666554321 223466666555333333 7889999
Q ss_pred HhhhhhcCCeeEEeCh--hhHHHHHHHHhhhcCCCccChHHHHHHHHHHhhcCC
Q 019981 126 KEELMWEGSSVVMLSS--AEQKFLEASMAYVAGKPIMSDEEYDKLKQKLKMEGS 177 (333)
Q Consensus 126 kEeL~weGSsvv~L~~--~Eq~fLEA~~AY~sGkPimsDeeFD~LK~kLk~~GS 177 (333)
+|. +|-|+.+..- -.|+||||--+- .-+-+++++.+|+|-
T Consensus 214 ren---e~GS~FF~~~VSHNrD~LeA~ss~---------Rr~~Eveqrirw~~I 255 (284)
T PF03317_consen 214 REN---EEGSVFFRDVVSHNRDFLEAESSA---------RRCLEVEQRIRWEEI 255 (284)
T ss_pred Hhc---CCCcEEeHhhhhccHhHHHHHhhh---------hHHHHHHHHhhhhhh
Confidence 998 7777766554 459999996543 346788999999873
No 149
>smart00400 ZnF_CHCC zinc finger.
Probab=26.25 E-value=39 Score=24.59 Aligned_cols=27 Identities=30% Similarity=0.578 Sum_probs=19.8
Q ss_pred ecCCCCCccccceeccccccccCCCCccceeCCC
Q 019981 272 KGPCPNCGTENVSFFGTILSISSGGTTNTINCSN 305 (333)
Q Consensus 272 KG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~ 305 (333)
+|.||-+++..-+|- | +...|...|..
T Consensus 2 ~~~cPfh~d~~pSf~-----v--~~~kn~~~Cf~ 28 (55)
T smart00400 2 KGLCPFHGEKTPSFS-----V--SPDKQFFHCFG 28 (55)
T ss_pred cccCcCCCCCCCCEE-----E--ECCCCEEEEeC
Confidence 578999999999884 2 33456677765
No 150
>cd07114 ALDH_DhaS Uncharacterized Candidatus pelagibacter aldehyde dehydrogenase, DhaS-like. Uncharacterized aldehyde dehydrogenase from Candidatus pelagibacter (DhaS) and other related sequences are present in this CD.
Probab=26.06 E-value=1.8e+02 Score=29.08 Aligned_cols=69 Identities=16% Similarity=0.421 Sum_probs=47.5
Q ss_pred ccChHHHhhhHhhhhh-----cCCee-----EEeCh-hhHHHHHHHHhhh----cC---------CCccChHHHHHHHHH
Q 019981 116 VMSNEEFDNLKEELMW-----EGSSV-----VMLSS-AEQKFLEASMAYV----AG---------KPIMSDEEYDKLKQK 171 (333)
Q Consensus 116 ~~sdeefd~LkEeL~w-----eGSsv-----v~L~~-~Eq~fLEA~~AY~----sG---------kPimsDeeFD~LK~k 171 (333)
++.|.+.|..=+.+.| .|..| |.+-+ --.+|++++.... -| -|+++.+.+|+++..
T Consensus 237 V~~dAdl~~aa~~i~~~~~~~~GQ~C~a~~~v~V~~~v~~~f~~~l~~~~~~~~~g~p~~~~~~~gpli~~~~~~~~~~~ 316 (457)
T cd07114 237 VFDDADLDAAVNGVVAGIFAAAGQTCVAGSRLLVQRSIYDEFVERLVARARAIRVGDPLDPETQMGPLATERQLEKVERY 316 (457)
T ss_pred ECCCCCHHHHHHHHHHHHHhccCCCCCCCceEEEcHHHHHHHHHHHHHHHHhCCCCCCCCCCCCCCCCcCHHHHHHHHHH
Confidence 4568888888888777 55544 34433 3367888876533 33 378899999999998
Q ss_pred Hhhc---CCceeeecC
Q 019981 172 LKME---GSEIVVEGP 184 (333)
Q Consensus 172 Lk~~---GS~VVvk~P 184 (333)
+... |.+++.-|.
T Consensus 317 i~~a~~~ga~~l~gg~ 332 (457)
T cd07114 317 VARAREEGARVLTGGE 332 (457)
T ss_pred HHHHHHCCCEEEeCCC
Confidence 8754 887766543
No 151
>PRK09401 reverse gyrase; Reviewed
Probab=25.67 E-value=28 Score=40.04 Aligned_cols=16 Identities=44% Similarity=0.960 Sum_probs=12.6
Q ss_pred eeeecCCCCCccccce
Q 019981 269 LILKGPCPNCGTENVS 284 (333)
Q Consensus 269 ~iLKG~CPNCGeEv~a 284 (333)
.+-++.|||||-++-+
T Consensus 4 ~~y~~~cpnc~g~i~~ 19 (1176)
T PRK09401 4 AIYKNSCPNCGGDISD 19 (1176)
T ss_pred hhhcccCCCCCCcCcH
Confidence 3568899999988764
No 152
>COG5257 GCD11 Translation initiation factor 2, gamma subunit (eIF-2gamma; GTPase) [Translation, ribosomal structure and biogenesis]
Probab=25.66 E-value=54 Score=33.91 Aligned_cols=44 Identities=23% Similarity=0.562 Sum_probs=30.9
Q ss_pred hcceeeeecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccceeeecCC
Q 019981 265 VRESLILKGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTRLITLPE 329 (333)
Q Consensus 265 ~~D~~iLKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R~itl~~ 329 (333)
.-|..|- -||+|+.. -.| +-+-+|++ ||+..++-+..-.+..|+
T Consensus 52 YAd~~i~--kC~~c~~~-~~y------------~~~~~C~~------cg~~~~l~R~VSfVDaPG 95 (415)
T COG5257 52 YADAKIY--KCPECYRP-ECY------------TTEPKCPN------CGAETELVRRVSFVDAPG 95 (415)
T ss_pred cccCceE--eCCCCCCC-ccc------------ccCCCCCC------CCCCccEEEEEEEeeCCc
Confidence 3344443 59999876 222 22458998 999999998888877775
No 153
>COG1326 Uncharacterized archaeal Zn-finger protein [General function prediction only]
Probab=25.16 E-value=30 Score=32.82 Aligned_cols=41 Identities=22% Similarity=0.465 Sum_probs=21.2
Q ss_pred eecCCCCCc-cccceeccccccccCCCCccceeCCCCccc-cccCceeEEe
Q 019981 271 LKGPCPNCG-TENVSFFGTILSISSGGTTNTINCSNLTFC-FSCGTTMVYD 319 (333)
Q Consensus 271 LKG~CPNCG-eEv~aFfgtilsv~s~~~~n~vkC~~~aeC-HVC~t~L~f~ 319 (333)
..-.||+|| +|+..=+ |...+..-.++|.+ | ||=-..+.+.
T Consensus 5 iy~~Cp~Cg~eev~hEV-----ik~~g~~~lvrC~e---CG~V~~~~i~~~ 47 (201)
T COG1326 5 IYIECPSCGSEEVSHEV-----IKERGREPLVRCEE---CGTVHPAIIKTP 47 (201)
T ss_pred EEEECCCCCcchhhHHH-----HHhcCCceEEEccC---CCcEeeceeecc
Confidence 345799999 4442211 11222336789964 6 3443344444
No 154
>TIGR00398 metG methionyl-tRNA synthetase. The methionyl-tRNA synthetase (metG) is a class I amino acyl-tRNA ligase. This model appears to recognize the methionyl-tRNA synthetase of every species, including eukaryotic cytosolic and mitochondrial forms. The UPGMA difference tree calculated after search and alignment according to this model shows an unusual deep split between two families of MetG. One family contains forms from the Archaea, yeast cytosol, spirochetes, and E. coli, among others. The other family includes forms from yeast mitochondrion, Synechocystis sp., Bacillus subtilis, the Mycoplasmas, Aquifex aeolicus, and Helicobacter pylori. The E. coli enzyme is homodimeric, although monomeric forms can be prepared that are fully active. Activity of this enzyme in bacteria includes aminoacylation of fMet-tRNA with Met; subsequent formylation of the Met to fMet is catalyzed by a separate enzyme. Note that the protein from Aquifex aeolicus is split into an alpha (large) and beta (sma
Probab=25.03 E-value=57 Score=33.35 Aligned_cols=48 Identities=23% Similarity=0.398 Sum_probs=22.9
Q ss_pred ecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEEeccce
Q 019981 272 KGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNTR 323 (333)
Q Consensus 272 KG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~R 323 (333)
+|-||.||.+- -+|++- +..+...+..=.....|-.||++++++....
T Consensus 136 ~g~cp~c~~~~--~~g~~c--e~cg~~~~~~~l~~p~~~~~~~~~e~~~~~~ 183 (530)
T TIGR00398 136 EGTCPKCGSED--ARGDHC--EVCGRHLEPTELINPRCKICGAKPELRDSEH 183 (530)
T ss_pred cCCCCCCCCcc--cccchh--hhccccCCHHHhcCCccccCCCcceEEecce
Confidence 58899999862 223321 0011111000001123555899998885543
No 155
>cd07078 ALDH NAD(P)+ dependent aldehyde dehydrogenase family. The aldehyde dehydrogenase family (ALDH) of NAD(P)+ dependent enzymes, in general, oxidize a wide range of endogenous and exogenous aliphatic and aromatic aldehydes to their corresponding carboxylic acids and play an important role in detoxification. Besides aldehyde detoxification, many ALDH isozymes possess multiple additional catalytic and non-catalytic functions such as participating in metabolic pathways, or as binding proteins, or as osmoregulants, to mention a few. The enzyme has three domains, a NAD(P)+ cofactor-binding domain, a catalytic domain, and a bridging domain; and the active enzyme is generally either homodimeric or homotetrameric. The catalytic mechanism is proposed to involve cofactor binding, resulting in a conformational change and activation of an invariant catalytic cysteine nucleophile. The cysteine and aldehyde substrate form an oxyanion thiohemiacetal intermediate resulting in hydride transfer
Probab=24.95 E-value=2.2e+02 Score=27.86 Aligned_cols=68 Identities=19% Similarity=0.447 Sum_probs=44.2
Q ss_pred cChHHHhhhHhhhhh-----cCC-----eeEEe-ChhhHHHHHHHHh----hhcCC---------CccChHHHHHHHHHH
Q 019981 117 MSNEEFDNLKEELMW-----EGS-----SVVML-SSAEQKFLEASMA----YVAGK---------PIMSDEEYDKLKQKL 172 (333)
Q Consensus 117 ~sdeefd~LkEeL~w-----eGS-----svv~L-~~~Eq~fLEA~~A----Y~sGk---------PimsDeeFD~LK~kL 172 (333)
+.+.+++..-+.+.| .|- ..+.+ +....+|++++.+ +.-|. |+++.+.+++++..+
T Consensus 215 ~~~ad~~~aa~~i~~~~~~~~Gq~C~a~~~i~v~~~~~~~~~~~L~~~l~~~~~g~p~~~~~~~~~~~~~~~~~~~~~~i 294 (432)
T cd07078 215 FDDADLDAAVKGAVFGAFGNAGQVCTAASRLLVHESIYDEFVERLVERVKALKVGNPLDPDTDMGPLISAAQLDRVLAYI 294 (432)
T ss_pred CCCCCHHHHHHHHHHHHHhccCCCccCCceEEEcHHHHHHHHHHHHHHHHccCcCCCCCCCCCCCCCCCHHHHHHHHHHH
Confidence 556677776666554 453 23333 3344678887643 55454 488999999999888
Q ss_pred hh---cCCceeeecC
Q 019981 173 KM---EGSEIVVEGP 184 (333)
Q Consensus 173 k~---~GS~VVvk~P 184 (333)
.. .|.+++.-++
T Consensus 295 ~~~~~~g~~~~~gg~ 309 (432)
T cd07078 295 EDAKAEGAKLLCGGK 309 (432)
T ss_pred HHHHhCCCEEEeCCc
Confidence 76 5777776443
No 156
>PRK03564 formate dehydrogenase accessory protein FdhE; Provisional
Probab=24.89 E-value=53 Score=32.74 Aligned_cols=24 Identities=33% Similarity=0.354 Sum_probs=12.6
Q ss_pred HHHHHHHHhhhcCCCccChHHHHHHH
Q 019981 144 QKFLEASMAYVAGKPIMSDEEYDKLK 169 (333)
Q Consensus 144 q~fLEA~~AY~sGkPimsDeeFD~LK 169 (333)
++.|.++.+=. +|.+++.....|+
T Consensus 104 ~~~L~~Ll~~l--~~~~~~~~~~~l~ 127 (309)
T PRK03564 104 QKLLMALIAEL--KPEASGPALAVIE 127 (309)
T ss_pred HHHHHHHHHHh--cccCCHHHHHHHH
Confidence 45566655522 4456666654443
No 157
>PF05193 Peptidase_M16_C: Peptidase M16 inactive domain; InterPro: IPR007863 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Metalloproteases are the most diverse of the four main types of protease, with more than 50 families identified to date. In these enzymes, a divalent cation, usually zinc, activates the water molecule. The metal ion is held in place by amino acid ligands, usually three in number. The known metal ligands are His, Glu, Asp or Lys and at least one other residue is required for catalysis, which may play an electrophillic role. Of the known metalloproteases, around half contain an HEXXH motif, which has been shown in crystallographic studies to form part of the metal-binding site []. The HEXXH motif is relatively common, but can be more stringently defined for metalloproteases as 'abXHEbbHbc', where 'a' is most often valine or threonine and forms part of the S1' subsite in thermolysin and neprilysin, 'b' is an uncharged residue, and 'c' a hydrophobic residue. Proline is never found in this site, possibly because it would break the helical structure adopted by this motif in metalloproteases []. These metallopeptidases belong to MEROPS peptidase family M16 (clan ME). They include proteins, which are classified as non-peptidase homologues either have been found experimentally to be without peptidase activity, or lack amino acid residues that are believed to be essential for the catalytic activity. The peptidases in this group of sequences include: Insulinase, insulin-degrading enzyme (3.4.24.56 from EC) Mitochondrial processing peptidase alpha subunit, (Alpha-MPP, 3.4.24.64 from EC) Pitrlysin, Protease III precursor (3.4.24.55 from EC) Nardilysin, (3.4.24.61 from EC) Ubiquinol-cytochrome C reductase complex core protein I,mitochondrial precursor (1.10.2.2 from EC) Coenzyme PQQ synthesis protein F (3.4.99 from EC) These proteins do not share many regions of sequence similarity; the most noticeable is in the N-terminal section. This region includes a conserved histidine followed, two residues later by a glutamate and another histidine. In pitrilysin, it has been shown [] that this H-x-x-E-H motif is involved in enzymatic activity; the two histidines bind zinc and the glutamate is necessary for catalytic activity. The mitochondrial processing peptidase consists of two structurally related domains. One is the active peptidase whereas the other, the C-terminal region, is inactive. The two domains hold the substrate like a clamp [].; GO: 0004222 metalloendopeptidase activity, 0008270 zinc ion binding, 0006508 proteolysis; PDB: 1BE3_B 1PP9_B 2A06_B 1SQB_B 1SQP_B 1L0N_B 1SQX_B 1NU1_B 1L0L_B 2FYU_B ....
Probab=24.87 E-value=1.2e+02 Score=24.36 Aligned_cols=33 Identities=33% Similarity=0.493 Sum_probs=26.1
Q ss_pred chhHHHHHHHHHHHHhhhcCccccChHHHhhhHhhh
Q 019981 94 SLGELEQEFLQALQAFYYEGKAVMSNEEFDNLKEEL 129 (333)
Q Consensus 94 slge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL 129 (333)
.+.+.+..+++-+...=.+| ++++||++.|+.|
T Consensus 152 ~~~~~~~~~~~~l~~l~~~~---~s~~el~~~k~~L 184 (184)
T PF05193_consen 152 NLDEAIEAILQELKRLREGG---ISEEELERAKNQL 184 (184)
T ss_dssp GHHHHHHHHHHHHHHHHHHC---S-HHHHHHHHHHH
T ss_pred cHHHHHHHHHHHHHHHHHcC---CCHHHHHHHHhcC
Confidence 56777778888887777765 9999999999876
No 158
>PF03884 DUF329: Domain of unknown function (DUF329); InterPro: IPR005584 The biological function of these short proteins is unknown, but they contain four conserved cysteines, suggesting that they all bind zinc. YacG (Q5X8H6 from SWISSPROT) from Escherichia coli has been shown to bind zinc and contains the structural motifs typical of zinc-binding proteins []. The conserved four cysteine motif in these proteins (-C-X(2)-C-X(15)-C-X(3)-C-) is not found in other zinc-binding proteins with known structures.; GO: 0008270 zinc ion binding; PDB: 1LV3_A.
Probab=24.77 E-value=26 Score=27.00 Aligned_cols=13 Identities=31% Similarity=0.585 Sum_probs=5.9
Q ss_pred ecCCCCCccccce
Q 019981 272 KGPCPNCGTENVS 284 (333)
Q Consensus 272 KG~CPNCGeEv~a 284 (333)
+-.||.||.++..
T Consensus 2 ~v~CP~C~k~~~~ 14 (57)
T PF03884_consen 2 TVKCPICGKPVEW 14 (57)
T ss_dssp EEE-TTT--EEE-
T ss_pred cccCCCCCCeecc
Confidence 4567888776643
No 159
>COG1779 C4-type Zn-finger protein [General function prediction only]
Probab=24.74 E-value=43 Score=31.82 Aligned_cols=30 Identities=30% Similarity=0.661 Sum_probs=18.6
Q ss_pred eeeecCCCCCccccce--------eccccccccCCCCccceeCCC
Q 019981 269 LILKGPCPNCGTENVS--------FFGTILSISSGGTTNTINCSN 305 (333)
Q Consensus 269 ~iLKG~CPNCGeEv~a--------Ffgtilsv~s~~~~n~vkC~~ 305 (333)
..-...||.||....+ |||.++ -.+.-|.+
T Consensus 11 ~~~~~~CPvCg~~l~~~~~~~~IPyFG~V~-------i~t~~C~~ 48 (201)
T COG1779 11 FETRIDCPVCGGTLKAHMYLYDIPYFGEVL-------ISTGVCER 48 (201)
T ss_pred eeeeecCCcccceeeEEEeeecCCccceEE-------EEEEEccc
Confidence 3456789999985443 666654 12456766
No 160
>PTZ00381 aldehyde dehydrogenase family protein; Provisional
Probab=24.71 E-value=1.7e+02 Score=30.24 Aligned_cols=67 Identities=24% Similarity=0.471 Sum_probs=47.6
Q ss_pred ccChHHHhhhHhhhhhc-----CCeeE-----Ee-ChhhHHHHHHHH----hhhcCC---------CccChHHHHHHHHH
Q 019981 116 VMSNEEFDNLKEELMWE-----GSSVV-----ML-SSAEQKFLEASM----AYVAGK---------PIMSDEEYDKLKQK 171 (333)
Q Consensus 116 ~~sdeefd~LkEeL~we-----GSsvv-----~L-~~~Eq~fLEA~~----AY~sGk---------PimsDeeFD~LK~k 171 (333)
++.|.+.|.--+.+.|. |-.|+ .+ .....+|++++. .++ |. |+++++.|++++.-
T Consensus 224 V~~dAdl~~Aa~~i~~g~~~naGQ~C~A~~~vlV~~~i~d~f~~~l~~~~~~~~-g~~~~~~~~~gpli~~~~~~ri~~~ 302 (493)
T PTZ00381 224 VDKSCNLKVAARRIAWGKFLNAGQTCVAPDYVLVHRSIKDKFIEALKEAIKEFF-GEDPKKSEDYSRIVNEFHTKRLAEL 302 (493)
T ss_pred EcCCCCHHHHHHHHHHHHHhhcCCcCCCCCEEEEeHHHHHHHHHHHHHHHHHHh-CCCCccCCCcCCCCCHHHHHHHHHH
Confidence 55688888888888883 54433 33 334567888764 344 43 67999999999999
Q ss_pred HhhcCCceeeec
Q 019981 172 LKMEGSEIVVEG 183 (333)
Q Consensus 172 Lk~~GS~VVvk~ 183 (333)
++.+|.+++.-|
T Consensus 303 i~~~ga~~~~gG 314 (493)
T PTZ00381 303 IKDHGGKVVYGG 314 (493)
T ss_pred HHhCCCcEEECC
Confidence 988898887643
No 161
>cd02661 Peptidase_C19E A subfamily of Peptidase C19. Peptidase C19 contains ubiquitinyl hydrolases. They are intracellular peptidases that remove ubiquitin molecules from polyubiquinated peptides by cleavage of isopeptide bonds. They hydrolyze bonds involving the carboxyl group of the C-terminal Gly residue of ubiquitin. The purpose of the de-ubiquitination is thought to be editing of the ubiquitin conjugates, which could rescue them from degradation, as well as recycling of the ubiquitin. The ubiquitin/proteasome system is responsible for most protein turnover in the mammalian cell, and with over 50 members, family C19 is one of the largest families of peptidases in the human genome.
Probab=24.65 E-value=79 Score=28.62 Aligned_cols=24 Identities=13% Similarity=0.390 Sum_probs=14.4
Q ss_pred ceeCCCCccccccCceeEEeccceeeecCC
Q 019981 300 TINCSNLTFCFSCGTTMVYDSNTRLITLPE 329 (333)
Q Consensus 300 ~vkC~~~aeCHVC~t~L~f~tk~R~itl~~ 329 (333)
..+|++ |+..-......++.++|+
T Consensus 182 ~~~C~~------C~~~~~~~~~~~i~~~P~ 205 (304)
T cd02661 182 KYKCER------CKKKVKASKQLTIHRAPN 205 (304)
T ss_pred CeeCCC------CCCccceEEEEEEecCCc
Confidence 346666 887665555555556664
No 162
>PF06750 DiS_P_DiS: Bacterial Peptidase A24 N-terminal domain; InterPro: IPR010627 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Aspartic endopeptidases 3.4.23. from EC of vertebrate, fungal and retroviral origin have been characterised []. More recently, aspartic endopeptidases associated with the processing of bacterial type 4 prepilin [] and archaean preflagellin have been described [, ]. Structurally, aspartic endopeptidases are bilobal enzymes, each lobe contributing a catalytic Asp residue, with an extended active site cleft localised between the two lobes of the molecule. One lobe has probably evolved from the other through a gene duplication event in the distant past. In modern-day enzymes, although the three-dimensional structures are very similar, the amino acid sequences are more divergent, except for the catalytic site motif, which is very conserved. The presence and position of disulphide bridges are other conserved features of aspartic peptidases. All or most aspartate peptidases are endopeptidases. These enzymes have been assigned into clans (proteins which are evolutionary related), and further sub-divided into families, largely on the basis of their tertiary structure. This domain is found at the N terminus of bacterial aspartic peptidases belonging to MEROPS peptidase family A24 (clan AD), subfamily A24A (type IV prepilin peptidase, IPR000045 from INTERPRO). It's function has not been specifically determined; however some of the family have been characterised as bifunctional [], and this domain may contain the N-methylation activity. The domain consists of an intracellular region between a pair of transmembrane domains. This intracellular region contains an invariant proline and four conserved cysteines. These Cys residues are arranged in a two-pair motif, with the Cys residues of a pair separated (usually) by 2 aa and with each pair separated by 21 largely hydrophilic residues (C-X-X-C...X21...C-X-X-C); they have been shown to be essential to the overall function of the enzyme [, ]. The bifunctional enzyme prepilin peptidase (PilD) from Pseudomonas aeruginosa is a key determinant in both type-IV pilus biogenesis and extracellular protein secretion, in its roles as a leader peptidase and methyl transferase (MTase). It is responsible for endopeptidic cleavage of the unique leader peptides that characterise type-IV pilin precursors, as well as proteins with homologous leader sequences that are essential components of the general secretion pathway found in a variety of Gram-negative pathogens. Following removal of the leader peptides, the same enzyme is responsible for the second posttranslational modification that characterises the type-IV pilins and their homologues, namely N-methylation of the newly exposed N-terminal amino acid residue [].
Probab=24.54 E-value=40 Score=27.63 Aligned_cols=26 Identities=35% Similarity=0.999 Sum_probs=18.1
Q ss_pred HHHhhhhHHHHHHHHHHhhhhcceeeeecCCCCCccccc
Q 019981 245 TWFAAVPLIVYLSQSLTKLIVRESLILKGPCPNCGTENV 283 (333)
Q Consensus 245 t~~~~~P~i~~~a~~Lt~l~~~D~~iLKG~CPNCGeEv~ 283 (333)
.|.-..|+++++. +||-|.+|++.+-
T Consensus 44 ~~~~lIPi~S~l~-------------lrGrCr~C~~~I~ 69 (92)
T PF06750_consen 44 SWWDLIPILSYLL-------------LRGRCRYCGAPIP 69 (92)
T ss_pred cccccchHHHHHH-------------hCCCCcccCCCCC
Confidence 3556677777654 6788888887764
No 163
>TIGR00100 hypA hydrogenase nickel insertion protein HypA. In Hpylori, hypA mutant abolished hydrogenase activity and decrease in urease activity. Nickel supplementation in media restored urease activity and partial hydrogenase activity. HypA probably involved in inserting Ni in enzymes.
Probab=24.08 E-value=44 Score=28.23 Aligned_cols=14 Identities=21% Similarity=0.577 Sum_probs=10.6
Q ss_pred ceeeeecCCCCCcc
Q 019981 267 ESLILKGPCPNCGT 280 (333)
Q Consensus 267 D~~iLKG~CPNCGe 280 (333)
+.+-+.+-|++||.
T Consensus 65 ~~~p~~~~C~~Cg~ 78 (115)
T TIGR00100 65 EDEPVECECEDCSE 78 (115)
T ss_pred EeeCcEEEcccCCC
Confidence 34557789999993
No 164
>PF12207 DUF3600: Domain of unknown function (DUF3600); InterPro: IPR022019 This family of proteins is found in bacteria. Proteins in this family are approximately 230 amino acids in length. This domain is the C-terminal of the putative ecf-type sigma factor negative effector. ; PDB: 3FGG_A 3FH3_A.
Probab=23.98 E-value=49 Score=30.47 Aligned_cols=66 Identities=32% Similarity=0.427 Sum_probs=38.5
Q ss_pred cccchhHHHHHH--HHHHHHhhh------cCccccChHHHhhhHhhhhh-------cCCe-eEEeCh----hhHHHHHHH
Q 019981 91 EKKSLGELEQEF--LQALQAFYY------EGKAVMSNEEFDNLKEELMW-------EGSS-VVMLSS----AEQKFLEAS 150 (333)
Q Consensus 91 ~k~slge~E~~f--l~Al~~fY~------~gk~~~sdeefd~LkEeL~w-------eGSs-vv~L~~----~Eq~fLEA~ 150 (333)
++-|.+|.|+.= .--||-||. .-|.+|+++|||.-+|.||= .||+ -+++.. .-++|++|.
T Consensus 72 e~ls~~eqee~k~~~~eLqPYFdKLN~~~SsK~vlt~~E~d~y~eALm~~e~v~vk~~~~~~~~ve~vpe~~~e~f~~a~ 151 (162)
T PF12207_consen 72 EKLSKEEQEEYKKLTMELQPYFDKLNGHKSSKEVLTQEEYDQYIEALMTYETVRVKTKSSGGITVEEVPEAYKERFIKAE 151 (162)
T ss_dssp GGS-HHHHHHHHHHHHHHHHHHHHHTT---HHHHS-HHHHHHHHHHHHHHHHHHHHCT-SS---GGGS-HHHHHHHHHHH
T ss_pred HhCCHHHHHHHHHHHHhcchHHHHhcCCcchhhhcCHHHHHHHHHHHhhhheeeeeccCCCCCcHHhccHHHHHHHHHHH
Confidence 455666666432 223566664 34669999999999999986 3433 333332 347899885
Q ss_pred H--hhhcC
Q 019981 151 M--AYVAG 156 (333)
Q Consensus 151 ~--AY~sG 156 (333)
+ -|.++
T Consensus 152 ~~~~yv~~ 159 (162)
T PF12207_consen 152 QFMEYVNE 159 (162)
T ss_dssp HHHHHHHH
T ss_pred HHHHHHHh
Confidence 4 47654
No 165
>TIGR01391 dnaG DNA primase, catalytic core. This protein contains a CHC2 zinc finger (Pfam:PF01807) and a Toprim domain (Pfam:PF01751).
Probab=23.92 E-value=48 Score=33.46 Aligned_cols=28 Identities=21% Similarity=0.383 Sum_probs=19.8
Q ss_pred eecCCCCCccccceeccccccccCCCCccceeCCC
Q 019981 271 LKGPCPNCGTENVSFFGTILSISSGGTTNTINCSN 305 (333)
Q Consensus 271 LKG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~ 305 (333)
.+|.||-|++..-+|. | +...+..+|..
T Consensus 33 ~~~~CPfh~ek~pSf~-----v--~~~k~~~~Cf~ 60 (415)
T TIGR01391 33 YVGLCPFHHEKTPSFS-----V--SPEKQFYHCFG 60 (415)
T ss_pred eEeeCCCCCCCCCeEE-----E--EcCCCcEEECC
Confidence 4589999999998886 2 23455566654
No 166
>PRK00448 polC DNA polymerase III PolC; Validated
Probab=23.77 E-value=46 Score=39.36 Aligned_cols=36 Identities=39% Similarity=0.597 Sum_probs=26.8
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEE
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVY 318 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f 318 (333)
-||||. ++=|-++-++.||-+--.-+||+ ||+.|.=
T Consensus 910 ~C~~C~---~~ef~~~~~~~sG~Dlpdk~Cp~------Cg~~~~k 945 (1437)
T PRK00448 910 VCPNCK---YSEFFTDGSVGSGFDLPDKDCPK------CGTKLKK 945 (1437)
T ss_pred cCcccc---cccccccccccccccCccccCcc------ccccccc
Confidence 399995 44444556777888887888888 9998753
No 167
>PF14485 DUF4431: Domain of unknown function (DUF4431)
Probab=23.69 E-value=67 Score=23.75 Aligned_cols=22 Identities=41% Similarity=0.768 Sum_probs=18.3
Q ss_pred ccChHHHHHHHHHHhhcCCceeeec
Q 019981 159 IMSDEEYDKLKQKLKMEGSEIVVEG 183 (333)
Q Consensus 159 imsDeeFD~LK~kLk~~GS~VVvk~ 183 (333)
++++++|+.++. +.|..|.|.|
T Consensus 5 ~l~~~~~~~~~~---~~Gk~V~V~G 26 (48)
T PF14485_consen 5 ILSEEDYSYLKS---LLGKRVSVTG 26 (48)
T ss_pred EeChhhhHHHHH---hcCCeEEEEE
Confidence 458999999887 6899999875
No 168
>cd07120 ALDH_PsfA-ACA09737 Pseudomonas putida aldehyde dehydrogenase PsfA (ACA09737)-like. Included in this CD is the aldehyde dehydrogenase (PsfA, locus ACA09737) of Pseudomonas putida involved in furoic acid metabolism. Transcription of psfA was induced in response to 2-furoic acid, furfuryl alcohol, and furfural.
Probab=23.66 E-value=2.3e+02 Score=28.77 Aligned_cols=70 Identities=20% Similarity=0.379 Sum_probs=48.2
Q ss_pred ccChHHHhhhHhhhhh-----cCCe-----eEEeC-hhhHHHHHHHHh----hhcCC---------CccChHHHHHHHHH
Q 019981 116 VMSNEEFDNLKEELMW-----EGSS-----VVMLS-SAEQKFLEASMA----YVAGK---------PIMSDEEYDKLKQK 171 (333)
Q Consensus 116 ~~sdeefd~LkEeL~w-----eGSs-----vv~L~-~~Eq~fLEA~~A----Y~sGk---------PimsDeeFD~LK~k 171 (333)
++.|.+.|..-+.+.| .|-. .|.+- ..-.+|++++.+ ..-|. |+++.+.+++++.-
T Consensus 236 V~~daDl~~aa~~i~~~~~~~~GQ~C~a~~rv~V~~~i~~~f~~~l~~~~~~l~~G~p~~~~~~~gpli~~~~~~~~~~~ 315 (455)
T cd07120 236 VFDDADLDAALPKLERALTIFAGQFCMAGSRVLVQRSIADEVRDRLAARLAAVKVGPGLDPASDMGPLIDRANVDRVDRM 315 (455)
T ss_pred ECCCCCHHHHHHHHHHHHHHhCCCCCCCCeEEEEcHHHHHHHHHHHHHHHHhcCcCCCCCCCCCcCCccCHHHHHHHHHH
Confidence 5568889998888888 3532 33443 344678888754 33343 68999999999987
Q ss_pred Hhh---cCCceeeecCe
Q 019981 172 LKM---EGSEIVVEGPR 185 (333)
Q Consensus 172 Lk~---~GS~VVvk~Pr 185 (333)
+.. +|.+++..|.+
T Consensus 316 i~~a~~~ga~~~~~g~~ 332 (455)
T cd07120 316 VERAIAAGAEVVLRGGP 332 (455)
T ss_pred HHHHHHCCCEEEeCCcc
Confidence 765 68888876643
No 169
>PF13058 DUF3920: Protein of unknown function (DUF3920)
Probab=23.38 E-value=36 Score=30.09 Aligned_cols=47 Identities=32% Similarity=0.531 Sum_probs=35.6
Q ss_pred cccccccc--ccchhHHHHHHHHHHHHhhhcCccc---cChHHHhhhHhhhh
Q 019981 84 YCSIDKKE--KKSLGELEQEFLQALQAFYYEGKAV---MSNEEFDNLKEELM 130 (333)
Q Consensus 84 yCsiD~~~--k~slge~E~~fl~Al~~fY~~gk~~---~sdeefd~LkEeL~ 130 (333)
+|..=+.. -.+|||.|-+||.++..||-..|.+ .-=|||+.+=+.|.
T Consensus 30 FcdTc~an~vl~~LgeeeeefLf~~~g~y~kek~~iFv~~we~y~qvlktll 81 (126)
T PF13058_consen 30 FCDTCDANKVLLSLGEEEEEFLFPAGGFYHKEKQLIFVCMWEEYEQVLKTLL 81 (126)
T ss_pred EecccchhHHHHHhccchhhhccccchhhhccccEEEEEehHHHHHHHHHHH
Confidence 45443333 3389999999999999999998873 34788988877764
No 170
>KOG2589 consensus Histone tail methylase [Chromatin structure and dynamics]
Probab=23.35 E-value=37 Score=35.34 Aligned_cols=30 Identities=27% Similarity=0.516 Sum_probs=23.5
Q ss_pred ccccceeccccccccCCCCccceeCCCCccccccCceeE
Q 019981 279 GTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMV 317 (333)
Q Consensus 279 GeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~ 317 (333)
|+|+.-|+|. +.=..++..| |||.||+..+
T Consensus 229 GeEITcFYgs-----~fFG~~N~~C----eC~TCER~g~ 258 (453)
T KOG2589|consen 229 GEEITCFYGS-----GFFGENNEEC----ECVTCERRGT 258 (453)
T ss_pred CceeEEeecc-----cccCCCCcee----EEeecccccc
Confidence 8999999986 3445566677 7999998765
No 171
>PRK15398 aldehyde dehydrogenase EutE; Provisional
Probab=23.33 E-value=1.7e+02 Score=30.13 Aligned_cols=59 Identities=12% Similarity=0.294 Sum_probs=44.4
Q ss_pred ccChHHHhhhHhhhhh-----cCCeeE------EeChhhHHHHHHHHhhhcCCCccChHHHHHHHHHHhhcC
Q 019981 116 VMSNEEFDNLKEELMW-----EGSSVV------MLSSAEQKFLEASMAYVAGKPIMSDEEYDKLKQKLKMEG 176 (333)
Q Consensus 116 ~~sdeefd~LkEeL~w-----eGSsvv------~L~~~Eq~fLEA~~AY~sGkPimsDeeFD~LK~kLk~~G 176 (333)
++.|.+.|.--+.+.| .|..|. +=...-.+|++++.+. +.|+++.+++|+++.-+...|
T Consensus 247 V~~dADld~Aa~~i~~g~~~n~GQ~C~A~~rvlV~~si~d~f~~~l~~~--~~~li~~~~~~~v~~~l~~~~ 316 (465)
T PRK15398 247 VDETADIEKAARDIVKGASFDNNLPCIAEKEVIVVDSVADELMRLMEKN--GAVLLTAEQAEKLQKVVLKNG 316 (465)
T ss_pred EecCCCHHHHHHHHHHhcccCCCCcCCCCceEEEeHHHHHHHHHHHHHc--CCccCCHHHHHHHHHHHhhcc
Confidence 4457788888888888 465443 3333446799999887 789999999999998777554
No 172
>PRK06427 bifunctional hydroxy-methylpyrimidine kinase/ hydroxy-phosphomethylpyrimidine kinase; Reviewed
Probab=23.04 E-value=3.1e+02 Score=25.03 Aligned_cols=54 Identities=22% Similarity=0.388 Sum_probs=36.2
Q ss_pred hhhHhhhhhcCCeeEEeChhhHHHHHHHHhhhcCCCccChHH-HHHHHHHHhhcCC-ceeeecC
Q 019981 123 DNLKEELMWEGSSVVMLSSAEQKFLEASMAYVAGKPIMSDEE-YDKLKQKLKMEGS-EIVVEGP 184 (333)
Q Consensus 123 d~LkEeL~weGSsvv~L~~~Eq~fLEA~~AY~sGkPimsDee-FD~LK~kLk~~GS-~VVvk~P 184 (333)
+.++++| .....+++.+..|-+.| .|.++-++++ ..+.-.+|...|- .|++++-
T Consensus 124 ~~~~~~l-l~~~dvitpN~~Ea~~L-------~g~~~~~~~~~~~~~a~~l~~~g~~~Vvit~g 179 (266)
T PRK06427 124 AALRERL-LPLATLITPNLPEAEAL-------TGLPIADTEDEMKAAARALHALGCKAVLIKGG 179 (266)
T ss_pred HHHHHhh-hCcCeEEcCCHHHHHHH-------hCCCCCCcHHHHHHHHHHHHhcCCCEEEEcCC
Confidence 4566654 35677999999998877 3666655554 5566667777775 4666653
No 173
>PF00412 LIM: LIM domain; InterPro: IPR001781 Zinc finger (Znf) domains are relatively small protein motifs which contain multiple finger-like protrusions that make tandem contacts with their target molecule. Some of these domains bind zinc, but many do not; instead binding other metals such as iron, or no metal at all. For example, some family members form salt bridges to stabilise the finger-like folds. They were first identified as a DNA-binding motif in transcription factor TFIIIA from Xenopus laevis (African clawed frog), however they are now recognised to bind DNA, RNA, protein and/or lipid substrates [, , , , ]. Their binding properties depend on the amino acid sequence of the finger domains and of the linker between fingers, as well as on the higher-order structures and the number of fingers. Znf domains are often found in clusters, where fingers can have different binding specificities. There are many superfamilies of Znf motifs, varying in both sequence and structure. They display considerable versatility in binding modes, even between members of the same class (e.g. some bind DNA, others protein), suggesting that Znf motifs are stable scaffolds that have evolved specialised functions. For example, Znf-containing proteins function in gene transcription, translation, mRNA trafficking, cytoskeleton organisation, epithelial development, cell adhesion, protein folding, chromatin remodelling and zinc sensing, to name but a few []. Zinc-binding motifs are stable structures, and they rarely undergo conformational changes upon binding their target. This entry represents LIM-type zinc finger (Znf) domains. LIM domains coordinate one or more zinc atoms, and are named after the three proteins (LIN-11, Isl1 and MEC-3) in which they were first found. They consist of two zinc-binding motifs that resemble GATA-like Znf's, however the residues holding the zinc atom(s) are variable, involving Cys, His, Asp or Glu residues. LIM domains are involved in proteins with differing functions, including gene expression, and cytoskeleton organisation and development [, ]. Protein containing LIM Znf domains include: Caenorhabditis elegans mec-3; a protein required for the differentiation of the set of six touch receptor neurons in this nematode. C. elegans. lin-11; a protein required for the asymmetric division of vulval blast cells. Vertebrate insulin gene enhancer binding protein isl-1. Isl-1 binds to one of the two cis-acting protein-binding domains of the insulin gene. Vertebrate homeobox proteins lim-1, lim-2 (lim-5) and lim3. Vertebrate lmx-1, which acts as a transcriptional activator by binding to the FLAT element; a beta-cell-specific transcriptional enhancer found in the insulin gene. Mammalian LH-2, a transcriptional regulatory protein involved in the control of cell differentiation in developing lymphoid and neural cell types. Drosophila melanogaster (Fruit fly) protein apterous, required for the normal development of the wing and halter imaginal discs. Vertebrate protein kinases LIMK-1 and LIMK-2. Mammalian rhombotins. Rhombotin 1 (RBTN1 or TTG-1) and rhombotin-2 (RBTN2 or TTG-2) are proteins of about 160 amino acids whose genes are disrupted by chromosomal translocations in T-cell leukemia. Mammalian and avian cysteine-rich protein (CRP), a 192 amino-acid protein of unknown function. Seems to interact with zyxin. Mammalian cysteine-rich intestinal protein (CRIP), a small protein which seems to have a role in zinc absorption and may function as an intracellular zinc transport protein. Vertebrate paxillin, a cytoskeletal focal adhesion protein. Mus musculus (Mouse) testin which should not be confused with rat testin which is a thiol protease homologue (see IPR000169 from INTERPRO). Helianthus annuus (Common sunflower) pollen specific protein SF3. Chicken zyxin. Zyxin is a low-abundance adhesion plaque protein which has been shown to interact with CRP. Yeast protein LRG1 which is involved in sporulation []. Saccharomyces cerevisiae (Baker's yeast) rho-type GTPase activating protein RGA1/DBM1. C. elegans homeobox protein ceh-14. C. elegans homeobox protein unc-97. S. cerevisiae hypothetical protein YKR090w. C. elegans hypothetical proteins C28H8.6. These proteins generally contain two tandem copies of the LIM domain in their N-terminal section. Zyxin and paxillin are exceptions in that they contain respectively three and four LIM domains at their C-terminal extremity. In apterous, isl-1, LH-2, lin-11, lim-1 to lim-3, lmx-1 and ceh-14 and mec-3 there is a homeobox domain some 50 to 95 amino acids after the LIM domains. LIM domains contain seven conserved cysteine residues and a histidine. The arrangement followed by these conserved residues is: C-x(2)-C-x(16,23)-H-x(2)-[CH]-x(2)-C-x(2)-C-x(16,21)-C-x(2,3)-[CHD] LIM domains bind two zinc ions []. LIM does not bind DNA, rather it seems to act as an interface for protein-protein interaction. More information about these proteins can be found at Protein of the Month: Zinc Fingers [].; GO: 0008270 zinc ion binding; PDB: 2CO8_A 2EGQ_A 2CUR_A 3IXE_B 1CTL_A 1B8T_A 1X62_A 2DFY_C 1IML_A 2CUQ_A ....
Probab=23.03 E-value=34 Score=24.10 Aligned_cols=37 Identities=30% Similarity=0.611 Sum_probs=19.4
Q ss_pred CCCCccccceeccccccccCCCCccceeCCCCccccccCceeE
Q 019981 275 CPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMV 317 (333)
Q Consensus 275 CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~ 317 (333)
|+.|+++|. ++...+...+..--..|= -|..|++.|.
T Consensus 1 C~~C~~~I~---~~~~~~~~~~~~~H~~Cf---~C~~C~~~l~ 37 (58)
T PF00412_consen 1 CARCGKPIY---GTEIVIKAMGKFWHPECF---KCSKCGKPLN 37 (58)
T ss_dssp BTTTSSBES---SSSEEEEETTEEEETTTS---BETTTTCBTT
T ss_pred CCCCCCCcc---CcEEEEEeCCcEEEcccc---ccCCCCCccC
Confidence 788998887 333332222222223342 3666887763
No 174
>TIGR01405 polC_Gram_pos DNA polymerase III, alpha chain, Gram-positive type. The N-terminal region of about 200 amino acids is rich in low-complexity sequence, poorly alignable, and not included n this model.
Probab=22.99 E-value=48 Score=38.49 Aligned_cols=36 Identities=36% Similarity=0.581 Sum_probs=25.8
Q ss_pred CCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEE
Q 019981 274 PCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVY 318 (333)
Q Consensus 274 ~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f 318 (333)
-||||. ++-|-++-++.||-+--.-+|++ ||+.|.=
T Consensus 685 ~c~~c~---~~ef~~~~~~~sg~dlp~k~cp~------c~~~~~~ 720 (1213)
T TIGR01405 685 LCPNCK---YSEFITDGSVGSGFDLPDKDCPK------CGAPLKK 720 (1213)
T ss_pred cCcccc---cccccccccccccccCccccCcc------ccccccc
Confidence 499995 43344556777777777778888 9998753
No 175
>PRK03824 hypA hydrogenase nickel incorporation protein; Provisional
Probab=22.89 E-value=76 Score=27.56 Aligned_cols=12 Identities=33% Similarity=0.545 Sum_probs=9.2
Q ss_pred eeeecCCCCCcc
Q 019981 269 LILKGPCPNCGT 280 (333)
Q Consensus 269 ~iLKG~CPNCGe 280 (333)
+-..+-|++||.
T Consensus 67 ~p~~~~C~~CG~ 78 (135)
T PRK03824 67 EEAVLKCRNCGN 78 (135)
T ss_pred cceEEECCCCCC
Confidence 336788999993
No 176
>COG3464 Transposase and inactivated derivatives [DNA replication, recombination, and repair]
Probab=22.86 E-value=62 Score=32.81 Aligned_cols=47 Identities=21% Similarity=0.476 Sum_probs=24.9
Q ss_pred eeecCCCCCccccceecc----ccccccCCCCccceeCCCCcc-ccccCcee
Q 019981 270 ILKGPCPNCGTENVSFFG----TILSISSGGTTNTINCSNLTF-CFSCGTTM 316 (333)
Q Consensus 270 iLKG~CPNCGeEv~aFfg----tilsv~s~~~~n~vkC~~~ae-CHVC~t~L 316 (333)
+.|..||.||+..-.--+ -|--++=++-+-.+.+..+.. |+.|++..
T Consensus 36 ~~~~~CP~Cg~~~~~~~~~~~~~I~~L~~~~~~~~L~~r~rR~~c~~c~~~~ 87 (402)
T COG3464 36 PRKHRCPECGQRTIRRHGWRIRKIQDLPLFEVPVYLFLRKRRYKCCRCGKRF 87 (402)
T ss_pred cccCCCCCCCCcceeccccceeeeeecccCCeeEEEEeccceeecccCCCCc
Confidence 444999999999733211 112222233333444444333 66698875
No 177
>PF08976 DUF1880: Domain of unknown function (DUF1880); InterPro: IPR015070 This entry represents EF-hand calcium-binding domain-containing protein 6 that negatively regulates the androgen receptor by recruiting histone deacetylase complex, and protein DJ-1 antagonises this inhibition by abrogation of this complex [].; PDB: 1WLZ_C.
Probab=22.54 E-value=39 Score=29.70 Aligned_cols=19 Identities=32% Similarity=0.623 Sum_probs=11.3
Q ss_pred cccChHHHhhhHhh--hhhcC
Q 019981 115 AVMSNEEFDNLKEE--LMWEG 133 (333)
Q Consensus 115 ~~~sdeefd~LkEe--L~weG 133 (333)
++|+||+||+|=.| |+|.|
T Consensus 2 qiLtDeQFdrLW~e~Pvn~~G 22 (118)
T PF08976_consen 2 QILTDEQFDRLWNEMPVNAKG 22 (118)
T ss_dssp ----HHHHHHHHTTS-B-TTS
T ss_pred ccccHHHhhhhhhhCcCCccC
Confidence 58999999999777 35555
No 178
>PF13913 zf-C2HC_2: zinc-finger of a C2HC-type
Probab=22.31 E-value=40 Score=21.54 Aligned_cols=8 Identities=63% Similarity=1.788 Sum_probs=5.6
Q ss_pred CCCCCccc
Q 019981 274 PCPNCGTE 281 (333)
Q Consensus 274 ~CPNCGeE 281 (333)
+||+||..
T Consensus 4 ~C~~CgR~ 11 (25)
T PF13913_consen 4 PCPICGRK 11 (25)
T ss_pred cCCCCCCE
Confidence 68888753
No 179
>KOG3457 consensus Sec61 protein translocation complex, beta subunit [Posttranslational modification, protein turnover, chaperones]
Probab=22.16 E-value=47 Score=27.91 Aligned_cols=30 Identities=30% Similarity=0.527 Sum_probs=20.6
Q ss_pred hhhcccccccceeeeeccCCCCchhHHHHH
Q 019981 218 LFFFLDDITGFEITYLLELPEPFSFIFTWF 247 (333)
Q Consensus 218 l~~~ldd~~Gf~i~~~~~lpep~gfi~t~~ 247 (333)
|.|+-||+.|+.|....-|-=+++||+.++
T Consensus 46 lkfYTDda~GlKV~PvvVLvmSvgFIasV~ 75 (88)
T KOG3457|consen 46 LKFYTDDAPGLKVDPVVVLVMSVGFIASVF 75 (88)
T ss_pred eEEeecCCCCceeCCeeehhhhHHHHHHHH
Confidence 456679999999866555544567776544
No 180
>PF05907 DUF866: Eukaryotic protein of unknown function (DUF866); InterPro: IPR008584 This family consists of a number of hypothetical eukaryotic proteins of unknown function with an average length of around 165 residues.; PDB: 1ZSO_B.
Probab=22.16 E-value=67 Score=28.85 Aligned_cols=45 Identities=20% Similarity=0.448 Sum_probs=21.6
Q ss_pred eeeecCCCCCccccc--eeccc--cccccCCCCccc--eeCCCCccccccCceeEEe
Q 019981 269 LILKGPCPNCGTENV--SFFGT--ILSISSGGTTNT--INCSNLTFCFSCGTTMVYD 319 (333)
Q Consensus 269 ~iLKG~CPNCGeEv~--aFfgt--ilsv~s~~~~n~--vkC~~~aeCHVC~t~L~f~ 319 (333)
-.+|-.|.||||+.- .++-. ...+++++..++ .||.. |++....+
T Consensus 27 ~~fkvkCt~CgE~~~k~V~i~~~e~~e~~gsrG~aNfv~KCk~------C~re~si~ 77 (161)
T PF05907_consen 27 WFFKVKCTSCGEVHPKWVYINRFEKHEIPGSRGTANFVMKCKF------CKRESSID 77 (161)
T ss_dssp EEEEEEETTSS--EEEEEEE-TT-BEE-TTSS-EESEEE--SS------SS--EEEE
T ss_pred EEEEEEECCCCCccCcceEeecceEEecCCCccceEeEecCcC------cCCccEEE
Confidence 457888999999754 44331 122444444333 48887 99977664
No 181
>KOG1372 consensus GDP-mannose 4,6 dehydratase [Carbohydrate transport and metabolism]
Probab=21.95 E-value=47 Score=33.39 Aligned_cols=39 Identities=28% Similarity=0.444 Sum_probs=28.7
Q ss_pred HHHHHHHHHHhhhcCccc---------cChHHHhh-----hHhhhhhcCCeeE
Q 019981 99 EQEFLQALQAFYYEGKAV---------MSNEEFDN-----LKEELMWEGSSVV 137 (333)
Q Consensus 99 E~~fl~Al~~fY~~gk~~---------~sdeefd~-----LkEeL~weGSsvv 137 (333)
-.+|.+|||.--...+|. -|-+||++ +-|+|+|+|..|=
T Consensus 257 A~dYVEAMW~mLQ~d~PdDfViATge~hsVrEF~~~aF~~ig~~l~Weg~gv~ 309 (376)
T KOG1372|consen 257 AGDYVEAMWLMLQQDSPDDFVIATGEQHSVREFCNLAFAEIGEVLNWEGEGVD 309 (376)
T ss_pred hHHHHHHHHHHHhcCCCCceEEecCCcccHHHHHHHHHHhhCcEEeecccccc
Confidence 358999999988877762 24556655 6799999987654
No 182
>PF10751 DUF2535: Protein of unknown function (DUF2535); InterPro: IPR019687 This entry represents proteins with unknown function, and appear to be restricted to Bacillus spp.
Probab=21.87 E-value=88 Score=26.12 Aligned_cols=40 Identities=25% Similarity=0.213 Sum_probs=28.0
Q ss_pred eEEeChhh------HHHHHHHHh--hhcCCCccChHHHHHHHHHHhhc
Q 019981 136 VVMLSSAE------QKFLEASMA--YVAGKPIMSDEEYDKLKQKLKME 175 (333)
Q Consensus 136 vv~L~~~E------q~fLEA~~A--Y~sGkPimsDeeFD~LK~kLk~~ 175 (333)
+++|.+++ |.-||+.++ |.+..|--+=.-=|-||+.|||.
T Consensus 21 IPVL~ed~p~~Fmi~~rLq~fi~~vy~~~~~~~vYSFreYlKr~lKW~ 68 (83)
T PF10751_consen 21 IPVLEEDNPYYFMIQLRLQLFIAKVYNSKSPRKVYSFREYLKRVLKWP 68 (83)
T ss_pred cceecCCCceEeeHHHHHHHHHHHHHhCCCCCceeeHHHHHHHhcCcH
Confidence 34555555 567888776 77766666655666799999996
No 183
>PF14952 zf-tcix: Putative treble-clef, zinc-finger, Zn-binding
Probab=21.82 E-value=41 Score=25.10 Aligned_cols=9 Identities=67% Similarity=1.516 Sum_probs=8.0
Q ss_pred CCCCCcccc
Q 019981 274 PCPNCGTEN 282 (333)
Q Consensus 274 ~CPNCGeEv 282 (333)
.||.||+.|
T Consensus 13 kCp~CGt~N 21 (44)
T PF14952_consen 13 KCPKCGTYN 21 (44)
T ss_pred cCCcCcCcc
Confidence 599999977
No 184
>TIGR00777 ahpD alkylhydroperoxidase, AhpD family. Members of this family are alkylhydroperoxidases, which catalyze the reduction of peroxides to their corresponding alcohols via oxidation of cysteine residues. In these alkylhydroperoxidases, the cysteines are located in a conserved -CXXC- motif located towards the COOH terminus. In Mycobacterium tuberculosis, two non-homologous alkylhydroperoxidases, AhpD and AhpC, are found in the same operon.
Probab=21.80 E-value=47 Score=30.82 Aligned_cols=26 Identities=23% Similarity=0.569 Sum_probs=23.0
Q ss_pred HHh-hhcCCCccChHHHHHHHHHHhhc
Q 019981 150 SMA-YVAGKPIMSDEEYDKLKQKLKME 175 (333)
Q Consensus 150 ~~A-Y~sGkPimsDeeFD~LK~kLk~~ 175 (333)
|.- ||+...+++|++|+.++.+||..
T Consensus 79 mnNv~Yr~~hl~~~~~y~~~pa~lrmn 105 (177)
T TIGR00777 79 MNNVFYRGRHLLEGARYDDLRPGLRMN 105 (177)
T ss_pred hhhHHHHhHhhcccchhhcCCccchhH
Confidence 444 99999999999999999999876
No 185
>PRK12380 hydrogenase nickel incorporation protein HybF; Provisional
Probab=21.61 E-value=51 Score=27.78 Aligned_cols=14 Identities=14% Similarity=0.190 Sum_probs=10.7
Q ss_pred ceeeeecCCCCCcc
Q 019981 267 ESLILKGPCPNCGT 280 (333)
Q Consensus 267 D~~iLKG~CPNCGe 280 (333)
+.+-+.+-|++||.
T Consensus 65 ~~vp~~~~C~~Cg~ 78 (113)
T PRK12380 65 VYKPAQAWCWDCSQ 78 (113)
T ss_pred EeeCcEEEcccCCC
Confidence 44557889999994
No 186
>COG2960 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=21.51 E-value=1.1e+02 Score=26.43 Aligned_cols=51 Identities=27% Similarity=0.346 Sum_probs=38.3
Q ss_pred ccchhHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhhhcCCeeEEeChhhHHHHHHHH
Q 019981 92 KKSLGELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELMWEGSSVVMLSSAEQKFLEASM 151 (333)
Q Consensus 92 k~slge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~weGSsvv~L~~~Eq~fLEA~~ 151 (333)
+..-+|.|.-|-+-+|+.++ .....+.||||..++.|.= .+++-+-|||..
T Consensus 32 ~~~~~evE~~~r~~~q~~ln-kLDlVsREEFdvq~qvl~r--------tR~kl~~Leari 82 (103)
T COG2960 32 QEVRAEVEKAFRAQLQRQLN-KLDLVSREEFDVQRQVLLR--------TREKLAALEARI 82 (103)
T ss_pred hhhHHHHHHHHHHHHHHHHh-hhhhhhHHHHHHHHHHHHH--------HHHHHHHHHHHH
Confidence 34558999999999999986 5789999999999998543 234445555543
No 187
>PF09845 DUF2072: Zn-ribbon containing protein (DUF2072); InterPro: IPR018645 This archaeal Zinc-ribbon containing proteins have no known function.
Probab=21.45 E-value=44 Score=29.79 Aligned_cols=21 Identities=33% Similarity=0.822 Sum_probs=17.6
Q ss_pred eeeeecCCCCCccccceecccc
Q 019981 268 SLILKGPCPNCGTENVSFFGTI 289 (333)
Q Consensus 268 ~~iLKG~CPNCGeEv~aFfgti 289 (333)
..||+| ||+||---|.|+...
T Consensus 16 ~eil~G-CP~CGg~kF~yv~~~ 36 (131)
T PF09845_consen 16 KEILSG-CPECGGNKFQYVPEE 36 (131)
T ss_pred HHHHcc-CcccCCcceEEcCCC
Confidence 357777 999999999999764
No 188
>PRK03564 formate dehydrogenase accessory protein FdhE; Provisional
Probab=21.45 E-value=1.2e+02 Score=30.34 Aligned_cols=13 Identities=38% Similarity=0.853 Sum_probs=10.6
Q ss_pred eecCCCCCccccc
Q 019981 271 LKGPCPNCGTENV 283 (333)
Q Consensus 271 LKG~CPNCGeEv~ 283 (333)
-+|-||.||..=.
T Consensus 186 ~~~~CPvCGs~P~ 198 (309)
T PRK03564 186 QRQFCPVCGSMPV 198 (309)
T ss_pred CCCCCCCCCCcch
Confidence 5799999998743
No 189
>cd07092 ALDH_ABALDH-YdcW Escherichia coli NAD+-dependent gamma-aminobutyraldehyde dehydrogenase YdcW-like. NAD+-dependent, tetrameric, gamma-aminobutyraldehyde dehydrogenase (ABALDH), YdcW of Escherichia coli K12, catalyzes the oxidation of gamma-aminobutyraldehyde to gamma-aminobutyric acid. ABALDH can also oxidize n-alkyl medium-chain aldehydes, but with a lower catalytic efficiency.
Probab=21.38 E-value=2.8e+02 Score=27.67 Aligned_cols=68 Identities=16% Similarity=0.322 Sum_probs=46.0
Q ss_pred ccChHHHhhhHhhhhh-----cCCee-----EEeC-hhhHHHHHHHHh----hhcCC---------CccChHHHHHHHHH
Q 019981 116 VMSNEEFDNLKEELMW-----EGSSV-----VMLS-SAEQKFLEASMA----YVAGK---------PIMSDEEYDKLKQK 171 (333)
Q Consensus 116 ~~sdeefd~LkEeL~w-----eGSsv-----v~L~-~~Eq~fLEA~~A----Y~sGk---------PimsDeeFD~LK~k 171 (333)
++.|.+.|..=+.+.| .|-.| |.+- ..-.+|++++.+ +.-|. |+++.+.+++++.-
T Consensus 235 V~~dAdl~~aa~~iv~~~~~~~GQ~C~a~~~v~V~~~i~~~f~~~l~~~~~~~~~g~p~~~~~~~gpli~~~~~~~i~~~ 314 (450)
T cd07092 235 VFDDADLDAAVAGIATAGYYNAGQDCTAACRVYVHESVYDEFVAALVEAVSAIRVGDPDDEDTEMGPLNSAAQRERVAGF 314 (450)
T ss_pred ECCCCCHHHHHHHHHHHHHhhCCCCCCCCcEEEEeHHHHHHHHHHHHHHHhhCCcCCCCCCCCccCcccCHHHHHHHHHH
Confidence 4568888888888888 44433 3343 344689988754 33453 57888999999986
Q ss_pred Hhhc--CCceeeec
Q 019981 172 LKME--GSEIVVEG 183 (333)
Q Consensus 172 Lk~~--GS~VVvk~ 183 (333)
+... |.+++.-|
T Consensus 315 i~~a~~ga~~~~gg 328 (450)
T cd07092 315 VERAPAHARVLTGG 328 (450)
T ss_pred HHHHHcCCEEEeCC
Confidence 6543 77766544
No 190
>PF04475 DUF555: Protein of unknown function (DUF555); InterPro: IPR007564 This is a family of uncharacterised, hypothetical archaeal proteins.
Probab=21.37 E-value=44 Score=28.80 Aligned_cols=15 Identities=40% Similarity=0.740 Sum_probs=11.5
Q ss_pred eeecCCCCCccccce
Q 019981 270 ILKGPCPNCGTENVS 284 (333)
Q Consensus 270 iLKG~CPNCGeEv~a 284 (333)
+=.-.||.||+|..+
T Consensus 45 vG~~~cP~Cge~~~~ 59 (102)
T PF04475_consen 45 VGDTICPKCGEELDS 59 (102)
T ss_pred cCcccCCCCCCccCc
Confidence 344579999999873
No 191
>COG3024 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=21.31 E-value=44 Score=26.72 Aligned_cols=14 Identities=43% Similarity=1.030 Sum_probs=11.3
Q ss_pred eeecCCCCCccccc
Q 019981 270 ILKGPCPNCGTENV 283 (333)
Q Consensus 270 iLKG~CPNCGeEv~ 283 (333)
.+.-+||-||.+|-
T Consensus 5 ~~~v~CP~Cgkpv~ 18 (65)
T COG3024 5 RITVPCPTCGKPVV 18 (65)
T ss_pred cccccCCCCCCccc
Confidence 45678999999875
No 192
>smart00709 Zpr1 Duplicated domain in the epidermal growth factor- and elongation factor-1alpha-binding protein Zpr1. Also present in archaeal proteins.
Probab=21.29 E-value=46 Score=29.93 Aligned_cols=6 Identities=67% Similarity=2.073 Sum_probs=4.0
Q ss_pred CCCCCc
Q 019981 274 PCPNCG 279 (333)
Q Consensus 274 ~CPNCG 279 (333)
.|||||
T Consensus 2 ~Cp~C~ 7 (160)
T smart00709 2 DCPSCG 7 (160)
T ss_pred cCCCCC
Confidence 477775
No 193
>PF04328 DUF466: Protein of unknown function (DUF466); InterPro: IPR007423 This is a small bacterial protein of unknown function.
Probab=21.20 E-value=1.8e+02 Score=22.76 Aligned_cols=34 Identities=18% Similarity=0.332 Sum_probs=28.3
Q ss_pred HHHHHHHHhhhcCCCccChHHHHHHHHHHhhcCC
Q 019981 144 QKFLEASMAYVAGKPIMSDEEYDKLKQKLKMEGS 177 (333)
Q Consensus 144 q~fLEA~~AY~sGkPimsDeeFD~LK~kLk~~GS 177 (333)
..+|+=..+--.|+|+||-+||-+-..+=++.|-
T Consensus 26 e~Yv~H~~~~HP~~p~ms~~eF~r~r~~~r~~~~ 59 (65)
T PF04328_consen 26 ERYVEHMRRHHPDEPPMSEREFFRERQDARYGNP 59 (65)
T ss_pred HHHHHHHHHHCcCCCCCCHHHHHHHHHHHHhcCC
Confidence 5688888888899999999999988777776553
No 194
>PF09334 tRNA-synt_1g: tRNA synthetases class I (M); InterPro: IPR015413 The aminoacyl-tRNA synthetases (6.1.1. from EC) catalyse the attachment of an amino acid to its cognate transfer RNA molecule in a highly specific two-step reaction. These proteins differ widely in size and oligomeric state, and have limited sequence homology []. The 20 aminoacyl-tRNA synthetases are divided into two classes, I and II. Class I aminoacyl-tRNA synthetases contain a characteristic Rossman fold catalytic domain and are mostly monomeric []. Class II aminoacyl-tRNA synthetases share an anti-parallel beta-sheet fold flanked by alpha-helices [], and are mostly dimeric or multimeric, containing at least three conserved regions [, , ]. However, tRNA binding involves an alpha-helical structure that is conserved between class I and class II synthetases. In reactions catalysed by the class I aminoacyl-tRNA synthetases, the aminoacyl group is coupled to the 2'-hydroxyl of the tRNA, while, in class II reactions, the 3'-hydroxyl site is preferred. The synthetases specific for arginine, cysteine, glutamic acid, glutamine, isoleucine, leucine, methionine, tyrosine, tryptophan and valine belong to class I synthetases. The synthetases specific for alanine, asparagine, aspartic acid, glycine, histidine, lysine, phenylalanine, proline, serine, and threonine belong to class-II synthetases []. Based on their mode of binding to the tRNA acceptor stem, both classes of tRNA synthetases have been subdivided into three subclasses, designated 1a, 1b, 1c and 2a, 2b, 2c. This domain is found in methionyl and leucyl tRNA synthetases. ; GO: 0000166 nucleotide binding, 0004812 aminoacyl-tRNA ligase activity, 0005524 ATP binding, 0006418 tRNA aminoacylation for protein translation, 0005737 cytoplasm; PDB: 2D5B_A 1A8H_A 1WOY_A 2D54_A 4DLP_A 2CT8_B 2CSX_A 1MED_A 1PFU_A 1PFW_A ....
Probab=21.11 E-value=62 Score=32.46 Aligned_cols=9 Identities=56% Similarity=1.682 Sum_probs=6.1
Q ss_pred eecCCCCCc
Q 019981 271 LKGPCPNCG 279 (333)
Q Consensus 271 LKG~CPNCG 279 (333)
++|.||.||
T Consensus 135 v~g~CP~C~ 143 (391)
T PF09334_consen 135 VEGTCPYCG 143 (391)
T ss_dssp ETCEETTT-
T ss_pred eeccccCcC
Confidence 468888887
No 195
>TIGR00310 ZPR1_znf ZPR1 zinc finger domain.
Probab=21.10 E-value=47 Score=30.82 Aligned_cols=20 Identities=25% Similarity=0.416 Sum_probs=17.8
Q ss_pred hhcceeeeecCCCCCccccc
Q 019981 264 IVRESLILKGPCPNCGTENV 283 (333)
Q Consensus 264 ~~~D~~iLKG~CPNCGeEv~ 283 (333)
.+++.+|+...||+||-.+.
T Consensus 22 ~F~evii~sf~C~~CGyr~~ 41 (192)
T TIGR00310 22 YFGEVLETSTICEHCGYRSN 41 (192)
T ss_pred CcceEEEEEEECCCCCCccc
Confidence 38999999999999998776
No 196
>cd02674 Peptidase_C19R A subfamily of peptidase C19. Peptidase C19 contains ubiquitinyl hydrolases. They are intracellular peptidases that remove ubiquitin molecules from polyubiquinated peptides by cleavage of isopeptide bonds. They hydrolyze bonds involving the carboxyl group of the C-terminal Gly residue of ubiquitin. The purpose of the de-ubiquitination is thought to be editing of the ubiquitin conjugates, which could rescue them from degradation, as well as recycling of the ubiquitin. The ubiquitin/proteasome system is responsible for most protein turnover in the mammalian cell, and with over 50 members, family C19 is one of the largest families of peptidases in the human genome.
Probab=20.77 E-value=99 Score=27.01 Aligned_cols=25 Identities=20% Similarity=0.456 Sum_probs=16.3
Q ss_pred cceeCCCCccccccCceeEEeccceeeecCC
Q 019981 299 NTINCSNLTFCFSCGTTMVYDSNTRLITLPE 329 (333)
Q Consensus 299 n~vkC~~~aeCHVC~t~L~f~tk~R~itl~~ 329 (333)
+..+|+. |+..-...+..++.++|+
T Consensus 103 ~~~~C~~------C~~~~~~~~~~~i~~lP~ 127 (230)
T cd02674 103 NAWKCPK------CKKKRKATKKLTISRLPK 127 (230)
T ss_pred CceeCCC------CCCccceEEEEEEecCCh
Confidence 4566776 887766666666666664
No 197
>cd07105 ALDH_SaliADH Salicylaldehyde dehydrogenase, DoxF-like. Salicylaldehyde dehydrogenase (DoxF, SaliADH, EC=1.2.1.65) involved in the upper naphthalene catabolic pathway of Pseudomonas strain C18 and other similar sequences are present in this CD.
Probab=20.76 E-value=2.7e+02 Score=27.73 Aligned_cols=70 Identities=21% Similarity=0.446 Sum_probs=46.8
Q ss_pred ccChHHHhhhHhhhhh-----cCCee-----EEeCh-hhHHHHHHHHh----hhcC----CCccChHHHHHHHHHHhh--
Q 019981 116 VMSNEEFDNLKEELMW-----EGSSV-----VMLSS-AEQKFLEASMA----YVAG----KPIMSDEEYDKLKQKLKM-- 174 (333)
Q Consensus 116 ~~sdeefd~LkEeL~w-----eGSsv-----v~L~~-~Eq~fLEA~~A----Y~sG----kPimsDeeFD~LK~kLk~-- 174 (333)
++.|.+.|.--+.+.| .|-.| +.+-+ .-.+|+|++.+ +.-| -|+++...+++++.-+..
T Consensus 219 V~~dadl~~aa~~i~~~~~~~~GQ~C~a~~~v~V~~~i~~~f~~~l~~~~~~~~~g~~~~gp~i~~~~~~~~~~~i~~a~ 298 (432)
T cd07105 219 VLEDADLDAAANAALFGAFLNSGQICMSTERIIVHESIADEFVEKLKAAAEKLFAGPVVLGSLVSAAAADRVKELVDDAL 298 (432)
T ss_pred ECCCCCHHHHHHHHHHHHHhcCCCCCcCCceEEEcHHHHHHHHHHHHHHHHhhcCCCCcccccCCHHHHHHHHHHHHHHH
Confidence 4557788888777777 35433 33333 34578888754 3322 389999999999988765
Q ss_pred -cCCceeeecCe
Q 019981 175 -EGSEIVVEGPR 185 (333)
Q Consensus 175 -~GS~VVvk~Pr 185 (333)
.|.+++.-|.+
T Consensus 299 ~~ga~~~~gg~~ 310 (432)
T cd07105 299 SKGAKLVVGGLA 310 (432)
T ss_pred HCCCEEEeCCCc
Confidence 57787775543
No 198
>PF11682 DUF3279: Protein of unknown function (DUF3279); InterPro: IPR021696 This family of proteins with unknown function appears to be restricted to Enterobacteriaceae.
Probab=20.70 E-value=53 Score=29.04 Aligned_cols=15 Identities=27% Similarity=0.727 Sum_probs=11.9
Q ss_pred ccccccCceeEEecc
Q 019981 307 TFCFSCGTTMVYDSN 321 (333)
Q Consensus 307 aeCHVC~t~L~f~tk 321 (333)
-.||.||++|.|...
T Consensus 29 ~tC~~Cg~~L~lh~~ 43 (128)
T PF11682_consen 29 WTCHSCGCPLILHPG 43 (128)
T ss_pred EEEecCCceEEEecC
Confidence 367889999999843
No 199
>PHA02942 putative transposase; Provisional
Probab=20.70 E-value=69 Score=32.27 Aligned_cols=27 Identities=30% Similarity=0.843 Sum_probs=18.4
Q ss_pred cCCCCCccccceeccccccccCCCCccceeCCCCccccccCcee
Q 019981 273 GPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTM 316 (333)
Q Consensus 273 G~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L 316 (333)
--||+||..+.. + + ...++|++ ||..+
T Consensus 326 q~Cs~CG~~~~~-----l---~---~r~f~C~~------CG~~~ 352 (383)
T PHA02942 326 VSCPKCGHKMVE-----I---A---HRYFHCPS------CGYEN 352 (383)
T ss_pred ccCCCCCCccCc-----C---C---CCEEECCC------CCCEe
Confidence 349999987631 1 1 23689988 99865
No 200
>PF04423 Rad50_zn_hook: Rad50 zinc hook motif; InterPro: IPR007517 The Mre11 complex (Mre11 Rad50 Nbs1) is central to chromosomal maintenance and functions in homologous recombination, telomere maintenance and sister chromatid association. The Rad50 coiled-coil region contains a dimer interface at the apex of the coiled coils in which pairs of conserved Cys-X-X-Cys motifs form interlocking hooks that bind one Zn ion. This alignment includes the zinc hook motif and a short stretch of coiled-coil on either side.; GO: 0004518 nuclease activity, 0005524 ATP binding, 0008270 zinc ion binding, 0006281 DNA repair; PDB: 1L8D_B.
Probab=20.70 E-value=43 Score=24.42 Aligned_cols=6 Identities=33% Similarity=0.993 Sum_probs=1.8
Q ss_pred cCceeE
Q 019981 312 CGTTMV 317 (333)
Q Consensus 312 C~t~L~ 317 (333)
|++.|.
T Consensus 26 C~r~l~ 31 (54)
T PF04423_consen 26 CGRPLD 31 (54)
T ss_dssp T--EE-
T ss_pred CCCCCC
Confidence 666553
No 201
>PF08278 DnaG_DnaB_bind: DNA primase DnaG DnaB-binding ; InterPro: IPR013173 Eubacterial DnaG primases interact with several factors to form the replisome. One of these factors is DnaB, a helicase. This domain has been demonstrated to be responsible for the interaction between DnaG and DnaB []. This domain has a multi-helical structure that forms an orthogonal bundle [].; GO: 0003896 DNA primase activity, 0006269 DNA replication, synthesis of RNA primer; PDB: 2HAJ_A 1T3W_B.
Probab=20.57 E-value=1.8e+02 Score=23.81 Aligned_cols=48 Identities=25% Similarity=0.329 Sum_probs=32.9
Q ss_pred chhHHHHHHHHHHHHhhhcCccccChHHHhhhHhhhhhcCCeeEEeChhhHHHHHHH
Q 019981 94 SLGELEQEFLQALQAFYYEGKAVMSNEEFDNLKEELMWEGSSVVMLSSAEQKFLEAS 150 (333)
Q Consensus 94 slge~E~~fl~Al~~fY~~gk~~~sdeefd~LkEeL~weGSsvv~L~~~Eq~fLEA~ 150 (333)
+-.+.|++|.+++.+..... -+.+++.|+....=+| |+..|++.|-.+
T Consensus 79 ~~~~~~~ef~d~l~~L~~~~----~~~~i~~L~~k~~~~~-----Lt~eEk~el~~L 126 (127)
T PF08278_consen 79 DEEDIEQEFQDALARLQEQA----LERRIEELKAKPRRGG-----LTDEEKQELRRL 126 (127)
T ss_dssp HHHHHHHHHHHHHHHHHHHH----HHHHHHHHHHHHTTT--------HHHHHHHHHH
T ss_pred CchhHHHHHHHHHHHHHHHH----HHHHHHHHHHhhccCC-----cCHHHHHHHHHh
Confidence 66789999999999988765 5678888887744322 677776665443
No 202
>PF09297 zf-NADH-PPase: NADH pyrophosphatase zinc ribbon domain; InterPro: IPR015376 This domain has a zinc ribbon structure and is often found between two NUDIX domains.; GO: 0016787 hydrolase activity, 0046872 metal ion binding; PDB: 1VK6_A 2GB5_A.
Probab=20.45 E-value=34 Score=22.64 Aligned_cols=10 Identities=50% Similarity=1.351 Sum_probs=4.7
Q ss_pred CCCCCccccc
Q 019981 274 PCPNCGTENV 283 (333)
Q Consensus 274 ~CPNCGeEv~ 283 (333)
-||+||.+.|
T Consensus 23 ~C~~Cg~~~y 32 (32)
T PF09297_consen 23 RCPSCGHEHY 32 (32)
T ss_dssp EESSSS-EE-
T ss_pred ECCCCcCEeC
Confidence 3666666543
No 203
>PF12207 DUF3600: Domain of unknown function (DUF3600); InterPro: IPR022019 This family of proteins is found in bacteria. Proteins in this family are approximately 230 amino acids in length. This domain is the C-terminal of the putative ecf-type sigma factor negative effector. ; PDB: 3FGG_A 3FH3_A.
Probab=20.45 E-value=83 Score=28.99 Aligned_cols=65 Identities=29% Similarity=0.406 Sum_probs=42.0
Q ss_pred cCccccChHHHhhhHhhhhh-----------cCCe-eEEeChhhHHHHHHH----Hhhhc-------CCCccChHHHHHH
Q 019981 112 EGKAVMSNEEFDNLKEELMW-----------EGSS-VVMLSSAEQKFLEAS----MAYVA-------GKPIMSDEEYDKL 168 (333)
Q Consensus 112 ~gk~~~sdeefd~LkEeL~w-----------eGSs-vv~L~~~Eq~fLEA~----~AY~s-------GkPimsDeeFD~L 168 (333)
.-|..|+.+||..-++.|-= +|-- -=-|++.||+....+ +-||+ -|-|++++|||.-
T Consensus 35 qAK~~lgeeEfeef~~lLK~lt~~kLkygD~NGnidye~ls~~eqee~k~~~~eLqPYFdKLN~~~SsK~vlt~~E~d~y 114 (162)
T PF12207_consen 35 QAKGELGEEEFEEFKELLKKLTNAKLKYGDKNGNIDYEKLSKEEQEEYKKLTMELQPYFDKLNGHKSSKEVLTQEEYDQY 114 (162)
T ss_dssp HHHHCS-HHHHHHHHHHHHHHHHHHHHHB-TTS-B-GGGS-HHHHHHHHHHHHHHHHHHHHHTT---HHHHS-HHHHHHH
T ss_pred HHHHhhhHHHHHHHHHHHHHHHHhHHhhcccCCCcCHHhCCHHHHHHHHHHHHhcchHHHHhcCCcchhhhcCHHHHHHH
Confidence 45789999999888877632 3322 224677888776664 33764 4668999999998
Q ss_pred HHHHhhcC
Q 019981 169 KQKLKMEG 176 (333)
Q Consensus 169 K~kLk~~G 176 (333)
++-|+.+-
T Consensus 115 ~eALm~~e 122 (162)
T PF12207_consen 115 IEALMTYE 122 (162)
T ss_dssp HHHHHHHH
T ss_pred HHHHhhhh
Confidence 88886653
No 204
>PF09930 DUF2162: Predicted transporter (DUF2162); InterPro: IPR017199 This group represents a predicted membrane transporter, MTH672 type.
Probab=20.43 E-value=79 Score=30.11 Aligned_cols=35 Identities=34% Similarity=0.371 Sum_probs=23.6
Q ss_pred HhhhhHHHHHHHHHHhhh--------hcceeeeecCCCCCcccc
Q 019981 247 FAAVPLIVYLSQSLTKLI--------VRESLILKGPCPNCGTEN 282 (333)
Q Consensus 247 ~~~~P~i~~~a~~Lt~l~--------~~D~~iLKG~CPNCGeEv 282 (333)
.+++=++++.-..+.+ | .+..+++--|||.|-+-+
T Consensus 73 imal~li~~Gi~ti~~-W~~~~~~~s~~t~lal~~PCPvCl~Ai 115 (224)
T PF09930_consen 73 IMALLLIYAGIYTIKK-WKKSGKDSSRRTFLALSLPCPVCLTAI 115 (224)
T ss_pred HHHHHHHHHHHHHHHH-HcccCCCCcccchhhhhcCchHHHHHH
Confidence 4455556665555544 4 555789999999997654
No 205
>PF04216 FdhE: Protein involved in formate dehydrogenase formation; InterPro: IPR006452 This family of sequences describe an accessory protein required for the assembly of formate dehydrogenase of certain proteobacteria although not present in the final complex []. The exact nature of the function of FdhE in the assembly of the complex is unknown, but considering the presence of selenocysteine, molybdopterin, iron-sulphur clusters and cytochrome b556, it is likely to be involved in the insertion of cofactors. ; GO: 0005737 cytoplasm; PDB: 2FIY_B.
Probab=20.40 E-value=58 Score=31.06 Aligned_cols=15 Identities=27% Similarity=0.822 Sum_probs=6.3
Q ss_pred eeecCCCCCccccce
Q 019981 270 ILKGPCPNCGTENVS 284 (333)
Q Consensus 270 iLKG~CPNCGeEv~a 284 (333)
..-.-||+||++.-.
T Consensus 209 ~~R~~Cp~Cg~~~~~ 223 (290)
T PF04216_consen 209 FVRIKCPYCGNTDHE 223 (290)
T ss_dssp --TTS-TTT---SS-
T ss_pred ecCCCCcCCCCCCCc
Confidence 345679999998764
No 206
>PLN02278 succinic semialdehyde dehydrogenase
Probab=20.18 E-value=2.9e+02 Score=28.42 Aligned_cols=68 Identities=19% Similarity=0.434 Sum_probs=46.5
Q ss_pred ccChHHHhhhHhhhhh-----cCCe------eEEeChhhHHHHHHHHhh----hcCC---------CccChHHHHHHHHH
Q 019981 116 VMSNEEFDNLKEELMW-----EGSS------VVMLSSAEQKFLEASMAY----VAGK---------PIMSDEEYDKLKQK 171 (333)
Q Consensus 116 ~~sdeefd~LkEeL~w-----eGSs------vv~L~~~Eq~fLEA~~AY----~sGk---------PimsDeeFD~LK~k 171 (333)
++.|.+.|.--+.+.| .|-. +++-...-.+|+|++.+. .-|. |+++...+|+++.-
T Consensus 278 V~~dAdl~~aa~~i~~~~f~~~GQ~C~a~~rv~V~~~i~~~f~~~L~~~~~~l~~G~p~~~~~~~Gpli~~~~~~~v~~~ 357 (498)
T PLN02278 278 VFDDADLDVAVKGALASKFRNSGQTCVCANRILVQEGIYDKFAEAFSKAVQKLVVGDGFEEGVTQGPLINEAAVQKVESH 357 (498)
T ss_pred ECCCCCHHHHHHHHHHHHhccCCCCCcCCcEEEEeHHHHHHHHHHHHHHHHhcCCCCCCCCCCcCCCccCHHHHHHHHHH
Confidence 5568888887788777 3433 333333457899987653 3343 68999999999987
Q ss_pred Hh---hcCCceeeec
Q 019981 172 LK---MEGSEIVVEG 183 (333)
Q Consensus 172 Lk---~~GS~VVvk~ 183 (333)
+. .+|.+++.-|
T Consensus 358 i~~a~~~Ga~vl~gG 372 (498)
T PLN02278 358 VQDAVSKGAKVLLGG 372 (498)
T ss_pred HHHHHhCCCEEEeCC
Confidence 65 4788887654
No 207
>PF09855 DUF2082: Nucleic-acid-binding protein containing Zn-ribbon domain (DUF2082); InterPro: IPR018652 This family of proteins contains various hypothetical prokaryotic proteins as well as some Zn-ribbon nucleic-acid-binding proteins.
Probab=20.17 E-value=1.1e+02 Score=24.06 Aligned_cols=42 Identities=33% Similarity=0.781 Sum_probs=29.1
Q ss_pred CCCCCccccce---------eccccccccCCCCccceeCCCCccccccCceeEEeccc
Q 019981 274 PCPNCGTENVS---------FFGTILSISSGGTTNTINCSNLTFCFSCGTTMVYDSNT 322 (333)
Q Consensus 274 ~CPNCGeEv~a---------Ffgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f~tk~ 322 (333)
-||-||.+.+. .|+.++.|+.++ --.+-|.+ ||=.=.|++++
T Consensus 2 ~C~KCg~~~~e~~~v~~tgg~~skiFdvq~~~-f~~v~C~~------CGYTE~Y~~~~ 52 (64)
T PF09855_consen 2 KCPKCGNEEYESGEVRATGGGLSKIFDVQNKK-FTTVSCTN------CGYTEFYKAKT 52 (64)
T ss_pred CCCCCCCcceecceEEccCCeeEEEEEecCcE-EEEEECCC------CCCEEEEeecC
Confidence 49999987653 355566665553 34677888 99987777654
No 208
>TIGR02159 PA_CoA_Oxy4 phenylacetate-CoA oxygenase, PaaJ subunit. Phenylacetate-CoA oxygenase is comprised of a five gene complex responsible for the hydroxylation of phenylacetate-CoA (PA-CoA) as the second catabolic step in phenylacetic acid (PA) degradation. Although the exact function of this enzyme has not been determined, it has been shown to be required for phenylacetic acid degradation and has been proposed to function in a multicomponent oxygenase acting on phenylacetate-CoA.
Probab=20.07 E-value=34 Score=30.36 Aligned_cols=38 Identities=32% Similarity=0.699 Sum_probs=22.2
Q ss_pred ecCCCCCccccceeccccccccCCCCccceeCCCCccccccCceeEE
Q 019981 272 KGPCPNCGTENVSFFGTILSISSGGTTNTINCSNLTFCFSCGTTMVY 318 (333)
Q Consensus 272 KG~CPNCGeEv~aFfgtilsv~s~~~~n~vkC~~~aeCHVC~t~L~f 318 (333)
.-+||.||..+..-. +-.+++ -|...--|.-|..+.+|
T Consensus 105 ~~~cp~c~s~~t~~~------s~fg~t---~cka~~~c~~c~epf~~ 142 (146)
T TIGR02159 105 SVQCPRCGSADTTIT------SIFGPT---ACKALYRCRACKEPFEY 142 (146)
T ss_pred CCcCCCCCCCCcEee------cCCCCh---hhHHHhhhhhhCCcHhh
Confidence 368999998876443 223333 24434455558876654
Done!