Query 017471
Match_columns 371
No_of_seqs 416 out of 3016
Neff 7.8
Searched_HMMs 46136
Date Fri Mar 29 08:51:48 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/017471.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/017471hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PRK10139 serine endoprotease; 100.0 2.5E-48 5.4E-53 390.1 29.1 309 5-342 82-414 (455)
2 TIGR02037 degP_htrA_DO peripla 100.0 3.1E-47 6.7E-52 381.6 29.6 308 10-342 55-386 (428)
3 PRK10942 serine endoprotease; 100.0 8.1E-46 1.7E-50 373.6 27.7 300 10-342 108-432 (473)
4 TIGR02038 protease_degS peripl 100.0 7.2E-44 1.6E-48 347.9 26.1 248 12-278 77-349 (351)
5 PRK10898 serine endoprotease; 100.0 2.6E-43 5.6E-48 343.8 27.0 250 11-279 76-351 (353)
6 COG0265 DegQ Trypsin-like seri 100.0 2.3E-34 4.9E-39 281.2 23.0 234 25-277 103-340 (347)
7 KOG1320 Serine protease [Postt 99.9 2.9E-26 6.3E-31 225.8 11.0 320 3-340 77-420 (473)
8 KOG1421 Predicted signaling-as 99.9 4.7E-24 1E-28 212.4 16.5 303 4-333 75-418 (955)
9 KOG1320 Serine protease [Postt 99.9 1.3E-21 2.7E-26 193.1 17.5 236 28-277 211-468 (473)
10 KOG1421 Predicted signaling-as 99.8 3E-18 6.6E-23 171.3 20.7 271 27-333 585-877 (955)
11 PF13180 PDZ_2: PDZ domain; PD 99.6 3E-15 6.4E-20 115.8 7.8 81 173-275 1-82 (82)
12 TIGR03279 cyano_FeS_chp putati 99.5 2.4E-15 5.2E-20 147.7 1.4 102 202-363 2-107 (433)
13 cd00987 PDZ_serine_protease PD 99.4 6.8E-13 1.5E-17 103.8 9.0 88 173-272 1-89 (90)
14 cd00986 PDZ_LON_protease PDZ d 99.3 9.9E-12 2.1E-16 95.2 8.4 72 197-278 7-78 (79)
15 cd00991 PDZ_archaeal_metallopr 99.3 1.1E-11 2.4E-16 95.1 8.0 69 196-274 8-77 (79)
16 cd00990 PDZ_glycyl_aminopeptid 99.3 2.3E-11 5.1E-16 93.1 8.3 77 173-276 1-78 (80)
17 TIGR01713 typeII_sec_gspC gene 99.2 5.4E-11 1.2E-15 111.4 10.4 101 153-275 158-259 (259)
18 TIGR02037 degP_htrA_DO peripla 99.1 1.8E-10 4E-15 115.8 9.1 90 172-272 337-427 (428)
19 PRK10779 zinc metallopeptidase 99.1 9.9E-11 2.2E-15 118.4 6.9 117 200-342 128-245 (449)
20 cd00989 PDZ_metalloprotease PD 99.1 2.8E-10 6.1E-15 86.7 6.6 66 198-274 12-78 (79)
21 cd00988 PDZ_CTP_protease PDZ d 99.0 9.3E-10 2E-14 85.1 8.5 68 197-275 12-83 (85)
22 PF13365 Trypsin_2: Trypsin-li 98.8 1.4E-08 2.9E-13 83.0 8.3 87 18-134 31-120 (120)
23 cd00136 PDZ PDZ domain, also c 98.8 8E-09 1.7E-13 76.8 6.1 65 174-262 2-69 (70)
24 TIGR00054 RIP metalloprotease 98.7 2.3E-08 5E-13 100.4 7.0 69 198-277 203-272 (420)
25 smart00228 PDZ Domain present 98.7 6.4E-08 1.4E-12 74.2 7.3 73 173-266 12-85 (85)
26 PRK10779 zinc metallopeptidase 98.6 5.4E-08 1.2E-12 98.6 7.2 68 199-277 222-290 (449)
27 TIGR00225 prc C-terminal pepti 98.6 1.2E-07 2.5E-12 92.5 8.0 71 198-279 62-135 (334)
28 TIGR00054 RIP metalloprotease 98.6 6.1E-08 1.3E-12 97.3 5.4 66 197-274 127-193 (420)
29 PF00595 PDZ: PDZ domain (Also 98.5 2.3E-07 4.9E-12 71.2 5.7 72 172-263 9-81 (81)
30 PRK10139 serine endoprotease; 98.5 3E-07 6.5E-12 93.1 7.3 64 198-273 390-454 (455)
31 PLN00049 carboxyl-terminal pro 98.5 5.9E-07 1.3E-11 89.3 9.3 69 198-275 102-171 (389)
32 TIGR02860 spore_IV_B stage IV 98.4 5.1E-07 1.1E-11 88.9 6.4 69 197-276 104-181 (402)
33 PRK10942 serine endoprotease; 98.4 6.3E-07 1.4E-11 91.2 6.6 64 198-273 408-472 (473)
34 cd00992 PDZ_signaling PDZ doma 98.4 8.7E-07 1.9E-11 67.7 5.7 52 173-235 12-66 (82)
35 PF14685 Tricorn_PDZ: Tricorn 98.3 4.3E-06 9.3E-11 65.2 9.4 65 197-272 11-87 (88)
36 COG0793 Prc Periplasmic protea 98.3 1.5E-06 3.2E-11 86.7 8.2 83 172-278 99-184 (406)
37 COG3480 SdrC Predicted secrete 98.2 4E-06 8.6E-11 78.9 7.4 72 197-278 129-201 (342)
38 PRK09681 putative type II secr 98.0 9.4E-06 2E-10 76.2 6.8 68 198-275 204-275 (276)
39 PF04495 GRASP55_65: GRASP55/6 98.0 1.1E-05 2.4E-10 68.3 5.0 86 172-276 25-114 (138)
40 PF00089 Trypsin: Trypsin; In 97.9 6.4E-05 1.4E-09 67.4 9.5 120 40-159 86-219 (220)
41 PRK11186 carboxy-terminal prot 97.8 5.7E-05 1.2E-09 79.5 8.5 71 198-274 255-332 (667)
42 COG3975 Predicted protease wit 97.8 4.3E-05 9.4E-10 76.5 7.0 86 175-280 439-527 (558)
43 KOG3129 26S proteasome regulat 97.7 7.7E-05 1.7E-09 66.3 6.5 73 199-279 140-213 (231)
44 KOG3553 Tax interaction protei 97.4 0.00017 3.6E-09 56.6 3.4 35 197-231 58-93 (124)
45 PF12812 PDZ_1: PDZ-like domai 97.3 0.00033 7.2E-09 53.4 4.5 60 173-236 9-69 (78)
46 COG3031 PulC Type II secretory 97.3 0.00019 4.1E-09 65.1 3.4 66 199-274 208-274 (275)
47 COG3591 V8-like Glu-specific e 97.0 0.0036 7.7E-08 58.1 8.5 87 67-163 159-249 (251)
48 PF00863 Peptidase_C4: Peptida 96.9 0.015 3.2E-07 53.6 11.6 112 40-163 81-196 (235)
49 cd00190 Tryp_SPc Trypsin-like 96.3 0.03 6.6E-07 50.2 9.9 100 40-139 88-208 (232)
50 PF10459 Peptidase_S46: Peptid 96.2 0.0035 7.7E-08 66.5 3.7 56 108-163 623-686 (698)
51 KOG3209 WW domain-containing p 96.0 0.0066 1.4E-07 62.9 4.3 55 202-266 782-838 (984)
52 KOG3580 Tight junction protein 96.0 0.0043 9.2E-08 63.0 2.7 59 198-264 429-488 (1027)
53 smart00020 Tryp_SPc Trypsin-li 95.9 0.042 9.1E-07 49.5 8.5 100 40-139 88-208 (229)
54 KOG3550 Receptor targeting pro 95.5 0.028 6.1E-07 47.6 5.3 37 197-233 114-152 (207)
55 KOG3532 Predicted protein kina 95.4 0.02 4.4E-07 59.2 4.9 50 174-236 387-437 (1051)
56 PF08192 Peptidase_S64: Peptid 95.3 0.17 3.6E-06 52.8 11.2 119 38-163 540-688 (695)
57 PF00949 Peptidase_S7: Peptida 95.1 0.025 5.5E-07 47.4 3.9 33 108-140 87-119 (132)
58 PF00947 Pico_P2A: Picornaviru 94.9 0.16 3.5E-06 42.0 7.8 97 34-138 13-109 (127)
59 KOG3834 Golgi reassembly stack 94.8 0.05 1.1E-06 53.6 5.6 132 197-362 14-151 (462)
60 KOG3209 WW domain-containing p 94.7 0.044 9.4E-07 57.0 5.0 58 198-266 923-982 (984)
61 PF00548 Peptidase_C3: 3C cyst 94.3 0.67 1.4E-05 40.8 11.0 103 33-138 61-170 (172)
62 KOG3605 Beta amyloid precursor 94.0 0.078 1.7E-06 54.7 5.1 102 118-231 680-790 (829)
63 KOG3542 cAMP-regulated guanine 93.9 0.05 1.1E-06 56.3 3.5 37 197-233 561-598 (1283)
64 KOG3651 Protein kinase C, alph 93.8 0.089 1.9E-06 49.6 4.6 39 198-236 30-70 (429)
65 KOG1892 Actin filament-binding 93.8 0.082 1.8E-06 56.7 4.8 64 194-267 956-1021(1629)
66 COG0750 Predicted membrane-ass 93.7 0.11 2.3E-06 51.3 5.5 57 202-269 133-194 (375)
67 KOG2921 Intramembrane metallop 93.6 0.051 1.1E-06 53.0 2.7 40 196-235 218-259 (484)
68 KOG3580 Tight junction protein 93.4 0.092 2E-06 53.7 4.4 61 197-268 39-100 (1027)
69 KOG3552 FERM domain protein FR 92.5 0.12 2.7E-06 55.2 3.9 57 198-264 75-131 (1298)
70 PF00944 Peptidase_S3: Alphavi 91.4 0.19 4.2E-06 41.9 3.1 26 114-139 102-127 (158)
71 KOG3606 Cell polarity protein 91.1 0.25 5.5E-06 45.9 3.8 81 147-232 146-230 (358)
72 KOG3571 Dishevelled 3 and rela 90.9 0.27 5.8E-06 49.5 4.1 37 197-233 276-314 (626)
73 KOG3551 Syntrophins (type beta 88.9 0.33 7.1E-06 47.4 2.8 72 173-266 96-172 (506)
74 KOG3549 Syntrophins (type gamm 88.3 0.5 1.1E-05 45.5 3.6 55 199-263 81-137 (505)
75 PF05579 Peptidase_S32: Equine 86.4 0.51 1.1E-05 44.0 2.5 23 116-138 206-228 (297)
76 KOG0609 Calcium/calmodulin-dep 86.0 1.3 2.9E-05 45.1 5.4 57 199-265 147-205 (542)
77 PF02907 Peptidase_S29: Hepati 84.3 0.83 1.8E-05 38.2 2.5 25 116-140 106-130 (148)
78 cd00987 PDZ_serine_protease PD 83.7 1.3 2.9E-05 33.6 3.4 47 296-343 2-49 (90)
79 KOG3605 Beta amyloid precursor 80.5 1.7 3.6E-05 45.4 3.5 61 203-271 678-740 (829)
80 PF01732 DUF31: Putative pepti 76.6 1.8 4E-05 42.8 2.5 26 112-137 349-374 (374)
81 KOG0606 Microtubule-associated 76.5 2.9 6.3E-05 46.3 4.1 34 200-233 660-694 (1205)
82 PF05580 Peptidase_S55: SpoIVB 76.1 2.4 5.1E-05 38.5 2.8 39 114-155 176-214 (218)
83 PF03761 DUF316: Domain of unk 73.4 31 0.00068 32.3 10.0 89 40-138 160-254 (282)
84 COG1625 Fe-S oxidoreductase, r 72.3 2.8 6E-05 41.7 2.5 35 201-235 4-40 (414)
85 KOG3834 Golgi reassembly stack 68.3 3.9 8.5E-05 40.7 2.5 65 201-276 112-178 (462)
86 KOG3938 RGS-GAIP interacting p 64.2 5.5 0.00012 37.3 2.5 67 190-264 138-209 (334)
87 PF03510 Peptidase_C24: 2C end 64.0 38 0.00083 27.2 7.0 52 11-68 3-59 (105)
88 PF12812 PDZ_1: PDZ-like domai 63.6 15 0.00031 27.9 4.4 39 291-332 5-44 (78)
89 PF13180 PDZ_2: PDZ domain; PD 60.3 6.5 0.00014 29.5 2.0 37 296-342 2-38 (82)
90 cd01735 LSm12_N LSm12 belongs 55.1 38 0.00082 24.5 5.0 34 17-50 6-39 (61)
91 TIGR02038 protease_degS peripl 55.0 9.9 0.00021 37.3 2.8 48 296-344 256-304 (351)
92 cd01726 LSm6 The eukaryotic Sm 51.7 29 0.00062 25.3 4.1 33 17-49 10-42 (67)
93 PF02122 Peptidase_S39: Peptid 50.0 21 0.00045 32.3 3.7 46 108-154 137-182 (203)
94 TIGR03000 plancto_dom_1 Planct 49.8 54 0.0012 24.7 5.3 50 217-275 10-63 (75)
95 cd00600 Sm_like The eukaryotic 48.2 42 0.0009 23.6 4.5 33 19-51 8-40 (63)
96 cd01722 Sm_F The eukaryotic Sm 47.7 33 0.00072 25.0 3.9 33 17-49 11-43 (68)
97 PRK00737 small nuclear ribonuc 47.2 43 0.00093 24.8 4.5 34 17-50 14-47 (72)
98 cd01717 Sm_B The eukaryotic Sm 47.0 38 0.00082 25.5 4.3 31 20-50 13-43 (79)
99 cd01730 LSm3 The eukaryotic Sm 46.3 35 0.00076 25.9 4.0 29 20-48 14-42 (82)
100 KOG3553 Tax interaction protei 45.0 32 0.0007 27.4 3.6 27 316-342 57-83 (124)
101 TIGR02860 spore_IV_B stage IV 44.7 20 0.00043 35.9 3.0 37 114-153 356-392 (402)
102 PRK10898 serine endoprotease; 43.4 22 0.00049 34.9 3.2 48 296-344 257-305 (353)
103 cd01732 LSm5 The eukaryotic Sm 43.3 45 0.00098 25.1 4.1 31 18-48 14-44 (76)
104 PF11874 DUF3394: Domain of un 43.3 19 0.00041 32.0 2.4 28 197-224 121-149 (183)
105 COG0298 HypC Hydrogenase matur 42.4 47 0.001 25.3 4.0 47 30-78 5-52 (82)
106 cd01729 LSm7 The eukaryotic Sm 42.3 53 0.0012 24.9 4.5 30 20-49 15-44 (81)
107 cd01720 Sm_D2 The eukaryotic S 42.2 50 0.0011 25.6 4.3 31 19-49 16-46 (87)
108 cd01731 archaeal_Sm1 The archa 41.9 53 0.0011 23.9 4.3 33 18-50 11-43 (68)
109 KOG1738 Membrane-associated gu 41.4 21 0.00045 37.4 2.7 34 200-233 227-262 (638)
110 cd06168 LSm9 The eukaryotic Sm 40.9 60 0.0013 24.3 4.5 31 19-49 12-42 (75)
111 PF00571 CBS: CBS domain CBS d 39.5 26 0.00057 23.7 2.3 20 118-137 29-48 (57)
112 COG0260 PepB Leucyl aminopepti 39.0 44 0.00096 34.3 4.6 58 189-251 291-348 (485)
113 smart00651 Sm snRNP Sm protein 37.9 70 0.0015 22.8 4.4 33 18-50 9-41 (67)
114 PF01455 HupF_HypC: HupF/HypC 36.2 59 0.0013 23.9 3.7 41 30-75 5-47 (68)
115 cd01728 LSm1 The eukaryotic Sm 36.1 82 0.0018 23.5 4.5 31 19-49 14-44 (74)
116 cd01721 Sm_D3 The eukaryotic S 34.3 82 0.0018 23.1 4.3 35 16-50 9-43 (70)
117 cd01727 LSm8 The eukaryotic Sm 32.2 92 0.002 23.0 4.3 31 20-50 12-42 (74)
118 KOG3627 Trypsin [Amino acid tr 30.8 33 0.00072 31.3 2.1 99 41-139 106-228 (256)
119 cd01719 Sm_G The eukaryotic Sm 30.5 1.1E+02 0.0024 22.6 4.5 30 20-49 13-42 (72)
120 PRK05015 aminopeptidase B; Pro 29.0 90 0.0019 31.5 4.8 39 190-230 230-268 (424)
121 PF09465 LBR_tudor: Lamin-B re 28.3 2E+02 0.0043 20.3 5.0 37 15-51 7-44 (55)
122 COG1958 LSM1 Small nuclear rib 27.1 1.2E+02 0.0026 22.7 4.2 32 19-50 19-50 (79)
123 PF11874 DUF3394: Domain of un 27.0 1.2E+02 0.0026 27.0 4.7 71 247-336 67-140 (183)
124 PF15483 DUF4641: Domain of un 25.8 45 0.00097 33.3 2.0 21 344-364 417-438 (445)
125 COG2524 Predicted transcriptio 24.5 2.1E+02 0.0045 27.1 6.0 20 117-137 201-220 (294)
126 cd01723 LSm4 The eukaryotic Sm 24.4 1.7E+02 0.0038 21.7 4.6 33 17-49 11-43 (76)
127 PF12381 Peptidase_C3G: Tungro 23.0 96 0.0021 28.3 3.4 54 107-163 169-228 (231)
128 PF01423 LSM: LSM domain ; In 22.5 1.9E+02 0.0041 20.5 4.4 33 19-51 10-42 (67)
129 KOG2561 Adaptor protein NUB1, 22.3 28 0.0006 35.1 -0.2 23 328-350 195-223 (568)
130 PF10049 DUF2283: Protein of u 21.8 62 0.0013 22.1 1.6 11 126-136 36-46 (50)
131 COG5233 GRH1 Peripheral Golgi 21.7 47 0.001 32.1 1.2 30 201-230 66-96 (417)
132 PF14827 Cache_3: Sensory doma 21.7 77 0.0017 25.4 2.4 17 123-139 95-111 (116)
133 TIGR00074 hypC_hupF hydrogenas 21.6 1.6E+02 0.0035 22.1 3.9 41 30-75 5-45 (76)
134 PF15436 PGBA_N: Plasminogen-b 21.5 3.4E+02 0.0074 24.8 6.7 59 14-74 28-88 (218)
135 cd00991 PDZ_archaeal_metallopr 21.4 55 0.0012 24.2 1.4 27 316-342 8-34 (79)
136 PF11730 DUF3297: Protein of u 21.3 58 0.0013 23.9 1.4 59 204-274 5-64 (71)
137 PRK09570 rpoH DNA-directed RNA 20.8 59 0.0013 24.7 1.4 19 204-222 39-59 (79)
138 cd00433 Peptidase_M17 Cytosol 20.8 1.1E+02 0.0023 31.5 3.6 28 203-230 290-317 (468)
139 PRK00913 multifunctional amino 20.7 1.1E+02 0.0024 31.5 3.7 27 204-230 305-331 (483)
140 cd01718 Sm_E The eukaryotic Sm 20.6 1.9E+02 0.0041 22.0 4.1 30 20-49 21-52 (79)
141 cd01725 LSm2 The eukaryotic Sm 20.5 2E+02 0.0043 21.8 4.3 33 17-49 11-43 (81)
142 cd01724 Sm_D1 The eukaryotic S 20.5 1.8E+02 0.004 22.5 4.2 35 16-50 10-44 (90)
No 1
>PRK10139 serine endoprotease; Provisional
Probab=100.00 E-value=2.5e-48 Score=390.08 Aligned_cols=309 Identities=21% Similarity=0.272 Sum_probs=252.1
Q ss_pred eccCCeeeeeccEEEEE---EEecCCCe-----------------EeEEEEEecCCCCEEEEEEecCCCcCCccceecCC
Q 017471 5 WSTTRRLNSRNEALILS---TWLLCSPS-----------------APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGE 64 (371)
Q Consensus 5 ~~~~~~~~~~gsg~vi~---~~~~~~~~-----------------~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~ 64 (371)
|...+...+.||||+++ ++++++.| ++|++++.|+++||||||++... ++++++|++
T Consensus 82 ~~~~~~~~~~GSG~ii~~~~g~IlTn~HVv~~a~~i~V~~~dg~~~~a~vvg~D~~~DlAvlkv~~~~---~l~~~~lg~ 158 (455)
T PRK10139 82 DQPAQPFEGLGSGVIIDAAKGYVLTNNHVINQAQKISIQLNDGREFDAKLIGSDDQSDIALLQIQNPS---KLTQIAIAD 158 (455)
T ss_pred ccccccccceEEEEEEECCCCEEEeChHHhCCCCEEEEEECCCCEEEEEEEEEcCCCCEEEEEecCCC---CCceeEecC
Confidence 44444556789999996 57766665 79999999999999999998653 789999998
Q ss_pred CCC--CCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccC
Q 017471 65 LPA--LQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE 142 (371)
Q Consensus 65 s~~--lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~ 142 (371)
+.. +||+|+++|||++... +++.|+||+.++..... .....+||+|++||+|||||||+|.+||||||+++.+..+
T Consensus 159 s~~~~~G~~V~aiG~P~g~~~-tvt~GivS~~~r~~~~~-~~~~~~iqtda~in~GnSGGpl~n~~G~vIGi~~~~~~~~ 236 (455)
T PRK10139 159 SDKLRVGDFAVAVGNPFGLGQ-TATSGIISALGRSGLNL-EGLENFIQTDASINRGNSGGALLNLNGELIGINTAILAPG 236 (455)
T ss_pred ccccCCCCEEEEEecCCCCCC-ceEEEEEccccccccCC-CCcceEEEECCccCCCCCcceEECCCCeEEEEEEEEEcCC
Confidence 764 6899999999999776 89999999998743222 1234689999999999999999999999999999987543
Q ss_pred -CccceeccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCE
Q 017471 143 -DVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDI 220 (371)
Q Consensus 143 -~~~~~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDv 220 (371)
+..+++||||++.+++++++|.++|++. ++|||+.++++ +++.++.+|++ ...|++|.+|.++|||++ |||+||+
T Consensus 237 ~~~~gigfaIP~~~~~~v~~~l~~~g~v~-r~~LGv~~~~l-~~~~~~~lgl~-~~~Gv~V~~V~~~SpA~~AGL~~GDv 313 (455)
T PRK10139 237 GGSVGIGFAIPSNMARTLAQQLIDFGEIK-RGLLGIKGTEM-SADIAKAFNLD-VQRGAFVSEVLPNSGSAKAGVKAGDI 313 (455)
T ss_pred CCccceEEEEEhHHHHHHHHHHhhcCccc-ccceeEEEEEC-CHHHHHhcCCC-CCCceEEEEECCCChHHHCCCCCCCE
Confidence 4578999999999999999999999999 99999999999 89999999997 467999999999999999 9999999
Q ss_pred EEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccccccCCCCCCCCCCCceeeccEE
Q 017471 221 ILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFV 300 (371)
Q Consensus 221 Il~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~~~~~~~~~~~~~~~Gl~ 300 (371)
|++|||++|.++.++. ..+....+|++++++|.|+|+.+++++++...+......... .+ .+.|+.
T Consensus 314 Il~InG~~V~s~~dl~----------~~l~~~~~g~~v~l~V~R~G~~~~l~v~~~~~~~~~~~~~~~-~~---~~~g~~ 379 (455)
T PRK10139 314 ITSLNGKPLNSFAELR----------SRIATTEPGTKVKLGLLRNGKPLEVEVTLDTSTSSSASAEMI-TP---ALQGAT 379 (455)
T ss_pred EEEECCEECCCHHHHH----------HHHHhcCCCCEEEEEEEECCEEEEEEEEECCCCCcccccccc-cc---cccccE
Confidence 9999999999999875 567666789999999999999999999985433221111100 01 134655
Q ss_pred EechHHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhH
Q 017471 301 FSRCLYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLL 342 (371)
Q Consensus 301 ~~~l~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~ 342 (371)
+++. + ......|+++..+.++|||+.--.+.+|+|
T Consensus 380 l~~~-----~--~~~~~~Gv~V~~V~~~spA~~aGL~~GD~I 414 (455)
T PRK10139 380 LSDG-----Q--LKDGTKGIKIDEVVKGSPAAQAGLQKDDVI 414 (455)
T ss_pred eccc-----c--cccCCCceEEEEeCCCChHHHcCCCCCCEE
Confidence 5431 1 122346899999999999998888888775
No 2
>TIGR02037 degP_htrA_DO periplasmic serine protease, Do/DeqQ family. This family consists of a set proteins various designated DegP, heat shock protein HtrA, and protease DO. The ortholog in Pseudomonas aeruginosa is designated MucD and is found in an operon that controls mucoid phenotype. This family also includes the DegQ (HhoA) paralog in E. coli which can rescue a DegP mutant, but not the smaller DegS paralog, which cannot. Members of this family are located in the periplasm and have separable functions as both protease and chaperone. Members have a trypsin domain and two copies of a PDZ domain. This protein protects bacteria from thermal and other stresses and may be important for the survival of bacterial pathogens.// The chaperone function is dominant at low temperatures, whereas the proteolytic activity is turned on at elevated temperatures.
Probab=100.00 E-value=3.1e-47 Score=381.60 Aligned_cols=308 Identities=22% Similarity=0.307 Sum_probs=263.1
Q ss_pred eeeeeccEEEEE--EEecCCCe-----------------EeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCC--CC
Q 017471 10 RLNSRNEALILS--TWLLCSPS-----------------APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELP--AL 68 (371)
Q Consensus 10 ~~~~~gsg~vi~--~~~~~~~~-----------------~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~--~l 68 (371)
...+.||||+++ ++++++.| ++|++++.|+.+|||+||++... ++++++++++. .+
T Consensus 55 ~~~~~GSGfii~~~G~IlTn~Hvv~~~~~i~V~~~~~~~~~a~vv~~d~~~DlAllkv~~~~---~~~~~~l~~~~~~~~ 131 (428)
T TIGR02037 55 KVRGLGSGVIISADGYILTNNHVVDGADEITVTLSDGREFKAKLVGKDPRTDIAVLKIDAKK---NLPVIKLGDSDKLRV 131 (428)
T ss_pred cccceeeEEEECCCCEEEEcHHHcCCCCeEEEEeCCCCEEEEEEEEecCCCCEEEEEecCCC---CceEEEccCCCCCCC
Confidence 456789999997 56665555 79999999999999999999763 68999999765 57
Q ss_pred CCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccC-Cccce
Q 017471 69 QDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-DVENI 147 (371)
Q Consensus 69 gd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~-~~~~~ 147 (371)
||+|+++|||++... +++.|+||+..+.. .....+..++|+|+++++|||||||+|.+|+||||+++.+... +..++
T Consensus 132 G~~v~aiG~p~g~~~-~~t~G~vs~~~~~~-~~~~~~~~~i~tda~i~~GnSGGpl~n~~G~viGI~~~~~~~~g~~~g~ 209 (428)
T TIGR02037 132 GDWVLAIGNPFGLGQ-TVTSGIVSALGRSG-LGIGDYENFIQTDAAINPGNSGGPLVNLRGEVIGINTAIYSPSGGNVGI 209 (428)
T ss_pred CCEEEEEECCCcCCC-cEEEEEEEecccCc-cCCCCccceEEECCCCCCCCCCCceECCCCeEEEEEeEEEcCCCCccce
Confidence 899999999999766 89999999988642 1222334589999999999999999999999999999877543 45689
Q ss_pred eccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECC
Q 017471 148 GYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDG 226 (371)
Q Consensus 148 ~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG 226 (371)
+||||++.+++++++|+++|++. ++|||+.++++ +++.++.+|++. ..|++|.+|.++|||++ ||++||+|++|||
T Consensus 210 ~faiP~~~~~~~~~~l~~~g~~~-~~~lGi~~~~~-~~~~~~~lgl~~-~~Gv~V~~V~~~spA~~aGL~~GDvI~~Vng 286 (428)
T TIGR02037 210 GFAIPSNMAKNVVDQLIEGGKVQ-RGWLGVTIQEV-TSDLAKSLGLEK-QRGALVAQVLPGSPAEKAGLKAGDVILSVNG 286 (428)
T ss_pred EEEEEhHHHHHHHHHHHhcCcCc-CCcCceEeecC-CHHHHHHcCCCC-CCceEEEEccCCCChHHcCCCCCCEEEEECC
Confidence 99999999999999999999998 99999999999 899999999984 57999999999999999 9999999999999
Q ss_pred EEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccccccCCCCCCCCCCCceeeccEEEech-H
Q 017471 227 IDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRC-L 305 (371)
Q Consensus 227 ~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~~~~~~~~~~~~~~~Gl~~~~l-~ 305 (371)
++|.++.++. ..+....+|+++++++.|+|+.+++++++...+...+ ++...++|+.++++ +
T Consensus 287 ~~i~~~~~~~----------~~l~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~-------~~~~~~lGi~~~~l~~ 349 (428)
T TIGR02037 287 KPISSFADLR----------RAIGTLKPGKKVTLGILRKGKEKTITVTLGASPEEQA-------SSSNPFLGLTVANLSP 349 (428)
T ss_pred EEcCCHHHHH----------HHHHhcCCCCEEEEEEEECCEEEEEEEEECcCCCccc-------cccccccceEEecCCH
Confidence 9999998865 6676777899999999999999999999876543211 12334799999998 7
Q ss_pred HHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhH
Q 017471 306 YLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLL 342 (371)
Q Consensus 306 ~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~ 342 (371)
..++.++++....|++++.+.++|||+..-...+|+|
T Consensus 350 ~~~~~~~l~~~~~Gv~V~~V~~~SpA~~aGL~~GDvI 386 (428)
T TIGR02037 350 EIRKELRLKGDVKGVVVTKVVSGSPAARAGLQPGDVI 386 (428)
T ss_pred HHHHHcCCCcCcCceEEEEeCCCCHHHHcCCCCCCEE
Confidence 7777888876668999999999999998877777765
No 3
>PRK10942 serine endoprotease; Provisional
Probab=100.00 E-value=8.1e-46 Score=373.56 Aligned_cols=300 Identities=20% Similarity=0.267 Sum_probs=247.5
Q ss_pred eeeeeccEEEEE---EEecCCCe-----------------EeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCC--C
Q 017471 10 RLNSRNEALILS---TWLLCSPS-----------------APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELP--A 67 (371)
Q Consensus 10 ~~~~~gsg~vi~---~~~~~~~~-----------------~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~--~ 67 (371)
...+.||||+++ ++++++.| ++|++++.|+.+||||||++... ++++++++++. +
T Consensus 108 ~~~~~GSG~ii~~~~G~IlTn~HVv~~a~~i~V~~~dg~~~~a~vv~~D~~~DlAvlki~~~~---~l~~~~lg~s~~l~ 184 (473)
T PRK10942 108 KFMALGSGVIIDADKGYVVTNNHVVDNATKIKVQLSDGRKFDAKVVGKDPRSDIALIQLQNPK---NLTAIKMADSDALR 184 (473)
T ss_pred cccceEEEEEEECCCCEEEeChhhcCCCCEEEEEECCCCEEEEEEEEecCCCCEEEEEecCCC---CCceeEecCccccC
Confidence 346789999997 46665555 79999999999999999998643 78999999876 4
Q ss_pred CCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccC-Cccc
Q 017471 68 LQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-DVEN 146 (371)
Q Consensus 68 lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~-~~~~ 146 (371)
+||+|+++|||++... +++.|+||+..+..... ..+..+||+|+++++|||||||+|.+||||||+++.+..+ +..+
T Consensus 185 ~G~~V~aiG~P~g~~~-tvt~GiVs~~~r~~~~~-~~~~~~iqtda~i~~GnSGGpL~n~~GeviGI~t~~~~~~g~~~g 262 (473)
T PRK10942 185 VGDYTVAIGNPYGLGE-TVTSGIVSALGRSGLNV-ENYENFIQTDAAINRGNSGGALVNLNGELIGINTAILAPDGGNIG 262 (473)
T ss_pred CCCEEEEEcCCCCCCc-ceeEEEEEEeecccCCc-ccccceEEeccccCCCCCcCccCCCCCeEEEEEEEEEcCCCCccc
Confidence 6899999999999766 89999999998642221 1234689999999999999999999999999999887654 3468
Q ss_pred eeccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEEC
Q 017471 147 IGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFD 225 (371)
Q Consensus 147 ~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vn 225 (371)
++||||++.+++++++|.+.|++. |+|+|+.++++ ++++++.++++ ...|++|.+|.++|||++ |||+||+|++||
T Consensus 263 ~gfaIP~~~~~~v~~~l~~~g~v~-rg~lGv~~~~l-~~~~a~~~~l~-~~~GvlV~~V~~~SpA~~AGL~~GDvIl~In 339 (473)
T PRK10942 263 IGFAIPSNMVKNLTSQMVEYGQVK-RGELGIMGTEL-NSELAKAMKVD-AQRGAFVSQVLPNSSAAKAGIKAGDVITSLN 339 (473)
T ss_pred EEEEEEHHHHHHHHHHHHhccccc-cceeeeEeeec-CHHHHHhcCCC-CCCceEEEEECCCChHHHcCCCCCCEEEEEC
Confidence 999999999999999999999999 99999999999 88999999998 467999999999999999 999999999999
Q ss_pred CEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccccccCCCCCCCCCCCceeeccEEEech-
Q 017471 226 GIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRC- 304 (371)
Q Consensus 226 G~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~~~~~~~~~~~~~~~Gl~~~~l- 304 (371)
|++|.++.++. ..+....+|++++++|.|+|+.+++++++...+..... +...++|+...++
T Consensus 340 G~~V~s~~dl~----------~~l~~~~~g~~v~l~v~R~G~~~~v~v~l~~~~~~~~~-------~~~~~lGl~g~~l~ 402 (473)
T PRK10942 340 GKPISSFAALR----------AQVGTMPVGSKLTLGLLRDGKPVNVNVELQQSSQNQVD-------SSNIFNGIEGAELS 402 (473)
T ss_pred CEECCCHHHHH----------HHHHhcCCCCEEEEEEEECCeEEEEEEEeCcCcccccc-------cccccccceeeecc
Confidence 99999999875 67777788999999999999999999998664221111 1112356544333
Q ss_pred HHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhH
Q 017471 305 LYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLL 342 (371)
Q Consensus 305 ~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~ 342 (371)
+. ....++++..+.++|||+..-...+|+|
T Consensus 403 ~~--------~~~~gvvV~~V~~~S~A~~aGL~~GDvI 432 (473)
T PRK10942 403 NK--------GGDKGVVVDNVKPGTPAAQIGLKKGDVI 432 (473)
T ss_pred cc--------cCCCCeEEEEeCCCChHHHcCCCCCCEE
Confidence 11 1125899999999999998777777765
No 4
>TIGR02038 protease_degS periplasmic serine pepetdase DegS. This family consists of the periplasmic serine protease DegS (HhoB), a shorter paralog of protease DO (HtrA, DegP) and DegQ (HhoA). It is found in E. coli and several other Proteobacteria of the gamma subdivision. It contains a trypsin domain and a single copy of PDZ domain (in contrast to DegP with two copies). A critical role of this DegS is to sense stress in the periplasm and partially degrade an inhibitor of sigma(E).
Probab=100.00 E-value=7.2e-44 Score=347.85 Aligned_cols=248 Identities=24% Similarity=0.380 Sum_probs=215.3
Q ss_pred eeeccEEEEE--EEecCCCe-----------------EeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCC--CCCC
Q 017471 12 NSRNEALILS--TWLLCSPS-----------------APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELP--ALQD 70 (371)
Q Consensus 12 ~~~gsg~vi~--~~~~~~~~-----------------~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~--~lgd 70 (371)
.+.||||+++ ++++++.| ++|++++.|+++|||+||++.. ++++++++++. ++||
T Consensus 77 ~~~GSG~vi~~~G~IlTn~HVV~~~~~i~V~~~dg~~~~a~vv~~d~~~DlAvlkv~~~----~~~~~~l~~s~~~~~G~ 152 (351)
T TIGR02038 77 QGLGSGVIMSKEGYILTNYHVIKKADQIVVALQDGRKFEAELVGSDPLTDLAVLKIEGD----NLPTIPVNLDRPPHVGD 152 (351)
T ss_pred cceEEEEEEeCCeEEEecccEeCCCCEEEEEECCCCEEEEEEEEecCCCCEEEEEecCC----CCceEeccCcCccCCCC
Confidence 4568999997 66766665 6899999999999999999986 57888998764 6789
Q ss_pred eEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccC---Cccce
Q 017471 71 AVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE---DVENI 147 (371)
Q Consensus 71 ~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~---~~~~~ 147 (371)
+|+++|||++... +++.|+||+.++.... ......+||+|+++++|||||||+|.+||||||+++.+... ...++
T Consensus 153 ~V~aiG~P~~~~~-s~t~GiIs~~~r~~~~-~~~~~~~iqtda~i~~GnSGGpl~n~~G~vIGI~~~~~~~~~~~~~~g~ 230 (351)
T TIGR02038 153 VVLAIGNPYNLGQ-TITQGIISATGRNGLS-SVGRQNFIQTDAAINAGNSGGALINTNGELVGINTASFQKGGDEGGEGI 230 (351)
T ss_pred EEEEEeCCCCCCC-cEEEEEEEeccCcccC-CCCcceEEEECCccCCCCCcceEECCCCeEEEEEeeeecccCCCCccce
Confidence 9999999998766 8999999999874332 22335689999999999999999999999999998766432 23688
Q ss_pred eccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECC
Q 017471 148 GYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDG 226 (371)
Q Consensus 148 ~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG 226 (371)
+|+||++.++++++++.++|++. ++|||+.++++ ++..++.+|++ ...|++|.+|.++|||++ ||++||+|++|||
T Consensus 231 ~faIP~~~~~~vl~~l~~~g~~~-r~~lGv~~~~~-~~~~~~~lgl~-~~~Gv~V~~V~~~spA~~aGL~~GDvI~~Ing 307 (351)
T TIGR02038 231 NFAIPIKLAHKIMGKIIRDGRVI-RGYIGVSGEDI-NSVVAQGLGLP-DLRGIVITGVDPNGPAARAGILVRDVILKYDG 307 (351)
T ss_pred EEEecHHHHHHHHHHHhhcCccc-ceEeeeEEEEC-CHHHHHhcCCC-ccccceEeecCCCChHHHCCCCCCCEEEEECC
Confidence 99999999999999999999998 99999999999 88889999997 457999999999999999 9999999999999
Q ss_pred EEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccc
Q 017471 227 IDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATH 278 (371)
Q Consensus 227 ~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~ 278 (371)
++|.++.++. ..+...++|++++++|.|+|+.+++++++.+.
T Consensus 308 ~~V~s~~dl~----------~~l~~~~~g~~v~l~v~R~g~~~~~~v~l~~~ 349 (351)
T TIGR02038 308 KDVIGAEELM----------DRIAETRPGSKVMVTVLRQGKQLELPVTIDEK 349 (351)
T ss_pred EEcCCHHHHH----------HHHHhcCCCCEEEEEEEECCEEEEEEEEecCC
Confidence 9999998865 56766678999999999999999999988654
No 5
>PRK10898 serine endoprotease; Provisional
Probab=100.00 E-value=2.6e-43 Score=343.81 Aligned_cols=250 Identities=20% Similarity=0.342 Sum_probs=214.3
Q ss_pred eeeeccEEEEE--EEecCCCe-----------------EeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCC--CCC
Q 017471 11 LNSRNEALILS--TWLLCSPS-----------------APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELP--ALQ 69 (371)
Q Consensus 11 ~~~~gsg~vi~--~~~~~~~~-----------------~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~--~lg 69 (371)
..+.||||+++ ++++++.| ++|++++.|+.+||||||++.. ++++++++++. .+|
T Consensus 76 ~~~~GSGfvi~~~G~IlTn~HVv~~a~~i~V~~~dg~~~~a~vv~~d~~~DlAvl~v~~~----~l~~~~l~~~~~~~~G 151 (353)
T PRK10898 76 IRTLGSGVIMDQRGYILTNKHVINDADQIIVALQDGRVFEALLVGSDSLTDLAVLKINAT----NLPVIPINPKRVPHIG 151 (353)
T ss_pred ccceeeEEEEeCCeEEEecccEeCCCCEEEEEeCCCCEEEEEEEEEcCCCCEEEEEEcCC----CCCeeeccCcCcCCCC
Confidence 34679999997 66766666 7999999999999999999986 57889998764 578
Q ss_pred CeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccCC----cc
Q 017471 70 DAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHED----VE 145 (371)
Q Consensus 70 d~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~----~~ 145 (371)
|+|+++|||++... +++.|+||+.++..... .....+||+|+++++|||||||+|.+||||||+++.+...+ ..
T Consensus 152 ~~V~aiG~P~g~~~-~~t~Giis~~~r~~~~~-~~~~~~iqtda~i~~GnSGGPl~n~~G~vvGI~~~~~~~~~~~~~~~ 229 (353)
T PRK10898 152 DVVLAIGNPYNLGQ-TITQGIISATGRIGLSP-TGRQNFLQTDASINHGNSGGALVNSLGELMGINTLSFDKSNDGETPE 229 (353)
T ss_pred CEEEEEeCCCCcCC-CcceeEEEeccccccCC-ccccceEEeccccCCCCCcceEECCCCeEEEEEEEEecccCCCCccc
Confidence 99999999998765 89999999987643222 12246899999999999999999999999999998765332 25
Q ss_pred ceeccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEE
Q 017471 146 NIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSF 224 (371)
Q Consensus 146 ~~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~v 224 (371)
+++||||++.+++++++|+++|++. ++|||+..+++ ++..+..++++ ...|++|.+|.++|||++ ||++||+|++|
T Consensus 230 g~~faIP~~~~~~~~~~l~~~G~~~-~~~lGi~~~~~-~~~~~~~~~~~-~~~Gv~V~~V~~~spA~~aGL~~GDvI~~I 306 (353)
T PRK10898 230 GIGFAIPTQLATKIMDKLIRDGRVI-RGYIGIGGREI-APLHAQGGGID-QLQGIVVNEVSPDGPAAKAGIQVNDLIISV 306 (353)
T ss_pred ceEEEEchHHHHHHHHHHhhcCccc-ccccceEEEEC-CHHHHHhcCCC-CCCeEEEEEECCCChHHHcCCCCCCEEEEE
Confidence 7899999999999999999999998 89999999998 77777777876 357999999999999999 99999999999
Q ss_pred CCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEecccc
Q 017471 225 DGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHR 279 (371)
Q Consensus 225 nG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~ 279 (371)
||++|.++.++. +.+....+|++++++|.|+|+.+++++++.+.+
T Consensus 307 ng~~V~s~~~l~----------~~l~~~~~g~~v~l~v~R~g~~~~~~v~l~~~p 351 (353)
T PRK10898 307 NNKPAISALETM----------DQVAEIRPGSVIPVVVMRDDKQLTLQVTIQEYP 351 (353)
T ss_pred CCEEcCCHHHHH----------HHHHhcCCCCEEEEEEEECCEEEEEEEEeccCC
Confidence 999999988864 566666789999999999999999999887554
No 6
>COG0265 DegQ Trypsin-like serine proteases, typically periplasmic, contain C-terminal PDZ domain [Posttranslational modification, protein turnover, chaperones]
Probab=100.00 E-value=2.3e-34 Score=281.25 Aligned_cols=234 Identities=26% Similarity=0.402 Sum_probs=205.2
Q ss_pred cCCCeEeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCCC--CCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCC
Q 017471 25 LCSPSAPSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELPA--LQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHG 102 (371)
Q Consensus 25 ~~~~~~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~~--lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~ 102 (371)
.+++.++|++++.|+..|+|++|++... .++.+.++++.. +||+++++|+|++... +++.|+||..++......
T Consensus 103 ~dg~~~~a~~vg~d~~~dlavlki~~~~---~~~~~~~~~s~~l~vg~~v~aiGnp~g~~~-tvt~Givs~~~r~~v~~~ 178 (347)
T COG0265 103 ADGREVPAKLVGKDPISDLAVLKIDGAG---GLPVIALGDSDKLRVGDVVVAIGNPFGLGQ-TVTSGIVSALGRTGVGSA 178 (347)
T ss_pred CCCCEEEEEEEecCCccCEEEEEeccCC---CCceeeccCCCCcccCCEEEEecCCCCccc-ceeccEEeccccccccCc
Confidence 4666689999999999999999999874 378889998875 5699999999999665 999999999998622222
Q ss_pred CeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccCC-ccceeccccCcchhHhHhhhhhCCeeecccccceeeeE
Q 017471 103 STELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHED-VENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQK 181 (371)
Q Consensus 103 ~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~-~~~~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~ 181 (371)
.....+||+|+++|+||||||++|.+|++|||+++.+...+ ..+++|++|++.++.+++++...|++. ++++|+.+.+
T Consensus 179 ~~~~~~IqtdAain~gnsGgpl~n~~g~~iGint~~~~~~~~~~gigfaiP~~~~~~v~~~l~~~G~v~-~~~lgv~~~~ 257 (347)
T COG0265 179 GGYVNFIQTDAAINPGNSGGPLVNIDGEVVGINTAIIAPSGGSSGIGFAIPVNLVAPVLDELISKGKVV-RGYLGVIGEP 257 (347)
T ss_pred ccccchhhcccccCCCCCCCceEcCCCcEEEEEEEEecCCCCcceeEEEecHHHHHHHHHHHHHcCCcc-ccccceEEEE
Confidence 23567899999999999999999999999999998886553 466999999999999999999888888 9999999998
Q ss_pred cCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEE
Q 017471 182 MENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAV 260 (371)
Q Consensus 182 ~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l 260 (371)
+ +...+ +|++ ...|++|.+|.+++||++ |+++||+|+++||+++.+..++. ..+....+|+++.+
T Consensus 258 ~-~~~~~--~g~~-~~~G~~V~~v~~~spa~~agi~~Gdii~~vng~~v~~~~~l~----------~~v~~~~~g~~v~~ 323 (347)
T COG0265 258 L-TADIA--LGLP-VAAGAVVLGVLPGSPAAKAGIKAGDIITAVNGKPVASLSDLV----------AAVASNRPGDEVAL 323 (347)
T ss_pred c-ccccc--cCCC-CCCceEEEecCCCChHHHcCCCCCCEEEEECCEEccCHHHHH----------HHHhccCCCCEEEE
Confidence 8 65555 7877 667999999999999999 99999999999999999998875 67777779999999
Q ss_pred EEEECCEEEEEEEEecc
Q 017471 261 KVLRDSKILNFNITLAT 277 (371)
Q Consensus 261 ~v~R~g~~~~~~v~l~~ 277 (371)
++.|+|+.+++.+++..
T Consensus 324 ~~~r~g~~~~~~v~l~~ 340 (347)
T COG0265 324 KLLRGGKERELAVTLGD 340 (347)
T ss_pred EEEECCEEEEEEEEecC
Confidence 99999999999999976
No 7
>KOG1320 consensus Serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=99.93 E-value=2.9e-26 Score=225.77 Aligned_cols=320 Identities=34% Similarity=0.498 Sum_probs=276.7
Q ss_pred eeeccCCeeeeeccEEEEE-EEecCCCe---------------------EeEEEEEecCCCCEEEEEEecCCCcCCccce
Q 017471 3 ILWSTTRRLNSRNEALILS-TWLLCSPS---------------------APSATLVTADICIYTMLTVEDDEFWEGVLPV 60 (371)
Q Consensus 3 ~~~~~~~~~~~~gsg~vi~-~~~~~~~~---------------------~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~ 60 (371)
.+|++.++..+.++||.+. ..+++++| +.|++...-.++|+|++.++..+||+++.|+
T Consensus 77 ~pw~~~~q~~~~~s~f~i~~~~lltn~~~v~~~~~~~~v~v~~~gs~~k~~~~v~~~~~~cd~Avv~Ie~~~f~~~~~~~ 156 (473)
T KOG1320|consen 77 LPWQRTRQFSSGGSGFAIYGKKLLTNAHVVAPNNDHKFVTVKKHGSPRKYKAFVAAVFEECDLAVVYIESEEFWKGMNPF 156 (473)
T ss_pred CcceeeehhcccccchhhcccceeecCccccccccccccccccCCCchhhhhhHHHhhhcccceEEEEeeccccCCCccc
Confidence 5899999999999999998 44444444 4677777778899999999999999999999
Q ss_pred ecCCCCCCCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecc
Q 017471 61 EFGELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLK 140 (371)
Q Consensus 61 ~l~~s~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~ 140 (371)
++++.+.+.+-++++| ++...+|.|.|++.+...|.++...+..+|+|+++++|+||+|.+...+++.|+++..+.
T Consensus 157 e~~~ip~l~~S~~Vv~----gd~i~VTnghV~~~~~~~y~~~~~~l~~vqi~aa~~~~~s~ep~i~g~d~~~gvA~l~ik 232 (473)
T KOG1320|consen 157 ELGDIPSLNGSGFVVG----GDGIIVTNGHVVRVEPRIYAHSSTVLLRVQIDAAIGPGNSGEPVIVGVDKVAGVAFLKIK 232 (473)
T ss_pred ccCCCcccCccEEEEc----CCcEEEEeeEEEEEEeccccCCCcceeeEEEEEeecCCccCCCeEEccccccceEEEEEe
Confidence 9999999999999998 455699999999999888888888888999999999999999999988999999998874
Q ss_pred cCCccceeccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCccccCCCCCCE
Q 017471 141 HEDVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESEVLKPSDI 220 (371)
Q Consensus 141 ~~~~~~~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~GL~~GDv 220 (371)
..+ ++.+.+|...+.++.......+.+.+++.++...+.+++.+.++.+.|..+ +|+.+.++.+.++|.+-+++||.
T Consensus 233 ~~~--~i~~~i~~~~~~~~~~G~~~~a~~~~f~~~nt~t~g~vs~~~R~~~~lg~~-~g~~i~~~~qtd~ai~~~nsg~~ 309 (473)
T KOG1320|consen 233 TPE--NILYVIPLGVSSHFRTGVEVSAIGNGFGLLNTLTQGMVSGQLRKSFKLGLE-TGVLISKINQTDAAINPGNSGGP 309 (473)
T ss_pred cCC--cccceeecceeeeecccceeeccccCceeeeeeeecccccccccccccCcc-cceeeeeecccchhhhcccCCCc
Confidence 322 788999999999998888788888889999999999989999999999877 89999999999988888999999
Q ss_pred EEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccccccCCCCCCCCCCCceeeccEE
Q 017471 221 ILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFV 300 (371)
Q Consensus 221 Il~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~~~~~~~~~~~~~~~Gl~ 300 (371)
|+++||..|. +.+++.+|+.|++.+..+.+++++...+.|.+ ++.++++......|.+.+.+.|.|++..|++
T Consensus 310 ll~~DG~~Ig----Vn~~~~~ri~~~~~iSf~~p~d~vl~~v~r~~---e~~~~lr~~~~~~p~~~~~g~~s~~i~~g~v 382 (473)
T KOG1320|consen 310 LLNLDGEVIG----VNTRKVTRIGFSHGISFKIPIDTVLVIVLRLG---EFQISLRPVKPLVPVHQYIGLPSYYIFAGLV 382 (473)
T ss_pred EEEecCcEee----eeeeeeEEeeccccceeccCchHhhhhhhhhh---hhceeeccccCcccccccCCceeEEEecceE
Confidence 9999999997 55677889999999999999999999999998 5677777778888888999999999999999
Q ss_pred Eech--HHHHHHcCcccccCceEEEEEecCChHhHHhhhhhh
Q 017471 301 FSRC--LYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSS 340 (371)
Q Consensus 301 ~~~l--~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~ 340 (371)
|+++ ++...+ ....++++..+-++||+.--..+-||
T Consensus 383 f~~~~~~~~~~~----~~~q~v~is~Vlp~~~~~~~~~~~g~ 420 (473)
T KOG1320|consen 383 FVPLTKSYIFPS----GVVQLVLVSQVLPGSINGGYGLKPGD 420 (473)
T ss_pred EeecCCCccccc----cceeEEEEEEeccCCCcccccccCCC
Confidence 9876 333211 12268999999999998765555544
No 8
>KOG1421 consensus Predicted signaling-associated protein (contains a PDZ domain) [General function prediction only]
Probab=99.91 E-value=4.7e-24 Score=212.40 Aligned_cols=303 Identities=16% Similarity=0.241 Sum_probs=240.4
Q ss_pred eeccCCeeeeeccEEEEE---EEecCCCeE------------------eEEEEEecCCCCEEEEEEecCCC-cCCcccee
Q 017471 4 LWSTTRRLNSRNEALILS---TWLLCSPSA------------------PSATLVTADICIYTMLTVEDDEF-WEGVLPVE 61 (371)
Q Consensus 4 ~~~~~~~~~~~gsg~vi~---~~~~~~~~~------------------~A~vv~~d~~~DlAlLkv~~~~~-~~~l~~~~ 61 (371)
++++.-...+.++||+++ +++++++|+ +--.++.||.||+.+++++++.. ...+..+.
T Consensus 75 ~fdtesag~~~atgfvvd~~~gyiLtnrhvv~pgP~va~avf~n~ee~ei~pvyrDpVhdfGf~r~dps~ir~s~vt~i~ 154 (955)
T KOG1421|consen 75 AFDTESAGESEATGFVVDKKLGYILTNRHVVAPGPFVASAVFDNHEEIEIYPVYRDPVHDFGFFRYDPSTIRFSIVTEIC 154 (955)
T ss_pred ecccccccccceeEEEEecccceEEEeccccCCCCceeEEEecccccCCcccccCCchhhcceeecChhhcceeeeeccc
Confidence 356677788899999999 888999883 22345778889999999998753 12344555
Q ss_pred cCCC-CCCCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCC-----CeEEeEEEEEeeccCCCCCCceecCCCcEEEEE
Q 017471 62 FGEL-PALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHG-----STELLGLQIDAAINSGNSGGPAFNDKGKCVGIA 135 (371)
Q Consensus 62 l~~s-~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~-----~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~ 135 (371)
+... .++|.++.++||..+. ..++-.|.+|++++.....+ .....++|.-+....|.||.|++|.+|..|.++
T Consensus 155 lap~~akvgseirvvgNDagE-klsIlagflSrldr~apdyg~~~yndfnTfy~QaasstsggssgspVv~i~gyAVAl~ 233 (955)
T KOG1421|consen 155 LAPELAKVGSEIRVVGNDAGE-KLSILAGFLSRLDRNAPDYGEDTYNDFNTFYIQAASSTSGGSSGSPVVDIPGYAVALN 233 (955)
T ss_pred cCccccccCCceEEecCCccc-eEEeehhhhhhccCCCccccccccccccceeeeehhcCCCCCCCCceecccceEEeee
Confidence 5543 3688999999998664 45889999999987533222 122346899999999999999999999999998
Q ss_pred eeecccCCccceeccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCC-----------CCCc-eEEE
Q 017471 136 FQSLKHEDVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKA-----------DQKG-VRIR 203 (371)
Q Consensus 136 ~~~~~~~~~~~~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~-----------~~~g-v~V~ 203 (371)
.... .....+|++|.+.+.+.|.-++++.-.. |+.|-+++..- .-+..+.+||+. ...| ++|.
T Consensus 234 agg~---~ssas~ffLpLdrV~RaL~clq~n~PIt-RGtLqvefl~k-~~de~rrlGL~sE~eqv~r~k~P~~tgmLvV~ 308 (955)
T KOG1421|consen 234 AGGS---ISSASDFFLPLDRVVRALRCLQNNTPIT-RGTLQVEFLHK-LFDECRRLGLSSEWEQVVRTKFPERTGMLVVE 308 (955)
T ss_pred cCCc---ccccccceeeccchhhhhhhhhcCCCcc-cceEEEEEehh-hhHHHHhcCCcHHHHHHHHhcCcccceeEEEE
Confidence 7644 3456689999999999999997666666 89998888776 677888899864 2345 4577
Q ss_pred EECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccccccCC
Q 017471 204 RVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIP 283 (371)
Q Consensus 204 ~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~ 283 (371)
.|.++|||++-|++||++++||+.-+.++.++. .+.....|+.++|+|+|+|++.+++++........|
T Consensus 309 ~vL~~gpa~k~Le~GDillavN~t~l~df~~l~-----------~iLDegvgk~l~LtI~Rggqelel~vtvqdlh~itp 377 (955)
T KOG1421|consen 309 TVLPEGPAEKKLEPGDILLAVNSTCLNDFEALE-----------QILDEGVGKNLELTIQRGGQELELTVTVQDLHGITP 377 (955)
T ss_pred EeccCCchhhccCCCcEEEEEcceehHHHHHHH-----------HHHhhccCceEEEEEEeCCEEEEEEEEeccccCCCC
Confidence 899999999999999999999999998888764 444556899999999999999999999987766544
Q ss_pred CCCCCCCCCceeeccEEEech-HHHHHHcCcccccCceEEEEEecCChHhH
Q 017471 284 SHNKGRPPSYYIIAGFVFSRC-LYLISVLSMERIMNMKLRSSFWTSSCIQC 333 (371)
Q Consensus 284 ~~~~~~~~~~~~~~Gl~~~~l-~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~ 333 (371)
. ++..|+|.+|+++ +.++..|.++.. |+|++++. +||+.-
T Consensus 378 ~-------R~levcGav~hdlsyq~ar~y~lP~~--GvyVa~~~-gsf~~~ 418 (955)
T KOG1421|consen 378 D-------RFLEVCGAVFHDLSYQLARLYALPVE--GVYVASPG-GSFRHR 418 (955)
T ss_pred c-------eEEEEcceEecCCCHHHHhhcccccC--cEEEccCC-CCcccc
Confidence 4 4666999999999 888888888744 99999998 888643
No 9
>KOG1320 consensus Serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=99.88 E-value=1.3e-21 Score=193.07 Aligned_cols=236 Identities=21% Similarity=0.223 Sum_probs=185.6
Q ss_pred CeEeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCCC--CCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCC--
Q 017471 28 PSAPSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELPA--LQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGS-- 103 (371)
Q Consensus 28 ~~~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~~--lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~-- 103 (371)
..+.+.+++.|+..|+|+++++.++ ...++++++.+.. .|+++.++|+|++..+ +.+.|+++...|..+..+.
T Consensus 211 ~s~ep~i~g~d~~~gvA~l~ik~~~--~i~~~i~~~~~~~~~~G~~~~a~~~~f~~~n-t~t~g~vs~~~R~~~~lg~~~ 287 (473)
T KOG1320|consen 211 NSGEPVIVGVDKVAGVAFLKIKTPE--NILYVIPLGVSSHFRTGVEVSAIGNGFGLLN-TLTQGMVSGQLRKSFKLGLET 287 (473)
T ss_pred ccCCCeEEccccccceEEEEEecCC--cccceeecceeeeecccceeeccccCceeee-eeeecccccccccccccCccc
Confidence 6678999999999999999998664 2378888887665 5799999999999998 9999999999887554332
Q ss_pred --eEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccCC-ccceeccccCcchhHhHhhhhhCC---eee-----cc
Q 017471 104 --TELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHED-VENIGYVIPTPVIMHFIQDYEKNG---AYT-----GF 172 (371)
Q Consensus 104 --~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~-~~~~~~aiP~~~i~~~l~~L~~~g---~~~-----g~ 172 (371)
.-.+++|+|+++++||||||++|.+|++||++++...+-+ ...++|++|.+.+..++.+..+.. +.. .+
T Consensus 288 g~~i~~~~qtd~ai~~~nsg~~ll~~DG~~IgVn~~~~~ri~~~~~iSf~~p~d~vl~~v~r~~e~~~~lr~~~~~~p~~ 367 (473)
T KOG1320|consen 288 GVLISKINQTDAAINPGNSGGPLLNLDGEVIGVNTRKVTRIGFSHGISFKIPIDTVLVIVLRLGEFQISLRPVKPLVPVH 367 (473)
T ss_pred ceeeeeecccchhhhcccCCCcEEEecCcEeeeeeeeeEEeeccccceeccCchHhhhhhhhhhhhceeeccccCccccc
Confidence 3356899999999999999999999999999988764322 357899999999988888763211 111 13
Q ss_pred cccceeeeEcCCHHHH-----HhccCCC-CCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchh
Q 017471 173 PLLGVEWQKMENPDLR-----VAMSMKA-DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGF 245 (371)
Q Consensus 173 ~~lGi~~~~~~~~~~~-----~~~gl~~-~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~ 245 (371)
.|+|.....+ ++.+. +.+-.+. ..++++|.+|.|++++.. ++++||+|.+|||++|.+..++.
T Consensus 368 ~~~g~~s~~i-~~g~vf~~~~~~~~~~~~~~q~v~is~Vlp~~~~~~~~~~~g~~V~~vng~~V~n~~~l~--------- 437 (473)
T KOG1320|consen 368 QYIGLPSYYI-FAGLVFVPLTKSYIFPSGVVQLVLVSQVLPGSINGGYGLKPGDQVVKVNGKPVKNLKHLY--------- 437 (473)
T ss_pred ccCCceeEEE-ecceEEeecCCCccccccceeEEEEEEeccCCCcccccccCCCEEEEECCEEeechHHHH---------
Confidence 4666665555 22211 1121221 125899999999999999 99999999999999999999986
Q ss_pred hhhhhccCCCCEEEEEEEECCEEEEEEEEecc
Q 017471 246 SYLVSQKYTGDSAAVKVLRDSKILNFNITLAT 277 (371)
Q Consensus 246 ~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~ 277 (371)
.++....+++++.+..+|+.|..++.+.++.
T Consensus 438 -~~i~~~~~~~~v~vl~~~~~e~~tl~Il~~~ 468 (473)
T KOG1320|consen 438 -ELIEECSTEDKVAVLDRRSAEDATLEILPEH 468 (473)
T ss_pred -HHHHhcCcCceEEEEEecCccceeEEecccc
Confidence 6788777889999999999999998887653
No 10
>KOG1421 consensus Predicted signaling-associated protein (contains a PDZ domain) [General function prediction only]
Probab=99.80 E-value=3e-18 Score=171.28 Aligned_cols=271 Identities=15% Similarity=0.138 Sum_probs=205.6
Q ss_pred CCeEeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCC-CCCCeEEEEEeCCCCCCc--eeeeeEEeeeeeeeccCCC
Q 017471 27 SPSAPSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELP-ALQDAVTVVGYPIGGDTI--SVTSGVVSRIEILSYVHGS 103 (371)
Q Consensus 27 ~~~~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~-~lgd~V~~iG~p~g~~~~--s~t~G~Vs~~~~~~~~~~~ 103 (371)
.-.++|++.+.|+.+++|.+|++++. ...+.|.+.. ..||++...|+....+.. ..+...||.........++
T Consensus 585 S~~i~a~~~fL~~t~n~a~~kydp~~----~~~~kl~~~~v~~gD~~~f~g~~~~~r~ltaktsv~dvs~~~~ps~~~pr 660 (955)
T KOG1421|consen 585 SDGIPANVSFLHPTENVASFKYDPAL----EVQLKLTDTTVLRGDECTFEGFTEDLRALTAKTSVTDVSVVIIPSSVMPR 660 (955)
T ss_pred cccccceeeEecCccceeEeccChhH----hhhhccceeeEecCCceeEecccccchhhcccceeeeeEEEEecCCCCcc
Confidence 33479999999999999999999973 3556666654 457999999998765531 2233334433332222222
Q ss_pred ---eEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccC-C--ccceeccccCcchhHhHhhhhhCCeeecccccce
Q 017471 104 ---TELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-D--VENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGV 177 (371)
Q Consensus 104 ---~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~-~--~~~~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi 177 (371)
..++.|.+++.+..++-.|-+.|.+|+|+|+|...+.+. + ...+.|.+.+..+.+.++.|+.+++.. ...+|+
T Consensus 661 ~r~~n~e~Is~~~nlsT~c~sg~ltdddg~vvalwl~~~ge~~~~kd~~y~~gl~~~~~l~vl~rlk~g~~~r-p~i~~v 739 (955)
T KOG1421|consen 661 FRATNLEVISFMDNLSTSCLSGRLTDDDGEVVALWLSVVGEDVGGKDYTYKYGLSMSYILPVLERLKLGPSAR-PTIAGV 739 (955)
T ss_pred eeecceEEEEEeccccccccceEEECCCCeEEEEEeeeeccccCCceeEEEeccchHHHHHHHHHHhcCCCCC-ceeecc
Confidence 346789999999999989999999999999998777543 2 234678899999999999998877776 677899
Q ss_pred eeeEcCCHHHHHhccCCCC------------CCceEEEEECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchh
Q 017471 178 EWQKMENPDLRVAMSMKAD------------QKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGF 245 (371)
Q Consensus 178 ~~~~~~~~~~~~~~gl~~~------------~~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~ 245 (371)
+|..+ +...++.+|++.+ .+-.+|++|.+.-+- -|..||+|+++||+-|+...|+.
T Consensus 740 ef~~i-~laqar~lglp~e~imk~e~es~~~~ql~~ishv~~~~~k--il~~gdiilsvngk~itr~~dl~--------- 807 (955)
T KOG1421|consen 740 EFSHI-TLAQARTLGLPSEFIMKSEEESTIPRQLYVISHVRPLLHK--ILGVGDIILSVNGKMITRLSDLH--------- 807 (955)
T ss_pred ceeeE-EeehhhccCCCHHHHhhhhhcCCCcceEEEEEeeccCccc--ccccccEEEEecCeEEeeehhhh---------
Confidence 99999 8888889999852 124668888775543 59999999999999999998874
Q ss_pred hhhhhccCCCCEEEEEEEECCEEEEEEEEeccccccCCCCCCCCCCCceeeccEEEech-HHHHHHcCcccccCceEEEE
Q 017471 246 SYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRC-LYLISVLSMERIMNMKLRSS 324 (371)
Q Consensus 246 ~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~~~~~~~~~~~~~~~Gl~~~~l-~~~~~~~~~~~~~~~~~v~~ 324 (371)
. +. .+...|+|+|.+++++++.-+..+ ..+..+|+|..+|++ ..+.++. .+-..|+|+.+
T Consensus 808 -d-~~------eid~~ilrdg~~~~ikipt~p~~e---------t~r~vi~~gailq~ph~av~~q~--edlp~gvyvt~ 868 (955)
T KOG1421|consen 808 -D-FE------EIDAVILRDGIEMEIKIPTYPEYE---------TSRAVIWMGAILQPPHSAVFEQV--EDLPEGVYVTS 868 (955)
T ss_pred -h-hh------hhheeeeecCcEEEEEeccccccc---------cceEEEEEeccccCchHHHHHHH--hccCCceEEee
Confidence 1 21 578899999999999887765432 224567999999998 6666553 34559999999
Q ss_pred EecCChHhH
Q 017471 325 FWTSSCIQC 333 (371)
Q Consensus 325 ~~~~Sp~~~ 333 (371)
..++|||.-
T Consensus 869 rg~gspalq 877 (955)
T KOG1421|consen 869 RGYGSPALQ 877 (955)
T ss_pred cccCChhHh
Confidence 999999854
No 11
>PF13180 PDZ_2: PDZ domain; PDB: 2L97_A 1Y8T_A 2Z9I_A 1LCY_A 2PZD_B 2P3W_A 1VCW_C 1TE0_B 1SOZ_C 1SOT_C ....
Probab=99.60 E-value=3e-15 Score=115.81 Aligned_cols=81 Identities=32% Similarity=0.525 Sum_probs=70.0
Q ss_pred cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc
Q 017471 173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ 251 (371)
Q Consensus 173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~ 251 (371)
||||+.+....+ ..|++|.+|.++|||++ ||++||+|++|||++|++..++. +.+..
T Consensus 1 ~~lGv~~~~~~~------------~~g~~V~~V~~~spA~~aGl~~GD~I~~ing~~v~~~~~~~----------~~l~~ 58 (82)
T PF13180_consen 1 GGLGVTVQNLSD------------TGGVVVVSVIPGSPAAKAGLQPGDIILAINGKPVNSSEDLV----------NILSK 58 (82)
T ss_dssp -E-SEEEEECSC------------SSSEEEEEESTTSHHHHTTS-TTEEEEEETTEESSSHHHHH----------HHHHC
T ss_pred CEECeEEEEccC------------CCeEEEEEeCCCCcHHHCCCCCCcEEEEECCEEcCCHHHHH----------HHHHh
Confidence 589999998821 46999999999999999 99999999999999999988865 67778
Q ss_pred cCCCCEEEEEEEECCEEEEEEEEe
Q 017471 252 KYTGDSAAVKVLRDSKILNFNITL 275 (371)
Q Consensus 252 ~~~g~~v~l~v~R~g~~~~~~v~l 275 (371)
..+|+++++++.|+|+.+++++++
T Consensus 59 ~~~g~~v~l~v~R~g~~~~~~v~l 82 (82)
T PF13180_consen 59 GKPGDTVTLTVLRDGEELTVEVTL 82 (82)
T ss_dssp SSTTSEEEEEEEETTEEEEEEEE-
T ss_pred CCCCCEEEEEEEECCEEEEEEEEC
Confidence 899999999999999999999875
No 12
>TIGR03279 cyano_FeS_chp putative FeS-containing Cyanobacterial-specific oxidoreductase. Members of this protein family are predicted FeS-containing oxidoreductases of unknown function, apparently restricted to and universal across the Cyanobacteria. The high trusted cutoff score for this model, 700 bits, excludes homologs from other lineages. This exclusion seems justified because a significant number of sequence positions are simultaneously unique to and invariant across the Cyanobacteria, suggesting a specialized, conserved function, perhaps related to photosynthesis. A distantly related protein family, TIGR03278, in universal in and restricted to archaeal methanogens, and may be linked to methanogenesis.
Probab=99.52 E-value=2.4e-15 Score=147.75 Aligned_cols=102 Identities=21% Similarity=0.272 Sum_probs=80.7
Q ss_pred EEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEE-ECCEEEEEEEEecccc
Q 017471 202 IRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVL-RDSKILNFNITLATHR 279 (371)
Q Consensus 202 V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~-R~g~~~~~~v~l~~~~ 279 (371)
|.+|.|+|||++ ||++||+|++|||++|.++.|+. ..+ .++.++++|. |+|+..+++++....+
T Consensus 2 I~~V~pgSpAe~AGLe~GD~IlsING~~V~Dw~D~~----------~~l----~~e~l~L~V~~rdGe~~~l~Ie~~~de 67 (433)
T TIGR03279 2 ISAVLPGSIAEELGFEPGDALVSINGVAPRDLIDYQ----------FLC----ADEELELEVLDANGESHQIEIEKDLDE 67 (433)
T ss_pred cCCcCCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHh----cCCcEEEEEEcCCCeEEEEEEecCCCC
Confidence 668999999999 99999999999999999998864 233 2467899997 8998877776653222
Q ss_pred ccCCCCCCCCCCCceeeccEEEech--HHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhHHHHHHHHHHHhhhhh
Q 017471 280 RLIPSHNKGRPPSYYIIAGFVFSRC--LYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLLWCLRCLWLILILDMR 357 (371)
Q Consensus 280 ~~~~~~~~~~~~~~~~~~Gl~~~~l--~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~~~~~~~~~~~~~~~~ 357 (371)
+ +|+.|.+. ...+ +..-. |.|||.-|||.+||
T Consensus 68 d----------------lG~~f~~~~~d~~~-~C~N~-----------------------------C~FCFidQlP~gmR 101 (433)
T TIGR03279 68 D----------------LGLEFTTALFDGLI-QCNNR-----------------------------CPFCFIDQQPPGKR 101 (433)
T ss_pred C----------------CcEEeccccCCccc-ccCCc-----------------------------CceEeccCCCCCCc
Confidence 2 79998654 2232 55555 99999999999999
Q ss_pred hHHHHH
Q 017471 358 RLLTLR 363 (371)
Q Consensus 358 ~~~~~~ 363 (371)
++||+|
T Consensus 102 ~sLY~K 107 (433)
T TIGR03279 102 ESLYLK 107 (433)
T ss_pred Ccceec
Confidence 999976
No 13
>cd00987 PDZ_serine_protease PDZ domain of tryspin-like serine proteases, such as DegP/HtrA, which are oligomeric proteins involved in heat-shock response, chaperone function, and apoptosis. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.43 E-value=6.8e-13 Score=103.81 Aligned_cols=88 Identities=35% Similarity=0.599 Sum_probs=74.9
Q ss_pred cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc
Q 017471 173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ 251 (371)
Q Consensus 173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~ 251 (371)
+|+|+.++++ +++.+..++++ ...|++|.+|.++|||++ ||++||+|++|||+++.++.++. ..+..
T Consensus 1 ~~~G~~~~~~-~~~~~~~~~~~-~~~g~~V~~v~~~s~a~~~gl~~GD~I~~Ing~~i~~~~~~~----------~~l~~ 68 (90)
T cd00987 1 PWLGVTVQDL-TPDLAEELGLK-DTKGVLVASVDPGSPAAKAGLKPGDVILAVNGKPVKSVADLR----------RALAE 68 (90)
T ss_pred CccceEEeEC-CHHHHHHcCCC-CCCEEEEEEECCCCHHHHcCCCcCCEEEEECCEECCCHHHHH----------HHHHh
Confidence 5899999999 77777767765 457999999999999998 99999999999999999988764 56665
Q ss_pred cCCCCEEEEEEEECCEEEEEE
Q 017471 252 KYTGDSAAVKVLRDSKILNFN 272 (371)
Q Consensus 252 ~~~g~~v~l~v~R~g~~~~~~ 272 (371)
...++.+.+++.|+|+..+++
T Consensus 69 ~~~~~~i~l~v~r~g~~~~~~ 89 (90)
T cd00987 69 LKPGDKVTLTVLRGGKELTVT 89 (90)
T ss_pred cCCCCEEEEEEEECCEEEEee
Confidence 556899999999999876654
No 14
>cd00986 PDZ_LON_protease PDZ domain of ATP-dependent LON serine proteases. Most PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this bacterial subfamily of protease-associated PDZ domains a C-terminal beta-strand is thought to form the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.30 E-value=9.9e-12 Score=95.22 Aligned_cols=72 Identities=28% Similarity=0.366 Sum_probs=63.5
Q ss_pred CCceEEEEECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEec
Q 017471 197 QKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLA 276 (371)
Q Consensus 197 ~~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~ 276 (371)
..|++|.+|.++|||++||++||+|++|||+++.++.++. ..+....+|+.+.+++.|+|+.+++++++.
T Consensus 7 ~~Gv~V~~V~~~s~A~~gL~~GD~I~~Ing~~v~~~~~~~----------~~l~~~~~~~~v~l~v~r~g~~~~~~v~l~ 76 (79)
T cd00986 7 YHGVYVTSVVEGMPAAGKLKAGDHIIAVDGKPFKEAEELI----------DYIQSKKEGDTVKLKVKREEKELPEDLILK 76 (79)
T ss_pred ecCEEEEEECCCCchhhCCCCCCEEEEECCEECCCHHHHH----------HHHHhCCCCCEEEEEEEECCEEEEEEEEEe
Confidence 4689999999999998799999999999999999988864 566655688999999999999999999987
Q ss_pred cc
Q 017471 277 TH 278 (371)
Q Consensus 277 ~~ 278 (371)
..
T Consensus 77 ~~ 78 (79)
T cd00986 77 TF 78 (79)
T ss_pred cc
Confidence 53
No 15
>cd00991 PDZ_archaeal_metalloprotease PDZ domain of archaeal zinc metalloprotases, presumably membrane-associated or integral membrane proteases, which may be involved in signalling and regulatory mechanisms. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.29 E-value=1.1e-11 Score=95.11 Aligned_cols=69 Identities=26% Similarity=0.273 Sum_probs=60.7
Q ss_pred CCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEE
Q 017471 196 DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT 274 (371)
Q Consensus 196 ~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~ 274 (371)
...|++|.+|.++|||++ ||++||+|++|||+++.++.++. ..+....+|+++.+++.|+|+..+++++
T Consensus 8 ~~~Gv~V~~V~~~spa~~aGL~~GDiI~~Ing~~v~~~~d~~----------~~l~~~~~g~~v~l~v~r~g~~~~~~~~ 77 (79)
T cd00991 8 AVAGVVIVGVIVGSPAENAVLHTGDVIYSINGTPITTLEDFM----------EALKPTKPGEVITVTVLPSTTKLTNVST 77 (79)
T ss_pred cCCcEEEEEECCCChHHhcCCCCCCEEEEECCEEcCCHHHHH----------HHHhcCCCCCEEEEEEEECCEEEEEEEE
Confidence 357999999999999998 99999999999999999998864 5666656789999999999999887765
No 16
>cd00990 PDZ_glycyl_aminopeptidase PDZ domain associated with archaeal and bacterial M61 glycyl-aminopeptidases. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand is presumed to form the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.26 E-value=2.3e-11 Score=93.11 Aligned_cols=77 Identities=22% Similarity=0.384 Sum_probs=63.8
Q ss_pred cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc
Q 017471 173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ 251 (371)
Q Consensus 173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~ 251 (371)
+|+|+.+..- ..|++|.+|.++|||++ ||++||+|++|||+++.++.+ .+..
T Consensus 1 ~~~G~~~~~~--------------~~~~~V~~V~~~s~a~~aGl~~GD~I~~Ing~~v~~~~~-------------~l~~ 53 (80)
T cd00990 1 PYLGLTLDKE--------------EGLGKVTFVRDDSPADKAGLVAGDELVAVNGWRVDALQD-------------RLKE 53 (80)
T ss_pred CcccEEEEcc--------------CCcEEEEEECCCChHHHhCCCCCCEEEEECCEEhHHHHH-------------HHHh
Confidence 5788877532 35799999999999999 999999999999999987443 3444
Q ss_pred cCCCCEEEEEEEECCEEEEEEEEec
Q 017471 252 KYTGDSAAVKVLRDSKILNFNITLA 276 (371)
Q Consensus 252 ~~~g~~v~l~v~R~g~~~~~~v~l~ 276 (371)
..+++.+.+++.|+|+..++++++.
T Consensus 54 ~~~~~~v~l~v~r~g~~~~~~v~~~ 78 (80)
T cd00990 54 YQAGDPVELTVFRDDRLIEVPLTLA 78 (80)
T ss_pred cCCCCEEEEEEEECCEEEEEEEEec
Confidence 4578899999999999988888764
No 17
>TIGR01713 typeII_sec_gspC general secretion pathway protein C. This model represents GspC, protein C of the main terminal branch of the general secretion pathway, also called type II secretion. This system transports folded proteins across the bacterial outer membrane and is widely distributed in Gram-negative pathogens.
Probab=99.22 E-value=5.4e-11 Score=111.38 Aligned_cols=101 Identities=15% Similarity=0.183 Sum_probs=86.9
Q ss_pred CcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecC
Q 017471 153 TPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAN 231 (371)
Q Consensus 153 ~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~ 231 (371)
...+.++++++.++++.. ++|+|+..... + ....|+.|..+.+++||++ |||+||+|++|||+++++
T Consensus 158 ~~~~~~v~~~l~~~g~~~-~~~lgi~p~~~-~----------g~~~G~~v~~v~~~s~a~~aGLr~GDvIv~ING~~i~~ 225 (259)
T TIGR01713 158 IVVSRRIIEELTKDPQKM-FDYIRLSPVMK-N----------DKLEGYRLNPGKDPSLFYKSGLQDGDIAVALNGLDLRD 225 (259)
T ss_pred hhhHHHHHHHHHHCHHhh-hheEeEEEEEe-C----------CceeEEEEEecCCCCHHHHcCCCCCCEEEEECCEEcCC
Confidence 346788899999989888 89999998655 2 1246999999999999999 999999999999999999
Q ss_pred CCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEe
Q 017471 232 DGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITL 275 (371)
Q Consensus 232 ~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l 275 (371)
+.++. ..+....++++++++|.|+|+.+++.+.+
T Consensus 226 ~~~~~----------~~l~~~~~~~~v~l~V~R~G~~~~i~v~~ 259 (259)
T TIGR01713 226 PEQAF----------QALQMLREETNLTLTVERDGQREDIYVRF 259 (259)
T ss_pred HHHHH----------HHHHhcCCCCeEEEEEEECCEEEEEEEEC
Confidence 98865 67777788899999999999998888764
No 18
>TIGR02037 degP_htrA_DO periplasmic serine protease, Do/DeqQ family. This family consists of a set proteins various designated DegP, heat shock protein HtrA, and protease DO. The ortholog in Pseudomonas aeruginosa is designated MucD and is found in an operon that controls mucoid phenotype. This family also includes the DegQ (HhoA) paralog in E. coli which can rescue a DegP mutant, but not the smaller DegS paralog, which cannot. Members of this family are located in the periplasm and have separable functions as both protease and chaperone. Members have a trypsin domain and two copies of a PDZ domain. This protein protects bacteria from thermal and other stresses and may be important for the survival of bacterial pathogens.// The chaperone function is dominant at low temperatures, whereas the proteolytic activity is turned on at elevated temperatures.
Probab=99.11 E-value=1.8e-10 Score=115.84 Aligned_cols=90 Identities=24% Similarity=0.473 Sum_probs=79.5
Q ss_pred ccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhh
Q 017471 172 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVS 250 (371)
Q Consensus 172 ~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~ 250 (371)
..++|+.+.++ +++.++.++++....|++|.+|.++|||++ ||++||+|++|||++|.++.++. +.+.
T Consensus 337 ~~~lGi~~~~l-~~~~~~~~~l~~~~~Gv~V~~V~~~SpA~~aGL~~GDvI~~Ing~~V~s~~d~~----------~~l~ 405 (428)
T TIGR02037 337 NPFLGLTVANL-SPEIRKELRLKGDVKGVVVTKVVSGSPAARAGLQPGDVILSVNQQPVSSVAELR----------KVLD 405 (428)
T ss_pred ccccceEEecC-CHHHHHHcCCCcCcCceEEEEeCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHH
Confidence 35799999999 888889999986567999999999999999 99999999999999999998865 6776
Q ss_pred ccCCCCEEEEEEEECCEEEEEE
Q 017471 251 QKYTGDSAAVKVLRDSKILNFN 272 (371)
Q Consensus 251 ~~~~g~~v~l~v~R~g~~~~~~ 272 (371)
..++|++++++|.|+|+...+.
T Consensus 406 ~~~~g~~v~l~v~R~g~~~~~~ 427 (428)
T TIGR02037 406 RAKKGGRVALLILRGGATIFVT 427 (428)
T ss_pred hcCCCCEEEEEEEECCEEEEEE
Confidence 6667999999999999987654
No 19
>PRK10779 zinc metallopeptidase RseP; Provisional
Probab=99.11 E-value=9.9e-11 Score=118.38 Aligned_cols=117 Identities=12% Similarity=0.072 Sum_probs=86.6
Q ss_pred eEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccc
Q 017471 200 VRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATH 278 (371)
Q Consensus 200 v~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~ 278 (371)
.+|++|.++|||++ |||+||+|+++||++|++++++. ..+....+|+++++++.|+|+.+++++++...
T Consensus 128 ~lV~~V~~~SpA~kAGLk~GDvI~~vnG~~V~~~~~l~----------~~v~~~~~g~~v~v~v~R~gk~~~~~v~l~~~ 197 (449)
T PRK10779 128 PVVGEIAPNSIAAQAQIAPGTELKAVDGIETPDWDAVR----------LALVSKIGDESTTITVAPFGSDQRRDKTLDLR 197 (449)
T ss_pred ccccccCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHhhccCCceEEEEEeCCccceEEEEeccc
Confidence 46899999999999 99999999999999999999976 56777778899999999999999999988654
Q ss_pred cccCCCCCCCCCCCceeeccEEEechHHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhH
Q 017471 279 RRLIPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLL 342 (371)
Q Consensus 279 ~~~~~~~~~~~~~~~~~~~Gl~~~~l~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~ 342 (371)
+...... ... .....|+. +. .....+.+..+.++|||+-.-.+-+|.|
T Consensus 198 ~~~~~~~--~~~--~~~~lGl~--~~----------~~~~~~vV~~V~~~SpA~~AGL~~GDvI 245 (449)
T PRK10779 198 HWAFEPD--KQD--PVSSLGIR--PR----------GPQIEPVLAEVQPNSAASKAGLQAGDRI 245 (449)
T ss_pred ccccCcc--ccc--hhhccccc--cc----------CCCcCcEEEeeCCCCHHHHcCCCCCCEE
Confidence 3221100 000 01123432 21 0112468899999999998777777765
No 20
>cd00989 PDZ_metalloprotease PDZ domain of bacterial and plant zinc metalloprotases, presumably membrane-associated or integral membrane proteases, which may be involved in signalling and regulatory mechanisms. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.08 E-value=2.8e-10 Score=86.72 Aligned_cols=66 Identities=24% Similarity=0.335 Sum_probs=55.9
Q ss_pred CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEE
Q 017471 198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT 274 (371)
Q Consensus 198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~ 274 (371)
..++|.+|.++|||++ ||++||+|++|||+++.++.++. ..+.. ..++.+.+++.|+|+..++.++
T Consensus 12 ~~~~V~~v~~~s~a~~~gl~~GD~I~~ing~~i~~~~~~~----------~~l~~-~~~~~~~l~v~r~~~~~~~~l~ 78 (79)
T cd00989 12 IEPVIGEVVPGSPAAKAGLKAGDRILAINGQKIKSWEDLV----------DAVQE-NPGKPLTLTVERNGETITLTLT 78 (79)
T ss_pred cCcEEEeECCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHHHH-CCCceEEEEEEECCEEEEEEec
Confidence 3588999999999998 99999999999999999988864 45544 3478899999999988777664
No 21
>cd00988 PDZ_CTP_protease PDZ domain of C-terminal processing-, tail-specific-, and tricorn proteases, which function in posttranslational protein processing, maturation, and disassembly or degradation, in Bacteria, Archaea, and plant chloroplasts. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.05 E-value=9.3e-10 Score=85.13 Aligned_cols=68 Identities=24% Similarity=0.342 Sum_probs=57.1
Q ss_pred CCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCC--CCccccccccchhhhhhhccCCCCEEEEEEEEC-CEEEEEE
Q 017471 197 QKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAND--GTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD-SKILNFN 272 (371)
Q Consensus 197 ~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~--~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~-g~~~~~~ 272 (371)
..+++|..|.++|||++ ||++||+|++|||+++.++ .++. ..+. ..+|+++.+++.|+ |+..+++
T Consensus 12 ~~~~~V~~v~~~s~a~~~gl~~GD~I~~vng~~i~~~~~~~~~----------~~l~-~~~~~~i~l~v~r~~~~~~~~~ 80 (85)
T cd00988 12 DGGLVITSVLPGSPAAKAGIKAGDIIVAIDGEPVDGLSLEDVV----------KLLR-GKAGTKVRLTLKRGDGEPREVT 80 (85)
T ss_pred CCeEEEEEecCCCCHHHcCCCCCCEEEEECCEEcCCCCHHHHH----------HHhc-CCCCCEEEEEEEcCCCCEEEEE
Confidence 36899999999999999 9999999999999999998 6643 3343 35688999999999 8888877
Q ss_pred EEe
Q 017471 273 ITL 275 (371)
Q Consensus 273 v~l 275 (371)
+++
T Consensus 81 ~~~ 83 (85)
T cd00988 81 LTR 83 (85)
T ss_pred EEE
Confidence 764
No 22
>PF13365 Trypsin_2: Trypsin-like peptidase domain; PDB: 1Y8T_A 2Z9I_A 3QO6_A 1L1J_A 1QY6_A 2O8L_A 3OTP_E 2ZLE_I 1KY9_A 3CS0_A ....
Probab=98.83 E-value=1.4e-08 Score=82.98 Aligned_cols=87 Identities=24% Similarity=0.325 Sum_probs=52.8
Q ss_pred EEEEEEecCCCeEe--EEEEEecCC-CCEEEEEEecCCCcCCccceecCCCCCCCCeEEEEEeCCCCCCceeeeeEEeee
Q 017471 18 LILSTWLLCSPSAP--SATLVTADI-CIYTMLTVEDDEFWEGVLPVEFGELPALQDAVTVVGYPIGGDTISVTSGVVSRI 94 (371)
Q Consensus 18 ~vi~~~~~~~~~~~--A~vv~~d~~-~DlAlLkv~~~~~~~~l~~~~l~~s~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~ 94 (371)
..+.+...++...+ |++++.|+. +|+|||+++ .....+.. ....+.....
T Consensus 31 ~~~~~~~~~~~~~~~~~~~~~~~~~~~D~All~v~---------------------~~~~~~~~------~~~~~~~~~~ 83 (120)
T PF13365_consen 31 SSVEVVFPDGRRVPPVAEVVYFDPDDYDLALLKVD---------------------PWTGVGGG------VRVPGSTSGV 83 (120)
T ss_dssp SEEEEEETTSCEEETEEEEEEEETT-TTEEEEEES---------------------CEEEEEEE------EEEEEEEEEE
T ss_pred CEEEEEecCCCEEeeeEEEEEECCccccEEEEEEe---------------------cccceeee------eEeeeecccc
Confidence 34446666777777 999999999 999999999 00000000 0111111111
Q ss_pred eeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEE
Q 017471 95 EILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGI 134 (371)
Q Consensus 95 ~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI 134 (371)
... ........++ +|+++.+|+||||++|.+|+||||
T Consensus 84 ~~~--~~~~~~~~~~-~~~~~~~G~SGgpv~~~~G~vvGi 120 (120)
T PF13365_consen 84 SPT--STNDNRMLYI-TDADTRPGSSGGPVFDSDGRVVGI 120 (120)
T ss_dssp EEE--EEEETEEEEE-ESSS-STTTTTSEEEETTSEEEEE
T ss_pred ccc--cCcccceeEe-eecccCCCcEeHhEECCCCEEEeC
Confidence 100 0001111125 899999999999999999999997
No 23
>cd00136 PDZ PDZ domain, also called DHR (Dlg homologous region) or GLGF (after a conserved sequence motif). Many PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. Heterodimerization through PDZ-PDZ domain interactions adds to the domain's versatility, and PDZ domain-mediated interactions may be modulated dynamically through target phosphorylation. Some PDZ domains play a role in scaffolding supramolecular complexes. PDZ domains are found in diverse signaling proteins in bacteria, archebacteria, and eurkayotes. This CD contains two distinct structural subgroups with either a N- or C-terminal beta-strand forming the peptide-binding groove base. The circular permutation placing the strand on the N-terminus appears to be found in Eumetazoa only, while the C-terminal variant is found in all three kingdoms of life, and seems to co-occur with protease domains. PDZ domains have been named after PSD95(pos
Probab=98.82 E-value=8e-09 Score=76.78 Aligned_cols=65 Identities=28% Similarity=0.454 Sum_probs=51.8
Q ss_pred ccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCC--CCccccccccchhhhhhh
Q 017471 174 LLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAND--GTVPFRHGERIGFSYLVS 250 (371)
Q Consensus 174 ~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~--~~l~~~~~~~~~~~~~~~ 250 (371)
++|+.+...+ ..+++|.+|.++|||++ ||++||+|++|||+++.++ .++. +.+.
T Consensus 2 ~~G~~~~~~~-------------~~~~~V~~v~~~s~a~~~gl~~GD~I~~Ing~~v~~~~~~~~~----------~~l~ 58 (70)
T cd00136 2 GLGFSIRGGT-------------EGGVVVLSVEPGSPAERAGLQAGDVILAVNGTDVKNLTLEDVA----------ELLK 58 (70)
T ss_pred CccEEEecCC-------------CCCEEEEEeCCCCHHHHcCCCCCCEEEEECCEECCCCCHHHHH----------HHHh
Confidence 5777776541 14899999999999999 9999999999999999998 5543 4444
Q ss_pred ccCCCCEEEEEE
Q 017471 251 QKYTGDSAAVKV 262 (371)
Q Consensus 251 ~~~~g~~v~l~v 262 (371)
. .+|+++++++
T Consensus 59 ~-~~g~~v~l~v 69 (70)
T cd00136 59 K-EVGEKVTLTV 69 (70)
T ss_pred h-CCCCeEEEEE
Confidence 4 3488888876
No 24
>TIGR00054 RIP metalloprotease RseP. A model that detects fragments as well matches a number of members of the PEPTIDASE FAMILY S2C. The region of match appears not to overlap the active site domain.
Probab=98.71 E-value=2.3e-08 Score=100.36 Aligned_cols=69 Identities=25% Similarity=0.323 Sum_probs=61.0
Q ss_pred CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEec
Q 017471 198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLA 276 (371)
Q Consensus 198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~ 276 (371)
.+++|.+|.++|||++ |||+||+|++|||++|++++|+. +.+.. .+++++++++.|+|+..++++++.
T Consensus 203 ~g~vV~~V~~~SpA~~aGL~~GD~Iv~Vng~~V~s~~dl~----------~~l~~-~~~~~v~l~v~R~g~~~~~~v~~~ 271 (420)
T TIGR00054 203 IEPVLSDVTPNSPAEKAGLKEGDYIQSINGEKLRSWTDFV----------SAVKE-NPGKSMDIKVERNGETLSISLTPE 271 (420)
T ss_pred cCcEEEEECCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHHHh-CCCCceEEEEEECCEEEEEEEEEc
Confidence 4799999999999999 99999999999999999999875 45544 578889999999999999998885
Q ss_pred c
Q 017471 277 T 277 (371)
Q Consensus 277 ~ 277 (371)
.
T Consensus 272 ~ 272 (420)
T TIGR00054 272 A 272 (420)
T ss_pred C
Confidence 3
No 25
>smart00228 PDZ Domain present in PSD-95, Dlg, and ZO-1/2. Also called DHR (Dlg homologous region) or GLGF (relatively well conserved tetrapeptide in these domains). Some PDZs have been shown to bind C-terminal polypeptides; others appear to bind internal (non-C-terminal) polypeptides. Different PDZs possess different binding specificities.
Probab=98.68 E-value=6.4e-08 Score=74.25 Aligned_cols=73 Identities=25% Similarity=0.340 Sum_probs=55.2
Q ss_pred cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc
Q 017471 173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ 251 (371)
Q Consensus 173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~ 251 (371)
..+|+.+....+ ...|++|..|.++|||++ ||++||+|++|||+++.+..+.. .....
T Consensus 12 ~~~G~~~~~~~~-----------~~~~~~i~~v~~~s~a~~~gl~~GD~I~~In~~~v~~~~~~~----------~~~~~ 70 (85)
T smart00228 12 GGLGFSLVGGKD-----------EGGGVVVSSVVPGSPAAKAGLKVGDVILEVNGTSVEGLTHLE----------AVDLL 70 (85)
T ss_pred CcccEEEECCCC-----------CCCCEEEEEECCCCHHHHcCCCCCCEEEEECCEECCCCCHHH----------HHHHH
Confidence 467888765411 116899999999999999 99999999999999999876643 22222
Q ss_pred cCCCCEEEEEEEECC
Q 017471 252 KYTGDSAAVKVLRDS 266 (371)
Q Consensus 252 ~~~g~~v~l~v~R~g 266 (371)
...++.+.+++.|++
T Consensus 71 ~~~~~~~~l~i~r~~ 85 (85)
T smart00228 71 KKAGGKVTLTVLRGG 85 (85)
T ss_pred HhCCCeEEEEEEeCC
Confidence 334668999999875
No 26
>PRK10779 zinc metallopeptidase RseP; Provisional
Probab=98.63 E-value=5.4e-08 Score=98.55 Aligned_cols=68 Identities=24% Similarity=0.324 Sum_probs=60.3
Q ss_pred ceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEecc
Q 017471 199 GVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLAT 277 (371)
Q Consensus 199 gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~ 277 (371)
+++|.+|.++|||++ ||++||+|++|||++|++++|+. +.+.. .+|+++.+++.|+|+..++++++..
T Consensus 222 ~~vV~~V~~~SpA~~AGL~~GDvIl~Ing~~V~s~~dl~----------~~l~~-~~~~~v~l~v~R~g~~~~~~v~~~~ 290 (449)
T PRK10779 222 EPVLAEVQPNSAASKAGLQAGDRIVKVDGQPLTQWQTFV----------TLVRD-NPGKPLALEIERQGSPLSLTLTPDS 290 (449)
T ss_pred CcEEEeeCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHh-CCCCEEEEEEEECCEEEEEEEEeee
Confidence 588999999999999 99999999999999999999875 45544 5788999999999999999988863
No 27
>TIGR00225 prc C-terminal peptidase (prc). A C-terminal peptidase with different substrates in different species including processing of D1 protein of the photosystem II reaction center in higher plants and cleavage of a peptide of 11 residues from the precursor form of penicillin-binding protein in E.coli E.coli and H influenza have the most distal branch of the tree and their proteins have an N-terminal 200 amino acids that show no homology to other proteins in the database.
Probab=98.59 E-value=1.2e-07 Score=92.52 Aligned_cols=71 Identities=21% Similarity=0.307 Sum_probs=58.0
Q ss_pred CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCC--CccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEE
Q 017471 198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDG--TVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT 274 (371)
Q Consensus 198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~--~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~ 274 (371)
.+++|.+|.++|||++ ||++||+|++|||++|.++. ++. . .....+|+++.+++.|+|+..+++++
T Consensus 62 ~~~~V~~V~~~spA~~aGL~~GD~I~~Ing~~v~~~~~~~~~----------~-~l~~~~g~~v~l~v~R~g~~~~~~v~ 130 (334)
T TIGR00225 62 GEIVIVSPFEGSPAEKAGIKPGDKIIKINGKSVAGMSLDDAV----------A-LIRGKKGTKVSLEILRAGKSKPLTFT 130 (334)
T ss_pred CEEEEEEeCCCChHHHcCCCCCCEEEEECCEECCCCCHHHHH----------H-hccCCCCCEEEEEEEeCCCCceEEEE
Confidence 4799999999999999 99999999999999999873 221 2 22335789999999999988888877
Q ss_pred ecccc
Q 017471 275 LATHR 279 (371)
Q Consensus 275 l~~~~ 279 (371)
+....
T Consensus 131 l~~~~ 135 (334)
T TIGR00225 131 LKRDR 135 (334)
T ss_pred EEEEE
Confidence 76543
No 28
>TIGR00054 RIP metalloprotease RseP. A model that detects fragments as well matches a number of members of the PEPTIDASE FAMILY S2C. The region of match appears not to overlap the active site domain.
Probab=98.57 E-value=6.1e-08 Score=97.30 Aligned_cols=66 Identities=23% Similarity=0.305 Sum_probs=55.5
Q ss_pred CCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEE
Q 017471 197 QKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT 274 (371)
Q Consensus 197 ~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~ 274 (371)
..|++|.+|.++|||++ |||+||+|+++||+++.++.++. ..+.... +++.+++.|+++..+++++
T Consensus 127 ~~g~~V~~V~~~SpA~~AGL~~GDvI~~vng~~v~~~~dl~----------~~ia~~~--~~v~~~I~r~g~~~~l~v~ 193 (420)
T TIGR00054 127 EVGPVIELLDKNSIALEAGIEPGDEILSVNGNKIPGFKDVR----------QQIADIA--GEPMVEILAERENWTFEVM 193 (420)
T ss_pred CCCceeeccCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHhhc--ccceEEEEEecCceEeccc
Confidence 36889999999999999 99999999999999999999875 4555444 6789999999887665443
No 29
>PF00595 PDZ: PDZ domain (Also known as DHR or GLGF) Coordinates are not yet available; InterPro: IPR001478 PDZ domains are found in diverse signalling proteins in bacteria, yeasts, plants, insects and vertebrates [, ]. PDZ domains can occur in one or multiple copies and are nearly always found in cytoplasmic proteins. They bind either the carboxyl-terminal sequences of proteins or internal peptide sequences []. In most cases, interaction between a PDZ domain and its target is constitutive, with a binding affinity of 1 to 10 microns. However, agonist-dependent activation of cell surface receptors is sometimes required to promote interaction with a PDZ protein. PDZ domain proteins are frequently associated with the plasma membrane, a compartment where high concentrations of phosphatidylinositol 4,5-bisphosphate (PIP2) are found. Direct interaction between PIP2 and a subset of class II PDZ domains (syntenin, CASK, Tiam-1) has been demonstrated. PDZ domains consist of 80 to 90 amino acids comprising six beta-strands (beta-A to beta-F) and two alpha-helices, A and B, compactly arranged in a globular structure. Peptide binding of the ligand takes place in an elongated surface groove as an anti-parallel beta-strand interacts with the beta-B strand and the B helix. The structure of PDZ domains allows binding to a free carboxylate group at the end of a peptide through a carboxylate-binding loop between the beta-A and beta-B strands.; GO: 0005515 protein binding; PDB: 3AXA_A 1WF8_A 1QAV_B 1QAU_A 1B8Q_A 1MC7_A 2KAW_A 1I16_A 1VB7_A 1WI4_A ....
Probab=98.50 E-value=2.3e-07 Score=71.18 Aligned_cols=72 Identities=24% Similarity=0.322 Sum_probs=52.8
Q ss_pred ccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhh
Q 017471 172 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVS 250 (371)
Q Consensus 172 ~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~ 250 (371)
...+|+.+....+ . ...+++|.+|.++|||++ ||++||.|++|||+++.++...+ ....+.
T Consensus 9 ~~~lG~~l~~~~~-~---------~~~~~~V~~v~~~~~a~~~gl~~GD~Il~INg~~v~~~~~~~--------~~~~l~ 70 (81)
T PF00595_consen 9 NGPLGFTLRGGSD-N---------DEKGVFVSSVVPGSPAERAGLKVGDRILEINGQSVRGMSHDE--------VVQLLK 70 (81)
T ss_dssp TSBSSEEEEEEST-S---------SSEEEEEEEECTTSHHHHHTSSTTEEEEEETTEESTTSBHHH--------HHHHHH
T ss_pred CCCcCEEEEecCC-C---------CcCCEEEEEEeCCChHHhcccchhhhhheeCCEeCCCCCHHH--------HHHHHH
Confidence 4568999887621 0 025899999999999999 99999999999999999886543 112222
Q ss_pred ccCCCCEEEEEEE
Q 017471 251 QKYTGDSAAVKVL 263 (371)
Q Consensus 251 ~~~~g~~v~l~v~ 263 (371)
. .+.+++|+|+
T Consensus 71 -~-~~~~v~L~V~ 81 (81)
T PF00595_consen 71 -S-ASNPVTLTVQ 81 (81)
T ss_dssp -H-STSEEEEEEE
T ss_pred -C-CCCcEEEEEC
Confidence 2 3448888774
No 30
>PRK10139 serine endoprotease; Provisional
Probab=98.47 E-value=3e-07 Score=93.11 Aligned_cols=64 Identities=16% Similarity=0.323 Sum_probs=56.0
Q ss_pred CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEE
Q 017471 198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNI 273 (371)
Q Consensus 198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v 273 (371)
.|++|.+|.++|||++ |||+||+|++|||++|.++.++. +.+.+ .+ +++.++|.|+|+...+.+
T Consensus 390 ~Gv~V~~V~~~spA~~aGL~~GD~I~~Ing~~v~~~~~~~----------~~l~~-~~-~~v~l~v~R~g~~~~~~~ 454 (455)
T PRK10139 390 KGIKIDEVVKGSPAAQAGLQKDDVIIGVNRDRVNSIAEMR----------KVLAA-KP-AIIALQIVRGNESIYLLL 454 (455)
T ss_pred CceEEEEeCCCChHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHh-CC-CeEEEEEEECCEEEEEEe
Confidence 5899999999999999 99999999999999999999875 56654 33 689999999999877665
No 31
>PLN00049 carboxyl-terminal processing protease; Provisional
Probab=98.46 E-value=5.9e-07 Score=89.33 Aligned_cols=69 Identities=20% Similarity=0.302 Sum_probs=54.9
Q ss_pred CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEe
Q 017471 198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITL 275 (371)
Q Consensus 198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l 275 (371)
.|++|..|.++|||++ ||++||+|++|||++|.++.... +...+ ....|+++.++|.|+|+..+++++-
T Consensus 102 ~g~~V~~V~~~SPA~~aGl~~GD~Iv~InG~~v~~~~~~~--------~~~~l-~g~~g~~v~ltv~r~g~~~~~~l~r 171 (389)
T PLN00049 102 AGLVVVAPAPGGPAARAGIRPGDVILAIDGTSTEGLSLYE--------AADRL-QGPEGSSVELTLRRGPETRLVTLTR 171 (389)
T ss_pred CcEEEEEeCCCChHHHcCCCCCCEEEEECCEECCCCCHHH--------HHHHH-hcCCCCEEEEEEEECCEEEEEEEEe
Confidence 3899999999999999 99999999999999998753211 11233 3457899999999999887776654
No 32
>TIGR02860 spore_IV_B stage IV sporulation protein B. SpoIVB, the stage IV sporulation protein B of endospore-forming bacteria such as Bacillus subtilis, is a serine proteinase, expressed in the spore (rather than mother cell) compartment, that participates in a proteolytic activation cascade for Sigma-K. It appears to be universal among endospore-forming bacteria and occurs nowhere else.
Probab=98.39 E-value=5.1e-07 Score=88.86 Aligned_cols=69 Identities=25% Similarity=0.364 Sum_probs=57.3
Q ss_pred CCceEEEEEC--------CCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCE
Q 017471 197 QKGVRIRRVD--------PTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK 267 (371)
Q Consensus 197 ~~gv~V~~V~--------~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~ 267 (371)
.+||+|.... .+|||++ |||+||+|++|||++|++++|+. +.+... .++++.+++.|+|+
T Consensus 104 t~GVlVvg~~~v~~~~g~~~SPAa~AGLq~GDiIvsING~~V~s~~DL~----------~iL~~~-~g~~V~LtV~R~Ge 172 (402)
T TIGR02860 104 TKGVLVVGFSDIETEKGKIHSPGEEAGIQIGDRILKINGEKIKNMDDLA----------NLINKA-GGEKLTLTIERGGK 172 (402)
T ss_pred cCEEEEEEEEcccccCCCCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHHHhC-CCCeEEEEEEECCE
Confidence 4689886653 2589998 99999999999999999999875 555544 58899999999999
Q ss_pred EEEEEEEec
Q 017471 268 ILNFNITLA 276 (371)
Q Consensus 268 ~~~~~v~l~ 276 (371)
..++++++.
T Consensus 173 ~~tv~V~Pv 181 (402)
T TIGR02860 173 IIETVIKPV 181 (402)
T ss_pred EEEEEEEEe
Confidence 998888754
No 33
>PRK10942 serine endoprotease; Provisional
Probab=98.36 E-value=6.3e-07 Score=91.20 Aligned_cols=64 Identities=20% Similarity=0.351 Sum_probs=55.7
Q ss_pred CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEE
Q 017471 198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNI 273 (371)
Q Consensus 198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v 273 (371)
.|++|.+|.++|||++ ||++||+|++|||++|.++.++. +.+.. .+ +.+.++|.|+|+.+.+.+
T Consensus 408 ~gvvV~~V~~~S~A~~aGL~~GDvIv~VNg~~V~s~~dl~----------~~l~~-~~-~~v~l~V~R~g~~~~v~~ 472 (473)
T PRK10942 408 KGVVVDNVKPGTPAAQIGLKKGDVIIGANQQPVKNIAELR----------KILDS-KP-SVLALNIQRGDSSIYLLM 472 (473)
T ss_pred CCeEEEEeCCCChHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHh-CC-CeEEEEEEECCEEEEEEe
Confidence 5899999999999999 99999999999999999999875 55554 33 789999999999877654
No 34
>cd00992 PDZ_signaling PDZ domain found in a variety of Eumetazoan signaling molecules, often in tandem arrangements. May be responsible for specific protein-protein interactions, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of PDZ domains an N-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in proteases.
Probab=98.35 E-value=8.7e-07 Score=67.65 Aligned_cols=52 Identities=23% Similarity=0.427 Sum_probs=42.1
Q ss_pred cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEec--CCCCc
Q 017471 173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIA--NDGTV 235 (371)
Q Consensus 173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~--~~~~l 235 (371)
..+|+.+....+ ...|++|.+|.++|||++ ||++||+|++|||+++. +..++
T Consensus 12 ~~~G~~~~~~~~-----------~~~~~~V~~v~~~s~a~~~gl~~GD~I~~ing~~i~~~~~~~~ 66 (82)
T cd00992 12 GGLGFSLRGGKD-----------SGGGIFVSRVEPGGPAERGGLRVGDRILEVNGVSVEGLTHEEA 66 (82)
T ss_pred CCcCEEEeCccc-----------CCCCeEEEEECCCChHHhCCCCCCCEEEEECCEEcCccCHHHH
Confidence 458888875511 135899999999999999 99999999999999998 44443
No 35
>PF14685 Tricorn_PDZ: Tricorn protease PDZ domain; PDB: 1N6F_D 1N6D_C 1N6E_C 1K32_A.
Probab=98.34 E-value=4.3e-06 Score=65.17 Aligned_cols=65 Identities=22% Similarity=0.330 Sum_probs=44.8
Q ss_pred CCceEEEEECCC--------Ccccc-C--CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEEC
Q 017471 197 QKGVRIRRVDPT--------APESE-V--LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD 265 (371)
Q Consensus 197 ~~gv~V~~V~~~--------spA~~-G--L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~ 265 (371)
..+..|.+|.++ ||..+ | +++||+|++|||++++...++. .+...+.|+++.|+|.+.
T Consensus 11 ~~~y~I~~I~~gd~~~~~~~sPL~~pGv~v~~GD~I~aInG~~v~~~~~~~-----------~lL~~~agk~V~Ltv~~~ 79 (88)
T PF14685_consen 11 NGGYRIARIYPGDPWNPNARSPLAQPGVDVREGDYILAINGQPVTADANPY-----------RLLEGKAGKQVLLTVNRK 79 (88)
T ss_dssp TTEEEEEEE-BS-TTSSS-B-GGGGGS----TT-EEEEETTEE-BTTB-HH-----------HHHHTTTTSEEEEEEE-S
T ss_pred CCEEEEEEEeCCCCCCccccCCccCCCCCCCCCCEEEEECCEECCCCCCHH-----------HHhcccCCCEEEEEEecC
Confidence 367889999875 78877 6 5699999999999999888763 455567899999999997
Q ss_pred C-EEEEEE
Q 017471 266 S-KILNFN 272 (371)
Q Consensus 266 g-~~~~~~ 272 (371)
+ +.+++.
T Consensus 80 ~~~~R~v~ 87 (88)
T PF14685_consen 80 PGGARTVV 87 (88)
T ss_dssp TT-EEEEE
T ss_pred CCCceEEE
Confidence 6 455554
No 36
>COG0793 Prc Periplasmic protease [Cell envelope biogenesis, outer membrane]
Probab=98.33 E-value=1.5e-06 Score=86.73 Aligned_cols=83 Identities=24% Similarity=0.390 Sum_probs=63.6
Q ss_pred ccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCC--ccccccccchhhhh
Q 017471 172 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGT--VPFRHGERIGFSYL 248 (371)
Q Consensus 172 ~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~--l~~~~~~~~~~~~~ 248 (371)
+..+|++++.- +..++.|.++.+++||++ ||++||+|++|||+++....- . -.
T Consensus 99 ~~GiG~~i~~~-------------~~~~~~V~s~~~~~PA~kagi~~GD~I~~IdG~~~~~~~~~~a-----------v~ 154 (406)
T COG0793 99 FGGIGIELQME-------------DIGGVKVVSPIDGSPAAKAGIKPGDVIIKIDGKSVGGVSLDEA-----------VK 154 (406)
T ss_pred ccceeEEEEEe-------------cCCCcEEEecCCCChHHHcCCCCCCEEEEECCEEccCCCHHHH-----------HH
Confidence 56688888754 126899999999999999 999999999999999987642 1 12
Q ss_pred hhccCCCCEEEEEEEECCEEEEEEEEeccc
Q 017471 249 VSQKYTGDSAAVKVLRDSKILNFNITLATH 278 (371)
Q Consensus 249 ~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~ 278 (371)
..+..+|..|+|++.|.+....+.+++.+.
T Consensus 155 ~irG~~Gt~V~L~i~r~~~~k~~~v~l~Re 184 (406)
T COG0793 155 LIRGKPGTKVTLTILRAGGGKPFTVTLTRE 184 (406)
T ss_pred HhCCCCCCeEEEEEEEcCCCceeEEEEEEE
Confidence 334578999999999985454555555543
No 37
>COG3480 SdrC Predicted secreted protein containing a PDZ domain [Signal transduction mechanisms]
Probab=98.20 E-value=4e-06 Score=78.87 Aligned_cols=72 Identities=25% Similarity=0.295 Sum_probs=65.4
Q ss_pred CCceEEEEECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEE-CCEEEEEEEEe
Q 017471 197 QKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLR-DSKILNFNITL 275 (371)
Q Consensus 197 ~~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R-~g~~~~~~v~l 275 (371)
..||++..+..+||+..-|++||.|++|||+++.+.+++. ..+...++|++|++++.| +++....++++
T Consensus 129 y~gvyv~~v~~~~~~~gkl~~gD~i~avdg~~f~s~~e~i----------~~v~~~k~Gd~VtI~~~r~~~~~~~~~~tl 198 (342)
T COG3480 129 YAGVYVLSVIDNSPFKGKLEAGDTIIAVDGEPFTSSDELI----------DYVSSKKPGDEVTIDYERHNETPEIVTITL 198 (342)
T ss_pred EeeEEEEEccCCcchhceeccCCeEEeeCCeecCCHHHHH----------HHHhccCCCCeEEEEEEeccCCCceEEEEE
Confidence 3699999999999998789999999999999999999976 788888999999999997 88888888888
Q ss_pred ccc
Q 017471 276 ATH 278 (371)
Q Consensus 276 ~~~ 278 (371)
...
T Consensus 199 ~~~ 201 (342)
T COG3480 199 IKN 201 (342)
T ss_pred Eee
Confidence 776
No 38
>PRK09681 putative type II secretion protein GspC; Provisional
Probab=98.05 E-value=9.4e-06 Score=76.16 Aligned_cols=68 Identities=25% Similarity=0.314 Sum_probs=55.2
Q ss_pred CceEEEEECCCCcc---cc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEE
Q 017471 198 KGVRIRRVDPTAPE---SE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNI 273 (371)
Q Consensus 198 ~gv~V~~V~~~spA---~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v 273 (371)
.|+.=-++.|+..+ .+ |||+||++++|||.++++.++.. .++........++++|+|||+..++.+
T Consensus 204 ~Gl~GYrl~Pgkd~~lF~~~GLq~GDva~sING~dL~D~~qa~----------~l~~~L~~~tei~ltVeRdGq~~~i~i 273 (276)
T PRK09681 204 EGIVGYAVKPGADRSLFDASGFKEGDIAIALNQQDFTDPRAMI----------ALMRQLPSMDSIQLTVLRKGARHDISI 273 (276)
T ss_pred CCceEEEECCCCcHHHHHHcCCCCCCEEEEeCCeeCCCHHHHH----------HHHHHhccCCeEEEEEEECCEEEEEEE
Confidence 35333467787544 45 99999999999999999888754 677777888999999999999999988
Q ss_pred Ee
Q 017471 274 TL 275 (371)
Q Consensus 274 ~l 275 (371)
.+
T Consensus 274 ~l 275 (276)
T PRK09681 274 AL 275 (276)
T ss_pred Ec
Confidence 75
No 39
>PF04495 GRASP55_65: GRASP55/65 PDZ-like domain ; InterPro: IPR007583 GRASP55 (Golgi reassembly stacking protein of 55 kDa) and GRASP65 (a 65 kDa) protein are highly homologous. GRASP55 is a component of the Golgi stacking machinery. GRASP65, an N-ethylmaleimide-sensitive membrane protein required for the stacking of Golgi cisternae in a cell-free system [].; PDB: 3RLE_A 4EDJ_A.
Probab=97.96 E-value=1.1e-05 Score=68.34 Aligned_cols=86 Identities=23% Similarity=0.331 Sum_probs=54.7
Q ss_pred ccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCC-CCEEEEECCEEecCCCCccccccccchhhhhh
Q 017471 172 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKP-SDIILSFDGIDIANDGTVPFRHGERIGFSYLV 249 (371)
Q Consensus 172 ~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~-GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~ 249 (371)
.+.||++++--.. +. ....+..|.+|.|+|||++ ||++ .|.|+.+|+..+++.+++. +.+
T Consensus 25 ~g~LG~sv~~~~~-~~-------~~~~~~~Vl~V~p~SPA~~AGL~p~~DyIig~~~~~l~~~~~l~----------~~v 86 (138)
T PF04495_consen 25 QGLLGISVRFESF-EG-------AEEEGWHVLRVAPNSPAAKAGLEPFFDYIIGIDGGLLDDEDDLF----------ELV 86 (138)
T ss_dssp SSSS-EEEEEEE--TT-------GCCCEEEEEEE-TTSHHHHTT--TTTEEEEEETTCE--STCHHH----------HHH
T ss_pred CCCCcEEEEEecc-cc-------cccceEEEeEecCCCHHHHCCccccccEEEEccceecCCHHHHH----------HHH
Confidence 4678887764411 10 1246899999999999999 9999 6999999999998776653 444
Q ss_pred hccCCCCEEEEEEEECC--EEEEEEEEec
Q 017471 250 SQKYTGDSAAVKVLRDS--KILNFNITLA 276 (371)
Q Consensus 250 ~~~~~g~~v~l~v~R~g--~~~~~~v~l~ 276 (371)
. .+.++++.+.|+... +.+++++++.
T Consensus 87 ~-~~~~~~l~L~Vyns~~~~vR~V~i~P~ 114 (138)
T PF04495_consen 87 E-ANENKPLQLYVYNSKTDSVREVTITPS 114 (138)
T ss_dssp H-HTTTS-EEEEEEETTTTCEEEEEE---
T ss_pred H-HcCCCcEEEEEEECCCCeEEEEEEEcC
Confidence 4 567889999999754 4455555543
No 40
>PF00089 Trypsin: Trypsin; InterPro: IPR001254 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine proteases belong to the MEROPS peptidase family S1 (chymotrypsin family, clan PA(S))and to peptidase family S6 (Hap serine peptidases). The chymotrypsin family is almost totally confined to animals, although trypsin-like enzymes are found in actinomycetes of the genera Streptomyces and Saccharopolyspora, and in the fungus Fusarium oxysporum []. The enzymes are inherently secreted, being synthesised with a signal peptide that targets them to the secretory pathway. Animal enzymes are either secreted directly, packaged into vesicles for regulated secretion, or are retained in leukocyte granules []. The Hap family, 'Haemophilus adhesion and penetration', are proteins that play a role in the interaction with human epithelial cells. The serine protease activity is localized at the N-terminal domain, whereas the binding domain is in the C-terminal region. ; GO: 0004252 serine-type endopeptidase activity, 0006508 proteolysis; PDB: 1SPJ_A 1A5I_A 2ZGH_A 2ZKS_A 2ZGJ_A 2ZGC_A 2ODP_A 2I6Q_A 2I6S_A 2ODQ_A ....
Probab=97.91 E-value=6.4e-05 Score=67.39 Aligned_cols=120 Identities=16% Similarity=0.165 Sum_probs=74.1
Q ss_pred CCCEEEEEEecC-CCcCCccceecCCCC---CCCCeEEEEEeCCCCCCc---eeeeeEEeeeeee---eccCCCeEEeEE
Q 017471 40 ICIYTMLTVEDD-EFWEGVLPVEFGELP---ALQDAVTVVGYPIGGDTI---SVTSGVVSRIEIL---SYVHGSTELLGL 109 (371)
Q Consensus 40 ~~DlAlLkv~~~-~~~~~l~~~~l~~s~---~lgd~V~~iG~p~g~~~~---s~t~G~Vs~~~~~---~~~~~~~~~~~i 109 (371)
.+|+|||+++.+ .+.+.+.++.+.... ..++.+.++|++...... ......+.-+... ...........+
T Consensus 86 ~~DiAll~L~~~~~~~~~~~~~~l~~~~~~~~~~~~~~~~G~~~~~~~~~~~~~~~~~~~~~~~~~c~~~~~~~~~~~~~ 165 (220)
T PF00089_consen 86 DNDIALLKLDRPITFGDNIQPICLPSAGSDPNVGTSCIVVGWGRTSDNGYSSNLQSVTVPVVSRKTCRSSYNDNLTPNMI 165 (220)
T ss_dssp TTSEEEEEESSSSEHBSSBEESBBTSTTHTTTTTSEEEEEESSBSSTTSBTSBEEEEEEEEEEHHHHHHHTTTTSTTTEE
T ss_pred cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccc
Confidence 479999999987 333466777777632 567999999998753321 3343444333321 111111111234
Q ss_pred EEEe----eccCCCCCCceecCCCcEEEEEeeecccCCccceeccccCcchhHh
Q 017471 110 QIDA----AINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHF 159 (371)
Q Consensus 110 ~~da----~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~~~~~~~aiP~~~i~~~ 159 (371)
..+. ....|+||||+++.++.++||.+.....+......+..+++...++
T Consensus 166 c~~~~~~~~~~~g~sG~pl~~~~~~lvGI~s~~~~c~~~~~~~v~~~v~~~~~W 219 (220)
T PF00089_consen 166 CAGSSGSGDACQGDSGGPLICNNNYLVGIVSFGENCGSPNYPGVYTRVSSYLDW 219 (220)
T ss_dssp EEETTSSSBGGTTTTTSEEEETTEEEEEEEEEESSSSBTTSEEEEEEGGGGHHH
T ss_pred cccccccccccccccccccccceeeecceeeecCCCCCCCcCEEEEEHHHhhcc
Confidence 4444 6789999999999888899999876432222234666777665554
No 41
>PRK11186 carboxy-terminal protease; Provisional
Probab=97.82 E-value=5.7e-05 Score=79.50 Aligned_cols=71 Identities=15% Similarity=0.135 Sum_probs=49.5
Q ss_pred CceEEEEECCCCcccc--CCCCCCEEEEEC--CEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEEC---CEEEE
Q 017471 198 KGVRIRRVDPTAPESE--VLKPSDIILSFD--GIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD---SKILN 270 (371)
Q Consensus 198 ~gv~V~~V~~~spA~~--GL~~GDvIl~vn--G~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~---g~~~~ 270 (371)
.+++|.+|.|||||++ ||++||+|++|| |.++.+...+. +.--..+.+..+|.+|.|+|.|+ ++..+
T Consensus 255 ~~~~V~~vipGsPA~ka~gLk~GD~IlaVn~~g~~~~dv~g~~------~~~vv~lirG~~Gt~V~LtV~r~~~~~~~~~ 328 (667)
T PRK11186 255 DYTVINSLVAGGPAAKSKKLSVGDKIVGVGQDGKPIVDVIGWR------LDDVVALIKGPKGSKVRLEILPAGKGTKTRI 328 (667)
T ss_pred CeEEEEEccCCChHHHhCCCCCCCEEEEECCCCCcccccccCC------HHHHHHHhcCCCCCEEEEEEEeCCCCCceEE
Confidence 4689999999999997 899999999999 56554433221 00012334456899999999994 44555
Q ss_pred EEEE
Q 017471 271 FNIT 274 (371)
Q Consensus 271 ~~v~ 274 (371)
++++
T Consensus 329 vtl~ 332 (667)
T PRK11186 329 VTLT 332 (667)
T ss_pred EEEE
Confidence 5544
No 42
>COG3975 Predicted protease with the C-terminal PDZ domain [General function prediction only]
Probab=97.81 E-value=4.3e-05 Score=76.50 Aligned_cols=86 Identities=21% Similarity=0.364 Sum_probs=66.4
Q ss_pred cceeeeEcCCHHHHHhccCCC--CCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc
Q 017471 175 LGVEWQKMENPDLRVAMSMKA--DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ 251 (371)
Q Consensus 175 lGi~~~~~~~~~~~~~~gl~~--~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~ 251 (371)
.|+.+.+...+ +-.+|++- +..+.+|+.|.++|||++ ||.+||.|++|||. + ..+.+
T Consensus 439 ~gL~~~~~~~~--~~~LGl~v~~~~g~~~i~~V~~~gPA~~AGl~~Gd~ivai~G~---s---------------~~l~~ 498 (558)
T COG3975 439 FGLTFTPKPRE--AYYLGLKVKSEGGHEKITFVFPGGPAYKAGLSPGDKIVAINGI---S---------------DQLDR 498 (558)
T ss_pred cceEEEecCCC--CcccceEecccCCeeEEEecCCCChhHhccCCCccEEEEEcCc---c---------------ccccc
Confidence 46777665222 34566543 345689999999999999 99999999999999 1 13445
Q ss_pred cCCCCEEEEEEEECCEEEEEEEEeccccc
Q 017471 252 KYTGDSAAVKVLRDSKILNFNITLATHRR 280 (371)
Q Consensus 252 ~~~g~~v~l~v~R~g~~~~~~v~l~~~~~ 280 (371)
.+.++.+++++.|.|+.+++.+++.....
T Consensus 499 ~~~~d~i~v~~~~~~~L~e~~v~~~~~~~ 527 (558)
T COG3975 499 YKVNDKIQVHVFREGRLREFLVKLGGDPT 527 (558)
T ss_pred cccccceEEEEccCCceEEeecccCCCcc
Confidence 67899999999999999999988876543
No 43
>KOG3129 consensus 26S proteasome regulatory complex, subunit PSMD9 [Posttranslational modification, protein turnover, chaperones]
Probab=97.72 E-value=7.7e-05 Score=66.26 Aligned_cols=73 Identities=22% Similarity=0.216 Sum_probs=59.1
Q ss_pred ceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEecc
Q 017471 199 GVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLAT 277 (371)
Q Consensus 199 gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~ 277 (371)
-++|.+|.|+|||+. ||+.||.|+++....-.++..++ =-..+.+...++.+.+++.|.|+...+.++++.
T Consensus 140 Fa~V~sV~~~SPA~~aGl~~gD~il~fGnV~sgn~~~lq--------~i~~~v~~~e~~~v~v~v~R~g~~v~L~ltP~~ 211 (231)
T KOG3129|consen 140 FAVVDSVVPGSPADEAGLCVGDEILKFGNVHSGNFLPLQ--------NIAAVVQSNEDQIVSVTVIREGQKVVLSLTPKK 211 (231)
T ss_pred eEEEeecCCCChhhhhCcccCceEEEecccccccchhHH--------HHHHHHHhccCcceeEEEecCCCEEEEEeCccc
Confidence 468999999999999 99999999999887776666543 012444567889999999999999999988875
Q ss_pred cc
Q 017471 278 HR 279 (371)
Q Consensus 278 ~~ 279 (371)
..
T Consensus 212 W~ 213 (231)
T KOG3129|consen 212 WQ 213 (231)
T ss_pred cc
Confidence 43
No 44
>KOG3553 consensus Tax interaction protein TIP1 [Cell wall/membrane/envelope biogenesis]
Probab=97.37 E-value=0.00017 Score=56.58 Aligned_cols=35 Identities=31% Similarity=0.484 Sum_probs=32.2
Q ss_pred CCceEEEEECCCCcccc-CCCCCCEEEEECCEEecC
Q 017471 197 QKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAN 231 (371)
Q Consensus 197 ~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~ 231 (371)
..|++|++|.++|||+. ||+.+|.|+.+||...+-
T Consensus 58 D~GiYvT~V~eGsPA~~AGLrihDKIlQvNG~DfTM 93 (124)
T KOG3553|consen 58 DKGIYVTRVSEGSPAEIAGLRIHDKILQVNGWDFTM 93 (124)
T ss_pred CccEEEEEeccCChhhhhcceecceEEEecCceeEE
Confidence 57999999999999999 999999999999987653
No 45
>PF12812 PDZ_1: PDZ-like domain
Probab=97.32 E-value=0.00033 Score=53.41 Aligned_cols=60 Identities=12% Similarity=0.057 Sum_probs=51.8
Q ss_pred cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCcc
Q 017471 173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVP 236 (371)
Q Consensus 173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~ 236 (371)
-|.|..++++ +-+.++.++++- |+++.....++++.. |+.+|-+|.+|||+++.+.+++.
T Consensus 9 ~~~Ga~f~~L-s~q~aR~~~~~~---~gv~v~~~~g~~~~~~~i~~g~iI~~Vn~kpt~~Ld~f~ 69 (78)
T PF12812_consen 9 EVCGAVFHDL-SYQQARQYGIPV---GGVYVAVSGGSLAFAGGISKGFIITSVNGKPTPDLDDFI 69 (78)
T ss_pred EEcCeecccC-CHHHHHHhCCCC---CEEEEEecCCChhhhCCCCCCeEEEeECCcCCcCHHHHH
Confidence 4789999999 899999999973 355666788999988 69999999999999999998864
No 46
>COG3031 PulC Type II secretory pathway, component PulC [Intracellular trafficking and secretion]
Probab=97.30 E-value=0.00019 Score=65.11 Aligned_cols=66 Identities=18% Similarity=0.223 Sum_probs=52.4
Q ss_pred ceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEE
Q 017471 199 GVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT 274 (371)
Q Consensus 199 gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~ 274 (371)
|..+.-..+++..++ |||.||+.+++|+..+++.+++. .++.....-+.+.++|.|+|++..+.|.
T Consensus 208 Gyr~~pgkd~slF~~sglq~GDIavaiNnldltdp~~m~----------~llq~l~~m~s~qlTv~R~G~rhdInV~ 274 (275)
T COG3031 208 GYRFEPGKDGSLFYKSGLQRGDIAVAINNLDLTDPEDMF----------RLLQMLRNMPSLQLTVIRRGKRHDINVR 274 (275)
T ss_pred EEEecCCCCcchhhhhcCCCcceEEEecCcccCCHHHHH----------HHHHhhhcCcceEEEEEecCccceeeec
Confidence 333333344566677 99999999999999999999864 6676666667899999999999988875
No 47
>COG3591 V8-like Glu-specific endopeptidase [Amino acid transport and metabolism]
Probab=96.98 E-value=0.0036 Score=58.13 Aligned_cols=87 Identities=28% Similarity=0.414 Sum_probs=57.0
Q ss_pred CCCCeEEEEEeCCCCCC---ceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccCC
Q 017471 67 ALQDAVTVVGYPIGGDT---ISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHED 143 (371)
Q Consensus 67 ~lgd~V~~iG~p~g~~~---~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~ 143 (371)
.+++.+-++|||.+... .-...+.|-.+.. ..++.|+.+.||+||.|+++.+.++||+.+......+
T Consensus 159 ~~~d~i~v~GYP~dk~~~~~~~e~t~~v~~~~~----------~~l~y~~dT~pG~SGSpv~~~~~~vigv~~~g~~~~~ 228 (251)
T COG3591 159 KANDRITVIGYPGDKPNIGTMWESTGKVNSIKG----------NKLFYDADTLPGSSGSPVLISKDEVIGVHYNGPGANG 228 (251)
T ss_pred ccCceeEEEeccCCCCcceeEeeecceeEEEec----------ceEEEEecccCCCCCCceEecCceEEEEEecCCCccc
Confidence 45688999999987552 2233444444331 2588999999999999999999999999987654333
Q ss_pred ccceeccc-cCcchhHhHhhh
Q 017471 144 VENIGYVI-PTPVIMHFIQDY 163 (371)
Q Consensus 144 ~~~~~~ai-P~~~i~~~l~~L 163 (371)
....++++ -...++++++++
T Consensus 229 ~~~~n~~vr~t~~~~~~I~~~ 249 (251)
T COG3591 229 GSLANNAVRLTPEILNFIQQN 249 (251)
T ss_pred ccccCcceEecHHHHHHHHHh
Confidence 33333332 233455555443
No 48
>PF00863 Peptidase_C4: Peptidase family C4 This family belongs to family C4 of the peptidase classification.; InterPro: IPR001730 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. Nuclear inclusion A (NIA) proteases from potyviruses are cysteine peptidases belong to the MEROPS peptidase family C4 (NIa protease family, clan PA(C)) [, ]. Potyviruses include plant viruses in which the single-stranded RNA encodes a polyprotein with NIA protease activity, where proteolytic cleavage is specific for Gln+Gly sites. The NIA protease acts on the polyprotein, releasing itself by Gln+Gly cleavage at both the N- and C-termini. It further processes the polyprotein by cleavage at five similar sites in the C-terminal half of the sequence. In addition to its C-terminal protease activity, the NIA protease contains an N-terminal domain that has been implicated in the transcription process []. This peptidase is present in the nuclear inclusion protein of potyviruses.; GO: 0008234 cysteine-type peptidase activity, 0006508 proteolysis; PDB: 3MMG_B 1Q31_B 1LVB_A 1LVM_A.
Probab=96.88 E-value=0.015 Score=53.58 Aligned_cols=112 Identities=19% Similarity=0.163 Sum_probs=53.3
Q ss_pred CCCEEEEEEecCCCcCCccceecC---CCCCCCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeecc
Q 017471 40 ICIYTMLTVEDDEFWEGVLPVEFG---ELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAIN 116 (371)
Q Consensus 40 ~~DlAlLkv~~~~~~~~l~~~~l~---~s~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~ 116 (371)
..|+.++|...+ +||.+-. ..|..+|.|..+|..+.....+.+...-|.+.. ...+. +..+-.+..
T Consensus 81 ~~DiviirmPkD-----fpPf~~kl~FR~P~~~e~v~mVg~~fq~k~~~s~vSesS~i~p--~~~~~----fWkHwIsTk 149 (235)
T PF00863_consen 81 GRDIVIIRMPKD-----FPPFPQKLKFRAPKEGERVCMVGSNFQEKSISSTVSESSWIYP--EENSH----FWKHWISTK 149 (235)
T ss_dssp CSSEEEEE--TT-----S----S---B----TT-EEEEEEEECSSCCCEEEEEEEEEEEE--ETTTT----EEEE-C---
T ss_pred CccEEEEeCCcc-----cCCcchhhhccCCCCCCEEEEEEEEEEcCCeeEEECCceEEee--cCCCC----eeEEEecCC
Confidence 589999998764 5554322 346678999999997754332333332222221 11122 344444556
Q ss_pred CCCCCCceec-CCCcEEEEEeeecccCCccceeccccCcchhHhHhhh
Q 017471 117 SGNSGGPAFN-DKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDY 163 (371)
Q Consensus 117 ~G~SGGPlvn-~~G~VIGI~~~~~~~~~~~~~~~aiP~~~i~~~l~~L 163 (371)
.|+-|.|+++ .+|.+|||.+.... ....++.-++|-+....+++..
T Consensus 150 ~G~CG~PlVs~~Dg~IVGiHsl~~~-~~~~N~F~~f~~~f~~~~l~~~ 196 (235)
T PF00863_consen 150 DGDCGLPLVSTKDGKIVGIHSLTSN-TSSRNYFTPFPDDFEEFYLENI 196 (235)
T ss_dssp TT-TT-EEEETTT--EEEEEEEEET-TTSSEEEEE--TTHHHHHCC-C
T ss_pred CCccCCcEEEcCCCcEEEEEcCccC-CCCeEEEEcCCHHHHHHHhccc
Confidence 7999999998 66999999876542 2344555556556555555443
No 49
>cd00190 Tryp_SPc Trypsin-like serine protease; Many of these are synthesized as inactive precursor zymogens that are cleaved during limited proteolysis to generate their active forms. Alignment contains also inactive enzymes that have substitutions of the catalytic triad residues.
Probab=96.33 E-value=0.03 Score=50.23 Aligned_cols=100 Identities=17% Similarity=0.145 Sum_probs=54.4
Q ss_pred CCCEEEEEEecCC-CcCCccceecCCC---CCCCCeEEEEEeCCCCCC----ceeeeeEEeeeeeee---ccC---C-Ce
Q 017471 40 ICIYTMLTVEDDE-FWEGVLPVEFGEL---PALQDAVTVVGYPIGGDT----ISVTSGVVSRIEILS---YVH---G-ST 104 (371)
Q Consensus 40 ~~DlAlLkv~~~~-~~~~l~~~~l~~s---~~lgd~V~~iG~p~g~~~----~s~t~G~Vs~~~~~~---~~~---~-~~ 104 (371)
.+|+|||+++.+. +...+.|+.+... ...++.+.+.|+...... .......+.-+.... ... . ..
T Consensus 88 ~~DiAll~L~~~~~~~~~v~picl~~~~~~~~~~~~~~~~G~g~~~~~~~~~~~~~~~~~~~~~~~~C~~~~~~~~~~~~ 167 (232)
T cd00190 88 DNDIALLKLKRPVTLSDNVRPICLPSSGYNLPAGTTCTVSGWGRTSEGGPLPDVLQEVNVPIVSNAECKRAYSYGGTITD 167 (232)
T ss_pred cCCEEEEEECCcccCCCcccceECCCccccCCCCCEEEEEeCCcCCCCCCCCceeeEEEeeeECHHHhhhhccCcccCCC
Confidence 4899999999653 2223677777654 245689999998654321 112222222221100 000 0 00
Q ss_pred EEeEEEE---EeeccCCCCCCceecCC---CcEEEEEeeec
Q 017471 105 ELLGLQI---DAAINSGNSGGPAFNDK---GKCVGIAFQSL 139 (371)
Q Consensus 105 ~~~~i~~---da~i~~G~SGGPlvn~~---G~VIGI~~~~~ 139 (371)
..-+... +...-+|+||||++... +.++||.+...
T Consensus 168 ~~~C~~~~~~~~~~c~gdsGgpl~~~~~~~~~lvGI~s~g~ 208 (232)
T cd00190 168 NMLCAGGLEGGKDACQGDSGGPLVCNDNGRGVLVGIVSWGS 208 (232)
T ss_pred ceEeeCCCCCCCccccCCCCCcEEEEeCCEEEEEEEEehhh
Confidence 0001111 23345699999999765 89999986543
No 50
>PF10459 Peptidase_S46: Peptidase S46; InterPro: IPR019500 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This entry represents S46 peptidases, where dipeptidyl-peptidase 7 (DPP-7) is the best-characterised member of this family. It is a serine peptidase that is located on the cell surface and is predicted to have two N-terminal transmembrane domains.
Probab=96.25 E-value=0.0035 Score=66.48 Aligned_cols=56 Identities=23% Similarity=0.391 Sum_probs=39.3
Q ss_pred EEEEEeeccCCCCCCceecCCCcEEEEEeeecccC--C----ccce--eccccCcchhHhHhhh
Q 017471 108 GLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE--D----VENI--GYVIPTPVIMHFIQDY 163 (371)
Q Consensus 108 ~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~--~----~~~~--~~aiP~~~i~~~l~~L 163 (371)
.+.++..|..||||+|++|.+||+||++|-...++ + .+.. +..|.+..+..+|+++
T Consensus 623 ~FlstnDitGGNSGSPvlN~~GeLVGl~FDgn~Esl~~D~~fdp~~~R~I~VDiRyvL~~ldkv 686 (698)
T PF10459_consen 623 NFLSTNDITGGNSGSPVLNAKGELVGLAFDGNWESLSGDIAFDPELNRTIHVDIRYVLWALDKV 686 (698)
T ss_pred EEEeccCcCCCCCCCccCCCCceEEEEeecCchhhcccccccccccceeEEEEHHHHHHHHHHH
Confidence 46788889999999999999999999998654332 1 1222 3445455566666654
No 51
>KOG3209 consensus WW domain-containing protein [General function prediction only]
Probab=96.04 E-value=0.0066 Score=62.87 Aligned_cols=55 Identities=27% Similarity=0.344 Sum_probs=42.9
Q ss_pred EEEECCCCcccc-C-CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECC
Q 017471 202 IRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDS 266 (371)
Q Consensus 202 V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g 266 (371)
|.+|.+||||++ | |+.||.|++|||+.|.+...-. .-.++ +..|-+|+|+|.-..
T Consensus 782 iGrIieGSPAdRCgkLkVGDrilAVNG~sI~~lsHad--------iv~LI--KdaGlsVtLtIip~e 838 (984)
T KOG3209|consen 782 IGRIIEGSPADRCGKLKVGDRILAVNGQSILNLSHAD--------IVSLI--KDAGLSVTLTIIPPE 838 (984)
T ss_pred ccccccCChhHhhccccccceEEEecCeeeeccCchh--------HHHHH--HhcCceEEEEEcChh
Confidence 778999999999 5 9999999999999998876532 00122 457889999987643
No 52
>KOG3580 consensus Tight junction proteins [Signal transduction mechanisms]
Probab=96.01 E-value=0.0043 Score=63.05 Aligned_cols=59 Identities=19% Similarity=0.286 Sum_probs=45.5
Q ss_pred CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEE
Q 017471 198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLR 264 (371)
Q Consensus 198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R 264 (371)
-|+.|..|.++|||++ ||+.||.||.||..+..+.--- +.+ ..+....+|+.+++.-.+
T Consensus 429 VGIFVaGvqegspA~~eGlqEGDQIL~VN~vdF~nl~RE-----eAV---lfLL~lPkGEevtilaQ~ 488 (1027)
T KOG3580|consen 429 VGIFVAGVQEGSPAEQEGLQEGDQILKVNTVDFRNLVRE-----EAV---LFLLELPKGEEVTILAQS 488 (1027)
T ss_pred eeEEEeecccCCchhhccccccceeEEeccccchhhhHH-----HHH---HHHhcCCCCcEEeehhhh
Confidence 4899999999999999 9999999999999988765211 111 245557789988886544
No 53
>smart00020 Tryp_SPc Trypsin-like serine protease. Many of these are synthesised as inactive precursor zymogens that are cleaved during limited proteolysis to generate their active forms. A few, however, are active as single chain molecules, and others are inactive due to substitutions of the catalytic triad residues.
Probab=95.86 E-value=0.042 Score=49.46 Aligned_cols=100 Identities=17% Similarity=0.137 Sum_probs=54.6
Q ss_pred CCCEEEEEEecCC-CcCCccceecCCC---CCCCCeEEEEEeCCCCCC-----ceeeeeEEeeeeeeecc---C-----C
Q 017471 40 ICIYTMLTVEDDE-FWEGVLPVEFGEL---PALQDAVTVVGYPIGGDT-----ISVTSGVVSRIEILSYV---H-----G 102 (371)
Q Consensus 40 ~~DlAlLkv~~~~-~~~~l~~~~l~~s---~~lgd~V~~iG~p~g~~~-----~s~t~G~Vs~~~~~~~~---~-----~ 102 (371)
.+|+|||+++.+. +.+.+.|+.+... ...++.+.+.|++..... .......+.-+...... . .
T Consensus 88 ~~DiAll~L~~~i~~~~~~~pi~l~~~~~~~~~~~~~~~~g~g~~~~~~~~~~~~~~~~~~~~~~~~~C~~~~~~~~~~~ 167 (229)
T smart00020 88 DNDIALLKLKSPVTLSDNVRPICLPSSNYNVPAGTTCTVSGWGRTSEGAGSLPDTLQEVNVPIVSNATCRRAYSGGGAIT 167 (229)
T ss_pred cCCEEEEEECcccCCCCceeeccCCCcccccCCCCEEEEEeCCCCCCCCCcCCCEeeEEEEEEeCHHHhhhhhccccccC
Confidence 5899999998763 2234667766653 345688999998765421 01122222222110000 0 0
Q ss_pred CeEEeEEE--EEeeccCCCCCCceecCCC--cEEEEEeeec
Q 017471 103 STELLGLQ--IDAAINSGNSGGPAFNDKG--KCVGIAFQSL 139 (371)
Q Consensus 103 ~~~~~~i~--~da~i~~G~SGGPlvn~~G--~VIGI~~~~~ 139 (371)
....-... .....-+|+||||++...+ .++||.+...
T Consensus 168 ~~~~C~~~~~~~~~~c~gdsG~pl~~~~~~~~l~Gi~s~g~ 208 (229)
T smart00020 168 DNMLCAGGLEGGKDACQGDSGGPLVCNDGRWVLVGIVSWGS 208 (229)
T ss_pred CCcEeecCCCCCCcccCCCCCCeeEEECCCEEEEEEEEECC
Confidence 00000001 1234557999999998765 9999987643
No 54
>KOG3550 consensus Receptor targeting protein Lin-7 [Extracellular structures]
Probab=95.52 E-value=0.028 Score=47.59 Aligned_cols=37 Identities=24% Similarity=0.424 Sum_probs=33.0
Q ss_pred CCceEEEEECCCCcccc--CCCCCCEEEEECCEEecCCC
Q 017471 197 QKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDG 233 (371)
Q Consensus 197 ~~gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~V~~~~ 233 (371)
..-++|+++.|++-|++ ||+.||.++++||..|....
T Consensus 114 nspiyisriipggvadrhgglkrgdqllsvngvsvege~ 152 (207)
T KOG3550|consen 114 NSPIYISRIIPGGVADRHGGLKRGDQLLSVNGVSVEGEH 152 (207)
T ss_pred CCceEEEeecCCccccccCcccccceeEeecceeecchh
Confidence 45699999999999998 89999999999999997543
No 55
>KOG3532 consensus Predicted protein kinase [General function prediction only]
Probab=95.43 E-value=0.02 Score=59.15 Aligned_cols=50 Identities=16% Similarity=0.458 Sum_probs=42.0
Q ss_pred ccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCcc
Q 017471 174 LLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVP 236 (371)
Q Consensus 174 ~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~ 236 (371)
-+|+.|..- ...-|.|..|.+++||.+ .+++||++++|||.||++..+..
T Consensus 387 ~ig~vf~~~-------------~~~~v~v~tv~~ns~a~k~~~~~gdvlvai~~~pi~s~~q~~ 437 (1051)
T KOG3532|consen 387 PIGLVFDKN-------------TNRAVKVCTVEDNSLADKAAFKPGDVLVAINNVPIRSERQAT 437 (1051)
T ss_pred ceeEEEecC-------------CceEEEEEEecCCChhhHhcCCCcceEEEecCccchhHHHHH
Confidence 477777643 245688999999999999 99999999999999999987764
No 56
>PF08192 Peptidase_S64: Peptidase family S64; InterPro: IPR012985 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This family of fungal proteins is involved in the processing of membrane bound transcription factor Stp1 [] and belongs to MEROPS petidase family S64 (clan PA). The processing causes the signalling domain of Stp1 to be passed to the nucleus where several permease genes are induced. The permeases are important for uptake of amino acids, and processing of tp1 only occurs in an amino acid-rich environment. This family is predicted to be distantly related to the trypsin family (MEROPS peptidase family S1) and to have a typical trypsin-like catalytic triad [].
Probab=95.33 E-value=0.17 Score=52.76 Aligned_cols=119 Identities=17% Similarity=0.285 Sum_probs=73.0
Q ss_pred cCCCCEEEEEEecCC-----CcCCc------cceecCCC--------CCCCCeEEEEEeCCCCCCceeeeeEEeeeeeee
Q 017471 38 ADICIYTMLTVEDDE-----FWEGV------LPVEFGEL--------PALQDAVTVVGYPIGGDTISVTSGVVSRIEILS 98 (371)
Q Consensus 38 d~~~DlAlLkv~~~~-----~~~~l------~~~~l~~s--------~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~ 98 (371)
....|+||++++... +.+++ |.+.+.+. ...|.+|+=+|+..+ .|.|+|.++....
T Consensus 540 ~~LsD~AIIkV~~~~~~~N~LGddi~f~~~dP~l~f~NlyV~~~~~~~~~G~~VfK~GrTTg-----yT~G~lNg~klvy 614 (695)
T PF08192_consen 540 KRLSDWAIIKVNKERKCQNYLGDDIQFNEPDPTLMFQNLYVREVVSNLVPGMEVFKVGRTTG-----YTTGILNGIKLVY 614 (695)
T ss_pred ccccceEEEEeCCCceecCCCCccccccCCCccccccccchhhhhhccCCCCeEEEecccCC-----ccceEecceEEEE
Confidence 345799999999653 22222 33344331 123678999988755 5667777765432
Q ss_pred ccCCCeE-EeEEEEE----eeccCCCCCCceecCCCc------EEEEEeeecccCCccceeccccCcchhHhHhhh
Q 017471 99 YVHGSTE-LLGLQID----AAINSGNSGGPAFNDKGK------CVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDY 163 (371)
Q Consensus 99 ~~~~~~~-~~~i~~d----a~i~~G~SGGPlvn~~G~------VIGI~~~~~~~~~~~~~~~aiP~~~i~~~l~~L 163 (371)
...+... .+++... +=...|+||.-+++.-++ |+||.++.- .....++++.|+..|.+-+++.
T Consensus 615 w~dG~i~s~efvV~s~~~~~Fa~~GDSGS~VLtk~~d~~~gLgvvGMlhsyd--ge~kqfglftPi~~il~rl~~v 688 (695)
T PF08192_consen 615 WADGKIQSSEFVVSSDNNPAFASGGDSGSWVLTKLEDNNKGLGVVGMLHSYD--GEQKQFGLFTPINEILDRLEEV 688 (695)
T ss_pred ecCCCeEEEEEEEecCCCccccCCCCcccEEEecccccccCceeeEEeeecC--CccceeeccCcHHHHHHHHHHh
Confidence 2233222 2344333 223479999999987555 999998743 2455788888887766555443
No 57
>PF00949 Peptidase_S7: Peptidase S7, Flavivirus NS3 serine protease ; InterPro: IPR001850 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This signature identifies serine peptidases belong to MEROPS peptidase family S7 (flavivirin family, clan PA(S)). The protein fold of the peptidase domain for members of this family resembles that of chymotrypsin, the type example for clan PA. Flaviviruses produce a polyprotein from the ssRNA genome. The N terminus of the NS3 protein (approx. 180 aa) is required for the processing of the polyprotein. NS3 also has conserved homology with NTP-binding proteins and DEAD family of RNA helicase [, , ].; GO: 0003723 RNA binding, 0003724 RNA helicase activity, 0005524 ATP binding; PDB: 2IJO_B 3E90_D 2GGV_B 2FP7_B 2WV9_A 3U1I_B 3U1J_B 2WZQ_A 2WHX_A 3L6P_A ....
Probab=95.14 E-value=0.025 Score=47.42 Aligned_cols=33 Identities=33% Similarity=0.511 Sum_probs=23.6
Q ss_pred EEEEEeeccCCCCCCceecCCCcEEEEEeeecc
Q 017471 108 GLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLK 140 (371)
Q Consensus 108 ~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~ 140 (371)
+...+....+|.||.|++|.+|+++|+......
T Consensus 87 ~~~~~~d~~~GsSGSpi~n~~g~ivGlYg~g~~ 119 (132)
T PF00949_consen 87 IGAIDLDFPKGSSGSPIFNQNGEIVGLYGNGVE 119 (132)
T ss_dssp EEEE---S-TTGTT-EEEETTSCEEEEEEEEEE
T ss_pred EEeeecccCCCCCCCceEcCCCcEEEEEcccee
Confidence 345566688999999999999999999877653
No 58
>PF00947 Pico_P2A: Picornavirus core protein 2A; InterPro: IPR000081 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This domain defines cysteine peptidases belong to MEROPS peptidase family C3 (picornain, clan PA(C)), subfamilies 3CA and 3CB. The protein fold of this peptidase domain for members of this family resembles that of the serine peptidase, chymotrypsin [], the type example for clan PA. Picornaviral proteins are expressed as a single polyprotein which is cleaved by the viral 3C cysteine protease []. The poliovirus polyprotein is selectively cleaved between the Gln-|-Gly bond. In other picornavirus reactions Glu may be substituted for Gln, and Ser or Thr for Gly. ; GO: 0008233 peptidase activity, 0006508 proteolysis, 0016032 viral reproduction; PDB: 2HRV_B 1Z8R_A.
Probab=94.86 E-value=0.16 Score=41.96 Aligned_cols=97 Identities=13% Similarity=-0.004 Sum_probs=53.2
Q ss_pred EEEecCCCCEEEEEEecCCCcCCccceecCCCCCCCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEe
Q 017471 34 TLVTADICIYTMLTVEDDEFWEGVLPVEFGELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDA 113 (371)
Q Consensus 34 vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da 113 (371)
.+..+...||+|.+.+... ...++.-+ -..-|+--=.-.-.-..+++.-..-.++..++.+.+...+.+.-..
T Consensus 13 ~v~~~~~rDL~V~~t~a~G----~D~I~~C~---Ct~GvYyCks~~k~yPV~~~~~~~~~i~~s~YYP~h~Q~~~l~g~G 85 (127)
T PF00947_consen 13 TVWEDYTRDLLVDRTTAHG----CDTIPRCD---CTTGVYYCKSKNKYYPVTVTGPTWYWIEESEYYPKHYQYNLLIGEG 85 (127)
T ss_dssp CCEEECCCTEEEEEECCEE----E--BB-------SEEEEEETTTTCEEEEEEEEECEEEE-SBTTB-SEEEECEEEEE-
T ss_pred ceehhhCCCEEEEecCCCC----CCcccCcc---CCCCEEEeeECCeEeeEEEeccceEEECCccCchhheecCceeecc
Confidence 4677889999999998763 22222221 0011111000000000122222233455556667777777888899
Q ss_pred eccCCCCCCceecCCCcEEEEEeee
Q 017471 114 AINSGNSGGPAFNDKGKCVGIAFQS 138 (371)
Q Consensus 114 ~i~~G~SGGPlvn~~G~VIGI~~~~ 138 (371)
+..||..||+|+-.. -||||.++.
T Consensus 86 p~~PGdCGg~L~C~H-GViGi~Tag 109 (127)
T PF00947_consen 86 PAEPGDCGGILRCKH-GVIGIVTAG 109 (127)
T ss_dssp SSSTT-TCSEEEETT-CEEEEEEEE
T ss_pred cCCCCCCCceeEeCC-CeEEEEEeC
Confidence 999999999999766 599999884
No 59
>KOG3834 consensus Golgi reassembly stacking protein GRASP65, contains PDZ domain [Intracellular trafficking, secretion, and vesicular transport]
Probab=94.82 E-value=0.05 Score=53.59 Aligned_cols=132 Identities=17% Similarity=0.246 Sum_probs=80.1
Q ss_pred CCceEEEEECCCCcccc-CCCC-CCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCE--EEEEE
Q 017471 197 QKGVRIRRVDPTAPESE-VLKP-SDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK--ILNFN 272 (371)
Q Consensus 197 ~~gv~V~~V~~~spA~~-GL~~-GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~--~~~~~ 272 (371)
..|.-|-+|.++|+|.+ ||.+ -|-|++|||..++...|.. +.+.+....+ |+++++--.. .+.++
T Consensus 14 teg~hvlkVqedSpa~~aglepffdFIvSI~g~rL~~dnd~L----------k~llk~~sek-Vkltv~n~kt~~~R~v~ 82 (462)
T KOG3834|consen 14 TEGYHVLKVQEDSPAHKAGLEPFFDFIVSINGIRLNKDNDTL----------KALLKANSEK-VKLTVYNSKTQEVRIVE 82 (462)
T ss_pred ceeEEEEEeecCChHHhcCcchhhhhhheeCcccccCchHHH----------HHHHHhcccc-eEEEEEecccceeEEEE
Confidence 46888999999999999 9888 5899999999999887753 3444433333 9999886432 33333
Q ss_pred EEeccccccCCCCCCCCCCCceeeccEEEechHHHHHHcCcccc-cCceEEEEEecCChHhHHhhh-hhhhHHHHHHHHH
Q 017471 273 ITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERI-MNMKLRSSFWTSSCIQCHNCQ-MSSLLWCLRCLWL 350 (371)
Q Consensus 273 v~l~~~~~~~~~~~~~~~~~~~~~~Gl~~~~l~~~~~~~~~~~~-~~~~~v~~~~~~Sp~~~~~~~-~~~~~~~~~~~~~ 350 (371)
|+........ ++|+.+.-. +...+ ..-.=+-+|.++|||++.... ..|.+= =-|-
T Consensus 83 I~ps~~wggq-------------llGvsvrFc-------sf~~A~~~vwHvl~V~p~SPaalAgl~~~~DYiv---G~~~ 139 (462)
T KOG3834|consen 83 IVPSNNWGGQ-------------LLGVSVRFC-------SFDGAVESVWHVLSVEPNSPAALAGLRPYTDYIV---GIWD 139 (462)
T ss_pred eccccccccc-------------ccceEEEec-------cCccchhheeeeeecCCCCHHHhcccccccceEe---cchh
Confidence 3332211100 356654211 11111 111225578999999998776 555432 1245
Q ss_pred HHhhhhhhHHHH
Q 017471 351 ILILDMRRLLTL 362 (371)
Q Consensus 351 ~~~~~~~~~~~~ 362 (371)
++-.+.+.+++|
T Consensus 140 ~~~~~~eDl~~l 151 (462)
T KOG3834|consen 140 AVMHEEEDLFTL 151 (462)
T ss_pred hhccchHHHHHH
Confidence 566666666554
No 60
>KOG3209 consensus WW domain-containing protein [General function prediction only]
Probab=94.70 E-value=0.044 Score=57.04 Aligned_cols=58 Identities=14% Similarity=0.161 Sum_probs=46.9
Q ss_pred CceEEEEECCCCcccc-C-CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECC
Q 017471 198 KGVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDS 266 (371)
Q Consensus 198 ~gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g 266 (371)
-+++|-++.+++||.+ | ++.||.|++|||++.++...- +++.-.+.|....+.++|.|
T Consensus 923 M~LfVLRlAeDGPA~rdGrm~VGDqi~eINGesTkgmtH~-----------rAIelIk~gg~~vll~Lr~g 982 (984)
T KOG3209|consen 923 MDLFVLRLAEDGPAIRDGRMRVGDQITEINGESTKGMTHD-----------RAIELIKQGGRRVLLLLRRG 982 (984)
T ss_pred cceEEEEeccCCCccccCceeecceEEEecCcccCCCcHH-----------HHHHHHHhCCeEEEEEeccC
Confidence 4699999999999999 6 999999999999999887653 35555556666667777765
No 61
>PF00548 Peptidase_C3: 3C cysteine protease (picornain 3C); InterPro: IPR000199 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This signature defines cysteine peptidases belong to MEROPS peptidase family C3 (picornain, clan PA(C)), subfamilies C3A and C3B. The protein fold of this peptidase domain for members of this family resembles that of the serine peptidase, chymotrypsin [], the type example for clan PA. Picornaviral proteins are expressed as a single polyprotein which is cleaved by the viral C3 cysteine protease. The poliovirus polyprotein is selectively cleaved between the Gln-|-Gly bond. In other picornavirus reactions Glu may be substituted for Gln, and Ser or Thr for Gly. ; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis; PDB: 3SJO_E 2H6M_A 1QA7_C 1HAV_B 2HAL_A 2H9H_A 3QZQ_B 3QZR_A 3R0F_B 3SJ9_A ....
Probab=94.28 E-value=0.67 Score=40.80 Aligned_cols=103 Identities=17% Similarity=0.281 Sum_probs=63.3
Q ss_pred EEEEecCC---CCEEEEEEecCCCcCC-ccceecCCCCCCCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeE
Q 017471 33 ATLVTADI---CIYTMLTVEDDEFWEG-VLPVEFGELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLG 108 (371)
Q Consensus 33 ~vv~~d~~---~DlAlLkv~~~~~~~~-l~~~~l~~s~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~ 108 (371)
.+.-+|.+ .|+++++++..+-+.+ .+.++ ...+...+.+.++-++ .........+.|+..+.. ...+......
T Consensus 61 ~~~lv~~~~~~~Dl~~v~l~~~~kfrDIrk~~~-~~~~~~~~~~l~v~~~-~~~~~~~~v~~v~~~~~i-~~~g~~~~~~ 137 (172)
T PF00548_consen 61 SVVLVDRDGVDTDLTLVKLPRNPKFRDIRKFFP-ESIPEYPECVLLVNST-KFPRMIVEVGFVTNFGFI-NLSGTTTPRS 137 (172)
T ss_dssp EEEEEETTSSEEEEEEEEEESSS-B--GGGGSB-SSGGTEEEEEEEEESS-SSTCEEEEEEEEEEEEEE-EETTEEEEEE
T ss_pred eEEEecCCCcceeEEEEEccCCcccCchhhhhc-cccccCCCcEEEEECC-CCccEEEEEEEEeecCcc-ccCCCEeeEE
Confidence 43445554 6999999976431222 22333 2222444666666544 333335566666666543 3334444567
Q ss_pred EEEEeeccCCCCCCceec---CCCcEEEEEeee
Q 017471 109 LQIDAAINSGNSGGPAFN---DKGKCVGIAFQS 138 (371)
Q Consensus 109 i~~da~i~~G~SGGPlvn---~~G~VIGI~~~~ 138 (371)
+..+++-.+|.-||||+. ..++++||..+.
T Consensus 138 ~~Y~~~t~~G~CG~~l~~~~~~~~~i~GiHvaG 170 (172)
T PF00548_consen 138 LKYKAPTKPGMCGSPLVSRIGGQGKIIGIHVAG 170 (172)
T ss_dssp EEEESEEETTGTTEEEEESCGGTTEEEEEEEEE
T ss_pred EEEccCCCCCccCCeEEEeeccCccEEEEEecc
Confidence 899999999999999985 358999998874
No 62
>KOG3605 consensus Beta amyloid precursor-binding protein [General function prediction only]
Probab=94.03 E-value=0.078 Score=54.71 Aligned_cols=102 Identities=21% Similarity=0.262 Sum_probs=68.3
Q ss_pred CCCCCce-----ecCCCcEEEEEeeecccCCccceeccccCcchhHhHhhhhhCCeeeccc---ccceeeeEcCCHHHHH
Q 017471 118 GNSGGPA-----FNDKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDYEKNGAYTGFP---LLGVEWQKMENPDLRV 189 (371)
Q Consensus 118 G~SGGPl-----vn~~G~VIGI~~~~~~~~~~~~~~~aiP~~~i~~~l~~L~~~g~~~g~~---~lGi~~~~~~~~~~~~ 189 (371)
=++|||. +|.--+++.||--.+ ..+|.+.+..++..+++.-.+. +. .--+.-..+..|+.+-
T Consensus 680 mm~~GpAarsgkLnIGDQiiaING~SL---------VGLPLstcQs~Ik~~KnQT~Vk-ltiV~cpPV~~V~I~RPd~ky 749 (829)
T KOG3605|consen 680 MMHGGPAARSGKLNIGDQIMSINGTSL---------VGLPLSTCQSIIKGLKNQTAVK-LNIVSCPPVTTVLIRRPDLRY 749 (829)
T ss_pred cccCChhhhcCCccccceeEeecCcee---------ccccHHHHHHHHhcccccceEE-EEEecCCCceEEEeecccchh
Confidence 3556664 444456666653222 4489999999888886544332 11 1112222233778888
Q ss_pred hccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecC
Q 017471 190 AMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAN 231 (371)
Q Consensus 190 ~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~ 231 (371)
.+|+. -..|+ |-+...|+-|++ |++.|-.|++|||+.|--
T Consensus 750 QLGFS-VQNGi-ICSLlRGGIAERGGVRVGHRIIEINgQSVVA 790 (829)
T KOG3605|consen 750 QLGFS-VQNGI-ICSLLRGGIAERGGVRVGHRIIEINGQSVVA 790 (829)
T ss_pred hccce-eeCcE-eehhhcccchhccCceeeeeEEEECCceEEe
Confidence 89997 34566 677889999999 999999999999999753
No 63
>KOG3542 consensus cAMP-regulated guanine nucleotide exchange factor [Signal transduction mechanisms]
Probab=93.91 E-value=0.05 Score=56.29 Aligned_cols=37 Identities=27% Similarity=0.339 Sum_probs=33.4
Q ss_pred CCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCC
Q 017471 197 QKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDG 233 (371)
Q Consensus 197 ~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~ 233 (371)
..|++|.+|.|++.|+. ||+.||.|++|||+...+..
T Consensus 561 GfgifV~~V~pgskAa~~GlKRgDqilEVNgQnfenis 598 (1283)
T KOG3542|consen 561 GFGIFVAEVFPGSKAAREGLKRGDQILEVNGQNFENIS 598 (1283)
T ss_pred cceeEEeeecCCchHHHhhhhhhhhhhhccccchhhhh
Confidence 45899999999999999 99999999999999887654
No 64
>KOG3651 consensus Protein kinase C, alpha binding protein [Signal transduction mechanisms]
Probab=93.78 E-value=0.089 Score=49.57 Aligned_cols=39 Identities=21% Similarity=0.318 Sum_probs=34.7
Q ss_pred CceEEEEECCCCcccc-C-CCCCCEEEEECCEEecCCCCcc
Q 017471 198 KGVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVP 236 (371)
Q Consensus 198 ~gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~ 236 (371)
.-++|.+|..++||++ | ++.||.|++|||..|+....+.
T Consensus 30 PClYiVQvFD~tPAa~dG~i~~GDEi~avNg~svKGktKve 70 (429)
T KOG3651|consen 30 PCLYIVQVFDKTPAAKDGRIRCGDEIVAVNGISVKGKTKVE 70 (429)
T ss_pred CeEEEEEeccCCchhccCccccCCeeEEecceeecCccHHH
Confidence 3589999999999999 6 9999999999999999876654
No 65
>KOG1892 consensus Actin filament-binding protein Afadin [Cytoskeleton]
Probab=93.76 E-value=0.082 Score=56.73 Aligned_cols=64 Identities=17% Similarity=0.212 Sum_probs=49.3
Q ss_pred CCCCCceEEEEECCCCcccc-C-CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCE
Q 017471 194 KADQKGVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK 267 (371)
Q Consensus 194 ~~~~~gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~ 267 (371)
-++.-|++|..|.+|++|+. | |+.||.+++|||+..-...+-. .+-...+.|..|.+.|...|.
T Consensus 956 Gq~klGIYvKsVV~GgaAd~DGRL~aGDQLLsVdG~SLiGisQEr----------AA~lmtrtg~vV~leVaKqgA 1021 (1629)
T KOG1892|consen 956 GQRKLGIYVKSVVEGGAADHDGRLEAGDQLLSVDGHSLIGISQER----------AARLMTRTGNVVHLEVAKQGA 1021 (1629)
T ss_pred CccccceEEEEeccCCccccccccccCceeeeecCcccccccHHH----------HHHHHhccCCeEEEehhhhhh
Confidence 33456999999999999998 6 9999999999999987766532 122224578899999877553
No 66
>COG0750 Predicted membrane-associated Zn-dependent proteases 1 [Cell envelope biogenesis, outer membrane]
Probab=93.71 E-value=0.11 Score=51.29 Aligned_cols=57 Identities=28% Similarity=0.424 Sum_probs=45.3
Q ss_pred EEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCE---EEEEEEE-CCEEE
Q 017471 202 IRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDS---AAVKVLR-DSKIL 269 (371)
Q Consensus 202 V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~---v~l~v~R-~g~~~ 269 (371)
+.++..+|+|+. |+++||.|+++|++++.+++++. +.. ....+.. +.+.+.| +++..
T Consensus 133 ~~~v~~~s~a~~a~l~~Gd~iv~~~~~~i~~~~~~~----------~~~-~~~~~~~~~~~~i~~~~~~~~~~ 194 (375)
T COG0750 133 VGEVAPKSAAALAGLRPGDRIVAVDGEKVASWDDVR----------RLL-VAAAGDVFNLLTILVIRLDGEAH 194 (375)
T ss_pred eeecCCCCHHHHcCCCCCCEEEeECCEEccCHHHHH----------HHH-HhccCCcccceEEEEEeccceee
Confidence 337899999999 99999999999999999998864 223 3334555 8899999 77763
No 67
>KOG2921 consensus Intramembrane metalloprotease (sterol-regulatory element-binding protein (SREBP) protease) [Posttranslational modification, protein turnover, chaperones]
Probab=93.55 E-value=0.051 Score=53.01 Aligned_cols=40 Identities=25% Similarity=0.341 Sum_probs=36.3
Q ss_pred CCCceEEEEECCCCcccc--CCCCCCEEEEECCEEecCCCCc
Q 017471 196 DQKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTV 235 (371)
Q Consensus 196 ~~~gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~V~~~~~l 235 (371)
...|+.|++|...||+.. ||++||+|+++||-+|.+.+|.
T Consensus 218 ~g~gV~Vtev~~~Spl~gprGL~vgdvitsldgcpV~~v~dW 259 (484)
T KOG2921|consen 218 HGEGVTVTEVPSVSPLFGPRGLSVGDVITSLDGCPVHKVSDW 259 (484)
T ss_pred cCceEEEEeccccCCCcCcccCCccceEEecCCcccCCHHHH
Confidence 357999999999999987 9999999999999999988774
No 68
>KOG3580 consensus Tight junction proteins [Signal transduction mechanisms]
Probab=93.44 E-value=0.092 Score=53.72 Aligned_cols=61 Identities=28% Similarity=0.464 Sum_probs=47.6
Q ss_pred CCceEEEEECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc-cCCCCEEEEEEEECCEE
Q 017471 197 QKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ-KYTGDSAAVKVLRDSKI 268 (371)
Q Consensus 197 ~~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~-~~~g~~v~l~v~R~g~~ 268 (371)
.+.++|++|.||+||+.-||.||.|+.|||.+..+.... | ++.+ .+.|+...++|.|..+.
T Consensus 39 etSiViSDVlpGGPAeG~LQenDrvvMVNGvsMenv~ha---------F--AvQqLrksgK~A~ItvkRprkv 100 (1027)
T KOG3580|consen 39 ETSIVISDVLPGGPAEGLLQENDRVVMVNGVSMENVLHA---------F--AVQQLRKSGKVAAITVKRPRKV 100 (1027)
T ss_pred ceeEEEeeccCCCCcccccccCCeEEEEcCcchhhhHHH---------H--HHHHHHhhccceeEEeccccee
Confidence 456999999999999878999999999999999876542 1 2222 34677888999886543
No 69
>KOG3552 consensus FERM domain protein FRM-8 [General function prediction only]
Probab=92.50 E-value=0.12 Score=55.25 Aligned_cols=57 Identities=26% Similarity=0.325 Sum_probs=42.3
Q ss_pred CceEEEEECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEE
Q 017471 198 KGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLR 264 (371)
Q Consensus 198 ~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R 264 (371)
.-|+|..|.+|+|+...|+|||.|++|||.+|+...-- +.| .++. ...+.|.|+|.+
T Consensus 75 rPviVr~VT~GGps~GKL~PGDQIl~vN~Epv~dapre-----rvI---dlvR--ace~sv~ltV~q 131 (1298)
T KOG3552|consen 75 RPVIVRFVTEGGPSIGKLQPGDQILAVNGEPVKDAPRE-----RVI---DLVR--ACESSVNLTVCQ 131 (1298)
T ss_pred CceEEEEecCCCCccccccCCCeEEEecCcccccccHH-----HHH---HHHH--HHhhhcceEEec
Confidence 45889999999999878999999999999999765321 111 1222 245678888887
No 70
>PF00944 Peptidase_S3: Alphavirus core protein ; InterPro: IPR000930 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. Togavirin, also known as Sindbis virus core endopeptidase, is a serine protease resident at the N terminus of the p130 polyprotein of togaviruses []. The endopeptidase signature identifies the peptidase as belonging to the MEROPS peptidase family S3 (togavirin family, clan PA(S)). The polyprotein also includes structural proteins for the nucleocapsid core and for the glycoprotein spikes []. Togavirin is only active while part of the polyprotein, cleavage at a Trp-Ser bond resulting in total lack of activity []. Mutagenesis studies have identified the location of the His-Asp-Ser catalytic triad, and X-ray studies have revealed the protein fold to be similar to that of chymotrypsin [, ].; GO: 0004252 serine-type endopeptidase activity, 0006508 proteolysis, 0016020 membrane; PDB: 2YEW_D 1EP5_A 3J0C_F 1EP6_C 1WYK_D 1DYL_A 1VCQ_B 1VCP_B 1LD4_D 1KXA_A ....
Probab=91.44 E-value=0.19 Score=41.88 Aligned_cols=26 Identities=31% Similarity=0.631 Sum_probs=22.5
Q ss_pred eccCCCCCCceecCCCcEEEEEeeec
Q 017471 114 AINSGNSGGPAFNDKGKCVGIAFQSL 139 (371)
Q Consensus 114 ~i~~G~SGGPlvn~~G~VIGI~~~~~ 139 (371)
.-.+|+||-|++|..|+||||+....
T Consensus 102 ~g~~GDSGRpi~DNsGrVVaIVLGG~ 127 (158)
T PF00944_consen 102 VGKPGDSGRPIFDNSGRVVAIVLGGA 127 (158)
T ss_dssp S-STTSTTEEEESTTSBEEEEEEEEE
T ss_pred CCCCCCCCCccCcCCCCEEEEEecCC
Confidence 45699999999999999999998765
No 71
>KOG3606 consensus Cell polarity protein PAR6 [Signal transduction mechanisms]
Probab=91.09 E-value=0.25 Score=45.95 Aligned_cols=81 Identities=19% Similarity=0.273 Sum_probs=52.0
Q ss_pred eeccccCcchhHhHhh--hhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-C-CCCCCEEE
Q 017471 147 IGYVIPTPVIMHFIQD--YEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-V-LKPSDIIL 222 (371)
Q Consensus 147 ~~~aiP~~~i~~~l~~--L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-G-L~~GDvIl 222 (371)
.+-.|.++.+-+.-.+ |.++|.-. .||+...+- +.--...-|+. ...|+.|++..||+-|+. | |..+|.++
T Consensus 146 VSsIIDVDivPEtHRRVRL~khG~ek---PLGFYIRDG-~SVRVtp~Gle-kvpGIFISRlVpGGLAeSTGLLaVnDEVl 220 (358)
T KOG3606|consen 146 VSSIIDVDIVPETHRRVRLHKHGSEK---PLGFYIRDG-TSVRVTPHGLE-KVPGIFISRLVPGGLAESTGLLAVNDEVL 220 (358)
T ss_pred eceeeeecccchhhhheehhhcCCCC---CceEEEecC-ceEEecccccc-ccCceEEEeecCCccccccceeeecceeE
Confidence 3344445444443333 33444432 477776544 11111224554 467999999999999999 7 78899999
Q ss_pred EECCEEecCC
Q 017471 223 SFDGIDIAND 232 (371)
Q Consensus 223 ~vnG~~V~~~ 232 (371)
+|||.+|...
T Consensus 221 EVNGIEVaGK 230 (358)
T KOG3606|consen 221 EVNGIEVAGK 230 (358)
T ss_pred EEcCEEeccc
Confidence 9999999764
No 72
>KOG3571 consensus Dishevelled 3 and related proteins [General function prediction only]
Probab=90.91 E-value=0.27 Score=49.51 Aligned_cols=37 Identities=16% Similarity=0.331 Sum_probs=32.0
Q ss_pred CCceEEEEECCCCcccc-C-CCCCCEEEEECCEEecCCC
Q 017471 197 QKGVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDG 233 (371)
Q Consensus 197 ~~gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~ 233 (371)
..|++|.+|.++++-+. | +.+||.||.||.....+..
T Consensus 276 DggIYVgsImkgGAVA~DGRIe~GDMiLQVNevsFENmS 314 (626)
T KOG3571|consen 276 DGGIYVGSIMKGGAVALDGRIEPGDMILQVNEVSFENMS 314 (626)
T ss_pred CCceEEeeeccCceeeccCccCccceEEEeeecchhhcC
Confidence 46899999999998877 6 9999999999998776654
No 73
>KOG3551 consensus Syntrophins (type beta) [Extracellular structures]
Probab=88.86 E-value=0.33 Score=47.39 Aligned_cols=72 Identities=21% Similarity=0.264 Sum_probs=49.4
Q ss_pred cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc--CCCCCCEEEEECCEEecCCCCccccccccchhhhhhh
Q 017471 173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVS 250 (371)
Q Consensus 173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~ 250 (371)
+-|||+++--. ++.--++|+.+.++-.|.+ -|..||.|++|||....+...- ..+.
T Consensus 96 gGLGISIKGGr-----------eNkMPIlISKIFkGlAADQt~aL~~gDaIlSVNG~dL~~AtHd-----------eAVq 153 (506)
T KOG3551|consen 96 GGLGISIKGGR-----------ENKMPILISKIFKGLAADQTGALFLGDAILSVNGEDLRDATHD-----------EAVQ 153 (506)
T ss_pred CcceEEeecCc-----------ccCCceehhHhccccccccccceeeccEEEEecchhhhhcchH-----------HHHH
Confidence 55788776431 1223589999999999998 5999999999999998766432 1222
Q ss_pred c-cCCCCEEEEEE--EECC
Q 017471 251 Q-KYTGDSAAVKV--LRDS 266 (371)
Q Consensus 251 ~-~~~g~~v~l~v--~R~g 266 (371)
. ++.|+.|.+.| .|+-
T Consensus 154 aLKraGkeV~levKy~REv 172 (506)
T KOG3551|consen 154 ALKRAGKEVLLEVKYMREV 172 (506)
T ss_pred HHHhhCceeeeeeeeehhc
Confidence 2 45687766544 4543
No 74
>KOG3549 consensus Syntrophins (type gamma) [Extracellular structures]
Probab=88.32 E-value=0.5 Score=45.54 Aligned_cols=55 Identities=22% Similarity=0.266 Sum_probs=42.2
Q ss_pred ceEEEEECCCCcccc-C-CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEE
Q 017471 199 GVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVL 263 (371)
Q Consensus 199 gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~ 263 (371)
-++|+.+.++-.|+. | |-.||-|+.|||..|+.-..- +.+ +.+ .+.|+.|+++|.
T Consensus 81 PvviSkI~kdQaAd~tG~LFvGDAilqvNGi~v~~c~He-----evV---~iL--RNAGdeVtlTV~ 137 (505)
T KOG3549|consen 81 PVVISKIYKDQAADITGQLFVGDAILQVNGIYVTACPHE-----EVV---NIL--RNAGDEVTLTVK 137 (505)
T ss_pred cEEeehhhhhhhhhhcCceEeeeeeEEeccEEeecCChH-----HHH---HHH--HhcCCEEEEEeH
Confidence 488999999999988 6 889999999999999865431 111 122 457999998885
No 75
>PF05579 Peptidase_S32: Equine arteritis virus serine endopeptidase S32; InterPro: IPR008760 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to MEROPS peptidase family S32 (clan PA(S)). The type example is equine arteritis virus serine endopeptidase (equine arteritis virus), which is involved in processing of nidovirus polyproteins [].; GO: 0004252 serine-type endopeptidase activity, 0016032 viral reproduction, 0019082 viral protein processing; PDB: 3FAN_A 3FAO_A 1MBM_A.
Probab=86.42 E-value=0.51 Score=44.02 Aligned_cols=23 Identities=30% Similarity=0.680 Sum_probs=18.6
Q ss_pred cCCCCCCceecCCCcEEEEEeee
Q 017471 116 NSGNSGGPAFNDKGKCVGIAFQS 138 (371)
Q Consensus 116 ~~G~SGGPlvn~~G~VIGI~~~~ 138 (371)
++|+||+|++..+|.+||+.+.+
T Consensus 206 ~~GDSGSPVVt~dg~liGVHTGS 228 (297)
T PF05579_consen 206 GPGDSGSPVVTEDGDLIGVHTGS 228 (297)
T ss_dssp -GGCTT-EEEETTC-EEEEEEEE
T ss_pred CCCCCCCccCcCCCCEEEEEecC
Confidence 58999999999999999999864
No 76
>KOG0609 consensus Calcium/calmodulin-dependent serine protein kinase/membrane-associated guanylate kinase [Signal transduction mechanisms]
Probab=86.02 E-value=1.3 Score=45.10 Aligned_cols=57 Identities=25% Similarity=0.336 Sum_probs=42.3
Q ss_pred ceEEEEECCCCcccc-C-CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEEC
Q 017471 199 GVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD 265 (371)
Q Consensus 199 gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~ 265 (371)
-++|+++..|+.+++ | |+.||.|+++||..+.+..--+ +..++.... | .+++.+.-.
T Consensus 147 ~~~vARI~~GG~~~r~glL~~GD~i~EvNGi~v~~~~~~e--------~q~~l~~~~-G-~itfkiiP~ 205 (542)
T KOG0609|consen 147 KVVVARIMHGGMADRQGLLHVGDEILEVNGISVANKSPEE--------LQELLRNSR-G-SITFKIIPS 205 (542)
T ss_pred ccEEeeeccCCcchhccceeeccchheecCeecccCCHHH--------HHHHHHhCC-C-cEEEEEccc
Confidence 589999999999998 6 9999999999999998763211 224444443 5 577777544
No 77
>PF02907 Peptidase_S29: Hepatitis C virus NS3 protease; InterPro: IPR004109 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This signature identifies the Hepatitis C virus NS3 protein as a serine protease which belongs to MEROPS peptidase family S29 (hepacivirin family, clan PA(S)), which has a trypsin-like fold. The non-structural (NS) protein NS3 is one of the NS proteins involved in replication of the HCV genome. The NS2 proteinase (IPR002518 from INTERPRO), a zinc-dependent enzyme, performs a single proteolytic cut to release the N terminus of NS3. The action of NS3 proteinase (NS3P), which resides in the N-terminal one-third of the NS3 protein, then yields all remaining non-structural proteins. The C-terminal two-thirds of the NS3 protein contain a helicase. The functional relationship between the proteinase and helicase domains is unknown. NS3 has a structural zinc-binding site and requires cofactor NS4. It has been suggested that the NS3 serine protease of hepatitus C is involved in cell transformation and that the ability to transform requires an active enzyme [].; GO: 0008236 serine-type peptidase activity, 0006508 proteolysis, 0019087 transformation of host cell by virus; PDB: 2QV1_B 3LOX_C 2OBQ_C 2OC1_C 2OC0_A 3LON_A 3KNX_A 2O8M_A 2OBO_A 2OC8_A ....
Probab=84.26 E-value=0.83 Score=38.22 Aligned_cols=25 Identities=32% Similarity=0.547 Sum_probs=19.6
Q ss_pred cCCCCCCceecCCCcEEEEEeeecc
Q 017471 116 NSGNSGGPAFNDKGKCVGIAFQSLK 140 (371)
Q Consensus 116 ~~G~SGGPlvn~~G~VIGI~~~~~~ 140 (371)
-.|.||||++-.+|.+|||-.+...
T Consensus 106 lkGSSGgPiLC~~GH~vG~f~aa~~ 130 (148)
T PF02907_consen 106 LKGSSGGPILCPSGHAVGMFRAAVC 130 (148)
T ss_dssp HTT-TT-EEEETTSEEEEEEEEEEE
T ss_pred EecCCCCcccCCCCCEEEEEEEEEE
Confidence 3799999999999999999766553
No 78
>cd00987 PDZ_serine_protease PDZ domain of tryspin-like serine proteases, such as DegP/HtrA, which are oligomeric proteins involved in heat-shock response, chaperone function, and apoptosis. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=83.66 E-value=1.3 Score=33.59 Aligned_cols=47 Identities=13% Similarity=-0.014 Sum_probs=37.2
Q ss_pred eccEEEech-HHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhHH
Q 017471 296 IAGFVFSRC-LYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLLW 343 (371)
Q Consensus 296 ~~Gl~~~~l-~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~~ 343 (371)
|+|+.++++ +..+..+++ ....|+.+..+.++||++......+|+|.
T Consensus 2 ~~G~~~~~~~~~~~~~~~~-~~~~g~~V~~v~~~s~a~~~gl~~GD~I~ 49 (90)
T cd00987 2 WLGVTVQDLTPDLAEELGL-KDTKGVLVASVDPGSPAAKAGLKPGDVIL 49 (90)
T ss_pred ccceEEeECCHHHHHHcCC-CCCCEEEEEEECCCCHHHHcCCCcCCEEE
Confidence 689999988 666655554 34578999999999999987777788763
No 79
>KOG3605 consensus Beta amyloid precursor-binding protein [General function prediction only]
Probab=80.46 E-value=1.7 Score=45.37 Aligned_cols=61 Identities=10% Similarity=0.149 Sum_probs=39.4
Q ss_pred EEECCCCcccc-C-CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEE
Q 017471 203 RRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNF 271 (371)
Q Consensus 203 ~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~ 271 (371)
.....++||++ | |-.||.|++|||..+-..---. -..++...+.-..|+++|.+=--..++
T Consensus 678 Anmm~~GpAarsgkLnIGDQiiaING~SLVGLPLst--------cQs~Ik~~KnQT~VkltiV~cpPV~~V 740 (829)
T KOG3605|consen 678 ANMMHGGPAARSGKLNIGDQIMSINGTSLVGLPLST--------CQSIIKGLKNQTAVKLNIVSCPPVTTV 740 (829)
T ss_pred HhcccCChhhhcCCccccceeEeecCceeccccHHH--------HHHHHhcccccceEEEEEecCCCceEE
Confidence 36677999999 5 9999999999998875432110 012333344445688888775444333
No 80
>PF01732 DUF31: Putative peptidase (DUF31); InterPro: IPR022382 This domain has no known function. It is found in various hypothetical proteins and putative lipoproteins from mycoplasmas.
Probab=76.63 E-value=1.8 Score=42.83 Aligned_cols=26 Identities=31% Similarity=0.543 Sum_probs=22.3
Q ss_pred EeeccCCCCCCceecCCCcEEEEEee
Q 017471 112 DAAINSGNSGGPAFNDKGKCVGIAFQ 137 (371)
Q Consensus 112 da~i~~G~SGGPlvn~~G~VIGI~~~ 137 (371)
+..+..|.||+.|+|.+|++|||.++
T Consensus 349 ~~~l~gGaSGS~V~n~~~~lvGIy~g 374 (374)
T PF01732_consen 349 NYSLGGGASGSMVINQNNELVGIYFG 374 (374)
T ss_pred ccCCCCCCCcCeEECCCCCEEEEeCC
Confidence 34667899999999999999999753
No 81
>KOG0606 consensus Microtubule-associated serine/threonine kinase and related proteins [Signal transduction mechanisms; General function prediction only]
Probab=76.50 E-value=2.9 Score=46.30 Aligned_cols=34 Identities=21% Similarity=0.258 Sum_probs=30.2
Q ss_pred eEEEEECCCCcccc-CCCCCCEEEEECCEEecCCC
Q 017471 200 VRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDG 233 (371)
Q Consensus 200 v~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~ 233 (371)
=+|..|.++|||.. |+++||.|+.+||+++....
T Consensus 660 h~v~sv~egsPA~~agls~~DlIthvnge~v~gl~ 694 (1205)
T KOG0606|consen 660 HSVGSVEEGSPAFEAGLSAGDLITHVNGEPVHGLV 694 (1205)
T ss_pred eeeeeecCCCCccccCCCccceeEeccCcccchhh
Confidence 45788999999988 99999999999999997654
No 82
>PF05580 Peptidase_S55: SpoIVB peptidase S55; InterPro: IPR008763 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to the MEROPS peptidase family S55 (SpoIVB peptidase family, clan PA(S)). The protein SpoIVB plays a key role in signalling in the final sigma-K checkpoint of Bacillus subtilis [, ].
Probab=76.05 E-value=2.4 Score=38.52 Aligned_cols=39 Identities=28% Similarity=0.390 Sum_probs=29.5
Q ss_pred eccCCCCCCceecCCCcEEEEEeeecccCCccceeccccCcc
Q 017471 114 AINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPV 155 (371)
Q Consensus 114 ~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~~~~~~~aiP~~~ 155 (371)
-|-.|+||+|++- +|++||=++..+. +.+..+|.+|++.
T Consensus 176 GIvqGMSGSPI~q-dGKLiGAVthvf~--~dp~~Gygi~ie~ 214 (218)
T PF05580_consen 176 GIVQGMSGSPIIQ-DGKLIGAVTHVFV--NDPTKGYGIFIEW 214 (218)
T ss_pred CEEecccCCCEEE-CCEEEEEEEEEEe--cCCCceeeecHHH
Confidence 4668999999985 8999998776553 4466777787643
No 83
>PF03761 DUF316: Domain of unknown function (DUF316) ; InterPro: IPR005514 This is a family of uncharacterised proteins from Caenorhabditis elegans.
Probab=73.42 E-value=31 Score=32.30 Aligned_cols=89 Identities=19% Similarity=0.281 Sum_probs=53.0
Q ss_pred CCCEEEEEEecCCCcCCccceecCCCCC---CCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeecc
Q 017471 40 ICIYTMLTVEDDEFWEGVLPVEFGELPA---LQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAIN 116 (371)
Q Consensus 40 ~~DlAlLkv~~~~~~~~l~~~~l~~s~~---lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~ 116 (371)
.++++||+++.+ +.....|+=+++++. .++.+.+.|+.... ......+.-.... . ....+..+....
T Consensus 160 ~~~~mIlEl~~~-~~~~~~~~Cl~~~~~~~~~~~~~~~yg~~~~~---~~~~~~~~i~~~~---~---~~~~~~~~~~~~ 229 (282)
T PF03761_consen 160 PYSPMILELEED-FSKNVSPPCLADSSTNWEKGDEVDVYGFNSTG---KLKHRKLKITNCT---K---CAYSICTKQYSC 229 (282)
T ss_pred ccceEEEEEccc-ccccCCCEEeCCCccccccCceEEEeecCCCC---eEEEEEEEEEEee---c---cceeEecccccC
Confidence 489999999987 333677777877653 35888899882221 1222222211110 0 112345555666
Q ss_pred CCCCCCcee---cCCCcEEEEEeee
Q 017471 117 SGNSGGPAF---NDKGKCVGIAFQS 138 (371)
Q Consensus 117 ~G~SGGPlv---n~~G~VIGI~~~~ 138 (371)
.|++|||++ |.+-.|||+.+..
T Consensus 230 ~~d~Gg~lv~~~~gr~tlIGv~~~~ 254 (282)
T PF03761_consen 230 KGDRGGPLVKNINGRWTLIGVGASG 254 (282)
T ss_pred CCCccCeEEEEECCCEEEEEEEccC
Confidence 899999997 4445688886543
No 84
>COG1625 Fe-S oxidoreductase, related to NifB/MoaA family [Energy production and conversion]
Probab=72.31 E-value=2.8 Score=41.69 Aligned_cols=35 Identities=14% Similarity=0.194 Sum_probs=30.2
Q ss_pred EEEEECCCCcccc-CCCCCCEEEEEC-CEEecCCCCc
Q 017471 201 RIRRVDPTAPESE-VLKPSDIILSFD-GIDIANDGTV 235 (371)
Q Consensus 201 ~V~~V~~~spA~~-GL~~GDvIl~vn-G~~V~~~~~l 235 (371)
.+..+.+++.++. |+.+||.+.+|| |.+.++-.+.
T Consensus 4 ~i~~v~~~~~~d~~Gfe~~~~l~~Vn~~~~~~~c~~~ 40 (414)
T COG1625 4 KISKVGGISGADCDGFEEGDYLLKVNPGFGCKDCIPY 40 (414)
T ss_pred ceeeccCCCcccccCccccceeeecCCCCCCCcCCCc
Confidence 4678889999999 999999999999 8888776654
No 85
>KOG3834 consensus Golgi reassembly stacking protein GRASP65, contains PDZ domain [Intracellular trafficking, secretion, and vesicular transport]
Probab=68.29 E-value=3.9 Score=40.71 Aligned_cols=65 Identities=14% Similarity=0.190 Sum_probs=45.9
Q ss_pred EEEEECCCCcccc-CCC-CCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEec
Q 017471 201 RIRRVDPTAPESE-VLK-PSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLA 276 (371)
Q Consensus 201 ~V~~V~~~spA~~-GL~-~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~ 276 (371)
-|-+|.++|||+. ||+ -+|.|+.+-.......+|+. .+...+.++.+++.|+--.....-+|++.
T Consensus 112 Hvl~V~p~SPaalAgl~~~~DYivG~~~~~~~~~eDl~-----------~lIeshe~kpLklyVYN~D~d~~ReVti~ 178 (462)
T KOG3834|consen 112 HVLSVEPNSPAALAGLRPYTDYIVGIWDAVMHEEEDLF-----------TLIESHEGKPLKLYVYNHDTDSCREVTIT 178 (462)
T ss_pred eeeecCCCCHHHhcccccccceEecchhhhccchHHHH-----------HHHHhccCCCcceeEeecCCCccceEEee
Confidence 3668999999999 999 68999999545555566653 34445678899999987554443444443
No 86
>KOG3938 consensus RGS-GAIP interacting protein GIPC, contains PDZ domain [Signal transduction mechanisms; Intracellular trafficking, secretion, and vesicular transport]
Probab=64.24 E-value=5.5 Score=37.25 Aligned_cols=67 Identities=13% Similarity=0.280 Sum_probs=49.5
Q ss_pred hccCCCCCCc---eEEEEECCCCcccc--CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEE
Q 017471 190 AMSMKADQKG---VRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLR 264 (371)
Q Consensus 190 ~~gl~~~~~g---v~V~~V~~~spA~~--GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R 264 (371)
.+|+.-...| ..|..+.++|--.+ -++.||.|-+|||+.|-.+..++ ..+.+.....|++.++.+.-
T Consensus 138 alGlTITDNG~GyAFIKrIkegsvidri~~i~VGd~IEaiNge~ivG~RHYe--------VArmLKel~rge~ftlrLie 209 (334)
T KOG3938|consen 138 ALGLTITDNGAGYAFIKRIKEGSVIDRIEAICVGDHIEAINGESIVGKRHYE--------VARMLKELPRGETFTLRLIE 209 (334)
T ss_pred ccceEEeeCCcceeeeEeecCCchhhhhhheeHHhHHHhhcCccccchhHHH--------HHHHHHhcccCCeeEEEeec
Confidence 4555433333 56889999998887 79999999999999998887654 34566666778877776554
No 87
>PF03510 Peptidase_C24: 2C endopeptidase (C24) cysteine protease family; InterPro: IPR000317 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. The two signatures that defines this group of calivirus polyproteins identify a cysteine peptidase signature that belongs to MEROPS peptidase family C24 (clan PA(C)). Caliciviruses are positive-stranded ssRNA viruses that cause gastroenteritis. The calicivirus genome contains two open reading frames, ORF1 and ORF2. ORF2 encodes a structural protein []; while ORF1 encodes a non-structural polypeptide, which has RNA helicase, cysteine protease and RNA polymerase activity. The regions of the polyprotein in which these activities lie are similar to proteins produced by the picornaviruses. Two different families of caliciviruses can be distinguished on the basis of sequence similarity, namely those classified as small round structured viruses (SRSVs) and those classed as non-SRSVs. Calicivirus proteases from the non-SRSV group, which are members of the PA protease clan, constitute family C24 of the cysteine proteases (proteases from SRSVs belong to the C37 family). As mentioned above, the protease activity resides within a polyprotein. The enzyme cleaves the polyprotein at sites N-terminal to itself, liberating the polyprotein helicase.; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis
Probab=63.96 E-value=38 Score=27.24 Aligned_cols=52 Identities=8% Similarity=-0.049 Sum_probs=35.5
Q ss_pred eeeeccEEEEE-EEecCCCeE----eEEEEEecCCCCEEEEEEecCCCcCCccceecCCCCCC
Q 017471 11 LNSRNEALILS-TWLLCSPSA----PSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELPAL 68 (371)
Q Consensus 11 ~~~~gsg~vi~-~~~~~~~~~----~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~~l 68 (371)
.+..|.|..++ -|+...... +-+++. ..-|+|+++.+.. .+|..++++++.+
T Consensus 3 avHIGnG~~vt~tHva~~~~~v~g~~f~~~~--~~ge~~~v~~~~~----~~p~~~ig~g~Pv 59 (105)
T PF03510_consen 3 AVHIGNGRYVTVTHVAKSSDSVDGQPFKIVK--TDGELCWVQSPLV----HLPAAQIGTGKPV 59 (105)
T ss_pred eEEeCCCEEEEEEEEeccCceEcCcCcEEEE--eccCEEEEECCCC----CCCeeEeccCCCE
Confidence 35568888888 777776643 223333 4569999999887 4688888875543
No 88
>PF12812 PDZ_1: PDZ-like domain
Probab=63.57 E-value=15 Score=27.89 Aligned_cols=39 Identities=10% Similarity=0.012 Sum_probs=29.2
Q ss_pred CCceeeccEEEech-HHHHHHcCcccccCceEEEEEecCChHh
Q 017471 291 PSYYIIAGFVFSRC-LYLISVLSMERIMNMKLRSSFWTSSCIQ 332 (371)
Q Consensus 291 ~~~~~~~Gl~~~~l-~~~~~~~~~~~~~~~~~v~~~~~~Sp~~ 332 (371)
.++..|+|.+|+++ .....+++++ .++++.+...+||+.
T Consensus 5 ~r~v~~~Ga~f~~Ls~q~aR~~~~~---~~gv~v~~~~g~~~~ 44 (78)
T PF12812_consen 5 SRFVEVCGAVFHDLSYQQARQYGIP---VGGVYVAVSGGSLAF 44 (78)
T ss_pred CEEEEEcCeecccCCHHHHHHhCCC---CCEEEEEecCCChhh
Confidence 35677999999998 6667788766 336666778888865
No 89
>PF13180 PDZ_2: PDZ domain; PDB: 2L97_A 1Y8T_A 2Z9I_A 1LCY_A 2PZD_B 2P3W_A 1VCW_C 1TE0_B 1SOZ_C 1SOT_C ....
Probab=60.33 E-value=6.5 Score=29.53 Aligned_cols=37 Identities=11% Similarity=-0.081 Sum_probs=27.6
Q ss_pred eccEEEechHHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhH
Q 017471 296 IAGFVFSRCLYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLL 342 (371)
Q Consensus 296 ~~Gl~~~~l~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~ 342 (371)
++|+.++.. ....+++|.++.++|||+-.-.+.+|+|
T Consensus 2 ~lGv~~~~~----------~~~~g~~V~~V~~~spA~~aGl~~GD~I 38 (82)
T PF13180_consen 2 GLGVTVQNL----------SDTGGVVVVSVIPGSPAAKAGLQPGDII 38 (82)
T ss_dssp E-SEEEEEC----------SCSSSEEEEEESTTSHHHHTTS-TTEEE
T ss_pred EECeEEEEc----------cCCCeEEEEEeCCCCcHHHCCCCCCcEE
Confidence 678877543 1156899999999999999988888875
No 90
>cd01735 LSm12_N LSm12 belongs to a family of Sm-like proteins that associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet that associates with other Sm proteins to form hexameric and heptameric ring structures. In addition to the N-terminal Sm-like domain, LSm12 has a novel methyltransferase domain.
Probab=55.13 E-value=38 Score=24.46 Aligned_cols=34 Identities=6% Similarity=-0.088 Sum_probs=27.4
Q ss_pred EEEEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471 17 ALILSTWLLCSPSAPSATLVTADICIYTMLTVED 50 (371)
Q Consensus 17 g~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~ 50 (371)
|-.++...-.|.++.++|+++|+...+.+||-.+
T Consensus 6 Gs~V~~kTc~g~~ieGEV~afD~~tk~lIlk~~s 39 (61)
T cd01735 6 GSQVSCRTCFEQRLQGEVVAFDYPSKMLILKCPS 39 (61)
T ss_pred ccEEEEEecCCceEEEEEEEecCCCcEEEEECcc
Confidence 3445555556999999999999999999999655
No 91
>TIGR02038 protease_degS periplasmic serine pepetdase DegS. This family consists of the periplasmic serine protease DegS (HhoB), a shorter paralog of protease DO (HtrA, DegP) and DegQ (HhoA). It is found in E. coli and several other Proteobacteria of the gamma subdivision. It contains a trypsin domain and a single copy of PDZ domain (in contrast to DegP with two copies). A critical role of this DegS is to sense stress in the periplasm and partially degrade an inhibitor of sigma(E).
Probab=54.96 E-value=9.9 Score=37.32 Aligned_cols=48 Identities=4% Similarity=-0.083 Sum_probs=39.8
Q ss_pred eccEEEech-HHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhHHH
Q 017471 296 IAGFVFSRC-LYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLLWC 344 (371)
Q Consensus 296 ~~Gl~~~~l-~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~~~ 344 (371)
|+|+.++++ +...+.++++. ..|++|..+.++||++-.-...+|+|..
T Consensus 256 ~lGv~~~~~~~~~~~~lgl~~-~~Gv~V~~V~~~spA~~aGL~~GDvI~~ 304 (351)
T TIGR02038 256 YIGVSGEDINSVVAQGLGLPD-LRGIVITGVDPNGPAARAGILVRDVILK 304 (351)
T ss_pred EeeeEEEECCHHHHHhcCCCc-cccceEeecCCCChHHHCCCCCCCEEEE
Confidence 899999888 77777888864 4799999999999999877777887753
No 92
>cd01726 LSm6 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm6 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=51.69 E-value=29 Score=25.26 Aligned_cols=33 Identities=6% Similarity=-0.150 Sum_probs=28.3
Q ss_pred EEEEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471 17 ALILSTWLLCSPSAPSATLVTADICIYTMLTVE 49 (371)
Q Consensus 17 g~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~ 49 (371)
|-.+.+.+.+|+++.+++.++|+..|+.+=...
T Consensus 10 ~~~V~V~Lk~g~~~~G~L~~~D~~mNlvL~~~~ 42 (67)
T cd01726 10 GRPVVVKLNSGVDYRGILACLDGYMNIALEQTE 42 (67)
T ss_pred CCeEEEEECCCCEEEEEEEEEccceeeEEeeEE
Confidence 446778999999999999999999999886654
No 93
>PF02122 Peptidase_S39: Peptidase S39; InterPro: IPR000382 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. ORF2 of Potato leafroll virus (PLrV) encodes a polyprotein which is translated following a -1 frameshift. The polyprotein has a putative linear arrangement of membrane achor-VPg-peptidase-polmerase domains. The serine peptidase domain which is found in this group of sequences belongs to MEROPS peptidase family S39 (clan PA(S)). It is likely that the peptidase domain is involved in the cleavage of the polyprotein []. The nucleotide sequence for the RNA of PLrV has been determined [, ]. The sequence contains six large open reading frames (ORFs). The 5' coding region encodes two polypeptides of 28K and 70K, which overlap in different reading frames; it is suggested that the third ORF in the 5' block is translated by frameshift readthrough near the end of the 70K protein, yielding a 118K polypeptide []. Segments of the predicted amino acid sequences of these ORFs resemble those of known viral RNA polymerases, ATP-binding proteins and viral genome-linked proteins. The nucleotide sequence of the genomic RNA of Beet western yellows virus (BWYV) has been determined []. The sequence contains six long ORFs. A cluster of three of these ORFs, including the coat protein cistron, display extensive amino acid sequence similarity to corresponding ORFs of a second luteovirus: Barley yellow dwarf virus [].; GO: 0004252 serine-type endopeptidase activity, 0022415 viral reproductive process, 0016021 integral to membrane; PDB: 1ZYO_A.
Probab=49.98 E-value=21 Score=32.35 Aligned_cols=46 Identities=26% Similarity=0.337 Sum_probs=18.1
Q ss_pred EEEEEeeccCCCCCCceecCCCcEEEEEeeecccCCccceeccccCc
Q 017471 108 GLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTP 154 (371)
Q Consensus 108 ~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~~~~~~~aiP~~ 154 (371)
+..+-+.-.+|.||-|.++.+ +++|+.....+....++.++.-|+.
T Consensus 137 ~~~vls~T~~G~SGtp~y~g~-~vvGvH~G~~~~~~~~n~n~~spip 182 (203)
T PF02122_consen 137 FASVLSNTSPGWSGTPYYSGK-NVVGVHTGSPSGSNRENNNRMSPIP 182 (203)
T ss_dssp EEEE-----TT-TT-EEE-SS--EEEEEEEE----------------
T ss_pred CCceEcCCCCCCCCCCeEECC-CceEeecCccccccccccccccccc
Confidence 456667778999999999998 9999988753333445666655554
No 94
>TIGR03000 plancto_dom_1 Planctomycetes uncharacterized domain TIGR03000. Domains described by this model are found, so far, only in the Planctomycetes (Pirellula sp. strain 1 and Gemmata obscuriglobus), in up to six proteins per genome, and may be duplicated within a protein. The function is unknown.
Probab=49.83 E-value=54 Score=24.69 Aligned_cols=50 Identities=28% Similarity=0.408 Sum_probs=33.8
Q ss_pred CCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCC----EEEEEEEECCEEEEEEEEe
Q 017471 217 PSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGD----SAAVKVLRDSKILNFNITL 275 (371)
Q Consensus 217 ~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~----~v~l~v~R~g~~~~~~v~l 275 (371)
|-|-.+.+||++.++.+..+- ..-....+|. ++..++.|||+..+.+-++
T Consensus 10 PadAkl~v~G~~t~~~G~~R~---------F~T~~L~~G~~y~Y~v~a~~~~dG~~~t~~~~V 63 (75)
T TIGR03000 10 PADAKLKVDGKETNGTGTVRT---------FTTPPLEAGKEYEYTVTAEYDRDGRILTRTRTV 63 (75)
T ss_pred CCCCEEEECCeEcccCccEEE---------EECCCCCCCCEEEEEEEEEEecCCcEEEEEEEE
Confidence 468899999999999887640 1112234565 4677888999876655444
No 95
>cd00600 Sm_like The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=48.21 E-value=42 Score=23.59 Aligned_cols=33 Identities=9% Similarity=-0.044 Sum_probs=28.2
Q ss_pred EEEEEecCCCeEeEEEEEecCCCCEEEEEEecC
Q 017471 19 ILSTWLLCSPSAPSATLVTADICIYTMLTVEDD 51 (371)
Q Consensus 19 vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~~ 51 (371)
.+.+.+.+++.+.|++.++|+..|+.+-.....
T Consensus 8 ~V~V~l~~g~~~~G~L~~~D~~~Ni~L~~~~~~ 40 (63)
T cd00600 8 TVRVELKDGRVLEGVLVAFDKYMNLVLDDVEET 40 (63)
T ss_pred EEEEEECCCcEEEEEEEEECCCCCEEECCEEEE
Confidence 456888999999999999999999988776543
No 96
>cd01722 Sm_F The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit F is capable of forming both homo- and hetero-heptamer ring structures. To form the hetero-heptamer, Sm subunit F initially binds subunits E and G to form a trimer which then assembles onto snRNA along with the D3/B and D1/D2 heterodimers.
Probab=47.66 E-value=33 Score=25.02 Aligned_cols=33 Identities=6% Similarity=-0.092 Sum_probs=28.0
Q ss_pred EEEEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471 17 ALILSTWLLCSPSAPSATLVTADICIYTMLTVE 49 (371)
Q Consensus 17 g~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~ 49 (371)
|-.+.+.+.+|+.+.+++.++|...|+.+=...
T Consensus 11 g~~V~V~Lk~g~~~~G~L~~~D~~mNi~L~~~~ 43 (68)
T cd01722 11 GKPVIVKLKWGMEYKGTLVSVDSYMNLQLANTE 43 (68)
T ss_pred CCEEEEEECCCcEEEEEEEEECCCEEEEEeeEE
Confidence 345678999999999999999999999886554
No 97
>PRK00737 small nuclear ribonucleoprotein; Provisional
Probab=47.20 E-value=43 Score=24.75 Aligned_cols=34 Identities=6% Similarity=-0.217 Sum_probs=28.8
Q ss_pred EEEEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471 17 ALILSTWLLCSPSAPSATLVTADICIYTMLTVED 50 (371)
Q Consensus 17 g~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~ 50 (371)
+-.+.+.+.+|+++.+++.++|+..|+.+=....
T Consensus 14 ~k~V~V~lk~g~~~~G~L~~~D~~mNlvL~d~~e 47 (72)
T PRK00737 14 NSPVLVRLKGGREFRGELQGYDIHMNLVLDNAEE 47 (72)
T ss_pred CCEEEEEECCCCEEEEEEEEEcccceeEEeeEEE
Confidence 3356688999999999999999999998877654
No 98
>cd01717 Sm_B The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit B heterodimerizes with subunit D3 and three such heterodimers form a hexameric ring structure with alternating B and D3 subunits. The D3 - B heterodimer also assembles into a heptameric ring containing D1, D2, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=47.04 E-value=38 Score=25.50 Aligned_cols=31 Identities=13% Similarity=-0.037 Sum_probs=26.2
Q ss_pred EEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471 20 LSTWLLCSPSAPSATLVTADICIYTMLTVED 50 (371)
Q Consensus 20 i~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~ 50 (371)
+.+++-+|+.+.|++.++|+..|+.|=....
T Consensus 13 V~V~l~dgR~~~G~L~~~D~~~NlVL~~~~E 43 (79)
T cd01717 13 LRVTLQDGRQFVGQFLAFDKHMNLVLSDCEE 43 (79)
T ss_pred EEEEECCCcEEEEEEEEEcCccCEEcCCEEE
Confidence 4578899999999999999999998765543
No 99
>cd01730 LSm3 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm3 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=46.29 E-value=35 Score=25.93 Aligned_cols=29 Identities=7% Similarity=-0.183 Sum_probs=25.1
Q ss_pred EEEEecCCCeEeEEEEEecCCCCEEEEEE
Q 017471 20 LSTWLLCSPSAPSATLVTADICIYTMLTV 48 (371)
Q Consensus 20 i~~~~~~~~~~~A~vv~~d~~~DlAlLkv 48 (371)
+.+.+.+|+.+.|++.++|.+.|+.+=..
T Consensus 14 V~V~l~~gr~~~G~L~~fD~~mNlvL~d~ 42 (82)
T cd01730 14 VYVKLRGDRELRGRLHAYDQHLNMILGDV 42 (82)
T ss_pred EEEEECCCCEEEEEEEEEccceEEeccce
Confidence 45788999999999999999999987544
No 100
>KOG3553 consensus Tax interaction protein TIP1 [Cell wall/membrane/envelope biogenesis]
Probab=45.01 E-value=32 Score=27.40 Aligned_cols=27 Identities=4% Similarity=-0.121 Sum_probs=22.5
Q ss_pred ccCceEEEEEecCChHhHHhhhhhhhH
Q 017471 316 IMNMKLRSSFWTSSCIQCHNCQMSSLL 342 (371)
Q Consensus 316 ~~~~~~v~~~~~~Sp~~~~~~~~~~~~ 342 (371)
...|+||+.+.+||||+.--+...|-|
T Consensus 57 tD~GiYvT~V~eGsPA~~AGLrihDKI 83 (124)
T KOG3553|consen 57 TDKGIYVTRVSEGSPAEIAGLRIHDKI 83 (124)
T ss_pred CCccEEEEEeccCChhhhhcceecceE
Confidence 357999999999999999877776643
No 101
>TIGR02860 spore_IV_B stage IV sporulation protein B. SpoIVB, the stage IV sporulation protein B of endospore-forming bacteria such as Bacillus subtilis, is a serine proteinase, expressed in the spore (rather than mother cell) compartment, that participates in a proteolytic activation cascade for Sigma-K. It appears to be universal among endospore-forming bacteria and occurs nowhere else.
Probab=44.68 E-value=20 Score=35.90 Aligned_cols=37 Identities=27% Similarity=0.447 Sum_probs=26.7
Q ss_pred eccCCCCCCceecCCCcEEEEEeeecccCCccceeccccC
Q 017471 114 AINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPT 153 (371)
Q Consensus 114 ~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~~~~~~~aiP~ 153 (371)
-|-.|+||+|++- +|++||=.+.-+-+ .+..+|+|-+
T Consensus 356 GivqGMSGSPi~q-~gkliGAvtHVfvn--dpt~GYGi~i 392 (402)
T TIGR02860 356 GIVQGMSGSPIIQ-NGKVIGAVTHVFVN--DPTSGYGVYI 392 (402)
T ss_pred CEEecccCCCEEE-CCEEEEEEEEEEec--CCCcceeehH
Confidence 4567999999994 79999987766643 3455566644
No 102
>PRK10898 serine endoprotease; Provisional
Probab=43.41 E-value=22 Score=34.88 Aligned_cols=48 Identities=6% Similarity=-0.088 Sum_probs=37.3
Q ss_pred eccEEEech-HHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhHHH
Q 017471 296 IAGFVFSRC-LYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLLWC 344 (371)
Q Consensus 296 ~~Gl~~~~l-~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~~~ 344 (371)
|+|+..+++ +.....+++. ...|++|..+.++||++-.-...+|+|..
T Consensus 257 ~lGi~~~~~~~~~~~~~~~~-~~~Gv~V~~V~~~spA~~aGL~~GDvI~~ 305 (353)
T PRK10898 257 YIGIGGREIAPLHAQGGGID-QLQGIVVNEVSPDGPAAKAGIQVNDLIIS 305 (353)
T ss_pred ccceEEEECCHHHHHhcCCC-CCCeEEEEEECCCChHHHcCCCCCCEEEE
Confidence 799988877 5555555554 34899999999999999887888887653
No 103
>cd01732 LSm5 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm4 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=43.28 E-value=45 Score=25.06 Aligned_cols=31 Identities=10% Similarity=-0.095 Sum_probs=26.2
Q ss_pred EEEEEEecCCCeEeEEEEEecCCCCEEEEEE
Q 017471 18 LILSTWLLCSPSAPSATLVTADICIYTMLTV 48 (371)
Q Consensus 18 ~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv 48 (371)
--+.+++.+++.+.|++.++|...|+.+=..
T Consensus 14 ~~V~V~l~~gr~~~G~L~g~D~~mNlvL~da 44 (76)
T cd01732 14 SRIWIVMKSDKEFVGTLLGFDDYVNMVLEDV 44 (76)
T ss_pred CEEEEEECCCeEEEEEEEEeccceEEEEccE
Confidence 3455788999999999999999999987554
No 104
>PF11874 DUF3394: Domain of unknown function (DUF3394); InterPro: IPR021814 This domain is functionally uncharacterised. This domain is found in bacteria. This presumed domain is about 190 amino acids in length. This domain is found associated with PF06808 from PFAM.
Probab=43.26 E-value=19 Score=32.00 Aligned_cols=28 Identities=14% Similarity=0.051 Sum_probs=25.1
Q ss_pred CCceEEEEECCCCcccc-CCCCCCEEEEE
Q 017471 197 QKGVRIRRVDPTAPESE-VLKPSDIILSF 224 (371)
Q Consensus 197 ~~gv~V~~V~~~spA~~-GL~~GDvIl~v 224 (371)
...+.|..|..+|||++ |+..|+.|+++
T Consensus 121 ~~~~~Vd~v~fgS~A~~~g~d~d~~I~~v 149 (183)
T PF11874_consen 121 GGKVIVDEVEFGSPAEKAGIDFDWEITEV 149 (183)
T ss_pred CCEEEEEecCCCCHHHHcCCCCCcEEEEE
Confidence 35689999999999999 99999998877
No 105
>COG0298 HypC Hydrogenase maturation factor [Posttranslational modification, protein turnover, chaperones]
Probab=42.41 E-value=47 Score=25.32 Aligned_cols=47 Identities=23% Similarity=0.220 Sum_probs=32.4
Q ss_pred EeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCCCCCCeEEE-EEeC
Q 017471 30 APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELPALQDAVTV-VGYP 78 (371)
Q Consensus 30 ~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~~lgd~V~~-iG~p 78 (371)
+|++|+.+|...++|++.+-.-. .+..---+++...+||+|++ +||.
T Consensus 5 iPgqI~~I~~~~~~A~Vd~gGvk--reV~l~Lv~~~v~~GdyVLVHvGfA 52 (82)
T COG0298 5 IPGQIVEIDDNNHLAIVDVGGVK--REVNLDLVGEEVKVGDYVLVHVGFA 52 (82)
T ss_pred cccEEEEEeCCCceEEEEeccEe--EEEEeeeecCccccCCEEEEEeeEE
Confidence 68999999998889999987642 11222223336688999887 5554
No 106
>cd01729 LSm7 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm7 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=42.28 E-value=53 Score=24.93 Aligned_cols=30 Identities=0% Similarity=-0.170 Sum_probs=25.5
Q ss_pred EEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471 20 LSTWLLCSPSAPSATLVTADICIYTMLTVE 49 (371)
Q Consensus 20 i~~~~~~~~~~~A~vv~~d~~~DlAlLkv~ 49 (371)
+.+.+.+|+++.+++.++|...|+.+=...
T Consensus 15 V~V~l~~gr~~~G~L~~~D~~mNlvL~~~~ 44 (81)
T cd01729 15 IRVKFQGGREVTGILKGYDQLLNLVLDDTV 44 (81)
T ss_pred EEEEECCCcEEEEEEEEEcCcccEEecCEE
Confidence 447788999999999999999999885543
No 107
>cd01720 Sm_D2 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit D2 heterodimerizes with subunit D1 and three such heterodimers form a hexameric ring structure with alternating D1 and D2 subunits. The D1 - D2 heterodimer also assembles into a heptameric ring containing D2, D3, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=42.15 E-value=50 Score=25.57 Aligned_cols=31 Identities=6% Similarity=-0.090 Sum_probs=26.9
Q ss_pred EEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471 19 ILSTWLLCSPSAPSATLVTADICIYTMLTVE 49 (371)
Q Consensus 19 vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~ 49 (371)
.+.+++-+++.+.|++.++|.+.|+.+=...
T Consensus 16 ~V~V~lr~~r~~~G~L~~fD~hmNlvL~d~~ 46 (87)
T cd01720 16 QVLINCRNNKKLLGRVKAFDRHCNMVLENVK 46 (87)
T ss_pred EEEEEEcCCCEEEEEEEEecCccEEEEcceE
Confidence 4568899999999999999999999976654
No 108
>cd01731 archaeal_Sm1 The archaeal sm1 proteins: The Sm proteins are conserved in all three domains of life and are always associated with U-rich RNA sequences. They function to mediate RNA-RNA interactions and RNA biogenesis. All Sm proteins contain a common sequence motif in two segments, Sm1 and Sm2, separated by a short variable linker. Eukaryotic Sm proteins form part of specific small nuclear ribonucleoproteins (snRNPs) that are involved in the processing of pre-mRNAs to mature mRNAs, and are a major component of the eukaryotic spliceosome. Most snRNPs consist of seven Sm proteins (B/B', D1, D2, D3, E, F and G) arranged in a ring on a uridine-rich sequence (Sm site), plus a small nuclear RNA (snRNA) (either U1, U2, U5 or U4/6). Since archaebacteria do not have any splicing apparatus, Sm proteins of archaebacteria may play a more general role. Archaeal Lsm proteins are likely to represent the ancestral Sm domain.
Probab=41.91 E-value=53 Score=23.85 Aligned_cols=33 Identities=6% Similarity=-0.146 Sum_probs=28.7
Q ss_pred EEEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471 18 LILSTWLLCSPSAPSATLVTADICIYTMLTVED 50 (371)
Q Consensus 18 ~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~ 50 (371)
-.+.+.+.+|+.+.+++.++|+..|+.+-....
T Consensus 11 ~~V~V~l~~g~~~~G~L~~~D~~mNlvL~~~~e 43 (68)
T cd01731 11 KPVLVKLKGGKEVRGRLKSYDQHMNLVLEDAEE 43 (68)
T ss_pred CEEEEEECCCCEEEEEEEEECCcceEEEeeEEE
Confidence 356688899999999999999999999887754
No 109
>KOG1738 consensus Membrane-associated guanylate kinase-interacting protein/connector enhancer of KSR-like [Nucleotide transport and metabolism]
Probab=41.36 E-value=21 Score=37.39 Aligned_cols=34 Identities=9% Similarity=0.085 Sum_probs=30.4
Q ss_pred eEEEEECCCCcccc--CCCCCCEEEEECCEEecCCC
Q 017471 200 VRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDG 233 (371)
Q Consensus 200 v~V~~V~~~spA~~--GL~~GDvIl~vnG~~V~~~~ 233 (371)
.+|+++.++|||.. -|..||.|+.||++.|-.|.
T Consensus 227 h~~s~~~e~Spad~~~kI~dgdEv~qiN~qtvVgwq 262 (638)
T KOG1738|consen 227 HVTSKIFEQSPADYRQKILDGDEVLQINEQTVVGWQ 262 (638)
T ss_pred eeccccccCChHHHhhcccCccceeeecccccccch
Confidence 56788999999988 69999999999999988775
No 110
>cd06168 LSm9 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm9 proteins have a single Sm-like domain structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=40.86 E-value=60 Score=24.34 Aligned_cols=31 Identities=10% Similarity=0.071 Sum_probs=26.6
Q ss_pred EEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471 19 ILSTWLLCSPSAPSATLVTADICIYTMLTVE 49 (371)
Q Consensus 19 vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~ 49 (371)
.+.+++.||+.+.+++.++|+..|+.+=...
T Consensus 12 ~v~V~l~dgR~~~G~l~~~D~~~NivL~~~~ 42 (75)
T cd06168 12 TMRIHMTDGRTLVGVFLCTDRDCNIILGSAQ 42 (75)
T ss_pred eEEEEEcCCeEEEEEEEEEcCCCcEEecCcE
Confidence 4568999999999999999999999875554
No 111
>PF00571 CBS: CBS domain CBS domain web page. Mutations in the CBS domain of Swiss:P35520 lead to homocystinuria.; InterPro: IPR000644 CBS (cystathionine-beta-synthase) domains are small intracellular modules, mostly found in two or four copies within a protein, that occur in a variety of proteins in bacteria, archaea, and eukaryotes [, ]. Tandem pairs of CBS domains can act as binding domains for adenosine derivatives and may regulate the activity of attached enzymatic or other domains []. In some cases, CBS domains may act as sensors of cellular energy status by being activated by AMP and inhibited by ATP []. In chloride ion channels, the CBS domains have been implicated in intracellular targeting and trafficking, as well as in protein-protein interactions, but results vary with different channels: in the CLC-5 channel, the CBS domain was shown to be required for trafficking [], while in the CLC-1 channel, the CBS domain was shown to be critical for channel function, but not necessary for trafficking []. Recent experiments revealing that CBS domains can bind adenosine-containing ligands such ATP, AMP, or S-adenosylmethionine have led to the hypothesis that CBS domains function as sensors of intracellular metabolites [, ]. Crystallographic studies of CBS domains have shown that pairs of CBS sequences form a globular domain where each CBS unit adopts a beta-alpha-beta-beta-alpha pattern []. Crystal structure of the CBS domains of the AMP-activated protein kinase in complexes with AMP and ATP shows that the phosphate groups of AMP/ATP lie in a surface pocket at the interface of two CBS domains, which is lined with basic residues, many of which are associated with disease-causing mutations []. In humans, mutations in conserved residues within CBS domains cause a variety of human hereditary diseases, including (with the gene mutated in parentheses): homocystinuria (cystathionine beta-synthase); Wolff-Parkinson-White syndrome (gamma 2 subunit of AMP-activated protein kinase); retinitis pigmentosa (IMP dehydrogenase-1); congenital myotonia, idiopathic generalized epilepsy, hypercalciuric nephrolithiasis, and classic Bartter syndrome (CLC chloride channel family members).; GO: 0005515 protein binding; PDB: 3JTF_A 3TE5_C 3TDH_C 3T4N_C 2QLV_C 3OI8_A 3LV9_A 2QH1_B 1PVM_B 3LQN_A ....
Probab=39.53 E-value=26 Score=23.68 Aligned_cols=20 Identities=40% Similarity=0.567 Sum_probs=16.3
Q ss_pred CCCCCceecCCCcEEEEEee
Q 017471 118 GNSGGPAFNDKGKCVGIAFQ 137 (371)
Q Consensus 118 G~SGGPlvn~~G~VIGI~~~ 137 (371)
+.+.-|++|.+|+++|+.+.
T Consensus 29 ~~~~~~V~d~~~~~~G~is~ 48 (57)
T PF00571_consen 29 GISRLPVVDEDGKLVGIISR 48 (57)
T ss_dssp TSSEEEEESTTSBEEEEEEH
T ss_pred CCcEEEEEecCCEEEEEEEH
Confidence 45667899999999999764
No 112
>COG0260 PepB Leucyl aminopeptidase [Amino acid transport and metabolism]
Probab=39.03 E-value=44 Score=34.35 Aligned_cols=58 Identities=16% Similarity=0.251 Sum_probs=34.7
Q ss_pred HhccCCCCCCceEEEEECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc
Q 017471 189 VAMSMKADQKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ 251 (371)
Q Consensus 189 ~~~gl~~~~~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~ 251 (371)
..++++.+ -+.|....++.|.....||||||++.||+.|.=...= ..+|+.+.+.+..
T Consensus 291 a~l~l~vn--v~~vl~~~ENm~~g~A~rPGDVits~~GkTVEV~NTD---AEGRLVLADaLtY 348 (485)
T COG0260 291 AELKLPVN--VVGVLPAVENMPSGNAYRPGDVITSMNGKTVEVLNTD---AEGRLVLADALTY 348 (485)
T ss_pred HHcCCCce--EEEEEeeeccCCCCCCCCCCCeEEecCCcEEEEcccC---ccHHHHHHHHHHH
Confidence 34556532 2334445566666556899999999999987422110 1266666665544
No 113
>smart00651 Sm snRNP Sm proteins. small nuclear ribonucleoprotein particles (snRNPs) involved in pre-mRNA splicing
Probab=37.92 E-value=70 Score=22.80 Aligned_cols=33 Identities=9% Similarity=-0.125 Sum_probs=27.8
Q ss_pred EEEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471 18 LILSTWLLCSPSAPSATLVTADICIYTMLTVED 50 (371)
Q Consensus 18 ~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~ 50 (371)
-.+.+++.+|+.+.|++.++|+..|+-+=....
T Consensus 9 ~~V~V~l~~g~~~~G~L~~~D~~~NlvL~~~~e 41 (67)
T smart00651 9 KRVLVELKNGREYRGTLKGFDQFMNLVLEDVEE 41 (67)
T ss_pred cEEEEEECCCcEEEEEEEEECccccEEEccEEE
Confidence 356688899999999999999999998866654
No 114
>PF01455 HupF_HypC: HupF/HypC family; InterPro: IPR001109 The large subunit of [NiFe]-hydrogenase, as well as other nickel metalloenzymes, is synthesised as a precursor devoid of the metalloenzyme active site. This precursor then undergoes a complex post-translational maturation process that requires a number of accessory proteins. The hydrogenase expression/formation proteins (HupF/HypC) form a family of small proteins that are hydrogenase precursor-specific chaperones required for this maturation process []. They are believed to keep the hydrogenase precursor in a conformation accessible for metal incorporation [, ].; PDB: 3D3R_A 2Z1C_C 2OT2_A.
Probab=36.19 E-value=59 Score=23.89 Aligned_cols=41 Identities=12% Similarity=-0.006 Sum_probs=29.1
Q ss_pred EeEEEEEecCCCCEEEEEEecCCCcCCcccee--cCCCCCCCCeEEEE
Q 017471 30 APSATLVTADICIYTMLTVEDDEFWEGVLPVE--FGELPALQDAVTVV 75 (371)
Q Consensus 30 ~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~--l~~s~~lgd~V~~i 75 (371)
+|++|+.++.....|++.+.... ..+. +-+..++||+|++-
T Consensus 5 iP~~Vv~v~~~~~~A~v~~~G~~-----~~V~~~lv~~v~~Gd~VLVH 47 (68)
T PF01455_consen 5 IPGRVVEVDEDGGMAVVDFGGVR-----REVSLALVPDVKVGDYVLVH 47 (68)
T ss_dssp EEEEEEEEETTTTEEEEEETTEE-----EEEEGTTCTSB-TT-EEEEE
T ss_pred ccEEEEEEeCCCCEEEEEcCCcE-----EEEEEEEeCCCCCCCEEEEe
Confidence 68999999888999999888642 3333 33446789999886
No 115
>cd01728 LSm1 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm1 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=36.08 E-value=82 Score=23.53 Aligned_cols=31 Identities=3% Similarity=-0.170 Sum_probs=26.2
Q ss_pred EEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471 19 ILSTWLLCSPSAPSATLVTADICIYTMLTVE 49 (371)
Q Consensus 19 vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~ 49 (371)
.+.+.+.+|+.+.|++.++|+..|+.+=...
T Consensus 14 ~v~V~l~~gr~~~G~L~~fD~~~NlvL~d~~ 44 (74)
T cd01728 14 KVVVLLRDGRKLIGILRSFDQFANLVLQDTV 44 (74)
T ss_pred EEEEEEcCCeEEEEEEEEECCcccEEecceE
Confidence 3447888999999999999999999886654
No 116
>cd01721 Sm_D3 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit D3 heterodimerizes with subunit B and three such heterodimers form a hexameric ring structure with alternating B and D3 subunits. The D3 - B heterodimer also assembles into a heptameric ring containing D1, D2, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=34.27 E-value=82 Score=23.10 Aligned_cols=35 Identities=14% Similarity=0.016 Sum_probs=29.9
Q ss_pred cEEEEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471 16 EALILSTWLLCSPSAPSATLVTADICIYTMLTVED 50 (371)
Q Consensus 16 sg~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~ 50 (371)
.|-.+.+.+.+|..+.+++..+|...|+.+-....
T Consensus 9 ~g~~V~VeLk~g~~~~G~L~~~D~~MNl~L~~~~~ 43 (70)
T cd01721 9 EGHIVTVELKTGEVYRGKLIEAEDNMNCQLKDVTV 43 (70)
T ss_pred CCCEEEEEECCCcEEEEEEEEEcCCceeEEEEEEE
Confidence 44566788899999999999999999999988753
No 117
>cd01727 LSm8 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm8 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=32.24 E-value=92 Score=23.05 Aligned_cols=31 Identities=3% Similarity=-0.166 Sum_probs=26.3
Q ss_pred EEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471 20 LSTWLLCSPSAPSATLVTADICIYTMLTVED 50 (371)
Q Consensus 20 i~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~ 50 (371)
+.+.+-+++++.+++.++|+..|+.+=....
T Consensus 12 V~V~l~dgr~~~G~L~~~D~~~NlvL~~~~E 42 (74)
T cd01727 12 VSVITVDGRVIVGTLKGFDQATNLILDDSHE 42 (74)
T ss_pred EEEEECCCcEEEEEEEEEccccCEEccceEE
Confidence 4477889999999999999999998877543
No 118
>KOG3627 consensus Trypsin [Amino acid transport and metabolism]
Probab=30.83 E-value=33 Score=31.28 Aligned_cols=99 Identities=17% Similarity=0.148 Sum_probs=49.8
Q ss_pred CCEEEEEEecC-CCcCCccceecCCCC----CCC-CeEEEEEeCCCC----C-CceeeeeEEeeeeee----eccCC---
Q 017471 41 CIYTMLTVEDD-EFWEGVLPVEFGELP----ALQ-DAVTVVGYPIGG----D-TISVTSGVVSRIEIL----SYVHG--- 102 (371)
Q Consensus 41 ~DlAlLkv~~~-~~~~~l~~~~l~~s~----~lg-d~V~~iG~p~g~----~-~~s~t~G~Vs~~~~~----~~~~~--- 102 (371)
.|+|+|+++.+ .|-+.+.|+.+.... ..+ ..+.+.|+.... . ........+.-+... .+...
T Consensus 106 nDiall~l~~~v~~~~~i~piclp~~~~~~~~~~~~~~~v~GWG~~~~~~~~~~~~L~~~~v~i~~~~~C~~~~~~~~~~ 185 (256)
T KOG3627|consen 106 NDIALLRLSEPVTFSSHIQPICLPSSADPYFPPGGTTCLVSGWGRTESGGGPLPDTLQEVDVPIISNSECRRAYGGLGTI 185 (256)
T ss_pred CCEEEEEECCCcccCCcccccCCCCCcccCCCCCCCEEEEEeCCCcCCCCCCCCceeEEEEEeEcChhHhcccccCcccc
Confidence 79999999974 344456666664222 223 677777764321 1 111121122111110 11100
Q ss_pred CeEEeEE---EEEeeccCCCCCCceecCC---CcEEEEEeeec
Q 017471 103 STELLGL---QIDAAINSGNSGGPAFNDK---GKCVGIAFQSL 139 (371)
Q Consensus 103 ~~~~~~i---~~da~i~~G~SGGPlvn~~---G~VIGI~~~~~ 139 (371)
....-+. .-....-.|+|||||+-.+ ..++||++...
T Consensus 186 ~~~~~Ca~~~~~~~~~C~GDSGGPLv~~~~~~~~~~GivS~G~ 228 (256)
T KOG3627|consen 186 TDTMLCAGGPEGGKDACQGDSGGPLVCEDNGRWVLVGIVSWGS 228 (256)
T ss_pred CCCEEeeCccCCCCccccCCCCCeEEEeeCCcEEEEEEEEecC
Confidence 0001011 1112234699999998765 69999987754
No 119
>cd01719 Sm_G The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit G binds subunits E and F to form a trimer which then assembles onto snRNA along with the D1/D2 and D3/B heterodimers forming a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=30.47 E-value=1.1e+02 Score=22.57 Aligned_cols=30 Identities=10% Similarity=-0.165 Sum_probs=25.2
Q ss_pred EEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471 20 LSTWLLCSPSAPSATLVTADICIYTMLTVE 49 (371)
Q Consensus 20 i~~~~~~~~~~~A~vv~~d~~~DlAlLkv~ 49 (371)
+.+.+-+|+.+.+++.++|...|+.+=...
T Consensus 13 V~V~L~~g~~~~G~L~~~D~~mNlvL~~~~ 42 (72)
T cd01719 13 LSLKLNGNRKVSGILRGFDPFMNLVLDDAV 42 (72)
T ss_pred EEEEECCCeEEEEEEEEEcccccEEeccEE
Confidence 346788999999999999999999886554
No 120
>PRK05015 aminopeptidase B; Provisional
Probab=28.96 E-value=90 Score=31.52 Aligned_cols=39 Identities=13% Similarity=0.075 Sum_probs=26.6
Q ss_pred hccCCCCCCceEEEEECCCCccccCCCCCCEEEEECCEEec
Q 017471 190 AMSMKADQKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIA 230 (371)
Q Consensus 190 ~~gl~~~~~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~ 230 (371)
.++++.+ -..|.-..++.+.....+|||||.+-||+.|.
T Consensus 230 ~~~l~~n--V~~il~~aENmisg~A~kpgDVIt~~nGkTVE 268 (424)
T PRK05015 230 TRGLNKR--VKLFLCCAENLISGNAFKLGDIITYRNGKTVE 268 (424)
T ss_pred hcCCCce--EEEEEEecccCCCCCCCCCCCEEEecCCcEEe
Confidence 3455522 22344455666665678999999999999874
No 121
>PF09465 LBR_tudor: Lamin-B receptor of TUDOR domain; InterPro: IPR019023 The Lamin-B receptor is a chromatin and lamin binding protein in the inner nuclear membrane. It is one of the integral inner nuclear envelope membrane proteins responsible for targeting nuclear membranes to chromatin, being a downstream effector of Ran, a small Ras-like nuclear GTPase which regulates NE assembly. Lamin-B receptor interacts with importin beta, a Ran-binding protein, thereby directly contributing to the fusion of membrane vesicles and the formation of the nuclear envelope []. ; PDB: 2L8D_A 2DIG_A.
Probab=28.33 E-value=2e+02 Score=20.35 Aligned_cols=37 Identities=11% Similarity=-0.101 Sum_probs=28.7
Q ss_pred ccEEEEEEEecCCCeE-eEEEEEecCCCCEEEEEEecC
Q 017471 15 NEALILSTWLLCSPSA-PSATLVTADICIYTMLTVEDD 51 (371)
Q Consensus 15 gsg~vi~~~~~~~~~~-~A~vv~~d~~~DlAlLkv~~~ 51 (371)
..|=++.++.+.+..+ +|+|...|...+++-+++++-
T Consensus 7 ~~Ge~V~~rWP~s~lYYe~kV~~~d~~~~~y~V~Y~DG 44 (55)
T PF09465_consen 7 AIGEVVMVRWPGSSLYYEGKVLSYDSKSDRYTVLYEDG 44 (55)
T ss_dssp -SS-EEEEE-TTTS-EEEEEEEEEETTTTEEEEEETTS
T ss_pred cCCCEEEEECCCCCcEEEEEEEEecccCceEEEEEcCC
Confidence 4566778888887775 999999999999999999875
No 122
>COG1958 LSM1 Small nuclear ribonucleoprotein (snRNP) homolog [Transcription]
Probab=27.12 E-value=1.2e+02 Score=22.69 Aligned_cols=32 Identities=9% Similarity=-0.089 Sum_probs=27.3
Q ss_pred EEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471 19 ILSTWLLCSPSAPSATLVTADICIYTMLTVED 50 (371)
Q Consensus 19 vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~ 50 (371)
.+.+.+-+|+.+.|++.++|+..|+.+--+..
T Consensus 19 ~V~V~lk~g~~~~G~L~~~D~~mNlvL~d~~e 50 (79)
T COG1958 19 RVLVKLKNGREYRGTLVGFDQYMNLVLDDVEE 50 (79)
T ss_pred EEEEEECCCCEEEEEEEEEccceeEEEeceEE
Confidence 44578899999999999999999998876655
No 123
>PF11874 DUF3394: Domain of unknown function (DUF3394); InterPro: IPR021814 This domain is functionally uncharacterised. This domain is found in bacteria. This presumed domain is about 190 amino acids in length. This domain is found associated with PF06808 from PFAM.
Probab=26.97 E-value=1.2e+02 Score=26.99 Aligned_cols=71 Identities=11% Similarity=-0.010 Sum_probs=40.8
Q ss_pred hhhhccCCCCEEEEEEEE---CCEEEEEEEEeccccccCCCCCCCCCCCceeeccEEEechHHHHHHcCcccccCceEEE
Q 017471 247 YLVSQKYTGDSAAVKVLR---DSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERIMNMKLRS 323 (371)
Q Consensus 247 ~~~~~~~~g~~v~l~v~R---~g~~~~~~v~l~~~~~~~~~~~~~~~~~~~~~~Gl~~~~l~~~~~~~~~~~~~~~~~v~ 323 (371)
..+.+..+|+.+.+.|.+ .|+..+.++.+.-.+....... ..-.|+.+. +....+.+.
T Consensus 67 ~~~~~~~~g~~lrl~V~G~~~~G~~~~k~v~lpl~~~~~g~eR-------L~~~GL~l~------------~e~~~~~Vd 127 (183)
T PF11874_consen 67 QVAEQLPPGSSLRLRVEGPDFEGDPVTKTVLLPLGDGADGEER-------LEAAGLTLM------------EEGGKVIVD 127 (183)
T ss_pred HHHhcCCCCCEEEEEEEccCCCCCceEEEEEEEcCCCCCHHHH-------HHhCCCEEE------------eeCCEEEEE
Confidence 456667789999999988 3555444444433222111000 001455542 223557899
Q ss_pred EEecCChHhHHhh
Q 017471 324 SFWTSSCIQCHNC 336 (371)
Q Consensus 324 ~~~~~Sp~~~~~~ 336 (371)
++..|||++-...
T Consensus 128 ~v~fgS~A~~~g~ 140 (183)
T PF11874_consen 128 EVEFGSPAEKAGI 140 (183)
T ss_pred ecCCCCHHHHcCC
Confidence 9999999876543
No 124
>PF15483 DUF4641: Domain of unknown function (DUF4641)
Probab=25.82 E-value=45 Score=33.33 Aligned_cols=21 Identities=43% Similarity=0.605 Sum_probs=17.3
Q ss_pred HHHHHHHHH-hhhhhhHHHHHH
Q 017471 344 CLRCLWLIL-ILDMRRLLTLRF 364 (371)
Q Consensus 344 ~~~~~~~~~-~~~~~~~~~~~~ 364 (371)
|+||+|||- |.|.|+-|-.+-
T Consensus 417 CpRC~~LQkEIedLreQLaamq 438 (445)
T PF15483_consen 417 CPRCLVLQKEIEDLREQLAAMQ 438 (445)
T ss_pred CcccHHHHHHHHHHHHHHHHHH
Confidence 999999985 889998776543
No 125
>COG2524 Predicted transcriptional regulator, contains C-terminal CBS domains [Transcription]
Probab=24.51 E-value=2.1e+02 Score=27.06 Aligned_cols=20 Identities=40% Similarity=0.625 Sum_probs=17.3
Q ss_pred CCCCCCceecCCCcEEEEEee
Q 017471 117 SGNSGGPAFNDKGKCVGIAFQ 137 (371)
Q Consensus 117 ~G~SGGPlvn~~G~VIGI~~~ 137 (371)
.|-.|.|++|.+ +++||.+.
T Consensus 201 ~~i~GaPVvd~d-k~vGiit~ 220 (294)
T COG2524 201 KGIRGAPVVDDD-KIVGIITL 220 (294)
T ss_pred cCccCCceecCC-ceEEEEEH
Confidence 688999999977 99999754
No 126
>cd01723 LSm4 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm4 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=24.40 E-value=1.7e+02 Score=21.69 Aligned_cols=33 Identities=6% Similarity=-0.158 Sum_probs=28.1
Q ss_pred EEEEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471 17 ALILSTWLLCSPSAPSATLVTADICIYTMLTVE 49 (371)
Q Consensus 17 g~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~ 49 (371)
|-.+.+.+.+|..+.+++..+|...|+.+-...
T Consensus 11 g~~V~VeLkng~~~~G~L~~~D~~mNi~L~~~~ 43 (76)
T cd01723 11 NHPMLVELKNGETYNGHLVNCDNWMNIHLREVI 43 (76)
T ss_pred CCEEEEEECCCCEEEEEEEEEcCCCceEEEeEE
Confidence 345667888999999999999999999987764
No 127
>PF12381 Peptidase_C3G: Tungro spherical virus-type peptidase; InterPro: IPR024387 This entry represents a rice tungro spherical waikavirus-type peptidase that belongs to MEROPS peptidase family C3G. It is a picornain 3C-type protease, and is responsible for the self-cleavage of the positive single-stranded polyproteins of a number of plant viral genomes. The location of the protease activity of the polyprotein is at the C-terminal end, adjacent and N-terminal to the putative RNA polymerase [, ].
Probab=22.99 E-value=96 Score=28.34 Aligned_cols=54 Identities=24% Similarity=0.413 Sum_probs=37.1
Q ss_pred eEEEEEeeccCCCCCCceecCC----CcEEEEEeeecccCCccceeccccCc--chhHhHhhh
Q 017471 107 LGLQIDAAINSGNSGGPAFNDK----GKCVGIAFQSLKHEDVENIGYVIPTP--VIMHFIQDY 163 (371)
Q Consensus 107 ~~i~~da~i~~G~SGGPlvn~~----G~VIGI~~~~~~~~~~~~~~~aiP~~--~i~~~l~~L 163 (371)
..++..++-..|+-|+|++=.+ -+++||..+.. .....+||-++. .+++.++.|
T Consensus 169 ~gleY~~~t~~GdCGs~i~~~~t~~~RKIvGiHVAG~---~~~~~gYAe~itQEDL~~A~~~l 228 (231)
T PF12381_consen 169 QGLEYQMPTMNGDCGSPIVRNNTQMVRKIVGIHVAGS---ANHAMGYAESITQEDLMRAINKL 228 (231)
T ss_pred eeeeEECCCcCCCccceeeEcchhhhhhhheeeeccc---ccccceehhhhhHHHHHHHHHhh
Confidence 3577888889999999986322 58999998865 235667777663 344444443
No 128
>PF01423 LSM: LSM domain ; InterPro: IPR001163 This family is found in Lsm (like-Sm) proteins and in bacterial Lsm-related Hfq proteins. In each case, the domain adopts a core structure consisting of an open beta-barrel with an SH3-like topology. Lsm (like-Sm) proteins have diverse functions, and are thought to be important modulators of RNA biogenesis and function [, ]. The Sm proteins form part of specific small nuclear ribonucleoproteins (snRNPs) that are involved in the processing of pre-mRNAs to mature mRNAs, and are a major component of the eukaryotic spliceosome. Most snRNPs consist of seven Sm proteins (B/B', D1, D2, D3, E, F and G) arranged in a ring on a uridine-rich sequence (Sm site), plus a small nuclear RNA (snRNA) (either U1, U2, U5 or U4/6) []. All Sm proteins contain a common sequence motif in two segments, Sm1 and Sm2, separated by a short variable linker []. In other snRNPs, certain Sm proteins are replaced with different Lsm proteins, such as with U7 snRNPs, in which the D1 and D2 Sm proteins are replaced with U7-specific Lsm10 and Lsm11 proteins, where Lsm11 plays a role in histone U7-specific RNA processing []. Lsm proteins are also found in archaebacteria, which do not have any splicing apparatus suggesting a more general role for Lsm proteins. The pleiotropic translational regulator Hfq (host factor Q) is a bacterial Lsm-like protein, which modulates the structure of numerous RNA molecules by binding preferentially to A/U-rich sequences in RNA []. Hfq forms an Lsm-like fold, however, unlike the heptameric Sm proteins, Hfq forms a homo-hexameric ring.; PDB: 1D3B_K 2Y9D_D 2Y9A_D 2Y9C_R 3VRI_C 2Y9B_K 3QUI_D 3M4G_H 3INZ_E 1U1S_C ....
Probab=22.54 E-value=1.9e+02 Score=20.48 Aligned_cols=33 Identities=6% Similarity=-0.040 Sum_probs=28.4
Q ss_pred EEEEEecCCCeEeEEEEEecCCCCEEEEEEecC
Q 017471 19 ILSTWLLCSPSAPSATLVTADICIYTMLTVEDD 51 (371)
Q Consensus 19 vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~~ 51 (371)
.+.+.+.+|+.+.|++.++|...|+.+-.....
T Consensus 10 ~V~V~l~~g~~~~G~L~~~D~~~Nl~L~~~~~~ 42 (67)
T PF01423_consen 10 RVRVELKNGRTYRGTLVSFDQFMNLVLSDVTET 42 (67)
T ss_dssp EEEEEETTSEEEEEEEEEEETTEEEEEEEEEEE
T ss_pred EEEEEEeCCEEEEEEEEEeechheEEeeeEEEE
Confidence 355788899999999999999999998887754
No 129
>KOG2561 consensus Adaptor protein NUB1, contains UBA domain [Posttranslational modification, protein turnover, chaperones; Signal transduction mechanisms]
Probab=22.25 E-value=28 Score=35.15 Aligned_cols=23 Identities=13% Similarity=0.144 Sum_probs=15.2
Q ss_pred CChHhHHhhhhh------hhHHHHHHHHH
Q 017471 328 SSCIQCHNCQMS------SLLWCLRCLWL 350 (371)
Q Consensus 328 ~Sp~~~~~~~~~------~~~~~~~~~~~ 350 (371)
-.+-++++.+=| ||.||-|||.-
T Consensus 195 ~Cd~klLe~VDNyallnLDIVWCYfrLkn 223 (568)
T KOG2561|consen 195 LCDSKLLELVDNYALLNLDIVWCYFRLKN 223 (568)
T ss_pred hhhHHHHHhhcchhhhhcchhheehhhcc
Confidence 344556655433 99999998853
No 130
>PF10049 DUF2283: Protein of unknown function (DUF2283); InterPro: IPR019270 Members of this family of hypothetical proteins have no known function.
Probab=21.80 E-value=62 Score=22.09 Aligned_cols=11 Identities=36% Similarity=0.866 Sum_probs=8.4
Q ss_pred cCCCcEEEEEe
Q 017471 126 NDKGKCVGIAF 136 (371)
Q Consensus 126 n~~G~VIGI~~ 136 (371)
|.+|++|||-.
T Consensus 36 d~~G~ivGIEI 46 (50)
T PF10049_consen 36 DEDGRIVGIEI 46 (50)
T ss_pred CCCCCEEEEEE
Confidence 46799999843
No 131
>COG5233 GRH1 Peripheral Golgi membrane protein [Intracellular trafficking and secretion]
Probab=21.71 E-value=47 Score=32.09 Aligned_cols=30 Identities=23% Similarity=0.424 Sum_probs=26.6
Q ss_pred EEEEECCCCcccc-CCCCCCEEEEECCEEec
Q 017471 201 RIRRVDPTAPESE-VLKPSDIILSFDGIDIA 230 (371)
Q Consensus 201 ~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~ 230 (371)
-+-+|.+.+||++ |.-.||.|+.+|+-++.
T Consensus 66 ~~lrv~~~~~~e~~~~~~~dyilg~n~Dp~~ 96 (417)
T COG5233 66 EVLRVNPESPAEKAGMVVGDYILGINEDPLR 96 (417)
T ss_pred hheeccccChhHhhccccceeEEeecCCcHH
Confidence 4568899999999 99999999999988875
No 132
>PF14827 Cache_3: Sensory domain of two-component sensor kinase; PDB: 1OJG_A 3BY8_A 1P0Z_I 2V9A_A 2J80_B.
Probab=21.67 E-value=77 Score=25.40 Aligned_cols=17 Identities=24% Similarity=0.718 Sum_probs=11.7
Q ss_pred ceecCCCcEEEEEeeec
Q 017471 123 PAFNDKGKCVGIAFQSL 139 (371)
Q Consensus 123 Plvn~~G~VIGI~~~~~ 139 (371)
|+.|.+|+++|++.-.+
T Consensus 95 PV~d~~g~viG~V~VG~ 111 (116)
T PF14827_consen 95 PVYDSDGKVIGVVSVGV 111 (116)
T ss_dssp EEE-TTS-EEEEEEEEE
T ss_pred eeECCCCcEEEEEEEEE
Confidence 66788999999986543
No 133
>TIGR00074 hypC_hupF hydrogenase assembly chaperone HypC/HupF. An additional proposed function is to shuttle the iron atom that has been liganded at the HypC/HypD complex to the precursor of the large hydrogenase (HycE) subunit. PubMed:12441107.
Probab=21.59 E-value=1.6e+02 Score=22.14 Aligned_cols=41 Identities=10% Similarity=0.063 Sum_probs=27.0
Q ss_pred EeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCCCCCCeEEEE
Q 017471 30 APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELPALQDAVTVV 75 (371)
Q Consensus 30 ~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~~lgd~V~~i 75 (371)
+|++|+.++. +.|++.+.... .--.+.+-+..++||+|++-
T Consensus 5 iP~~V~~i~~--~~A~v~~~G~~---~~v~l~lv~~~~vGD~VLVH 45 (76)
T TIGR00074 5 IPGQVVEIDE--NIALVEFCGIK---RDVSLDLVGEVKVGDYVLVH 45 (76)
T ss_pred cceEEEEEcC--CEEEEEcCCeE---EEEEEEeeCCCCCCCEEEEe
Confidence 6899999876 57888887542 11122333456789998874
No 134
>PF15436 PGBA_N: Plasminogen-binding protein pgbA N-terminal
Probab=21.52 E-value=3.4e+02 Score=24.84 Aligned_cols=59 Identities=8% Similarity=0.042 Sum_probs=37.1
Q ss_pred eccEEEEEEEecCCCeEeEEEEEecCCCCEEEEEEecCCCcC--CccceecCCCCCCCCeEEE
Q 017471 14 RNEALILSTWLLCSPSAPSATLVTADICIYTMLTVEDDEFWE--GVLPVEFGELPALQDAVTV 74 (371)
Q Consensus 14 ~gsg~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~~~~~~--~l~~~~l~~s~~lgd~V~~ 74 (371)
+.||+|+.-.--+-..+-|+++....+.+.|.+|+.+-+-.+ .+|.... .++.||+|+.
T Consensus 28 G~SGiV~h~~~~~~~~IiA~a~V~~~~~g~A~~kf~~fd~L~Q~aLP~p~~--~pk~GD~vil 88 (218)
T PF15436_consen 28 GESGIVVHKFDKDHSSIIARAVVISKKNGVAKAKFSVFDSLKQDALPTPKM--VPKKGDEVIL 88 (218)
T ss_pred CCceEEEEEecCCcceeeeEEEEEEecCCeeEEEEeehhhhhhhcCCCCcc--ccCCCCEEEE
Confidence 567777764445666677887777778999999998653111 2222222 3567877664
No 135
>cd00991 PDZ_archaeal_metalloprotease PDZ domain of archaeal zinc metalloprotases, presumably membrane-associated or integral membrane proteases, which may be involved in signalling and regulatory mechanisms. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=21.37 E-value=55 Score=24.23 Aligned_cols=27 Identities=4% Similarity=-0.089 Sum_probs=20.7
Q ss_pred ccCceEEEEEecCChHhHHhhhhhhhH
Q 017471 316 IMNMKLRSSFWTSSCIQCHNCQMSSLL 342 (371)
Q Consensus 316 ~~~~~~v~~~~~~Sp~~~~~~~~~~~~ 342 (371)
...|+++..+.++||++-.-.+.+|+|
T Consensus 8 ~~~Gv~V~~V~~~spa~~aGL~~GDiI 34 (79)
T cd00991 8 AVAGVVIVGVIVGSPAENAVLHTGDVI 34 (79)
T ss_pred cCCcEEEEEECCCChHHhcCCCCCCEE
Confidence 346888899999999987767767664
No 136
>PF11730 DUF3297: Protein of unknown function (DUF3297); InterPro: IPR021724 This family is expressed in Proteobacteria and Actinobacteria. The function is not known.
Probab=21.35 E-value=58 Score=23.91 Aligned_cols=59 Identities=17% Similarity=0.291 Sum_probs=38.6
Q ss_pred EECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEE
Q 017471 204 RVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT 274 (371)
Q Consensus 204 ~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~ 274 (371)
++.|.||-.. -+-.-|+=+.+||+.=++.+++-+..| ..+-..|+ ...|.|+.++++++
T Consensus 5 S~~P~Sp~~~~~~l~~~iGIrfng~Er~nVeEYciSEG--------Wvrv~~gk----a~DR~G~Pl~iklk 64 (71)
T PF11730_consen 5 SINPRSPHYDAEVLERGIGIRFNGKERTNVEEYCISEG--------WVRVAAGK----ALDRRGNPLTIKLK 64 (71)
T ss_pred ccCCCChhhHHHHHhcCcceEECCeEcccceeEeccCC--------EEEeecCc----ccccCCCeeEEEEc
Confidence 4678898877 667778889999999999988752211 11122232 33577777666553
No 137
>PRK09570 rpoH DNA-directed RNA polymerase subunit H; Reviewed
Probab=20.77 E-value=59 Score=24.74 Aligned_cols=19 Identities=26% Similarity=0.529 Sum_probs=14.0
Q ss_pred EECCCCcccc--CCCCCCEEE
Q 017471 204 RVDPTAPESE--VLKPSDIIL 222 (371)
Q Consensus 204 ~V~~~spA~~--GL~~GDvIl 222 (371)
.+...-|+++ |+++||+|-
T Consensus 39 ~I~~~DPv~r~~g~k~GdVvk 59 (79)
T PRK09570 39 KIKASDPVVKAIGAKPGDVIK 59 (79)
T ss_pred ceeccChhhhhcCCCCCCEEE
Confidence 3445567766 999999984
No 138
>cd00433 Peptidase_M17 Cytosol aminopeptidase family, N-terminal and catalytic domains. Family M17 contains zinc- and manganese-dependent exopeptidases ( EC 3.4.11.1), including leucine aminopeptidase. They catalyze removal of amino acids from the N-terminus of a protein and play a key role in protein degradation and in the metabolism of biologically active peptides. They do not contain HEXXH motif (which is used as one of the signature patterns to group the peptidase families) in the metal-binding site. The two associated zinc ions and the active site are entirely enclosed within the C-terminal catalytic domain in leucine aminopeptidase. The enzyme is a hexamer, with the catalytic domains clustered around the three-fold axis, and the two trimers related to one another by a two-fold rotation. The N-terminal domain is structurally similar to the ADP-ribose binding Macro domain. This family includes proteins from bacteria, archaea, animals and plants.
Probab=20.75 E-value=1.1e+02 Score=31.55 Aligned_cols=28 Identities=18% Similarity=0.299 Sum_probs=21.9
Q ss_pred EEECCCCccccCCCCCCEEEEECCEEec
Q 017471 203 RRVDPTAPESEVLKPSDIILSFDGIDIA 230 (371)
Q Consensus 203 ~~V~~~spA~~GL~~GDvIl~vnG~~V~ 230 (371)
.-..+|.+.....+|||||.+.||+.|.
T Consensus 290 ~~~~EN~is~~A~rPgDVi~s~~GkTVE 317 (468)
T cd00433 290 LPLAENMISGNAYRPGDVITSRSGKTVE 317 (468)
T ss_pred EEeeecCCCCCCCCCCCEeEeCCCcEEE
Confidence 3445666666678999999999999874
No 139
>PRK00913 multifunctional aminopeptidase A; Provisional
Probab=20.67 E-value=1.1e+02 Score=31.55 Aligned_cols=27 Identities=22% Similarity=0.465 Sum_probs=21.7
Q ss_pred EECCCCccccCCCCCCEEEEECCEEec
Q 017471 204 RVDPTAPESEVLKPSDIILSFDGIDIA 230 (371)
Q Consensus 204 ~V~~~spA~~GL~~GDvIl~vnG~~V~ 230 (371)
-..++.|...-.+|||||++.||+.|.
T Consensus 305 ~l~ENm~~~~A~rPgDVi~~~~GkTVE 331 (483)
T PRK00913 305 AACENMPSGNAYRPGDVLTSMSGKTIE 331 (483)
T ss_pred EeeccCCCCCCCCCCCEEEECCCcEEE
Confidence 345667766678999999999999875
No 140
>cd01718 Sm_E The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit E binds subunits F and G to form a trimer which then assembles onto snRNA along with the D1/D2 and D3/B heterodimers forming a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=20.61 E-value=1.9e+02 Score=21.97 Aligned_cols=30 Identities=10% Similarity=0.152 Sum_probs=24.0
Q ss_pred EEEEec--CCCeEeEEEEEecCCCCEEEEEEe
Q 017471 20 LSTWLL--CSPSAPSATLVTADICIYTMLTVE 49 (371)
Q Consensus 20 i~~~~~--~~~~~~A~vv~~d~~~DlAlLkv~ 49 (371)
+.+++. +++++.+++.++|...|+.+=...
T Consensus 21 V~V~l~~~~g~~~~G~L~gfD~~mNlvL~d~~ 52 (79)
T cd01718 21 VQIWLYEQTDLRIEGVIIGFDEYMNLVLDDAE 52 (79)
T ss_pred EEEEEEeCCCcEEEEEEEEEccceeEEEcCEE
Confidence 345554 899999999999999999876543
No 141
>cd01725 LSm2 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm2 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=20.49 E-value=2e+02 Score=21.77 Aligned_cols=33 Identities=6% Similarity=-0.092 Sum_probs=28.6
Q ss_pred EEEEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471 17 ALILSTWLLCSPSAPSATLVTADICIYTMLTVE 49 (371)
Q Consensus 17 g~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~ 49 (371)
|-.+.+.+.+|..+.+++..+|...|+-+=.+.
T Consensus 11 g~~V~VeLKng~~~~G~L~~vD~~MNi~L~n~~ 43 (81)
T cd01725 11 GKEVTVELKNDLSIRGTLHSVDQYLNIKLTNIS 43 (81)
T ss_pred CCEEEEEECCCcEEEEEEEEECCCcccEEEEEE
Confidence 445668889999999999999999999888775
No 142
>cd01724 Sm_D1 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit D1 heterodimerizes with subunit D2 and three such heterodimers form a hexameric ring structure with alternating D1 and D2 subunits. The D1 - D2 heterodimer also assembles into a heptameric ring containing DB, D3, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=20.49 E-value=1.8e+02 Score=22.51 Aligned_cols=35 Identities=6% Similarity=-0.175 Sum_probs=30.1
Q ss_pred cEEEEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471 16 EALILSTWLLCSPSAPSATLVTADICIYTMLTVED 50 (371)
Q Consensus 16 sg~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~ 50 (371)
.|-.+.+.+.+|..+.+++..+|...|+.+-.+..
T Consensus 10 ~g~~V~VeLKng~~~~G~L~~vD~~MNl~L~~a~~ 44 (90)
T cd01724 10 TNETVTIELKNGTIVHGTITGVDPSMNTHLKNVKL 44 (90)
T ss_pred CCCEEEEEECCCCEEEEEEEEEcCceeEEEEEEEE
Confidence 45567788999999999999999999999988753
Done!