Query 009784
Match_columns 526
No_of_seqs 483 out of 3258
Neff 7.3
Searched_HMMs 46136
Date Thu Mar 28 17:17:35 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/009784.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/009784hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PRK10139 serine endoprotease; 100.0 2.7E-51 5.9E-56 438.4 38.6 367 111-503 42-431 (455)
2 TIGR02037 degP_htrA_DO peripla 100.0 1.4E-48 3E-53 417.1 37.6 330 147-502 56-402 (428)
3 PRK10942 serine endoprotease; 100.0 2.2E-48 4.7E-53 418.0 36.2 329 148-502 110-448 (473)
4 TIGR02038 protease_degS peripl 100.0 1.6E-47 3.4E-52 398.1 34.5 296 111-433 47-349 (351)
5 PRK10898 serine endoprotease; 100.0 7.6E-47 1.6E-51 392.8 34.0 297 111-434 47-351 (353)
6 COG0265 DegQ Trypsin-like seri 100.0 8.5E-37 1.8E-41 317.9 28.5 300 111-433 35-341 (347)
7 KOG1421 Predicted signaling-as 100.0 1.7E-31 3.6E-36 281.0 20.4 322 112-463 52-395 (955)
8 KOG1320 Serine protease [Postt 100.0 1.2E-29 2.7E-34 265.5 15.8 372 118-503 56-439 (473)
9 KOG1320 Serine protease [Postt 99.9 8.3E-25 1.8E-29 229.3 20.3 301 118-432 134-468 (473)
10 KOG1421 Predicted signaling-as 99.8 2E-17 4.2E-22 175.5 24.2 327 124-492 530-891 (955)
11 PF13365 Trypsin_2: Trypsin-li 99.6 6E-15 1.3E-19 128.7 14.1 108 151-289 1-120 (120)
12 PF13180 PDZ_2: PDZ domain; PD 99.5 4E-14 8.7E-19 116.6 7.4 81 328-430 1-82 (82)
13 PF00089 Trypsin: Trypsin; In 99.4 1.5E-12 3.2E-17 125.1 14.4 185 131-315 4-220 (220)
14 cd00190 Tryp_SPc Trypsin-like 99.3 5E-11 1.1E-15 115.2 16.8 161 134-294 7-208 (232)
15 cd00987 PDZ_serine_protease PD 99.3 1.5E-11 3.3E-16 102.3 8.5 88 328-427 1-89 (90)
16 smart00020 Tryp_SPc Trypsin-li 99.2 1.5E-10 3.3E-15 112.2 13.7 164 131-294 5-208 (229)
17 cd00986 PDZ_LON_protease PDZ d 99.2 8.6E-11 1.9E-15 95.9 8.0 72 352-433 7-78 (79)
18 cd00991 PDZ_archaeal_metallopr 99.1 1.2E-10 2.6E-15 95.2 7.8 69 351-429 8-77 (79)
19 TIGR01713 typeII_sec_gspC gene 99.1 4E-10 8.6E-15 112.5 11.4 100 309-430 159-259 (259)
20 cd00990 PDZ_glycyl_aminopeptid 99.1 4.5E-10 9.8E-15 91.5 8.9 77 328-431 1-78 (80)
21 PRK10779 zinc metallopeptidase 99.0 8.9E-10 1.9E-14 118.9 9.0 132 355-503 128-262 (449)
22 TIGR02037 degP_htrA_DO peripla 98.9 3.3E-09 7.2E-14 113.9 8.6 90 327-427 337-427 (428)
23 cd00989 PDZ_metalloprotease PD 98.9 4.1E-09 8.8E-14 85.5 6.8 66 353-429 12-78 (79)
24 cd00988 PDZ_CTP_protease PDZ d 98.9 1E-08 2.2E-13 84.4 8.3 68 352-430 12-83 (85)
25 COG3591 V8-like Glu-specific e 98.8 9.2E-08 2E-12 94.1 14.8 161 149-319 64-250 (251)
26 cd00136 PDZ PDZ domain, also c 98.6 8.4E-08 1.8E-12 75.9 5.8 55 353-418 13-70 (70)
27 TIGR00054 RIP metalloprotease 98.5 1.7E-07 3.6E-12 100.4 7.6 113 352-501 127-242 (420)
28 TIGR00054 RIP metalloprotease 98.5 2.4E-07 5.2E-12 99.2 7.2 69 353-432 203-272 (420)
29 smart00228 PDZ Domain present 98.5 4.8E-07 1E-11 73.8 7.2 59 353-421 26-85 (85)
30 PRK10779 zinc metallopeptidase 98.4 6E-07 1.3E-11 97.0 7.2 68 354-432 222-290 (449)
31 TIGR00225 prc C-terminal pepti 98.3 1.3E-06 2.8E-11 90.8 8.3 70 353-433 62-134 (334)
32 PF00595 PDZ: PDZ domain (Also 98.3 1.5E-06 3.2E-11 71.0 5.4 72 327-418 9-81 (81)
33 PRK10139 serine endoprotease; 98.2 2.1E-06 4.6E-11 92.8 6.7 64 353-428 390-454 (455)
34 TIGR03279 cyano_FeS_chp putati 98.2 1.9E-06 4.2E-11 90.9 6.2 62 357-432 2-65 (433)
35 KOG3627 Trypsin [Amino acid tr 98.2 9.9E-05 2.2E-09 73.2 17.3 167 128-295 13-229 (256)
36 PLN00049 carboxyl-terminal pro 98.1 4.9E-06 1.1E-10 88.3 8.0 69 353-430 102-171 (389)
37 PF14685 Tricorn_PDZ: Tricorn 98.1 1.6E-05 3.5E-10 66.1 9.1 65 352-427 11-87 (88)
38 TIGR02860 spore_IV_B stage IV 98.1 4E-06 8.6E-11 88.0 6.3 69 352-431 104-181 (402)
39 PRK10942 serine endoprotease; 98.1 5.5E-06 1.2E-10 90.0 6.8 64 353-428 408-472 (473)
40 cd00992 PDZ_signaling PDZ doma 98.0 9.4E-06 2E-10 65.9 5.6 52 328-390 12-66 (82)
41 PF00863 Peptidase_C4: Peptida 98.0 0.00061 1.3E-08 66.8 18.2 169 119-317 14-195 (235)
42 COG3480 SdrC Predicted secrete 97.9 1.4E-05 3.1E-10 80.0 6.1 72 352-433 129-201 (342)
43 COG0793 Prc Periplasmic protea 97.9 3E-05 6.6E-10 82.5 8.6 84 327-434 99-185 (406)
44 PF05579 Peptidase_S32: Equine 97.8 0.00024 5.1E-09 69.8 11.7 115 149-294 112-229 (297)
45 PRK09681 putative type II secr 97.8 3.9E-05 8.4E-10 76.8 6.0 62 359-430 210-275 (276)
46 KOG3129 26S proteasome regulat 97.7 6.7E-05 1.5E-09 70.9 6.6 73 354-434 140-213 (231)
47 PF04495 GRASP55_65: GRASP55/6 97.5 0.00034 7.4E-09 63.3 7.2 87 327-432 25-115 (138)
48 PF00548 Peptidase_C3: 3C cyst 97.4 0.0047 1E-07 58.2 14.2 138 147-293 23-170 (172)
49 PRK11186 carboxy-terminal prot 97.4 0.00068 1.5E-08 76.2 9.9 71 353-429 255-332 (667)
50 COG3975 Predicted protease wit 97.3 0.00056 1.2E-08 73.1 7.7 85 330-434 439-526 (558)
51 COG3031 PulC Type II secretory 97.0 0.00067 1.5E-08 65.5 3.7 66 354-429 208-274 (275)
52 PF03761 DUF316: Domain of unk 96.8 0.039 8.4E-07 55.8 15.2 107 195-312 160-272 (282)
53 KOG3553 Tax interaction protei 96.7 0.0018 3.8E-08 54.3 3.5 35 351-385 57-92 (124)
54 PF08192 Peptidase_S64: Peptid 96.5 0.03 6.4E-07 61.8 12.5 117 195-318 542-688 (695)
55 COG5640 Secreted trypsin-like 96.4 0.062 1.3E-06 55.3 13.1 59 149-207 61-135 (413)
56 PF12812 PDZ_1: PDZ-like domai 96.3 0.0044 9.6E-08 50.5 3.8 60 328-391 9-69 (78)
57 PF10459 Peptidase_S46: Peptid 96.3 0.0048 1E-07 69.8 5.4 55 264-318 624-686 (698)
58 PF02122 Peptidase_S39: Peptid 96.2 0.0035 7.5E-08 60.4 3.0 137 159-310 42-183 (203)
59 KOG3209 WW domain-containing p 95.8 0.031 6.6E-07 61.7 8.5 140 351-502 672-819 (984)
60 KOG3209 WW domain-containing p 95.8 0.0092 2E-07 65.6 4.5 152 357-518 782-980 (984)
61 KOG3580 Tight junction protein 95.5 0.011 2.3E-07 63.9 3.5 61 351-419 427-488 (1027)
62 KOG3580 Tight junction protein 94.7 0.039 8.4E-07 59.7 4.9 74 345-429 212-287 (1027)
63 PF00949 Peptidase_S7: Peptida 94.6 0.043 9.3E-07 49.2 4.2 29 266-294 90-118 (132)
64 KOG3550 Receptor targeting pro 94.0 0.1 2.2E-06 47.2 5.2 36 352-387 114-151 (207)
65 KOG3834 Golgi reassembly stack 93.3 0.36 7.8E-06 50.8 8.5 114 352-489 14-137 (462)
66 PF09342 DUF1986: Domain of un 93.2 1.2 2.6E-05 43.9 11.4 99 135-234 12-131 (267)
67 KOG3532 Predicted protein kina 93.0 0.1 2.2E-06 57.4 4.2 39 353-391 398-437 (1051)
68 KOG2921 Intramembrane metallop 91.8 0.12 2.5E-06 53.8 2.7 40 351-390 218-259 (484)
69 KOG3605 Beta amyloid precursor 91.1 0.37 8.1E-06 53.0 5.8 103 272-386 679-790 (829)
70 PF00944 Peptidase_S3: Alphavi 90.8 0.18 3.9E-06 44.8 2.5 27 268-294 101-127 (158)
71 PF05580 Peptidase_S55: SpoIVB 90.4 6.6 0.00014 38.1 12.9 41 268-311 175-215 (218)
72 COG0750 Predicted membrane-ass 90.4 0.47 1E-05 49.9 5.8 57 358-425 134-195 (375)
73 KOG1892 Actin filament-binding 90.3 0.34 7.4E-06 55.4 4.7 61 351-421 958-1020(1629)
74 PF02907 Peptidase_S29: Hepati 90.3 0.37 8.1E-06 42.9 4.0 41 270-311 105-146 (148)
75 KOG3542 cAMP-regulated guanine 90.2 0.21 4.6E-06 54.9 3.0 42 347-388 556-598 (1283)
76 KOG3552 FERM domain protein FR 89.4 0.32 6.9E-06 55.5 3.6 57 354-420 76-132 (1298)
77 KOG3571 Dishevelled 3 and rela 87.5 0.66 1.4E-05 49.8 4.3 38 352-389 276-315 (626)
78 KOG3551 Syntrophins (type beta 87.3 0.4 8.6E-06 49.8 2.5 56 355-420 112-171 (506)
79 PF00947 Pico_P2A: Picornaviru 86.7 2 4.4E-05 38.0 6.2 31 262-293 79-109 (127)
80 KOG3605 Beta amyloid precursor 84.9 2.3 4.9E-05 47.2 6.8 119 359-505 679-799 (829)
81 PF10459 Peptidase_S46: Peptid 84.1 0.53 1.2E-05 53.6 1.8 22 149-170 47-69 (698)
82 PF02395 Peptidase_S6: Immunog 81.4 7.2 0.00016 45.1 9.5 161 151-318 67-266 (769)
83 KOG3606 Cell polarity protein 81.3 1.8 3.9E-05 43.0 4.0 64 318-386 164-229 (358)
84 PF12812 PDZ_1: PDZ-like domai 80.6 2.8 6E-05 34.1 4.3 56 446-501 5-69 (78)
85 KOG3549 Syntrophins (type gamm 80.3 1.8 3.9E-05 44.4 3.7 55 354-418 81-137 (505)
86 KOG3651 Protein kinase C, alph 79.4 2.7 5.9E-05 42.4 4.6 37 354-390 31-69 (429)
87 PF03510 Peptidase_C24: 2C end 75.7 9.6 0.00021 32.8 6.3 54 153-220 3-56 (105)
88 PF01732 DUF31: Putative pepti 71.0 3.2 7E-05 43.9 2.9 24 269-292 351-374 (374)
89 KOG0609 Calcium/calmodulin-dep 70.2 6.7 0.00014 42.8 5.0 57 354-420 147-205 (542)
90 KOG0606 Microtubule-associated 69.6 4.9 0.00011 47.4 4.1 34 355-388 660-694 (1205)
91 TIGR02860 spore_IV_B stage IV 68.0 3.6 7.9E-05 43.8 2.5 42 268-312 355-396 (402)
92 KOG1924 RhoA GTPase effector D 54.4 42 0.00091 38.5 7.6 10 197-206 719-728 (1102)
93 PF05416 Peptidase_C37: Southa 54.3 95 0.0021 33.3 9.8 135 149-294 379-527 (535)
94 smart00384 AT_hook DNA binding 46.0 13 0.00028 23.6 1.2 16 5-20 1-16 (26)
95 PF12381 Peptidase_C3G: Tungro 45.5 30 0.00065 33.6 4.3 54 263-319 170-229 (231)
96 PF13180 PDZ_2: PDZ domain; PD 42.8 63 0.0014 25.7 5.3 50 452-502 3-54 (82)
97 KOG1924 RhoA GTPase effector D 40.7 80 0.0017 36.4 7.1 9 12-20 502-510 (1102)
98 KOG3938 RGS-GAIP interacting p 39.7 15 0.00032 36.8 1.2 58 355-420 151-210 (334)
99 cd01720 Sm_D2 The eukaryotic S 37.1 63 0.0014 26.8 4.4 37 167-204 10-46 (87)
100 cd00600 Sm_like The eukaryotic 34.5 1.1E+02 0.0023 23.0 5.1 33 172-205 7-39 (63)
101 KOG3834 Golgi reassembly stack 31.5 37 0.00081 36.2 2.7 65 356-431 112-180 (462)
102 PF02178 AT_hook: AT hook moti 31.2 21 0.00045 19.0 0.4 11 5-15 1-11 (13)
103 PF09465 LBR_tudor: Lamin-B re 30.5 2.1E+02 0.0046 21.7 5.8 38 169-206 7-44 (55)
104 TIGR03000 plancto_dom_1 Planct 30.4 1.5E+02 0.0032 24.0 5.3 49 372-429 10-62 (75)
105 PF00571 CBS: CBS domain CBS d 29.8 45 0.00099 24.1 2.3 21 272-292 28-48 (57)
106 cd01726 LSm6 The eukaryotic Sm 29.3 1.2E+02 0.0026 23.5 4.7 32 172-204 11-42 (67)
107 PRK00737 small nuclear ribonuc 28.9 1.2E+02 0.0027 23.9 4.8 33 172-205 15-47 (72)
108 cd01731 archaeal_Sm1 The archa 28.8 1.3E+02 0.0028 23.4 4.8 33 172-205 11-43 (68)
109 cd01722 Sm_F The eukaryotic Sm 28.5 1.2E+02 0.0025 23.7 4.5 32 172-204 12-43 (68)
110 cd01730 LSm3 The eukaryotic Sm 27.1 1.1E+02 0.0024 24.9 4.3 31 172-203 12-42 (82)
111 cd01717 Sm_B The eukaryotic Sm 26.2 1.3E+02 0.0029 24.1 4.6 32 172-204 11-42 (79)
112 COG0298 HypC Hydrogenase matur 26.0 1.5E+02 0.0033 24.3 4.7 47 185-234 5-53 (82)
113 cd06168 LSm9 The eukaryotic Sm 25.6 1.6E+02 0.0035 23.6 4.9 32 172-204 11-42 (75)
114 cd01732 LSm5 The eukaryotic Sm 24.7 1.5E+02 0.0032 23.9 4.5 31 172-203 14-44 (76)
115 cd01729 LSm7 The eukaryotic Sm 24.6 1.6E+02 0.0034 24.0 4.8 32 172-204 13-44 (81)
116 cd01735 LSm12_N LSm12 belongs 24.6 2.6E+02 0.0055 21.7 5.6 33 172-205 7-39 (61)
117 PF11874 DUF3394: Domain of un 24.0 71 0.0015 30.4 2.9 28 352-379 121-149 (183)
118 COG2524 Predicted transcriptio 23.4 2.2E+02 0.0048 28.7 6.2 94 193-292 110-220 (294)
119 cd01719 Sm_G The eukaryotic Sm 22.7 1.9E+02 0.0042 22.9 4.8 32 172-204 11-42 (72)
120 PF00595 PDZ: PDZ domain (Also 22.5 1.8E+02 0.004 22.8 4.8 52 453-504 13-67 (81)
121 cd01727 LSm8 The eukaryotic Sm 21.9 3.4E+02 0.0073 21.5 6.1 33 172-205 10-42 (74)
122 cd01721 Sm_D3 The eukaryotic S 21.7 2.1E+02 0.0046 22.4 4.9 32 172-204 11-42 (70)
123 smart00651 Sm snRNP Sm protein 21.3 2.2E+02 0.0049 21.6 4.9 33 172-205 9-41 (67)
124 cd01728 LSm1 The eukaryotic Sm 20.8 2.2E+02 0.0047 22.8 4.7 32 172-204 13-44 (74)
125 PF01423 LSM: LSM domain ; In 20.7 1.7E+02 0.0037 22.3 4.1 34 172-206 9-42 (67)
126 COG0260 PepB Leucyl aminopepti 20.2 1.3E+02 0.0028 33.1 4.4 46 358-406 303-348 (485)
No 1
>PRK10139 serine endoprotease; Provisional
Probab=100.00 E-value=2.7e-51 Score=438.40 Aligned_cols=367 Identities=23% Similarity=0.326 Sum_probs=290.3
Q ss_pred ChhhhhhhcccCCCeEEEEeeeeCCC-------------CCCccccCCCcceEEEEEEEe--CCEEEecccccCCCCeEE
Q 009784 111 VEPGVARVVPAMDAVVKVFCVHTEPN-------------FSLPWQRKRQYSSSSSGFAIG--GRRVLTNAHSVEHYTQVK 175 (526)
Q Consensus 111 ~~~~v~~~~~~~~SVV~I~~~~~~~~-------------~~~P~~~~~~~~~~GSGfvI~--~g~ILT~aHvV~~~~~i~ 175 (526)
+.+.++++.| |||.|.+...... ...||.......+.||||+|+ +||||||+|||.++..+.
T Consensus 42 ~~~~~~~~~p---avV~i~~~~~~~~~~~~~~~~~~~f~~~~~~~~~~~~~~~GSG~ii~~~~g~IlTn~HVv~~a~~i~ 118 (455)
T PRK10139 42 LAPMLEKVLP---AVVSVRVEGTASQGQKIPEEFKKFFGDDLPDQPAQPFEGLGSGVIIDAAKGYVLTNNHVINQAQKIS 118 (455)
T ss_pred HHHHHHHhCC---cEEEEEEEEeecccccCchhHHHhccccCCccccccccceEEEEEEECCCCEEEeChHHhCCCCEEE
Confidence 4455555555 9999987653221 011333333345789999997 589999999999999999
Q ss_pred EEEcCCCcEEEEEEEEeccCCCeEEEEecccccccCceeeecCCCC--cCCCcEEEEeeCCCCCceeEEEEEEeceeeee
Q 009784 176 LKKRGSDTKYLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGELP--ALQDAVTVVGYPIGGDTISVTSGVVSRIEILS 253 (526)
Q Consensus 176 V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~~--~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~ 253 (526)
|++. |++.++|++++.|+.+||||||++... .+++++|+++. ++|++|+++|||++... +++.|+||+..+..
T Consensus 119 V~~~-dg~~~~a~vvg~D~~~DlAvlkv~~~~---~l~~~~lg~s~~~~~G~~V~aiG~P~g~~~-tvt~GivS~~~r~~ 193 (455)
T PRK10139 119 IQLN-DGREFDAKLIGSDDQSDIALLQIQNPS---KLTQIAIADSDKLRVGDFAVAVGNPFGLGQ-TATSGIISALGRSG 193 (455)
T ss_pred EEEC-CCCEEEEEEEEEcCCCCEEEEEecCCC---CCceeEecCccccCCCCEEEEEecCCCCCC-ceEEEEEccccccc
Confidence 9997 999999999999999999999998643 67899999765 57999999999999776 89999999987643
Q ss_pred ccCCceeeeEEEEcccCCCCCCCCeeecCCCeEEEEEeeccccC-ccccccccccHHHHHHHHHHHHHcCceeeccccCc
Q 009784 254 YVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-DVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGV 332 (526)
Q Consensus 254 ~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~-~~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi 332 (526)
... .....+||+|+++++|||||||+|.+|+||||+++.+... +..+++||||++.+++++++|+++|++. ++|||+
T Consensus 194 ~~~-~~~~~~iqtda~in~GnSGGpl~n~~G~vIGi~~~~~~~~~~~~gigfaIP~~~~~~v~~~l~~~g~v~-r~~LGv 271 (455)
T PRK10139 194 LNL-EGLENFIQTDASINRGNSGGALLNLNGELIGINTAILAPGGGSVGIGFAIPSNMARTLAQQLIDFGEIK-RGLLGI 271 (455)
T ss_pred cCC-CCcceEEEECCccCCCCCcceEECCCCeEEEEEEEEEcCCCCccceEEEEEhHHHHHHHHHHhhcCccc-ccceeE
Confidence 221 1234589999999999999999999999999999877543 3578999999999999999999999998 999999
Q ss_pred eeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCC
Q 009784 333 EWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGD 411 (526)
Q Consensus 333 ~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~ 411 (526)
.++++ +++.++.+|++ ...|++|.+|.++|||++ |||+||+|++|||++|.++.++. ..+....+|+
T Consensus 272 ~~~~l-~~~~~~~lgl~-~~~Gv~V~~V~~~SpA~~AGL~~GDvIl~InG~~V~s~~dl~----------~~l~~~~~g~ 339 (455)
T PRK10139 272 KGTEM-SADIAKAFNLD-VQRGAFVSEVLPNSGSAKAGVKAGDIITSLNGKPLNSFAELR----------SRIATTEPGT 339 (455)
T ss_pred EEEEC-CHHHHHhcCCC-CCCceEEEEECCCChHHHCCCCCCCEEEEECCEECCCHHHHH----------HHHHhcCCCC
Confidence 99999 88999999997 467999999999999999 99999999999999999999874 6666667899
Q ss_pred EEEEEEEECCEEEEEEEEecccccccCCCCCCCCCceEEEeeEEEEecccc--ceeeeeeeecc--hhhhccccccceee
Q 009784 412 SAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLYL--ISVLSMERIMN--MKLRSSFWTSSCIQ 487 (526)
Q Consensus 412 ~v~l~v~R~G~~~~~~v~l~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt~~--~~~~~~~~i~~--~~~~sg~~~~~~~~ 487 (526)
++.++|+|+|+.+++++++...+...... ....+ .+.|+.+.+.... ...+.+..+.+ .+.++||+.||.|.
T Consensus 340 ~v~l~V~R~G~~~~l~v~~~~~~~~~~~~-~~~~~---~~~g~~l~~~~~~~~~~Gv~V~~V~~~spA~~aGL~~GD~I~ 415 (455)
T PRK10139 340 KVKLGLLRNGKPLEVEVTLDTSTSSSASA-EMITP---ALQGATLSDGQLKDGTKGIKIDEVVKGSPAAQAGLQKDDVII 415 (455)
T ss_pred EEEEEEEECCEEEEEEEEECCCCCccccc-ccccc---cccccEecccccccCCCceEEEEeCCCChHHHcCCCCCCEEE
Confidence 99999999999999999985443211110 00111 1234444432110 12234445544 45679999999999
Q ss_pred eecccchhhhHHHHHH
Q 009784 488 CHNCQMSSLLWCLRCL 503 (526)
Q Consensus 488 ~~~~~~~~~~~~~~~~ 503 (526)
..|.+..++...|+.+
T Consensus 416 ~Ing~~v~~~~~~~~~ 431 (455)
T PRK10139 416 GVNRDRVNSIAEMRKV 431 (455)
T ss_pred EECCEEcCCHHHHHHH
Confidence 9999998888877654
No 2
>TIGR02037 degP_htrA_DO periplasmic serine protease, Do/DeqQ family. This family consists of a set proteins various designated DegP, heat shock protein HtrA, and protease DO. The ortholog in Pseudomonas aeruginosa is designated MucD and is found in an operon that controls mucoid phenotype. This family also includes the DegQ (HhoA) paralog in E. coli which can rescue a DegP mutant, but not the smaller DegS paralog, which cannot. Members of this family are located in the periplasm and have separable functions as both protease and chaperone. Members have a trypsin domain and two copies of a PDZ domain. This protein protects bacteria from thermal and other stresses and may be important for the survival of bacterial pathogens.// The chaperone function is dominant at low temperatures, whereas the proteolytic activity is turned on at elevated temperatures.
Probab=100.00 E-value=1.4e-48 Score=417.12 Aligned_cols=330 Identities=25% Similarity=0.340 Sum_probs=278.5
Q ss_pred cceEEEEEEEe-CCEEEecccccCCCCeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecccccccCceeeecCCCC--cC
Q 009784 147 YSSSSSGFAIG-GRRVLTNAHSVEHYTQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGELP--AL 223 (526)
Q Consensus 147 ~~~~GSGfvI~-~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~~--~~ 223 (526)
..+.||||+|+ +||||||+||+.++..+.|++. +++.++|++++.|+.+||||||++... .+++++|+++. +.
T Consensus 56 ~~~~GSGfii~~~G~IlTn~Hvv~~~~~i~V~~~-~~~~~~a~vv~~d~~~DlAllkv~~~~---~~~~~~l~~~~~~~~ 131 (428)
T TIGR02037 56 VRGLGSGVIISADGYILTNNHVVDGADEITVTLS-DGREFKAKLVGKDPRTDIAVLKIDAKK---NLPVIKLGDSDKLRV 131 (428)
T ss_pred ccceeeEEEECCCCEEEEcHHHcCCCCeEEEEeC-CCCEEEEEEEEecCCCCEEEEEecCCC---CceEEEccCCCCCCC
Confidence 45789999999 7899999999999999999998 899999999999999999999998753 68999998654 67
Q ss_pred CCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecCCCeEEEEEeeccccC-ccccc
Q 009784 224 QDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-DVENI 302 (526)
Q Consensus 224 g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~-~~~~~ 302 (526)
|++|+++|||++... +++.|+|+...+... ....+..++++|+++++|||||||+|.+|+||||+++.+... +..++
T Consensus 132 G~~v~aiG~p~g~~~-~~t~G~vs~~~~~~~-~~~~~~~~i~tda~i~~GnSGGpl~n~~G~viGI~~~~~~~~g~~~g~ 209 (428)
T TIGR02037 132 GDWVLAIGNPFGLGQ-TVTSGIVSALGRSGL-GIGDYENFIQTDAAINPGNSGGPLVNLRGEVIGINTAIYSPSGGNVGI 209 (428)
T ss_pred CCEEEEEECCCcCCC-cEEEEEEEecccCcc-CCCCccceEEECCCCCCCCCCCceECCCCeEEEEEeEEEcCCCCccce
Confidence 999999999999775 899999998876432 122234579999999999999999999999999998876543 34678
Q ss_pred cccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECC
Q 009784 303 GYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDG 381 (526)
Q Consensus 303 ~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG 381 (526)
+|+||++.+++++++|+++|++. ++|||+.++.+ +++.++.+|++. ..|++|.+|.++|||++ |||+||+|++|||
T Consensus 210 ~faiP~~~~~~~~~~l~~~g~~~-~~~lGi~~~~~-~~~~~~~lgl~~-~~Gv~V~~V~~~spA~~aGL~~GDvI~~Vng 286 (428)
T TIGR02037 210 GFAIPSNMAKNVVDQLIEGGKVQ-RGWLGVTIQEV-TSDLAKSLGLEK-QRGALVAQVLPGSPAEKAGLKAGDVILSVNG 286 (428)
T ss_pred EEEEEhHHHHHHHHHHHhcCcCc-CCcCceEeecC-CHHHHHHcCCCC-CCceEEEEccCCCChHHcCCCCCCEEEEECC
Confidence 99999999999999999999998 99999999999 889999999974 57999999999999999 9999999999999
Q ss_pred EEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEecccccccCCCCCCCCCceEEEeeEEEEeccc
Q 009784 382 IDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLY 461 (526)
Q Consensus 382 ~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt~ 461 (526)
++|.++.++. ..+....+|++++++|+|+|+.+++++++...+...+ .+...+.|+.+++++.
T Consensus 287 ~~i~~~~~~~----------~~l~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~-------~~~~~~lGi~~~~l~~ 349 (428)
T TIGR02037 287 KPISSFADLR----------RAIGTLKPGKKVTLGILRKGKEKTITVTLGASPEEQA-------SSSNPFLGLTVANLSP 349 (428)
T ss_pred EEcCCHHHHH----------HHHHhcCCCCEEEEEEEECCEEEEEEEEECcCCCccc-------cccccccceEEecCCH
Confidence 9999988864 6676777899999999999999999999876543211 1233467889988762
Q ss_pred cc----------eeeeeeeecc--hhhhccccccceeeeecccchhhhHHHHH
Q 009784 462 LI----------SVLSMERIMN--MKLRSSFWTSSCIQCHNCQMSSLLWCLRC 502 (526)
Q Consensus 462 ~~----------~~~~~~~i~~--~~~~sg~~~~~~~~~~~~~~~~~~~~~~~ 502 (526)
.. ..+.+..+.+ .+.++||+.||+|...|.+-..+...++-
T Consensus 350 ~~~~~~~l~~~~~Gv~V~~V~~~SpA~~aGL~~GDvI~~Ing~~V~s~~d~~~ 402 (428)
T TIGR02037 350 EIRKELRLKGDVKGVVVTKVVSGSPAARAGLQPGDVILSVNQQPVSSVAELRK 402 (428)
T ss_pred HHHHHcCCCcCcCceEEEEeCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHHH
Confidence 11 3455555554 34578999999999999988887766553
No 3
>PRK10942 serine endoprotease; Provisional
Probab=100.00 E-value=2.2e-48 Score=417.96 Aligned_cols=329 Identities=24% Similarity=0.302 Sum_probs=270.9
Q ss_pred ceEEEEEEEe--CCEEEecccccCCCCeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecccccccCceeeecCCCC--cC
Q 009784 148 SSSSSGFAIG--GRRVLTNAHSVEHYTQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGELP--AL 223 (526)
Q Consensus 148 ~~~GSGfvI~--~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~~--~~ 223 (526)
.+.||||+|+ +||||||+|||.++++++|++. |++.++|++++.|+.+||||||++... .+++++|+++. ++
T Consensus 110 ~~~GSG~ii~~~~G~IlTn~HVv~~a~~i~V~~~-dg~~~~a~vv~~D~~~DlAvlki~~~~---~l~~~~lg~s~~l~~ 185 (473)
T PRK10942 110 MALGSGVIIDADKGYVVTNNHVVDNATKIKVQLS-DGRKFDAKVVGKDPRSDIALIQLQNPK---NLTAIKMADSDALRV 185 (473)
T ss_pred cceEEEEEEECCCCEEEeChhhcCCCCEEEEEEC-CCCEEEEEEEEecCCCCEEEEEecCCC---CCceeEecCccccCC
Confidence 4689999998 4899999999999999999998 999999999999999999999997543 67899998765 67
Q ss_pred CCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecCCCeEEEEEeeccccC-ccccc
Q 009784 224 QDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-DVENI 302 (526)
Q Consensus 224 g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~-~~~~~ 302 (526)
|++|+++|+|++... +++.|+|++..+..... ..+..+||+|+++++|||||||+|.+|+||||+++.+... +..++
T Consensus 186 G~~V~aiG~P~g~~~-tvt~GiVs~~~r~~~~~-~~~~~~iqtda~i~~GnSGGpL~n~~GeviGI~t~~~~~~g~~~g~ 263 (473)
T PRK10942 186 GDYTVAIGNPYGLGE-TVTSGIVSALGRSGLNV-ENYENFIQTDAAINRGNSGGALVNLNGELIGINTAILAPDGGNIGI 263 (473)
T ss_pred CCEEEEEcCCCCCCc-ceeEEEEEEeecccCCc-ccccceEEeccccCCCCCcCccCCCCCeEEEEEEEEEcCCCCcccE
Confidence 999999999998766 89999999887642211 1234579999999999999999999999999999877544 34679
Q ss_pred cccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECC
Q 009784 303 GYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDG 381 (526)
Q Consensus 303 ~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG 381 (526)
+|+||++.+++++++|+++|++. |+|||+.++.+ ++++++.++++ ...|++|.+|.++|||++ |||+||+|++|||
T Consensus 264 gfaIP~~~~~~v~~~l~~~g~v~-rg~lGv~~~~l-~~~~a~~~~l~-~~~GvlV~~V~~~SpA~~AGL~~GDvIl~InG 340 (473)
T PRK10942 264 GFAIPSNMVKNLTSQMVEYGQVK-RGELGIMGTEL-NSELAKAMKVD-AQRGAFVSQVLPNSSAAKAGIKAGDVITSLNG 340 (473)
T ss_pred EEEEEHHHHHHHHHHHHhccccc-cceeeeEeeec-CHHHHHhcCCC-CCCceEEEEECCCChHHHcCCCCCCEEEEECC
Confidence 99999999999999999999998 99999999999 78899999997 467999999999999999 9999999999999
Q ss_pred EEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEecccccccCCCCCCCCCceEEEeeEEEEeccc
Q 009784 382 IDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLY 461 (526)
Q Consensus 382 ~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt~ 461 (526)
++|.++.++. ..+....+|+++.++|+|+|+.+++++++...+..... .... +.|+....++-
T Consensus 341 ~~V~s~~dl~----------~~l~~~~~g~~v~l~v~R~G~~~~v~v~l~~~~~~~~~----~~~~---~lGl~g~~l~~ 403 (473)
T PRK10942 341 KPISSFAALR----------AQVGTMPVGSKLTLGLLRDGKPVNVNVELQQSSQNQVD----SSNI---FNGIEGAELSN 403 (473)
T ss_pred EECCCHHHHH----------HHHHhcCCCCEEEEEEEECCeEEEEEEEeCcCcccccc----cccc---cccceeeeccc
Confidence 9999999875 66777778999999999999999999988664221110 1111 22333333321
Q ss_pred --cceeeeeeeecc--hhhhccccccceeeeecccchhhhHHHHH
Q 009784 462 --LISVLSMERIMN--MKLRSSFWTSSCIQCHNCQMSSLLWCLRC 502 (526)
Q Consensus 462 --~~~~~~~~~i~~--~~~~sg~~~~~~~~~~~~~~~~~~~~~~~ 502 (526)
....+.+..+.+ .+.++||+.||+|...|.+-..+..+|+-
T Consensus 404 ~~~~~gvvV~~V~~~S~A~~aGL~~GDvIv~VNg~~V~s~~dl~~ 448 (473)
T PRK10942 404 KGGDKGVVVDNVKPGTPAAQIGLKKGDVIIGANQQPVKNIAELRK 448 (473)
T ss_pred ccCCCCeEEEEeCCCChHHHcCCCCCCEEEEECCEEcCCHHHHHH
Confidence 112344545543 44579999999999999999988887765
No 4
>TIGR02038 protease_degS periplasmic serine pepetdase DegS. This family consists of the periplasmic serine protease DegS (HhoB), a shorter paralog of protease DO (HtrA, DegP) and DegQ (HhoA). It is found in E. coli and several other Proteobacteria of the gamma subdivision. It contains a trypsin domain and a single copy of PDZ domain (in contrast to DegP with two copies). A critical role of this DegS is to sense stress in the periplasm and partially degrade an inhibitor of sigma(E).
Probab=100.00 E-value=1.6e-47 Score=398.05 Aligned_cols=296 Identities=25% Similarity=0.390 Sum_probs=249.7
Q ss_pred ChhhhhhhcccCCCeEEEEeeeeCCCCCCccccCCCcceEEEEEEEe-CCEEEecccccCCCCeEEEEEcCCCcEEEEEE
Q 009784 111 VEPGVARVVPAMDAVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIG-GRRVLTNAHSVEHYTQVKLKKRGSDTKYLATV 189 (526)
Q Consensus 111 ~~~~v~~~~~~~~SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~-~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~v 189 (526)
+.+.++++. +|||.|.+.....+. + ......+.||||+|+ +||||||+|||.++..+.|++. ||+.++|++
T Consensus 47 ~~~~~~~~~---psVV~I~~~~~~~~~---~-~~~~~~~~GSG~vi~~~G~IlTn~HVV~~~~~i~V~~~-dg~~~~a~v 118 (351)
T TIGR02038 47 FNKAVRRAA---PAVVNIYNRSISQNS---L-NQLSIQGLGSGVIMSKEGYILTNYHVIKKADQIVVALQ-DGRKFEAEL 118 (351)
T ss_pred HHHHHHhcC---CcEEEEEeEeccccc---c-ccccccceEEEEEEeCCeEEEecccEeCCCCEEEEEEC-CCCEEEEEE
Confidence 344455555 599999986543321 1 112345689999999 7899999999999999999997 899999999
Q ss_pred EEeccCCCeEEEEecccccccCceeeecCCC--CcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEc
Q 009784 190 LAIGTECDIAMLTVEDDEFWEGVLPVEFGEL--PALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQID 267 (526)
Q Consensus 190 v~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~--~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~d 267 (526)
++.|+.+||||||++.. .+++++++++ .+.|++|+++|||.+... +++.|+|+...+..... .....+||+|
T Consensus 119 v~~d~~~DlAvlkv~~~----~~~~~~l~~s~~~~~G~~V~aiG~P~~~~~-s~t~GiIs~~~r~~~~~-~~~~~~iqtd 192 (351)
T TIGR02038 119 VGSDPLTDLAVLKIEGD----NLPTIPVNLDRPPHVGDVVLAIGNPYNLGQ-TITQGIISATGRNGLSS-VGRQNFIQTD 192 (351)
T ss_pred EEecCCCCEEEEEecCC----CCceEeccCcCccCCCCEEEEEeCCCCCCC-cEEEEEEEeccCcccCC-CCcceEEEEC
Confidence 99999999999999976 3677778754 478999999999998775 89999999987643321 2234689999
Q ss_pred ccCCCCCCCCeeecCCCeEEEEEeeccccC---ccccccccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHHh
Q 009784 268 AAINSGNSGGPAFNDKGKCVGIAFQSLKHE---DVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRV 344 (526)
Q Consensus 268 a~i~~G~SGGPlvn~~G~VVGI~~~~~~~~---~~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~ 344 (526)
+.+++|||||||+|.+|+||||+++.+... ...+++|+||++.+++++++|+++|++. ++|||+.++++ ++..++
T Consensus 193 a~i~~GnSGGpl~n~~G~vIGI~~~~~~~~~~~~~~g~~faIP~~~~~~vl~~l~~~g~~~-r~~lGv~~~~~-~~~~~~ 270 (351)
T TIGR02038 193 AAINAGNSGGALINTNGELVGINTASFQKGGDEGGEGINFAIPIKLAHKIMGKIIRDGRVI-RGYIGVSGEDI-NSVVAQ 270 (351)
T ss_pred CccCCCCCcceEECCCCeEEEEEeeeecccCCCCccceEEEecHHHHHHHHHHHhhcCccc-ceEeeeEEEEC-CHHHHH
Confidence 999999999999999999999998765432 2367899999999999999999999998 99999999998 788888
Q ss_pred hhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEE
Q 009784 345 AMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKI 423 (526)
Q Consensus 345 ~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~ 423 (526)
.+|++ ...|++|.+|.++|||++ ||++||+|++|||++|.++.++. ..+...++|+++.++|+|+|+.
T Consensus 271 ~lgl~-~~~Gv~V~~V~~~spA~~aGL~~GDvI~~Ing~~V~s~~dl~----------~~l~~~~~g~~v~l~v~R~g~~ 339 (351)
T TIGR02038 271 GLGLP-DLRGIVITGVDPNGPAARAGILVRDVILKYDGKDVIGAEELM----------DRIAETRPGSKVMVTVLRQGKQ 339 (351)
T ss_pred hcCCC-ccccceEeecCCCChHHHCCCCCCCEEEEECCEEcCCHHHHH----------HHHHhcCCCCEEEEEEEECCEE
Confidence 99997 357999999999999999 99999999999999999998864 6666667899999999999999
Q ss_pred EEEEEEeccc
Q 009784 424 LNFNITLATH 433 (526)
Q Consensus 424 ~~~~v~l~~~ 433 (526)
+++++++.+.
T Consensus 340 ~~~~v~l~~~ 349 (351)
T TIGR02038 340 LELPVTIDEK 349 (351)
T ss_pred EEEEEEecCC
Confidence 9999988654
No 5
>PRK10898 serine endoprotease; Provisional
Probab=100.00 E-value=7.6e-47 Score=392.79 Aligned_cols=297 Identities=23% Similarity=0.360 Sum_probs=247.9
Q ss_pred ChhhhhhhcccCCCeEEEEeeeeCCCCCCccccCCCcceEEEEEEEe-CCEEEecccccCCCCeEEEEEcCCCcEEEEEE
Q 009784 111 VEPGVARVVPAMDAVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIG-GRRVLTNAHSVEHYTQVKLKKRGSDTKYLATV 189 (526)
Q Consensus 111 ~~~~v~~~~~~~~SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~-~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~v 189 (526)
..+.++++.+ |||.|.+....... .......+.||||+|+ +||||||+|||.++..+.|++. ||+.++|++
T Consensus 47 ~~~~~~~~~p---svV~v~~~~~~~~~----~~~~~~~~~GSGfvi~~~G~IlTn~HVv~~a~~i~V~~~-dg~~~~a~v 118 (353)
T PRK10898 47 YNQAVRRAAP---AVVNVYNRSLNSTS----HNQLEIRTLGSGVIMDQRGYILTNKHVINDADQIIVALQ-DGRVFEALL 118 (353)
T ss_pred HHHHHHHhCC---cEEEEEeEeccccC----cccccccceeeEEEEeCCeEEEecccEeCCCCEEEEEeC-CCCEEEEEE
Confidence 3445555555 99999986643221 1222344789999999 7899999999999999999997 899999999
Q ss_pred EEeccCCCeEEEEecccccccCceeeecCCC--CcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEc
Q 009784 190 LAIGTECDIAMLTVEDDEFWEGVLPVEFGEL--PALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQID 267 (526)
Q Consensus 190 v~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~--~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~d 267 (526)
++.|+.+||||||++.. .+++++++++ .+.|++|+++|||.+... +++.|+|+...+..... .....+||+|
T Consensus 119 v~~d~~~DlAvl~v~~~----~l~~~~l~~~~~~~~G~~V~aiG~P~g~~~-~~t~Giis~~~r~~~~~-~~~~~~iqtd 192 (353)
T PRK10898 119 VGSDSLTDLAVLKINAT----NLPVIPINPKRVPHIGDVVLAIGNPYNLGQ-TITQGIISATGRIGLSP-TGRQNFLQTD 192 (353)
T ss_pred EEEcCCCCEEEEEEcCC----CCCeeeccCcCcCCCCCEEEEEeCCCCcCC-CcceeEEEeccccccCC-ccccceEEec
Confidence 99999999999999875 4677788764 468999999999998765 79999999877643221 1223579999
Q ss_pred ccCCCCCCCCeeecCCCeEEEEEeeccccCc----cccccccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHH
Q 009784 268 AAINSGNSGGPAFNDKGKCVGIAFQSLKHED----VENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLR 343 (526)
Q Consensus 268 a~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~----~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~ 343 (526)
+++++|||||||+|.+|+||||+++.+...+ ..+++|+||++.+++++++|+++|++. ++|||+..+.+ ++..+
T Consensus 193 a~i~~GnSGGPl~n~~G~vvGI~~~~~~~~~~~~~~~g~~faIP~~~~~~~~~~l~~~G~~~-~~~lGi~~~~~-~~~~~ 270 (353)
T PRK10898 193 ASINHGNSGGALVNSLGELMGINTLSFDKSNDGETPEGIGFAIPTQLATKIMDKLIRDGRVI-RGYIGIGGREI-APLHA 270 (353)
T ss_pred cccCCCCCcceEECCCCeEEEEEEEEecccCCCCcccceEEEEchHHHHHHHHHHhhcCccc-ccccceEEEEC-CHHHH
Confidence 9999999999999999999999998764322 257899999999999999999999998 99999999988 56666
Q ss_pred hhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCE
Q 009784 344 VAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK 422 (526)
Q Consensus 344 ~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~ 422 (526)
..++++ ...|++|.+|.++|||++ ||++||+|++|||++|.++.++. ..+....+|++++++|+|+|+
T Consensus 271 ~~~~~~-~~~Gv~V~~V~~~spA~~aGL~~GDvI~~Ing~~V~s~~~l~----------~~l~~~~~g~~v~l~v~R~g~ 339 (353)
T PRK10898 271 QGGGID-QLQGIVVNEVSPDGPAAKAGIQVNDLIISVNNKPAISALETM----------DQVAEIRPGSVIPVVVMRDDK 339 (353)
T ss_pred HhcCCC-CCCeEEEEEECCCChHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHhcCCCCEEEEEEEECCE
Confidence 677775 347999999999999999 99999999999999999988864 666666789999999999999
Q ss_pred EEEEEEEecccc
Q 009784 423 ILNFNITLATHR 434 (526)
Q Consensus 423 ~~~~~v~l~~~~ 434 (526)
.+++++++.+.+
T Consensus 340 ~~~~~v~l~~~p 351 (353)
T PRK10898 340 QLTLQVTIQEYP 351 (353)
T ss_pred EEEEEEEeccCC
Confidence 999999887653
No 6
>COG0265 DegQ Trypsin-like serine proteases, typically periplasmic, contain C-terminal PDZ domain [Posttranslational modification, protein turnover, chaperones]
Probab=100.00 E-value=8.5e-37 Score=317.92 Aligned_cols=300 Identities=26% Similarity=0.389 Sum_probs=249.6
Q ss_pred ChhhhhhhcccCCCeEEEEeeeeCCCCC-CccccCCC-cceEEEEEEEe-CCEEEecccccCCCCeEEEEEcCCCcEEEE
Q 009784 111 VEPGVARVVPAMDAVVKVFCVHTEPNFS-LPWQRKRQ-YSSSSSGFAIG-GRRVLTNAHSVEHYTQVKLKKRGSDTKYLA 187 (526)
Q Consensus 111 ~~~~v~~~~~~~~SVV~I~~~~~~~~~~-~P~~~~~~-~~~~GSGfvI~-~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a 187 (526)
+...++++.+ +||.|.......... ++-..... ..+.||||+++ +|||+||.||+.++.++.+.+. +|+.+++
T Consensus 35 ~~~~~~~~~~---~vV~~~~~~~~~~~~~~~~~~~~~~~~~~gSg~i~~~~g~ivTn~hVi~~a~~i~v~l~-dg~~~~a 110 (347)
T COG0265 35 FATAVEKVAP---AVVSIATGLTAKLRSFFPSDPPLRSAEGLGSGFIISSDGYIVTNNHVIAGAEEITVTLA-DGREVPA 110 (347)
T ss_pred HHHHHHhcCC---cEEEEEeeeeecchhcccCCcccccccccccEEEEcCCeEEEecceecCCcceEEEEeC-CCCEEEE
Confidence 3445555555 999999876544200 00000000 14789999999 9999999999999999999996 9999999
Q ss_pred EEEEeccCCCeEEEEecccccccCceeeecCCCC--cCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEE
Q 009784 188 TVLAIGTECDIAMLTVEDDEFWEGVLPVEFGELP--ALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQ 265 (526)
Q Consensus 188 ~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~~--~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~ 265 (526)
++++.|+..|+|+||++... .++.+.++++. .+|++++++|+|++... +++.|+|+...+...........+||
T Consensus 111 ~~vg~d~~~dlavlki~~~~---~~~~~~~~~s~~l~vg~~v~aiGnp~g~~~-tvt~Givs~~~r~~v~~~~~~~~~Iq 186 (347)
T COG0265 111 KLVGKDPISDLAVLKIDGAG---GLPVIALGDSDKLRVGDVVVAIGNPFGLGQ-TVTSGIVSALGRTGVGSAGGYVNFIQ 186 (347)
T ss_pred EEEecCCccCEEEEEeccCC---CCceeeccCCCCcccCCEEEEecCCCCccc-ceeccEEeccccccccCcccccchhh
Confidence 99999999999999999875 26777888765 46899999999999665 89999999998762222122556899
Q ss_pred EcccCCCCCCCCeeecCCCeEEEEEeeccccCc-cccccccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHHh
Q 009784 266 IDAAINSGNSGGPAFNDKGKCVGIAFQSLKHED-VENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRV 344 (526)
Q Consensus 266 ~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~-~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~ 344 (526)
+|+++++||||||++|.+|++|||+++.....+ ..+++|+||++.++.+++++.+.|++. ++++|+.+.++ +.+.+
T Consensus 187 tdAain~gnsGgpl~n~~g~~iGint~~~~~~~~~~gigfaiP~~~~~~v~~~l~~~G~v~-~~~lgv~~~~~-~~~~~- 263 (347)
T COG0265 187 TDAAINPGNSGGPLVNIDGEVVGINTAIIAPSGGSSGIGFAIPVNLVAPVLDELISKGKVV-RGYLGVIGEPL-TADIA- 263 (347)
T ss_pred cccccCCCCCCCceEcCCCcEEEEEEEEecCCCCcceeEEEecHHHHHHHHHHHHHcCCcc-ccccceEEEEc-ccccc-
Confidence 999999999999999999999999999886543 456899999999999999999988887 99999999988 55544
Q ss_pred hhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEE
Q 009784 345 AMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKI 423 (526)
Q Consensus 345 ~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~ 423 (526)
+|++ ...|++|.+|.+++||++ |++.||+|+++||+++.+..++. ..+....+|+++.++++|+|+.
T Consensus 264 -~g~~-~~~G~~V~~v~~~spa~~agi~~Gdii~~vng~~v~~~~~l~----------~~v~~~~~g~~v~~~~~r~g~~ 331 (347)
T COG0265 264 -LGLP-VAAGAVVLGVLPGSPAAKAGIKAGDIITAVNGKPVASLSDLV----------AAVASNRPGDEVALKLLRGGKE 331 (347)
T ss_pred -cCCC-CCCceEEEecCCCChHHHcCCCCCCEEEEECCEEccCHHHHH----------HHHhccCCCCEEEEEEEECCEE
Confidence 7776 678899999999999999 99999999999999999998865 6677777999999999999999
Q ss_pred EEEEEEeccc
Q 009784 424 LNFNITLATH 433 (526)
Q Consensus 424 ~~~~v~l~~~ 433 (526)
+++.+++.+.
T Consensus 332 ~~~~v~l~~~ 341 (347)
T COG0265 332 RELAVTLGDR 341 (347)
T ss_pred EEEEEEecCc
Confidence 9999998773
No 7
>KOG1421 consensus Predicted signaling-associated protein (contains a PDZ domain) [General function prediction only]
Probab=99.98 E-value=1.7e-31 Score=281.02 Aligned_cols=322 Identities=17% Similarity=0.264 Sum_probs=266.3
Q ss_pred hhhhhhhcccCCCeEEEEeeeeCCCCCCccccCCCcceEEEEEEEe--CCEEEecccccCCCCe-EEEEEcCCCcEEEEE
Q 009784 112 EPGVARVVPAMDAVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIG--GRRVLTNAHSVEHYTQ-VKLKKRGSDTKYLAT 188 (526)
Q Consensus 112 ~~~v~~~~~~~~SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~--~g~ILT~aHvV~~~~~-i~V~~~~~g~~~~a~ 188 (526)
..|...++.+.+|||.|.+..... |.......+.++||+++ .||||||+|++..... -.+.+. +..+.+.-
T Consensus 52 e~w~~~ia~VvksvVsI~~S~v~~-----fdtesag~~~atgfvvd~~~gyiLtnrhvv~pgP~va~avf~-n~ee~ei~ 125 (955)
T KOG1421|consen 52 EDWRNTIANVVKSVVSIRFSAVRA-----FDTESAGESEATGFVVDKKLGYILTNRHVVAPGPFVASAVFD-NHEEIEIY 125 (955)
T ss_pred hhhhhhhhhhcccEEEEEehheee-----cccccccccceeEEEEecccceEEEeccccCCCCceeEEEec-ccccCCcc
Confidence 356666677777999999877543 34455667889999999 6999999999985544 455554 77788888
Q ss_pred EEEeccCCCeEEEEeccccc-ccCceeeecCC-CCcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCc-----eee
Q 009784 189 VLAIGTECDIAMLTVEDDEF-WEGVLPVEFGE-LPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGS-----TEL 261 (526)
Q Consensus 189 vv~~d~~~DlAlLkv~~~~~-~~~~~pl~l~~-~~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~-----~~~ 261 (526)
.++.|+.||+.+++.+++.+ +..+..+++.. ..++|.++.++|+..+. ..++..|.++++++.....+. .+.
T Consensus 126 pvyrDpVhdfGf~r~dps~ir~s~vt~i~lap~~akvgseirvvgNDagE-klsIlagflSrldr~apdyg~~~yndfnT 204 (955)
T KOG1421|consen 126 PVYRDPVHDFGFFRYDPSTIRFSIVTEICLAPELAKVGSEIRVVGNDAGE-KLSILAGFLSRLDRNAPDYGEDTYNDFNT 204 (955)
T ss_pred cccCCchhhcceeecChhhcceeeeeccccCccccccCCceEEecCCccc-eEEeehhhhhhccCCCccccccccccccc
Confidence 89999999999999998754 33456666663 45789999999998774 458999999999876544433 222
Q ss_pred eEEEEcccCCCCCCCCeeecCCCeEEEEEeeccccCccccccccccHHHHHHHHHHHHHcCceeeccccCceeeeccChh
Q 009784 262 LGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPD 341 (526)
Q Consensus 262 ~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~ 341 (526)
.++|..+....|.||+|++|.+|..|.++.++. .....+|++|++.+.+.|..++.+..++ |+.|.++|... ..+
T Consensus 205 fy~QaasstsggssgspVv~i~gyAVAl~agg~---~ssas~ffLpLdrV~RaL~clq~n~PIt-RGtLqvefl~k-~~d 279 (955)
T KOG1421|consen 205 FYIQAASSTSGGSSGSPVVDIPGYAVALNAGGS---ISSASDFFLPLDRVVRALRCLQNNTPIT-RGTLQVEFLHK-LFD 279 (955)
T ss_pred eeeeehhcCCCCCCCCceecccceEEeeecCCc---ccccccceeeccchhhhhhhhhcCCCcc-cceEEEEEehh-hhH
Confidence 368888899999999999999999999998765 4567799999999999999999877777 99999999988 889
Q ss_pred HHhhhcCCC-----------Cccce-EEEEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCC
Q 009784 342 LRVAMSMKA-----------DQKGV-RIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYT 409 (526)
Q Consensus 342 ~~~~lgl~~-----------~~~Gv-~V~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~ 409 (526)
.++.+||+. ...|+ +|..|.++|||++.|++||++++||+.-+.++..+. ..++ ...
T Consensus 280 e~rrlGL~sE~eqv~r~k~P~~tgmLvV~~vL~~gpa~k~Le~GDillavN~t~l~df~~l~----------~iLD-egv 348 (955)
T KOG1421|consen 280 ECRRLGLSSEWEQVVRTKFPERTGMLVVETVLPEGPAEKKLEPGDILLAVNSTCLNDFEALE----------QILD-EGV 348 (955)
T ss_pred HHHhcCCcHHHHHHHHhcCcccceeEEEEEeccCCchhhccCCCcEEEEEcceehHHHHHHH----------HHHh-hcc
Confidence 999999975 23454 567889999999999999999999999999988763 4444 458
Q ss_pred CCEEEEEEEECCEEEEEEEEecccccccCCCCCCCCCceEEEeeEEEEeccccc
Q 009784 410 GDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLYLI 463 (526)
Q Consensus 410 G~~v~l~v~R~G~~~~~~v~l~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt~~~ 463 (526)
|+.++|+|+|+|+++++++++...+...|. ||+.++|++||+++|+.
T Consensus 349 gk~l~LtI~Rggqelel~vtvqdlh~itp~-------R~levcGav~hdlsyq~ 395 (955)
T KOG1421|consen 349 GKNLELTIQRGGQELELTVTVQDLHGITPD-------RFLEVCGAVFHDLSYQL 395 (955)
T ss_pred CceEEEEEEeCCEEEEEEEEeccccCCCCc-------eEEEEcceEecCCCHHH
Confidence 999999999999999999999999988887 99999999999999763
No 8
>KOG1320 consensus Serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=99.96 E-value=1.2e-29 Score=265.52 Aligned_cols=372 Identities=40% Similarity=0.557 Sum_probs=328.2
Q ss_pred hcccCCCeEEEEeeeeCCCCCCccccCCCcceEEEEEEEeCCEEEecccccC---CCCeEEEEEcCCCcEEEEEEEEecc
Q 009784 118 VVPAMDAVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIGGRRVLTNAHSVE---HYTQVKLKKRGSDTKYLATVLAIGT 194 (526)
Q Consensus 118 ~~~~~~SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~~g~ILT~aHvV~---~~~~i~V~~~~~g~~~~a~vv~~d~ 194 (526)
......|++.+.+....+.+..||+...+....|+||.+....++||+|++. +...+.+...+.-+.|.+++...-.
T Consensus 56 ~~~~~~s~~~v~~~~~~~~~~~pw~~~~q~~~~~s~f~i~~~~lltn~~~v~~~~~~~~v~v~~~gs~~k~~~~v~~~~~ 135 (473)
T KOG1320|consen 56 VDLALQSVVKVFSVSTEPSSVLPWQRTRQFSSGGSGFAIYGKKLLTNAHVVAPNNDHKFVTVKKHGSPRKYKAFVAAVFE 135 (473)
T ss_pred ccccccceeEEEeecccccccCcceeeehhcccccchhhcccceeecCccccccccccccccccCCCchhhhhhHHHhhh
Confidence 3445569999999999999999999999888999999999999999999999 6667777766566788999998889
Q ss_pred CCCeEEEEecccccccCceeeecCCCCcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCC
Q 009784 195 ECDIAMLTVEDDEFWEGVLPVEFGELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGN 274 (526)
Q Consensus 195 ~~DlAlLkv~~~~~~~~~~pl~l~~~~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~ 274 (526)
+.|+|+|.++..+|+....|+++++.+.+.+.++++| ++..+++.|.|++.....+..+......+++|+++++|+
T Consensus 136 ~cd~Avv~Ie~~~f~~~~~~~e~~~ip~l~~S~~Vv~----gd~i~VTnghV~~~~~~~y~~~~~~l~~vqi~aa~~~~~ 211 (473)
T KOG1320|consen 136 ECDLAVVYIESEEFWKGMNPFELGDIPSLNGSGFVVG----GDGIIVTNGHVVRVEPRIYAHSSTVLLRVQIDAAIGPGN 211 (473)
T ss_pred cccceEEEEeeccccCCCcccccCCCcccCccEEEEc----CCcEEEEeeEEEEEEeccccCCCcceeeEEEEEeecCCc
Confidence 9999999999999988888999999999999999999 345699999999999888888877788899999999999
Q ss_pred CCCeeecCCCeEEEEEeeccccCccccccccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHHhhhcCCCCccc
Q 009784 275 SGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKG 354 (526)
Q Consensus 275 SGGPlvn~~G~VVGI~~~~~~~~~~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~~lgl~~~~~G 354 (526)
||+|.+...+++.|+.+...+..+ ++++.||.-.+.+++....+.+.+.++++++...+.+.+.+.++.+.|..+ .|
T Consensus 212 s~ep~i~g~d~~~gvA~l~ik~~~--~i~~~i~~~~~~~~~~G~~~~a~~~~f~~~nt~t~g~vs~~~R~~~~lg~~-~g 288 (473)
T KOG1320|consen 212 SGEPVIVGVDKVAGVAFLKIKTPE--NILYVIPLGVSSHFRTGVEVSAIGNGFGLLNTLTQGMVSGQLRKSFKLGLE-TG 288 (473)
T ss_pred cCCCeEEccccccceEEEEEecCC--cccceeecceeeeecccceeeccccCceeeeeeeecccccccccccccCcc-cc
Confidence 999999888899999998874322 789999999999999998888988899999999999999999999999877 99
Q ss_pred eEEEEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEecccc
Q 009784 355 VRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHR 434 (526)
Q Consensus 355 v~V~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~~~ 434 (526)
+.+.++.+.+.|.+-++.||+|+.+||+.|. +.++..+|+.|++.+..+.++|++.+.+.|.+ ++++.++...
T Consensus 289 ~~i~~~~qtd~ai~~~nsg~~ll~~DG~~Ig----Vn~~~~~ri~~~~~iSf~~p~d~vl~~v~r~~---e~~~~lr~~~ 361 (473)
T KOG1320|consen 289 VLISKINQTDAAINPGNSGGPLLNLDGEVIG----VNTRKVTRIGFSHGISFKIPIDTVLVIVLRLG---EFQISLRPVK 361 (473)
T ss_pred eeeeeecccchhhhcccCCCcEEEecCcEee----eeeeeeEEeeccccceeccCchHhhhhhhhhh---hhceeecccc
Confidence 9999999999999988999999999999998 55667889999999999999999999999998 6777788888
Q ss_pred cccCCCCCCCCCceEEEeeEEEEeccc--cce-----eeeeeeecch--hhhccccccceeeeecccchhhhHHHHHH
Q 009784 435 RLIPSHNKGRPPSYYIIAGFVFSRCLY--LIS-----VLSMERIMNM--KLRSSFWTSSCIQCHNCQMSSLLWCLRCL 503 (526)
Q Consensus 435 ~~~p~~~~~~~p~~~i~gG~~f~~lt~--~~~-----~~~~~~i~~~--~~~sg~~~~~~~~~~~~~~~~~~~~~~~~ 503 (526)
.+.|.+.+...+.|++++|++|++++. ... .+.+..+++. ..+.|+..+|.+...|.+...++-+|+-+
T Consensus 362 ~~~p~~~~~g~~s~~i~~g~vf~~~~~~~~~~~~~~q~v~is~Vlp~~~~~~~~~~~g~~V~~vng~~V~n~~~l~~~ 439 (473)
T KOG1320|consen 362 PLVPVHQYIGLPSYYIFAGLVFVPLTKSYIFPSGVVQLVLVSQVLPGSINGGYGLKPGDQVVKVNGKPVKNLKHLYEL 439 (473)
T ss_pred CcccccccCCceeEEEecceEEeecCCCccccccceeEEEEEEeccCCCcccccccCCCEEEEECCEEeechHHHHHH
Confidence 888889999999999999999999973 222 2445555553 35678889999999999999999988764
No 9
>KOG1320 consensus Serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=99.93 E-value=8.3e-25 Score=229.33 Aligned_cols=301 Identities=21% Similarity=0.204 Sum_probs=225.6
Q ss_pred hcccCCCeEEEEeeeeCCCCCCccccCCCcceEEEEEEEe-CCEEEecccccCCCC-----------eEEEEEcC-CCcE
Q 009784 118 VVPAMDAVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIG-GRRVLTNAHSVEHYT-----------QVKLKKRG-SDTK 184 (526)
Q Consensus 118 ~~~~~~SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~-~g~ILT~aHvV~~~~-----------~i~V~~~~-~g~~ 184 (526)
..+...+||.|....--.. ..|+....-....|||||++ +|+++||+||+.... .+.|.... .+..
T Consensus 134 ~~~cd~Avv~Ie~~~f~~~-~~~~e~~~ip~l~~S~~Vv~gd~i~VTnghV~~~~~~~y~~~~~~l~~vqi~aa~~~~~s 212 (473)
T KOG1320|consen 134 FEECDLAVVYIESEEFWKG-MNPFELGDIPSLNGSGFVVGGDGIIVTNGHVVRVEPRIYAHSSTVLLRVQIDAAIGPGNS 212 (473)
T ss_pred hhcccceEEEEeeccccCC-CcccccCCCcccCccEEEEcCCcEEEEeeEEEEEEeccccCCCcceeeEEEEEeecCCcc
Confidence 3444558999887432111 12455555666789999999 999999999997432 26666652 2488
Q ss_pred EEEEEEEeccCCCeEEEEecccccccCceeeecCC--CCcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCC----c
Q 009784 185 YLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGE--LPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHG----S 258 (526)
Q Consensus 185 ~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~--~~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~----~ 258 (526)
+++.+++.|+..|+|+++++.+. ....+++++- ....|+++.++|.|++..+ +++.|+++...|..+.-+ .
T Consensus 213 ~ep~i~g~d~~~gvA~l~ik~~~--~i~~~i~~~~~~~~~~G~~~~a~~~~f~~~n-t~t~g~vs~~~R~~~~lg~~~g~ 289 (473)
T KOG1320|consen 213 GEPVIVGVDKVAGVAFLKIKTPE--NILYVIPLGVSSHFRTGVEVSAIGNGFGLLN-TLTQGMVSGQLRKSFKLGLETGV 289 (473)
T ss_pred CCCeEEccccccceEEEEEecCC--cccceeecceeeeecccceeeccccCceeee-eeeecccccccccccccCcccce
Confidence 89999999999999999997553 1356666664 3456899999999999887 799999998877655422 3
Q ss_pred eeeeEEEEcccCCCCCCCCeeecCCCeEEEEEeeccccC-ccccccccccHHHHHHHHHHHHHcCc---ee-----eccc
Q 009784 259 TELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-DVENIGYVIPTPVIMHFIQDYEKNGA---YT-----GFPL 329 (526)
Q Consensus 259 ~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~-~~~~~~~aIP~~~i~~~l~~l~~~g~---~~-----~~~~ 329 (526)
....++|+|++++.|+||||++|.+|++||++++..... -..+++|++|.+.++.++.+..+... .. .+.|
T Consensus 290 ~i~~~~qtd~ai~~~nsg~~ll~~DG~~IgVn~~~~~ri~~~~~iSf~~p~d~vl~~v~r~~e~~~~lr~~~~~~p~~~~ 369 (473)
T KOG1320|consen 290 LISKINQTDAAINPGNSGGPLLNLDGEVIGVNTRKVTRIGFSHGISFKIPIDTVLVIVLRLGEFQISLRPVKPLVPVHQY 369 (473)
T ss_pred eeeeecccchhhhcccCCCcEEEecCcEeeeeeeeeEEeeccccceeccCchHhhhhhhhhhhhceeeccccCccccccc
Confidence 445689999999999999999999999999998876322 23678999999999988888743221 11 1346
Q ss_pred cCceeeeccChhH----HhhhcCC-CCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhh
Q 009784 330 LGVEWQKMENPDL----RVAMSMK-ADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYL 403 (526)
Q Consensus 330 LGi~~~~~~~~~~----~~~lgl~-~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~ 403 (526)
+|.....+...-. .+.+-.+ ...++++|.+|.+++++.. ++++||+|++|||++|.+..+|. ++
T Consensus 370 ~g~~s~~i~~g~vf~~~~~~~~~~~~~~q~v~is~Vlp~~~~~~~~~~~g~~V~~vng~~V~n~~~l~----------~~ 439 (473)
T KOG1320|consen 370 IGLPSYYIFAGLVFVPLTKSYIFPSGVVQLVLVSQVLPGSINGGYGLKPGDQVVKVNGKPVKNLKHLY----------EL 439 (473)
T ss_pred CCceeEEEecceEEeecCCCccccccceeEEEEEEeccCCCcccccccCCCEEEEECCEEeechHHHH----------HH
Confidence 6666554421111 1111122 2346899999999999999 99999999999999999999975 78
Q ss_pred hhhcCCCCEEEEEEEECCEEEEEEEEecc
Q 009784 404 VSQKYTGDSAAVKVLRDSKILNFNITLAT 432 (526)
Q Consensus 404 l~~~~~G~~v~l~v~R~G~~~~~~v~l~~ 432 (526)
+....++++|.+..+|+.|..++.+....
T Consensus 440 i~~~~~~~~v~vl~~~~~e~~tl~Il~~~ 468 (473)
T KOG1320|consen 440 IEECSTEDKVAVLDRRSAEDATLEILPEH 468 (473)
T ss_pred HHhcCcCceEEEEEecCccceeEEecccc
Confidence 88888889999999999999999887654
No 10
>KOG1421 consensus Predicted signaling-associated protein (contains a PDZ domain) [General function prediction only]
Probab=99.79 E-value=2e-17 Score=175.49 Aligned_cols=327 Identities=13% Similarity=0.122 Sum_probs=240.0
Q ss_pred CeEEEEeeeeCCCCCCccccCCCcceEEEEEEEe--CCEEEecccccC-CCCeEEEEEcCCCcEEEEEEEEeccCCCeEE
Q 009784 124 AVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIG--GRRVLTNAHSVE-HYTQVKLKKRGSDTKYLATVLAIGTECDIAM 200 (526)
Q Consensus 124 SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~--~g~ILT~aHvV~-~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAl 200 (526)
+.|.+.......-.++ ......|||.|++ +|++++++.+|. ++.+..|+.. +...++|.+.+.++..++|.
T Consensus 530 ~~~~v~~~~~~~l~g~-----s~~i~kgt~~i~d~~~g~~vvsr~~vp~d~~d~~vt~~-dS~~i~a~~~fL~~t~n~a~ 603 (955)
T KOG1421|consen 530 CLVDVEPMMPVNLDGV-----SSDIYKGTALIMDTSKGLGVVSRSVVPSDAKDQRVTEA-DSDGIPANVSFLHPTENVAS 603 (955)
T ss_pred hhhhheeceeeccccc-----hhhhhcCceEEEEccCCceeEecccCCchhhceEEeec-ccccccceeeEecCccceeE
Confidence 6666666554333221 1123579999999 799999999997 6778999987 78889999999999999999
Q ss_pred EEecccccccCceeeecCCCC-cCCCcEEEEeeCCCCCceeEEEEEEece-----eeee-ccCCceeeeEEEEcccCCCC
Q 009784 201 LTVEDDEFWEGVLPVEFGELP-ALQDAVTVVGYPIGGDTISVTSGVVSRI-----EILS-YVHGSTELLGLQIDAAINSG 273 (526)
Q Consensus 201 Lkv~~~~~~~~~~pl~l~~~~-~~g~~V~~iG~p~~~~~~sv~~GiVs~~-----~~~~-~~~~~~~~~~i~~da~i~~G 273 (526)
+|+++.. ...++|.+.. ..|+++...|+....... .....|..+ .... ......+...|.+++.+.-+
T Consensus 604 ~kydp~~----~~~~kl~~~~v~~gD~~~f~g~~~~~r~l-taktsv~dvs~~~~ps~~~pr~r~~n~e~Is~~~nlsT~ 678 (955)
T KOG1421|consen 604 FKYDPAL----EVQLKLTDTTVLRGDECTFEGFTEDLRAL-TAKTSVTDVSVVIIPSSVMPRFRATNLEVISFMDNLSTS 678 (955)
T ss_pred eccChhH----hhhhccceeeEecCCceeEecccccchhh-cccceeeeeEEEEecCCCCcceeecceEEEEEecccccc
Confidence 9999874 3455565433 568999999998765431 111122221 1111 11223556778888887777
Q ss_pred CCCCeeecCCCeEEEEEeeccccC-c--cccccccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHHhhhcCCC
Q 009784 274 NSGGPAFNDKGKCVGIAFQSLKHE-D--VENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKA 350 (526)
Q Consensus 274 ~SGGPlvn~~G~VVGI~~~~~~~~-~--~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~~lgl~~ 350 (526)
+--|-+.|.+|+|+|++...+.+. + ...+-|.+.+..++++|++|+.++... ...+|++|..+ +...++.+|++.
T Consensus 679 c~sg~ltdddg~vvalwl~~~ge~~~~kd~~y~~gl~~~~~l~vl~rlk~g~~~r-p~i~~vef~~i-~laqar~lglp~ 756 (955)
T KOG1421|consen 679 CLSGRLTDDDGEVVALWLSVVGEDVGGKDYTYKYGLSMSYILPVLERLKLGPSAR-PTIAGVEFSHI-TLAQARTLGLPS 756 (955)
T ss_pred ccceEEECCCCeEEEEEeeeeccccCCceeEEEeccchHHHHHHHHHHhcCCCCC-ceeeccceeeE-EeehhhccCCCH
Confidence 777789999999999998776543 1 223456788899999999998777765 66789999999 888899999985
Q ss_pred ------------CccceEEEEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEE
Q 009784 351 ------------DQKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVL 418 (526)
Q Consensus 351 ------------~~~Gv~V~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~ 418 (526)
..+-.+|++|.+..+- -|..||||+++||+-|+...||. + +. .++.+|+
T Consensus 757 e~imk~e~es~~~~ql~~ishv~~~~~k--il~~gdiilsvngk~itr~~dl~----------d-~~------eid~~il 817 (955)
T KOG1421|consen 757 EFIMKSEEESTIPRQLYVISHVRPLLHK--ILGVGDIILSVNGKMITRLSDLH----------D-FE------EIDAVIL 817 (955)
T ss_pred HHHhhhhhcCCCcceEEEEEeeccCccc--ccccccEEEEecCeEEeeehhhh----------h-hh------hhheeee
Confidence 2345678888876543 59999999999999999998874 2 21 6899999
Q ss_pred ECCEEEEEEEEecccccccCCCCCCCCCceEEEeeEEEEecc---------ccceeeeeeeecc-hhhhccccccceeee
Q 009784 419 RDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCL---------YLISVLSMERIMN-MKLRSSFWTSSCIQC 488 (526)
Q Consensus 419 R~G~~~~~~v~l~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt---------~~~~~~~~~~i~~-~~~~sg~~~~~~~~~ 488 (526)
|+|..+++++++.+.. ++. |.++|.|..+|+-- .-++++.+-|-.. .+++ ++.+-..|..
T Consensus 818 rdg~~~~ikipt~p~~--et~-------r~vi~~gailq~ph~av~~q~edlp~gvyvt~rg~gspalq-~l~aa~fita 887 (955)
T KOG1421|consen 818 RDGIEMEIKIPTYPEY--ETS-------RAVIWMGAILQPPHSAVFEQVEDLPEGVYVTSRGYGSPALQ-MLRAAHFITA 887 (955)
T ss_pred ecCcEEEEEecccccc--ccc-------eEEEEEeccccCchHHHHHHHhccCCceEEeecccCChhHh-hcchheeEEE
Confidence 9999999999877654 333 89999999887632 2356666666555 4564 8888888888
Q ss_pred eccc
Q 009784 489 HNCQ 492 (526)
Q Consensus 489 ~~~~ 492 (526)
.|.-
T Consensus 888 vng~ 891 (955)
T KOG1421|consen 888 VNGH 891 (955)
T ss_pred eccc
Confidence 8773
No 11
>PF13365 Trypsin_2: Trypsin-like peptidase domain; PDB: 1Y8T_A 2Z9I_A 3QO6_A 1L1J_A 1QY6_A 2O8L_A 3OTP_E 2ZLE_I 1KY9_A 3CS0_A ....
Probab=99.63 E-value=6e-15 Score=128.71 Aligned_cols=108 Identities=33% Similarity=0.488 Sum_probs=72.7
Q ss_pred EEEEEEe-CCEEEecccccC--------CCCeEEEEEcCCCcEEE--EEEEEeccC-CCeEEEEecccccccCceeeecC
Q 009784 151 SSGFAIG-GRRVLTNAHSVE--------HYTQVKLKKRGSDTKYL--ATVLAIGTE-CDIAMLTVEDDEFWEGVLPVEFG 218 (526)
Q Consensus 151 GSGfvI~-~g~ILT~aHvV~--------~~~~i~V~~~~~g~~~~--a~vv~~d~~-~DlAlLkv~~~~~~~~~~pl~l~ 218 (526)
||||+|+ +|+||||+||+. ....+.+... ++..+. +++++.++. +|+|||+++..
T Consensus 1 GTGf~i~~~g~ilT~~Hvv~~~~~~~~~~~~~~~~~~~-~~~~~~~~~~~~~~~~~~~D~All~v~~~------------ 67 (120)
T PF13365_consen 1 GTGFLIGPDGYILTAAHVVEDWNDGKQPDNSSVEVVFP-DGRRVPPVAEVVYFDPDDYDLALLKVDPW------------ 67 (120)
T ss_dssp EEEEEEETTTEEEEEHHHHTCCTT--G-TCSEEEEEET-TSCEEETEEEEEEEETT-TTEEEEEESCE------------
T ss_pred CEEEEEcCCceEEEchhheecccccccCCCCEEEEEec-CCCEEeeeEEEEEECCccccEEEEEEecc------------
Confidence 8999999 559999999998 4567888887 677777 999999999 99999999910
Q ss_pred CCCcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecCCCeEEEE
Q 009784 219 ELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGI 289 (526)
Q Consensus 219 ~~~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI 289 (526)
...+.. ....+.......... .......+ +++.+.+|+|||||||.+|+||||
T Consensus 68 ---------~~~~~~------~~~~~~~~~~~~~~~--~~~~~~~~-~~~~~~~G~SGgpv~~~~G~vvGi 120 (120)
T PF13365_consen 68 ---------TGVGGG------VRVPGSTSGVSPTST--NDNRMLYI-TDADTRPGSSGGPVFDSDGRVVGI 120 (120)
T ss_dssp ---------EEEEEE------EEEEEEEEEEEEEEE--EETEEEEE-ESSS-STTTTTSEEEETTSEEEEE
T ss_pred ---------cceeee------eEeeeeccccccccC--cccceeEe-eecccCCCcEeHhEECCCCEEEeC
Confidence 000000 000000011000000 00111124 899999999999999999999997
No 12
>PF13180 PDZ_2: PDZ domain; PDB: 2L97_A 1Y8T_A 2Z9I_A 1LCY_A 2PZD_B 2P3W_A 1VCW_C 1TE0_B 1SOZ_C 1SOT_C ....
Probab=99.50 E-value=4e-14 Score=116.57 Aligned_cols=81 Identities=33% Similarity=0.535 Sum_probs=69.7
Q ss_pred cccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhh
Q 009784 328 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ 406 (526)
Q Consensus 328 ~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~ 406 (526)
||||+.+.... ...|++|.+|.++|||++ |||+||+|++|||++|.+..++. ..+..
T Consensus 1 ~~lGv~~~~~~------------~~~g~~V~~V~~~spA~~aGl~~GD~I~~ing~~v~~~~~~~----------~~l~~ 58 (82)
T PF13180_consen 1 GGLGVTVQNLS------------DTGGVVVVSVIPGSPAAKAGLQPGDIILAINGKPVNSSEDLV----------NILSK 58 (82)
T ss_dssp -E-SEEEEECS------------CSSSEEEEEESTTSHHHHTTS-TTEEEEEETTEESSSHHHHH----------HHHHC
T ss_pred CEECeEEEEcc------------CCCeEEEEEeCCCCcHHHCCCCCCcEEEEECCEEcCCHHHHH----------HHHHh
Confidence 58999999872 246999999999999999 99999999999999999988864 77778
Q ss_pred cCCCCEEEEEEEECCEEEEEEEEe
Q 009784 407 KYTGDSAAVKVLRDSKILNFNITL 430 (526)
Q Consensus 407 ~~~G~~v~l~v~R~G~~~~~~v~l 430 (526)
..+|++++|+|+|+|+.++++++|
T Consensus 59 ~~~g~~v~l~v~R~g~~~~~~v~l 82 (82)
T PF13180_consen 59 GKPGDTVTLTVLRDGEELTVEVTL 82 (82)
T ss_dssp SSTTSEEEEEEEETTEEEEEEEE-
T ss_pred CCCCCEEEEEEEECCEEEEEEEEC
Confidence 889999999999999999999875
No 13
>PF00089 Trypsin: Trypsin; InterPro: IPR001254 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine proteases belong to the MEROPS peptidase family S1 (chymotrypsin family, clan PA(S))and to peptidase family S6 (Hap serine peptidases). The chymotrypsin family is almost totally confined to animals, although trypsin-like enzymes are found in actinomycetes of the genera Streptomyces and Saccharopolyspora, and in the fungus Fusarium oxysporum []. The enzymes are inherently secreted, being synthesised with a signal peptide that targets them to the secretory pathway. Animal enzymes are either secreted directly, packaged into vesicles for regulated secretion, or are retained in leukocyte granules []. The Hap family, 'Haemophilus adhesion and penetration', are proteins that play a role in the interaction with human epithelial cells. The serine protease activity is localized at the N-terminal domain, whereas the binding domain is in the C-terminal region. ; GO: 0004252 serine-type endopeptidase activity, 0006508 proteolysis; PDB: 1SPJ_A 1A5I_A 2ZGH_A 2ZKS_A 2ZGJ_A 2ZGC_A 2ODP_A 2I6Q_A 2I6S_A 2ODQ_A ....
Probab=99.44 E-value=1.5e-12 Score=125.06 Aligned_cols=185 Identities=22% Similarity=0.269 Sum_probs=118.2
Q ss_pred eeeCCCCCCccccCCCc---ceEEEEEEEeCCEEEecccccCCCCeEEEEEcC------CC--cEEEEEEEEecc-----
Q 009784 131 VHTEPNFSLPWQRKRQY---SSSSSGFAIGGRRVLTNAHSVEHYTQVKLKKRG------SD--TKYLATVLAIGT----- 194 (526)
Q Consensus 131 ~~~~~~~~~P~~~~~~~---~~~GSGfvI~~g~ILT~aHvV~~~~~i~V~~~~------~g--~~~~a~vv~~d~----- 194 (526)
+......++||...... ...|+|++|++.+|||++||+.+..++.+.+.. ++ ..+..+-+..++
T Consensus 4 g~~~~~~~~p~~v~i~~~~~~~~C~G~li~~~~vLTaahC~~~~~~~~v~~g~~~~~~~~~~~~~~~v~~~~~h~~~~~~ 83 (220)
T PF00089_consen 4 GDPASPGEFPWVVSIRYSNGRFFCTGTLISPRWVLTAAHCVDGASDIKVRLGTYSIRNSDGSEQTIKVSKIIIHPKYDPS 83 (220)
T ss_dssp SEECGTTSSTTEEEEEETTTEEEEEEEEEETTEEEEEGGGHTSGGSEEEEESESBTTSTTTTSEEEEEEEEEEETTSBTT
T ss_pred CEECCCCCCCeEEEEeeCCCCeeEeEEecccccccccccccccccccccccccccccccccccccccccccccccccccc
Confidence 34455566677654322 468999999999999999999996667665431 22 234444443432
Q ss_pred --CCCeEEEEeccc-ccccCceeeecCCCC---cCCCcEEEEeeCCCCCce---eE---EEEEEeceeeeeccCCceeee
Q 009784 195 --ECDIAMLTVEDD-EFWEGVLPVEFGELP---ALQDAVTVVGYPIGGDTI---SV---TSGVVSRIEILSYVHGSTELL 262 (526)
Q Consensus 195 --~~DlAlLkv~~~-~~~~~~~pl~l~~~~---~~g~~V~~iG~p~~~~~~---sv---~~GiVs~~~~~~~~~~~~~~~ 262 (526)
.+|+|||+++.+ .+...+.++.+.... ..++.+.++||+...... .+ ...+++...+...........
T Consensus 84 ~~~~DiAll~L~~~~~~~~~~~~~~l~~~~~~~~~~~~~~~~G~~~~~~~~~~~~~~~~~~~~~~~~~c~~~~~~~~~~~ 163 (220)
T PF00089_consen 84 TYDNDIALLKLDRPITFGDNIQPICLPSAGSDPNVGTSCIVVGWGRTSDNGYSSNLQSVTVPVVSRKTCRSSYNDNLTPN 163 (220)
T ss_dssp TTTTSEEEEEESSSSEHBSSBEESBBTSTTHTTTTTSEEEEEESSBSSTTSBTSBEEEEEEEEEEHHHHHHHTTTTSTTT
T ss_pred cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccc
Confidence 479999999987 345578888887632 578899999998753221 23 333333333322111111123
Q ss_pred EEEEcc----cCCCCCCCCeeecCCCeEEEEEeeccccCccccccccccHHHHHHHH
Q 009784 263 GLQIDA----AINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFI 315 (526)
Q Consensus 263 ~i~~da----~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~~~~~~~aIP~~~i~~~l 315 (526)
.+++.. ..+.|+|||||++.++.++||++.+.........++..++....++|
T Consensus 164 ~~c~~~~~~~~~~~g~sG~pl~~~~~~lvGI~s~~~~c~~~~~~~v~~~v~~~~~WI 220 (220)
T PF00089_consen 164 MICAGSSGSGDACQGDSGGPLICNNNYLVGIVSFGENCGSPNYPGVYTRVSSYLDWI 220 (220)
T ss_dssp EEEEETTSSSBGGTTTTTSEEEETTEEEEEEEEEESSSSBTTSEEEEEEGGGGHHHH
T ss_pred cccccccccccccccccccccccceeeecceeeecCCCCCCCcCEEEEEHHHhhccC
Confidence 466655 78899999999998777999998874322222346777776655543
No 14
>cd00190 Tryp_SPc Trypsin-like serine protease; Many of these are synthesized as inactive precursor zymogens that are cleaved during limited proteolysis to generate their active forms. Alignment contains also inactive enzymes that have substitutions of the catalytic triad residues.
Probab=99.32 E-value=5e-11 Score=115.21 Aligned_cols=161 Identities=24% Similarity=0.236 Sum_probs=98.6
Q ss_pred CCCCCCccccCCC---cceEEEEEEEeCCEEEecccccCCC--CeEEEEEcCC--------CcEEEEEEEEec-------
Q 009784 134 EPNFSLPWQRKRQ---YSSSSSGFAIGGRRVLTNAHSVEHY--TQVKLKKRGS--------DTKYLATVLAIG------- 193 (526)
Q Consensus 134 ~~~~~~P~~~~~~---~~~~GSGfvI~~g~ILT~aHvV~~~--~~i~V~~~~~--------g~~~~a~vv~~d------- 193 (526)
.....+||..... ....|+|++|++.+|||+|||+.+. ..+.|.+... ...+..+-+..+
T Consensus 7 ~~~~~~Pw~v~i~~~~~~~~C~GtlIs~~~VLTaAhC~~~~~~~~~~v~~g~~~~~~~~~~~~~~~v~~~~~hp~y~~~~ 86 (232)
T cd00190 7 AKIGSFPWQVSLQYTGGRHFCGGSLISPRWVLTAAHCVYSSAPSNYTVRLGSHDLSSNEGGGQVIKVKKVIVHPNYNPST 86 (232)
T ss_pred CCCCCCCCEEEEEccCCcEEEEEEEeeCCEEEECHHhcCCCCCccEEEEeCcccccCCCCceEEEEEEEEEECCCCCCCC
Confidence 3444556655332 3478999999999999999999875 5666665311 122334444444
Q ss_pred cCCCeEEEEecccc-cccCceeeecCCC---CcCCCcEEEEeeCCCCCc-------eeEEEEEEeceeeeeccC--Ccee
Q 009784 194 TECDIAMLTVEDDE-FWEGVLPVEFGEL---PALQDAVTVVGYPIGGDT-------ISVTSGVVSRIEILSYVH--GSTE 260 (526)
Q Consensus 194 ~~~DlAlLkv~~~~-~~~~~~pl~l~~~---~~~g~~V~~iG~p~~~~~-------~sv~~GiVs~~~~~~~~~--~~~~ 260 (526)
..+|||||+++.+. +...+.|+.|... ...++.+.+.||...... ......+++...+..... ....
T Consensus 87 ~~~DiAll~L~~~~~~~~~v~picl~~~~~~~~~~~~~~~~G~g~~~~~~~~~~~~~~~~~~~~~~~~C~~~~~~~~~~~ 166 (232)
T cd00190 87 YDNDIALLKLKRPVTLSDNVRPICLPSSGYNLPAGTTCTVSGWGRTSEGGPLPDVLQEVNVPIVSNAECKRAYSYGGTIT 166 (232)
T ss_pred CcCCEEEEEECCcccCCCcccceECCCccccCCCCCEEEEEeCCcCCCCCCCCceeeEEEeeeECHHHhhhhccCcccCC
Confidence 35899999999763 2334788888754 345789999998765321 112223333322221111 0111
Q ss_pred eeEEEE-----cccCCCCCCCCeeecCC---CeEEEEEeecc
Q 009784 261 LLGLQI-----DAAINSGNSGGPAFNDK---GKCVGIAFQSL 294 (526)
Q Consensus 261 ~~~i~~-----da~i~~G~SGGPlvn~~---G~VVGI~~~~~ 294 (526)
...++. +...+.|+|||||+... +.++||.+.+.
T Consensus 167 ~~~~C~~~~~~~~~~c~gdsGgpl~~~~~~~~~lvGI~s~g~ 208 (232)
T cd00190 167 DNMLCAGGLEGGKDACQGDSGGPLVCNDNGRGVLVGIVSWGS 208 (232)
T ss_pred CceEeeCCCCCCCccccCCCCCcEEEEeCCEEEEEEEEehhh
Confidence 122333 33578899999999864 78999998765
No 15
>cd00987 PDZ_serine_protease PDZ domain of tryspin-like serine proteases, such as DegP/HtrA, which are oligomeric proteins involved in heat-shock response, chaperone function, and apoptosis. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.27 E-value=1.5e-11 Score=102.31 Aligned_cols=88 Identities=35% Similarity=0.599 Sum_probs=74.0
Q ss_pred cccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhh
Q 009784 328 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ 406 (526)
Q Consensus 328 ~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~ 406 (526)
+|+|+.++.+ ++..++.++++ ...|++|.+|.++|||++ ||++||+|++|||++|.++.++. ..+..
T Consensus 1 ~~~G~~~~~~-~~~~~~~~~~~-~~~g~~V~~v~~~s~a~~~gl~~GD~I~~Ing~~i~~~~~~~----------~~l~~ 68 (90)
T cd00987 1 PWLGVTVQDL-TPDLAEELGLK-DTKGVLVASVDPGSPAAKAGLKPGDVILAVNGKPVKSVADLR----------RALAE 68 (90)
T ss_pred CccceEEeEC-CHHHHHHcCCC-CCCEEEEEEECCCCHHHHcCCCcCCEEEEECCEECCCHHHHH----------HHHHh
Confidence 5899999999 66666666664 457999999999999998 99999999999999999988764 56666
Q ss_pred cCCCCEEEEEEEECCEEEEEE
Q 009784 407 KYTGDSAAVKVLRDSKILNFN 427 (526)
Q Consensus 407 ~~~G~~v~l~v~R~G~~~~~~ 427 (526)
...|+.+.+++.|+|+..+++
T Consensus 69 ~~~~~~i~l~v~r~g~~~~~~ 89 (90)
T cd00987 69 LKPGDKVTLTVLRGGKELTVT 89 (90)
T ss_pred cCCCCEEEEEEEECCEEEEee
Confidence 556899999999999876654
No 16
>smart00020 Tryp_SPc Trypsin-like serine protease. Many of these are synthesised as inactive precursor zymogens that are cleaved during limited proteolysis to generate their active forms. A few, however, are active as single chain molecules, and others are inactive due to substitutions of the catalytic triad residues.
Probab=99.22 E-value=1.5e-10 Score=112.18 Aligned_cols=164 Identities=24% Similarity=0.239 Sum_probs=99.9
Q ss_pred eeeCCCCCCccccCCC---cceEEEEEEEeCCEEEecccccCCCC--eEEEEEcCCC-------cEEEEEEEEec-----
Q 009784 131 VHTEPNFSLPWQRKRQ---YSSSSSGFAIGGRRVLTNAHSVEHYT--QVKLKKRGSD-------TKYLATVLAIG----- 193 (526)
Q Consensus 131 ~~~~~~~~~P~~~~~~---~~~~GSGfvI~~g~ILT~aHvV~~~~--~i~V~~~~~g-------~~~~a~vv~~d----- 193 (526)
........+||..... ....|+|++|++.+|||+|||+.+.. .+.|.+.... ..+...-+..+
T Consensus 5 G~~~~~~~~Pw~~~i~~~~~~~~C~GtlIs~~~VLTaahC~~~~~~~~~~v~~g~~~~~~~~~~~~~~v~~~~~~p~~~~ 84 (229)
T smart00020 5 GSEANIGSFPWQVSLQYRGGRHFCGGSLISPRWVLTAAHCVYGSDPSNIRVRLGSHDLSSGEEGQVIKVSKVIIHPNYNP 84 (229)
T ss_pred CCcCCCCCCCcEEEEEEcCCCcEEEEEEecCCEEEECHHHcCCCCCcceEEEeCcccCCCCCCceEEeeEEEEECCCCCC
Confidence 3344455566655322 24579999999999999999998753 6777764221 23344444433
Q ss_pred --cCCCeEEEEecccc-cccCceeeecCCC---CcCCCcEEEEeeCCCCCc-----ee---EEEEEEeceeeeeccCC--
Q 009784 194 --TECDIAMLTVEDDE-FWEGVLPVEFGEL---PALQDAVTVVGYPIGGDT-----IS---VTSGVVSRIEILSYVHG-- 257 (526)
Q Consensus 194 --~~~DlAlLkv~~~~-~~~~~~pl~l~~~---~~~g~~V~~iG~p~~~~~-----~s---v~~GiVs~~~~~~~~~~-- 257 (526)
...|||||+++.+. +...+.|+.+... ...++.+.+.||+..... .. ...-+++...+......
T Consensus 85 ~~~~~DiAll~L~~~i~~~~~~~pi~l~~~~~~~~~~~~~~~~g~g~~~~~~~~~~~~~~~~~~~~~~~~~C~~~~~~~~ 164 (229)
T smart00020 85 STYDNDIALLKLKSPVTLSDNVRPICLPSSNYNVPAGTTCTVSGWGRTSEGAGSLPDTLQEVNVPIVSNATCRRAYSGGG 164 (229)
T ss_pred CCCcCCEEEEEECcccCCCCceeeccCCCcccccCCCCEEEEEeCCCCCCCCCcCCCEeeEEEEEEeCHHHhhhhhcccc
Confidence 45899999998763 2345788888753 345788999998776430 01 11222222122111100
Q ss_pred ceeeeEEEE-----cccCCCCCCCCeeecCCC--eEEEEEeecc
Q 009784 258 STELLGLQI-----DAAINSGNSGGPAFNDKG--KCVGIAFQSL 294 (526)
Q Consensus 258 ~~~~~~i~~-----da~i~~G~SGGPlvn~~G--~VVGI~~~~~ 294 (526)
......++. ....++|+||||++...+ .++||++.+.
T Consensus 165 ~~~~~~~C~~~~~~~~~~c~gdsG~pl~~~~~~~~l~Gi~s~g~ 208 (229)
T smart00020 165 AITDNMLCAGGLEGGKDACQGDSGGPLVCNDGRWVLVGIVSWGS 208 (229)
T ss_pred ccCCCcEeecCCCCCCcccCCCCCCeeEEECCCEEEEEEEEECC
Confidence 011112333 345788999999998654 8999998764
No 17
>cd00986 PDZ_LON_protease PDZ domain of ATP-dependent LON serine proteases. Most PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this bacterial subfamily of protease-associated PDZ domains a C-terminal beta-strand is thought to form the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.17 E-value=8.6e-11 Score=95.87 Aligned_cols=72 Identities=28% Similarity=0.366 Sum_probs=63.7
Q ss_pred ccceEEEEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEec
Q 009784 352 QKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLA 431 (526)
Q Consensus 352 ~~Gv~V~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~ 431 (526)
..|++|.+|.++|||+.||++||+|++|||+++.++.++. ..+....+|+.+.+++.|+|+..++++++.
T Consensus 7 ~~Gv~V~~V~~~s~A~~gL~~GD~I~~Ing~~v~~~~~~~----------~~l~~~~~~~~v~l~v~r~g~~~~~~v~l~ 76 (79)
T cd00986 7 YHGVYVTSVVEGMPAAGKLKAGDHIIAVDGKPFKEAEELI----------DYIQSKKEGDTVKLKVKREEKELPEDLILK 76 (79)
T ss_pred ecCEEEEEECCCCchhhCCCCCCEEEEECCEECCCHHHHH----------HHHHhCCCCCEEEEEEEECCEEEEEEEEEe
Confidence 4689999999999998899999999999999999988864 566655679999999999999999999987
Q ss_pred cc
Q 009784 432 TH 433 (526)
Q Consensus 432 ~~ 433 (526)
..
T Consensus 77 ~~ 78 (79)
T cd00986 77 TF 78 (79)
T ss_pred cc
Confidence 64
No 18
>cd00991 PDZ_archaeal_metalloprotease PDZ domain of archaeal zinc metalloprotases, presumably membrane-associated or integral membrane proteases, which may be involved in signalling and regulatory mechanisms. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.15 E-value=1.2e-10 Score=95.22 Aligned_cols=69 Identities=26% Similarity=0.273 Sum_probs=60.8
Q ss_pred CccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEE
Q 009784 351 DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT 429 (526)
Q Consensus 351 ~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~ 429 (526)
...|++|.+|.++|||++ |||+||+|++|||+++.++.++. ..+....+|+++.+++.|+|+..+++++
T Consensus 8 ~~~Gv~V~~V~~~spa~~aGL~~GDiI~~Ing~~v~~~~d~~----------~~l~~~~~g~~v~l~v~r~g~~~~~~~~ 77 (79)
T cd00991 8 AVAGVVIVGVIVGSPAENAVLHTGDVIYSINGTPITTLEDFM----------EALKPTKPGEVITVTVLPSTTKLTNVST 77 (79)
T ss_pred cCCcEEEEEECCCChHHhcCCCCCCEEEEECCEEcCCHHHHH----------HHHhcCCCCCEEEEEEEECCEEEEEEEE
Confidence 357999999999999998 99999999999999999988864 5666656799999999999999888775
No 19
>TIGR01713 typeII_sec_gspC general secretion pathway protein C. This model represents GspC, protein C of the main terminal branch of the general secretion pathway, also called type II secretion. This system transports folded proteins across the bacterial outer membrane and is widely distributed in Gram-negative pathogens.
Probab=99.11 E-value=4e-10 Score=112.52 Aligned_cols=100 Identities=14% Similarity=0.178 Sum_probs=86.3
Q ss_pred HHHHHHHHHHHHcCceeeccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCC
Q 009784 309 PVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAND 387 (526)
Q Consensus 309 ~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~ 387 (526)
..++++++++.++++.. +.|+|+...... ....|++|..+.++++|++ |||+||+|++|||+++.++
T Consensus 159 ~~~~~v~~~l~~~g~~~-~~~lgi~p~~~~-----------g~~~G~~v~~v~~~s~a~~aGLr~GDvIv~ING~~i~~~ 226 (259)
T TIGR01713 159 VVSRRIIEELTKDPQKM-FDYIRLSPVMKN-----------DKLEGYRLNPGKDPSLFYKSGLQDGDIAVALNGLDLRDP 226 (259)
T ss_pred hhHHHHHHHHHHCHHhh-hheEeEEEEEeC-----------CceeEEEEEecCCCCHHHHcCCCCCCEEEEECCEEcCCH
Confidence 45778899999999888 899999876541 2347999999999999999 9999999999999999999
Q ss_pred CCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEe
Q 009784 388 GTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITL 430 (526)
Q Consensus 388 ~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l 430 (526)
.++. ..+.....+++++|+|+|+|+.+++++.+
T Consensus 227 ~~~~----------~~l~~~~~~~~v~l~V~R~G~~~~i~v~~ 259 (259)
T TIGR01713 227 EQAF----------QALQMLREETNLTLTVERDGQREDIYVRF 259 (259)
T ss_pred HHHH----------HHHHhcCCCCeEEEEEEECCEEEEEEEEC
Confidence 8864 67777778899999999999998888764
No 20
>cd00990 PDZ_glycyl_aminopeptidase PDZ domain associated with archaeal and bacterial M61 glycyl-aminopeptidases. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand is presumed to form the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.09 E-value=4.5e-10 Score=91.54 Aligned_cols=77 Identities=22% Similarity=0.384 Sum_probs=63.4
Q ss_pred cccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhh
Q 009784 328 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ 406 (526)
Q Consensus 328 ~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~ 406 (526)
+|+|+.+... ..|++|.+|.++|||++ ||++||+|++|||+++.++.+ .+..
T Consensus 1 ~~~G~~~~~~--------------~~~~~V~~V~~~s~a~~aGl~~GD~I~~Ing~~v~~~~~-------------~l~~ 53 (80)
T cd00990 1 PYLGLTLDKE--------------EGLGKVTFVRDDSPADKAGLVAGDELVAVNGWRVDALQD-------------RLKE 53 (80)
T ss_pred CcccEEEEcc--------------CCcEEEEEECCCChHHHhCCCCCCEEEEECCEEhHHHHH-------------HHHh
Confidence 4778777532 35799999999999999 999999999999999987443 3444
Q ss_pred cCCCCEEEEEEEECCEEEEEEEEec
Q 009784 407 KYTGDSAAVKVLRDSKILNFNITLA 431 (526)
Q Consensus 407 ~~~G~~v~l~v~R~G~~~~~~v~l~ 431 (526)
...|+.+.+++.|+|+..++++++.
T Consensus 54 ~~~~~~v~l~v~r~g~~~~~~v~~~ 78 (80)
T cd00990 54 YQAGDPVELTVFRDDRLIEVPLTLA 78 (80)
T ss_pred cCCCCEEEEEEEECCEEEEEEEEec
Confidence 4578899999999999988888764
No 21
>PRK10779 zinc metallopeptidase RseP; Provisional
Probab=99.00 E-value=8.9e-10 Score=118.86 Aligned_cols=132 Identities=12% Similarity=0.080 Sum_probs=94.6
Q ss_pred eEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEeccc
Q 009784 355 VRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATH 433 (526)
Q Consensus 355 v~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~~ 433 (526)
.+|.+|.++|||++ |||+||+|++|||++|.+++++. ..+....+|++++++|+|+|+.+++++++...
T Consensus 128 ~lV~~V~~~SpA~kAGLk~GDvI~~vnG~~V~~~~~l~----------~~v~~~~~g~~v~v~v~R~gk~~~~~v~l~~~ 197 (449)
T PRK10779 128 PVVGEIAPNSIAAQAQIAPGTELKAVDGIETPDWDAVR----------LALVSKIGDESTTITVAPFGSDQRRDKTLDLR 197 (449)
T ss_pred ccccccCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHhhccCCceEEEEEeCCccceEEEEeccc
Confidence 46899999999999 99999999999999999999975 56667778899999999999999998888654
Q ss_pred ccccCCCCCCCCCceEEEeeEEEEeccccceeeeeeeecc--hhhhccccccceeeeecccchhhhHHHHHH
Q 009784 434 RRLIPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERIMN--MKLRSSFWTSSCIQCHNCQMSSLLWCLRCL 503 (526)
Q Consensus 434 ~~~~p~~~~~~~p~~~i~gG~~f~~lt~~~~~~~~~~i~~--~~~~sg~~~~~~~~~~~~~~~~~~~~~~~~ 503 (526)
+....... .. .. ...| +.+++-.. ...+..+.+ .+.++|++.||+|...|.+-.++...++-+
T Consensus 198 ~~~~~~~~--~~-~~-~~lG--l~~~~~~~-~~vV~~V~~~SpA~~AGL~~GDvIl~Ing~~V~s~~dl~~~ 262 (449)
T PRK10779 198 HWAFEPDK--QD-PV-SSLG--IRPRGPQI-EPVLAEVQPNSAASKAGLQAGDRIVKVDGQPLTQWQTFVTL 262 (449)
T ss_pred ccccCccc--cc-hh-hccc--ccccCCCc-CcEEEeeCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHHHH
Confidence 32111000 00 00 0123 23333211 123444544 456799999999999999888776666543
No 22
>TIGR02037 degP_htrA_DO periplasmic serine protease, Do/DeqQ family. This family consists of a set proteins various designated DegP, heat shock protein HtrA, and protease DO. The ortholog in Pseudomonas aeruginosa is designated MucD and is found in an operon that controls mucoid phenotype. This family also includes the DegQ (HhoA) paralog in E. coli which can rescue a DegP mutant, but not the smaller DegS paralog, which cannot. Members of this family are located in the periplasm and have separable functions as both protease and chaperone. Members have a trypsin domain and two copies of a PDZ domain. This protein protects bacteria from thermal and other stresses and may be important for the survival of bacterial pathogens.// The chaperone function is dominant at low temperatures, whereas the proteolytic activity is turned on at elevated temperatures.
Probab=98.89 E-value=3.3e-09 Score=113.88 Aligned_cols=90 Identities=24% Similarity=0.473 Sum_probs=78.6
Q ss_pred ccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhh
Q 009784 327 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVS 405 (526)
Q Consensus 327 ~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~ 405 (526)
..++|+.+..+ ++..++.++++....|++|.+|.++|||++ ||++||+|++|||++|.++.++. +.+.
T Consensus 337 ~~~lGi~~~~l-~~~~~~~~~l~~~~~Gv~V~~V~~~SpA~~aGL~~GDvI~~Ing~~V~s~~d~~----------~~l~ 405 (428)
T TIGR02037 337 NPFLGLTVANL-SPEIRKELRLKGDVKGVVVTKVVSGSPAARAGLQPGDVILSVNQQPVSSVAELR----------KVLD 405 (428)
T ss_pred ccccceEEecC-CHHHHHHcCCCcCcCceEEEEeCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHH
Confidence 35789999988 788888889886668999999999999999 99999999999999999988864 6676
Q ss_pred hcCCCCEEEEEEEECCEEEEEE
Q 009784 406 QKYTGDSAAVKVLRDSKILNFN 427 (526)
Q Consensus 406 ~~~~G~~v~l~v~R~G~~~~~~ 427 (526)
..+.|+.+.|+|+|+|+...+.
T Consensus 406 ~~~~g~~v~l~v~R~g~~~~~~ 427 (428)
T TIGR02037 406 RAKKGGRVALLILRGGATIFVT 427 (428)
T ss_pred hcCCCCEEEEEEEECCEEEEEE
Confidence 6667999999999999987654
No 23
>cd00989 PDZ_metalloprotease PDZ domain of bacterial and plant zinc metalloprotases, presumably membrane-associated or integral membrane proteases, which may be involved in signalling and regulatory mechanisms. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=98.88 E-value=4.1e-09 Score=85.49 Aligned_cols=66 Identities=24% Similarity=0.346 Sum_probs=55.7
Q ss_pred cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEE
Q 009784 353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT 429 (526)
Q Consensus 353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~ 429 (526)
..++|.+|.++|||++ ||++||+|++|||+++.++.++. ..+... .++.+.+++.|+|+..+++++
T Consensus 12 ~~~~V~~v~~~s~a~~~gl~~GD~I~~ing~~i~~~~~~~----------~~l~~~-~~~~~~l~v~r~~~~~~~~l~ 78 (79)
T cd00989 12 IEPVIGEVVPGSPAAKAGLKAGDRILAINGQKIKSWEDLV----------DAVQEN-PGKPLTLTVERNGETITLTLT 78 (79)
T ss_pred cCcEEEeECCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHHHHC-CCceEEEEEEECCEEEEEEec
Confidence 3488999999999998 99999999999999999988763 455443 478999999999988777664
No 24
>cd00988 PDZ_CTP_protease PDZ domain of C-terminal processing-, tail-specific-, and tricorn proteases, which function in posttranslational protein processing, maturation, and disassembly or degradation, in Bacteria, Archaea, and plant chloroplasts. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=98.85 E-value=1e-08 Score=84.40 Aligned_cols=68 Identities=24% Similarity=0.343 Sum_probs=57.0
Q ss_pred ccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCC--CCcccccccchhhhhhhhhcCCCCEEEEEEEEC-CEEEEEE
Q 009784 352 QKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAND--GTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD-SKILNFN 427 (526)
Q Consensus 352 ~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~--~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~-G~~~~~~ 427 (526)
..+++|..|.+++||++ ||++||+|++|||+++.++ .++. ..+.. ..|+.+.+++.|+ |+..+++
T Consensus 12 ~~~~~V~~v~~~s~a~~~gl~~GD~I~~vng~~i~~~~~~~~~----------~~l~~-~~~~~i~l~v~r~~~~~~~~~ 80 (85)
T cd00988 12 DGGLVITSVLPGSPAAKAGIKAGDIIVAIDGEPVDGLSLEDVV----------KLLRG-KAGTKVRLTLKRGDGEPREVT 80 (85)
T ss_pred CCeEEEEEecCCCCHHHcCCCCCCEEEEECCEEcCCCCHHHHH----------HHhcC-CCCCEEEEEEEcCCCCEEEEE
Confidence 36799999999999999 9999999999999999998 5642 44433 4689999999999 8888877
Q ss_pred EEe
Q 009784 428 ITL 430 (526)
Q Consensus 428 v~l 430 (526)
+++
T Consensus 81 ~~~ 83 (85)
T cd00988 81 LTR 83 (85)
T ss_pred EEE
Confidence 654
No 25
>COG3591 V8-like Glu-specific endopeptidase [Amino acid transport and metabolism]
Probab=98.82 E-value=9.2e-08 Score=94.07 Aligned_cols=161 Identities=22% Similarity=0.237 Sum_probs=96.2
Q ss_pred eEEEEEEEeCCEEEecccccCCCC----eEEEEEc---CC-CcEEE--EEEEE-ecc---CCCeEEEEecccccc-----
Q 009784 149 SSSSGFAIGGRRVLTNAHSVEHYT----QVKLKKR---GS-DTKYL--ATVLA-IGT---ECDIAMLTVEDDEFW----- 209 (526)
Q Consensus 149 ~~GSGfvI~~g~ILT~aHvV~~~~----~i~V~~~---~~-g~~~~--a~vv~-~d~---~~DlAlLkv~~~~~~----- 209 (526)
..|++|+|.+..+||++||+.... ++.+..+ ++ +..+. ..... ... +.|.+...+.+..+.
T Consensus 64 ~~~~~~lI~pntvLTa~Hc~~s~~~G~~~~~~~p~g~~~~~~~~~~~~~~~~~~~~g~~~~~d~~~~~v~~~~~~~g~~~ 143 (251)
T COG3591 64 LCTAATLIGPNTVLTAGHCIYSPDYGEDDIAAAPPGVNSDGGPFYGITKIEIRVYPGELYKEDGASYDVGEAALESGINI 143 (251)
T ss_pred ceeeEEEEcCceEEEeeeEEecCCCChhhhhhcCCcccCCCCCCCceeeEEEEecCCceeccCCceeeccHHHhccCCCc
Confidence 455669999999999999997433 2222221 11 21221 11111 122 456666666543321
Q ss_pred -cCce--eeecCCCCcCCCcEEEEeeCCCCCc---eeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecCC
Q 009784 210 -EGVL--PVEFGELPALQDAVTVVGYPIGGDT---ISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDK 283 (526)
Q Consensus 210 -~~~~--pl~l~~~~~~g~~V~~iG~p~~~~~---~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~ 283 (526)
.... ...+....++++.+.++|||.+... .-.+.+.|..+.. ..++.++.+.+|+||+|+++.+
T Consensus 144 ~~~~~~~~~~~~~~~~~~d~i~v~GYP~dk~~~~~~~e~t~~v~~~~~----------~~l~y~~dT~pG~SGSpv~~~~ 213 (251)
T COG3591 144 GDVVNYLKRNTASEAKANDRITVIGYPGDKPNIGTMWESTGKVNSIKG----------NKLFYDADTLPGSSGSPVLISK 213 (251)
T ss_pred cccccccccccccccccCceeEEEeccCCCCcceeEeeecceeEEEec----------ceEEEEecccCCCCCCceEecC
Confidence 1111 2223334467788999999987652 1223344433321 2578889999999999999998
Q ss_pred CeEEEEEeeccccCcccccccc-ccHHHHHHHHHHHH
Q 009784 284 GKCVGIAFQSLKHEDVENIGYV-IPTPVIMHFIQDYE 319 (526)
Q Consensus 284 G~VVGI~~~~~~~~~~~~~~~a-IP~~~i~~~l~~l~ 319 (526)
.++||+++.+....+....+++ .-...++++++++.
T Consensus 214 ~~vigv~~~g~~~~~~~~~n~~vr~t~~~~~~I~~~~ 250 (251)
T COG3591 214 DEVIGVHYNGPGANGGSLANNAVRLTPEILNFIQQNI 250 (251)
T ss_pred ceEEEEEecCCCcccccccCcceEecHHHHHHHHHhh
Confidence 8999999987754333344433 44566777777653
No 26
>cd00136 PDZ PDZ domain, also called DHR (Dlg homologous region) or GLGF (after a conserved sequence motif). Many PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. Heterodimerization through PDZ-PDZ domain interactions adds to the domain's versatility, and PDZ domain-mediated interactions may be modulated dynamically through target phosphorylation. Some PDZ domains play a role in scaffolding supramolecular complexes. PDZ domains are found in diverse signaling proteins in bacteria, archebacteria, and eurkayotes. This CD contains two distinct structural subgroups with either a N- or C-terminal beta-strand forming the peptide-binding groove base. The circular permutation placing the strand on the N-terminus appears to be found in Eumetazoa only, while the C-terminal variant is found in all three kingdoms of life, and seems to co-occur with protease domains. PDZ domains have been named after PSD95(pos
Probab=98.59 E-value=8.4e-08 Score=75.87 Aligned_cols=55 Identities=29% Similarity=0.527 Sum_probs=46.2
Q ss_pred cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCC--CCcccccccchhhhhhhhhcCCCCEEEEEEE
Q 009784 353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAND--GTVPFRHGERIGFSYLVSQKYTGDSAAVKVL 418 (526)
Q Consensus 353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~--~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~ 418 (526)
.|++|.+|.+++||++ ||++||+|++|||+++.++ .++. ..+... .|++++|+++
T Consensus 13 ~~~~V~~v~~~s~a~~~gl~~GD~I~~Ing~~v~~~~~~~~~----------~~l~~~-~g~~v~l~v~ 70 (70)
T cd00136 13 GGVVVLSVEPGSPAERAGLQAGDVILAVNGTDVKNLTLEDVA----------ELLKKE-VGEKVTLTVR 70 (70)
T ss_pred CCEEEEEeCCCCHHHHcCCCCCCEEEEECCEECCCCCHHHHH----------HHHhhC-CCCeEEEEEC
Confidence 4899999999999999 9999999999999999998 5543 555554 4889998763
No 27
>TIGR00054 RIP metalloprotease RseP. A model that detects fragments as well matches a number of members of the PEPTIDASE FAMILY S2C. The region of match appears not to overlap the active site domain.
Probab=98.52 E-value=1.7e-07 Score=100.45 Aligned_cols=113 Identities=18% Similarity=0.161 Sum_probs=82.6
Q ss_pred ccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEe
Q 009784 352 QKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITL 430 (526)
Q Consensus 352 ~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l 430 (526)
..|.+|.+|.++|||++ |||+||+|+++||+++.++.++. ..+.... +++.+++.|+|+..++++++
T Consensus 127 ~~g~~V~~V~~~SpA~~AGL~~GDvI~~vng~~v~~~~dl~----------~~ia~~~--~~v~~~I~r~g~~~~l~v~l 194 (420)
T TIGR00054 127 EVGPVIELLDKNSIALEAGIEPGDEILSVNGNKIPGFKDVR----------QQIADIA--GEPMVEILAERENWTFEVMK 194 (420)
T ss_pred CCCceeeccCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHhhc--ccceEEEEEecCceEecccc
Confidence 36789999999999999 99999999999999999999875 4455543 68999999999887665543
Q ss_pred cccccccCCCCCCCCCceEEEeeEEEEeccccceeeeeeeecc--hhhhccccccceeeeecccchhhhHHHH
Q 009784 431 ATHRRLIPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERIMN--MKLRSSFWTSSCIQCHNCQMSSLLWCLR 501 (526)
Q Consensus 431 ~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt~~~~~~~~~~i~~--~~~~sg~~~~~~~~~~~~~~~~~~~~~~ 501 (526)
.-.+. .+. ....+..+.+ .+.++|++.||+|...|.+-.++...++
T Consensus 195 ~~~~~---------~~~----------------~g~vV~~V~~~SpA~~aGL~~GD~Iv~Vng~~V~s~~dl~ 242 (420)
T TIGR00054 195 ELIPR---------GPK----------------IEPVLSDVTPNSPAEKAGLKEGDYIQSINGEKLRSWTDFV 242 (420)
T ss_pred cceec---------CCC----------------cCcEEEEECCCCHHHHcCCCCCCEEEEECCEECCCHHHHH
Confidence 31110 000 0122233333 4557999999999999998877655544
No 28
>TIGR00054 RIP metalloprotease RseP. A model that detects fragments as well matches a number of members of the PEPTIDASE FAMILY S2C. The region of match appears not to overlap the active site domain.
Probab=98.47 E-value=2.4e-07 Score=99.22 Aligned_cols=69 Identities=25% Similarity=0.323 Sum_probs=60.8
Q ss_pred cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEec
Q 009784 353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLA 431 (526)
Q Consensus 353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~ 431 (526)
.|++|.+|.++|||++ |||+||+|++|||++|.+++|+. ..+.. .+|+++.++++|+|+..++++++.
T Consensus 203 ~g~vV~~V~~~SpA~~aGL~~GD~Iv~Vng~~V~s~~dl~----------~~l~~-~~~~~v~l~v~R~g~~~~~~v~~~ 271 (420)
T TIGR00054 203 IEPVLSDVTPNSPAEKAGLKEGDYIQSINGEKLRSWTDFV----------SAVKE-NPGKSMDIKVERNGETLSISLTPE 271 (420)
T ss_pred cCcEEEEECCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHHHh-CCCCceEEEEEECCEEEEEEEEEc
Confidence 4799999999999999 99999999999999999999874 45544 578899999999999999988875
Q ss_pred c
Q 009784 432 T 432 (526)
Q Consensus 432 ~ 432 (526)
.
T Consensus 272 ~ 272 (420)
T TIGR00054 272 A 272 (420)
T ss_pred C
Confidence 3
No 29
>smart00228 PDZ Domain present in PSD-95, Dlg, and ZO-1/2. Also called DHR (Dlg homologous region) or GLGF (relatively well conserved tetrapeptide in these domains). Some PDZs have been shown to bind C-terminal polypeptides; others appear to bind internal (non-C-terminal) polypeptides. Different PDZs possess different binding specificities.
Probab=98.46 E-value=4.8e-07 Score=73.84 Aligned_cols=59 Identities=27% Similarity=0.402 Sum_probs=47.9
Q ss_pred cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECC
Q 009784 353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDS 421 (526)
Q Consensus 353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G 421 (526)
.|++|..|.+++||+. ||++||+|++|||+++.+..+.. ........++.+.+++.|++
T Consensus 26 ~~~~i~~v~~~s~a~~~gl~~GD~I~~In~~~v~~~~~~~----------~~~~~~~~~~~~~l~i~r~~ 85 (85)
T smart00228 26 GGVVVSSVVPGSPAAKAGLKVGDVILEVNGTSVEGLTHLE----------AVDLLKKAGGKVTLTVLRGG 85 (85)
T ss_pred CCEEEEEECCCCHHHHcCCCCCCEEEEECCEECCCCCHHH----------HHHHHHhCCCeEEEEEEeCC
Confidence 6899999999999999 99999999999999999776542 22222334679999999975
No 30
>PRK10779 zinc metallopeptidase RseP; Provisional
Probab=98.37 E-value=6e-07 Score=97.04 Aligned_cols=68 Identities=24% Similarity=0.324 Sum_probs=60.1
Q ss_pred ceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEecc
Q 009784 354 GVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLAT 432 (526)
Q Consensus 354 Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~ 432 (526)
+++|.+|.++|||++ |||+||+|++|||++|.++.|+. ..+.. ..|+.+.+++.|+|+..++++++..
T Consensus 222 ~~vV~~V~~~SpA~~AGL~~GDvIl~Ing~~V~s~~dl~----------~~l~~-~~~~~v~l~v~R~g~~~~~~v~~~~ 290 (449)
T PRK10779 222 EPVLAEVQPNSAASKAGLQAGDRIVKVDGQPLTQWQTFV----------TLVRD-NPGKPLALEIERQGSPLSLTLTPDS 290 (449)
T ss_pred CcEEEeeCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHh-CCCCEEEEEEEECCEEEEEEEEeee
Confidence 588999999999999 99999999999999999999874 45544 5788999999999999999988753
No 31
>TIGR00225 prc C-terminal peptidase (prc). A C-terminal peptidase with different substrates in different species including processing of D1 protein of the photosystem II reaction center in higher plants and cleavage of a peptide of 11 residues from the precursor form of penicillin-binding protein in E.coli E.coli and H influenza have the most distal branch of the tree and their proteins have an N-terminal 200 amino acids that show no homology to other proteins in the database.
Probab=98.33 E-value=1.3e-06 Score=90.83 Aligned_cols=70 Identities=21% Similarity=0.327 Sum_probs=57.5
Q ss_pred cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCC--CcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEE
Q 009784 353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDG--TVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT 429 (526)
Q Consensus 353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~--~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~ 429 (526)
.+++|.+|.++|||++ ||++||+|++|||++|.++. ++ ...+ ....|+++.+++.|+|+..+++++
T Consensus 62 ~~~~V~~V~~~spA~~aGL~~GD~I~~Ing~~v~~~~~~~~----------~~~l-~~~~g~~v~l~v~R~g~~~~~~v~ 130 (334)
T TIGR00225 62 GEIVIVSPFEGSPAEKAGIKPGDKIIKINGKSVAGMSLDDA----------VALI-RGKKGTKVSLEILRAGKSKPLTFT 130 (334)
T ss_pred CEEEEEEeCCCChHHHcCCCCCCEEEEECCEECCCCCHHHH----------HHhc-cCCCCCEEEEEEEeCCCCceEEEE
Confidence 5799999999999999 99999999999999998863 22 1222 335789999999999988888877
Q ss_pred eccc
Q 009784 430 LATH 433 (526)
Q Consensus 430 l~~~ 433 (526)
+...
T Consensus 131 l~~~ 134 (334)
T TIGR00225 131 LKRD 134 (334)
T ss_pred EEEE
Confidence 7654
No 32
>PF00595 PDZ: PDZ domain (Also known as DHR or GLGF) Coordinates are not yet available; InterPro: IPR001478 PDZ domains are found in diverse signalling proteins in bacteria, yeasts, plants, insects and vertebrates [, ]. PDZ domains can occur in one or multiple copies and are nearly always found in cytoplasmic proteins. They bind either the carboxyl-terminal sequences of proteins or internal peptide sequences []. In most cases, interaction between a PDZ domain and its target is constitutive, with a binding affinity of 1 to 10 microns. However, agonist-dependent activation of cell surface receptors is sometimes required to promote interaction with a PDZ protein. PDZ domain proteins are frequently associated with the plasma membrane, a compartment where high concentrations of phosphatidylinositol 4,5-bisphosphate (PIP2) are found. Direct interaction between PIP2 and a subset of class II PDZ domains (syntenin, CASK, Tiam-1) has been demonstrated. PDZ domains consist of 80 to 90 amino acids comprising six beta-strands (beta-A to beta-F) and two alpha-helices, A and B, compactly arranged in a globular structure. Peptide binding of the ligand takes place in an elongated surface groove as an anti-parallel beta-strand interacts with the beta-B strand and the B helix. The structure of PDZ domains allows binding to a free carboxylate group at the end of a peptide through a carboxylate-binding loop between the beta-A and beta-B strands.; GO: 0005515 protein binding; PDB: 3AXA_A 1WF8_A 1QAV_B 1QAU_A 1B8Q_A 1MC7_A 2KAW_A 1I16_A 1VB7_A 1WI4_A ....
Probab=98.26 E-value=1.5e-06 Score=71.04 Aligned_cols=72 Identities=24% Similarity=0.308 Sum_probs=53.0
Q ss_pred ccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhh
Q 009784 327 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVS 405 (526)
Q Consensus 327 ~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~ 405 (526)
...+|+.+...... ...+++|.+|.++|||++ ||++||.|++|||+.+.++.... ...++.
T Consensus 9 ~~~lG~~l~~~~~~----------~~~~~~V~~v~~~~~a~~~gl~~GD~Il~INg~~v~~~~~~~--------~~~~l~ 70 (81)
T PF00595_consen 9 NGPLGFTLRGGSDN----------DEKGVFVSSVVPGSPAERAGLKVGDRILEINGQSVRGMSHDE--------VVQLLK 70 (81)
T ss_dssp TSBSSEEEEEESTS----------SSEEEEEEEECTTSHHHHHTSSTTEEEEEETTEESTTSBHHH--------HHHHHH
T ss_pred CCCcCEEEEecCCC----------CcCCEEEEEEeCCChHHhcccchhhhhheeCCEeCCCCCHHH--------HHHHHH
Confidence 45688888866210 126899999999999999 99999999999999999886532 112333
Q ss_pred hcCCCCEEEEEEE
Q 009784 406 QKYTGDSAAVKVL 418 (526)
Q Consensus 406 ~~~~G~~v~l~v~ 418 (526)
. .+.+++|+|+
T Consensus 71 ~--~~~~v~L~V~ 81 (81)
T PF00595_consen 71 S--ASNPVTLTVQ 81 (81)
T ss_dssp H--STSEEEEEEE
T ss_pred C--CCCcEEEEEC
Confidence 3 3448888874
No 33
>PRK10139 serine endoprotease; Provisional
Probab=98.20 E-value=2.1e-06 Score=92.79 Aligned_cols=64 Identities=17% Similarity=0.342 Sum_probs=55.9
Q ss_pred cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEE
Q 009784 353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNI 428 (526)
Q Consensus 353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v 428 (526)
.|++|.+|.++|||++ |||+||+|++|||++|.++.++. +.+... .+++.|+|+|+|+.+.+.+
T Consensus 390 ~Gv~V~~V~~~spA~~aGL~~GD~I~~Ing~~v~~~~~~~----------~~l~~~--~~~v~l~v~R~g~~~~~~~ 454 (455)
T PRK10139 390 KGIKIDEVVKGSPAAQAGLQKDDVIIGVNRDRVNSIAEMR----------KVLAAK--PAIIALQIVRGNESIYLLL 454 (455)
T ss_pred CceEEEEeCCCChHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHhC--CCeEEEEEEECCEEEEEEe
Confidence 5899999999999999 99999999999999999999874 556543 3789999999999877654
No 34
>TIGR03279 cyano_FeS_chp putative FeS-containing Cyanobacterial-specific oxidoreductase. Members of this protein family are predicted FeS-containing oxidoreductases of unknown function, apparently restricted to and universal across the Cyanobacteria. The high trusted cutoff score for this model, 700 bits, excludes homologs from other lineages. This exclusion seems justified because a significant number of sequence positions are simultaneously unique to and invariant across the Cyanobacteria, suggesting a specialized, conserved function, perhaps related to photosynthesis. A distantly related protein family, TIGR03278, in universal in and restricted to archaeal methanogens, and may be linked to methanogenesis.
Probab=98.20 E-value=1.9e-06 Score=90.94 Aligned_cols=62 Identities=19% Similarity=0.303 Sum_probs=51.7
Q ss_pred EEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEE-ECCEEEEEEEEecc
Q 009784 357 IRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVL-RDSKILNFNITLAT 432 (526)
Q Consensus 357 V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~-R~G~~~~~~v~l~~ 432 (526)
|.+|.|+|+|++ ||++||+|++|||++|.++.|+. ..+ .++.+.++|. |+|+..++++....
T Consensus 2 I~~V~pgSpAe~AGLe~GD~IlsING~~V~Dw~D~~----------~~l----~~e~l~L~V~~rdGe~~~l~Ie~~~ 65 (433)
T TIGR03279 2 ISAVLPGSIAEELGFEPGDALVSINGVAPRDLIDYQ----------FLC----ADEELELEVLDANGESHQIEIEKDL 65 (433)
T ss_pred cCCcCCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHh----cCCcEEEEEEcCCCeEEEEEEecCC
Confidence 678999999999 99999999999999999998863 333 2467999997 89988888877643
No 35
>KOG3627 consensus Trypsin [Amino acid transport and metabolism]
Probab=98.16 E-value=9.9e-05 Score=73.19 Aligned_cols=167 Identities=22% Similarity=0.233 Sum_probs=97.6
Q ss_pred EEeeeeCCCCCCccccCCCc----ceEEEEEEEeCCEEEecccccCCCC--eEEEEEcC--------CC---cEE-EEEE
Q 009784 128 VFCVHTEPNFSLPWQRKRQY----SSSSSGFAIGGRRVLTNAHSVEHYT--QVKLKKRG--------SD---TKY-LATV 189 (526)
Q Consensus 128 I~~~~~~~~~~~P~~~~~~~----~~~GSGfvI~~g~ILT~aHvV~~~~--~i~V~~~~--------~g---~~~-~a~v 189 (526)
|.........++||+..... ...|.|.+|++.||+|++||+.+.. .+.|.+.. .+ ... ..++
T Consensus 13 i~~g~~~~~~~~Pw~~~l~~~~~~~~~Cggsli~~~~vltaaHC~~~~~~~~~~V~~G~~~~~~~~~~~~~~~~~~v~~~ 92 (256)
T KOG3627|consen 13 IVGGTEAEPGSFPWQVSLQYGGNGRHLCGGSLISPRWVLTAAHCVKGASASLYTVRLGEHDINLSVSEGEEQLVGDVEKI 92 (256)
T ss_pred EeCCccCCCCCCCCEEEEEECCCcceeeeeEEeeCCEEEEChhhCCCCCCcceEEEECccccccccccCchhhhceeeEE
Confidence 34444444557788754433 2378888888889999999999865 66666521 11 111 1122
Q ss_pred EEecc-------C-CCeEEEEeccc-ccccCceeeecCCCC----cCC-CcEEEEeeCCCC----C-c---eeEEEEEEe
Q 009784 190 LAIGT-------E-CDIAMLTVEDD-EFWEGVLPVEFGELP----ALQ-DAVTVVGYPIGG----D-T---ISVTSGVVS 247 (526)
Q Consensus 190 v~~d~-------~-~DlAlLkv~~~-~~~~~~~pl~l~~~~----~~g-~~V~~iG~p~~~----~-~---~sv~~GiVs 247 (526)
+ .++ . .|||||+++.+ .+.+.+.|+.+.... ..+ ..+++.||.... . . ..+..-+++
T Consensus 93 i-~H~~y~~~~~~~nDiall~l~~~v~~~~~i~piclp~~~~~~~~~~~~~~~v~GWG~~~~~~~~~~~~L~~~~v~i~~ 171 (256)
T KOG3627|consen 93 I-VHPNYNPRTLENNDIALLRLSEPVTFSSHIQPICLPSSADPYFPPGGTTCLVSGWGRTESGGGPLPDTLQEVDVPIIS 171 (256)
T ss_pred E-ECCCCCCCCCCCCCEEEEEECCCcccCCcccccCCCCCcccCCCCCCCEEEEEeCCCcCCCCCCCCceeEEEEEeEcC
Confidence 2 222 2 79999999975 355678888875322 223 677788875431 1 1 111233333
Q ss_pred ceeeeeccCCc--eeeeEEEEc-----ccCCCCCCCCeeecCC---CeEEEEEeeccc
Q 009784 248 RIEILSYVHGS--TELLGLQID-----AAINSGNSGGPAFNDK---GKCVGIAFQSLK 295 (526)
Q Consensus 248 ~~~~~~~~~~~--~~~~~i~~d-----a~i~~G~SGGPlvn~~---G~VVGI~~~~~~ 295 (526)
...+....... .....+++. ...|.|+|||||+..+ ..++||++.+..
T Consensus 172 ~~~C~~~~~~~~~~~~~~~Ca~~~~~~~~~C~GDSGGPLv~~~~~~~~~~GivS~G~~ 229 (256)
T KOG3627|consen 172 NSECRRAYGGLGTITDTMLCAGGPEGGKDACQGDSGGPLVCEDNGRWVLVGIVSWGSG 229 (256)
T ss_pred hhHhcccccCccccCCCEEeeCccCCCCccccCCCCCeEEEeeCCcEEEEEEEEecCC
Confidence 32332221110 111135554 2468899999999764 699999988753
No 36
>PLN00049 carboxyl-terminal processing protease; Provisional
Probab=98.15 E-value=4.9e-06 Score=88.32 Aligned_cols=69 Identities=20% Similarity=0.302 Sum_probs=54.9
Q ss_pred cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEe
Q 009784 353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITL 430 (526)
Q Consensus 353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l 430 (526)
.|++|..|.++|||++ ||++||+|++|||++|.++.... +...+ ....|+.+.|+|.|+|+..+++++-
T Consensus 102 ~g~~V~~V~~~SPA~~aGl~~GD~Iv~InG~~v~~~~~~~--------~~~~l-~g~~g~~v~ltv~r~g~~~~~~l~r 171 (389)
T PLN00049 102 AGLVVVAPAPGGPAARAGIRPGDVILAIDGTSTEGLSLYE--------AADRL-QGPEGSSVELTLRRGPETRLVTLTR 171 (389)
T ss_pred CcEEEEEeCCCChHHHcCCCCCCEEEEECCEECCCCCHHH--------HHHHH-hcCCCCEEEEEEEECCEEEEEEEEe
Confidence 3899999999999999 99999999999999998753210 11333 3457899999999999887776654
No 37
>PF14685 Tricorn_PDZ: Tricorn protease PDZ domain; PDB: 1N6F_D 1N6D_C 1N6E_C 1K32_A.
Probab=98.14 E-value=1.6e-05 Score=66.12 Aligned_cols=65 Identities=23% Similarity=0.385 Sum_probs=43.8
Q ss_pred ccceEEEEeCCC--------CcccC-C--CCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEEC
Q 009784 352 QKGVRIRRVDPT--------APESE-V--LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD 420 (526)
Q Consensus 352 ~~Gv~V~~V~~~--------spA~~-G--L~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~ 420 (526)
..+..|.+|.++ ||..+ | +++||+|++|||+++....++ +.+...+.|+.|.|+|.+.
T Consensus 11 ~~~y~I~~I~~gd~~~~~~~sPL~~pGv~v~~GD~I~aInG~~v~~~~~~-----------~~lL~~~agk~V~Ltv~~~ 79 (88)
T PF14685_consen 11 NGGYRIARIYPGDPWNPNARSPLAQPGVDVREGDYILAINGQPVTADANP-----------YRLLEGKAGKQVLLTVNRK 79 (88)
T ss_dssp TTEEEEEEE-BS-TTSSS-B-GGGGGS----TT-EEEEETTEE-BTTB-H-----------HHHHHTTTTSEEEEEEE-S
T ss_pred CCEEEEEEEeCCCCCCccccCCccCCCCCCCCCCEEEEECCEECCCCCCH-----------HHHhcccCCCEEEEEEecC
Confidence 367889999875 67666 5 569999999999999988776 3444556899999999996
Q ss_pred C-EEEEEE
Q 009784 421 S-KILNFN 427 (526)
Q Consensus 421 G-~~~~~~ 427 (526)
+ +.+++.
T Consensus 80 ~~~~R~v~ 87 (88)
T PF14685_consen 80 PGGARTVV 87 (88)
T ss_dssp TT-EEEEE
T ss_pred CCCceEEE
Confidence 6 455554
No 38
>TIGR02860 spore_IV_B stage IV sporulation protein B. SpoIVB, the stage IV sporulation protein B of endospore-forming bacteria such as Bacillus subtilis, is a serine proteinase, expressed in the spore (rather than mother cell) compartment, that participates in a proteolytic activation cascade for Sigma-K. It appears to be universal among endospore-forming bacteria and occurs nowhere else.
Probab=98.11 E-value=4e-06 Score=88.03 Aligned_cols=69 Identities=25% Similarity=0.364 Sum_probs=57.1
Q ss_pred ccceEEEEeC--------CCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCE
Q 009784 352 QKGVRIRRVD--------PTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK 422 (526)
Q Consensus 352 ~~Gv~V~~V~--------~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~ 422 (526)
..||+|.... .+|||++ |||+||+|++|||++|.++.++. +.+... .|+++.++|.|+|+
T Consensus 104 t~GVlVvg~~~v~~~~g~~~SPAa~AGLq~GDiIvsING~~V~s~~DL~----------~iL~~~-~g~~V~LtV~R~Ge 172 (402)
T TIGR02860 104 TKGVLVVGFSDIETEKGKIHSPGEEAGIQIGDRILKINGEKIKNMDDLA----------NLINKA-GGEKLTLTIERGGK 172 (402)
T ss_pred cCEEEEEEEEcccccCCCCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHHHhC-CCCeEEEEEEECCE
Confidence 4688886552 2589988 99999999999999999999874 555554 48999999999999
Q ss_pred EEEEEEEec
Q 009784 423 ILNFNITLA 431 (526)
Q Consensus 423 ~~~~~v~l~ 431 (526)
..++++++.
T Consensus 173 ~~tv~V~Pv 181 (402)
T TIGR02860 173 IIETVIKPV 181 (402)
T ss_pred EEEEEEEEe
Confidence 998888754
No 39
>PRK10942 serine endoprotease; Provisional
Probab=98.08 E-value=5.5e-06 Score=90.01 Aligned_cols=64 Identities=20% Similarity=0.343 Sum_probs=55.5
Q ss_pred cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEE
Q 009784 353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNI 428 (526)
Q Consensus 353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v 428 (526)
.|++|.+|.++|+|++ ||++||+|++|||++|.++.++. +.+.. . ++.+.|+|+|+|+.+.+.+
T Consensus 408 ~gvvV~~V~~~S~A~~aGL~~GDvIv~VNg~~V~s~~dl~----------~~l~~-~-~~~v~l~V~R~g~~~~v~~ 472 (473)
T PRK10942 408 KGVVVDNVKPGTPAAQIGLKKGDVIIGANQQPVKNIAELR----------KILDS-K-PSVLALNIQRGDSSIYLLM 472 (473)
T ss_pred CCeEEEEeCCCChHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHh-C-CCeEEEEEEECCEEEEEEe
Confidence 5899999999999998 99999999999999999999874 55554 3 4789999999999877654
No 40
>cd00992 PDZ_signaling PDZ domain found in a variety of Eumetazoan signaling molecules, often in tandem arrangements. May be responsible for specific protein-protein interactions, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of PDZ domains an N-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in proteases.
Probab=98.02 E-value=9.4e-06 Score=65.93 Aligned_cols=52 Identities=23% Similarity=0.427 Sum_probs=41.3
Q ss_pred cccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEec--CCCCc
Q 009784 328 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIA--NDGTV 390 (526)
Q Consensus 328 ~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~--~~~~l 390 (526)
..+|+.+....+ ...|++|.+|.++|||++ ||++||+|++|||+++. +..++
T Consensus 12 ~~~G~~~~~~~~-----------~~~~~~V~~v~~~s~a~~~gl~~GD~I~~ing~~i~~~~~~~~ 66 (82)
T cd00992 12 GGLGFSLRGGKD-----------SGGGIFVSRVEPGGPAERGGLRVGDRILEVNGVSVEGLTHEEA 66 (82)
T ss_pred CCcCEEEeCccc-----------CCCCeEEEEECCCChHHhCCCCCCCEEEEECCEEcCccCHHHH
Confidence 457777765411 135899999999999999 99999999999999998 44443
No 41
>PF00863 Peptidase_C4: Peptidase family C4 This family belongs to family C4 of the peptidase classification.; InterPro: IPR001730 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. Nuclear inclusion A (NIA) proteases from potyviruses are cysteine peptidases belong to the MEROPS peptidase family C4 (NIa protease family, clan PA(C)) [, ]. Potyviruses include plant viruses in which the single-stranded RNA encodes a polyprotein with NIA protease activity, where proteolytic cleavage is specific for Gln+Gly sites. The NIA protease acts on the polyprotein, releasing itself by Gln+Gly cleavage at both the N- and C-termini. It further processes the polyprotein by cleavage at five similar sites in the C-terminal half of the sequence. In addition to its C-terminal protease activity, the NIA protease contains an N-terminal domain that has been implicated in the transcription process []. This peptidase is present in the nuclear inclusion protein of potyviruses.; GO: 0008234 cysteine-type peptidase activity, 0006508 proteolysis; PDB: 3MMG_B 1Q31_B 1LVB_A 1LVM_A.
Probab=97.98 E-value=0.00061 Score=66.77 Aligned_cols=169 Identities=20% Similarity=0.279 Sum_probs=84.2
Q ss_pred cccCCCeEEEEeeeeCCCCCCccccCCCcceEEEEEEEe-CCEEEecccccCC-CCeEEEEEcCCCcEEEEE-----EEE
Q 009784 119 VPAMDAVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIG-GRRVLTNAHSVEH-YTQVKLKKRGSDTKYLAT-----VLA 191 (526)
Q Consensus 119 ~~~~~SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~-~g~ILT~aHvV~~-~~~i~V~~~~~g~~~~a~-----vv~ 191 (526)
.+....|+++...... ...+=+-|. ..+|+||+|.... ...++|... .| .|... -+.
T Consensus 14 n~Ia~~ic~l~n~s~~--------------~~~~l~gigyG~~iItn~HLf~~nng~L~i~s~-hG-~f~v~nt~~lkv~ 77 (235)
T PF00863_consen 14 NPIASNICRLTNESDG--------------GTRSLYGIGYGSYIITNAHLFKRNNGELTIKSQ-HG-EFTVPNTTQLKVH 77 (235)
T ss_dssp HHHHTTEEEEEEEETT--------------EEEEEEEEEETTEEEEEGGGGSSTTCEEEEEET-TE-EEEECEGGGSEEE
T ss_pred chhhheEEEEEEEeCC--------------CeEEEEEEeECCEEEEChhhhccCCCeEEEEeC-ce-EEEcCCccccceE
Confidence 3445578888764411 122223333 6699999999964 456777764 33 33321 133
Q ss_pred eccCCCeEEEEecccccccCceeeecC---CCCcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcc
Q 009784 192 IGTECDIAMLTVEDDEFWEGVLPVEFG---ELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDA 268 (526)
Q Consensus 192 ~d~~~DlAlLkv~~~~~~~~~~pl~l~---~~~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da 268 (526)
.=+..||.++|++.+ ++|.+-- ..+..++.|..+|.-+.....+ ..||.........+.. +...-.
T Consensus 78 ~i~~~DiviirmPkD-----fpPf~~kl~FR~P~~~e~v~mVg~~fq~k~~~---s~vSesS~i~p~~~~~---fWkHwI 146 (235)
T PF00863_consen 78 PIEGRDIVIIRMPKD-----FPPFPQKLKFRAPKEGERVCMVGSNFQEKSIS---STVSESSWIYPEENSH---FWKHWI 146 (235)
T ss_dssp E-TCSSEEEEE--TT-----S----S---B----TT-EEEEEEEECSSCCCE---EEEEEEEEEEEETTTT---EEEE-C
T ss_pred EeCCccEEEEeCCcc-----cCCcchhhhccCCCCCCEEEEEEEEEEcCCee---EEECCceEEeecCCCC---eeEEEe
Confidence 345799999999874 3554321 3567789999999866543321 2233322211111111 223333
Q ss_pred cCCCCCCCCeeecC-CCeEEEEEeeccccCccccccccccH--HHHHHHHHH
Q 009784 269 AINSGNSGGPAFND-KGKCVGIAFQSLKHEDVENIGYVIPT--PVIMHFIQD 317 (526)
Q Consensus 269 ~i~~G~SGGPlvn~-~G~VVGI~~~~~~~~~~~~~~~aIP~--~~i~~~l~~ 317 (526)
....|+=|+|+++. +|++|||++... .....+|+.|+ +.+..+++.
T Consensus 147 sTk~G~CG~PlVs~~Dg~IVGiHsl~~---~~~~~N~F~~f~~~f~~~~l~~ 195 (235)
T PF00863_consen 147 STKDGDCGLPLVSTKDGKIVGIHSLTS---NTSSRNYFTPFPDDFEEFYLEN 195 (235)
T ss_dssp ---TT-TT-EEEETTT--EEEEEEEEE---TTTSSEEEEE--TTHHHHHCC-
T ss_pred cCCCCccCCcEEEcCCCcEEEEEcCcc---CCCCeEEEEcCCHHHHHHHhcc
Confidence 44579999999986 899999999765 23445566554 444444433
No 42
>COG3480 SdrC Predicted secreted protein containing a PDZ domain [Signal transduction mechanisms]
Probab=97.94 E-value=1.4e-05 Score=79.99 Aligned_cols=72 Identities=25% Similarity=0.295 Sum_probs=65.7
Q ss_pred ccceEEEEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEE-CCEEEEEEEEe
Q 009784 352 QKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLR-DSKILNFNITL 430 (526)
Q Consensus 352 ~~Gv~V~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R-~G~~~~~~v~l 430 (526)
-.|+++..|..++|+...|+.||.|++|||+++.+.+++. ..+..+++||+|++++.| +++....++++
T Consensus 129 y~gvyv~~v~~~~~~~gkl~~gD~i~avdg~~f~s~~e~i----------~~v~~~k~Gd~VtI~~~r~~~~~~~~~~tl 198 (342)
T COG3480 129 YAGVYVLSVIDNSPFKGKLEAGDTIIAVDGEPFTSSDELI----------DYVSSKKPGDEVTIDYERHNETPEIVTITL 198 (342)
T ss_pred EeeEEEEEccCCcchhceeccCCeEEeeCCeecCCHHHHH----------HHHhccCCCCeEEEEEEeccCCCceEEEEE
Confidence 4699999999999998899999999999999999999965 888889999999999997 88888888888
Q ss_pred ccc
Q 009784 431 ATH 433 (526)
Q Consensus 431 ~~~ 433 (526)
...
T Consensus 199 ~~~ 201 (342)
T COG3480 199 IKN 201 (342)
T ss_pred Eee
Confidence 776
No 43
>COG0793 Prc Periplasmic protease [Cell envelope biogenesis, outer membrane]
Probab=97.91 E-value=3e-05 Score=82.52 Aligned_cols=84 Identities=24% Similarity=0.387 Sum_probs=63.8
Q ss_pred ccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCC--cccccccchhhhhh
Q 009784 327 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGT--VPFRHGERIGFSYL 403 (526)
Q Consensus 327 ~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~--l~~~~~~~~~~~~~ 403 (526)
+..+|++++.. +..++.|.++.+++||++ ||++||+|++|||+++....- . ..
T Consensus 99 ~~GiG~~i~~~-------------~~~~~~V~s~~~~~PA~kagi~~GD~I~~IdG~~~~~~~~~~a-----------v~ 154 (406)
T COG0793 99 FGGIGIELQME-------------DIGGVKVVSPIDGSPAAKAGIKPGDVIIKIDGKSVGGVSLDEA-----------VK 154 (406)
T ss_pred ccceeEEEEEe-------------cCCCcEEEecCCCChHHHcCCCCCCEEEEECCEEccCCCHHHH-----------HH
Confidence 55677777644 126789999999999999 999999999999999987641 1 12
Q ss_pred hhhcCCCCEEEEEEEECCEEEEEEEEecccc
Q 009784 404 VSQKYTGDSAAVKVLRDSKILNFNITLATHR 434 (526)
Q Consensus 404 l~~~~~G~~v~l~v~R~G~~~~~~v~l~~~~ 434 (526)
..+..+|.+|+|++.|.|....+++++.+..
T Consensus 155 ~irG~~Gt~V~L~i~r~~~~k~~~v~l~Re~ 185 (406)
T COG0793 155 LIRGKPGTKVTLTILRAGGGKPFTVTLTREE 185 (406)
T ss_pred HhCCCCCCeEEEEEEEcCCCceeEEEEEEEE
Confidence 3445689999999999865555666665543
No 44
>PF05579 Peptidase_S32: Equine arteritis virus serine endopeptidase S32; InterPro: IPR008760 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to MEROPS peptidase family S32 (clan PA(S)). The type example is equine arteritis virus serine endopeptidase (equine arteritis virus), which is involved in processing of nidovirus polyproteins [].; GO: 0004252 serine-type endopeptidase activity, 0016032 viral reproduction, 0019082 viral protein processing; PDB: 3FAN_A 3FAO_A 1MBM_A.
Probab=97.80 E-value=0.00024 Score=69.76 Aligned_cols=115 Identities=19% Similarity=0.226 Sum_probs=63.0
Q ss_pred eEEEEEEEe-C--CEEEecccccCCCCeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecccccccCceeeecCCCCcCCC
Q 009784 149 SSSSGFAIG-G--RRVLTNAHSVEHYTQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGELPALQD 225 (526)
Q Consensus 149 ~~GSGfvI~-~--g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~~~~g~ 225 (526)
..|||=++. + -.|+|+.||+. .+...|.. .+.... ..++..-|+|.-.++.-. ...|.++++... .|.
T Consensus 112 s~Gsggvft~~~~~vvvTAtHVlg-~~~a~v~~--~g~~~~---~tF~~~GDfA~~~~~~~~--G~~P~~k~a~~~-~Gr 182 (297)
T PF05579_consen 112 SVGSGGVFTIGGNTVVVTATHVLG-GNTARVSG--VGTRRM---LTFKKNGDFAEADITNWP--GAAPKYKFAQNY-TGR 182 (297)
T ss_dssp SEEEEEEEECTTEEEEEEEHHHCB-TTEEEEEE--TTEEEE---EEEEEETTEEEEEETTS---S---B--B-TT--SEE
T ss_pred cccccceEEECCeEEEEEEEEEcC-CCeEEEEe--cceEEE---EEEeccCcEEEEECCCCC--CCCCceeecCCc-ccc
Confidence 355655555 3 38999999998 55666665 343333 335567799999884321 256667776221 111
Q ss_pred cEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecCCCeEEEEEeecc
Q 009784 226 AVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSL 294 (526)
Q Consensus 226 ~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~ 294 (526)
.-+ ... .-+..|.|....+ ++ -..+|+||+|++..+|.+|||++++.
T Consensus 183 AyW---~t~----tGvE~G~ig~~~~------------~~---fT~~GDSGSPVVt~dg~liGVHTGSn 229 (297)
T PF05579_consen 183 AYW---LTS----TGVEPGFIGGGGA------------VC---FTGPGDSGSPVVTEDGDLIGVHTGSN 229 (297)
T ss_dssp EEE---EET----TEEEEEEEETTEE------------EE---SS-GGCTT-EEEETTC-EEEEEEEEE
T ss_pred eEE---Ecc----cCcccceecCceE------------EE---EcCCCCCCCccCcCCCCEEEEEecCC
Confidence 100 011 1245566554332 22 23589999999999999999999865
No 45
>PRK09681 putative type II secretion protein GspC; Provisional
Probab=97.77 E-value=3.9e-05 Score=76.82 Aligned_cols=62 Identities=26% Similarity=0.330 Sum_probs=51.7
Q ss_pred EeCCCCcc---cC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEe
Q 009784 359 RVDPTAPE---SE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITL 430 (526)
Q Consensus 359 ~V~~~spA---~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l 430 (526)
.+.|+..+ ++ |||+||++++|||..+++.++.. .++.......+++|+|+|||+..++.+.+
T Consensus 210 rl~Pgkd~~lF~~~GLq~GDva~sING~dL~D~~qa~----------~l~~~L~~~tei~ltVeRdGq~~~i~i~l 275 (276)
T PRK09681 210 AVKPGADRSLFDASGFKEGDIAIALNQQDFTDPRAMI----------ALMRQLPSMDSIQLTVLRKGARHDISIAL 275 (276)
T ss_pred EECCCCcHHHHHHcCCCCCCEEEEeCCeeCCCHHHHH----------HHHHHhccCCeEEEEEEECCEEEEEEEEc
Confidence 55677543 45 99999999999999999887753 66777778899999999999999998875
No 46
>KOG3129 consensus 26S proteasome regulatory complex, subunit PSMD9 [Posttranslational modification, protein turnover, chaperones]
Probab=97.73 E-value=6.7e-05 Score=70.91 Aligned_cols=73 Identities=22% Similarity=0.216 Sum_probs=61.2
Q ss_pred ceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEecc
Q 009784 354 GVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLAT 432 (526)
Q Consensus 354 Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~ 432 (526)
-++|.+|.|+|||+. ||+.||.|++++...-.++..|+ -...+..+..++.+.++|.|.|+.+.+.++++.
T Consensus 140 Fa~V~sV~~~SPA~~aGl~~gD~il~fGnV~sgn~~~lq--------~i~~~v~~~e~~~v~v~v~R~g~~v~L~ltP~~ 211 (231)
T KOG3129|consen 140 FAVVDSVVPGSPADEAGLCVGDEILKFGNVHSGNFLPLQ--------NIAAVVQSNEDQIVSVTVIREGQKVVLSLTPKK 211 (231)
T ss_pred eEEEeecCCCChhhhhCcccCceEEEecccccccchhHH--------HHHHHHHhccCcceeEEEecCCCEEEEEeCccc
Confidence 578999999999999 99999999999988888777653 113455567899999999999999999998876
Q ss_pred cc
Q 009784 433 HR 434 (526)
Q Consensus 433 ~~ 434 (526)
+.
T Consensus 212 W~ 213 (231)
T KOG3129|consen 212 WQ 213 (231)
T ss_pred cc
Confidence 53
No 47
>PF04495 GRASP55_65: GRASP55/65 PDZ-like domain ; InterPro: IPR007583 GRASP55 (Golgi reassembly stacking protein of 55 kDa) and GRASP65 (a 65 kDa) protein are highly homologous. GRASP55 is a component of the Golgi stacking machinery. GRASP65, an N-ethylmaleimide-sensitive membrane protein required for the stacking of Golgi cisternae in a cell-free system [].; PDB: 3RLE_A 4EDJ_A.
Probab=97.47 E-value=0.00034 Score=63.32 Aligned_cols=87 Identities=23% Similarity=0.325 Sum_probs=55.2
Q ss_pred ccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCC-CcEEEEECCEEecCCCCcccccccchhhhhhh
Q 009784 327 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKP-SDIILSFDGIDIANDGTVPFRHGERIGFSYLV 404 (526)
Q Consensus 327 ~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~-GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l 404 (526)
.+.||+.++--. ..- ....+.-|.+|.|+|||++ ||++ .|.|+.+|+..+.+.++|. ..+
T Consensus 25 ~g~LG~sv~~~~-~~~-------~~~~~~~Vl~V~p~SPA~~AGL~p~~DyIig~~~~~l~~~~~l~----------~~v 86 (138)
T PF04495_consen 25 QGLLGISVRFES-FEG-------AEEEGWHVLRVAPNSPAAKAGLEPFFDYIIGIDGGLLDDEDDLF----------ELV 86 (138)
T ss_dssp SSSS-EEEEEEE--TT-------GCCCEEEEEEE-TTSHHHHTT--TTTEEEEEETTCE--STCHHH----------HHH
T ss_pred CCCCcEEEEEec-ccc-------cccceEEEeEecCCCHHHHCCccccccEEEEccceecCCHHHHH----------HHH
Confidence 466787776441 110 1356899999999999998 9999 6999999998888766542 555
Q ss_pred hhcCCCCEEEEEEEEC--CEEEEEEEEecc
Q 009784 405 SQKYTGDSAAVKVLRD--SKILNFNITLAT 432 (526)
Q Consensus 405 ~~~~~G~~v~l~v~R~--G~~~~~~v~l~~ 432 (526)
. .+.++++.|.|+.. .+.+++++++..
T Consensus 87 ~-~~~~~~l~L~Vyns~~~~vR~V~i~P~~ 115 (138)
T PF04495_consen 87 E-ANENKPLQLYVYNSKTDSVREVTITPSR 115 (138)
T ss_dssp H-HTTTS-EEEEEEETTTTCEEEEEE---T
T ss_pred H-HcCCCcEEEEEEECCCCeEEEEEEEcCC
Confidence 4 45789999999984 455666666553
No 48
>PF00548 Peptidase_C3: 3C cysteine protease (picornain 3C); InterPro: IPR000199 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This signature defines cysteine peptidases belong to MEROPS peptidase family C3 (picornain, clan PA(C)), subfamilies C3A and C3B. The protein fold of this peptidase domain for members of this family resembles that of the serine peptidase, chymotrypsin [], the type example for clan PA. Picornaviral proteins are expressed as a single polyprotein which is cleaved by the viral C3 cysteine protease. The poliovirus polyprotein is selectively cleaved between the Gln-|-Gly bond. In other picornavirus reactions Glu may be substituted for Gln, and Ser or Thr for Gly. ; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis; PDB: 3SJO_E 2H6M_A 1QA7_C 1HAV_B 2HAL_A 2H9H_A 3QZQ_B 3QZR_A 3R0F_B 3SJ9_A ....
Probab=97.39 E-value=0.0047 Score=58.15 Aligned_cols=138 Identities=17% Similarity=0.269 Sum_probs=83.7
Q ss_pred cceEEEEEEEeCCEEEecccccCCCCeEEEEEcCCCcEEEE--EEEEecc---CCCeEEEEecccccccCc-eeeecCCC
Q 009784 147 YSSSSSGFAIGGRRVLTNAHSVEHYTQVKLKKRGSDTKYLA--TVLAIGT---ECDIAMLTVEDDEFWEGV-LPVEFGEL 220 (526)
Q Consensus 147 ~~~~GSGfvI~~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a--~vv~~d~---~~DlAlLkv~~~~~~~~~-~pl~l~~~ 220 (526)
....++|+.|.+.++|.+.| .....++.+ +|..++. .+...+. ..||++++++...-+.++ +.+. +.
T Consensus 23 g~~t~l~~gi~~~~~lvp~H---~~~~~~i~i--~g~~~~~~d~~~lv~~~~~~~Dl~~v~l~~~~kfrDIrk~~~--~~ 95 (172)
T PF00548_consen 23 GEFTMLALGIYDRYFLVPTH---EEPEDTIYI--DGVEYKVDDSVVLVDRDGVDTDLTLVKLPRNPKFRDIRKFFP--ES 95 (172)
T ss_dssp EEEEEEEEEEEBTEEEEEGG---GGGCSEEEE--TTEEEEEEEEEEEEETTSSEEEEEEEEEESSS-B--GGGGSB--SS
T ss_pred ceEEEecceEeeeEEEEECc---CCCcEEEEE--CCEEEEeeeeEEEecCCCcceeEEEEEccCCcccCchhhhhc--cc
Confidence 44678999999999999999 223334444 4555542 3333444 469999999765422222 2222 22
Q ss_pred C-cCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecC---CCeEEEEEeec
Q 009784 221 P-ALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFND---KGKCVGIAFQS 293 (526)
Q Consensus 221 ~-~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~---~G~VVGI~~~~ 293 (526)
. ...+.+.++-.+ ......+..+.++..+.. ...+......+..+++..+|+-||||+.. .++++||+.++
T Consensus 96 ~~~~~~~~l~v~~~-~~~~~~~~v~~v~~~~~i-~~~g~~~~~~~~Y~~~t~~G~CG~~l~~~~~~~~~i~GiHvaG 170 (172)
T PF00548_consen 96 IPEYPECVLLVNST-KFPRMIVEVGFVTNFGFI-NLSGTTTPRSLKYKAPTKPGMCGSPLVSRIGGQGKIIGIHVAG 170 (172)
T ss_dssp GGTEEEEEEEEESS-SSTCEEEEEEEEEEEEEE-EETTEEEEEEEEEESEEETTGTTEEEEESCGGTTEEEEEEEEE
T ss_pred cccCCCcEEEEECC-CCccEEEEEEEEeecCcc-ccCCCEeeEEEEEccCCCCCccCCeEEEeeccCccEEEEEecc
Confidence 2 233444444333 333334555556555443 23344444578888999999999999863 58999999885
No 49
>PRK11186 carboxy-terminal protease; Provisional
Probab=97.39 E-value=0.00068 Score=76.21 Aligned_cols=71 Identities=15% Similarity=0.135 Sum_probs=49.1
Q ss_pred cceEEEEeCCCCcccC--CCCCCcEEEEEC--CEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEEC---CEEEE
Q 009784 353 KGVRIRRVDPTAPESE--VLKPSDIILSFD--GIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD---SKILN 425 (526)
Q Consensus 353 ~Gv~V~~V~~~spA~~--GL~~GDvIl~vn--G~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~---G~~~~ 425 (526)
.+++|.+|.+||||++ ||++||+|++|| |+++.+...+. +.-...+.....|.+|.|+|.|+ |+..+
T Consensus 255 ~~~~V~~vipGsPA~ka~gLk~GD~IlaVn~~g~~~~dv~g~~------~~~vv~lirG~~Gt~V~LtV~r~~~~~~~~~ 328 (667)
T PRK11186 255 DYTVINSLVAGGPAAKSKKLSVGDKIVGVGQDGKPIVDVIGWR------LDDVVALIKGPKGSKVRLEILPAGKGTKTRI 328 (667)
T ss_pred CeEEEEEccCCChHHHhCCCCCCCEEEEECCCCCcccccccCC------HHHHHHHhcCCCCCEEEEEEEeCCCCCceEE
Confidence 4688999999999997 899999999999 55554432221 11112233455799999999994 45555
Q ss_pred EEEE
Q 009784 426 FNIT 429 (526)
Q Consensus 426 ~~v~ 429 (526)
++++
T Consensus 329 vtl~ 332 (667)
T PRK11186 329 VTLT 332 (667)
T ss_pred EEEE
Confidence 5543
No 50
>COG3975 Predicted protease with the C-terminal PDZ domain [General function prediction only]
Probab=97.32 E-value=0.00056 Score=73.09 Aligned_cols=85 Identities=21% Similarity=0.363 Sum_probs=65.3
Q ss_pred cCceeeeccChhHHhhhcCCC--CccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhh
Q 009784 330 LGVEWQKMENPDLRVAMSMKA--DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ 406 (526)
Q Consensus 330 LGi~~~~~~~~~~~~~lgl~~--~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~ 406 (526)
.|+.+..... ..-++|+.- +..+.+|..|.++|||++ ||.+||.|++|||. + ..+.+
T Consensus 439 ~gL~~~~~~~--~~~~LGl~v~~~~g~~~i~~V~~~gPA~~AGl~~Gd~ivai~G~------s------------~~l~~ 498 (558)
T COG3975 439 FGLTFTPKPR--EAYYLGLKVKSEGGHEKITFVFPGGPAYKAGLSPGDKIVAINGI------S------------DQLDR 498 (558)
T ss_pred cceEEEecCC--CCcccceEecccCCeeEEEecCCCChhHhccCCCccEEEEEcCc------c------------ccccc
Confidence 3555555521 134566543 345688999999999999 99999999999999 1 23556
Q ss_pred cCCCCEEEEEEEECCEEEEEEEEecccc
Q 009784 407 KYTGDSAAVKVLRDSKILNFNITLATHR 434 (526)
Q Consensus 407 ~~~G~~v~l~v~R~G~~~~~~v~l~~~~ 434 (526)
.+.++.|++.+.|.|+.+++.+++....
T Consensus 499 ~~~~d~i~v~~~~~~~L~e~~v~~~~~~ 526 (558)
T COG3975 499 YKVNDKIQVHVFREGRLREFLVKLGGDP 526 (558)
T ss_pred cccccceEEEEccCCceEEeecccCCCc
Confidence 6789999999999999999988877654
No 51
>COG3031 PulC Type II secretory pathway, component PulC [Intracellular trafficking and secretion]
Probab=96.96 E-value=0.00067 Score=65.53 Aligned_cols=66 Identities=18% Similarity=0.223 Sum_probs=52.1
Q ss_pred ceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEE
Q 009784 354 GVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT 429 (526)
Q Consensus 354 Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~ 429 (526)
|..+.-..+++-.++ |||.||+.+++|+..+++.+++. .++.....-+.++++|+|+|+..++.|.
T Consensus 208 Gyr~~pgkd~slF~~sglq~GDIavaiNnldltdp~~m~----------~llq~l~~m~s~qlTv~R~G~rhdInV~ 274 (275)
T COG3031 208 GYRFEPGKDGSLFYKSGLQRGDIAVAINNLDLTDPEDMF----------RLLQMLRNMPSLQLTVIRRGKRHDINVR 274 (275)
T ss_pred EEEecCCCCcchhhhhcCCCcceEEEecCcccCCHHHHH----------HHHHhhhcCcceEEEEEecCccceeeec
Confidence 444444445566677 99999999999999999988853 5666666667899999999999988874
No 52
>PF03761 DUF316: Domain of unknown function (DUF316) ; InterPro: IPR005514 This is a family of uncharacterised proteins from Caenorhabditis elegans.
Probab=96.78 E-value=0.039 Score=55.79 Aligned_cols=107 Identities=17% Similarity=0.241 Sum_probs=62.6
Q ss_pred CCCeEEEEecccccccCceeeecCCCCc---CCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCC
Q 009784 195 ECDIAMLTVEDDEFWEGVLPVEFGELPA---LQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAIN 271 (526)
Q Consensus 195 ~~DlAlLkv~~~~~~~~~~pl~l~~~~~---~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~ 271 (526)
..+++||+++.+ +.....|+.|+++.. .++.+.+.|+... .. +....+.-..... ....+......+
T Consensus 160 ~~~~mIlEl~~~-~~~~~~~~Cl~~~~~~~~~~~~~~~yg~~~~-~~--~~~~~~~i~~~~~------~~~~~~~~~~~~ 229 (282)
T PF03761_consen 160 PYSPMILELEED-FSKNVSPPCLADSSTNWEKGDEVDVYGFNST-GK--LKHRKLKITNCTK------CAYSICTKQYSC 229 (282)
T ss_pred ccceEEEEEccc-ccccCCCEEeCCCccccccCceEEEeecCCC-Ce--EEEEEEEEEEeec------cceeEecccccC
Confidence 479999999988 334788889987543 4688889998222 11 2222222111100 112355566778
Q ss_pred CCCCCCeeecC-CC--eEEEEEeeccccCccccccccccHHHHH
Q 009784 272 SGNSGGPAFND-KG--KCVGIAFQSLKHEDVENIGYVIPTPVIM 312 (526)
Q Consensus 272 ~G~SGGPlvn~-~G--~VVGI~~~~~~~~~~~~~~~aIP~~~i~ 312 (526)
.|++|||++.. +| .||||.+..... ...+..+++.+...+
T Consensus 230 ~~d~Gg~lv~~~~gr~tlIGv~~~~~~~-~~~~~~~f~~v~~~~ 272 (282)
T PF03761_consen 230 KGDRGGPLVKNINGRWTLIGVGASGNYE-CNKNNSYFFNVSWYQ 272 (282)
T ss_pred CCCccCeEEEEECCCEEEEEEEccCCCc-ccccccEEEEHHHhh
Confidence 99999999832 44 699998764321 111244555554443
No 53
>KOG3553 consensus Tax interaction protein TIP1 [Cell wall/membrane/envelope biogenesis]
Probab=96.66 E-value=0.0018 Score=54.28 Aligned_cols=35 Identities=31% Similarity=0.505 Sum_probs=32.3
Q ss_pred CccceEEEEeCCCCcccC-CCCCCcEEEEECCEEec
Q 009784 351 DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIA 385 (526)
Q Consensus 351 ~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~ 385 (526)
.+.|++|++|..+|||+. ||+.+|.|+.+||...+
T Consensus 57 tD~GiYvT~V~eGsPA~~AGLrihDKIlQvNG~DfT 92 (124)
T KOG3553|consen 57 TDKGIYVTRVSEGSPAEIAGLRIHDKILQVNGWDFT 92 (124)
T ss_pred CCccEEEEEeccCChhhhhcceecceEEEecCceeE
Confidence 368999999999999999 99999999999998765
No 54
>PF08192 Peptidase_S64: Peptidase family S64; InterPro: IPR012985 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This family of fungal proteins is involved in the processing of membrane bound transcription factor Stp1 [] and belongs to MEROPS petidase family S64 (clan PA). The processing causes the signalling domain of Stp1 to be passed to the nucleus where several permease genes are induced. The permeases are important for uptake of amino acids, and processing of tp1 only occurs in an amino acid-rich environment. This family is predicted to be distantly related to the trypsin family (MEROPS peptidase family S1) and to have a typical trypsin-like catalytic triad [].
Probab=96.49 E-value=0.03 Score=61.81 Aligned_cols=117 Identities=19% Similarity=0.288 Sum_probs=74.2
Q ss_pred CCCeEEEEecccc-----cccCc------eeeecCCC--------CcCCCcEEEEeeCCCCCceeEEEEEEeceeeeecc
Q 009784 195 ECDIAMLTVEDDE-----FWEGV------LPVEFGEL--------PALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYV 255 (526)
Q Consensus 195 ~~DlAlLkv~~~~-----~~~~~------~pl~l~~~--------~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~ 255 (526)
-.|+||++++... +.+++ +.+.+.+. ...|.+|+=+|...+ .+.|.|.++....+.
T Consensus 542 LsD~AIIkV~~~~~~~N~LGddi~f~~~dP~l~f~NlyV~~~~~~~~~G~~VfK~GrTTg-----yT~G~lNg~klvyw~ 616 (695)
T PF08192_consen 542 LSDWAIIKVNKERKCQNYLGDDIQFNEPDPTLMFQNLYVREVVSNLVPGMEVFKVGRTTG-----YTTGILNGIKLVYWA 616 (695)
T ss_pred ccceEEEEeCCCceecCCCCccccccCCCccccccccchhhhhhccCCCCeEEEecccCC-----ccceEecceEEEEec
Confidence 3699999998653 12222 23344331 123678999998766 456888877655455
Q ss_pred CCcee-eeEEEEc----ccCCCCCCCCeeecCCCe------EEEEEeeccccCccccccccccHHHHHHHHHHH
Q 009784 256 HGSTE-LLGLQID----AAINSGNSGGPAFNDKGK------CVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDY 318 (526)
Q Consensus 256 ~~~~~-~~~i~~d----a~i~~G~SGGPlvn~~G~------VVGI~~~~~~~~~~~~~~~aIP~~~i~~~l~~l 318 (526)
+|... .+++... .-...|+||+-|++.-+. |+||..+.- .....+|++.|+..|.+-|++.
T Consensus 617 dG~i~s~efvV~s~~~~~Fa~~GDSGS~VLtk~~d~~~gLgvvGMlhsyd--ge~kqfglftPi~~il~rl~~v 688 (695)
T PF08192_consen 617 DGKIQSSEFVVSSDNNPAFASGGDSGSWVLTKLEDNNKGLGVVGMLHSYD--GEQKQFGLFTPINEILDRLEEV 688 (695)
T ss_pred CCCeEEEEEEEecCCCccccCCCCcccEEEecccccccCceeeEEeeecC--CccceeeccCcHHHHHHHHHHh
Confidence 54422 2233333 233579999999986444 999998743 2345789999998887766554
No 55
>COG5640 Secreted trypsin-like serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=96.37 E-value=0.062 Score=55.28 Aligned_cols=59 Identities=22% Similarity=0.304 Sum_probs=36.8
Q ss_pred eEEEEEEEeCCEEEecccccCCCC-----eEEE--EEcC--CCcEEEEEEEEec-------cCCCeEEEEecccc
Q 009784 149 SSSSGFAIGGRRVLTNAHSVEHYT-----QVKL--KKRG--SDTKYLATVLAIG-------TECDIAMLTVEDDE 207 (526)
Q Consensus 149 ~~GSGfvI~~g~ILT~aHvV~~~~-----~i~V--~~~~--~g~~~~a~vv~~d-------~~~DlAlLkv~~~~ 207 (526)
..|-|-++...||||+|||+.+.. .+.| .+.+ .+.....+.++.+ ...|+|++++....
T Consensus 61 tfCGgs~l~~RYvLTAAHC~~~~s~is~d~~~vv~~l~d~Sq~~rg~vr~i~~~efY~~~n~~ND~Av~~l~~~a 135 (413)
T COG5640 61 TFCGGSKLGGRYVLTAAHCADASSPISSDVNRVVVDLNDSSQAERGHVRTIYVHEFYSPGNLGNDIAVLELARAA 135 (413)
T ss_pred eEeccceecceEEeeehhhccCCCCccccceEEEecccccccccCcceEEEeeecccccccccCcceeecccccc
Confidence 357788888779999999998654 1222 2221 1222334444433 35799999998754
No 56
>PF12812 PDZ_1: PDZ-like domain
Probab=96.31 E-value=0.0044 Score=50.46 Aligned_cols=60 Identities=12% Similarity=0.047 Sum_probs=50.3
Q ss_pred cccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcc
Q 009784 328 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVP 391 (526)
Q Consensus 328 ~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~ 391 (526)
-+.|..++++ +-+.++.++++ -|+++.....++++.. |+..|-+|++|||+++.+.+++.
T Consensus 9 ~~~Ga~f~~L-s~q~aR~~~~~---~~gv~v~~~~g~~~~~~~i~~g~iI~~Vn~kpt~~Ld~f~ 69 (78)
T PF12812_consen 9 EVCGAVFHDL-SYQQARQYGIP---VGGVYVAVSGGSLAFAGGISKGFIITSVNGKPTPDLDDFI 69 (78)
T ss_pred EEcCeecccC-CHHHHHHhCCC---CCEEEEEecCCChhhhCCCCCCeEEEeECCcCCcCHHHHH
Confidence 3789999998 88899999988 3355556788899888 69999999999999999988753
No 57
>PF10459 Peptidase_S46: Peptidase S46; InterPro: IPR019500 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This entry represents S46 peptidases, where dipeptidyl-peptidase 7 (DPP-7) is the best-characterised member of this family. It is a serine peptidase that is located on the cell surface and is predicted to have two N-terminal transmembrane domains.
Probab=96.31 E-value=0.0048 Score=69.80 Aligned_cols=55 Identities=25% Similarity=0.349 Sum_probs=31.2
Q ss_pred EEEcccCCCCCCCCeeecCCCeEEEEEeeccccCc--------cccccccccHHHHHHHHHHH
Q 009784 264 LQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHED--------VENIGYVIPTPVIMHFIQDY 318 (526)
Q Consensus 264 i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~--------~~~~~~aIP~~~i~~~l~~l 318 (526)
+.++..|..||||||++|.+|++||+++-+.-++- ..+.+..+.+..|+.+|+++
T Consensus 624 FlstnDitGGNSGSPvlN~~GeLVGl~FDgn~Esl~~D~~fdp~~~R~I~VDiRyvL~~ldkv 686 (698)
T PF10459_consen 624 FLSTNDITGGNSGSPVLNAKGELVGLAFDGNWESLSGDIAFDPELNRTIHVDIRYVLWALDKV 686 (698)
T ss_pred EEeccCcCCCCCCCccCCCCceEEEEeecCchhhcccccccccccceeEEEEHHHHHHHHHHH
Confidence 44556667777777777777777777764432111 11223345555566666554
No 58
>PF02122 Peptidase_S39: Peptidase S39; InterPro: IPR000382 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. ORF2 of Potato leafroll virus (PLrV) encodes a polyprotein which is translated following a -1 frameshift. The polyprotein has a putative linear arrangement of membrane achor-VPg-peptidase-polmerase domains. The serine peptidase domain which is found in this group of sequences belongs to MEROPS peptidase family S39 (clan PA(S)). It is likely that the peptidase domain is involved in the cleavage of the polyprotein []. The nucleotide sequence for the RNA of PLrV has been determined [, ]. The sequence contains six large open reading frames (ORFs). The 5' coding region encodes two polypeptides of 28K and 70K, which overlap in different reading frames; it is suggested that the third ORF in the 5' block is translated by frameshift readthrough near the end of the 70K protein, yielding a 118K polypeptide []. Segments of the predicted amino acid sequences of these ORFs resemble those of known viral RNA polymerases, ATP-binding proteins and viral genome-linked proteins. The nucleotide sequence of the genomic RNA of Beet western yellows virus (BWYV) has been determined []. The sequence contains six long ORFs. A cluster of three of these ORFs, including the coat protein cistron, display extensive amino acid sequence similarity to corresponding ORFs of a second luteovirus: Barley yellow dwarf virus [].; GO: 0004252 serine-type endopeptidase activity, 0022415 viral reproductive process, 0016021 integral to membrane; PDB: 1ZYO_A.
Probab=96.19 E-value=0.0035 Score=60.38 Aligned_cols=137 Identities=22% Similarity=0.223 Sum_probs=49.3
Q ss_pred CEEEecccccCCCCeEEEEEcCCCcEEE---EEEEEeccCCCeEEEEeccccc-ccCceeeecCCCCcCC-CcEEEEeeC
Q 009784 159 RRVLTNAHSVEHYTQVKLKKRGSDTKYL---ATVLAIGTECDIAMLTVEDDEF-WEGVLPVEFGELPALQ-DAVTVVGYP 233 (526)
Q Consensus 159 g~ILT~aHvV~~~~~i~V~~~~~g~~~~---a~vv~~d~~~DlAlLkv~~~~~-~~~~~pl~l~~~~~~g-~~V~~iG~p 233 (526)
..++|++||......+.... +|+.++ -+.+..+...|++||+..+.-. .-.++.+.+.....+. ..+.+.++.
T Consensus 42 ~~L~ta~Hv~~~~~~~~~~k--~g~kipl~~f~~~~~~~~~D~~il~~P~n~~s~Lg~k~~~~~~~~~~~~g~~~~y~~~ 119 (203)
T PF02122_consen 42 DALLTARHVWSRPSKVTSLK--TGEKIPLAEFTDLLESRIADFVILRGPPNWESKLGVKAAQLSQNSQLAKGPVSFYGFS 119 (203)
T ss_dssp EEEEE-HHHHTSSS---EEE--TTEEEE--S-EEEEE-TTT-EEEEE--HHHHHHHT-----B----SEEEEESSTTSEE
T ss_pred cceecccccCCCccceeEcC--CCCcccchhChhhhCCCccCEEEEecCcCHHHHhCcccccccchhhhCCCCeeeeeec
Confidence 49999999999866665554 455443 3556678899999999984310 1134444443222110 000011110
Q ss_pred CCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecCCCeEEEEEeeccccCccccccccccHHH
Q 009784 234 IGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPV 310 (526)
Q Consensus 234 ~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~~~~~~~aIP~~~ 310 (526)
.+ ....+..-|. ...+ .+...-+...+|.||.|+++.+ +++|++.+..+....++.++..|+.-
T Consensus 120 ~~--~~~~~sa~i~------g~~~----~~~~vls~T~~G~SGtp~y~g~-~vvGvH~G~~~~~~~~n~n~~spip~ 183 (203)
T PF02122_consen 120 SG--EWPCSSAKIP------GTEG----KFASVLSNTSPGWSGTPYYSGK-NVVGVHTGSPSGSNRENNNRMSPIPP 183 (203)
T ss_dssp EE--EEEEEE-S----------ST----TEEEE-----TT-TT-EEE-SS--EEEEEEEE-----------------
T ss_pred CC--CceeccCccc------cccC----cCCceEcCCCCCCCCCCeEECC-CceEeecCcccccccccccccccccc
Confidence 00 0111111111 1111 1345556778999999999987 99999998643345566676655433
No 59
>KOG3209 consensus WW domain-containing protein [General function prediction only]
Probab=95.85 E-value=0.031 Score=61.68 Aligned_cols=140 Identities=16% Similarity=0.091 Sum_probs=85.9
Q ss_pred CccceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEE
Q 009784 351 DQKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNI 428 (526)
Q Consensus 351 ~~~Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v 428 (526)
..+-++|..|.+.+.|++ -|++||-|+.|||.+|.....-. -+ .++.......-|.|+|.|.-..-.
T Consensus 672 p~qpi~iG~Iv~lGaAe~DGRL~~gDElv~iDG~pV~GksH~~-----vv---~Lm~~AArnghV~LtVRRkv~~~~--- 740 (984)
T KOG3209|consen 672 PGQPIYIGAIVPLGAAEEDGRLREGDELVCIDGIPVEGKSHSE-----VV---DLMEAAARNGHVNLTVRRKVRTGP--- 740 (984)
T ss_pred CCCeeEEeeeeecccccccCcccCCCeEEEecCeeccCccHHH-----HH---HHHHHHHhcCceEEEEeeeeeecc---
Confidence 456699999999999998 49999999999999998765421 11 333333334569999988421110
Q ss_pred Eecccc--cccCCCCCCCCCceEEEeeEEEEecccc-ceeeeeeeecchh--hh-ccccccceeeeecccchhhhHHHHH
Q 009784 429 TLATHR--RLIPSHNKGRPPSYYIIAGFVFSRCLYL-ISVLSMERIMNMK--LR-SSFWTSSCIQCHNCQMSSLLWCLRC 502 (526)
Q Consensus 429 ~l~~~~--~~~p~~~~~~~p~~~i~gG~~f~~lt~~-~~~~~~~~i~~~~--~~-sg~~~~~~~~~~~~~~~~~~~~~~~ 502 (526)
-...+ ...+...++...+.....||-|.-++.+ .-.-.++||++.+ -| .-|++||.|..+|.+-.-.+.+--.
T Consensus 741 -~~rsp~~s~~~~~~yDV~lhR~ENeGFGFVi~sS~~kp~sgiGrIieGSPAdRCgkLkVGDrilAVNG~sI~~lsHadi 819 (984)
T KOG3209|consen 741 -ARRSPRNSAAPSGPYDVVLHRKENEGFGFVIMSSQNKPESGIGRIIEGSPADRCGKLKVGDRILAVNGQSILNLSHADI 819 (984)
T ss_pred -ccCCcccccCCCCCeeeEEecccCCceeEEEEecccCCCCCccccccCChhHhhccccccceEEEecCeeeeccCchhH
Confidence 01111 0111112222222223467777777633 2222388998854 23 3589999999999987666555443
No 60
>KOG3209 consensus WW domain-containing protein [General function prediction only]
Probab=95.83 E-value=0.0092 Score=65.59 Aligned_cols=152 Identities=20% Similarity=0.119 Sum_probs=88.8
Q ss_pred EEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEE-------
Q 009784 357 IRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFN------- 427 (526)
Q Consensus 357 V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~------- 427 (526)
|.+|.++|||+. .|+.||.|++|||+.|.+...-. ...++. ..|-+|+|+|.-..+.-..+
T Consensus 782 iGrIieGSPAdRCgkLkVGDrilAVNG~sI~~lsHad--------iv~LIK--daGlsVtLtIip~ee~~~~~~~~sa~~ 851 (984)
T KOG3209|consen 782 IGRIIEGSPADRCGKLKVGDRILAVNGQSILNLSHAD--------IVSLIK--DAGLSVTLTIIPPEEAGPPTSMTSAEK 851 (984)
T ss_pred ccccccCChhHhhccccccceEEEecCeeeeccCchh--------HHHHHH--hcCceEEEEEcChhccCCCCCCcchhh
Confidence 778999999999 69999999999999999876531 113333 36899999997644321111
Q ss_pred ---EEec----ccccccCC----CCCCCCC----------ceEEEeeEEEEecc--------------ccceeeeeeeec
Q 009784 428 ---ITLA----THRRLIPS----HNKGRPP----------SYYIIAGFVFSRCL--------------YLISVLSMERIM 472 (526)
Q Consensus 428 ---v~l~----~~~~~~p~----~~~~~~p----------~~~i~gG~~f~~lt--------------~~~~~~~~~~i~ 472 (526)
++.. ..-.+.+. .....+| +.-..+++.-.+|. ....-|.+-|+.
T Consensus 852 ~s~~t~~~~~~q~~glp~~~~s~~~~~pqpdt~~~~~~~~r~~qn~~~~~VelErG~kGFGFSiRGGreynM~LfVLRlA 931 (984)
T KOG3209|consen 852 QSPFTQNGPYEQQYGLPGPRPSVYEEHPQPDTFQGLSINDRMSQNGDLYTVELERGAKGFGFSIRGGREYNMDLFVLRLA 931 (984)
T ss_pred cCcccccCCHhHccCCCCCCccccccCCCCccccceeccccccccCCeeEEEeeccccccceEeecccccccceEEEEec
Confidence 1100 00000000 0001111 11111222222222 112233455554
Q ss_pred c--hhhhcc-ccccceeeeecccchhhhHHHHHHHHHHHHHhhhHHHHH
Q 009784 473 N--MKLRSS-FWTSSCIQCHNCQMSSLLWCLRCLWLILILDMRRLLTLR 518 (526)
Q Consensus 473 ~--~~~~sg-~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 518 (526)
+ .++|.| +++||-|...|.+--...-+-|.+=||----||-||.||
T Consensus 932 eDGPA~rdGrm~VGDqi~eINGesTkgmtH~rAIelIk~gg~~vll~Lr 980 (984)
T KOG3209|consen 932 EDGPAIRDGRMRVGDQITEINGESTKGMTHDRAIELIKQGGRRVLLLLR 980 (984)
T ss_pred cCCCccccCceeecceEEEecCcccCCCcHHHHHHHHHhCCeEEEEEec
Confidence 4 555665 567999999999988888888988888766666555444
No 61
>KOG3580 consensus Tight junction proteins [Signal transduction mechanisms]
Probab=95.51 E-value=0.011 Score=63.89 Aligned_cols=61 Identities=18% Similarity=0.262 Sum_probs=46.8
Q ss_pred CccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEE
Q 009784 351 DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLR 419 (526)
Q Consensus 351 ~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R 419 (526)
++-|+.|..|..++||++ |||.||.||+||.++..+.--= + ...++....+|+.++|.-.+
T Consensus 427 NDVGIFVaGvqegspA~~eGlqEGDQIL~VN~vdF~nl~RE-----e---AVlfLL~lPkGEevtilaQ~ 488 (1027)
T KOG3580|consen 427 NDVGIFVAGVQEGSPAEQEGLQEGDQILKVNTVDFRNLVRE-----E---AVLFLLELPKGEEVTILAQS 488 (1027)
T ss_pred CceeEEEeecccCCchhhccccccceeEEeccccchhhhHH-----H---HHHHHhcCCCCcEEeehhhh
Confidence 467999999999999999 9999999999999987764210 0 11345566789988886544
No 62
>KOG3580 consensus Tight junction proteins [Signal transduction mechanisms]
Probab=94.72 E-value=0.039 Score=59.75 Aligned_cols=74 Identities=23% Similarity=0.279 Sum_probs=52.4
Q ss_pred hhcCCCCccceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCE
Q 009784 345 AMSMKADQKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK 422 (526)
Q Consensus 345 ~lgl~~~~~Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~ 422 (526)
.||+. -.+-+.|.++...+-|++ +||.||+||+|||....|..-- + . ..+..+..| ++.|.|+||.+
T Consensus 212 EyGlr-LgSqIFvKeit~~gLAardgnlqEGDiiLkINGtvteNmSLt---D-----a-r~LIEkS~G-KL~lvVlRD~~ 280 (1027)
T KOG3580|consen 212 EYGLR-LGSQIFVKEITRTGLAARDGNLQEGDIILKINGTVTENMSLT---D-----A-RKLIEKSRG-KLQLVVLRDSQ 280 (1027)
T ss_pred hhccc-ccchhhhhhhcccchhhccCCcccccEEEEECcEeeccccch---h-----H-HHHHHhccC-ceEEEEEecCC
Confidence 45554 234578899988887776 8999999999999988775421 1 1 233444445 69999999987
Q ss_pred EEEEEEE
Q 009784 423 ILNFNIT 429 (526)
Q Consensus 423 ~~~~~v~ 429 (526)
..-++|.
T Consensus 281 qtLiNiP 287 (1027)
T KOG3580|consen 281 QTLINIP 287 (1027)
T ss_pred ceeeecC
Confidence 7666665
No 63
>PF00949 Peptidase_S7: Peptidase S7, Flavivirus NS3 serine protease ; InterPro: IPR001850 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This signature identifies serine peptidases belong to MEROPS peptidase family S7 (flavivirin family, clan PA(S)). The protein fold of the peptidase domain for members of this family resembles that of chymotrypsin, the type example for clan PA. Flaviviruses produce a polyprotein from the ssRNA genome. The N terminus of the NS3 protein (approx. 180 aa) is required for the processing of the polyprotein. NS3 also has conserved homology with NTP-binding proteins and DEAD family of RNA helicase [, , ].; GO: 0003723 RNA binding, 0003724 RNA helicase activity, 0005524 ATP binding; PDB: 2IJO_B 3E90_D 2GGV_B 2FP7_B 2WV9_A 3U1I_B 3U1J_B 2WZQ_A 2WHX_A 3L6P_A ....
Probab=94.61 E-value=0.043 Score=49.20 Aligned_cols=29 Identities=38% Similarity=0.672 Sum_probs=21.6
Q ss_pred EcccCCCCCCCCeeecCCCeEEEEEeecc
Q 009784 266 IDAAINSGNSGGPAFNDKGKCVGIAFQSL 294 (526)
Q Consensus 266 ~da~i~~G~SGGPlvn~~G~VVGI~~~~~ 294 (526)
.+..+.+|+||+|+||.+|++|||...+.
T Consensus 90 ~~~d~~~GsSGSpi~n~~g~ivGlYg~g~ 118 (132)
T PF00949_consen 90 IDLDFPKGSSGSPIFNQNGEIVGLYGNGV 118 (132)
T ss_dssp E---S-TTGTT-EEEETTSCEEEEEEEEE
T ss_pred eecccCCCCCCCceEcCCCcEEEEEccce
Confidence 34447799999999999999999987765
No 64
>KOG3550 consensus Receptor targeting protein Lin-7 [Extracellular structures]
Probab=94.04 E-value=0.1 Score=47.22 Aligned_cols=36 Identities=25% Similarity=0.448 Sum_probs=32.4
Q ss_pred ccceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCC
Q 009784 352 QKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIAND 387 (526)
Q Consensus 352 ~~Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~ 387 (526)
.+-++|+.+.|++-|+. ||+-||.+++|||..+...
T Consensus 114 nspiyisriipggvadrhgglkrgdqllsvngvsvege 151 (207)
T KOG3550|consen 114 NSPIYISRIIPGGVADRHGGLKRGDQLLSVNGVSVEGE 151 (207)
T ss_pred CCceEEEeecCCccccccCcccccceeEeecceeecch
Confidence 45699999999999988 9999999999999998753
No 65
>KOG3834 consensus Golgi reassembly stacking protein GRASP65, contains PDZ domain [Intracellular trafficking, secretion, and vesicular transport]
Probab=93.27 E-value=0.36 Score=50.84 Aligned_cols=114 Identities=16% Similarity=0.238 Sum_probs=73.6
Q ss_pred ccceEEEEeCCCCcccC-CCCC-CcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECC--EEEEEE
Q 009784 352 QKGVRIRRVDPTAPESE-VLKP-SDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDS--KILNFN 427 (526)
Q Consensus 352 ~~Gv~V~~V~~~spA~~-GL~~-GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G--~~~~~~ 427 (526)
..|.-|-+|..+|++++ ||++ -|.|++|||..++...|.. +.+.+.+.. +|+++|+-.. ..++++
T Consensus 14 teg~hvlkVqedSpa~~aglepffdFIvSI~g~rL~~dnd~L----------k~llk~~se-kVkltv~n~kt~~~R~v~ 82 (462)
T KOG3834|consen 14 TEGYHVLKVQEDSPAHKAGLEPFFDFIVSINGIRLNKDNDTL----------KALLKANSE-KVKLTVYNSKTQEVRIVE 82 (462)
T ss_pred ceeEEEEEeecCChHHhcCcchhhhhhheeCcccccCchHHH----------HHHHHhccc-ceEEEEEecccceeEEEE
Confidence 45788999999999999 9988 5899999999999877642 333333333 3999987643 334444
Q ss_pred EEecccccccCCCCCCCCCceEEEeeEEEEeccccceeeeeeeecc-----hhhhcccc-ccceeeee
Q 009784 428 ITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERIMN-----MKLRSSFW-TSSCIQCH 489 (526)
Q Consensus 428 v~l~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt~~~~~~~~~~i~~-----~~~~sg~~-~~~~~~~~ 489 (526)
|+...... . + +-|....-.++...++.+|-|+. .+..+||. ++|-|.-.
T Consensus 83 I~ps~~wg--------g--q---llGvsvrFcsf~~A~~~vwHvl~V~p~SPaalAgl~~~~DYivG~ 137 (462)
T KOG3834|consen 83 IVPSNNWG--------G--Q---LLGVSVRFCSFDGAVESVWHVLSVEPNSPAALAGLRPYTDYIVGI 137 (462)
T ss_pred eccccccc--------c--c---ccceEEEeccCccchhheeeeeecCCCCHHHhcccccccceEecc
Confidence 44333211 0 1 23555555555556666666554 34459999 78877765
No 66
>PF09342 DUF1986: Domain of unknown function (DUF1986); InterPro: IPR015420 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This domain is found in serine endopeptidases belonging to MEROPS peptidase family S1A (clan PA). It is found in unusual mosaic proteins, which are encoded by the Drosophila nudel gene (see P98159 from SWISSPROT). Nudel is involved in defining embryonic dorsoventral polarity. Three proteases; ndl, gd and snk process easter to create active easter. Active easter defines cell identities along the dorsal-ventral continuum by activating the spz ligand for the Tl receptor in the ventral region of the embryo. Nudel, pipe and windbeutel together trigger the protease cascade within the extraembryonic perivitelline compartment which induces dorsoventral polarity of the Drosophila embryo [].
Probab=93.19 E-value=1.2 Score=43.90 Aligned_cols=99 Identities=19% Similarity=0.277 Sum_probs=69.7
Q ss_pred CCCCCccccCC--CcceEEEEEEEeCCEEEecccccCCC----CeEEEEEcCCCcEEE------EEEEEec-----cCCC
Q 009784 135 PNFSLPWQRKR--QYSSSSSGFAIGGRRVLTNAHSVEHY----TQVKLKKRGSDTKYL------ATVLAIG-----TECD 197 (526)
Q Consensus 135 ~~~~~P~~~~~--~~~~~GSGfvI~~g~ILT~aHvV~~~----~~i~V~~~~~g~~~~------a~vv~~d-----~~~D 197 (526)
..+..||.... .+...|+|++|+..|||++..|+.+- ..+.+.+. .++.+. -++..+| ++.+
T Consensus 12 e~y~WPWlA~IYvdG~~~CsgvLlD~~WlLvsssCl~~I~L~~~YvsallG-~~Kt~~~v~Gp~EQI~rVD~~~~V~~S~ 90 (267)
T PF09342_consen 12 EDYHWPWLADIYVDGRYWCSGVLLDPHWLLVSSSCLRGISLSHHYVSALLG-GGKTYLSVDGPHEQISRVDCFKDVPESN 90 (267)
T ss_pred ccccCcceeeEEEcCeEEEEEEEeccceEEEeccccCCcccccceEEEEec-CcceecccCCChheEEEeeeeeeccccc
Confidence 35667887743 45568999999999999999999863 45677774 565443 2444444 5789
Q ss_pred eEEEEecccc-cccCceeeecCCC---CcCCCcEEEEeeCC
Q 009784 198 IAMLTVEDDE-FWEGVLPVEFGEL---PALQDAVTVVGYPI 234 (526)
Q Consensus 198 lAlLkv~~~~-~~~~~~pl~l~~~---~~~g~~V~~iG~p~ 234 (526)
++||.++.+. |...+.|+-+.+. ....+.++++|...
T Consensus 91 v~LLHL~~~~~fTr~VlP~flp~~~~~~~~~~~CVAVg~d~ 131 (267)
T PF09342_consen 91 VLLLHLEQPANFTRYVLPTFLPETSNENESDDECVAVGHDD 131 (267)
T ss_pred eeeeeecCcccceeeecccccccccCCCCCCCceEEEEccc
Confidence 9999998764 4455677656541 12346899999877
No 67
>KOG3532 consensus Predicted protein kinase [General function prediction only]
Probab=92.99 E-value=0.1 Score=57.41 Aligned_cols=39 Identities=15% Similarity=0.474 Sum_probs=35.2
Q ss_pred cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcc
Q 009784 353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVP 391 (526)
Q Consensus 353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~ 391 (526)
.-|.|-.|.++++|.+ .|++|||+++|||.+|.+..+..
T Consensus 398 ~~v~v~tv~~ns~a~k~~~~~gdvlvai~~~pi~s~~q~~ 437 (1051)
T KOG3532|consen 398 RAVKVCTVEDNSLADKAAFKPGDVLVAINNVPIRSERQAT 437 (1051)
T ss_pred eEEEEEEecCCChhhHhcCCCcceEEEecCccchhHHHHH
Confidence 4577889999999999 99999999999999999987753
No 68
>KOG2921 consensus Intramembrane metalloprotease (sterol-regulatory element-binding protein (SREBP) protease) [Posttranslational modification, protein turnover, chaperones]
Probab=91.76 E-value=0.12 Score=53.76 Aligned_cols=40 Identities=25% Similarity=0.341 Sum_probs=36.8
Q ss_pred CccceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCc
Q 009784 351 DQKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTV 390 (526)
Q Consensus 351 ~~~Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l 390 (526)
+..|+.|++|...||+.. ||++||+|+++||-+|++.+|.
T Consensus 218 ~g~gV~Vtev~~~Spl~gprGL~vgdvitsldgcpV~~v~dW 259 (484)
T KOG2921|consen 218 HGEGVTVTEVPSVSPLFGPRGLSVGDVITSLDGCPVHKVSDW 259 (484)
T ss_pred cCceEEEEeccccCCCcCcccCCccceEEecCCcccCCHHHH
Confidence 467999999999999987 9999999999999999998775
No 69
>KOG3605 consensus Beta amyloid precursor-binding protein [General function prediction only]
Probab=91.12 E-value=0.37 Score=53.03 Aligned_cols=103 Identities=20% Similarity=0.263 Sum_probs=69.9
Q ss_pred CCCCCCeee-----cCCCeEEEEEeeccccCccccccccccHHHHHHHHHHHHHcCceeeccc---cCceeeeccChhHH
Q 009784 272 SGNSGGPAF-----NDKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDYEKNGAYTGFPL---LGVEWQKMENPDLR 343 (526)
Q Consensus 272 ~G~SGGPlv-----n~~G~VVGI~~~~~~~~~~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~---LGi~~~~~~~~~~~ 343 (526)
.=++|||.- |.-.+++.|+-..+ ..+|.+..+.+++.++..-.+. +-. --+.-..+..|+.+
T Consensus 679 nmm~~GpAarsgkLnIGDQiiaING~SL---------VGLPLstcQs~Ik~~KnQT~Vk-ltiV~cpPV~~V~I~RPd~k 748 (829)
T KOG3605|consen 679 NMMHGGPAARSGKLNIGDQIMSINGTSL---------VGLPLSTCQSIIKGLKNQTAVK-LNIVSCPPVTTVLIRRPDLR 748 (829)
T ss_pred hcccCChhhhcCCccccceeEeecCcee---------ccccHHHHHHHHhcccccceEE-EEEecCCCceEEEeecccch
Confidence 456777753 44445666552221 3489999999999886533332 111 11222233478888
Q ss_pred hhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecC
Q 009784 344 VAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAN 386 (526)
Q Consensus 344 ~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~ 386 (526)
..||++ .+.|+ |=.+..++-|++ |++.|-.|++|||+.|--
T Consensus 749 yQLGFS-VQNGi-ICSLlRGGIAERGGVRVGHRIIEINgQSVVA 790 (829)
T KOG3605|consen 749 YQLGFS-VQNGI-ICSLLRGGIAERGGVRVGHRIIEINGQSVVA 790 (829)
T ss_pred hhccce-eeCcE-eehhhcccchhccCceeeeeEEEECCceEEe
Confidence 889997 67787 456889999999 999999999999998854
No 70
>PF00944 Peptidase_S3: Alphavirus core protein ; InterPro: IPR000930 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. Togavirin, also known as Sindbis virus core endopeptidase, is a serine protease resident at the N terminus of the p130 polyprotein of togaviruses []. The endopeptidase signature identifies the peptidase as belonging to the MEROPS peptidase family S3 (togavirin family, clan PA(S)). The polyprotein also includes structural proteins for the nucleocapsid core and for the glycoprotein spikes []. Togavirin is only active while part of the polyprotein, cleavage at a Trp-Ser bond resulting in total lack of activity []. Mutagenesis studies have identified the location of the His-Asp-Ser catalytic triad, and X-ray studies have revealed the protein fold to be similar to that of chymotrypsin [, ].; GO: 0004252 serine-type endopeptidase activity, 0006508 proteolysis, 0016020 membrane; PDB: 2YEW_D 1EP5_A 3J0C_F 1EP6_C 1WYK_D 1DYL_A 1VCQ_B 1VCP_B 1LD4_D 1KXA_A ....
Probab=90.80 E-value=0.18 Score=44.82 Aligned_cols=27 Identities=30% Similarity=0.614 Sum_probs=23.3
Q ss_pred ccCCCCCCCCeeecCCCeEEEEEeecc
Q 009784 268 AAINSGNSGGPAFNDKGKCVGIAFQSL 294 (526)
Q Consensus 268 a~i~~G~SGGPlvn~~G~VVGI~~~~~ 294 (526)
..-.+|+||-|++|..|+||||+.++.
T Consensus 101 g~g~~GDSGRpi~DNsGrVVaIVLGG~ 127 (158)
T PF00944_consen 101 GVGKPGDSGRPIFDNSGRVVAIVLGGA 127 (158)
T ss_dssp TS-STTSTTEEEESTTSBEEEEEEEEE
T ss_pred CCCCCCCCCCccCcCCCCEEEEEecCC
Confidence 345689999999999999999998865
No 71
>PF05580 Peptidase_S55: SpoIVB peptidase S55; InterPro: IPR008763 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to the MEROPS peptidase family S55 (SpoIVB peptidase family, clan PA(S)). The protein SpoIVB plays a key role in signalling in the final sigma-K checkpoint of Bacillus subtilis [, ].
Probab=90.44 E-value=6.6 Score=38.13 Aligned_cols=41 Identities=27% Similarity=0.394 Sum_probs=32.4
Q ss_pred ccCCCCCCCCeeecCCCeEEEEEeeccccCccccccccccHHHH
Q 009784 268 AAINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPVI 311 (526)
Q Consensus 268 a~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~~~~~~~aIP~~~i 311 (526)
..+-.|+||+|++- +|++||-++..+ -+....||.++++..
T Consensus 175 GGIvqGMSGSPI~q-dGKLiGAVthvf--~~dp~~Gygi~ie~M 215 (218)
T PF05580_consen 175 GGIVQGMSGSPIIQ-DGKLIGAVTHVF--VNDPTKGYGIFIEWM 215 (218)
T ss_pred CCEEecccCCCEEE-CCEEEEEEEEEE--ecCCCceeeecHHHH
Confidence 35678999999985 899999987766 345778899987653
No 72
>COG0750 Predicted membrane-associated Zn-dependent proteases 1 [Cell envelope biogenesis, outer membrane]
Probab=90.42 E-value=0.47 Score=49.93 Aligned_cols=57 Identities=28% Similarity=0.407 Sum_probs=45.2
Q ss_pred EEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCE---EEEEEEE-CCEEEE
Q 009784 358 RRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDS---AAVKVLR-DSKILN 425 (526)
Q Consensus 358 ~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~---v~l~v~R-~G~~~~ 425 (526)
.++..++++.. |+++||.|+++|++++.++.++. ..+. ...|.. +.+.+.| +++...
T Consensus 134 ~~v~~~s~a~~a~l~~Gd~iv~~~~~~i~~~~~~~----------~~~~-~~~~~~~~~~~i~~~~~~~~~~~ 195 (375)
T COG0750 134 GEVAPKSAAALAGLRPGDRIVAVDGEKVASWDDVR----------RLLV-AAAGDVFNLLTILVIRLDGEAHA 195 (375)
T ss_pred eecCCCCHHHHcCCCCCCEEEeECCEEccCHHHHH----------HHHH-hccCCcccceEEEEEeccceeee
Confidence 37889999999 99999999999999999998863 3333 334555 8999999 777743
No 73
>KOG1892 consensus Actin filament-binding protein Afadin [Cytoskeleton]
Probab=90.31 E-value=0.34 Score=55.36 Aligned_cols=61 Identities=20% Similarity=0.299 Sum_probs=47.2
Q ss_pred CccceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECC
Q 009784 351 DQKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDS 421 (526)
Q Consensus 351 ~~~Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G 421 (526)
+.-|++|..|.+|++|+. -|+.||.+++|||+.+-...+=. ..+++ ...|..|.|+|-..|
T Consensus 958 ~klGIYvKsVV~GgaAd~DGRL~aGDQLLsVdG~SLiGisQEr--------AA~lm--trtg~vV~leVaKqg 1020 (1629)
T KOG1892|consen 958 RKLGIYVKSVVEGGAADHDGRLEAGDQLLSVDGHSLIGISQER--------AARLM--TRTGNVVHLEVAKQG 1020 (1629)
T ss_pred cccceEEEEeccCCccccccccccCceeeeecCcccccccHHH--------HHHHH--hccCCeEEEehhhhh
Confidence 456999999999999987 59999999999999887665521 11222 346889999987655
No 74
>PF02907 Peptidase_S29: Hepatitis C virus NS3 protease; InterPro: IPR004109 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This signature identifies the Hepatitis C virus NS3 protein as a serine protease which belongs to MEROPS peptidase family S29 (hepacivirin family, clan PA(S)), which has a trypsin-like fold. The non-structural (NS) protein NS3 is one of the NS proteins involved in replication of the HCV genome. The NS2 proteinase (IPR002518 from INTERPRO), a zinc-dependent enzyme, performs a single proteolytic cut to release the N terminus of NS3. The action of NS3 proteinase (NS3P), which resides in the N-terminal one-third of the NS3 protein, then yields all remaining non-structural proteins. The C-terminal two-thirds of the NS3 protein contain a helicase. The functional relationship between the proteinase and helicase domains is unknown. NS3 has a structural zinc-binding site and requires cofactor NS4. It has been suggested that the NS3 serine protease of hepatitus C is involved in cell transformation and that the ability to transform requires an active enzyme [].; GO: 0008236 serine-type peptidase activity, 0006508 proteolysis, 0019087 transformation of host cell by virus; PDB: 2QV1_B 3LOX_C 2OBQ_C 2OC1_C 2OC0_A 3LON_A 3KNX_A 2O8M_A 2OBO_A 2OC8_A ....
Probab=90.28 E-value=0.37 Score=42.93 Aligned_cols=41 Identities=29% Similarity=0.528 Sum_probs=25.8
Q ss_pred CCCCCCCCeeecCCCeEEEEEeeccccCcc-ccccccccHHHH
Q 009784 270 INSGNSGGPAFNDKGKCVGIAFQSLKHEDV-ENIGYVIPTPVI 311 (526)
Q Consensus 270 i~~G~SGGPlvn~~G~VVGI~~~~~~~~~~-~~~~~aIP~~~i 311 (526)
...|+||||++..+|.+|||..+.....+. ..+-| +|.+.+
T Consensus 105 ~lkGSSGgPiLC~~GH~vG~f~aa~~trgvak~i~f-~P~e~l 146 (148)
T PF02907_consen 105 DLKGSSGGPILCPSGHAVGMFRAAVCTRGVAKAIDF-IPVETL 146 (148)
T ss_dssp HHTT-TT-EEEETTSEEEEEEEEEEEETTEEEEEEE-EEHHHH
T ss_pred EEecCCCCcccCCCCCEEEEEEEEEEcCCceeeEEE-Eeeeec
Confidence 347999999999999999997665432222 23333 376543
No 75
>KOG3542 consensus cAMP-regulated guanine nucleotide exchange factor [Signal transduction mechanisms]
Probab=90.21 E-value=0.21 Score=54.94 Aligned_cols=42 Identities=24% Similarity=0.279 Sum_probs=35.6
Q ss_pred cCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCC
Q 009784 347 SMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDG 388 (526)
Q Consensus 347 gl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~ 388 (526)
|=.+...|++|.+|.|++.|+. ||+-||.|++|||+...+..
T Consensus 556 GGsEkGfgifV~~V~pgskAa~~GlKRgDqilEVNgQnfenis 598 (1283)
T KOG3542|consen 556 GGSEKGFGIFVAEVFPGSKAAREGLKRGDQILEVNGQNFENIS 598 (1283)
T ss_pred cCccccceeEEeeecCCchHHHhhhhhhhhhhhccccchhhhh
Confidence 3344567999999999999998 99999999999999776543
No 76
>KOG3552 consensus FERM domain protein FRM-8 [General function prediction only]
Probab=89.43 E-value=0.32 Score=55.49 Aligned_cols=57 Identities=28% Similarity=0.353 Sum_probs=43.0
Q ss_pred ceEEEEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEEC
Q 009784 354 GVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD 420 (526)
Q Consensus 354 Gv~V~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~ 420 (526)
-|+|..|.+|+|+...|++||.|+.|||+++...-- +| ..+++.. -.+.|.|+|.+-
T Consensus 76 PviVr~VT~GGps~GKL~PGDQIl~vN~Epv~dapr------er--vIdlvRa--ce~sv~ltV~qP 132 (1298)
T KOG3552|consen 76 PVIVRFVTEGGPSIGKLQPGDQILAVNGEPVKDAPR------ER--VIDLVRA--CESSVNLTVCQP 132 (1298)
T ss_pred ceEEEEecCCCCccccccCCCeEEEecCcccccccH------HH--HHHHHHH--HhhhcceEEecc
Confidence 388999999999999999999999999999975431 11 1133333 356789998884
No 77
>KOG3571 consensus Dishevelled 3 and related proteins [General function prediction only]
Probab=87.51 E-value=0.66 Score=49.78 Aligned_cols=38 Identities=16% Similarity=0.338 Sum_probs=32.7
Q ss_pred ccceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCC
Q 009784 352 QKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGT 389 (526)
Q Consensus 352 ~~Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~ 389 (526)
+.|++|.+|.+++.-+. -+++||.||.||.....++..
T Consensus 276 DggIYVgsImkgGAVA~DGRIe~GDMiLQVNevsFENmSN 315 (626)
T KOG3571|consen 276 DGGIYVGSIMKGGAVALDGRIEPGDMILQVNEVSFENMSN 315 (626)
T ss_pred CCceEEeeeccCceeeccCccCccceEEEeeecchhhcCc
Confidence 57999999999998666 599999999999988777653
No 78
>KOG3551 consensus Syntrophins (type beta) [Extracellular structures]
Probab=87.27 E-value=0.4 Score=49.78 Aligned_cols=56 Identities=20% Similarity=0.280 Sum_probs=40.7
Q ss_pred eEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEE--EEEC
Q 009784 355 VRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVK--VLRD 420 (526)
Q Consensus 355 v~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~--v~R~ 420 (526)
++|+++-++-.|++ .|..||.|++|||..+.+...-+ ..-.-++.|++|.++ +.|+
T Consensus 112 IlISKIFkGlAADQt~aL~~gDaIlSVNG~dL~~AtHde----------AVqaLKraGkeV~levKy~RE 171 (506)
T KOG3551|consen 112 ILISKIFKGLAADQTGALFLGDAILSVNGEDLRDATHDE----------AVQALKRAGKEVLLEVKYMRE 171 (506)
T ss_pred eehhHhccccccccccceeeccEEEEecchhhhhcchHH----------HHHHHHhhCceeeeeeeeehh
Confidence 88999999999988 79999999999999887654311 111224468876654 4554
No 79
>PF00947 Pico_P2A: Picornavirus core protein 2A; InterPro: IPR000081 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This domain defines cysteine peptidases belong to MEROPS peptidase family C3 (picornain, clan PA(C)), subfamilies 3CA and 3CB. The protein fold of this peptidase domain for members of this family resembles that of the serine peptidase, chymotrypsin [], the type example for clan PA. Picornaviral proteins are expressed as a single polyprotein which is cleaved by the viral 3C cysteine protease []. The poliovirus polyprotein is selectively cleaved between the Gln-|-Gly bond. In other picornavirus reactions Glu may be substituted for Gln, and Ser or Thr for Gly. ; GO: 0008233 peptidase activity, 0006508 proteolysis, 0016032 viral reproduction; PDB: 2HRV_B 1Z8R_A.
Probab=86.68 E-value=2 Score=38.02 Aligned_cols=31 Identities=19% Similarity=0.209 Sum_probs=23.9
Q ss_pred eEEEEcccCCCCCCCCeeecCCCeEEEEEeec
Q 009784 262 LGLQIDAAINSGNSGGPAFNDKGKCVGIAFQS 293 (526)
Q Consensus 262 ~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~ 293 (526)
.++....+..||+-||+|+... -||||++++
T Consensus 79 ~~l~g~Gp~~PGdCGg~L~C~H-GViGi~Tag 109 (127)
T PF00947_consen 79 NLLIGEGPAEPGDCGGILRCKH-GVIGIVTAG 109 (127)
T ss_dssp CEEEEE-SSSTT-TCSEEEETT-CEEEEEEEE
T ss_pred CceeecccCCCCCCCceeEeCC-CeEEEEEeC
Confidence 4566667889999999999865 499999985
No 80
>KOG3605 consensus Beta amyloid precursor-binding protein [General function prediction only]
Probab=84.88 E-value=2.3 Score=47.15 Aligned_cols=119 Identities=15% Similarity=0.123 Sum_probs=70.1
Q ss_pred EeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEecccccc
Q 009784 359 RVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRL 436 (526)
Q Consensus 359 ~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~~~~~ 436 (526)
....++||++ .|-.||.|++|||..+-..---. .+.++...+.-..|+|+|.+=--..++.|. + +..
T Consensus 679 nmm~~GpAarsgkLnIGDQiiaING~SLVGLPLst--------cQs~Ik~~KnQT~VkltiV~cpPV~~V~I~--R-Pd~ 747 (829)
T KOG3605|consen 679 NMMHGGPAARSGKLNIGDQIMSINGTSLVGLPLST--------CQSIIKGLKNQTAVKLNIVSCPPVTTVLIR--R-PDL 747 (829)
T ss_pred hcccCChhhhcCCccccceeEeecCceeccccHHH--------HHHHHhcccccceEEEEEecCCCceEEEee--c-ccc
Confidence 5567899998 69999999999998775321111 123455554445688888875444443332 1 111
Q ss_pred cCCCCCCCCCceEEEeeEEEEeccccceeeeeeeecchhhhccccccceeeeecccchhhhHHHHHHHH
Q 009784 437 IPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLLWCLRCLWL 505 (526)
Q Consensus 437 ~p~~~~~~~p~~~i~gG~~f~~lt~~~~~~~~~~i~~~~~~sg~~~~~~~~~~~~~~~~~~~~~~~~~~ 505 (526)
+| -.||.+++= -...+.|- .-+.|-|+++|-.|-..|.|.+--.-|=|..-|
T Consensus 748 ----------ky--QLGFSVQNG----iICSLlRG-GIAERGGVRVGHRIIEINgQSVVA~pHekIV~l 799 (829)
T KOG3605|consen 748 ----------RY--QLGFSVQNG----IICSLLRG-GIAERGGVRVGHRIIEINGQSVVATPHEKIVQL 799 (829)
T ss_pred ----------hh--hccceeeCc----Eeehhhcc-cchhccCceeeeeEEEECCceEEeccHHHHHHH
Confidence 11 224443321 01112221 346689999999999999987766666565443
No 81
>PF10459 Peptidase_S46: Peptidase S46; InterPro: IPR019500 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This entry represents S46 peptidases, where dipeptidyl-peptidase 7 (DPP-7) is the best-characterised member of this family. It is a serine peptidase that is located on the cell surface and is predicted to have two N-terminal transmembrane domains.
Probab=84.06 E-value=0.53 Score=53.60 Aligned_cols=22 Identities=32% Similarity=0.253 Sum_probs=19.6
Q ss_pred eEEEEEEEe-CCEEEecccccCC
Q 009784 149 SSSSGFAIG-GRRVLTNAHSVEH 170 (526)
Q Consensus 149 ~~GSGfvI~-~g~ILT~aHvV~~ 170 (526)
+.|||-+|+ +|+||||.||+.+
T Consensus 47 gGCSgsfVS~~GLvlTNHHC~~~ 69 (698)
T PF10459_consen 47 GGCSGSFVSPDGLVLTNHHCGYG 69 (698)
T ss_pred CceeEEEEcCCceEEecchhhhh
Confidence 459999999 8999999999864
No 82
>PF02395 Peptidase_S6: Immunoglobulin A1 protease Serine protease Prosite pattern; InterPro: IPR000710 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to the MEROPS peptidase family S6 (clan PA(S)). The type sample being the IgA1-specific serine endopeptidase from Neisseria gonorrhoeae []. These cleave prolyl bonds in the hinge regions of immunoglobulin A heavy chains. Similar specificity is shown by the unrelated family of M26 metalloendopeptidases.; GO: 0004252 serine-type endopeptidase activity, 0006508 proteolysis; PDB: 3SZE_A 3H09_B 3SYJ_A 1WXR_A 3AK5_B.
Probab=81.36 E-value=7.2 Score=45.11 Aligned_cols=161 Identities=23% Similarity=0.249 Sum_probs=74.5
Q ss_pred EEEEEEeCCEEEecccccCCCCeEEEEEcCCCcEEEEEEEEecc--CCCeEEEEecccccccCceeeecCCCC----cC-
Q 009784 151 SSGFAIGGRRVLTNAHSVEHYTQVKLKKRGSDTKYLATVLAIGT--ECDIAMLTVEDDEFWEGVLPVEFGELP----AL- 223 (526)
Q Consensus 151 GSGfvI~~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~vv~~d~--~~DlAlLkv~~~~~~~~~~pl~l~~~~----~~- 223 (526)
|...+|++.||+|.+|+..+...+..--. +...| +++..+. ..|+.+-|++.-- ..+.|++..... ..
T Consensus 67 G~aTLigpqYiVSV~HN~~gy~~v~FG~~-g~~~Y--~iV~RNn~~~~Df~~pRLnK~V--TEvaP~~~t~~~~~~~~y~ 141 (769)
T PF02395_consen 67 GVATLIGPQYIVSVKHNGKGYNSVSFGNE-GQNTY--KIVDRNNYPSGDFHMPRLNKFV--TEVAPAEMTTAGSDSNTYN 141 (769)
T ss_dssp SS-EEEETTEEEBETTG-TSCCEECESCS-STCEE--EEEEEEBETTSTEBEEEESS-----SS----BBSSTTSTTGGG
T ss_pred ceEEEecCCeEEEEEccCCCcCceeeccc-CCceE--EEEEccCCCCcccceeecCceE--EEEeccccccccccccccc
Confidence 67889999999999999855544433221 22333 4444433 3699999998632 234555443321 00
Q ss_pred ---C-CcEEEEe-------eCCCCCc-------eeEEEEEEeceeeeeccCCceeee-----EEEEc----ccCCCCCCC
Q 009784 224 ---Q-DAVTVVG-------YPIGGDT-------ISVTSGVVSRIEILSYVHGSTELL-----GLQID----AAINSGNSG 276 (526)
Q Consensus 224 ---g-~~V~~iG-------~p~~~~~-------~sv~~GiVs~~~~~~~~~~~~~~~-----~i~~d----a~i~~G~SG 276 (526)
. ...+=+| +..+... ...+.|.+..... +..+..... ....+ ....+|+||
T Consensus 142 d~~rY~~f~R~GsG~Q~i~~~~g~~~~~~~~ay~yltgGt~~~~~~--~~n~~~~~~~~~~~~~~~~~pL~n~~~~GDSG 219 (769)
T PF02395_consen 142 DKERYPAFVRVGSGTQYIKDRNGNGTTILGGAYNYLTGGTVYNLPG--YGNGSMILSGDLKKFNSYNGPLPNYGSPGDSG 219 (769)
T ss_dssp HTTTC-EEEEEESSSEEEEECCEEEEEEEEETTSCEEEEEESSEEE--EECTCEEEEESTTTCCCCCSSSBEB--TT-TT
T ss_pred cchhchheeecCCceEEEEcCCCCeeEEEEeccceecCCccccccc--cccceEEEecccccccccCCccccccccCcCC
Confidence 0 1111122 2221100 0123344333110 001100000 01111 234689999
Q ss_pred Ceee--cC---CCeEEEEEeeccccCccccccccccHHHHHHHHHHH
Q 009784 277 GPAF--ND---KGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDY 318 (526)
Q Consensus 277 GPlv--n~---~G~VVGI~~~~~~~~~~~~~~~aIP~~~i~~~l~~l 318 (526)
+||| |. +.-++|+.+....-.+..+....+|.+.+.++.++.
T Consensus 220 SPlF~YD~~~kKWvl~Gv~~~~~~~~g~~~~~~~~~~~f~~~~~~~d 266 (769)
T PF02395_consen 220 SPLFAYDKEKKKWVLVGVLSGGNGYNGKGNWWNVIPPDFINQIKQND 266 (769)
T ss_dssp -EEEEEETTTTEEEEEEEEEEECCCCHSEEEEEEECHHHHHHHHHHC
T ss_pred CceEEEEccCCeEEEEEEEccccccCCccceeEEecHHHHHHHHhhh
Confidence 9987 43 345999987765332333445568888887777664
No 83
>KOG3606 consensus Cell polarity protein PAR6 [Signal transduction mechanisms]
Probab=81.29 E-value=1.8 Score=43.02 Aligned_cols=64 Identities=22% Similarity=0.314 Sum_probs=44.0
Q ss_pred HHHcCceeeccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-C-CCCCcEEEEECCEEecC
Q 009784 318 YEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIAN 386 (526)
Q Consensus 318 l~~~g~~~~~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~I~~ 386 (526)
|.++|.-. -||+....-.+.. ..-.|+. +..|+.|.+..|++-|+. | |-..|-+++|||.+|..
T Consensus 164 L~khG~ek---PLGFYIRDG~SVR-Vtp~Gle-kvpGIFISRlVpGGLAeSTGLLaVnDEVlEVNGIEVaG 229 (358)
T KOG3606|consen 164 LHKHGSEK---PLGFYIRDGTSVR-VTPHGLE-KVPGIFISRLVPGGLAESTGLLAVNDEVLEVNGIEVAG 229 (358)
T ss_pred hhhcCCCC---CceEEEecCceEE-ecccccc-ccCceEEEeecCCccccccceeeecceeEEEcCEEecc
Confidence 44555432 3666665431111 1123554 567999999999999999 6 56799999999999964
No 84
>PF12812 PDZ_1: PDZ-like domain
Probab=80.61 E-value=2.8 Score=34.09 Aligned_cols=56 Identities=13% Similarity=-0.023 Sum_probs=41.6
Q ss_pred CceEEEeeEEEEeccccc--------eeeeeeeecchhhhcc-ccccceeeeecccchhhhHHHH
Q 009784 446 PSYYIIAGFVFSRCLYLI--------SVLSMERIMNMKLRSS-FWTSSCIQCHNCQMSSLLWCLR 501 (526)
Q Consensus 446 p~~~i~gG~~f~~lt~~~--------~~~~~~~i~~~~~~sg-~~~~~~~~~~~~~~~~~~~~~~ 501 (526)
-+++.++|++|++|+|+. +++.+.+-..+...+| +..+-.|...|.+...+++.+-
T Consensus 5 ~r~v~~~Ga~f~~Ls~q~aR~~~~~~~gv~v~~~~g~~~~~~~i~~g~iI~~Vn~kpt~~Ld~f~ 69 (78)
T PF12812_consen 5 SRFVEVCGAVFHDLSYQQARQYGIPVGGVYVAVSGGSLAFAGGISKGFIITSVNGKPTPDLDDFI 69 (78)
T ss_pred CEEEEEcCeecccCCHHHHHHhCCCCCEEEEEecCCChhhhCCCCCCeEEEeECCcCCcCHHHHH
Confidence 389999999999999653 2333433333333455 9999999999999998888764
No 85
>KOG3549 consensus Syntrophins (type gamma) [Extracellular structures]
Probab=80.31 E-value=1.8 Score=44.43 Aligned_cols=55 Identities=22% Similarity=0.266 Sum_probs=41.0
Q ss_pred ceEEEEeCCCCcccC-C-CCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEE
Q 009784 354 GVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVL 418 (526)
Q Consensus 354 Gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~ 418 (526)
-++|.++..+-.|+. | |-.||-|++|||..|..-..= +-+ .++ .+.||.|+|+|.
T Consensus 81 PvviSkI~kdQaAd~tG~LFvGDAilqvNGi~v~~c~He-----evV---~iL--RNAGdeVtlTV~ 137 (505)
T KOG3549|consen 81 PVVISKIYKDQAADITGQLFVGDAILQVNGIYVTACPHE-----EVV---NIL--RNAGDEVTLTVK 137 (505)
T ss_pred cEEeehhhhhhhhhhcCceEeeeeeEEeccEEeecCChH-----HHH---HHH--HhcCCEEEEEeH
Confidence 388999998888888 5 789999999999999875321 111 222 347999998874
No 86
>KOG3651 consensus Protein kinase C, alpha binding protein [Signal transduction mechanisms]
Probab=79.41 E-value=2.7 Score=42.41 Aligned_cols=37 Identities=22% Similarity=0.356 Sum_probs=32.4
Q ss_pred ceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCc
Q 009784 354 GVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTV 390 (526)
Q Consensus 354 Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l 390 (526)
=++|.+|-.++||++ -++.||-|++|||..|.....+
T Consensus 31 ClYiVQvFD~tPAa~dG~i~~GDEi~avNg~svKGktKv 69 (429)
T KOG3651|consen 31 CLYIVQVFDKTPAAKDGRIRCGDEIVAVNGISVKGKTKV 69 (429)
T ss_pred eEEEEEeccCCchhccCccccCCeeEEecceeecCccHH
Confidence 378999999999998 5999999999999999876554
No 87
>PF03510 Peptidase_C24: 2C endopeptidase (C24) cysteine protease family; InterPro: IPR000317 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. The two signatures that defines this group of calivirus polyproteins identify a cysteine peptidase signature that belongs to MEROPS peptidase family C24 (clan PA(C)). Caliciviruses are positive-stranded ssRNA viruses that cause gastroenteritis. The calicivirus genome contains two open reading frames, ORF1 and ORF2. ORF2 encodes a structural protein []; while ORF1 encodes a non-structural polypeptide, which has RNA helicase, cysteine protease and RNA polymerase activity. The regions of the polyprotein in which these activities lie are similar to proteins produced by the picornaviruses. Two different families of caliciviruses can be distinguished on the basis of sequence similarity, namely those classified as small round structured viruses (SRSVs) and those classed as non-SRSVs. Calicivirus proteases from the non-SRSV group, which are members of the PA protease clan, constitute family C24 of the cysteine proteases (proteases from SRSVs belong to the C37 family). As mentioned above, the protease activity resides within a polyprotein. The enzyme cleaves the polyprotein at sites N-terminal to itself, liberating the polyprotein helicase.; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis
Probab=75.75 E-value=9.6 Score=32.80 Aligned_cols=54 Identities=15% Similarity=0.250 Sum_probs=33.8
Q ss_pred EEEEeCCEEEecccccCCCCeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecccccccCceeeecCCC
Q 009784 153 GFAIGGRRVLTNAHSVEHYTQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGEL 220 (526)
Q Consensus 153 GfvI~~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~ 220 (526)
++-|.+|.++|+.||.+.++.+. |..+ +++. ..-|+++++.+... ++..++++.
T Consensus 3 avHIGnG~~vt~tHva~~~~~v~------g~~f--~~~~--~~ge~~~v~~~~~~----~p~~~ig~g 56 (105)
T PF03510_consen 3 AVHIGNGRYVTVTHVAKSSDSVD------GQPF--KIVK--TDGELCWVQSPLVH----LPAAQIGTG 56 (105)
T ss_pred eEEeCCCEEEEEEEEeccCceEc------CcCc--EEEE--eccCEEEEECCCCC----CCeeEeccC
Confidence 45566899999999998765431 2221 1222 34599999988753 455566543
No 88
>PF01732 DUF31: Putative peptidase (DUF31); InterPro: IPR022382 This domain has no known function. It is found in various hypothetical proteins and putative lipoproteins from mycoplasmas.
Probab=71.03 E-value=3.2 Score=43.89 Aligned_cols=24 Identities=33% Similarity=0.588 Sum_probs=21.2
Q ss_pred cCCCCCCCCeeecCCCeEEEEEee
Q 009784 269 AINSGNSGGPAFNDKGKCVGIAFQ 292 (526)
Q Consensus 269 ~i~~G~SGGPlvn~~G~VVGI~~~ 292 (526)
.+..|.||+.|+|.+|++|||.++
T Consensus 351 ~l~gGaSGS~V~n~~~~lvGIy~g 374 (374)
T PF01732_consen 351 SLGGGASGSMVINQNNELVGIYFG 374 (374)
T ss_pred CCCCCCCcCeEECCCCCEEEEeCC
Confidence 556899999999999999999753
No 89
>KOG0609 consensus Calcium/calmodulin-dependent serine protein kinase/membrane-associated guanylate kinase [Signal transduction mechanisms]
Probab=70.18 E-value=6.7 Score=42.81 Aligned_cols=57 Identities=25% Similarity=0.336 Sum_probs=42.2
Q ss_pred ceEEEEeCCCCcccC-C-CCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEEC
Q 009784 354 GVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD 420 (526)
Q Consensus 354 Gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~ 420 (526)
-++|..+..|+-+++ | |+.||.|++|||..+.+..--+ +..++.... | .|+++|.=.
T Consensus 147 ~~~vARI~~GG~~~r~glL~~GD~i~EvNGi~v~~~~~~e--------~q~~l~~~~-G-~itfkiiP~ 205 (542)
T KOG0609|consen 147 KVVVARIMHGGMADRQGLLHVGDEILEVNGISVANKSPEE--------LQELLRNSR-G-SITFKIIPS 205 (542)
T ss_pred ccEEeeeccCCcchhccceeeccchheecCeecccCCHHH--------HHHHHHhCC-C-cEEEEEccc
Confidence 588999999999988 4 8999999999999998753211 224555543 4 688887543
No 90
>KOG0606 consensus Microtubule-associated serine/threonine kinase and related proteins [Signal transduction mechanisms; General function prediction only]
Probab=69.61 E-value=4.9 Score=47.40 Aligned_cols=34 Identities=21% Similarity=0.258 Sum_probs=30.0
Q ss_pred eEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCC
Q 009784 355 VRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDG 388 (526)
Q Consensus 355 v~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~ 388 (526)
-.|..|.+++||.. ||++||.|+.+||+++....
T Consensus 660 h~v~sv~egsPA~~agls~~DlIthvnge~v~gl~ 694 (1205)
T KOG0606|consen 660 HSVGSVEEGSPAFEAGLSAGDLITHVNGEPVHGLV 694 (1205)
T ss_pred eeeeeecCCCCccccCCCccceeEeccCcccchhh
Confidence 44788999999988 99999999999999997643
No 91
>TIGR02860 spore_IV_B stage IV sporulation protein B. SpoIVB, the stage IV sporulation protein B of endospore-forming bacteria such as Bacillus subtilis, is a serine proteinase, expressed in the spore (rather than mother cell) compartment, that participates in a proteolytic activation cascade for Sigma-K. It appears to be universal among endospore-forming bacteria and occurs nowhere else.
Probab=68.01 E-value=3.6 Score=43.79 Aligned_cols=42 Identities=24% Similarity=0.423 Sum_probs=32.2
Q ss_pred ccCCCCCCCCeeecCCCeEEEEEeeccccCccccccccccHHHHH
Q 009784 268 AAINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPVIM 312 (526)
Q Consensus 268 a~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~~~~~~~aIP~~~i~ 312 (526)
..+-.|+||+|++- +|++||=++--+. +....||.|-++...
T Consensus 355 gGivqGMSGSPi~q-~gkliGAvtHVfv--ndpt~GYGi~ie~Ml 396 (402)
T TIGR02860 355 GGIVQGMSGSPIIQ-NGKVIGAVTHVFV--NDPTSGYGVYIEWML 396 (402)
T ss_pred CCEEecccCCCEEE-CCEEEEEEEEEEe--cCCCcceeehHHHHH
Confidence 35678999999995 8999998877663 456778888776643
No 92
>KOG1924 consensus RhoA GTPase effector DIA/Diaphanous [Signal transduction mechanisms; Cytoskeleton]
Probab=54.44 E-value=42 Score=38.52 Aligned_cols=10 Identities=30% Similarity=0.531 Sum_probs=5.6
Q ss_pred CeEEEEeccc
Q 009784 197 DIAMLTVEDD 206 (526)
Q Consensus 197 DlAlLkv~~~ 206 (526)
-++||+++.+
T Consensus 719 k~~ILevne~ 728 (1102)
T KOG1924|consen 719 KNVILEVNED 728 (1102)
T ss_pred HHHHhhccHH
Confidence 4566666544
No 93
>PF05416 Peptidase_C37: Southampton virus-type processing peptidase; InterPro: IPR001665 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This group of cysteine peptidases belong to the MEROPS peptidase family C37, (clan PA(C)). The type example is calicivirin from Southampton virus, an endopeptidase that cleaves the polyprotein at sites N-terminal to itself, liberating the polyprotein helicase. Southampton virus is a positive-stranded ssRNA virus belonging to the Caliciviruses, which are viruses that cause gastroenteritis. The calicivirus genome contains two open reading frames, ORF1 and ORF2. ORF1 encodes a non-structural polypeptide, which has RNA helicase, cysteine protease and RNA polymerase activity []. The regions of the polyprotein in which these activities lie are similar to proteins produced by the picornaviruses []. ORF2 encodes a structural, capsid protein. Two different families of caliciviruses can be distinguished on the basis of sequence similarity, namely the Norwalk-like viruses or small round structured viruses (SRSVs), and those classed as non-SRSVs.; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis; PDB: 2FYQ_A 2FYR_A 1WQS_D 4ASH_A 2IPH_B.
Probab=54.29 E-value=95 Score=33.29 Aligned_cols=135 Identities=16% Similarity=0.215 Sum_probs=65.8
Q ss_pred eEEEEEEEeCCEEEecccccCCC-CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecccccccCceeeecCCCCcCCCcE
Q 009784 149 SSSSGFAIGGRRVLTNAHSVEHY-TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGELPALQDAV 227 (526)
Q Consensus 149 ~~GSGfvI~~g~ILT~aHvV~~~-~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~~~~g~~V 227 (526)
+.|=||-++..+++|+-||+... .++. | .+-.-+.++..-+++-+++..+- -.++.-+-|.+...-|.-+
T Consensus 379 GsGWGfWVS~~lfITttHViP~g~~E~F------G--v~i~~i~vh~sGeF~~~rFpk~i-RPDvtgmiLEeGapEGtV~ 449 (535)
T PF05416_consen 379 GSGWGFWVSPTLFITTTHVIPPGAKEAF------G--VPISQIQVHKSGEFCRFRFPKPI-RPDVTGMILEEGAPEGTVC 449 (535)
T ss_dssp TTEEEEESSSSEEEEEGGGS-STTSEET------T--EECGGEEEEEETTEEEEEESS-S-STTS---EE-SS--TT-EE
T ss_pred CCceeeeecceEEEEeeeecCCcchhhh------C--CChhHeEEeeccceEEEecCCCC-CCCccceeeccCCCCceEE
Confidence 46778999999999999999743 2210 0 01111233344566666666542 1234444443333334222
Q ss_pred -EEEeeCCCCC-ceeEEEEEEeceeeee-ccCCceeeeEEEE-------cccCCCCCCCCeeecCCC---eEEEEEeecc
Q 009784 228 -TVVGYPIGGD-TISVTSGVVSRIEILS-YVHGSTELLGLQI-------DAAINSGNSGGPAFNDKG---KCVGIAFQSL 294 (526)
Q Consensus 228 -~~iG~p~~~~-~~sv~~GiVs~~~~~~-~~~~~~~~~~i~~-------da~i~~G~SGGPlvn~~G---~VVGI~~~~~ 294 (526)
+.|-.+.|.. .+.+..|......... ...| ...++-+ |-.+.||+-|.|-+-..| -|+|++++..
T Consensus 450 siLiKR~sGEllpLAvRMgt~AsmkIqgr~v~G--Q~GMLLTGaNAK~mDLGT~PGDCGcPYvyKrgNd~VV~GVH~AAt 527 (535)
T PF05416_consen 450 SILIKRPSGELLPLAVRMGTHASMKIQGRTVHG--QMGMLLTGANAKGMDLGTIPGDCGCPYVYKRGNDWVVIGVHAAAT 527 (535)
T ss_dssp EEEEE-TTSBEEEEEEEEEEEEEEEETTEEEEE--EEEEETTSTT-SSTTTS--TTGTT-EEEEEETTEEEEEEEEEEE-
T ss_pred EEEEEcCCccchhhhhhhccceeEEEcceeecc--eeeeeeecCCccccccCCCCCCCCCceeeecCCcEEEEEEEehhc
Confidence 3355555532 2456677666543210 0111 1123323 335679999999886655 4999998865
No 94
>smart00384 AT_hook DNA binding domain with preference for A/T rich regions. Small DNA-binding motif first described in the high mobility group non-histone chromosomal protein HMG-I(Y).
Probab=45.96 E-value=13 Score=23.58 Aligned_cols=16 Identities=50% Similarity=0.667 Sum_probs=12.6
Q ss_pred hhccCCCCCCCCcccc
Q 009784 5 KRKRGRKPKIPDAEKT 20 (526)
Q Consensus 5 ~~~~~~~~~~~~~~~~ 20 (526)
+|||||-+|.+.....
T Consensus 1 kRkRGRPrK~~~~~~~ 16 (26)
T smart00384 1 KRKRGRPRKAPKDXXX 16 (26)
T ss_pred CCCCCCCCCCCCcccc
Confidence 6899999998876543
No 95
>PF12381 Peptidase_C3G: Tungro spherical virus-type peptidase; InterPro: IPR024387 This entry represents a rice tungro spherical waikavirus-type peptidase that belongs to MEROPS peptidase family C3G. It is a picornain 3C-type protease, and is responsible for the self-cleavage of the positive single-stranded polyproteins of a number of plant viral genomes. The location of the protease activity of the polyprotein is at the C-terminal end, adjacent and N-terminal to the putative RNA polymerase [, ].
Probab=45.51 E-value=30 Score=33.63 Aligned_cols=54 Identities=26% Similarity=0.426 Sum_probs=39.6
Q ss_pred EEEEcccCCCCCCCCeeecC----CCeEEEEEeeccccCccccccccccH--HHHHHHHHHHH
Q 009784 263 GLQIDAAINSGNSGGPAFND----KGKCVGIAFQSLKHEDVENIGYVIPT--PVIMHFIQDYE 319 (526)
Q Consensus 263 ~i~~da~i~~G~SGGPlvn~----~G~VVGI~~~~~~~~~~~~~~~aIP~--~~i~~~l~~l~ 319 (526)
.++..+....|+=|||++-. .-+++||+.++. .+...+||-++ +.+++.+..|.
T Consensus 170 gleY~~~t~~GdCGs~i~~~~t~~~RKIvGiHVAG~---~~~~~gYAe~itQEDL~~A~~~l~ 229 (231)
T PF12381_consen 170 GLEYQMPTMNGDCGSPIVRNNTQMVRKIVGIHVAGS---ANHAMGYAESITQEDLMRAINKLE 229 (231)
T ss_pred eeeEECCCcCCCccceeeEcchhhhhhhheeeeccc---ccccceehhhhhHHHHHHHHHhhc
Confidence 46677788899999997632 368999999876 34567888554 55777766664
No 96
>PF13180 PDZ_2: PDZ domain; PDB: 2L97_A 1Y8T_A 2Z9I_A 1LCY_A 2PZD_B 2P3W_A 1VCW_C 1TE0_B 1SOZ_C 1SOT_C ....
Probab=42.76 E-value=63 Score=25.73 Aligned_cols=50 Identities=8% Similarity=-0.073 Sum_probs=34.1
Q ss_pred eeEEEEeccccceeeeeeeecc--hhhhccccccceeeeecccchhhhHHHHH
Q 009784 452 AGFVFSRCLYLISVLSMERIMN--MKLRSSFWTSSCIQCHNCQMSSLLWCLRC 502 (526)
Q Consensus 452 gG~~f~~lt~~~~~~~~~~i~~--~~~~sg~~~~~~~~~~~~~~~~~~~~~~~ 502 (526)
-|+.|...+...+ +.+..+.+ .+.++|+..||+|...+..-..+...|.-
T Consensus 3 lGv~~~~~~~~~g-~~V~~V~~~spA~~aGl~~GD~I~~ing~~v~~~~~~~~ 54 (82)
T PF13180_consen 3 LGVTVQNLSDTGG-VVVVSVIPGSPAAKAGLQPGDIILAINGKPVNSSEDLVN 54 (82)
T ss_dssp -SEEEEECSCSSS-EEEEEESTTSHHHHTTS-TTEEEEEETTEESSSHHHHHH
T ss_pred ECeEEEEccCCCe-EEEEEeCCCCcHHHCCCCCCcEEEEECCEEcCCHHHHHH
Confidence 4667776665323 33334444 55689999999999999988888777763
No 97
>KOG1924 consensus RhoA GTPase effector DIA/Diaphanous [Signal transduction mechanisms; Cytoskeleton]
Probab=40.69 E-value=80 Score=36.37 Aligned_cols=9 Identities=33% Similarity=0.150 Sum_probs=4.5
Q ss_pred CCCCCcccc
Q 009784 12 PKIPDAEKT 20 (526)
Q Consensus 12 ~~~~~~~~~ 20 (526)
+||++-|-+
T Consensus 502 ~Ki~~l~ae 510 (1102)
T KOG1924|consen 502 EKIKLLEAE 510 (1102)
T ss_pred hhcccCchh
Confidence 555554443
No 98
>KOG3938 consensus RGS-GAIP interacting protein GIPC, contains PDZ domain [Signal transduction mechanisms; Intracellular trafficking, secretion, and vesicular transport]
Probab=39.74 E-value=15 Score=36.77 Aligned_cols=58 Identities=12% Similarity=0.228 Sum_probs=43.9
Q ss_pred eEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEEC
Q 009784 355 VRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD 420 (526)
Q Consensus 355 v~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~ 420 (526)
..|..+.++|--.. -++.||.|-+|||+.|-....++ ..+.+.....|++.+|++.--
T Consensus 151 AFIKrIkegsvidri~~i~VGd~IEaiNge~ivG~RHYe--------VArmLKel~rge~ftlrLieP 210 (334)
T KOG3938|consen 151 AFIKRIKEGSVIDRIEAICVGDHIEAINGESIVGKRHYE--------VARMLKELPRGETFTLRLIEP 210 (334)
T ss_pred eeeEeecCCchhhhhhheeHHhHHHhhcCccccchhHHH--------HHHHHHhcccCCeeEEEeecc
Confidence 56778888887776 79999999999999998766543 225566666788887776543
No 99
>cd01720 Sm_D2 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit D2 heterodimerizes with subunit D1 and three such heterodimers form a hexameric ring structure with alternating D1 and D2 subunits. The D1 - D2 heterodimer also assembles into a heptameric ring containing D2, D3, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=37.07 E-value=63 Score=26.82 Aligned_cols=37 Identities=30% Similarity=0.493 Sum_probs=30.5
Q ss_pred ccCCCCeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784 167 SVEHYTQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE 204 (526)
Q Consensus 167 vV~~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~ 204 (526)
++.....+.|.+. +++.+.+++.++|.+.++.|=...
T Consensus 10 ~~~~~~~V~V~lr-~~r~~~G~L~~fD~hmNlvL~d~~ 46 (87)
T cd01720 10 AVKNNTQVLINCR-NNKKLLGRVKAFDRHCNMVLENVK 46 (87)
T ss_pred HHcCCCEEEEEEc-CCCEEEEEEEEecCccEEEEcceE
Confidence 3444578999997 899999999999999999876654
No 100
>cd00600 Sm_like The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=34.50 E-value=1.1e+02 Score=23.04 Aligned_cols=33 Identities=12% Similarity=0.238 Sum_probs=27.8
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecc
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVED 205 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~ 205 (526)
..+.|.+. +|+.+.+.+..+|...++.|-....
T Consensus 7 ~~V~V~l~-~g~~~~G~L~~~D~~~Ni~L~~~~~ 39 (63)
T cd00600 7 KTVRVELK-DGRVLEGVLVAFDKYMNLVLDDVEE 39 (63)
T ss_pred CEEEEEEC-CCcEEEEEEEEECCCCCEEECCEEE
Confidence 46888887 9999999999999998888766543
No 101
>KOG3834 consensus Golgi reassembly stacking protein GRASP65, contains PDZ domain [Intracellular trafficking, secretion, and vesicular transport]
Probab=31.54 E-value=37 Score=36.24 Aligned_cols=65 Identities=17% Similarity=0.232 Sum_probs=45.3
Q ss_pred EEEEeCCCCcccC-CCC-CCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCE--EEEEEEEec
Q 009784 356 RIRRVDPTAPESE-VLK-PSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK--ILNFNITLA 431 (526)
Q Consensus 356 ~V~~V~~~spA~~-GL~-~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~--~~~~~v~l~ 431 (526)
-|-+|.++|||+. ||. -+|-|+-+-.......+||. .++. .+.++.+++-|+.-.. .++++++..
T Consensus 112 Hvl~V~p~SPaalAgl~~~~DYivG~~~~~~~~~eDl~----------~lIe-she~kpLklyVYN~D~d~~ReVti~pn 180 (462)
T KOG3834|consen 112 HVLSVEPNSPAALAGLRPYTDYIVGIWDAVMHEEEDLF----------TLIE-SHEGKPLKLYVYNHDTDSCREVTITPN 180 (462)
T ss_pred eeeecCCCCHHHhcccccccceEecchhhhccchHHHH----------HHHH-hccCCCcceeEeecCCCccceEEeecc
Confidence 3668999999999 999 68999998444455566652 4444 4568999999987443 344555543
No 102
>PF02178 AT_hook: AT hook motif; InterPro: IPR017956 AT hooks are DNA-binding motifs with a preference for A/T rich regions. These motifs are found in a variety of proteins, including the high mobility group (HMG) proteins [], in DNA-binding proteins from plants [] and in hBRG1 protein, a central ATPase of the human switching/sucrose non-fermenting (SWI/SNF) remodeling complex []. High mobility group (HMG) proteins are a family of relatively low molecular weight non-histone components in chromatin []. HMG-I and HMG-Y (HMGA) are proteins of about 100 amino acid residues which are produced by the alternative splicing of a single gene. HMG-I/Y proteins bind preferentially to the minor groove of AT-rich regions in double-stranded DNA in a non-sequence specific manner [, ]. It is suggested that these proteins could function in nucleosome phasing and in the 3' end processing of mRNA transcripts. They are also involved in the transcription regulation of genes containing, or in close proximity to, AT-rich regions. ; GO: 0003677 DNA binding; PDB: 2EZE_A 2EZD_A 2EZF_A 2EZG_A.
Probab=31.17 E-value=21 Score=18.98 Aligned_cols=11 Identities=64% Similarity=0.797 Sum_probs=3.8
Q ss_pred hhccCCCCCCC
Q 009784 5 KRKRGRKPKIP 15 (526)
Q Consensus 5 ~~~~~~~~~~~ 15 (526)
+|+|||-+|-.
T Consensus 1 ~r~RGRP~k~~ 11 (13)
T PF02178_consen 1 KRKRGRPRKNA 11 (13)
T ss_dssp S--SS--TT--
T ss_pred CCcCCCCcccc
Confidence 57899887753
No 103
>PF09465 LBR_tudor: Lamin-B receptor of TUDOR domain; InterPro: IPR019023 The Lamin-B receptor is a chromatin and lamin binding protein in the inner nuclear membrane. It is one of the integral inner nuclear envelope membrane proteins responsible for targeting nuclear membranes to chromatin, being a downstream effector of Ran, a small Ras-like nuclear GTPase which regulates NE assembly. Lamin-B receptor interacts with importin beta, a Ran-binding protein, thereby directly contributing to the fusion of membrane vesicles and the formation of the nuclear envelope []. ; PDB: 2L8D_A 2DIG_A.
Probab=30.53 E-value=2.1e+02 Score=21.69 Aligned_cols=38 Identities=24% Similarity=0.211 Sum_probs=29.5
Q ss_pred CCCCeEEEEEcCCCcEEEEEEEEeccCCCeEEEEeccc
Q 009784 169 EHYTQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDD 206 (526)
Q Consensus 169 ~~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~ 206 (526)
.....+.++.+++..-|++++..+|...++.-++++..
T Consensus 7 ~~Ge~V~~rWP~s~lYYe~kV~~~d~~~~~y~V~Y~DG 44 (55)
T PF09465_consen 7 AIGEVVMVRWPGSSLYYEGKVLSYDSKSDRYTVLYEDG 44 (55)
T ss_dssp -SS-EEEEE-TTTS-EEEEEEEEEETTTTEEEEEETTS
T ss_pred cCCCEEEEECCCCCcEEEEEEEEecccCceEEEEEcCC
Confidence 34567899999777788999999999999999999764
No 104
>TIGR03000 plancto_dom_1 Planctomycetes uncharacterized domain TIGR03000. Domains described by this model are found, so far, only in the Planctomycetes (Pirellula sp. strain 1 and Gemmata obscuriglobus), in up to six proteins per genome, and may be duplicated within a protein. The function is unknown.
Probab=30.44 E-value=1.5e+02 Score=24.04 Aligned_cols=49 Identities=29% Similarity=0.415 Sum_probs=32.0
Q ss_pred CCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCC----EEEEEEEECCEEEEEEEE
Q 009784 372 PSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGD----SAAVKVLRDSKILNFNIT 429 (526)
Q Consensus 372 ~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~----~v~l~v~R~G~~~~~~v~ 429 (526)
|-|-.+.+||++..+.+... . ..-.....|. ++..++.|||+..+.+-+
T Consensus 10 PadAkl~v~G~~t~~~G~~R------~---F~T~~L~~G~~y~Y~v~a~~~~dG~~~t~~~~ 62 (75)
T TIGR03000 10 PADAKLKVDGKETNGTGTVR------T---FTTPPLEAGKEYEYTVTAEYDRDGRILTRTRT 62 (75)
T ss_pred CCCCEEEECCeEcccCccEE------E---EECCCCCCCCEEEEEEEEEEecCCcEEEEEEE
Confidence 46788999999999888752 0 1111233455 467778899987665433
No 105
>PF00571 CBS: CBS domain CBS domain web page. Mutations in the CBS domain of Swiss:P35520 lead to homocystinuria.; InterPro: IPR000644 CBS (cystathionine-beta-synthase) domains are small intracellular modules, mostly found in two or four copies within a protein, that occur in a variety of proteins in bacteria, archaea, and eukaryotes [, ]. Tandem pairs of CBS domains can act as binding domains for adenosine derivatives and may regulate the activity of attached enzymatic or other domains []. In some cases, CBS domains may act as sensors of cellular energy status by being activated by AMP and inhibited by ATP []. In chloride ion channels, the CBS domains have been implicated in intracellular targeting and trafficking, as well as in protein-protein interactions, but results vary with different channels: in the CLC-5 channel, the CBS domain was shown to be required for trafficking [], while in the CLC-1 channel, the CBS domain was shown to be critical for channel function, but not necessary for trafficking []. Recent experiments revealing that CBS domains can bind adenosine-containing ligands such ATP, AMP, or S-adenosylmethionine have led to the hypothesis that CBS domains function as sensors of intracellular metabolites [, ]. Crystallographic studies of CBS domains have shown that pairs of CBS sequences form a globular domain where each CBS unit adopts a beta-alpha-beta-beta-alpha pattern []. Crystal structure of the CBS domains of the AMP-activated protein kinase in complexes with AMP and ATP shows that the phosphate groups of AMP/ATP lie in a surface pocket at the interface of two CBS domains, which is lined with basic residues, many of which are associated with disease-causing mutations []. In humans, mutations in conserved residues within CBS domains cause a variety of human hereditary diseases, including (with the gene mutated in parentheses): homocystinuria (cystathionine beta-synthase); Wolff-Parkinson-White syndrome (gamma 2 subunit of AMP-activated protein kinase); retinitis pigmentosa (IMP dehydrogenase-1); congenital myotonia, idiopathic generalized epilepsy, hypercalciuric nephrolithiasis, and classic Bartter syndrome (CLC chloride channel family members).; GO: 0005515 protein binding; PDB: 3JTF_A 3TE5_C 3TDH_C 3T4N_C 2QLV_C 3OI8_A 3LV9_A 2QH1_B 1PVM_B 3LQN_A ....
Probab=29.79 E-value=45 Score=24.12 Aligned_cols=21 Identities=38% Similarity=0.555 Sum_probs=17.4
Q ss_pred CCCCCCeeecCCCeEEEEEee
Q 009784 272 SGNSGGPAFNDKGKCVGIAFQ 292 (526)
Q Consensus 272 ~G~SGGPlvn~~G~VVGI~~~ 292 (526)
.+.+.-|++|.+|+++|+++.
T Consensus 28 ~~~~~~~V~d~~~~~~G~is~ 48 (57)
T PF00571_consen 28 NGISRLPVVDEDGKLVGIISR 48 (57)
T ss_dssp HTSSEEEEESTTSBEEEEEEH
T ss_pred cCCcEEEEEecCCEEEEEEEH
Confidence 356678999999999999875
No 106
>cd01726 LSm6 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm6 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=29.28 E-value=1.2e+02 Score=23.52 Aligned_cols=32 Identities=22% Similarity=0.255 Sum_probs=27.4
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE 204 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~ 204 (526)
..+.|.+. +|+.|.+++.++|+..++.|=...
T Consensus 11 ~~V~V~Lk-~g~~~~G~L~~~D~~mNlvL~~~~ 42 (67)
T cd01726 11 RPVVVKLN-SGVDYRGILACLDGYMNIALEQTE 42 (67)
T ss_pred CeEEEEEC-CCCEEEEEEEEEccceeeEEeeEE
Confidence 46889997 899999999999999998886553
No 107
>PRK00737 small nuclear ribonucleoprotein; Provisional
Probab=28.94 E-value=1.2e+02 Score=23.87 Aligned_cols=33 Identities=6% Similarity=0.227 Sum_probs=28.3
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecc
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVED 205 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~ 205 (526)
..+.|.+. +|+.+.+++.++|...++.|=....
T Consensus 15 k~V~V~lk-~g~~~~G~L~~~D~~mNlvL~d~~e 47 (72)
T PRK00737 15 SPVLVRLK-GGREFRGELQGYDIHMNLVLDNAEE 47 (72)
T ss_pred CEEEEEEC-CCCEEEEEEEEEcccceeEEeeEEE
Confidence 46888887 8999999999999999998877643
No 108
>cd01731 archaeal_Sm1 The archaeal sm1 proteins: The Sm proteins are conserved in all three domains of life and are always associated with U-rich RNA sequences. They function to mediate RNA-RNA interactions and RNA biogenesis. All Sm proteins contain a common sequence motif in two segments, Sm1 and Sm2, separated by a short variable linker. Eukaryotic Sm proteins form part of specific small nuclear ribonucleoproteins (snRNPs) that are involved in the processing of pre-mRNAs to mature mRNAs, and are a major component of the eukaryotic spliceosome. Most snRNPs consist of seven Sm proteins (B/B', D1, D2, D3, E, F and G) arranged in a ring on a uridine-rich sequence (Sm site), plus a small nuclear RNA (snRNA) (either U1, U2, U5 or U4/6). Since archaebacteria do not have any splicing apparatus, Sm proteins of archaebacteria may play a more general role. Archaeal Lsm proteins are likely to represent the ancestral Sm domain.
Probab=28.75 E-value=1.3e+02 Score=23.39 Aligned_cols=33 Identities=9% Similarity=0.172 Sum_probs=28.7
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecc
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVED 205 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~ 205 (526)
..+.|.+. +|+.+.+++.++|...++.|-....
T Consensus 11 ~~V~V~l~-~g~~~~G~L~~~D~~mNlvL~~~~e 43 (68)
T cd01731 11 KPVLVKLK-GGKEVRGRLKSYDQHMNLVLEDAEE 43 (68)
T ss_pred CEEEEEEC-CCCEEEEEEEEECCcceEEEeeEEE
Confidence 56888897 8999999999999999998877654
No 109
>cd01722 Sm_F The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit F is capable of forming both homo- and hetero-heptamer ring structures. To form the hetero-heptamer, Sm subunit F initially binds subunits E and G to form a trimer which then assembles onto snRNA along with the D3/B and D1/D2 heterodimers.
Probab=28.45 E-value=1.2e+02 Score=23.73 Aligned_cols=32 Identities=16% Similarity=0.306 Sum_probs=27.2
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE 204 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~ 204 (526)
..+.|.+. +|+.+.+++.++|...++.+=...
T Consensus 12 ~~V~V~Lk-~g~~~~G~L~~~D~~mNi~L~~~~ 43 (68)
T cd01722 12 KPVIVKLK-WGMEYKGTLVSVDSYMNLQLANTE 43 (68)
T ss_pred CEEEEEEC-CCcEEEEEEEEECCCEEEEEeeEE
Confidence 46889997 999999999999999888875553
No 110
>cd01730 LSm3 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm3 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=27.05 E-value=1.1e+02 Score=24.86 Aligned_cols=31 Identities=19% Similarity=0.225 Sum_probs=26.5
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEe
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTV 203 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv 203 (526)
..+.|.+. +|+.+.+++.++|...+|.|=..
T Consensus 12 k~V~V~l~-~gr~~~G~L~~fD~~mNlvL~d~ 42 (82)
T cd01730 12 ERVYVKLR-GDRELRGRLHAYDQHLNMILGDV 42 (82)
T ss_pred CEEEEEEC-CCCEEEEEEEEEccceEEeccce
Confidence 56888887 89999999999999998887544
No 111
>cd01717 Sm_B The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit B heterodimerizes with subunit D3 and three such heterodimers form a hexameric ring structure with alternating B and D3 subunits. The D3 - B heterodimer also assembles into a heptameric ring containing D1, D2, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=26.17 E-value=1.3e+02 Score=24.10 Aligned_cols=32 Identities=9% Similarity=0.327 Sum_probs=27.3
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE 204 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~ 204 (526)
..+.|.+. +|+.+.+.+.++|...+|.|=...
T Consensus 11 ~~V~V~l~-dgR~~~G~L~~~D~~~NlVL~~~~ 42 (79)
T cd01717 11 YRLRVTLQ-DGRQFVGQFLAFDKHMNLVLSDCE 42 (79)
T ss_pred CEEEEEEC-CCcEEEEEEEEEcCccCEEcCCEE
Confidence 46888887 999999999999999998876554
No 112
>COG0298 HypC Hydrogenase maturation factor [Posttranslational modification, protein turnover, chaperones]
Probab=26.04 E-value=1.5e+02 Score=24.25 Aligned_cols=47 Identities=21% Similarity=0.376 Sum_probs=31.8
Q ss_pred EEEEEEEeccCCCeEEEEecccccccCceeeec-CCCCcCCCcEEE-EeeCC
Q 009784 185 YLATVLAIGTECDIAMLTVEDDEFWEGVLPVEF-GELPALQDAVTV-VGYPI 234 (526)
Q Consensus 185 ~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l-~~~~~~g~~V~~-iG~p~ 234 (526)
++++++..+...++|++.+-.-. .---+.+ ....++|++|.+ +||..
T Consensus 5 iPgqI~~I~~~~~~A~Vd~gGvk---reV~l~Lv~~~v~~GdyVLVHvGfAi 53 (82)
T COG0298 5 IPGQIVEIDDNNHLAIVDVGGVK---REVNLDLVGEEVKVGDYVLVHVGFAM 53 (82)
T ss_pred cccEEEEEeCCCceEEEEeccEe---EEEEeeeecCccccCCEEEEEeeEEE
Confidence 57889999988889999987643 1112222 236688998876 67653
No 113
>cd06168 LSm9 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm9 proteins have a single Sm-like domain structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=25.59 E-value=1.6e+02 Score=23.58 Aligned_cols=32 Identities=9% Similarity=0.354 Sum_probs=27.2
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE 204 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~ 204 (526)
..+.|.+. ||+.+.+++..+|...+|.|=...
T Consensus 11 ~~v~V~l~-dgR~~~G~l~~~D~~~NivL~~~~ 42 (75)
T cd06168 11 RTMRIHMT-DGRTLVGVFLCTDRDCNIILGSAQ 42 (75)
T ss_pred CeEEEEEc-CCeEEEEEEEEEcCCCcEEecCcE
Confidence 46888997 999999999999999998775554
No 114
>cd01732 LSm5 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm4 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=24.68 E-value=1.5e+02 Score=23.90 Aligned_cols=31 Identities=16% Similarity=0.388 Sum_probs=26.7
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEe
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTV 203 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv 203 (526)
..+.|.+. +|+.+.+++.++|...++.|=..
T Consensus 14 ~~V~V~l~-~gr~~~G~L~g~D~~mNlvL~da 44 (76)
T cd01732 14 SRIWIVMK-SDKEFVGTLLGFDDYVNMVLEDV 44 (76)
T ss_pred CEEEEEEC-CCeEEEEEEEEeccceEEEEccE
Confidence 57888887 89999999999999999887554
No 115
>cd01729 LSm7 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm7 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=24.64 E-value=1.6e+02 Score=23.96 Aligned_cols=32 Identities=3% Similarity=0.074 Sum_probs=26.9
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE 204 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~ 204 (526)
..+.|.+. +|+.+.+++.++|...+|.|=...
T Consensus 13 k~V~V~l~-~gr~~~G~L~~~D~~mNlvL~~~~ 44 (81)
T cd01729 13 KKIRVKFQ-GGREVTGILKGYDQLLNLVLDDTV 44 (81)
T ss_pred CeEEEEEC-CCcEEEEEEEEEcCcccEEecCEE
Confidence 46888887 899999999999999988875543
No 116
>cd01735 LSm12_N LSm12 belongs to a family of Sm-like proteins that associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet that associates with other Sm proteins to form hexameric and heptameric ring structures. In addition to the N-terminal Sm-like domain, LSm12 has a novel methyltransferase domain.
Probab=24.63 E-value=2.6e+02 Score=21.69 Aligned_cols=33 Identities=15% Similarity=0.234 Sum_probs=27.2
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecc
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVED 205 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~ 205 (526)
..+.+... .|..++++|+.+|....+.+||-+.
T Consensus 7 s~V~~kTc-~g~~ieGEV~afD~~tk~lIlk~~s 39 (61)
T cd01735 7 SQVSCRTC-FEQRLQGEVVAFDYPSKMLILKCPS 39 (61)
T ss_pred cEEEEEec-CCceEEEEEEEecCCCcEEEEECcc
Confidence 44556665 6899999999999999999999665
No 117
>PF11874 DUF3394: Domain of unknown function (DUF3394); InterPro: IPR021814 This domain is functionally uncharacterised. This domain is found in bacteria. This presumed domain is about 190 amino acids in length. This domain is found associated with PF06808 from PFAM.
Probab=24.03 E-value=71 Score=30.37 Aligned_cols=28 Identities=14% Similarity=0.051 Sum_probs=25.1
Q ss_pred ccceEEEEeCCCCcccC-CCCCCcEEEEE
Q 009784 352 QKGVRIRRVDPTAPESE-VLKPSDIILSF 379 (526)
Q Consensus 352 ~~Gv~V~~V~~~spA~~-GL~~GDvIl~v 379 (526)
...+.|..|..+|||++ |+.-|+.|+++
T Consensus 121 ~~~~~Vd~v~fgS~A~~~g~d~d~~I~~v 149 (183)
T PF11874_consen 121 GGKVIVDEVEFGSPAEKAGIDFDWEITEV 149 (183)
T ss_pred CCEEEEEecCCCCHHHHcCCCCCcEEEEE
Confidence 45689999999999999 99999988887
No 118
>COG2524 Predicted transcriptional regulator, contains C-terminal CBS domains [Transcription]
Probab=23.39 E-value=2.2e+02 Score=28.72 Aligned_cols=94 Identities=26% Similarity=0.321 Sum_probs=47.9
Q ss_pred ccCCCeEEEEecccccccCceeeecCCCCcCC----CcEEEEeeCCCCCc----eeE-EEEEEeceeeeeccCCceeeeE
Q 009784 193 GTECDIAMLTVEDDEFWEGVLPVEFGELPALQ----DAVTVVGYPIGGDT----ISV-TSGVVSRIEILSYVHGSTELLG 263 (526)
Q Consensus 193 d~~~DlAlLkv~~~~~~~~~~pl~l~~~~~~g----~~V~~iG~p~~~~~----~sv-~~GiVs~~~~~~~~~~~~~~~~ 263 (526)
++..--|.+++..+ +..+..+++..+| .++.+.|-=.+.+. ..+ ...++|--..............
T Consensus 110 ~p~~c~a~i~v~Gd-----i~~~~~gD~VrVGPtP~~klvv~G~V~g~Dd~~~~ilidi~~m~siPk~~V~~~~s~~~i~ 184 (294)
T COG2524 110 NPDACRAVIRVVGD-----IRKINIGDSVRVGPTPVNKLVVEGKVIGRDDTANEILIDISKMVSIPKEKVKNLMSKKLIT 184 (294)
T ss_pred CCcccceEEEEEec-----cccCCCCCeEEECCcccceEEEEeEEecccccCCeEEEEEeeeeecCcchhhhhccCCceE
Confidence 45556778888764 5666677766665 33555554333221 111 1222221110000001111222
Q ss_pred EEEccc--------CCCCCCCCeeecCCCeEEEEEee
Q 009784 264 LQIDAA--------INSGNSGGPAFNDKGKCVGIAFQ 292 (526)
Q Consensus 264 i~~da~--------i~~G~SGGPlvn~~G~VVGI~~~ 292 (526)
+..|++ ...|-.|.|++|.+ ++||+.+.
T Consensus 185 v~~d~tl~eaak~f~~~~i~GaPVvd~d-k~vGiit~ 220 (294)
T COG2524 185 VRPDDTLREAAKLFYEKGIRGAPVVDDD-KIVGIITL 220 (294)
T ss_pred ecCCccHHHHHHHHHHcCccCCceecCC-ceEEEEEH
Confidence 333443 24799999999965 99999875
No 119
>cd01719 Sm_G The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit G binds subunits E and F to form a trimer which then assembles onto snRNA along with the D1/D2 and D3/B heterodimers forming a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=22.69 E-value=1.9e+02 Score=22.89 Aligned_cols=32 Identities=9% Similarity=0.104 Sum_probs=26.8
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE 204 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~ 204 (526)
..+.|.+. +|+.+.+++.++|...+|.|=...
T Consensus 11 k~V~V~L~-~g~~~~G~L~~~D~~mNlvL~~~~ 42 (72)
T cd01719 11 KKLSLKLN-GNRKVSGILRGFDPFMNLVLDDAV 42 (72)
T ss_pred CeEEEEEC-CCeEEEEEEEEEcccccEEeccEE
Confidence 46888886 899999999999999888875543
No 120
>PF00595 PDZ: PDZ domain (Also known as DHR or GLGF) Coordinates are not yet available; InterPro: IPR001478 PDZ domains are found in diverse signalling proteins in bacteria, yeasts, plants, insects and vertebrates [, ]. PDZ domains can occur in one or multiple copies and are nearly always found in cytoplasmic proteins. They bind either the carboxyl-terminal sequences of proteins or internal peptide sequences []. In most cases, interaction between a PDZ domain and its target is constitutive, with a binding affinity of 1 to 10 microns. However, agonist-dependent activation of cell surface receptors is sometimes required to promote interaction with a PDZ protein. PDZ domain proteins are frequently associated with the plasma membrane, a compartment where high concentrations of phosphatidylinositol 4,5-bisphosphate (PIP2) are found. Direct interaction between PIP2 and a subset of class II PDZ domains (syntenin, CASK, Tiam-1) has been demonstrated. PDZ domains consist of 80 to 90 amino acids comprising six beta-strands (beta-A to beta-F) and two alpha-helices, A and B, compactly arranged in a globular structure. Peptide binding of the ligand takes place in an elongated surface groove as an anti-parallel beta-strand interacts with the beta-B strand and the B helix. The structure of PDZ domains allows binding to a free carboxylate group at the end of a peptide through a carboxylate-binding loop between the beta-A and beta-B strands.; GO: 0005515 protein binding; PDB: 3AXA_A 1WF8_A 1QAV_B 1QAU_A 1B8Q_A 1MC7_A 2KAW_A 1I16_A 1VB7_A 1WI4_A ....
Probab=22.55 E-value=1.8e+02 Score=22.77 Aligned_cols=52 Identities=12% Similarity=-0.073 Sum_probs=34.7
Q ss_pred eEEEEeccccc-eeeeeeeecc--hhhhccccccceeeeecccchhhhHHHHHHH
Q 009784 453 GFVFSRCLYLI-SVLSMERIMN--MKLRSSFWTSSCIQCHNCQMSSLLWCLRCLW 504 (526)
Q Consensus 453 G~~f~~lt~~~-~~~~~~~i~~--~~~~sg~~~~~~~~~~~~~~~~~~~~~~~~~ 504 (526)
|+.+....... ....+..+.+ .+.++|++.||.|...|.+-.....+..+.=
T Consensus 13 G~~l~~~~~~~~~~~~V~~v~~~~~a~~~gl~~GD~Il~INg~~v~~~~~~~~~~ 67 (81)
T PF00595_consen 13 GFTLRGGSDNDEKGVFVSSVVPGSPAERAGLKVGDRILEINGQSVRGMSHDEVVQ 67 (81)
T ss_dssp SEEEEEESTSSSEEEEEEEECTTSHHHHHTSSTTEEEEEETTEESTTSBHHHHHH
T ss_pred CEEEEecCCCCcCCEEEEEEeCCChHHhcccchhhhhheeCCEeCCCCCHHHHHH
Confidence 55555544332 3444555555 4567899999999999998887776666543
No 121
>cd01727 LSm8 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm8 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=21.91 E-value=3.4e+02 Score=21.46 Aligned_cols=33 Identities=6% Similarity=0.104 Sum_probs=27.7
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecc
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVED 205 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~ 205 (526)
..+.|.+. +|+.+.+++.++|...++.|=....
T Consensus 10 ~~V~V~l~-dgr~~~G~L~~~D~~~NlvL~~~~E 42 (74)
T cd01727 10 KTVSVITV-DGRVIVGTLKGFDQATNLILDDSHE 42 (74)
T ss_pred CEEEEEEC-CCcEEEEEEEEEccccCEEccceEE
Confidence 46788886 9999999999999999988876543
No 122
>cd01721 Sm_D3 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit D3 heterodimerizes with subunit B and three such heterodimers form a hexameric ring structure with alternating B and D3 subunits. The D3 - B heterodimer also assembles into a heptameric ring containing D1, D2, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=21.74 E-value=2.1e+02 Score=22.41 Aligned_cols=32 Identities=9% Similarity=0.198 Sum_probs=28.3
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE 204 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~ 204 (526)
..+.|.+- +|..|.+++..+|...++.+-.+.
T Consensus 11 ~~V~VeLk-~g~~~~G~L~~~D~~MNl~L~~~~ 42 (70)
T cd01721 11 HIVTVELK-TGEVYRGKLIEAEDNMNCQLKDVT 42 (70)
T ss_pred CEEEEEEC-CCcEEEEEEEEEcCCceeEEEEEE
Confidence 56888887 899999999999999999988774
No 123
>smart00651 Sm snRNP Sm proteins. small nuclear ribonucleoprotein particles (snRNPs) involved in pre-mRNA splicing
Probab=21.30 E-value=2.2e+02 Score=21.56 Aligned_cols=33 Identities=15% Similarity=0.288 Sum_probs=27.4
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecc
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVED 205 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~ 205 (526)
..+.|.+. +|+.+.+.+..+|...++-|=....
T Consensus 9 ~~V~V~l~-~g~~~~G~L~~~D~~~NlvL~~~~e 41 (67)
T smart00651 9 KRVLVELK-NGREYRGTLKGFDQFMNLVLEDVEE 41 (67)
T ss_pred cEEEEEEC-CCcEEEEEEEEECccccEEEccEEE
Confidence 46888887 8999999999999998888766543
No 124
>cd01728 LSm1 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm1 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=20.75 E-value=2.2e+02 Score=22.81 Aligned_cols=32 Identities=9% Similarity=0.155 Sum_probs=27.2
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE 204 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~ 204 (526)
..+.|.+. +|+.+.+.+.++|+..++.|=...
T Consensus 13 k~v~V~l~-~gr~~~G~L~~fD~~~NlvL~d~~ 44 (74)
T cd01728 13 KKVVVLLR-DGRKLIGILRSFDQFANLVLQDTV 44 (74)
T ss_pred CEEEEEEc-CCeEEEEEEEEECCcccEEecceE
Confidence 46888887 899999999999999998876554
No 125
>PF01423 LSM: LSM domain ; InterPro: IPR001163 This family is found in Lsm (like-Sm) proteins and in bacterial Lsm-related Hfq proteins. In each case, the domain adopts a core structure consisting of an open beta-barrel with an SH3-like topology. Lsm (like-Sm) proteins have diverse functions, and are thought to be important modulators of RNA biogenesis and function [, ]. The Sm proteins form part of specific small nuclear ribonucleoproteins (snRNPs) that are involved in the processing of pre-mRNAs to mature mRNAs, and are a major component of the eukaryotic spliceosome. Most snRNPs consist of seven Sm proteins (B/B', D1, D2, D3, E, F and G) arranged in a ring on a uridine-rich sequence (Sm site), plus a small nuclear RNA (snRNA) (either U1, U2, U5 or U4/6) []. All Sm proteins contain a common sequence motif in two segments, Sm1 and Sm2, separated by a short variable linker []. In other snRNPs, certain Sm proteins are replaced with different Lsm proteins, such as with U7 snRNPs, in which the D1 and D2 Sm proteins are replaced with U7-specific Lsm10 and Lsm11 proteins, where Lsm11 plays a role in histone U7-specific RNA processing []. Lsm proteins are also found in archaebacteria, which do not have any splicing apparatus suggesting a more general role for Lsm proteins. The pleiotropic translational regulator Hfq (host factor Q) is a bacterial Lsm-like protein, which modulates the structure of numerous RNA molecules by binding preferentially to A/U-rich sequences in RNA []. Hfq forms an Lsm-like fold, however, unlike the heptameric Sm proteins, Hfq forms a homo-hexameric ring.; PDB: 1D3B_K 2Y9D_D 2Y9A_D 2Y9C_R 3VRI_C 2Y9B_K 3QUI_D 3M4G_H 3INZ_E 1U1S_C ....
Probab=20.65 E-value=1.7e+02 Score=22.28 Aligned_cols=34 Identities=12% Similarity=0.331 Sum_probs=29.1
Q ss_pred CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEeccc
Q 009784 172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDD 206 (526)
Q Consensus 172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~ 206 (526)
..+.|.+. +|+.+.+.+..+|...++.|-.....
T Consensus 9 ~~V~V~l~-~g~~~~G~L~~~D~~~Nl~L~~~~~~ 42 (67)
T PF01423_consen 9 KRVRVELK-NGRTYRGTLVSFDQFMNLVLSDVTET 42 (67)
T ss_dssp SEEEEEET-TSEEEEEEEEEEETTEEEEEEEEEEE
T ss_pred cEEEEEEe-CCEEEEEEEEEeechheEEeeeEEEE
Confidence 56889997 99999999999999988888777653
No 126
>COG0260 PepB Leucyl aminopeptidase [Amino acid transport and metabolism]
Probab=20.22 E-value=1.3e+02 Score=33.10 Aligned_cols=46 Identities=17% Similarity=0.246 Sum_probs=28.7
Q ss_pred EEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhh
Q 009784 358 RRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ 406 (526)
Q Consensus 358 ~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~ 406 (526)
.-...|.|.....+|||||++.||+.|.=...= -.+|+.+-+.+..
T Consensus 303 l~~~ENm~~g~A~rPGDVits~~GkTVEV~NTD---AEGRLVLADaLtY 348 (485)
T COG0260 303 LPAVENMPSGNAYRPGDVITSMNGKTVEVLNTD---AEGRLVLADALTY 348 (485)
T ss_pred EeeeccCCCCCCCCCCCeEEecCCcEEEEcccC---ccHHHHHHHHHHH
Confidence 344455555556799999999999887522110 1266666665543
Done!