Query 027022
Match_columns 229
No_of_seqs 175 out of 1517
Neff 7.3
Searched_HMMs 46136
Date Fri Mar 29 03:51:01 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/027022.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/027022hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PRK10139 serine endoprotease; 100.0 3.4E-28 7.4E-33 226.1 20.3 141 77-229 41-188 (455)
2 TIGR02038 protease_degS peripl 100.0 2.3E-27 5E-32 214.2 19.8 134 72-229 41-174 (351)
3 PRK10898 serine endoprotease; 100.0 2.2E-27 4.7E-32 214.5 19.5 133 73-229 42-174 (353)
4 PRK10942 serine endoprotease; 99.9 1.6E-26 3.6E-31 215.8 19.4 141 77-229 39-209 (473)
5 TIGR02037 degP_htrA_DO peripla 99.9 4.4E-25 9.5E-30 204.1 16.8 140 78-229 3-155 (428)
6 COG0265 DegQ Trypsin-like seri 99.8 1.6E-20 3.5E-25 169.1 14.7 137 76-229 33-169 (347)
7 KOG1320 Serine protease [Postt 99.6 2.6E-15 5.6E-20 138.6 10.0 148 72-229 124-275 (473)
8 PF13365 Trypsin_2: Trypsin-li 99.1 3.5E-10 7.7E-15 85.2 7.7 61 122-186 1-65 (120)
9 KOG1421 Predicted signaling-as 98.7 2.9E-08 6.4E-13 94.4 8.0 131 75-228 51-185 (955)
10 PF00089 Trypsin: Trypsin; In 98.2 5.6E-05 1.2E-09 62.1 13.1 92 119-218 24-131 (220)
11 cd00190 Tryp_SPc Trypsin-like 98.0 0.00023 4.9E-09 58.8 13.4 95 118-218 23-133 (232)
12 smart00020 Tryp_SPc Trypsin-li 97.9 0.00014 3.1E-09 60.3 10.7 94 119-218 25-133 (229)
13 PF10459 Peptidase_S46: Peptid 97.6 0.00036 7.7E-09 68.6 10.5 24 120-143 47-70 (698)
14 KOG1320 Serine protease [Postt 97.5 7.6E-05 1.7E-09 69.7 4.0 128 81-228 55-184 (473)
15 COG3591 V8-like Glu-specific e 95.6 0.14 2.9E-06 44.5 9.8 94 122-219 66-174 (251)
16 PF00863 Peptidase_C4: Peptida 94.8 0.29 6.3E-06 42.1 9.5 82 120-217 32-119 (235)
17 KOG1421 Predicted signaling-as 92.5 0.27 5.9E-06 48.1 5.8 87 118-217 548-635 (955)
18 PF05579 Peptidase_S32: Equine 88.2 2.1 4.6E-05 37.5 7.1 64 120-199 114-177 (297)
19 PF03510 Peptidase_C24: 2C end 82.2 3.1 6.6E-05 31.4 4.6 55 124-201 3-57 (105)
20 KOG3627 Trypsin [Amino acid tr 79.2 21 0.00045 30.0 9.5 86 121-214 39-148 (256)
21 PF09342 DUF1986: Domain of un 72.7 65 0.0014 28.1 11.2 99 117-223 25-137 (267)
22 PF01455 HupF_HypC: HupF/HypC 64.2 29 0.00062 23.9 5.6 43 167-212 5-47 (68)
23 COG0298 HypC Hydrogenase matur 64.0 14 0.00029 26.5 3.9 46 167-215 5-52 (82)
24 cd01735 LSm12_N LSm12 belongs 62.1 32 0.00068 23.3 5.3 33 150-186 6-38 (61)
25 PRK13922 rod shape-determining 58.7 1.2E+02 0.0026 26.2 9.9 51 178-228 190-240 (276)
26 PF01732 DUF31: Putative pepti 58.3 7 0.00015 35.7 2.2 24 118-141 34-67 (374)
27 COG5640 Secreted trypsin-like 54.9 25 0.00054 32.3 5.0 19 124-143 65-83 (413)
28 PRK10672 rare lipoprotein A; P 54.2 52 0.0011 30.2 7.0 29 111-141 85-113 (361)
29 PF02601 Exonuc_VII_L: Exonucl 49.3 21 0.00046 31.6 3.7 34 120-160 280-313 (319)
30 COG3338 Cah Carbonic anhydrase 48.2 32 0.00069 29.7 4.3 25 161-185 124-148 (250)
31 PF15436 PGBA_N: Plasminogen-b 46.1 1.7E+02 0.0036 25.0 8.4 74 128-211 12-88 (218)
32 PRK10413 hydrogenase 2 accesso 43.9 50 0.0011 23.7 4.2 45 167-211 5-51 (82)
33 PRK14864 putative biofilm stre 41.5 1.5E+02 0.0033 22.2 7.6 15 30-44 5-19 (104)
34 cd00600 Sm_like The eukaryotic 41.4 88 0.0019 20.2 5.0 33 150-186 6-38 (63)
35 PF08605 Rad9_Rad53_bind: Fung 41.3 43 0.00092 26.2 3.8 47 166-213 24-70 (131)
36 TIGR00074 hypC_hupF hydrogenas 37.2 80 0.0017 22.3 4.4 40 167-211 5-44 (76)
37 PF09465 LBR_tudor: Lamin-B re 37.1 1.2E+02 0.0027 20.0 5.1 36 149-187 8-43 (55)
38 PF00548 Peptidase_C3: 3C cyst 36.2 2.3E+02 0.005 22.9 8.0 56 118-187 23-81 (172)
39 PF03761 DUF316: Domain of unk 34.4 45 0.00098 28.7 3.4 40 175-214 158-199 (282)
40 cd01722 Sm_F The eukaryotic Sm 33.4 1.1E+02 0.0023 20.7 4.5 33 150-186 11-43 (68)
41 cd01726 LSm6 The eukaryotic Sm 32.3 1.2E+02 0.0026 20.4 4.6 33 150-186 10-42 (67)
42 cd01728 LSm1 The eukaryotic Sm 29.8 1.9E+02 0.0042 20.0 5.9 57 151-213 13-72 (74)
43 PRK00737 small nuclear ribonuc 29.7 1.9E+02 0.004 19.8 6.6 33 150-186 14-46 (72)
44 TIGR00219 mreC rod shape-deter 29.0 4E+02 0.0087 23.4 9.8 52 177-228 188-241 (283)
45 PF02122 Peptidase_S39: Peptid 28.8 88 0.0019 26.3 4.0 60 118-186 28-88 (203)
46 PF08758 Cadherin_pro: Cadheri 28.7 2.3E+02 0.0049 20.5 6.2 40 120-164 43-82 (90)
47 PRK09507 cspE cold shock prote 28.0 1.5E+02 0.0032 20.2 4.4 46 166-211 3-52 (69)
48 TIGR00237 xseA exodeoxyribonuc 26.9 64 0.0014 30.2 3.2 29 125-160 398-426 (432)
49 PRK10943 cold shock-like prote 25.9 1.5E+02 0.0033 20.1 4.2 45 167-211 4-52 (69)
50 COG5510 Predicted small secret 25.6 77 0.0017 20.0 2.3 22 30-51 2-23 (44)
51 cd01720 Sm_D2 The eukaryotic S 25.5 1.7E+02 0.0036 21.1 4.5 34 149-186 13-46 (87)
52 cd01731 archaeal_Sm1 The archa 24.6 2.2E+02 0.0048 19.0 6.5 33 150-186 10-42 (68)
53 PRK09890 cold shock protein Cs 24.4 2.1E+02 0.0045 19.5 4.7 45 167-211 5-53 (70)
54 cd01717 Sm_B The eukaryotic Sm 24.0 1.9E+02 0.0042 20.0 4.6 33 150-186 10-42 (79)
55 cd05701 S1_Rrp5_repeat_hs10 S1 23.8 1.6E+02 0.0034 20.3 3.8 39 171-212 17-55 (69)
56 cd01730 LSm3 The eukaryotic Sm 23.6 1.7E+02 0.0037 20.5 4.2 31 151-185 12-42 (82)
57 PRK10081 entericidin B membran 23.5 1.5E+02 0.0033 19.0 3.5 22 30-51 2-23 (48)
58 PRK00286 xseA exodeoxyribonucl 22.8 1.3E+02 0.0027 28.0 4.3 30 124-160 402-431 (438)
59 PRK10354 RNA chaperone/anti-te 22.3 2.5E+02 0.0053 19.1 4.7 43 168-210 6-52 (70)
60 TIGR03497 FliI_clade2 flagella 22.3 4.9E+02 0.011 24.3 8.1 38 166-216 33-70 (413)
61 COG1792 MreC Cell shape-determ 21.9 5.5E+02 0.012 22.6 9.2 34 196-229 206-239 (284)
62 PTZ00138 small nuclear ribonuc 21.7 2.4E+02 0.0052 20.5 4.7 36 149-186 25-60 (89)
63 PRK08927 fliI flagellum-specif 21.7 5.7E+02 0.012 24.2 8.4 51 122-185 17-70 (442)
64 PRK15464 cold shock-like prote 21.6 2.5E+02 0.0055 19.2 4.6 44 167-210 5-52 (70)
65 PF10844 DUF2577: Protein of u 21.5 2.5E+02 0.0055 20.5 5.0 56 124-218 37-92 (100)
66 PF04083 Abhydro_lipase: Parti 21.2 2.1E+02 0.0045 19.2 4.1 19 125-143 16-34 (63)
67 PRK10781 rcsF outer membrane l 21.1 1.1E+02 0.0023 24.1 2.9 15 121-136 100-114 (133)
68 cd06168 LSm9 The eukaryotic Sm 21.0 2.7E+02 0.0058 19.3 4.7 31 151-185 11-41 (75)
69 PRK15463 cold shock-like prote 21.0 2.4E+02 0.0051 19.3 4.4 45 167-211 5-53 (70)
70 cd01732 LSm5 The eukaryotic Sm 20.6 2.3E+02 0.0051 19.7 4.4 31 151-185 14-44 (76)
71 PF10518 TAT_signal: TAT (twin 20.6 1.5E+02 0.0032 16.3 2.7 19 28-46 2-20 (26)
72 TIGR00638 Mop molybdenum-pteri 20.2 2.5E+02 0.0055 18.2 4.4 46 166-212 8-58 (69)
73 PRK13684 Ycf48-like protein; P 20.1 2E+02 0.0044 25.6 5.0 8 131-138 102-109 (334)
No 1
>PRK10139 serine endoprotease; Provisional
Probab=99.96 E-value=3.4e-28 Score=226.12 Aligned_cols=141 Identities=28% Similarity=0.478 Sum_probs=115.5
Q ss_pred HHHHHHHHhCCceEEEEeeeeccCCCCCcchhhhhcc------ccccccceEEEEEEcC-CcEEEEccccccccccCCCC
Q 027022 77 RVVQLFQETSPSVVSIQDLELSKNPKSTSSELMLVDG------EYAKVEGTGSGFVWDK-FGHIVTNYHVVAKLATDTSG 149 (229)
Q Consensus 77 ~~~~~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~------~~~~~~~~GSGfiI~~-~G~IlTn~HVv~~~~~~~~~ 149 (229)
++.++++++.||||.|.+......+....+.|..+++ ......+.||||||++ +||||||+|||+ +
T Consensus 41 ~~~~~~~~~~pavV~i~~~~~~~~~~~~~~~~~~~f~~~~~~~~~~~~~~~GSG~ii~~~~g~IlTn~HVv~-------~ 113 (455)
T PRK10139 41 SLAPMLEKVLPAVVSVRVEGTASQGQKIPEEFKKFFGDDLPDQPAQPFEGLGSGVIIDAAKGYVLTNNHVIN-------Q 113 (455)
T ss_pred cHHHHHHHhCCcEEEEEEEEeecccccCchhHHHhccccCCccccccccceEEEEEEECCCCEEEeChHHhC-------C
Confidence 6899999999999999986654322111111211111 1233468999999985 799999999999 8
Q ss_pred cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEcC
Q 027022 150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTFQ 229 (229)
Q Consensus 150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVSa 229 (229)
++.+.|++.|++ +|.|++++.|+.+||||||++.. .++++++|++++.+++||+|++||||+|+..+++.||||+
T Consensus 114 a~~i~V~~~dg~----~~~a~vvg~D~~~DlAvlkv~~~-~~l~~~~lg~s~~~~~G~~V~aiG~P~g~~~tvt~GivS~ 188 (455)
T PRK10139 114 AQKISIQLNDGR----EFDAKLIGSDDQSDIALLQIQNP-SKLTQIAIADSDKLRVGDFAVAVGNPFGLGQTATSGIISA 188 (455)
T ss_pred CCEEEEEECCCC----EEEEEEEEEcCCCCEEEEEecCC-CCCceeEecCccccCCCCEEEEEecCCCCCCceEEEEEcc
Confidence 899999998754 78999999999999999999843 5799999999999999999999999999999999999985
No 2
>TIGR02038 protease_degS periplasmic serine pepetdase DegS. This family consists of the periplasmic serine protease DegS (HhoB), a shorter paralog of protease DO (HtrA, DegP) and DegQ (HhoA). It is found in E. coli and several other Proteobacteria of the gamma subdivision. It contains a trypsin domain and a single copy of PDZ domain (in contrast to DegP with two copies). A critical role of this DegS is to sense stress in the periplasm and partially degrade an inhibitor of sigma(E).
Probab=99.96 E-value=2.3e-27 Score=214.17 Aligned_cols=134 Identities=34% Similarity=0.527 Sum_probs=114.8
Q ss_pred ccchhHHHHHHHHhCCceEEEEeeeeccCCCCCcchhhhhccccccccceEEEEEEcCCcEEEEccccccccccCCCCcc
Q 027022 72 QLEEDRVVQLFQETSPSVVSIQDLELSKNPKSTSSELMLVDGEYAKVEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLH 151 (229)
Q Consensus 72 ~~~~~~~~~~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~ 151 (229)
.+.+..+.++++++.||||+|++.....+ ........+.||||||+++||||||+|||+ +++
T Consensus 41 ~~~~~~~~~~~~~~~psVV~I~~~~~~~~-----------~~~~~~~~~~GSG~vi~~~G~IlTn~HVV~-------~~~ 102 (351)
T TIGR02038 41 NTVEISFNKAVRRAAPAVVNIYNRSISQN-----------SLNQLSIQGLGSGVIMSKEGYILTNYHVIK-------KAD 102 (351)
T ss_pred cccchhHHHHHHhcCCcEEEEEeEecccc-----------ccccccccceEEEEEEeCCeEEEecccEeC-------CCC
Confidence 34455799999999999999998654321 011233467899999999999999999999 888
Q ss_pred eEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEcC
Q 027022 152 RCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTFQ 229 (229)
Q Consensus 152 ~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVSa 229 (229)
.+.|++.|+. .+.|+++++|+.+||||||++.. .++++++++++.+++||+|++||||+|+..+++.|+||+
T Consensus 103 ~i~V~~~dg~----~~~a~vv~~d~~~DlAvlkv~~~--~~~~~~l~~s~~~~~G~~V~aiG~P~~~~~s~t~GiIs~ 174 (351)
T TIGR02038 103 QIVVALQDGR----KFEAELVGSDPLTDLAVLKIEGD--NLPTIPVNLDRPPHVGDVVLAIGNPYNLGQTITQGIISA 174 (351)
T ss_pred EEEEEECCCC----EEEEEEEEecCCCCEEEEEecCC--CCceEeccCcCccCCCCEEEEEeCCCCCCCcEEEEEEEe
Confidence 9999998854 78999999999999999999853 589999999999999999999999999999999999985
No 3
>PRK10898 serine endoprotease; Provisional
Probab=99.95 E-value=2.2e-27 Score=214.48 Aligned_cols=133 Identities=29% Similarity=0.452 Sum_probs=113.9
Q ss_pred cchhHHHHHHHHhCCceEEEEeeeeccCCCCCcchhhhhccccccccceEEEEEEcCCcEEEEccccccccccCCCCcce
Q 027022 73 LEEDRVVQLFQETSPSVVSIQDLELSKNPKSTSSELMLVDGEYAKVEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHR 152 (229)
Q Consensus 73 ~~~~~~~~~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~ 152 (229)
..+..+.++++++.||||.|.+..... .........+.||||+|+++||||||+|||+ +++.
T Consensus 42 ~~~~~~~~~~~~~~psvV~v~~~~~~~-----------~~~~~~~~~~~GSGfvi~~~G~IlTn~HVv~-------~a~~ 103 (353)
T PRK10898 42 ETPASYNQAVRRAAPAVVNVYNRSLNS-----------TSHNQLEIRTLGSGVIMDQRGYILTNKHVIN-------DADQ 103 (353)
T ss_pred cccchHHHHHHHhCCcEEEEEeEeccc-----------cCcccccccceeeEEEEeCCeEEEecccEeC-------CCCE
Confidence 334578999999999999999865321 0111234457899999999999999999999 7889
Q ss_pred EEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEcC
Q 027022 153 CKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTFQ 229 (229)
Q Consensus 153 ~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVSa 229 (229)
+.|++.|+. +|.|+++++|+.+||||||++. .++++++|++++.+++||+|+++|||+|+..+++.|+||+
T Consensus 104 i~V~~~dg~----~~~a~vv~~d~~~DlAvl~v~~--~~l~~~~l~~~~~~~~G~~V~aiG~P~g~~~~~t~Giis~ 174 (353)
T PRK10898 104 IIVALQDGR----VFEALLVGSDSLTDLAVLKINA--TNLPVIPINPKRVPHIGDVVLAIGNPYNLGQTITQGIISA 174 (353)
T ss_pred EEEEeCCCC----EEEEEEEEEcCCCCEEEEEEcC--CCCCeeeccCcCcCCCCCEEEEEeCCCCcCCCcceeEEEe
Confidence 999998854 7899999999999999999984 4689999999999999999999999999999999999984
No 4
>PRK10942 serine endoprotease; Provisional
Probab=99.95 E-value=1.6e-26 Score=215.79 Aligned_cols=141 Identities=33% Similarity=0.470 Sum_probs=114.3
Q ss_pred HHHHHHHHhCCceEEEEeeeeccCC---C-CCcchhhhh---c---c-------------------ccccccceEEEEEE
Q 027022 77 RVVQLFQETSPSVVSIQDLELSKNP---K-STSSELMLV---D---G-------------------EYAKVEGTGSGFVW 127 (229)
Q Consensus 77 ~~~~~~~~~~psVV~I~~~~~~~~~---~-~~~~~~~~~---~---~-------------------~~~~~~~~GSGfiI 127 (229)
++.++++++.||||+|++......+ . .....||+. . + ......+.||||||
T Consensus 39 ~~~~~~~~~~pavv~i~~~~~~~~~~~~~~~~~~~ff~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSG~ii 118 (473)
T PRK10942 39 SLAPMLEKVMPSVVSINVEGSTTVNTPRMPRQFQQFFGDNSPFCQEGSPFQSSPFCQGGQGGNGGGQQQKFMALGSGVII 118 (473)
T ss_pred cHHHHHHHhCCceEEEEEEEeccccCCCCChhHHHhhcccccccccccccccccccccccccccccccccccceEEEEEE
Confidence 5999999999999999986644321 0 011222211 0 0 01123578999999
Q ss_pred cC-CcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCC
Q 027022 128 DK-FGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVG 206 (229)
Q Consensus 128 ~~-~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G 206 (229)
++ +||||||+|||. +++.++|++.|+. .|.|++++.|+.+||||||++. ..++++++|++++.+++|
T Consensus 119 ~~~~G~IlTn~HVv~-------~a~~i~V~~~dg~----~~~a~vv~~D~~~DlAvlki~~-~~~l~~~~lg~s~~l~~G 186 (473)
T PRK10942 119 DADKGYVVTNNHVVD-------NATKIKVQLSDGR----KFDAKVVGKDPRSDIALIQLQN-PKNLTAIKMADSDALRVG 186 (473)
T ss_pred ECCCCEEEeChhhcC-------CCCEEEEEECCCC----EEEEEEEEecCCCCEEEEEecC-CCCCceeEecCccccCCC
Confidence 86 599999999999 8899999998754 7899999999999999999974 357999999999999999
Q ss_pred CeEEEEecCCCCCCceeEeEEcC
Q 027022 207 QSCFAIGNPYGFEDTLTTGVTFQ 229 (229)
Q Consensus 207 ~~V~aiG~P~G~~~svt~GiVSa 229 (229)
|+|++||||+|+..+++.||||+
T Consensus 187 ~~V~aiG~P~g~~~tvt~GiVs~ 209 (473)
T PRK10942 187 DYTVAIGNPYGLGETVTSGIVSA 209 (473)
T ss_pred CEEEEEcCCCCCCcceeEEEEEE
Confidence 99999999999999999999985
No 5
>TIGR02037 degP_htrA_DO periplasmic serine protease, Do/DeqQ family. This family consists of a set proteins various designated DegP, heat shock protein HtrA, and protease DO. The ortholog in Pseudomonas aeruginosa is designated MucD and is found in an operon that controls mucoid phenotype. This family also includes the DegQ (HhoA) paralog in E. coli which can rescue a DegP mutant, but not the smaller DegS paralog, which cannot. Members of this family are located in the periplasm and have separable functions as both protease and chaperone. Members have a trypsin domain and two copies of a PDZ domain. This protein protects bacteria from thermal and other stresses and may be important for the survival of bacterial pathogens.// The chaperone function is dominant at low temperatures, whereas the proteolytic activity is turned on at elevated temperatures.
Probab=99.93 E-value=4.4e-25 Score=204.06 Aligned_cols=140 Identities=36% Similarity=0.514 Sum_probs=114.6
Q ss_pred HHHHHHHhCCceEEEEeeeeccCCCC---C---cchhhhh-c------cccccccceEEEEEEcCCcEEEEccccccccc
Q 027022 78 VVQLFQETSPSVVSIQDLELSKNPKS---T---SSELMLV-D------GEYAKVEGTGSGFVWDKFGHIVTNYHVVAKLA 144 (229)
Q Consensus 78 ~~~~~~~~~psVV~I~~~~~~~~~~~---~---~~~~~~~-~------~~~~~~~~~GSGfiI~~~G~IlTn~HVv~~~~ 144 (229)
+.++++++.||||.|.+......... . ...||+. . .......+.||||+|+++||||||+|||.
T Consensus 3 ~~~~~~~~~p~vv~i~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfii~~~G~IlTn~Hvv~--- 79 (428)
T TIGR02037 3 FAPLVEKVAPAVVNISVEGTVKRRNRPPALPPFFRQFFGDDMPNFPRQQRERKVRGLGSGVIISADGYILTNNHVVD--- 79 (428)
T ss_pred HHHHHHHhCCceEEEEEEEEecccCCCcccchhHHHhhcccccCcccccccccccceeeEEEECCCCEEEEcHHHcC---
Confidence 67999999999999998664432110 0 1122211 0 01234568899999999999999999999
Q ss_pred cCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeE
Q 027022 145 TDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTT 224 (229)
Q Consensus 145 ~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~ 224 (229)
+++.+.|++.+++ .|.|++++.|+.+||||||++.. .++++++|++++.+++||+|+++|||+|+..+++.
T Consensus 80 ----~~~~i~V~~~~~~----~~~a~vv~~d~~~DlAllkv~~~-~~~~~~~l~~~~~~~~G~~v~aiG~p~g~~~~~t~ 150 (428)
T TIGR02037 80 ----GADEITVTLSDGR----EFKAKLVGKDPRTDIAVLKIDAK-KNLPVIKLGDSDKLRVGDWVLAIGNPFGLGQTVTS 150 (428)
T ss_pred ----CCCeEEEEeCCCC----EEEEEEEEecCCCCEEEEEecCC-CCceEEEccCCCCCCCCCEEEEEECCCcCCCcEEE
Confidence 8889999998754 78999999999999999999853 57999999999999999999999999999999999
Q ss_pred eEEcC
Q 027022 225 GVTFQ 229 (229)
Q Consensus 225 GiVSa 229 (229)
|+||+
T Consensus 151 G~vs~ 155 (428)
T TIGR02037 151 GIVSA 155 (428)
T ss_pred EEEEe
Confidence 99984
No 6
>COG0265 DegQ Trypsin-like serine proteases, typically periplasmic, contain C-terminal PDZ domain [Posttranslational modification, protein turnover, chaperones]
Probab=99.85 E-value=1.6e-20 Score=169.09 Aligned_cols=137 Identities=41% Similarity=0.574 Sum_probs=113.5
Q ss_pred hHHHHHHHHhCCceEEEEeeeeccCCCCCcchhhhhccccccccceEEEEEEcCCcEEEEccccccccccCCCCcceEEE
Q 027022 76 DRVVQLFQETSPSVVSIQDLELSKNPKSTSSELMLVDGEYAKVEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKV 155 (229)
Q Consensus 76 ~~~~~~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V 155 (229)
..+.++++++.|+||++........ ..|+..........+.||||+++++|||+||+|||. +++++.+
T Consensus 33 ~~~~~~~~~~~~~vV~~~~~~~~~~-----~~~~~~~~~~~~~~~~gSg~i~~~~g~ivTn~hVi~-------~a~~i~v 100 (347)
T COG0265 33 LSFATAVEKVAPAVVSIATGLTAKL-----RSFFPSDPPLRSAEGLGSGFIISSDGYIVTNNHVIA-------GAEEITV 100 (347)
T ss_pred cCHHHHHHhcCCcEEEEEeeeeecc-----hhcccCCcccccccccccEEEEcCCeEEEecceecC-------CcceEEE
Confidence 5789999999999999998665431 111100000011158999999999999999999999 8899999
Q ss_pred EEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEcC
Q 027022 156 SLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTFQ 229 (229)
Q Consensus 156 ~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVSa 229 (229)
.+.| |+ ++.++++++|+..|||+||++.... ++.+.++++..+++||++++||||+|+..+++.||||+
T Consensus 101 ~l~d--g~--~~~a~~vg~d~~~dlavlki~~~~~-~~~~~~~~s~~l~vg~~v~aiGnp~g~~~tvt~Givs~ 169 (347)
T COG0265 101 TLAD--GR--EVPAKLVGKDPISDLAVLKIDGAGG-LPVIALGDSDKLRVGDVVVAIGNPFGLGQTVTSGIVSA 169 (347)
T ss_pred EeCC--CC--EEEEEEEecCCccCEEEEEeccCCC-CceeeccCCCCcccCCEEEEecCCCCcccceeccEEec
Confidence 9966 44 8899999999999999999986433 89999999999999999999999999999999999985
No 7
>KOG1320 consensus Serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=99.61 E-value=2.6e-15 Score=138.57 Aligned_cols=148 Identities=32% Similarity=0.308 Sum_probs=117.0
Q ss_pred ccchhHHHHHHHHhCCceEEEEeeeeccCCCCCcchhhhhccccccccceEEEEEEcCCcEEEEccccccccccCCC---
Q 027022 72 QLEEDRVVQLFQETSPSVVSIQDLELSKNPKSTSSELMLVDGEYAKVEGTGSGFVWDKFGHIVTNYHVVAKLATDTS--- 148 (229)
Q Consensus 72 ~~~~~~~~~~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~--- 148 (229)
......++++++++.+|+|.|+...--. + +..+....-....||||||+.+|+|+||+||+........
T Consensus 124 ~k~~~~v~~~~~~cd~Avv~Ie~~~f~~------~--~~~~e~~~ip~l~~S~~Vv~gd~i~VTnghV~~~~~~~y~~~~ 195 (473)
T KOG1320|consen 124 RKYKAFVAAVFEECDLAVVYIESEEFWK------G--MNPFELGDIPSLNGSGFVVGGDGIIVTNGHVVRVEPRIYAHSS 195 (473)
T ss_pred hhhhhhHHHhhhcccceEEEEeeccccC------C--CcccccCCCcccCccEEEEcCCcEEEEeeEEEEEEeccccCCC
Confidence 4556778899999999999999633211 0 1112223445678999999999999999999996433211
Q ss_pred -CcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEE
Q 027022 149 -GLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVT 227 (229)
Q Consensus 149 -~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiV 227 (229)
..-.+.|...++.|+ .+.+.+++.|+..|+|+++++.++..+++++++-+..++.|+++.++|+||++.++++.|+|
T Consensus 196 ~~l~~vqi~aa~~~~~--s~ep~i~g~d~~~gvA~l~ik~~~~i~~~i~~~~~~~~~~G~~~~a~~~~f~~~nt~t~g~v 273 (473)
T KOG1320|consen 196 TVLLRVQIDAAIGPGN--SGEPVIVGVDKVAGVAFLKIKTPENILYVIPLGVSSHFRTGVEVSAIGNGFGLLNTLTQGMV 273 (473)
T ss_pred cceeeEEEEEeecCCc--cCCCeEEccccccceEEEEEecCCcccceeecceeeeecccceeeccccCceeeeeeeeccc
Confidence 112477888887666 67999999999999999999866544899999999999999999999999999999999999
Q ss_pred cC
Q 027022 228 FQ 229 (229)
Q Consensus 228 Sa 229 (229)
|+
T Consensus 274 s~ 275 (473)
T KOG1320|consen 274 SG 275 (473)
T ss_pred cc
Confidence 74
No 8
>PF13365 Trypsin_2: Trypsin-like peptidase domain; PDB: 1Y8T_A 2Z9I_A 3QO6_A 1L1J_A 1QY6_A 2O8L_A 3OTP_E 2ZLE_I 1KY9_A 3CS0_A ....
Probab=99.09 E-value=3.5e-10 Score=85.18 Aligned_cols=61 Identities=36% Similarity=0.489 Sum_probs=46.7
Q ss_pred EEEEEEcCCcEEEEccccccccccCC-CCcceEEEEEecCCCCeeEEe--EEEEEEcCC-CcEEEEEEc
Q 027022 122 GSGFVWDKFGHIVTNYHVVAKLATDT-SGLHRCKVSLFDAKGNGFYRE--GKMVGCDPA-YDLAVLKVD 186 (229)
Q Consensus 122 GSGfiI~~~G~IlTn~HVv~~~~~~~-~~~~~~~V~~~~~~g~~~~~~--A~vv~~d~~-~DlAvLki~ 186 (229)
||||+|+++||||||+||+.+..... .....+.+...++ . .+. +++++.|+. .|+|||+++
T Consensus 1 GTGf~i~~~g~ilT~~Hvv~~~~~~~~~~~~~~~~~~~~~--~--~~~~~~~~~~~~~~~~D~All~v~ 65 (120)
T PF13365_consen 1 GTGFLIGPDGYILTAAHVVEDWNDGKQPDNSSVEVVFPDG--R--RVPPVAEVVYFDPDDYDLALLKVD 65 (120)
T ss_dssp EEEEEEETTTEEEEEHHHHTCCTT--G-TCSEEEEEETTS--C--EEETEEEEEEEETT-TTEEEEEES
T ss_pred CEEEEEcCCceEEEchhheecccccccCCCCEEEEEecCC--C--EEeeeEEEEEECCccccEEEEEEe
Confidence 89999999999999999999542110 0234566666554 3 556 999999999 999999997
No 9
>KOG1421 consensus Predicted signaling-associated protein (contains a PDZ domain) [General function prediction only]
Probab=98.73 E-value=2.9e-08 Score=94.44 Aligned_cols=131 Identities=25% Similarity=0.253 Sum_probs=98.5
Q ss_pred hhHHHHHHHHhCCceEEEEeeeeccCCCCCcchhhhhccccccccceEEEEEEcC-CcEEEEccccccccccCCCCcceE
Q 027022 75 EDRVVQLFQETSPSVVSIQDLELSKNPKSTSSELMLVDGEYAKVEGTGSGFVWDK-FGHIVTNYHVVAKLATDTSGLHRC 153 (229)
Q Consensus 75 ~~~~~~~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfiI~~-~G~IlTn~HVv~~~~~~~~~~~~~ 153 (229)
...+...+.++-++||.|+..+... .+......+.++||++++ .||||||+||+.. +.-..
T Consensus 51 ~e~w~~~ia~VvksvVsI~~S~v~~------------fdtesag~~~atgfvvd~~~gyiLtnrhvv~p------gP~va 112 (955)
T KOG1421|consen 51 SEDWRNTIANVVKSVVSIRFSAVRA------------FDTESAGESEATGFVVDKKLGYILTNRHVVAP------GPFVA 112 (955)
T ss_pred hhhhhhhhhhhcccEEEEEehheee------------cccccccccceeEEEEecccceEEEeccccCC------CCcee
Confidence 3488999999999999999866543 223344677899999997 5999999999995 44455
Q ss_pred EEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCC---CccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEc
Q 027022 154 KVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGF---ELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTF 228 (229)
Q Consensus 154 ~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~---~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVS 228 (229)
.+.+.+.. ..+--.++.|+-+|+.+++.+.... .+..+.+. .+.-++|.+++.+||-.|...++-.|.+|
T Consensus 113 ~avf~n~e----e~ei~pvyrDpVhdfGf~r~dps~ir~s~vt~i~la-p~~akvgseirvvgNDagEklsIlagflS 185 (955)
T KOG1421|consen 113 SAVFDNHE----EIEIYPVYRDPVHDFGFFRYDPSTIRFSIVTEICLA-PELAKVGSEIRVVGNDAGEKLSILAGFLS 185 (955)
T ss_pred EEEecccc----cCCcccccCCchhhcceeecChhhcceeeeeccccC-ccccccCCceEEecCCccceEEeehhhhh
Confidence 66665433 4566778999999999999984322 24555554 34579999999999988877777777655
No 10
>PF00089 Trypsin: Trypsin; InterPro: IPR001254 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine proteases belong to the MEROPS peptidase family S1 (chymotrypsin family, clan PA(S))and to peptidase family S6 (Hap serine peptidases). The chymotrypsin family is almost totally confined to animals, although trypsin-like enzymes are found in actinomycetes of the genera Streptomyces and Saccharopolyspora, and in the fungus Fusarium oxysporum []. The enzymes are inherently secreted, being synthesised with a signal peptide that targets them to the secretory pathway. Animal enzymes are either secreted directly, packaged into vesicles for regulated secretion, or are retained in leukocyte granules []. The Hap family, 'Haemophilus adhesion and penetration', are proteins that play a role in the interaction with human epithelial cells. The serine protease activity is localized at the N-terminal domain, whereas the binding domain is in the C-terminal region. ; GO: 0004252 serine-type endopeptidase activity, 0006508 proteolysis; PDB: 1SPJ_A 1A5I_A 2ZGH_A 2ZKS_A 2ZGJ_A 2ZGC_A 2ODP_A 2I6Q_A 2I6S_A 2ODQ_A ....
Probab=98.17 E-value=5.6e-05 Score=62.07 Aligned_cols=92 Identities=27% Similarity=0.344 Sum_probs=63.0
Q ss_pred cceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEec-----CCCCeeEEeEEEEE----EcC---CCcEEEEEEc
Q 027022 119 EGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFD-----AKGNGFYREGKMVG----CDP---AYDLAVLKVD 186 (229)
Q Consensus 119 ~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~-----~~g~~~~~~A~vv~----~d~---~~DlAvLki~ 186 (229)
...++|++|+++ +|||++|++. +.+++.+.+.. .++....+..+-+. ++. ..|+|||+++
T Consensus 24 ~~~C~G~li~~~-~vLTaahC~~-------~~~~~~v~~g~~~~~~~~~~~~~~~v~~~~~h~~~~~~~~~~DiAll~L~ 95 (220)
T PF00089_consen 24 RFFCTGTLISPR-WVLTAAHCVD-------GASDIKVRLGTYSIRNSDGSEQTIKVSKIIIHPKYDPSTYDNDIALLKLD 95 (220)
T ss_dssp EEEEEEEEEETT-EEEEEGGGHT-------SGGSEEEEESESBTTSTTTTSEEEEEEEEEEETTSBTTTTTTSEEEEEES
T ss_pred CeeEeEEecccc-cccccccccc-------cccccccccccccccccccccccccccccccccccccccccccccccccc
Confidence 567999999984 9999999999 54556664432 12211123332222 233 4699999998
Q ss_pred cC---CCCccceEcCCC-CCCCCCCeEEEEecCCCC
Q 027022 187 VE---GFELKPVVLGTS-HDLRVGQSCFAIGNPYGF 218 (229)
Q Consensus 187 ~~---~~~~~~l~lg~s-~~~~~G~~V~aiG~P~G~ 218 (229)
.+ ...+.++.+... ..++.|+.+.++|++...
T Consensus 96 ~~~~~~~~~~~~~l~~~~~~~~~~~~~~~~G~~~~~ 131 (220)
T PF00089_consen 96 RPITFGDNIQPICLPSAGSDPNVGTSCIVVGWGRTS 131 (220)
T ss_dssp SSSEHBSSBEESBBTSTTHTTTTTSEEEEEESSBSS
T ss_pred cccccccccccccccccccccccccccccccccccc
Confidence 65 345677778762 347999999999999863
No 11
>cd00190 Tryp_SPc Trypsin-like serine protease; Many of these are synthesized as inactive precursor zymogens that are cleaved during limited proteolysis to generate their active forms. Alignment contains also inactive enzymes that have substitutions of the catalytic triad residues.
Probab=97.98 E-value=0.00023 Score=58.81 Aligned_cols=95 Identities=21% Similarity=0.155 Sum_probs=62.4
Q ss_pred ccceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCC-----eeEEeEEEEEEc-------CCCcEEEEEE
Q 027022 118 VEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGN-----GFYREGKMVGCD-------PAYDLAVLKV 185 (229)
Q Consensus 118 ~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~-----~~~~~A~vv~~d-------~~~DlAvLki 185 (229)
.....+|++|++ .+|||++|.+.+. ....+.|.+...+.. ...+..+-+... ...|||||++
T Consensus 23 ~~~~C~GtlIs~-~~VLTaAhC~~~~-----~~~~~~v~~g~~~~~~~~~~~~~~~v~~~~~hp~y~~~~~~~DiAll~L 96 (232)
T cd00190 23 GRHFCGGSLISP-RWVLTAAHCVYSS-----APSNYTVRLGSHDLSSNEGGGQVIKVKKVIVHPNYNPSTYDNDIALLKL 96 (232)
T ss_pred CcEEEEEEEeeC-CEEEECHHhcCCC-----CCccEEEEeCcccccCCCCceEEEEEEEEEECCCCCCCCCcCCEEEEEE
Confidence 346799999997 6999999999842 124566665322211 112233333332 3579999999
Q ss_pred ccCC---CCccceEcCCCC-CCCCCCeEEEEecCCCC
Q 027022 186 DVEG---FELKPVVLGTSH-DLRVGQSCFAIGNPYGF 218 (229)
Q Consensus 186 ~~~~---~~~~~l~lg~s~-~~~~G~~V~aiG~P~G~ 218 (229)
+.+- ..+.|+.|.... .+..|+.+++.|+....
T Consensus 97 ~~~~~~~~~v~picl~~~~~~~~~~~~~~~~G~g~~~ 133 (232)
T cd00190 97 KRPVTLSDNVRPICLPSSGYNLPAGTTCTVSGWGRTS 133 (232)
T ss_pred CCcccCCCcccceECCCccccCCCCCEEEEEeCCcCC
Confidence 8542 226778886654 68899999999987653
No 12
>smart00020 Tryp_SPc Trypsin-like serine protease. Many of these are synthesised as inactive precursor zymogens that are cleaved during limited proteolysis to generate their active forms. A few, however, are active as single chain molecules, and others are inactive due to substitutions of the catalytic triad residues.
Probab=97.90 E-value=0.00014 Score=60.28 Aligned_cols=94 Identities=18% Similarity=0.142 Sum_probs=62.9
Q ss_pred cceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCe----eEEeEEEEE-------EcCCCcEEEEEEcc
Q 027022 119 EGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNG----FYREGKMVG-------CDPAYDLAVLKVDV 187 (229)
Q Consensus 119 ~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~----~~~~A~vv~-------~d~~~DlAvLki~~ 187 (229)
....+|.+|++ .+|||++|.+.+. ....+.|.+...+... ..+...-+. .....|||||+++.
T Consensus 25 ~~~C~GtlIs~-~~VLTaahC~~~~-----~~~~~~v~~g~~~~~~~~~~~~~~v~~~~~~p~~~~~~~~~DiAll~L~~ 98 (229)
T smart00020 25 RHFCGGSLISP-RWVLTAAHCVYGS-----DPSNIRVRLGSHDLSSGEEGQVIKVSKVIIHPNYNPSTYDNDIALLKLKS 98 (229)
T ss_pred CcEEEEEEecC-CEEEECHHHcCCC-----CCcceEEEeCcccCCCCCCceEEeeEEEEECCCCCCCCCcCCEEEEEECc
Confidence 56799999997 6999999999942 1246677775432211 123333333 23467999999985
Q ss_pred C---CCCccceEcCCC-CCCCCCCeEEEEecCCCC
Q 027022 188 E---GFELKPVVLGTS-HDLRVGQSCFAIGNPYGF 218 (229)
Q Consensus 188 ~---~~~~~~l~lg~s-~~~~~G~~V~aiG~P~G~ 218 (229)
+ ...+.++.|... ..+..|+.+.+.|++...
T Consensus 99 ~i~~~~~~~pi~l~~~~~~~~~~~~~~~~g~g~~~ 133 (229)
T smart00020 99 PVTLSDNVRPICLPSSNYNVPAGTTCTVSGWGRTS 133 (229)
T ss_pred ccCCCCceeeccCCCcccccCCCCEEEEEeCCCCC
Confidence 4 123667777553 367889999999987654
No 13
>PF10459 Peptidase_S46: Peptidase S46; InterPro: IPR019500 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This entry represents S46 peptidases, where dipeptidyl-peptidase 7 (DPP-7) is the best-characterised member of this family. It is a serine peptidase that is located on the cell surface and is predicted to have two N-terminal transmembrane domains.
Probab=97.65 E-value=0.00036 Score=68.55 Aligned_cols=24 Identities=29% Similarity=0.201 Sum_probs=21.3
Q ss_pred ceEEEEEEcCCcEEEEcccccccc
Q 027022 120 GTGSGFVWDKFGHIVTNYHVVAKL 143 (229)
Q Consensus 120 ~~GSGfiI~~~G~IlTn~HVv~~~ 143 (229)
+.+||-||+++|+|+||+|++-+.
T Consensus 47 gGCSgsfVS~~GLvlTNHHC~~~~ 70 (698)
T PF10459_consen 47 GGCSGSFVSPDGLVLTNHHCGYGA 70 (698)
T ss_pred CceeEEEEcCCceEEecchhhhhH
Confidence 359999999999999999998744
No 14
>KOG1320 consensus Serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=97.54 E-value=7.6e-05 Score=69.70 Aligned_cols=128 Identities=23% Similarity=0.217 Sum_probs=89.7
Q ss_pred HHHHhCCceEEEEeeeeccCCCCCcchhhhhccccccccceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecC
Q 027022 81 LFQETSPSVVSIQDLELSKNPKSTSSELMLVDGEYAKVEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDA 160 (229)
Q Consensus 81 ~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~ 160 (229)
.++....+++.+....++. .....|....+....||||.+.. ..++||+|++.... +.....+. ..
T Consensus 55 ~~~~~~~s~~~v~~~~~~~-------~~~~pw~~~~q~~~~~s~f~i~~-~~lltn~~~v~~~~----~~~~v~v~-~~- 120 (473)
T KOG1320|consen 55 VVDLALQSVVKVFSVSTEP-------SSVLPWQRTRQFSSGGSGFAIYG-KKLLTNAHVVAPNN----DHKFVTVK-KH- 120 (473)
T ss_pred CccccccceeEEEeecccc-------cccCcceeeehhcccccchhhcc-cceeecCccccccc----cccccccc-cC-
Confidence 3455566888888755443 22222444457788999999986 68999999999432 23333333 33
Q ss_pred CCCeeEEeEEEEEEcCCCcEEEEEEccC--CCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEc
Q 027022 161 KGNGFYREGKMVGCDPAYDLAVLKVDVE--GFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTF 228 (229)
Q Consensus 161 ~g~~~~~~A~vv~~d~~~DlAvLki~~~--~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVS 228 (229)
|....|.|++...-.+.|+|+|.++.. +....++.+++ -+.-.+-++++| |...++|-|.|+
T Consensus 121 -gs~~k~~~~v~~~~~~cd~Avv~Ie~~~f~~~~~~~e~~~--ip~l~~S~~Vv~---gd~i~VTnghV~ 184 (473)
T KOG1320|consen 121 -GSPRKYKAFVAAVFEECDLAVVYIESEEFWKGMNPFELGD--IPSLNGSGFVVG---GDGIIVTNGHVV 184 (473)
T ss_pred -CCchhhhhhHHHhhhcccceEEEEeeccccCCCcccccCC--CcccCccEEEEc---CCcEEEEeeEEE
Confidence 444478899999999999999999853 33333455544 577888899999 888899999986
No 15
>COG3591 V8-like Glu-specific endopeptidase [Amino acid transport and metabolism]
Probab=95.56 E-value=0.14 Score=44.54 Aligned_cols=94 Identities=13% Similarity=0.053 Sum_probs=51.4
Q ss_pred EEEEEEcCCcEEEEccccccccccCCCCcceEEEEEe--cCCCC-eeEEeEEEEEEcC----CCcEEEEEEccCC-----
Q 027022 122 GSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLF--DAKGN-GFYREGKMVGCDP----AYDLAVLKVDVEG----- 189 (229)
Q Consensus 122 GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~--~~~g~-~~~~~A~vv~~d~----~~DlAvLki~~~~----- 189 (229)
.++|+|++ ..||||.|++..... +-.++.+..+ .++|. .+.+........+ +.|.+...+..-.
T Consensus 66 ~~~~lI~p-ntvLTa~Hc~~s~~~---G~~~~~~~p~g~~~~~~~~~~~~~~~~~~~~g~~~~~d~~~~~v~~~~~~~g~ 141 (251)
T COG3591 66 TAATLIGP-NTVLTAGHCIYSPDY---GEDDIAAAPPGVNSDGGPFYGITKIEIRVYPGELYKEDGASYDVGEAALESGI 141 (251)
T ss_pred eeEEEEcC-ceEEEeeeEEecCCC---ChhhhhhcCCcccCCCCCCCceeeEEEEecCCceeccCCceeeccHHHhccCC
Confidence 45699998 699999999995431 1122222221 11222 1222222222222 3466666653211
Q ss_pred ---CCccceEcCCCCCCCCCCeEEEEecCCCCC
Q 027022 190 ---FELKPVVLGTSHDLRVGQSCFAIGNPYGFE 219 (229)
Q Consensus 190 ---~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~ 219 (229)
.......+......+.+|.+-.+|||.+..
T Consensus 142 ~~~~~~~~~~~~~~~~~~~~d~i~v~GYP~dk~ 174 (251)
T COG3591 142 NIGDVVNYLKRNTASEAKANDRITVIGYPGDKP 174 (251)
T ss_pred CccccccccccccccccccCceeEEEeccCCCC
Confidence 112222333456799999999999998855
No 16
>PF00863 Peptidase_C4: Peptidase family C4 This family belongs to family C4 of the peptidase classification.; InterPro: IPR001730 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. Nuclear inclusion A (NIA) proteases from potyviruses are cysteine peptidases belong to the MEROPS peptidase family C4 (NIa protease family, clan PA(C)) [, ]. Potyviruses include plant viruses in which the single-stranded RNA encodes a polyprotein with NIA protease activity, where proteolytic cleavage is specific for Gln+Gly sites. The NIA protease acts on the polyprotein, releasing itself by Gln+Gly cleavage at both the N- and C-termini. It further processes the polyprotein by cleavage at five similar sites in the C-terminal half of the sequence. In addition to its C-terminal protease activity, the NIA protease contains an N-terminal domain that has been implicated in the transcription process []. This peptidase is present in the nuclear inclusion protein of potyviruses.; GO: 0008234 cysteine-type peptidase activity, 0006508 proteolysis; PDB: 3MMG_B 1Q31_B 1LVB_A 1LVM_A.
Probab=94.79 E-value=0.29 Score=42.11 Aligned_cols=82 Identities=13% Similarity=0.173 Sum_probs=46.7
Q ss_pred ceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEE-----EEEEcCCCcEEEEEEccCCCCccc
Q 027022 120 GTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGK-----MVGCDPAYDLAVLKVDVEGFELKP 194 (229)
Q Consensus 120 ~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~-----vv~~d~~~DlAvLki~~~~~~~~~ 194 (229)
..-=||..++ |||||+|..+. +...+.|.... | .|... -+..=+..||.++|.. .++||
T Consensus 32 ~~l~gigyG~--~iItn~HLf~~------nng~L~i~s~h--G---~f~v~nt~~lkv~~i~~~DiviirmP---kDfpP 95 (235)
T PF00863_consen 32 RSLYGIGYGS--YIITNAHLFKR------NNGELTIKSQH--G---EFTVPNTTQLKVHPIEGRDIVIIRMP---KDFPP 95 (235)
T ss_dssp EEEEEEEETT--EEEEEGGGGSS------TTCEEEEEETT--E---EEEECEGGGSEEEE-TCSSEEEEE-----TTS--
T ss_pred EEEEEEeECC--EEEEChhhhcc------CCCeEEEEeCc--e---EEEcCCccccceEEeCCccEEEEeCC---cccCC
Confidence 3445677765 99999999985 33457776655 3 23222 2445568999999995 34666
Q ss_pred eEcC-CCCCCCCCCeEEEEecCCC
Q 027022 195 VVLG-TSHDLRVGQSCFAIGNPYG 217 (229)
Q Consensus 195 l~lg-~s~~~~~G~~V~aiG~P~G 217 (229)
.+-- ....++.||.|..+|.=+-
T Consensus 96 f~~kl~FR~P~~~e~v~mVg~~fq 119 (235)
T PF00863_consen 96 FPQKLKFRAPKEGERVCMVGSNFQ 119 (235)
T ss_dssp --S---B----TT-EEEEEEEECS
T ss_pred cchhhhccCCCCCCEEEEEEEEEE
Confidence 5431 2457999999999998554
No 17
>KOG1421 consensus Predicted signaling-associated protein (contains a PDZ domain) [General function prediction only]
Probab=92.47 E-value=0.27 Score=48.11 Aligned_cols=87 Identities=18% Similarity=0.128 Sum_probs=71.2
Q ss_pred ccceEEEEEEcC-CcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceE
Q 027022 118 VEGTGSGFVWDK-FGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVV 196 (229)
Q Consensus 118 ~~~~GSGfiI~~-~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~ 196 (229)
....|||.|++. .|+++++..+|.- +.....|++.|.. ...|++..-++...+|.+|.+.. ..-.++
T Consensus 548 ~i~kgt~~i~d~~~g~~vvsr~~vp~------d~~d~~vt~~dS~----~i~a~~~fL~~t~n~a~~kydp~--~~~~~k 615 (955)
T KOG1421|consen 548 DIYKGTALIMDTSKGLGVVSRSVVPS------DAKDQRVTEADSD----GIPANVSFLHPTENVASFKYDPA--LEVQLK 615 (955)
T ss_pred hhhcCceEEEEccCCceeEecccCCc------hhhceEEeecccc----cccceeeEecCccceeEeccChh--Hhhhhc
Confidence 456799999986 5999999999984 7888999998876 56899999999999999999843 234556
Q ss_pred cCCCCCCCCCCeEEEEecCCC
Q 027022 197 LGTSHDLRVGQSCFAIGNPYG 217 (229)
Q Consensus 197 lg~s~~~~~G~~V~aiG~P~G 217 (229)
|- ...++.||++-.+|+-..
T Consensus 616 l~-~~~v~~gD~~~f~g~~~~ 635 (955)
T KOG1421|consen 616 LT-DTTVLRGDECTFEGFTED 635 (955)
T ss_pred cc-eeeEecCCceeEeccccc
Confidence 63 456999999999998755
No 18
>PF05579 Peptidase_S32: Equine arteritis virus serine endopeptidase S32; InterPro: IPR008760 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to MEROPS peptidase family S32 (clan PA(S)). The type example is equine arteritis virus serine endopeptidase (equine arteritis virus), which is involved in processing of nidovirus polyproteins [].; GO: 0004252 serine-type endopeptidase activity, 0016032 viral reproduction, 0019082 viral protein processing; PDB: 3FAN_A 3FAO_A 1MBM_A.
Probab=88.18 E-value=2.1 Score=37.54 Aligned_cols=64 Identities=19% Similarity=0.150 Sum_probs=37.5
Q ss_pred ceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCC
Q 027022 120 GTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGT 199 (229)
Q Consensus 120 ~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~ 199 (229)
++|+.|-++.+-.|+|+.||+.+ +...|...+ . -+...++.+-|.|.-.++.-...+|.+++++
T Consensus 114 Gsggvft~~~~~vvvTAtHVlg~--------~~a~v~~~g---~-----~~~~tF~~~GDfA~~~~~~~~G~~P~~k~a~ 177 (297)
T PF05579_consen 114 GSGGVFTIGGNTVVVTATHVLGG--------NTARVSGVG---T-----RRMLTFKKNGDFAEADITNWPGAAPKYKFAQ 177 (297)
T ss_dssp EEEEEEECTTEEEEEEEHHHCBT--------TEEEEEETT---E-----EEEEEEEEETTEEEEEETTS-S---B--B-T
T ss_pred cccceEEECCeEEEEEEEEEcCC--------CeEEEEecc---e-----EEEEEEeccCcEEEEECCCCCCCCCceeecC
Confidence 34444545554589999999973 445555422 1 2456678889999999954445788888873
No 19
>PF03510 Peptidase_C24: 2C endopeptidase (C24) cysteine protease family; InterPro: IPR000317 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. The two signatures that defines this group of calivirus polyproteins identify a cysteine peptidase signature that belongs to MEROPS peptidase family C24 (clan PA(C)). Caliciviruses are positive-stranded ssRNA viruses that cause gastroenteritis. The calicivirus genome contains two open reading frames, ORF1 and ORF2. ORF2 encodes a structural protein []; while ORF1 encodes a non-structural polypeptide, which has RNA helicase, cysteine protease and RNA polymerase activity. The regions of the polyprotein in which these activities lie are similar to proteins produced by the picornaviruses. Two different families of caliciviruses can be distinguished on the basis of sequence similarity, namely those classified as small round structured viruses (SRSVs) and those classed as non-SRSVs. Calicivirus proteases from the non-SRSV group, which are members of the PA protease clan, constitute family C24 of the cysteine proteases (proteases from SRSVs belong to the C37 family). As mentioned above, the protease activity resides within a polyprotein. The enzyme cleaves the polyprotein at sites N-terminal to itself, liberating the polyprotein helicase.; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis
Probab=82.22 E-value=3.1 Score=31.39 Aligned_cols=55 Identities=24% Similarity=0.270 Sum_probs=36.1
Q ss_pred EEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCC
Q 027022 124 GFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSH 201 (229)
Q Consensus 124 GfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~ 201 (229)
++-|+. |.++|+.||.+ ..+.+ + |.. =+++. .+.|+|+++.+.. .+|.+++++..
T Consensus 3 avHIGn-G~~vt~tHva~-------~~~~v-----~--g~~----f~~~~--~~ge~~~v~~~~~--~~p~~~ig~g~ 57 (105)
T PF03510_consen 3 AVHIGN-GRYVTVTHVAK-------SSDSV-----D--GQP----FKIVK--TDGELCWVQSPLV--HLPAAQIGTGK 57 (105)
T ss_pred eEEeCC-CEEEEEEEEec-------cCceE-----c--CcC----cEEEE--eccCEEEEECCCC--CCCeeEeccCC
Confidence 456764 99999999999 44332 1 332 13333 4559999999743 47888887543
No 20
>KOG3627 consensus Trypsin [Amino acid transport and metabolism]
Probab=79.20 E-value=21 Score=29.99 Aligned_cols=86 Identities=22% Similarity=0.327 Sum_probs=47.6
Q ss_pred eEEEEEEcCCcEEEEccccccccccCCCCcc--eEEEEEecC-------CCC--eeEEeEEEE---EEcC---C-CcEEE
Q 027022 121 TGSGFVWDKFGHIVTNYHVVAKLATDTSGLH--RCKVSLFDA-------KGN--GFYREGKMV---GCDP---A-YDLAV 182 (229)
Q Consensus 121 ~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~--~~~V~~~~~-------~g~--~~~~~A~vv---~~d~---~-~DlAv 182 (229)
...|.+|++ .+|+|++|.+.+ .. ...|.+... .+. ......+++ .++. . .||||
T Consensus 39 ~Cggsli~~-~~vltaaHC~~~-------~~~~~~~V~~G~~~~~~~~~~~~~~~~~~v~~~i~H~~y~~~~~~~nDial 110 (256)
T KOG3627|consen 39 LCGGSLISP-RWVLTAAHCVKG-------ASASLYTVRLGEHDINLSVSEGEEQLVGDVEKIIVHPNYNPRTLENNDIAL 110 (256)
T ss_pred eeeeEEeeC-CEEEEChhhCCC-------CCCcceEEEECccccccccccCchhhhceeeEEEECCCCCCCCCCCCCEEE
Confidence 455667855 599999999994 22 455554210 010 111112333 1222 2 79999
Q ss_pred EEEccC---CCCccceEcCCCCC---CCCCCeEEEEec
Q 027022 183 LKVDVE---GFELKPVVLGTSHD---LRVGQSCFAIGN 214 (229)
Q Consensus 183 Lki~~~---~~~~~~l~lg~s~~---~~~G~~V~aiG~ 214 (229)
|+++.+ ...+.++.|-.... ...++.+++.|.
T Consensus 111 l~l~~~v~~~~~i~piclp~~~~~~~~~~~~~~~v~GW 148 (256)
T KOG3627|consen 111 LRLSEPVTFSSHIQPICLPSSADPYFPPGGTTCLVSGW 148 (256)
T ss_pred EEECCCcccCCcccccCCCCCcccCCCCCCCEEEEEeC
Confidence 999853 13355666632332 444588888884
No 21
>PF09342 DUF1986: Domain of unknown function (DUF1986); InterPro: IPR015420 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This domain is found in serine endopeptidases belonging to MEROPS peptidase family S1A (clan PA). It is found in unusual mosaic proteins, which are encoded by the Drosophila nudel gene (see P98159 from SWISSPROT). Nudel is involved in defining embryonic dorsoventral polarity. Three proteases; ndl, gd and snk process easter to create active easter. Active easter defines cell identities along the dorsal-ventral continuum by activating the spz ligand for the Tl receptor in the ventral region of the embryo. Nudel, pipe and windbeutel together trigger the protease cascade within the extraembryonic perivitelline compartment which induces dorsoventral polarity of the Drosophila embryo [].
Probab=72.68 E-value=65 Score=28.13 Aligned_cols=99 Identities=16% Similarity=0.247 Sum_probs=60.6
Q ss_pred cccceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEe----EEEEEEc-----CCCcEEEEEEcc
Q 027022 117 KVEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYRE----GKMVGCD-----PAYDLAVLKVDV 187 (229)
Q Consensus 117 ~~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~----A~vv~~d-----~~~DlAvLki~~ 187 (229)
.+.-..||++|++ .|||++...+.+... ....+.+.+ |.|+.+.+. -+|...| +.+++.||.++.
T Consensus 25 dG~~~CsgvLlD~-~WlLvsssCl~~I~L---~~~Yvsall--G~~Kt~~~v~Gp~EQI~rVD~~~~V~~S~v~LLHL~~ 98 (267)
T PF09342_consen 25 DGRYWCSGVLLDP-HWLLVSSSCLRGISL---SHHYVSALL--GGGKTYLSVDGPHEQISRVDCFKDVPESNVLLLHLEQ 98 (267)
T ss_pred cCeEEEEEEEecc-ceEEEeccccCCccc---ccceEEEEe--cCcceecccCCChheEEEeeeeeeccccceeeeeecC
Confidence 4567899999998 799999999996532 123444444 445533210 1333333 577999999985
Q ss_pred CCCC----ccceEcCC-CCCCCCCCeEEEEecCCCCCCcee
Q 027022 188 EGFE----LKPVVLGT-SHDLRVGQSCFAIGNPYGFEDTLT 223 (229)
Q Consensus 188 ~~~~----~~~l~lg~-s~~~~~G~~V~aiG~P~G~~~svt 223 (229)
+ .+ ..|+-+-+ +.+....+..+++|.-- .+.+.|
T Consensus 99 ~-~~fTr~VlP~flp~~~~~~~~~~~CVAVg~d~-~g~~kt 137 (267)
T PF09342_consen 99 P-ANFTRYVLPTFLPETSNENESDDECVAVGHDD-TGRIKT 137 (267)
T ss_pred c-ccceeeecccccccccCCCCCCCceEEEEccc-CCceee
Confidence 4 22 22222322 35667777999999754 333333
No 22
>PF01455 HupF_HypC: HupF/HypC family; InterPro: IPR001109 The large subunit of [NiFe]-hydrogenase, as well as other nickel metalloenzymes, is synthesised as a precursor devoid of the metalloenzyme active site. This precursor then undergoes a complex post-translational maturation process that requires a number of accessory proteins. The hydrogenase expression/formation proteins (HupF/HypC) form a family of small proteins that are hydrogenase precursor-specific chaperones required for this maturation process []. They are believed to keep the hydrogenase precursor in a conformation accessible for metal incorporation [, ].; PDB: 3D3R_A 2Z1C_C 2OT2_A.
Probab=64.19 E-value=29 Score=23.91 Aligned_cols=43 Identities=23% Similarity=0.297 Sum_probs=30.6
Q ss_pred EeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEE
Q 027022 167 REGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAI 212 (229)
Q Consensus 167 ~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~ai 212 (229)
.+++++..+.....|++.+.. ....+.+.=-.++++||+|++-
T Consensus 5 iP~~Vv~v~~~~~~A~v~~~G---~~~~V~~~lv~~v~~Gd~VLVH 47 (68)
T PF01455_consen 5 IPGRVVEVDEDGGMAVVDFGG---VRREVSLALVPDVKVGDYVLVH 47 (68)
T ss_dssp EEEEEEEEETTTTEEEEEETT---EEEEEEGTTCTSB-TT-EEEEE
T ss_pred ccEEEEEEeCCCCEEEEEcCC---cEEEEEEEEeCCCCCCCEEEEe
Confidence 578999998889999998863 3455555445569999999873
No 23
>COG0298 HypC Hydrogenase maturation factor [Posttranslational modification, protein turnover, chaperones]
Probab=64.03 E-value=14 Score=26.51 Aligned_cols=46 Identities=24% Similarity=0.376 Sum_probs=31.0
Q ss_pred EeEEEEEEcCCCcEEEEEEccCCCCccceEcCCC-CCCCCCCeEEE-EecC
Q 027022 167 REGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTS-HDLRVGQSCFA-IGNP 215 (229)
Q Consensus 167 ~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s-~~~~~G~~V~a-iG~P 215 (229)
.+++++..|.+.++|++.+-.-. .-+.+.=- ++++.||+|++ +||-
T Consensus 5 iPgqI~~I~~~~~~A~Vd~gGvk---reV~l~Lv~~~v~~GdyVLVHvGfA 52 (82)
T COG0298 5 IPGQIVEIDDNNHLAIVDVGGVK---REVNLDLVGEEVKVGDYVLVHVGFA 52 (82)
T ss_pred cccEEEEEeCCCceEEEEeccEe---EEEEeeeecCccccCCEEEEEeeEE
Confidence 46789999998889999885321 22222212 28999999986 5553
No 24
>cd01735 LSm12_N LSm12 belongs to a family of Sm-like proteins that associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet that associates with other Sm proteins to form hexameric and heptameric ring structures. In addition to the N-terminal Sm-like domain, LSm12 has a novel methyltransferase domain.
Probab=62.11 E-value=32 Score=23.30 Aligned_cols=33 Identities=15% Similarity=0.139 Sum_probs=26.8
Q ss_pred cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022 150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD 186 (229)
Q Consensus 150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~ 186 (229)
+..+.++...++ .+.++|+.+|....+.+|+-.
T Consensus 6 Gs~V~~kTc~g~----~ieGEV~afD~~tk~lIlk~~ 38 (61)
T cd01735 6 GSQVSCRTCFEQ----RLQGEVVAFDYPSKMLILKCP 38 (61)
T ss_pred ccEEEEEecCCc----eEEEEEEEecCCCcEEEEECc
Confidence 346677776655 889999999999999999854
No 25
>PRK13922 rod shape-determining protein MreC; Provisional
Probab=58.68 E-value=1.2e+02 Score=26.23 Aligned_cols=51 Identities=24% Similarity=0.167 Sum_probs=31.2
Q ss_pred CcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEc
Q 027022 178 YDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTF 228 (229)
Q Consensus 178 ~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVS 228 (229)
.+-++++=+.....+..--+....++++||.|+.=|.-.-+...+..|.|+
T Consensus 190 ~~~gi~~G~g~~~~l~l~~i~~~~~i~~GD~VvTSGl~g~fP~Gi~VG~V~ 240 (276)
T PRK13922 190 GIRGILSGNGSGDNLKLEFIPRSADIKVGDLVVTSGLGGIFPAGLPVGKVT 240 (276)
T ss_pred CceEEEEecCCCCceEEEecCCCCCCCCCCEEEECCCCCcCCCCCEEEEEE
Confidence 345666665321122333333456799999999999754456667777765
No 26
>PF01732 DUF31: Putative peptidase (DUF31); InterPro: IPR022382 This domain has no known function. It is found in various hypothetical proteins and putative lipoproteins from mycoplasmas.
Probab=58.28 E-value=7 Score=35.66 Aligned_cols=24 Identities=29% Similarity=0.451 Sum_probs=19.6
Q ss_pred ccceEEEEEEcC----Cc------EEEEcccccc
Q 027022 118 VEGTGSGFVWDK----FG------HIVTNYHVVA 141 (229)
Q Consensus 118 ~~~~GSGfiI~~----~G------~IlTn~HVv~ 141 (229)
....|||.|+|- ++ ||.||.||+.
T Consensus 34 ~~~~GT~WIlDy~~~~~~~~p~k~y~ATNlHVa~ 67 (374)
T PF01732_consen 34 SSVSGTGWILDYKKPEDNKYPTKWYFATNLHVAS 67 (374)
T ss_pred ccCcceEEEEEEeccCCCCCCeEEEEEechhhhc
Confidence 346899999972 23 9999999999
No 27
>COG5640 Secreted trypsin-like serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=54.94 E-value=25 Score=32.33 Aligned_cols=19 Identities=16% Similarity=0.059 Sum_probs=14.1
Q ss_pred EEEEcCCcEEEEcccccccc
Q 027022 124 GFVWDKFGHIVTNYHVVAKL 143 (229)
Q Consensus 124 GfiI~~~G~IlTn~HVv~~~ 143 (229)
|-++..+ ||||++|.+.+.
T Consensus 65 gs~l~~R-YvLTAAHC~~~~ 83 (413)
T COG5640 65 GSKLGGR-YVLTAAHCADAS 83 (413)
T ss_pred cceecce-EEeeehhhccCC
Confidence 3445554 999999999954
No 28
>PRK10672 rare lipoprotein A; Provisional
Probab=54.21 E-value=52 Score=30.19 Aligned_cols=29 Identities=24% Similarity=0.155 Sum_probs=19.4
Q ss_pred hccccccccceEEEEEEcCCcEEEEcccccc
Q 027022 111 VDGEYAKVEGTGSGFVWDKFGHIVTNYHVVA 141 (229)
Q Consensus 111 ~~~~~~~~~~~GSGfiI~~~G~IlTn~HVv~ 141 (229)
|.+.......+.+|=+++. +-+|++|---
T Consensus 85 wYg~~f~G~~TA~Ge~~~~--~~~tAAH~tL 113 (361)
T PRK10672 85 IYDAEAGSNLTASGERFDP--NALTAAHPTL 113 (361)
T ss_pred EeCCccCCCcCcCceeecC--CcCeeeccCC
Confidence 3344445566778888876 5789999544
No 29
>PF02601 Exonuc_VII_L: Exonuclease VII, large subunit; InterPro: IPR020579 Exonuclease VII 3.1.11.6 from EC is composed of two nonidentical subunits; one large subunit and 4 small ones []. Exonuclease VII catalyses exonucleolytic cleavage in either 5'-3' or 3'-5' direction to yield 5'-phosphomononucleotides. The large subunit also contains the OB-fold domains (IPR004365 from INTERPRO) that bind to nucleic acids at the N terminus. This entry represents Exonuclease VII, large subunit, C-terminal. ; GO: 0008855 exodeoxyribonuclease VII activity
Probab=49.29 E-value=21 Score=31.64 Aligned_cols=34 Identities=24% Similarity=0.305 Sum_probs=29.3
Q ss_pred ceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecC
Q 027022 120 GTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDA 160 (229)
Q Consensus 120 ~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~ 160 (229)
..|=.++-+++|.+||+..-++ .++.+.|.+.||
T Consensus 280 ~RGYaiv~~~~g~vI~s~~~l~-------~gd~i~i~l~DG 313 (319)
T PF02601_consen 280 KRGYAIVRDKDGKVITSVKQLK-------PGDEIEIRLADG 313 (319)
T ss_pred hCceEEEECCCCCEECCHHHCC-------CCCEEEEEEcce
Confidence 4466677778999999999999 889999999985
No 30
>COG3338 Cah Carbonic anhydrase [Inorganic ion transport and metabolism]
Probab=48.17 E-value=32 Score=29.66 Aligned_cols=25 Identities=40% Similarity=0.446 Sum_probs=23.1
Q ss_pred CCCeeEEeEEEEEEcCCCcEEEEEE
Q 027022 161 KGNGFYREGKMVGCDPAYDLAVLKV 185 (229)
Q Consensus 161 ~g~~~~~~A~vv~~d~~~DlAvLki 185 (229)
+|+.+.+.|..|..|++.+||||-+
T Consensus 124 ~Gk~~pmEaHFVHkd~~g~L~Vl~v 148 (250)
T COG3338 124 DGKSFPMEAHFVHKDAKGTLAVLAV 148 (250)
T ss_pred ccccccceeeeeecCCCCCEEEEEE
Confidence 6888889999999999999999987
No 31
>PF15436 PGBA_N: Plasminogen-binding protein pgbA N-terminal
Probab=46.05 E-value=1.7e+02 Score=24.99 Aligned_cols=74 Identities=16% Similarity=0.097 Sum_probs=42.3
Q ss_pred cCCcEEEE--ccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCC-CCccceEcCCCCCCC
Q 027022 128 DKFGHIVT--NYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEG-FELKPVVLGTSHDLR 204 (229)
Q Consensus 128 ~~~G~IlT--n~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~-~~~~~l~lg~s~~~~ 204 (229)
+.++.++| +.+..- +...+.+.-.+.+.. ..-|+.+-..-+...|.+|+.... ..-..++... -.++
T Consensus 12 d~~~~~i~~~~~~l~v-------G~SGiV~h~~~~~~~--~IiA~a~V~~~~~g~A~~kf~~fd~L~Q~aLP~p~-~~pk 81 (218)
T PF15436_consen 12 DDNNKIITFDAPDLKV-------GESGIVVHKFDKDHS--SIIARAVVISKKNGVAKAKFSVFDSLKQDALPTPK-MVPK 81 (218)
T ss_pred ecCCCEEEecCCcccc-------CCceEEEEEecCCcc--eeeeEEEEEEecCCeeEEEEeehhhhhhhcCCCCc-cccC
Confidence 44444555 444444 555666665543333 455666555557889999985321 1234444432 3689
Q ss_pred CCCeEEE
Q 027022 205 VGQSCFA 211 (229)
Q Consensus 205 ~G~~V~a 211 (229)
.||.|+.
T Consensus 82 ~GD~vil 88 (218)
T PF15436_consen 82 KGDEVIL 88 (218)
T ss_pred CCCEEEE
Confidence 9999873
No 32
>PRK10413 hydrogenase 2 accessory protein HypG; Provisional
Probab=43.89 E-value=50 Score=23.66 Aligned_cols=45 Identities=13% Similarity=0.112 Sum_probs=26.4
Q ss_pred EeEEEEEEcCCC-cEEEEEEccCCCCccceEcCCC-CCCCCCCeEEE
Q 027022 167 REGKMVGCDPAY-DLAVLKVDVEGFELKPVVLGTS-HDLRVGQSCFA 211 (229)
Q Consensus 167 ~~A~vv~~d~~~-DlAvLki~~~~~~~~~l~lg~s-~~~~~G~~V~a 211 (229)
.+++++..+.+. .+|++.+.....+....=+++. .++++||+|++
T Consensus 5 iP~kVi~i~~~~~~~A~vd~~Gv~r~V~l~Lv~~~~~~~~vGDyVLV 51 (82)
T PRK10413 5 VPGQVLAVGEDIHQLAQVEVCGIKRDVNIALICEGNPADLLGQWVLV 51 (82)
T ss_pred cceEEEEECCCCCcEEEEEcCCeEEEEEeeeeccCCcccccCCEEEE
Confidence 467888887653 6787777532212221112222 25789999987
No 33
>PRK14864 putative biofilm stress and motility protein A; Provisional
Probab=41.47 E-value=1.5e+02 Score=22.22 Aligned_cols=15 Identities=20% Similarity=0.191 Sum_probs=6.9
Q ss_pred chhHHHHHHHHHHHH
Q 027022 30 RRSSIGFGSSVILSS 44 (229)
Q Consensus 30 ~~~~~~~~~~~~~~a 44 (229)
+|+++.+++++++++
T Consensus 5 mk~~~~l~~~l~LS~ 19 (104)
T PRK14864 5 MRRFASLLLTLLLSA 19 (104)
T ss_pred HHHHHHHHHHHHHhh
Confidence 455544444444433
No 34
>cd00600 Sm_like The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=41.40 E-value=88 Score=20.24 Aligned_cols=33 Identities=27% Similarity=0.349 Sum_probs=26.6
Q ss_pred cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022 150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD 186 (229)
Q Consensus 150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~ 186 (229)
...+.|.+.++. .|.+.+.+.|+...+.+-...
T Consensus 6 g~~V~V~l~~g~----~~~G~L~~~D~~~Ni~L~~~~ 38 (63)
T cd00600 6 GKTVRVELKDGR----VLEGVLVAFDKYMNLVLDDVE 38 (63)
T ss_pred CCEEEEEECCCc----EEEEEEEEECCCCCEEECCEE
Confidence 357888887754 889999999999988876664
No 35
>PF08605 Rad9_Rad53_bind: Fungal Rad9-like Rad53-binding; InterPro: IPR013914 In Saccharomyces cerevisiae (Baker s yeast), the Rad9 is a key adaptor protein in DNA damage checkpoint pathways. DNA damage induces Rad9 phosphorylation, and Rad53 specifically associates with this region of Rad9, when phosphorylated, via the Rad53 IPR000253 from INTERPRO domain []. There is no clear higher eukaryotic ortholog to Rad9.
Probab=41.31 E-value=43 Score=26.21 Aligned_cols=47 Identities=23% Similarity=0.367 Sum_probs=33.8
Q ss_pred EEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEe
Q 027022 166 YREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIG 213 (229)
Q Consensus 166 ~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG 213 (229)
.|+|++++.+...+-.+++++....+.+.-.+ ..-++++||.|-.=|
T Consensus 24 yYPa~~~~~~~~~~~~~V~Fedg~~~i~~~dv-~~LDlRIGD~Vkv~~ 70 (131)
T PF08605_consen 24 YYPATCVGSGVDRDRSLVRFEDGTYEIKNEDV-KYLDLRIGDTVKVDG 70 (131)
T ss_pred EeeEEEEeecCCCCeEEEEEecCceEeCcccE-eeeeeecCCEEEECC
Confidence 78999999998888999999854322322222 123689999887766
No 36
>TIGR00074 hypC_hupF hydrogenase assembly chaperone HypC/HupF. An additional proposed function is to shuttle the iron atom that has been liganded at the HypC/HypD complex to the precursor of the large hydrogenase (HycE) subunit. PubMed:12441107.
Probab=37.16 E-value=80 Score=22.26 Aligned_cols=40 Identities=20% Similarity=0.307 Sum_probs=25.9
Q ss_pred EeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEE
Q 027022 167 REGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFA 211 (229)
Q Consensus 167 ~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~a 211 (229)
.+++++..+. +.|++.+... ...+.+.=-.++++||+|++
T Consensus 5 iP~~V~~i~~--~~A~v~~~G~---~~~v~l~lv~~~~vGD~VLV 44 (76)
T TIGR00074 5 IPGQVVEIDE--NIALVEFCGI---KRDVSLDLVGEVKVGDYVLV 44 (76)
T ss_pred cceEEEEEcC--CEEEEEcCCe---EEEEEEEeeCCCCCCCEEEE
Confidence 4678887766 4688877532 23333333357999999986
No 37
>PF09465 LBR_tudor: Lamin-B receptor of TUDOR domain; InterPro: IPR019023 The Lamin-B receptor is a chromatin and lamin binding protein in the inner nuclear membrane. It is one of the integral inner nuclear envelope membrane proteins responsible for targeting nuclear membranes to chromatin, being a downstream effector of Ran, a small Ras-like nuclear GTPase which regulates NE assembly. Lamin-B receptor interacts with importin beta, a Ran-binding protein, thereby directly contributing to the fusion of membrane vesicles and the formation of the nuclear envelope []. ; PDB: 2L8D_A 2DIG_A.
Probab=37.10 E-value=1.2e+02 Score=20.03 Aligned_cols=36 Identities=19% Similarity=0.075 Sum_probs=27.8
Q ss_pred CcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEcc
Q 027022 149 GLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDV 187 (229)
Q Consensus 149 ~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~ 187 (229)
.++.+.+..++.+ . .|.++|..+|...++.-++.+.
T Consensus 8 ~Ge~V~~rWP~s~--l-YYe~kV~~~d~~~~~y~V~Y~D 43 (55)
T PF09465_consen 8 IGEVVMVRWPGSS--L-YYEGKVLSYDSKSDRYTVLYED 43 (55)
T ss_dssp SS-EEEEE-TTTS----EEEEEEEEEETTTTEEEEEETT
T ss_pred CCCEEEEECCCCC--c-EEEEEEEEecccCceEEEEEcC
Confidence 5678888887643 2 5799999999999999999974
No 38
>PF00548 Peptidase_C3: 3C cysteine protease (picornain 3C); InterPro: IPR000199 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This signature defines cysteine peptidases belong to MEROPS peptidase family C3 (picornain, clan PA(C)), subfamilies C3A and C3B. The protein fold of this peptidase domain for members of this family resembles that of the serine peptidase, chymotrypsin [], the type example for clan PA. Picornaviral proteins are expressed as a single polyprotein which is cleaved by the viral C3 cysteine protease. The poliovirus polyprotein is selectively cleaved between the Gln-|-Gly bond. In other picornavirus reactions Glu may be substituted for Gln, and Ser or Thr for Gly. ; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis; PDB: 3SJO_E 2H6M_A 1QA7_C 1HAV_B 2HAL_A 2H9H_A 3QZQ_B 3QZR_A 3R0F_B 3SJ9_A ....
Probab=36.23 E-value=2.3e+02 Score=22.86 Aligned_cols=56 Identities=18% Similarity=0.081 Sum_probs=34.7
Q ss_pred ccceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcC---CCcEEEEEEcc
Q 027022 118 VEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDP---AYDLAVLKVDV 187 (229)
Q Consensus 118 ~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~---~~DlAvLki~~ 187 (229)
....++++-|-+ .++|-+.| -. ..+.+.+ + |+.+.....+...|. ..||++++++.
T Consensus 23 g~~t~l~~gi~~-~~~lvp~H-~~-------~~~~i~i---~--g~~~~~~d~~~lv~~~~~~~Dl~~v~l~~ 81 (172)
T PF00548_consen 23 GEFTMLALGIYD-RYFLVPTH-EE-------PEDTIYI---D--GVEYKVDDSVVLVDRDGVDTDLTLVKLPR 81 (172)
T ss_dssp EEEEEEEEEEEB-TEEEEEGG-GG-------GCSEEEE---T--TEEEEEEEEEEEEETTSSEEEEEEEEEES
T ss_pred ceEEEecceEee-eEEEEECc-CC-------CcEEEEE---C--CEEEEeeeeEEEecCCCcceeEEEEEccC
Confidence 456788888875 58999999 22 3334433 1 444333444434454 45999999964
No 39
>PF03761 DUF316: Domain of unknown function (DUF316) ; InterPro: IPR005514 This is a family of uncharacterised proteins from Caenorhabditis elegans.
Probab=34.44 E-value=45 Score=28.69 Aligned_cols=40 Identities=18% Similarity=0.326 Sum_probs=27.6
Q ss_pred cCCCcEEEEEEccC-CCCccceEcCCC-CCCCCCCeEEEEec
Q 027022 175 DPAYDLAVLKVDVE-GFELKPVVLGTS-HDLRVGQSCFAIGN 214 (229)
Q Consensus 175 d~~~DlAvLki~~~-~~~~~~l~lg~s-~~~~~G~~V~aiG~ 214 (229)
....++.||.++.+ .....++=|.++ ..+..|+.+.+.|+
T Consensus 158 ~~~~~~mIlEl~~~~~~~~~~~Cl~~~~~~~~~~~~~~~yg~ 199 (282)
T PF03761_consen 158 NRPYSPMILELEEDFSKNVSPPCLADSSTNWEKGDEVDVYGF 199 (282)
T ss_pred ccccceEEEEEcccccccCCCEEeCCCccccccCceEEEeec
Confidence 34568999999854 134455555554 45889999998888
No 40
>cd01722 Sm_F The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit F is capable of forming both homo- and hetero-heptamer ring structures. To form the hetero-heptamer, Sm subunit F initially binds subunits E and G to form a trimer which then assembles onto snRNA along with the D3/B and D1/D2 heterodimers.
Probab=33.41 E-value=1.1e+02 Score=20.74 Aligned_cols=33 Identities=18% Similarity=0.145 Sum_probs=26.6
Q ss_pred cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022 150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD 186 (229)
Q Consensus 150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~ 186 (229)
.+.+.|.+.++. .+..++.++|....|.+=...
T Consensus 11 g~~V~V~Lk~g~----~~~G~L~~~D~~mNi~L~~~~ 43 (68)
T cd01722 11 GKPVIVKLKWGM----EYKGTLVSVDSYMNLQLANTE 43 (68)
T ss_pred CCEEEEEECCCc----EEEEEEEEECCCEEEEEeeEE
Confidence 367888997754 889999999999988876553
No 41
>cd01726 LSm6 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm6 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=32.34 E-value=1.2e+02 Score=20.35 Aligned_cols=33 Identities=15% Similarity=0.086 Sum_probs=26.7
Q ss_pred cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022 150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD 186 (229)
Q Consensus 150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~ 186 (229)
.+.+.|.+.++. .|..++.++|+...|-+=...
T Consensus 10 ~~~V~V~Lk~g~----~~~G~L~~~D~~mNlvL~~~~ 42 (67)
T cd01726 10 GRPVVVKLNSGV----DYRGILACLDGYMNIALEQTE 42 (67)
T ss_pred CCeEEEEECCCC----EEEEEEEEEccceeeEEeeEE
Confidence 467889998754 889999999999988876654
No 42
>cd01728 LSm1 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm1 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=29.85 E-value=1.9e+02 Score=20.03 Aligned_cols=57 Identities=18% Similarity=0.164 Sum_probs=35.8
Q ss_pred ceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccC---CCCccceEcCCCCCCCCCCeEEEEe
Q 027022 151 HRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVE---GFELKPVVLGTSHDLRVGQSCFAIG 213 (229)
Q Consensus 151 ~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~---~~~~~~l~lg~s~~~~~G~~V~aiG 213 (229)
+++.|.+.++ + .+.+.+.++|+..-|.+=..... ........+|. -+-.|+.|+.+|
T Consensus 13 k~v~V~l~~g--r--~~~G~L~~fD~~~NlvL~d~~E~~~~~~~~~~~~lG~--~viRG~~V~~ig 72 (74)
T cd01728 13 KKVVVLLRDG--R--KLIGILRSFDQFANLVLQDTVERIYVGDKYGDIPRGI--FIIRGENVVLLG 72 (74)
T ss_pred CEEEEEEcCC--e--EEEEEEEEECCcccEEecceEEEEecCCccceeEeeE--EEEECCEEEEEE
Confidence 6788888774 4 78999999999987776544211 11112222221 255577888777
No 43
>PRK00737 small nuclear ribonucleoprotein; Provisional
Probab=29.70 E-value=1.9e+02 Score=19.78 Aligned_cols=33 Identities=18% Similarity=0.175 Sum_probs=26.7
Q ss_pred cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022 150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD 186 (229)
Q Consensus 150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~ 186 (229)
.+.+.|.+.++ + .|.+++.++|+...+-+-...
T Consensus 14 ~k~V~V~lk~g--~--~~~G~L~~~D~~mNlvL~d~~ 46 (72)
T PRK00737 14 NSPVLVRLKGG--R--EFRGELQGYDIHMNLVLDNAE 46 (72)
T ss_pred CCEEEEEECCC--C--EEEEEEEEEcccceeEEeeEE
Confidence 35688888774 4 789999999999988887764
No 44
>TIGR00219 mreC rod shape-determining protein MreC. MreC (murein formation C) is involved in the rod shape determination in E. coli, and more generally in cell shape determination of bacteria whether or not they are rod-shaped. Cells defective in MreC are round. Species with MreC include many of the Proteobacteria, Gram-positives, and spirochetes.
Probab=28.98 E-value=4e+02 Score=23.39 Aligned_cols=52 Identities=17% Similarity=0.142 Sum_probs=32.0
Q ss_pred CCcEEEEEEccCCC--CccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEc
Q 027022 177 AYDLAVLKVDVEGF--ELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTF 228 (229)
Q Consensus 177 ~~DlAvLki~~~~~--~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVS 228 (229)
..+-++++=...+. .+....+....++++||.|+.=|.-.-+...+..|.|+
T Consensus 188 t~~~gi~~G~~~g~~~~l~l~~~~~~~~v~~GD~VvTSGlgg~fP~Gl~VG~V~ 241 (283)
T TIGR00219 188 SDFRGLIEGNGYGKTLEMNLVNRPAEKDIKKGDLIVTSGLGGRFPEGYPIGVVT 241 (283)
T ss_pred CCceEEEEecCCCCCcEEEEEECCCCCCCCCCCEEEECCCCCcCCCCCEEEEEE
Confidence 34557777542111 12223344466899999999988765566667777664
No 45
>PF02122 Peptidase_S39: Peptidase S39; InterPro: IPR000382 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. ORF2 of Potato leafroll virus (PLrV) encodes a polyprotein which is translated following a -1 frameshift. The polyprotein has a putative linear arrangement of membrane achor-VPg-peptidase-polmerase domains. The serine peptidase domain which is found in this group of sequences belongs to MEROPS peptidase family S39 (clan PA(S)). It is likely that the peptidase domain is involved in the cleavage of the polyprotein []. The nucleotide sequence for the RNA of PLrV has been determined [, ]. The sequence contains six large open reading frames (ORFs). The 5' coding region encodes two polypeptides of 28K and 70K, which overlap in different reading frames; it is suggested that the third ORF in the 5' block is translated by frameshift readthrough near the end of the 70K protein, yielding a 118K polypeptide []. Segments of the predicted amino acid sequences of these ORFs resemble those of known viral RNA polymerases, ATP-binding proteins and viral genome-linked proteins. The nucleotide sequence of the genomic RNA of Beet western yellows virus (BWYV) has been determined []. The sequence contains six long ORFs. A cluster of three of these ORFs, including the coat protein cistron, display extensive amino acid sequence similarity to corresponding ORFs of a second luteovirus: Barley yellow dwarf virus [].; GO: 0004252 serine-type endopeptidase activity, 0022415 viral reproductive process, 0016021 integral to membrane; PDB: 1ZYO_A.
Probab=28.84 E-value=88 Score=26.32 Aligned_cols=60 Identities=17% Similarity=0.063 Sum_probs=24.0
Q ss_pred ccceEEEEE-EcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022 118 VEGTGSGFV-WDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD 186 (229)
Q Consensus 118 ~~~~GSGfi-I~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~ 186 (229)
..+.++.+- ++.+-.++|++||..+ ..... .+.++ .+.-.-.-+.+..+...|++||++.
T Consensus 28 hvGya~cv~l~~g~~~L~ta~Hv~~~-------~~~~~-~~k~g-~kipl~~f~~~~~~~~~D~~il~~P 88 (203)
T PF02122_consen 28 HVGYATCVRLFDGEDALLTARHVWSR-------PSKVT-SLKTG-EKIPLAEFTDLLESRIADFVILRGP 88 (203)
T ss_dssp ------EEEE----EEEEE-HHHHTS-------SS----EEETT-EEEE--S-EEEEE-TTT-EEEEE--
T ss_pred ccccceEEECcCCccceecccccCCC-------cccee-EcCCC-CcccchhChhhhCCCccCEEEEecC
Confidence 344555533 2333489999999993 22222 22232 1111122344556889999999996
No 46
>PF08758 Cadherin_pro: Cadherin prodomain like; InterPro: IPR014868 Cadherins are a group of proteins that mediate calcium dependent cell-cell adhesion. They are activated through cleavage of a prosequence in the late Golgi. This protein corresponds to the folded region of the prosequence, and is termed the prodomain. The prodomain shows structural resemblance to the cadherin domain, but lacks all the features known to be important for cadherin-cadherin interactions []. ; GO: 0007155 cell adhesion, 0016021 integral to membrane; PDB: 1OP4_A.
Probab=28.73 E-value=2.3e+02 Score=20.48 Aligned_cols=40 Identities=13% Similarity=0.104 Sum_probs=26.0
Q ss_pred ceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCe
Q 027022 120 GTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNG 164 (229)
Q Consensus 120 ~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~ 164 (229)
+.=+=|-|.+||-|.|..++..-. +.....|...|.+++.
T Consensus 43 ssDpdF~V~~DGsVy~~r~v~l~~-----~~~~F~V~a~D~~~~~ 82 (90)
T PF08758_consen 43 SSDPDFRVLEDGSVYAKRPVQLSS-----EQRSFTVHAWDSQTQE 82 (90)
T ss_dssp ---SEEEEETTTEEEEES--S-SS-----S-EEEEEEEEETTTTE
T ss_pred cCCCCEEEcCCCeEEEeeeEecCC-----CceEEEEEEECCCCCe
Confidence 334478899999999988888732 4457888888888774
No 47
>PRK09507 cspE cold shock protein CspE; Reviewed
Probab=27.96 E-value=1.5e+02 Score=20.19 Aligned_cols=46 Identities=9% Similarity=0.051 Sum_probs=30.3
Q ss_pred EEeEEEEEEcCCCcEEEEEEccCCCC----ccceEcCCCCCCCCCCeEEE
Q 027022 166 YREGKMVGCDPAYDLAVLKVDVEGFE----LKPVVLGTSHDLRVGQSCFA 211 (229)
Q Consensus 166 ~~~A~vv~~d~~~DlAvLki~~~~~~----~~~l~lg~s~~~~~G~~V~a 211 (229)
.+..+|..+|.+.+...++.+....+ +..+.-.....++.||.|--
T Consensus 3 ~~~G~Vk~f~~~kGyGFI~~~~g~~dvfvH~s~l~~~g~~~l~~G~~V~f 52 (69)
T PRK09507 3 KIKGNVKWFNESKGFGFITPEDGSKDVFVHFSAIQTNGFKTLAEGQRVEF 52 (69)
T ss_pred ccceEEEEEeCCCCcEEEecCCCCeeEEEEeecccccCCCCCCCCCEEEE
Confidence 35678888899999998888754322 23333222356899998854
No 48
>TIGR00237 xseA exodeoxyribonuclease VII, large subunit. This family consist of exodeoxyribonuclease VII, large subunit XseA which catalyses exonucleolytic cleavage in either the 5'-3' or 3'-5' direction to yield 5'-phosphomononucleotides. Exonuclease VII consists of one large subunit and four small subunits.
Probab=26.95 E-value=64 Score=30.21 Aligned_cols=29 Identities=17% Similarity=0.190 Sum_probs=25.0
Q ss_pred EEEcCCcEEEEccccccccccCCCCcceEEEEEecC
Q 027022 125 FVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDA 160 (229)
Q Consensus 125 fiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~ 160 (229)
++-+++|.|+++.+-+. ..+.+.+.+.||
T Consensus 398 i~~~~~g~~v~s~~~l~-------~gd~l~i~~~dG 426 (432)
T TIGR00237 398 IALNEKGKAIKSVKQVD-------RGDRLTTKLKDG 426 (432)
T ss_pred EEEecCCCEecCHHHCC-------CCCEEEEEECCe
Confidence 44467899999999999 789999999985
No 49
>PRK10943 cold shock-like protein CspC; Provisional
Probab=25.88 E-value=1.5e+02 Score=20.14 Aligned_cols=45 Identities=9% Similarity=0.036 Sum_probs=29.4
Q ss_pred EeEEEEEEcCCCcEEEEEEccCCCC----ccceEcCCCCCCCCCCeEEE
Q 027022 167 REGKMVGCDPAYDLAVLKVDVEGFE----LKPVVLGTSHDLRVGQSCFA 211 (229)
Q Consensus 167 ~~A~vv~~d~~~DlAvLki~~~~~~----~~~l~lg~s~~~~~G~~V~a 211 (229)
+..+|..+|.+.+...|+.+..+.+ ...+.-..-..+..||.|--
T Consensus 4 ~~G~Vk~f~~~kGfGFI~~~~g~~dvFvH~s~l~~~g~~~l~~G~~V~f 52 (69)
T PRK10943 4 IKGQVKWFNESKGFGFITPADGSKDVFVHFSAIQGNGFKTLAEGQNVEF 52 (69)
T ss_pred cceEEEEEeCCCCcEEEecCCCCeeEEEEhhHccccCCCCCCCCCEEEE
Confidence 4678888888888888888643222 34443322356889998753
No 50
>COG5510 Predicted small secreted protein [Function unknown]
Probab=25.56 E-value=77 Score=19.97 Aligned_cols=22 Identities=27% Similarity=0.389 Sum_probs=13.8
Q ss_pred chhHHHHHHHHHHHHHhhhccC
Q 027022 30 RRSSIGFGSSVILSSFLVNFCS 51 (229)
Q Consensus 30 ~~~~~~~~~~~~~~a~l~~~~~ 51 (229)
+++.+.+.+++++++.++.+|.
T Consensus 2 mk~t~l~i~~vll~s~llaaCN 23 (44)
T COG5510 2 MKKTILLIALVLLASTLLAACN 23 (44)
T ss_pred chHHHHHHHHHHHHHHHHHHhh
Confidence 4455555666666677767774
No 51
>cd01720 Sm_D2 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit D2 heterodimerizes with subunit D1 and three such heterodimers form a hexameric ring structure with alternating D1 and D2 subunits. The D1 - D2 heterodimer also assembles into a heptameric ring containing D2, D3, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=25.47 E-value=1.7e+02 Score=21.10 Aligned_cols=34 Identities=12% Similarity=0.139 Sum_probs=27.3
Q ss_pred CcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022 149 GLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD 186 (229)
Q Consensus 149 ~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~ 186 (229)
..+.+.|.+.++. .+.+++.++|...-|.|=..+
T Consensus 13 ~~~~V~V~lr~~r----~~~G~L~~fD~hmNlvL~d~~ 46 (87)
T cd01720 13 NNTQVLINCRNNK----KLLGRVKAFDRHCNMVLENVK 46 (87)
T ss_pred CCCEEEEEEcCCC----EEEEEEEEecCccEEEEcceE
Confidence 3468889997754 789999999999988876554
No 52
>cd01731 archaeal_Sm1 The archaeal sm1 proteins: The Sm proteins are conserved in all three domains of life and are always associated with U-rich RNA sequences. They function to mediate RNA-RNA interactions and RNA biogenesis. All Sm proteins contain a common sequence motif in two segments, Sm1 and Sm2, separated by a short variable linker. Eukaryotic Sm proteins form part of specific small nuclear ribonucleoproteins (snRNPs) that are involved in the processing of pre-mRNAs to mature mRNAs, and are a major component of the eukaryotic spliceosome. Most snRNPs consist of seven Sm proteins (B/B', D1, D2, D3, E, F and G) arranged in a ring on a uridine-rich sequence (Sm site), plus a small nuclear RNA (snRNA) (either U1, U2, U5 or U4/6). Since archaebacteria do not have any splicing apparatus, Sm proteins of archaebacteria may play a more general role. Archaeal Lsm proteins are likely to represent the ancestral Sm domain.
Probab=24.62 E-value=2.2e+02 Score=19.01 Aligned_cols=33 Identities=15% Similarity=0.187 Sum_probs=27.1
Q ss_pred cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022 150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD 186 (229)
Q Consensus 150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~ 186 (229)
.+++.|.+.++ + .+.+++.++|+...|.+-...
T Consensus 10 ~~~V~V~l~~g--~--~~~G~L~~~D~~mNlvL~~~~ 42 (68)
T cd01731 10 NKPVLVKLKGG--K--EVRGRLKSYDQHMNLVLEDAE 42 (68)
T ss_pred CCEEEEEECCC--C--EEEEEEEEECCcceEEEeeEE
Confidence 36788888774 4 789999999999999887764
No 53
>PRK09890 cold shock protein CspG; Provisional
Probab=24.43 E-value=2.1e+02 Score=19.49 Aligned_cols=45 Identities=9% Similarity=0.017 Sum_probs=29.4
Q ss_pred EeEEEEEEcCCCcEEEEEEccCCCC----ccceEcCCCCCCCCCCeEEE
Q 027022 167 REGKMVGCDPAYDLAVLKVDVEGFE----LKPVVLGTSHDLRVGQSCFA 211 (229)
Q Consensus 167 ~~A~vv~~d~~~DlAvLki~~~~~~----~~~l~lg~s~~~~~G~~V~a 211 (229)
+..+|..+|.+.+...|+.+..+.+ ...+.-.....++.||.|--
T Consensus 5 ~~G~Vk~f~~~kGfGFI~~~~g~~dvFvH~s~l~~~~~~~l~~G~~V~f 53 (70)
T PRK09890 5 MTGLVKWFNADKGFGFITPDDGSKDVFVHFTAIQSNEFRTLNENQKVEF 53 (70)
T ss_pred ceEEEEEEECCCCcEEEecCCCCceEEEEEeeeccCCCCCCCCCCEEEE
Confidence 3578888888888888888743222 33333333357899998854
No 54
>cd01717 Sm_B The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit B heterodimerizes with subunit D3 and three such heterodimers form a hexameric ring structure with alternating B and D3 subunits. The D3 - B heterodimer also assembles into a heptameric ring containing D1, D2, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=24.01 E-value=1.9e+02 Score=20.03 Aligned_cols=33 Identities=21% Similarity=0.296 Sum_probs=26.1
Q ss_pred cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022 150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD 186 (229)
Q Consensus 150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~ 186 (229)
.+.+.|.+.++ + .+.+.+.++|....|.|=...
T Consensus 10 ~~~V~V~l~dg--R--~~~G~L~~~D~~~NlVL~~~~ 42 (79)
T cd01717 10 NYRLRVTLQDG--R--QFVGQFLAFDKHMNLVLSDCE 42 (79)
T ss_pred CCEEEEEECCC--c--EEEEEEEEEcCccCEEcCCEE
Confidence 36788999874 4 789999999999988765553
No 55
>cd05701 S1_Rrp5_repeat_hs10 S1_Rrp5_repeat_hs10: Rrp5 is a trans-acting factor important for biogenesis of both the 40S and 60S eukaryotic ribosomal subunits. Rrp5 has two distinct regions, an N-terminal region containing tandemly repeated S1 RNA-binding domains (12 S1 repeats in Saccharomyces cerevisiae Rrp5 and 14 S1 repeats in Homo sapiens Rrp5) and a C-terminal region containing tetratricopeptide repeat (TPR) motifs thought to be involved in protein-protein interactions. Mutational studies have shown that each region represents a specific functional domain. Deletions within the S1-containing region inhibit pre-rRNA processing at either site A3 or A2, whereas deletions within the TPR region confer an inability to support cleavage of A0-A2. This CD includes H. sapiens S1 repeat 10 (hs10). Rrp5 is found in eukaryotes but not in prokaryotes or archaea.
Probab=23.85 E-value=1.6e+02 Score=20.29 Aligned_cols=39 Identities=26% Similarity=0.198 Sum_probs=21.4
Q ss_pred EEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEE
Q 027022 171 MVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAI 212 (229)
Q Consensus 171 vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~ai 212 (229)
|+.-+...||+.+.+.. .+--...-+++++++|+.+.+.
T Consensus 17 vvSL~~t~~L~a~p~~s---HLNdtfrf~seklkvG~~l~v~ 55 (69)
T cd05701 17 IVSLATTGDLAAFPTRS---HLNDTFRFDSEKLSVGQCLDVT 55 (69)
T ss_pred EEEeeccccEEEEEchh---hccccccccceeeeccceEEEE
Confidence 33444455555555532 1222223357889999988764
No 56
>cd01730 LSm3 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm3 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=23.60 E-value=1.7e+02 Score=20.53 Aligned_cols=31 Identities=19% Similarity=0.240 Sum_probs=24.6
Q ss_pred ceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEE
Q 027022 151 HRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKV 185 (229)
Q Consensus 151 ~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki 185 (229)
+.+.|.+.++ + .+.+++.++|....|.|=..
T Consensus 12 k~V~V~l~~g--r--~~~G~L~~fD~~mNlvL~d~ 42 (82)
T cd01730 12 ERVYVKLRGD--R--ELRGRLHAYDQHLNMILGDV 42 (82)
T ss_pred CEEEEEECCC--C--EEEEEEEEEccceEEeccce
Confidence 5788888774 4 78999999999988876433
No 57
>PRK10081 entericidin B membrane lipoprotein; Provisional
Probab=23.49 E-value=1.5e+02 Score=19.04 Aligned_cols=22 Identities=23% Similarity=0.297 Sum_probs=11.7
Q ss_pred chhHHHHHHHHHHHHHhhhccC
Q 027022 30 RRSSIGFGSSVILSSFLVNFCS 51 (229)
Q Consensus 30 ~~~~~~~~~~~~~~a~l~~~~~ 51 (229)
++|.+.+.++++++++++.+|.
T Consensus 2 mKk~i~~i~~~l~~~~~l~~Cn 23 (48)
T PRK10081 2 VKKTIAAIFSVLVLSTVLTACN 23 (48)
T ss_pred hHHHHHHHHHHHHHHHHHhhhh
Confidence 3455555455555455447774
No 58
>PRK00286 xseA exodeoxyribonuclease VII large subunit; Reviewed
Probab=22.81 E-value=1.3e+02 Score=28.05 Aligned_cols=30 Identities=20% Similarity=0.275 Sum_probs=25.1
Q ss_pred EEEEcCCcEEEEccccccccccCCCCcceEEEEEecC
Q 027022 124 GFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDA 160 (229)
Q Consensus 124 GfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~ 160 (229)
.++-+++|.++|+.+-++ ..+.+.+.+.||
T Consensus 402 a~v~~~~g~~i~s~~~~~-------~~d~i~i~~~dG 431 (438)
T PRK00286 402 AIVRDEDGKVIRSAKQLK-------PGDRLTIRLADG 431 (438)
T ss_pred EEEEeCCCCEeccHHHCC-------CCCEEEEEECCe
Confidence 344456799999999999 789999999985
No 59
>PRK10354 RNA chaperone/anti-terminator; Provisional
Probab=22.31 E-value=2.5e+02 Score=19.08 Aligned_cols=43 Identities=12% Similarity=0.094 Sum_probs=28.1
Q ss_pred eEEEEEEcCCCcEEEEEEccCCCC----ccceEcCCCCCCCCCCeEE
Q 027022 168 EGKMVGCDPAYDLAVLKVDVEGFE----LKPVVLGTSHDLRVGQSCF 210 (229)
Q Consensus 168 ~A~vv~~d~~~DlAvLki~~~~~~----~~~l~lg~s~~~~~G~~V~ 210 (229)
..+|..+|.+.+...|+.+....+ ...+.-.....++.||.|-
T Consensus 6 ~G~Vk~f~~~kGfGFI~~~~g~~dvfvH~s~l~~~g~~~l~~G~~V~ 52 (70)
T PRK10354 6 TGIVKWFNADKGFGFITPDDGSKDVFVHFSAIQNDGYKSLDEGQKVS 52 (70)
T ss_pred eEEEEEEeCCCCcEEEecCCCCccEEEEEeeccccCCCCCCCCCEEE
Confidence 577888888888888887643222 3333322235689999885
No 60
>TIGR03497 FliI_clade2 flagellar protein export ATPase FliI. Members of this protein family are the FliI protein of bacterial flagellum systems. This protein acts to drive protein export for flagellar biosynthesis. The most closely related family is the YscN family of bacterial type III secretion systems. This model represents one (of three) segment of the FliI family tree. These have been modeled separately in order to exclude the type III secretion ATPases more effectively.
Probab=22.28 E-value=4.9e+02 Score=24.29 Aligned_cols=38 Identities=24% Similarity=0.324 Sum_probs=20.8
Q ss_pred EEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCC
Q 027022 166 YREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPY 216 (229)
Q Consensus 166 ~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~ 216 (229)
...++|++.+ .|.+++.. +++...++.|+.|...|.|+
T Consensus 33 ~~~~eVi~~~--~~~v~l~~-----------~~~t~gl~~G~~V~~tg~~~ 70 (413)
T TIGR03497 33 PVLAEVVGFK--EENVLLMP-----------LGEVEGIGPGSLVIATGRPL 70 (413)
T ss_pred eEEEEEEEEc--CCeEEEEE-----------ccCccCCCCCCEEEEcCCee
Confidence 3467777777 33344443 34444555666666555544
No 61
>COG1792 MreC Cell shape-determining protein [Cell envelope biogenesis, outer membrane]
Probab=21.92 E-value=5.5e+02 Score=22.57 Aligned_cols=34 Identities=21% Similarity=0.205 Sum_probs=26.4
Q ss_pred EcCCCCCCCCCCeEEEEecCCCCCCceeEeEEcC
Q 027022 196 VLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTFQ 229 (229)
Q Consensus 196 ~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVSa 229 (229)
.+-...++++||.|.+-|.-.-+...+..|.|++
T Consensus 206 ~~~~~~~i~~GD~vvTSGlgg~fP~Gl~Vg~V~~ 239 (284)
T COG1792 206 YLPPNSDIKEGDLVVTSGLGGVFPAGLPVGEVSS 239 (284)
T ss_pred eccCCCCccCCCEEEecCCCCcCCCCcEEEEEEE
Confidence 3445678999999999888766777788887763
No 62
>PTZ00138 small nuclear ribonucleoprotein; Provisional
Probab=21.72 E-value=2.4e+02 Score=20.49 Aligned_cols=36 Identities=22% Similarity=0.371 Sum_probs=28.0
Q ss_pred CcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022 149 GLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD 186 (229)
Q Consensus 149 ~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~ 186 (229)
....+.|.+.++.++ .+...+.++|.-..|.+=...
T Consensus 25 ~~~~V~i~l~~~~~r--~~~G~L~gfD~~mNlVL~d~~ 60 (89)
T PTZ00138 25 EKTRVQIWLYDHPNL--RIEGKILGFDEYMNMVLDDAE 60 (89)
T ss_pred CCcEEEEEEEeCCCc--EEEEEEEEEcccceEEEccEE
Confidence 456788888886555 789999999999988776553
No 63
>PRK08927 fliI flagellum-specific ATP synthase; Validated
Probab=21.72 E-value=5.7e+02 Score=24.18 Aligned_cols=51 Identities=22% Similarity=0.035 Sum_probs=27.5
Q ss_pred EEEEEEcCCcEEEEcccc---ccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEE
Q 027022 122 GSGFVWDKFGHIVTNYHV---VAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKV 185 (229)
Q Consensus 122 GSGfiI~~~G~IlTn~HV---v~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki 185 (229)
-.|-|..-.|.++...-. +. -.+-+.|.. .+|+ ...++|++.+.+. ++|..
T Consensus 17 ~~g~v~~i~g~~i~v~g~~~~~~-------~ge~~~i~~--~~~~--~~~~eVv~~~~~~--~~l~~ 70 (442)
T PRK08927 17 IYGRVVAVRGLLVEVAGPIHALS-------VGARIVVET--RGGR--PVPCEVVGFRGDR--ALLMP 70 (442)
T ss_pred eeeEEEEEEccEEEEEecCCCCC-------cCCEEEEEc--CCCC--EEEEEEEEEcCCe--EEEEE
Confidence 345555555666665554 22 334455532 2243 3578999888774 44444
No 64
>PRK15464 cold shock-like protein CspH; Provisional
Probab=21.60 E-value=2.5e+02 Score=19.22 Aligned_cols=44 Identities=11% Similarity=-0.034 Sum_probs=29.2
Q ss_pred EeEEEEEEcCCCcEEEEEEccCCCC----ccceEcCCCCCCCCCCeEE
Q 027022 167 REGKMVGCDPAYDLAVLKVDVEGFE----LKPVVLGTSHDLRVGQSCF 210 (229)
Q Consensus 167 ~~A~vv~~d~~~DlAvLki~~~~~~----~~~l~lg~s~~~~~G~~V~ 210 (229)
+..+|..+|.+.....++.+....+ +..+.-.....++.||.|-
T Consensus 5 ~~G~Vk~fn~~KGfGFI~~~~g~~DvFvH~s~l~~~g~~~l~~G~~V~ 52 (70)
T PRK15464 5 MTGIVKTFDRKSGKGFIIPSDGRKEVQVHISAFTPRDAEVLIPGLRVE 52 (70)
T ss_pred ceEEEEEEECCCCeEEEccCCCCccEEEEehhehhcCCCCCCCCCEEE
Confidence 4678888888888888887653322 3444322334699999874
No 65
>PF10844 DUF2577: Protein of unknown function (DUF2577); InterPro: IPR022555 This family of proteins has no known function
Probab=21.53 E-value=2.5e+02 Score=20.48 Aligned_cols=56 Identities=14% Similarity=0.143 Sum_probs=34.2
Q ss_pred EEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCC
Q 027022 124 GFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDL 203 (229)
Q Consensus 124 GfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~ 203 (229)
+++++++-+++++ |+-+ ....+.+.....+.. .. +.+ .+.+
T Consensus 37 ~liL~~~~L~i~~-~l~~-------~~~~~~~~~~~~~~~----~~-------------------------i~~--~~~L 77 (100)
T PF10844_consen 37 KLILDKDFLIIPE-LLKD-------YTRDITIEHNSETDN----IT-------------------------ITF--TDGL 77 (100)
T ss_pred eEEEchHHEEeeh-hccc-------eEEEEEEeccccccc----ee-------------------------EEE--ecCC
Confidence 4888887788888 7766 444454444332110 11 344 3478
Q ss_pred CCCCeEEEEecCCCC
Q 027022 204 RVGQSCFAIGNPYGF 218 (229)
Q Consensus 204 ~~G~~V~aiG~P~G~ 218 (229)
++||.|+.+-.-.|.
T Consensus 78 k~GD~V~ll~~~~gQ 92 (100)
T PF10844_consen 78 KVGDKVLLLRVQGGQ 92 (100)
T ss_pred cCCCEEEEEEecCCC
Confidence 999999998765553
No 66
>PF04083 Abhydro_lipase: Partial alpha/beta-hydrolase lipase region; InterPro: IPR006693 The alpha/beta hydrolase fold is common to several hydrolytic enzymes of widely differing phylogenetic origin and catalytic function. The core of each enzyme is similar: an alpha/beta sheet, not barrel, of eight beta-sheets connected by alpha-helices []. This entry represents the N-terminal part of an alpha/beta hydrolase domain found in a number of lipases.; GO: 0006629 lipid metabolic process; PDB: 1K8Q_B 1HLG_B.
Probab=21.22 E-value=2.1e+02 Score=19.20 Aligned_cols=19 Identities=21% Similarity=-0.011 Sum_probs=13.9
Q ss_pred EEEcCCcEEEEcccccccc
Q 027022 125 FVWDKFGHIVTNYHVVAKL 143 (229)
Q Consensus 125 fiI~~~G~IlTn~HVv~~~ 143 (229)
.|.-+||||||-.++..+.
T Consensus 16 ~V~T~DGYiL~l~RIp~~~ 34 (63)
T PF04083_consen 16 EVTTEDGYILTLHRIPPGK 34 (63)
T ss_dssp EEE-TTSEEEEEEEE-SBT
T ss_pred EEEeCCCcEEEEEEccCCC
Confidence 3566899999999999864
No 67
>PRK10781 rcsF outer membrane lipoprotein; Reviewed
Probab=21.11 E-value=1.1e+02 Score=24.14 Aligned_cols=15 Identities=7% Similarity=0.180 Sum_probs=7.7
Q ss_pred eEEEEEEcCCcEEEEc
Q 027022 121 TGSGFVWDKFGHIVTN 136 (229)
Q Consensus 121 ~GSGfiI~~~G~IlTn 136 (229)
.|.|+||.+ +.++.+
T Consensus 100 gaN~Vvl~~-C~~~~~ 114 (133)
T PRK10781 100 KANAVLLHS-CEITSG 114 (133)
T ss_pred CCCEEEEEE-eeccCC
Confidence 355666654 444443
No 68
>cd06168 LSm9 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm9 proteins have a single Sm-like domain structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=20.98 E-value=2.7e+02 Score=19.35 Aligned_cols=31 Identities=10% Similarity=0.172 Sum_probs=25.3
Q ss_pred ceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEE
Q 027022 151 HRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKV 185 (229)
Q Consensus 151 ~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki 185 (229)
+.+.|.+.|+ + .+.+.+..+|....|-+=..
T Consensus 11 ~~v~V~l~dg--R--~~~G~l~~~D~~~NivL~~~ 41 (75)
T cd06168 11 RTMRIHMTDG--R--TLVGVFLCTDRDCNIILGSA 41 (75)
T ss_pred CeEEEEEcCC--e--EEEEEEEEEcCCCcEEecCc
Confidence 6788999884 4 78999999999988876544
No 69
>PRK15463 cold shock-like protein CspF; Provisional
Probab=20.96 E-value=2.4e+02 Score=19.31 Aligned_cols=45 Identities=11% Similarity=0.080 Sum_probs=29.9
Q ss_pred EeEEEEEEcCCCcEEEEEEccCCCC----ccceEcCCCCCCCCCCeEEE
Q 027022 167 REGKMVGCDPAYDLAVLKVDVEGFE----LKPVVLGTSHDLRVGQSCFA 211 (229)
Q Consensus 167 ~~A~vv~~d~~~DlAvLki~~~~~~----~~~l~lg~s~~~~~G~~V~a 211 (229)
+.++|..+|.+.....|..+....+ +..+.-.....|+.||.|--
T Consensus 5 ~~G~Vk~fn~~kGfGFI~~~~g~~DvFvH~sal~~~g~~~l~~G~~V~f 53 (70)
T PRK15463 5 MTGIVKTFDGKSGKGLITPSDGRKDVQVHISALNLRDAEELTTGLRVEF 53 (70)
T ss_pred ceEEEEEEeCCCceEEEecCCCCccEEEEehhhhhcCCCCCCCCCEEEE
Confidence 4678888888888888888654322 33443322457999998753
No 70
>cd01732 LSm5 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation. Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm4 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure. Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=20.64 E-value=2.3e+02 Score=19.69 Aligned_cols=31 Identities=16% Similarity=0.227 Sum_probs=24.8
Q ss_pred ceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEE
Q 027022 151 HRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKV 185 (229)
Q Consensus 151 ~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki 185 (229)
+++.|.+.+ |+ .+.+++.++|+..-|.+=..
T Consensus 14 ~~V~V~l~~--gr--~~~G~L~g~D~~mNlvL~da 44 (76)
T cd01732 14 SRIWIVMKS--DK--EFVGTLLGFDDYVNMVLEDV 44 (76)
T ss_pred CEEEEEECC--Ce--EEEEEEEEeccceEEEEccE
Confidence 678888877 44 78999999999998876544
No 71
>PF10518 TAT_signal: TAT (twin-arginine translocation) pathway signal sequence; InterPro: IPR019546 The twin-arginine translocation (Tat) pathway serves the role of transporting folded proteins across energy-transducing membranes []. Homologues of the genes that encode the transport apparatus occur in archaea, bacteria, chloroplasts, and plant mitochondria []. In bacteria, the Tat pathway catalyses the export of proteins from the cytoplasm across the inner/cytoplasmic membrane. In chloroplasts, the Tat components are found in the thylakoid membrane and direct the import of proteins from the stroma. The Tat pathway acts separately from the general secretory (Sec) pathway, which transports proteins in an unfolded state []. It is generally accepted that the primary role of the Tat system is to translocate fully folded proteins across membranes. An example of proteins that need to be exported in their 3D conformation are redox proteins that have acquired complex multi-atom cofactors in the bacterial cytoplasm (or the chloroplast stroma or mitochondrial matrix). They include hydrogenases, formate dehydrogenases, nitrate reductases, trimethylamine N-oxide (TMAO) reductases and dimethyl sulphoxide (DMSO) reductases [, ]. The Tat system can also export whole heteroligomeric complexes in which some proteins have no Tat signal. This is the case of the DMSO reductase or formate dehydrogenase complexes. But there are also other cases where the physiological rationale for targeting a protein to the Tat signal is less obvious. Indeed, there are examples of homologous proteins that are in some cases targeted to the Tat pathway and in other cases to the Sec apparatus. Some examples are: copper nitrite reductases, flavin domains of flavocytochrome c and N-acetylmuramoyl-L-alanine amidases []. In halophilic archaea such as Halobacterium almost all secreted proteins appear to be Tat targeted. It has been proposed to be a response to the difficulties these organisms would otherwise face in successfully folding proteins extracellularly at high ionic strength []. The Tat signal peptide consists of three motifs: the positively charged N-terminal motif, the hydrophobic region and the C-terminal region that generally ends with a consensus short motif (A-x-A) specifying cleavage by signal peptidase. Sequence analysis revealed that signal peptides capable of targeting the Tat protein contain the consensus sequence [ST]-R-R-x-F-L-K. The nearly invariant twin-arginine gave rise to the pathway's name. In addition the h-region of Tat signal peptides is typically less hydrophobic than that of Sec-specific signal peptides [, ].
Probab=20.57 E-value=1.5e+02 Score=16.26 Aligned_cols=19 Identities=21% Similarity=0.324 Sum_probs=11.6
Q ss_pred ccchhHHHHHHHHHHHHHh
Q 027022 28 ITRRSSIGFGSSVILSSFL 46 (229)
Q Consensus 28 ~~~~~~~~~~~~~~~~a~l 46 (229)
+.||.++..++++.+++.+
T Consensus 2 ~sRR~fLk~~~a~~a~~~~ 20 (26)
T PF10518_consen 2 LSRRQFLKGGAAAAAAAAL 20 (26)
T ss_pred CcHHHHHHHHHHHHHHHHh
Confidence 3566666666666665555
No 72
>TIGR00638 Mop molybdenum-pterin binding domain. This model describes a multigene family of molybdenum-pterin binding proteins of about 70 amino acids in Clostridium pasteurianum, as a tandemly-repeated domain C-terminal to an unrelated domain in ModE, a molybdate transport gene repressor of E. coli, and in single or tandemly paired domains in several related proteins.
Probab=20.22 E-value=2.5e+02 Score=18.17 Aligned_cols=46 Identities=22% Similarity=0.289 Sum_probs=27.2
Q ss_pred EEeEEEEEEcCCCcEEEEEEccCCC-CccceEcCC----CCCCCCCCeEEEE
Q 027022 166 YREGKMVGCDPAYDLAVLKVDVEGF-ELKPVVLGT----SHDLRVGQSCFAI 212 (229)
Q Consensus 166 ~~~A~vv~~d~~~DlAvLki~~~~~-~~~~l~lg~----s~~~~~G~~V~ai 212 (229)
.+.++|.......+.+-+.++..+. .+. ..+.. .-.+++|++|++.
T Consensus 8 ~l~g~I~~i~~~g~~~~v~l~~~~~~~l~-a~i~~~~~~~l~l~~G~~v~~~ 58 (69)
T TIGR00638 8 QLKGKVVAIEDGDVNAEVDLLLGGGTKLT-AVITLESVAELGLKPGKEVYAV 58 (69)
T ss_pred EEEEEEEEEEECCCeEEEEEEECCCCEEE-EEecHHHHhhCCCCCCCEEEEE
Confidence 5677777776666677666664332 111 11211 2257899999875
No 73
>PRK13684 Ycf48-like protein; Provisional
Probab=20.11 E-value=2e+02 Score=25.65 Aligned_cols=8 Identities=38% Similarity=0.397 Sum_probs=3.9
Q ss_pred cEEEEccc
Q 027022 131 GHIVTNYH 138 (229)
Q Consensus 131 G~IlTn~H 138 (229)
|+++....
T Consensus 102 ~~~~G~~g 109 (334)
T PRK13684 102 GWIVGQPS 109 (334)
T ss_pred EEEeCCCc
Confidence 45555444
Done!