Query         027022
Match_columns 229
No_of_seqs    175 out of 1517
Neff          7.3 
Searched_HMMs 46136
Date          Fri Mar 29 03:51:01 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/027022.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/027022hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PRK10139 serine endoprotease;  100.0 3.4E-28 7.4E-33  226.1  20.3  141   77-229    41-188 (455)
  2 TIGR02038 protease_degS peripl 100.0 2.3E-27   5E-32  214.2  19.8  134   72-229    41-174 (351)
  3 PRK10898 serine endoprotease;  100.0 2.2E-27 4.7E-32  214.5  19.5  133   73-229    42-174 (353)
  4 PRK10942 serine endoprotease;   99.9 1.6E-26 3.6E-31  215.8  19.4  141   77-229    39-209 (473)
  5 TIGR02037 degP_htrA_DO peripla  99.9 4.4E-25 9.5E-30  204.1  16.8  140   78-229     3-155 (428)
  6 COG0265 DegQ Trypsin-like seri  99.8 1.6E-20 3.5E-25  169.1  14.7  137   76-229    33-169 (347)
  7 KOG1320 Serine protease [Postt  99.6 2.6E-15 5.6E-20  138.6  10.0  148   72-229   124-275 (473)
  8 PF13365 Trypsin_2:  Trypsin-li  99.1 3.5E-10 7.7E-15   85.2   7.7   61  122-186     1-65  (120)
  9 KOG1421 Predicted signaling-as  98.7 2.9E-08 6.4E-13   94.4   8.0  131   75-228    51-185 (955)
 10 PF00089 Trypsin:  Trypsin;  In  98.2 5.6E-05 1.2E-09   62.1  13.1   92  119-218    24-131 (220)
 11 cd00190 Tryp_SPc Trypsin-like   98.0 0.00023 4.9E-09   58.8  13.4   95  118-218    23-133 (232)
 12 smart00020 Tryp_SPc Trypsin-li  97.9 0.00014 3.1E-09   60.3  10.7   94  119-218    25-133 (229)
 13 PF10459 Peptidase_S46:  Peptid  97.6 0.00036 7.7E-09   68.6  10.5   24  120-143    47-70  (698)
 14 KOG1320 Serine protease [Postt  97.5 7.6E-05 1.7E-09   69.7   4.0  128   81-228    55-184 (473)
 15 COG3591 V8-like Glu-specific e  95.6    0.14 2.9E-06   44.5   9.8   94  122-219    66-174 (251)
 16 PF00863 Peptidase_C4:  Peptida  94.8    0.29 6.3E-06   42.1   9.5   82  120-217    32-119 (235)
 17 KOG1421 Predicted signaling-as  92.5    0.27 5.9E-06   48.1   5.8   87  118-217   548-635 (955)
 18 PF05579 Peptidase_S32:  Equine  88.2     2.1 4.6E-05   37.5   7.1   64  120-199   114-177 (297)
 19 PF03510 Peptidase_C24:  2C end  82.2     3.1 6.6E-05   31.4   4.6   55  124-201     3-57  (105)
 20 KOG3627 Trypsin [Amino acid tr  79.2      21 0.00045   30.0   9.5   86  121-214    39-148 (256)
 21 PF09342 DUF1986:  Domain of un  72.7      65  0.0014   28.1  11.2   99  117-223    25-137 (267)
 22 PF01455 HupF_HypC:  HupF/HypC   64.2      29 0.00062   23.9   5.6   43  167-212     5-47  (68)
 23 COG0298 HypC Hydrogenase matur  64.0      14 0.00029   26.5   3.9   46  167-215     5-52  (82)
 24 cd01735 LSm12_N LSm12 belongs   62.1      32 0.00068   23.3   5.3   33  150-186     6-38  (61)
 25 PRK13922 rod shape-determining  58.7 1.2E+02  0.0026   26.2   9.9   51  178-228   190-240 (276)
 26 PF01732 DUF31:  Putative pepti  58.3       7 0.00015   35.7   2.2   24  118-141    34-67  (374)
 27 COG5640 Secreted trypsin-like   54.9      25 0.00054   32.3   5.0   19  124-143    65-83  (413)
 28 PRK10672 rare lipoprotein A; P  54.2      52  0.0011   30.2   7.0   29  111-141    85-113 (361)
 29 PF02601 Exonuc_VII_L:  Exonucl  49.3      21 0.00046   31.6   3.7   34  120-160   280-313 (319)
 30 COG3338 Cah Carbonic anhydrase  48.2      32 0.00069   29.7   4.3   25  161-185   124-148 (250)
 31 PF15436 PGBA_N:  Plasminogen-b  46.1 1.7E+02  0.0036   25.0   8.4   74  128-211    12-88  (218)
 32 PRK10413 hydrogenase 2 accesso  43.9      50  0.0011   23.7   4.2   45  167-211     5-51  (82)
 33 PRK14864 putative biofilm stre  41.5 1.5E+02  0.0033   22.2   7.6   15   30-44      5-19  (104)
 34 cd00600 Sm_like The eukaryotic  41.4      88  0.0019   20.2   5.0   33  150-186     6-38  (63)
 35 PF08605 Rad9_Rad53_bind:  Fung  41.3      43 0.00092   26.2   3.8   47  166-213    24-70  (131)
 36 TIGR00074 hypC_hupF hydrogenas  37.2      80  0.0017   22.3   4.4   40  167-211     5-44  (76)
 37 PF09465 LBR_tudor:  Lamin-B re  37.1 1.2E+02  0.0027   20.0   5.1   36  149-187     8-43  (55)
 38 PF00548 Peptidase_C3:  3C cyst  36.2 2.3E+02   0.005   22.9   8.0   56  118-187    23-81  (172)
 39 PF03761 DUF316:  Domain of unk  34.4      45 0.00098   28.7   3.4   40  175-214   158-199 (282)
 40 cd01722 Sm_F The eukaryotic Sm  33.4 1.1E+02  0.0023   20.7   4.5   33  150-186    11-43  (68)
 41 cd01726 LSm6 The eukaryotic Sm  32.3 1.2E+02  0.0026   20.4   4.6   33  150-186    10-42  (67)
 42 cd01728 LSm1 The eukaryotic Sm  29.8 1.9E+02  0.0042   20.0   5.9   57  151-213    13-72  (74)
 43 PRK00737 small nuclear ribonuc  29.7 1.9E+02   0.004   19.8   6.6   33  150-186    14-46  (72)
 44 TIGR00219 mreC rod shape-deter  29.0   4E+02  0.0087   23.4   9.8   52  177-228   188-241 (283)
 45 PF02122 Peptidase_S39:  Peptid  28.8      88  0.0019   26.3   4.0   60  118-186    28-88  (203)
 46 PF08758 Cadherin_pro:  Cadheri  28.7 2.3E+02  0.0049   20.5   6.2   40  120-164    43-82  (90)
 47 PRK09507 cspE cold shock prote  28.0 1.5E+02  0.0032   20.2   4.4   46  166-211     3-52  (69)
 48 TIGR00237 xseA exodeoxyribonuc  26.9      64  0.0014   30.2   3.2   29  125-160   398-426 (432)
 49 PRK10943 cold shock-like prote  25.9 1.5E+02  0.0033   20.1   4.2   45  167-211     4-52  (69)
 50 COG5510 Predicted small secret  25.6      77  0.0017   20.0   2.3   22   30-51      2-23  (44)
 51 cd01720 Sm_D2 The eukaryotic S  25.5 1.7E+02  0.0036   21.1   4.5   34  149-186    13-46  (87)
 52 cd01731 archaeal_Sm1 The archa  24.6 2.2E+02  0.0048   19.0   6.5   33  150-186    10-42  (68)
 53 PRK09890 cold shock protein Cs  24.4 2.1E+02  0.0045   19.5   4.7   45  167-211     5-53  (70)
 54 cd01717 Sm_B The eukaryotic Sm  24.0 1.9E+02  0.0042   20.0   4.6   33  150-186    10-42  (79)
 55 cd05701 S1_Rrp5_repeat_hs10 S1  23.8 1.6E+02  0.0034   20.3   3.8   39  171-212    17-55  (69)
 56 cd01730 LSm3 The eukaryotic Sm  23.6 1.7E+02  0.0037   20.5   4.2   31  151-185    12-42  (82)
 57 PRK10081 entericidin B membran  23.5 1.5E+02  0.0033   19.0   3.5   22   30-51      2-23  (48)
 58 PRK00286 xseA exodeoxyribonucl  22.8 1.3E+02  0.0027   28.0   4.3   30  124-160   402-431 (438)
 59 PRK10354 RNA chaperone/anti-te  22.3 2.5E+02  0.0053   19.1   4.7   43  168-210     6-52  (70)
 60 TIGR03497 FliI_clade2 flagella  22.3 4.9E+02   0.011   24.3   8.1   38  166-216    33-70  (413)
 61 COG1792 MreC Cell shape-determ  21.9 5.5E+02   0.012   22.6   9.2   34  196-229   206-239 (284)
 62 PTZ00138 small nuclear ribonuc  21.7 2.4E+02  0.0052   20.5   4.7   36  149-186    25-60  (89)
 63 PRK08927 fliI flagellum-specif  21.7 5.7E+02   0.012   24.2   8.4   51  122-185    17-70  (442)
 64 PRK15464 cold shock-like prote  21.6 2.5E+02  0.0055   19.2   4.6   44  167-210     5-52  (70)
 65 PF10844 DUF2577:  Protein of u  21.5 2.5E+02  0.0055   20.5   5.0   56  124-218    37-92  (100)
 66 PF04083 Abhydro_lipase:  Parti  21.2 2.1E+02  0.0045   19.2   4.1   19  125-143    16-34  (63)
 67 PRK10781 rcsF outer membrane l  21.1 1.1E+02  0.0023   24.1   2.9   15  121-136   100-114 (133)
 68 cd06168 LSm9 The eukaryotic Sm  21.0 2.7E+02  0.0058   19.3   4.7   31  151-185    11-41  (75)
 69 PRK15463 cold shock-like prote  21.0 2.4E+02  0.0051   19.3   4.4   45  167-211     5-53  (70)
 70 cd01732 LSm5 The eukaryotic Sm  20.6 2.3E+02  0.0051   19.7   4.4   31  151-185    14-44  (76)
 71 PF10518 TAT_signal:  TAT (twin  20.6 1.5E+02  0.0032   16.3   2.7   19   28-46      2-20  (26)
 72 TIGR00638 Mop molybdenum-pteri  20.2 2.5E+02  0.0055   18.2   4.4   46  166-212     8-58  (69)
 73 PRK13684 Ycf48-like protein; P  20.1   2E+02  0.0044   25.6   5.0    8  131-138   102-109 (334)

No 1  
>PRK10139 serine endoprotease; Provisional
Probab=99.96  E-value=3.4e-28  Score=226.12  Aligned_cols=141  Identities=28%  Similarity=0.478  Sum_probs=115.5

Q ss_pred             HHHHHHHHhCCceEEEEeeeeccCCCCCcchhhhhcc------ccccccceEEEEEEcC-CcEEEEccccccccccCCCC
Q 027022           77 RVVQLFQETSPSVVSIQDLELSKNPKSTSSELMLVDG------EYAKVEGTGSGFVWDK-FGHIVTNYHVVAKLATDTSG  149 (229)
Q Consensus        77 ~~~~~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~------~~~~~~~~GSGfiI~~-~G~IlTn~HVv~~~~~~~~~  149 (229)
                      ++.++++++.||||.|.+......+....+.|..+++      ......+.||||||++ +||||||+|||+       +
T Consensus        41 ~~~~~~~~~~pavV~i~~~~~~~~~~~~~~~~~~~f~~~~~~~~~~~~~~~GSG~ii~~~~g~IlTn~HVv~-------~  113 (455)
T PRK10139         41 SLAPMLEKVLPAVVSVRVEGTASQGQKIPEEFKKFFGDDLPDQPAQPFEGLGSGVIIDAAKGYVLTNNHVIN-------Q  113 (455)
T ss_pred             cHHHHHHHhCCcEEEEEEEEeecccccCchhHHHhccccCCccccccccceEEEEEEECCCCEEEeChHHhC-------C
Confidence            6899999999999999986654322111111211111      1233468999999985 799999999999       8


Q ss_pred             cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEcC
Q 027022          150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTFQ  229 (229)
Q Consensus       150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVSa  229 (229)
                      ++.+.|++.|++    +|.|++++.|+.+||||||++.. .++++++|++++.+++||+|++||||+|+..+++.||||+
T Consensus       114 a~~i~V~~~dg~----~~~a~vvg~D~~~DlAvlkv~~~-~~l~~~~lg~s~~~~~G~~V~aiG~P~g~~~tvt~GivS~  188 (455)
T PRK10139        114 AQKISIQLNDGR----EFDAKLIGSDDQSDIALLQIQNP-SKLTQIAIADSDKLRVGDFAVAVGNPFGLGQTATSGIISA  188 (455)
T ss_pred             CCEEEEEECCCC----EEEEEEEEEcCCCCEEEEEecCC-CCCceeEecCccccCCCCEEEEEecCCCCCCceEEEEEcc
Confidence            899999998754    78999999999999999999843 5799999999999999999999999999999999999985


No 2  
>TIGR02038 protease_degS periplasmic serine pepetdase DegS. This family consists of the periplasmic serine protease DegS (HhoB), a shorter paralog of protease DO (HtrA, DegP) and DegQ (HhoA). It is found in E. coli and several other Proteobacteria of the gamma subdivision. It contains a trypsin domain and a single copy of PDZ domain (in contrast to DegP with two copies). A critical role of this DegS is to sense stress in the periplasm and partially degrade an inhibitor of sigma(E).
Probab=99.96  E-value=2.3e-27  Score=214.17  Aligned_cols=134  Identities=34%  Similarity=0.527  Sum_probs=114.8

Q ss_pred             ccchhHHHHHHHHhCCceEEEEeeeeccCCCCCcchhhhhccccccccceEEEEEEcCCcEEEEccccccccccCCCCcc
Q 027022           72 QLEEDRVVQLFQETSPSVVSIQDLELSKNPKSTSSELMLVDGEYAKVEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLH  151 (229)
Q Consensus        72 ~~~~~~~~~~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~  151 (229)
                      .+.+..+.++++++.||||+|++.....+           ........+.||||||+++||||||+|||+       +++
T Consensus        41 ~~~~~~~~~~~~~~~psVV~I~~~~~~~~-----------~~~~~~~~~~GSG~vi~~~G~IlTn~HVV~-------~~~  102 (351)
T TIGR02038        41 NTVEISFNKAVRRAAPAVVNIYNRSISQN-----------SLNQLSIQGLGSGVIMSKEGYILTNYHVIK-------KAD  102 (351)
T ss_pred             cccchhHHHHHHhcCCcEEEEEeEecccc-----------ccccccccceEEEEEEeCCeEEEecccEeC-------CCC
Confidence            34455799999999999999998654321           011233467899999999999999999999       888


Q ss_pred             eEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEcC
Q 027022          152 RCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTFQ  229 (229)
Q Consensus       152 ~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVSa  229 (229)
                      .+.|++.|+.    .+.|+++++|+.+||||||++..  .++++++++++.+++||+|++||||+|+..+++.|+||+
T Consensus       103 ~i~V~~~dg~----~~~a~vv~~d~~~DlAvlkv~~~--~~~~~~l~~s~~~~~G~~V~aiG~P~~~~~s~t~GiIs~  174 (351)
T TIGR02038       103 QIVVALQDGR----KFEAELVGSDPLTDLAVLKIEGD--NLPTIPVNLDRPPHVGDVVLAIGNPYNLGQTITQGIISA  174 (351)
T ss_pred             EEEEEECCCC----EEEEEEEEecCCCCEEEEEecCC--CCceEeccCcCccCCCCEEEEEeCCCCCCCcEEEEEEEe
Confidence            9999998854    78999999999999999999853  589999999999999999999999999999999999985


No 3  
>PRK10898 serine endoprotease; Provisional
Probab=99.95  E-value=2.2e-27  Score=214.48  Aligned_cols=133  Identities=29%  Similarity=0.452  Sum_probs=113.9

Q ss_pred             cchhHHHHHHHHhCCceEEEEeeeeccCCCCCcchhhhhccccccccceEEEEEEcCCcEEEEccccccccccCCCCcce
Q 027022           73 LEEDRVVQLFQETSPSVVSIQDLELSKNPKSTSSELMLVDGEYAKVEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHR  152 (229)
Q Consensus        73 ~~~~~~~~~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~  152 (229)
                      ..+..+.++++++.||||.|.+.....           .........+.||||+|+++||||||+|||+       +++.
T Consensus        42 ~~~~~~~~~~~~~~psvV~v~~~~~~~-----------~~~~~~~~~~~GSGfvi~~~G~IlTn~HVv~-------~a~~  103 (353)
T PRK10898         42 ETPASYNQAVRRAAPAVVNVYNRSLNS-----------TSHNQLEIRTLGSGVIMDQRGYILTNKHVIN-------DADQ  103 (353)
T ss_pred             cccchHHHHHHHhCCcEEEEEeEeccc-----------cCcccccccceeeEEEEeCCeEEEecccEeC-------CCCE
Confidence            334578999999999999999865321           0111234457899999999999999999999       7889


Q ss_pred             EEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEcC
Q 027022          153 CKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTFQ  229 (229)
Q Consensus       153 ~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVSa  229 (229)
                      +.|++.|+.    +|.|+++++|+.+||||||++.  .++++++|++++.+++||+|+++|||+|+..+++.|+||+
T Consensus       104 i~V~~~dg~----~~~a~vv~~d~~~DlAvl~v~~--~~l~~~~l~~~~~~~~G~~V~aiG~P~g~~~~~t~Giis~  174 (353)
T PRK10898        104 IIVALQDGR----VFEALLVGSDSLTDLAVLKINA--TNLPVIPINPKRVPHIGDVVLAIGNPYNLGQTITQGIISA  174 (353)
T ss_pred             EEEEeCCCC----EEEEEEEEEcCCCCEEEEEEcC--CCCCeeeccCcCcCCCCCEEEEEeCCCCcCCCcceeEEEe
Confidence            999998854    7899999999999999999984  4689999999999999999999999999999999999984


No 4  
>PRK10942 serine endoprotease; Provisional
Probab=99.95  E-value=1.6e-26  Score=215.79  Aligned_cols=141  Identities=33%  Similarity=0.470  Sum_probs=114.3

Q ss_pred             HHHHHHHHhCCceEEEEeeeeccCC---C-CCcchhhhh---c---c-------------------ccccccceEEEEEE
Q 027022           77 RVVQLFQETSPSVVSIQDLELSKNP---K-STSSELMLV---D---G-------------------EYAKVEGTGSGFVW  127 (229)
Q Consensus        77 ~~~~~~~~~~psVV~I~~~~~~~~~---~-~~~~~~~~~---~---~-------------------~~~~~~~~GSGfiI  127 (229)
                      ++.++++++.||||+|++......+   . .....||+.   .   +                   ......+.||||||
T Consensus        39 ~~~~~~~~~~pavv~i~~~~~~~~~~~~~~~~~~~ff~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSG~ii  118 (473)
T PRK10942         39 SLAPMLEKVMPSVVSINVEGSTTVNTPRMPRQFQQFFGDNSPFCQEGSPFQSSPFCQGGQGGNGGGQQQKFMALGSGVII  118 (473)
T ss_pred             cHHHHHHHhCCceEEEEEEEeccccCCCCChhHHHhhcccccccccccccccccccccccccccccccccccceEEEEEE
Confidence            5999999999999999986644321   0 011222211   0   0                   01123578999999


Q ss_pred             cC-CcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCC
Q 027022          128 DK-FGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVG  206 (229)
Q Consensus       128 ~~-~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G  206 (229)
                      ++ +||||||+|||.       +++.++|++.|+.    .|.|++++.|+.+||||||++. ..++++++|++++.+++|
T Consensus       119 ~~~~G~IlTn~HVv~-------~a~~i~V~~~dg~----~~~a~vv~~D~~~DlAvlki~~-~~~l~~~~lg~s~~l~~G  186 (473)
T PRK10942        119 DADKGYVVTNNHVVD-------NATKIKVQLSDGR----KFDAKVVGKDPRSDIALIQLQN-PKNLTAIKMADSDALRVG  186 (473)
T ss_pred             ECCCCEEEeChhhcC-------CCCEEEEEECCCC----EEEEEEEEecCCCCEEEEEecC-CCCCceeEecCccccCCC
Confidence            86 599999999999       8899999998754    7899999999999999999974 357999999999999999


Q ss_pred             CeEEEEecCCCCCCceeEeEEcC
Q 027022          207 QSCFAIGNPYGFEDTLTTGVTFQ  229 (229)
Q Consensus       207 ~~V~aiG~P~G~~~svt~GiVSa  229 (229)
                      |+|++||||+|+..+++.||||+
T Consensus       187 ~~V~aiG~P~g~~~tvt~GiVs~  209 (473)
T PRK10942        187 DYTVAIGNPYGLGETVTSGIVSA  209 (473)
T ss_pred             CEEEEEcCCCCCCcceeEEEEEE
Confidence            99999999999999999999985


No 5  
>TIGR02037 degP_htrA_DO periplasmic serine protease, Do/DeqQ family. This family consists of a set proteins various designated DegP, heat shock protein HtrA, and protease DO. The ortholog in Pseudomonas aeruginosa is designated MucD and is found in an operon that controls mucoid phenotype. This family also includes the DegQ (HhoA) paralog in E. coli which can rescue a DegP mutant, but not the smaller DegS paralog, which cannot. Members of this family are located in the periplasm and have separable functions as both protease and chaperone. Members have a trypsin domain and two copies of a PDZ domain. This protein protects bacteria from thermal and other stresses and may be important for the survival of bacterial pathogens.// The chaperone function is dominant at low temperatures, whereas the proteolytic activity is turned on at elevated temperatures.
Probab=99.93  E-value=4.4e-25  Score=204.06  Aligned_cols=140  Identities=36%  Similarity=0.514  Sum_probs=114.6

Q ss_pred             HHHHHHHhCCceEEEEeeeeccCCCC---C---cchhhhh-c------cccccccceEEEEEEcCCcEEEEccccccccc
Q 027022           78 VVQLFQETSPSVVSIQDLELSKNPKS---T---SSELMLV-D------GEYAKVEGTGSGFVWDKFGHIVTNYHVVAKLA  144 (229)
Q Consensus        78 ~~~~~~~~~psVV~I~~~~~~~~~~~---~---~~~~~~~-~------~~~~~~~~~GSGfiI~~~G~IlTn~HVv~~~~  144 (229)
                      +.++++++.||||.|.+.........   .   ...||+. .      .......+.||||+|+++||||||+|||.   
T Consensus         3 ~~~~~~~~~p~vv~i~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfii~~~G~IlTn~Hvv~---   79 (428)
T TIGR02037         3 FAPLVEKVAPAVVNISVEGTVKRRNRPPALPPFFRQFFGDDMPNFPRQQRERKVRGLGSGVIISADGYILTNNHVVD---   79 (428)
T ss_pred             HHHHHHHhCCceEEEEEEEEecccCCCcccchhHHHhhcccccCcccccccccccceeeEEEECCCCEEEEcHHHcC---
Confidence            67999999999999998664432110   0   1122211 0      01234568899999999999999999999   


Q ss_pred             cCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeE
Q 027022          145 TDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTT  224 (229)
Q Consensus       145 ~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~  224 (229)
                          +++.+.|++.+++    .|.|++++.|+.+||||||++.. .++++++|++++.+++||+|+++|||+|+..+++.
T Consensus        80 ----~~~~i~V~~~~~~----~~~a~vv~~d~~~DlAllkv~~~-~~~~~~~l~~~~~~~~G~~v~aiG~p~g~~~~~t~  150 (428)
T TIGR02037        80 ----GADEITVTLSDGR----EFKAKLVGKDPRTDIAVLKIDAK-KNLPVIKLGDSDKLRVGDWVLAIGNPFGLGQTVTS  150 (428)
T ss_pred             ----CCCeEEEEeCCCC----EEEEEEEEecCCCCEEEEEecCC-CCceEEEccCCCCCCCCCEEEEEECCCcCCCcEEE
Confidence                8889999998754    78999999999999999999853 57999999999999999999999999999999999


Q ss_pred             eEEcC
Q 027022          225 GVTFQ  229 (229)
Q Consensus       225 GiVSa  229 (229)
                      |+||+
T Consensus       151 G~vs~  155 (428)
T TIGR02037       151 GIVSA  155 (428)
T ss_pred             EEEEe
Confidence            99984


No 6  
>COG0265 DegQ Trypsin-like serine proteases, typically periplasmic, contain C-terminal PDZ domain [Posttranslational modification, protein turnover, chaperones]
Probab=99.85  E-value=1.6e-20  Score=169.09  Aligned_cols=137  Identities=41%  Similarity=0.574  Sum_probs=113.5

Q ss_pred             hHHHHHHHHhCCceEEEEeeeeccCCCCCcchhhhhccccccccceEEEEEEcCCcEEEEccccccccccCCCCcceEEE
Q 027022           76 DRVVQLFQETSPSVVSIQDLELSKNPKSTSSELMLVDGEYAKVEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKV  155 (229)
Q Consensus        76 ~~~~~~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V  155 (229)
                      ..+.++++++.|+||++........     ..|+..........+.||||+++++|||+||+|||.       +++++.+
T Consensus        33 ~~~~~~~~~~~~~vV~~~~~~~~~~-----~~~~~~~~~~~~~~~~gSg~i~~~~g~ivTn~hVi~-------~a~~i~v  100 (347)
T COG0265          33 LSFATAVEKVAPAVVSIATGLTAKL-----RSFFPSDPPLRSAEGLGSGFIISSDGYIVTNNHVIA-------GAEEITV  100 (347)
T ss_pred             cCHHHHHHhcCCcEEEEEeeeeecc-----hhcccCCcccccccccccEEEEcCCeEEEecceecC-------CcceEEE
Confidence            5789999999999999998665431     111100000011158999999999999999999999       8899999


Q ss_pred             EEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEcC
Q 027022          156 SLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTFQ  229 (229)
Q Consensus       156 ~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVSa  229 (229)
                      .+.|  |+  ++.++++++|+..|||+||++.... ++.+.++++..+++||++++||||+|+..+++.||||+
T Consensus       101 ~l~d--g~--~~~a~~vg~d~~~dlavlki~~~~~-~~~~~~~~s~~l~vg~~v~aiGnp~g~~~tvt~Givs~  169 (347)
T COG0265         101 TLAD--GR--EVPAKLVGKDPISDLAVLKIDGAGG-LPVIALGDSDKLRVGDVVVAIGNPFGLGQTVTSGIVSA  169 (347)
T ss_pred             EeCC--CC--EEEEEEEecCCccCEEEEEeccCCC-CceeeccCCCCcccCCEEEEecCCCCcccceeccEEec
Confidence            9966  44  8899999999999999999986433 89999999999999999999999999999999999985


No 7  
>KOG1320 consensus Serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=99.61  E-value=2.6e-15  Score=138.57  Aligned_cols=148  Identities=32%  Similarity=0.308  Sum_probs=117.0

Q ss_pred             ccchhHHHHHHHHhCCceEEEEeeeeccCCCCCcchhhhhccccccccceEEEEEEcCCcEEEEccccccccccCCC---
Q 027022           72 QLEEDRVVQLFQETSPSVVSIQDLELSKNPKSTSSELMLVDGEYAKVEGTGSGFVWDKFGHIVTNYHVVAKLATDTS---  148 (229)
Q Consensus        72 ~~~~~~~~~~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~---  148 (229)
                      ......++++++++.+|+|.|+...--.      +  +..+....-....||||||+.+|+|+||+||+........   
T Consensus       124 ~k~~~~v~~~~~~cd~Avv~Ie~~~f~~------~--~~~~e~~~ip~l~~S~~Vv~gd~i~VTnghV~~~~~~~y~~~~  195 (473)
T KOG1320|consen  124 RKYKAFVAAVFEECDLAVVYIESEEFWK------G--MNPFELGDIPSLNGSGFVVGGDGIIVTNGHVVRVEPRIYAHSS  195 (473)
T ss_pred             hhhhhhHHHhhhcccceEEEEeeccccC------C--CcccccCCCcccCccEEEEcCCcEEEEeeEEEEEEeccccCCC
Confidence            4556778899999999999999633211      0  1112223445678999999999999999999996433211   


Q ss_pred             -CcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEE
Q 027022          149 -GLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVT  227 (229)
Q Consensus       149 -~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiV  227 (229)
                       ..-.+.|...++.|+  .+.+.+++.|+..|+|+++++.++..+++++++-+..++.|+++.++|+||++.++++.|+|
T Consensus       196 ~~l~~vqi~aa~~~~~--s~ep~i~g~d~~~gvA~l~ik~~~~i~~~i~~~~~~~~~~G~~~~a~~~~f~~~nt~t~g~v  273 (473)
T KOG1320|consen  196 TVLLRVQIDAAIGPGN--SGEPVIVGVDKVAGVAFLKIKTPENILYVIPLGVSSHFRTGVEVSAIGNGFGLLNTLTQGMV  273 (473)
T ss_pred             cceeeEEEEEeecCCc--cCCCeEEccccccceEEEEEecCCcccceeecceeeeecccceeeccccCceeeeeeeeccc
Confidence             112477888887666  67999999999999999999866544899999999999999999999999999999999999


Q ss_pred             cC
Q 027022          228 FQ  229 (229)
Q Consensus       228 Sa  229 (229)
                      |+
T Consensus       274 s~  275 (473)
T KOG1320|consen  274 SG  275 (473)
T ss_pred             cc
Confidence            74


No 8  
>PF13365 Trypsin_2:  Trypsin-like peptidase domain; PDB: 1Y8T_A 2Z9I_A 3QO6_A 1L1J_A 1QY6_A 2O8L_A 3OTP_E 2ZLE_I 1KY9_A 3CS0_A ....
Probab=99.09  E-value=3.5e-10  Score=85.18  Aligned_cols=61  Identities=36%  Similarity=0.489  Sum_probs=46.7

Q ss_pred             EEEEEEcCCcEEEEccccccccccCC-CCcceEEEEEecCCCCeeEEe--EEEEEEcCC-CcEEEEEEc
Q 027022          122 GSGFVWDKFGHIVTNYHVVAKLATDT-SGLHRCKVSLFDAKGNGFYRE--GKMVGCDPA-YDLAVLKVD  186 (229)
Q Consensus       122 GSGfiI~~~G~IlTn~HVv~~~~~~~-~~~~~~~V~~~~~~g~~~~~~--A~vv~~d~~-~DlAvLki~  186 (229)
                      ||||+|+++||||||+||+.+..... .....+.+...++  .  .+.  +++++.|+. .|+|||+++
T Consensus         1 GTGf~i~~~g~ilT~~Hvv~~~~~~~~~~~~~~~~~~~~~--~--~~~~~~~~~~~~~~~~D~All~v~   65 (120)
T PF13365_consen    1 GTGFLIGPDGYILTAAHVVEDWNDGKQPDNSSVEVVFPDG--R--RVPPVAEVVYFDPDDYDLALLKVD   65 (120)
T ss_dssp             EEEEEEETTTEEEEEHHHHTCCTT--G-TCSEEEEEETTS--C--EEETEEEEEEEETT-TTEEEEEES
T ss_pred             CEEEEEcCCceEEEchhheecccccccCCCCEEEEEecCC--C--EEeeeEEEEEECCccccEEEEEEe
Confidence            89999999999999999999542110 0234566666554  3  556  999999999 999999997


No 9  
>KOG1421 consensus Predicted signaling-associated protein (contains a PDZ domain) [General function prediction only]
Probab=98.73  E-value=2.9e-08  Score=94.44  Aligned_cols=131  Identities=25%  Similarity=0.253  Sum_probs=98.5

Q ss_pred             hhHHHHHHHHhCCceEEEEeeeeccCCCCCcchhhhhccccccccceEEEEEEcC-CcEEEEccccccccccCCCCcceE
Q 027022           75 EDRVVQLFQETSPSVVSIQDLELSKNPKSTSSELMLVDGEYAKVEGTGSGFVWDK-FGHIVTNYHVVAKLATDTSGLHRC  153 (229)
Q Consensus        75 ~~~~~~~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfiI~~-~G~IlTn~HVv~~~~~~~~~~~~~  153 (229)
                      ...+...+.++-++||.|+..+...            .+......+.++||++++ .||||||+||+..      +.-..
T Consensus        51 ~e~w~~~ia~VvksvVsI~~S~v~~------------fdtesag~~~atgfvvd~~~gyiLtnrhvv~p------gP~va  112 (955)
T KOG1421|consen   51 SEDWRNTIANVVKSVVSIRFSAVRA------------FDTESAGESEATGFVVDKKLGYILTNRHVVAP------GPFVA  112 (955)
T ss_pred             hhhhhhhhhhhcccEEEEEehheee------------cccccccccceeEEEEecccceEEEeccccCC------CCcee
Confidence            3488999999999999999866543            223344677899999997 5999999999995      44455


Q ss_pred             EEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCC---CccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEc
Q 027022          154 KVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGF---ELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTF  228 (229)
Q Consensus       154 ~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~---~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVS  228 (229)
                      .+.+.+..    ..+--.++.|+-+|+.+++.+....   .+..+.+. .+.-++|.+++.+||-.|...++-.|.+|
T Consensus       113 ~avf~n~e----e~ei~pvyrDpVhdfGf~r~dps~ir~s~vt~i~la-p~~akvgseirvvgNDagEklsIlagflS  185 (955)
T KOG1421|consen  113 SAVFDNHE----EIEIYPVYRDPVHDFGFFRYDPSTIRFSIVTEICLA-PELAKVGSEIRVVGNDAGEKLSILAGFLS  185 (955)
T ss_pred             EEEecccc----cCCcccccCCchhhcceeecChhhcceeeeeccccC-ccccccCCceEEecCCccceEEeehhhhh
Confidence            66665433    4566778999999999999984322   24555554 34579999999999988877777777655


No 10 
>PF00089 Trypsin:  Trypsin;  InterPro: IPR001254 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine proteases belong to the MEROPS peptidase family S1 (chymotrypsin family, clan PA(S))and to peptidase family S6 (Hap serine peptidases). The chymotrypsin family is almost totally confined to animals, although trypsin-like enzymes are found in actinomycetes of the genera Streptomyces and Saccharopolyspora, and in the fungus Fusarium oxysporum []. The enzymes are inherently secreted, being synthesised with a signal peptide that targets them to the secretory pathway. Animal enzymes are either secreted directly, packaged into vesicles for regulated secretion, or are retained in leukocyte granules []. The Hap family, 'Haemophilus adhesion and penetration', are proteins that play a role in the interaction with human epithelial cells. The serine protease activity is localized at the N-terminal domain, whereas the binding domain is in the C-terminal region. ; GO: 0004252 serine-type endopeptidase activity, 0006508 proteolysis; PDB: 1SPJ_A 1A5I_A 2ZGH_A 2ZKS_A 2ZGJ_A 2ZGC_A 2ODP_A 2I6Q_A 2I6S_A 2ODQ_A ....
Probab=98.17  E-value=5.6e-05  Score=62.07  Aligned_cols=92  Identities=27%  Similarity=0.344  Sum_probs=63.0

Q ss_pred             cceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEec-----CCCCeeEEeEEEEE----EcC---CCcEEEEEEc
Q 027022          119 EGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFD-----AKGNGFYREGKMVG----CDP---AYDLAVLKVD  186 (229)
Q Consensus       119 ~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~-----~~g~~~~~~A~vv~----~d~---~~DlAvLki~  186 (229)
                      ...++|++|+++ +|||++|++.       +.+++.+.+..     .++....+..+-+.    ++.   ..|+|||+++
T Consensus        24 ~~~C~G~li~~~-~vLTaahC~~-------~~~~~~v~~g~~~~~~~~~~~~~~~v~~~~~h~~~~~~~~~~DiAll~L~   95 (220)
T PF00089_consen   24 RFFCTGTLISPR-WVLTAAHCVD-------GASDIKVRLGTYSIRNSDGSEQTIKVSKIIIHPKYDPSTYDNDIALLKLD   95 (220)
T ss_dssp             EEEEEEEEEETT-EEEEEGGGHT-------SGGSEEEEESESBTTSTTTTSEEEEEEEEEEETTSBTTTTTTSEEEEEES
T ss_pred             CeeEeEEecccc-cccccccccc-------cccccccccccccccccccccccccccccccccccccccccccccccccc
Confidence            567999999984 9999999999       54556664432     12211123332222    233   4699999998


Q ss_pred             cC---CCCccceEcCCC-CCCCCCCeEEEEecCCCC
Q 027022          187 VE---GFELKPVVLGTS-HDLRVGQSCFAIGNPYGF  218 (229)
Q Consensus       187 ~~---~~~~~~l~lg~s-~~~~~G~~V~aiG~P~G~  218 (229)
                      .+   ...+.++.+... ..++.|+.+.++|++...
T Consensus        96 ~~~~~~~~~~~~~l~~~~~~~~~~~~~~~~G~~~~~  131 (220)
T PF00089_consen   96 RPITFGDNIQPICLPSAGSDPNVGTSCIVVGWGRTS  131 (220)
T ss_dssp             SSSEHBSSBEESBBTSTTHTTTTTSEEEEEESSBSS
T ss_pred             cccccccccccccccccccccccccccccccccccc
Confidence            65   345677778762 347999999999999863


No 11 
>cd00190 Tryp_SPc Trypsin-like serine protease; Many of these are synthesized as inactive precursor zymogens that are cleaved during limited proteolysis to generate their active forms. Alignment contains also inactive enzymes that have substitutions of the catalytic triad residues.
Probab=97.98  E-value=0.00023  Score=58.81  Aligned_cols=95  Identities=21%  Similarity=0.155  Sum_probs=62.4

Q ss_pred             ccceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCC-----eeEEeEEEEEEc-------CCCcEEEEEE
Q 027022          118 VEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGN-----GFYREGKMVGCD-------PAYDLAVLKV  185 (229)
Q Consensus       118 ~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~-----~~~~~A~vv~~d-------~~~DlAvLki  185 (229)
                      .....+|++|++ .+|||++|.+.+.     ....+.|.+...+..     ...+..+-+...       ...|||||++
T Consensus        23 ~~~~C~GtlIs~-~~VLTaAhC~~~~-----~~~~~~v~~g~~~~~~~~~~~~~~~v~~~~~hp~y~~~~~~~DiAll~L   96 (232)
T cd00190          23 GRHFCGGSLISP-RWVLTAAHCVYSS-----APSNYTVRLGSHDLSSNEGGGQVIKVKKVIVHPNYNPSTYDNDIALLKL   96 (232)
T ss_pred             CcEEEEEEEeeC-CEEEECHHhcCCC-----CCccEEEEeCcccccCCCCceEEEEEEEEEECCCCCCCCCcCCEEEEEE
Confidence            346799999997 6999999999842     124566665322211     112233333332       3579999999


Q ss_pred             ccCC---CCccceEcCCCC-CCCCCCeEEEEecCCCC
Q 027022          186 DVEG---FELKPVVLGTSH-DLRVGQSCFAIGNPYGF  218 (229)
Q Consensus       186 ~~~~---~~~~~l~lg~s~-~~~~G~~V~aiG~P~G~  218 (229)
                      +.+-   ..+.|+.|.... .+..|+.+++.|+....
T Consensus        97 ~~~~~~~~~v~picl~~~~~~~~~~~~~~~~G~g~~~  133 (232)
T cd00190          97 KRPVTLSDNVRPICLPSSGYNLPAGTTCTVSGWGRTS  133 (232)
T ss_pred             CCcccCCCcccceECCCccccCCCCCEEEEEeCCcCC
Confidence            8542   226778886654 68899999999987653


No 12 
>smart00020 Tryp_SPc Trypsin-like serine protease. Many of these are synthesised as inactive precursor zymogens that are cleaved during limited proteolysis to generate their active forms. A few, however, are active as single chain molecules, and others are inactive due to substitutions of the catalytic triad residues.
Probab=97.90  E-value=0.00014  Score=60.28  Aligned_cols=94  Identities=18%  Similarity=0.142  Sum_probs=62.9

Q ss_pred             cceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCe----eEEeEEEEE-------EcCCCcEEEEEEcc
Q 027022          119 EGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNG----FYREGKMVG-------CDPAYDLAVLKVDV  187 (229)
Q Consensus       119 ~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~----~~~~A~vv~-------~d~~~DlAvLki~~  187 (229)
                      ....+|.+|++ .+|||++|.+.+.     ....+.|.+...+...    ..+...-+.       .....|||||+++.
T Consensus        25 ~~~C~GtlIs~-~~VLTaahC~~~~-----~~~~~~v~~g~~~~~~~~~~~~~~v~~~~~~p~~~~~~~~~DiAll~L~~   98 (229)
T smart00020       25 RHFCGGSLISP-RWVLTAAHCVYGS-----DPSNIRVRLGSHDLSSGEEGQVIKVSKVIIHPNYNPSTYDNDIALLKLKS   98 (229)
T ss_pred             CcEEEEEEecC-CEEEECHHHcCCC-----CCcceEEEeCcccCCCCCCceEEeeEEEEECCCCCCCCCcCCEEEEEECc
Confidence            56799999997 6999999999942     1246677775432211    123333333       23467999999985


Q ss_pred             C---CCCccceEcCCC-CCCCCCCeEEEEecCCCC
Q 027022          188 E---GFELKPVVLGTS-HDLRVGQSCFAIGNPYGF  218 (229)
Q Consensus       188 ~---~~~~~~l~lg~s-~~~~~G~~V~aiG~P~G~  218 (229)
                      +   ...+.++.|... ..+..|+.+.+.|++...
T Consensus        99 ~i~~~~~~~pi~l~~~~~~~~~~~~~~~~g~g~~~  133 (229)
T smart00020       99 PVTLSDNVRPICLPSSNYNVPAGTTCTVSGWGRTS  133 (229)
T ss_pred             ccCCCCceeeccCCCcccccCCCCEEEEEeCCCCC
Confidence            4   123667777553 367889999999987654


No 13 
>PF10459 Peptidase_S46:  Peptidase S46;  InterPro: IPR019500 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This entry represents S46 peptidases, where dipeptidyl-peptidase 7 (DPP-7) is the best-characterised member of this family. It is a serine peptidase that is located on the cell surface and is predicted to have two N-terminal transmembrane domains. 
Probab=97.65  E-value=0.00036  Score=68.55  Aligned_cols=24  Identities=29%  Similarity=0.201  Sum_probs=21.3

Q ss_pred             ceEEEEEEcCCcEEEEcccccccc
Q 027022          120 GTGSGFVWDKFGHIVTNYHVVAKL  143 (229)
Q Consensus       120 ~~GSGfiI~~~G~IlTn~HVv~~~  143 (229)
                      +.+||-||+++|+|+||+|++-+.
T Consensus        47 gGCSgsfVS~~GLvlTNHHC~~~~   70 (698)
T PF10459_consen   47 GGCSGSFVSPDGLVLTNHHCGYGA   70 (698)
T ss_pred             CceeEEEEcCCceEEecchhhhhH
Confidence            359999999999999999998744


No 14 
>KOG1320 consensus Serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=97.54  E-value=7.6e-05  Score=69.70  Aligned_cols=128  Identities=23%  Similarity=0.217  Sum_probs=89.7

Q ss_pred             HHHHhCCceEEEEeeeeccCCCCCcchhhhhccccccccceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecC
Q 027022           81 LFQETSPSVVSIQDLELSKNPKSTSSELMLVDGEYAKVEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDA  160 (229)
Q Consensus        81 ~~~~~~psVV~I~~~~~~~~~~~~~~~~~~~~~~~~~~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~  160 (229)
                      .++....+++.+....++.       .....|....+....||||.+.. ..++||+|++....    +.....+. .. 
T Consensus        55 ~~~~~~~s~~~v~~~~~~~-------~~~~pw~~~~q~~~~~s~f~i~~-~~lltn~~~v~~~~----~~~~v~v~-~~-  120 (473)
T KOG1320|consen   55 VVDLALQSVVKVFSVSTEP-------SSVLPWQRTRQFSSGGSGFAIYG-KKLLTNAHVVAPNN----DHKFVTVK-KH-  120 (473)
T ss_pred             CccccccceeEEEeecccc-------cccCcceeeehhcccccchhhcc-cceeecCccccccc----cccccccc-cC-
Confidence            3455566888888755443       22222444457788999999986 68999999999432    23333333 33 


Q ss_pred             CCCeeEEeEEEEEEcCCCcEEEEEEccC--CCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEc
Q 027022          161 KGNGFYREGKMVGCDPAYDLAVLKVDVE--GFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTF  228 (229)
Q Consensus       161 ~g~~~~~~A~vv~~d~~~DlAvLki~~~--~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVS  228 (229)
                       |....|.|++...-.+.|+|+|.++..  +....++.+++  -+.-.+-++++|   |...++|-|.|+
T Consensus       121 -gs~~k~~~~v~~~~~~cd~Avv~Ie~~~f~~~~~~~e~~~--ip~l~~S~~Vv~---gd~i~VTnghV~  184 (473)
T KOG1320|consen  121 -GSPRKYKAFVAAVFEECDLAVVYIESEEFWKGMNPFELGD--IPSLNGSGFVVG---GDGIIVTNGHVV  184 (473)
T ss_pred             -CCchhhhhhHHHhhhcccceEEEEeeccccCCCcccccCC--CcccCccEEEEc---CCcEEEEeeEEE
Confidence             444478899999999999999999853  33333455544  577888899999   888899999986


No 15 
>COG3591 V8-like Glu-specific endopeptidase [Amino acid transport and metabolism]
Probab=95.56  E-value=0.14  Score=44.54  Aligned_cols=94  Identities=13%  Similarity=0.053  Sum_probs=51.4

Q ss_pred             EEEEEEcCCcEEEEccccccccccCCCCcceEEEEEe--cCCCC-eeEEeEEEEEEcC----CCcEEEEEEccCC-----
Q 027022          122 GSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLF--DAKGN-GFYREGKMVGCDP----AYDLAVLKVDVEG-----  189 (229)
Q Consensus       122 GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~--~~~g~-~~~~~A~vv~~d~----~~DlAvLki~~~~-----  189 (229)
                      .++|+|++ ..||||.|++.....   +-.++.+..+  .++|. .+.+........+    +.|.+...+..-.     
T Consensus        66 ~~~~lI~p-ntvLTa~Hc~~s~~~---G~~~~~~~p~g~~~~~~~~~~~~~~~~~~~~g~~~~~d~~~~~v~~~~~~~g~  141 (251)
T COG3591          66 TAATLIGP-NTVLTAGHCIYSPDY---GEDDIAAAPPGVNSDGGPFYGITKIEIRVYPGELYKEDGASYDVGEAALESGI  141 (251)
T ss_pred             eeEEEEcC-ceEEEeeeEEecCCC---ChhhhhhcCCcccCCCCCCCceeeEEEEecCCceeccCCceeeccHHHhccCC
Confidence            45699998 699999999995431   1122222221  11222 1222222222222    3466666653211     


Q ss_pred             ---CCccceEcCCCCCCCCCCeEEEEecCCCCC
Q 027022          190 ---FELKPVVLGTSHDLRVGQSCFAIGNPYGFE  219 (229)
Q Consensus       190 ---~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~  219 (229)
                         .......+......+.+|.+-.+|||.+..
T Consensus       142 ~~~~~~~~~~~~~~~~~~~~d~i~v~GYP~dk~  174 (251)
T COG3591         142 NIGDVVNYLKRNTASEAKANDRITVIGYPGDKP  174 (251)
T ss_pred             CccccccccccccccccccCceeEEEeccCCCC
Confidence               112222333456799999999999998855


No 16 
>PF00863 Peptidase_C4:  Peptidase family C4 This family belongs to family C4 of the peptidase classification.;  InterPro: IPR001730 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad [].  Nuclear inclusion A (NIA) proteases from potyviruses are cysteine peptidases belong to the MEROPS peptidase family C4 (NIa protease family, clan PA(C)) [, ].  Potyviruses include plant viruses in which the single-stranded RNA encodes a polyprotein with NIA protease activity, where proteolytic cleavage is specific for Gln+Gly sites. The NIA protease acts on the polyprotein, releasing itself by Gln+Gly cleavage at both the N- and C-termini. It further processes the polyprotein by cleavage at five similar sites in the C-terminal half of the sequence. In addition to its C-terminal protease activity, the NIA protease contains an N-terminal domain that has been implicated in the transcription process []. This peptidase is present in the nuclear inclusion protein of potyviruses.; GO: 0008234 cysteine-type peptidase activity, 0006508 proteolysis; PDB: 3MMG_B 1Q31_B 1LVB_A 1LVM_A.
Probab=94.79  E-value=0.29  Score=42.11  Aligned_cols=82  Identities=13%  Similarity=0.173  Sum_probs=46.7

Q ss_pred             ceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEE-----EEEEcCCCcEEEEEEccCCCCccc
Q 027022          120 GTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGK-----MVGCDPAYDLAVLKVDVEGFELKP  194 (229)
Q Consensus       120 ~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~-----vv~~d~~~DlAvLki~~~~~~~~~  194 (229)
                      ..-=||..++  |||||+|..+.      +...+.|....  |   .|...     -+..=+..||.++|..   .++||
T Consensus        32 ~~l~gigyG~--~iItn~HLf~~------nng~L~i~s~h--G---~f~v~nt~~lkv~~i~~~DiviirmP---kDfpP   95 (235)
T PF00863_consen   32 RSLYGIGYGS--YIITNAHLFKR------NNGELTIKSQH--G---EFTVPNTTQLKVHPIEGRDIVIIRMP---KDFPP   95 (235)
T ss_dssp             EEEEEEEETT--EEEEEGGGGSS------TTCEEEEEETT--E---EEEECEGGGSEEEE-TCSSEEEEE-----TTS--
T ss_pred             EEEEEEeECC--EEEEChhhhcc------CCCeEEEEeCc--e---EEEcCCccccceEEeCCccEEEEeCC---cccCC
Confidence            3445677765  99999999985      33457776655  3   23222     2445568999999995   34666


Q ss_pred             eEcC-CCCCCCCCCeEEEEecCCC
Q 027022          195 VVLG-TSHDLRVGQSCFAIGNPYG  217 (229)
Q Consensus       195 l~lg-~s~~~~~G~~V~aiG~P~G  217 (229)
                      .+-- ....++.||.|..+|.=+-
T Consensus        96 f~~kl~FR~P~~~e~v~mVg~~fq  119 (235)
T PF00863_consen   96 FPQKLKFRAPKEGERVCMVGSNFQ  119 (235)
T ss_dssp             --S---B----TT-EEEEEEEECS
T ss_pred             cchhhhccCCCCCCEEEEEEEEEE
Confidence            5431 2457999999999998554


No 17 
>KOG1421 consensus Predicted signaling-associated protein (contains a PDZ domain) [General function prediction only]
Probab=92.47  E-value=0.27  Score=48.11  Aligned_cols=87  Identities=18%  Similarity=0.128  Sum_probs=71.2

Q ss_pred             ccceEEEEEEcC-CcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceE
Q 027022          118 VEGTGSGFVWDK-FGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVV  196 (229)
Q Consensus       118 ~~~~GSGfiI~~-~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~  196 (229)
                      ....|||.|++. .|+++++..+|.-      +.....|++.|..    ...|++..-++...+|.+|.+..  ..-.++
T Consensus       548 ~i~kgt~~i~d~~~g~~vvsr~~vp~------d~~d~~vt~~dS~----~i~a~~~fL~~t~n~a~~kydp~--~~~~~k  615 (955)
T KOG1421|consen  548 DIYKGTALIMDTSKGLGVVSRSVVPS------DAKDQRVTEADSD----GIPANVSFLHPTENVASFKYDPA--LEVQLK  615 (955)
T ss_pred             hhhcCceEEEEccCCceeEecccCCc------hhhceEEeecccc----cccceeeEecCccceeEeccChh--Hhhhhc
Confidence            456799999986 5999999999984      7888999998876    56899999999999999999843  234556


Q ss_pred             cCCCCCCCCCCeEEEEecCCC
Q 027022          197 LGTSHDLRVGQSCFAIGNPYG  217 (229)
Q Consensus       197 lg~s~~~~~G~~V~aiG~P~G  217 (229)
                      |- ...++.||++-.+|+-..
T Consensus       616 l~-~~~v~~gD~~~f~g~~~~  635 (955)
T KOG1421|consen  616 LT-DTTVLRGDECTFEGFTED  635 (955)
T ss_pred             cc-eeeEecCCceeEeccccc
Confidence            63 456999999999998755


No 18 
>PF05579 Peptidase_S32:  Equine arteritis virus serine endopeptidase S32;  InterPro: IPR008760 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to MEROPS peptidase family S32 (clan PA(S)). The type example is equine arteritis virus serine endopeptidase (equine arteritis virus), which is involved in processing of nidovirus polyproteins [].; GO: 0004252 serine-type endopeptidase activity, 0016032 viral reproduction, 0019082 viral protein processing; PDB: 3FAN_A 3FAO_A 1MBM_A.
Probab=88.18  E-value=2.1  Score=37.54  Aligned_cols=64  Identities=19%  Similarity=0.150  Sum_probs=37.5

Q ss_pred             ceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCC
Q 027022          120 GTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGT  199 (229)
Q Consensus       120 ~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~  199 (229)
                      ++|+.|-++.+-.|+|+.||+.+        +...|...+   .     -+...++.+-|.|.-.++.-...+|.+++++
T Consensus       114 Gsggvft~~~~~vvvTAtHVlg~--------~~a~v~~~g---~-----~~~~tF~~~GDfA~~~~~~~~G~~P~~k~a~  177 (297)
T PF05579_consen  114 GSGGVFTIGGNTVVVTATHVLGG--------NTARVSGVG---T-----RRMLTFKKNGDFAEADITNWPGAAPKYKFAQ  177 (297)
T ss_dssp             EEEEEEECTTEEEEEEEHHHCBT--------TEEEEEETT---E-----EEEEEEEEETTEEEEEETTS-S---B--B-T
T ss_pred             cccceEEECCeEEEEEEEEEcCC--------CeEEEEecc---e-----EEEEEEeccCcEEEEECCCCCCCCCceeecC
Confidence            34444545554589999999973        445555422   1     2456678889999999954445788888873


No 19 
>PF03510 Peptidase_C24:  2C endopeptidase (C24) cysteine protease family;  InterPro: IPR000317 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad [].  The two signatures that defines this group of calivirus polyproteins identify a cysteine peptidase signature that belongs to MEROPS peptidase family C24 (clan PA(C)). Caliciviruses are positive-stranded ssRNA viruses that cause gastroenteritis. The calicivirus genome contains two open reading frames, ORF1 and ORF2. ORF2 encodes a structural protein []; while ORF1 encodes a non-structural polypeptide, which has RNA helicase, cysteine protease and RNA polymerase activity. The regions of the polyprotein in which these activities lie are similar to proteins produced by the picornaviruses. Two different families of caliciviruses can be distinguished on the basis of sequence similarity, namely those classified as small round structured viruses (SRSVs) and those classed as non-SRSVs. Calicivirus proteases from the non-SRSV group, which are members of the PA protease clan, constitute family C24 of the cysteine proteases (proteases from SRSVs belong to the C37 family). As mentioned above, the protease activity resides within a polyprotein. The enzyme cleaves the polyprotein at sites N-terminal to itself, liberating the polyprotein helicase.; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis
Probab=82.22  E-value=3.1  Score=31.39  Aligned_cols=55  Identities=24%  Similarity=0.270  Sum_probs=36.1

Q ss_pred             EEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCC
Q 027022          124 GFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSH  201 (229)
Q Consensus       124 GfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~  201 (229)
                      ++-|+. |.++|+.||.+       ..+.+     +  |..    =+++.  .+.|+|+++.+..  .+|.+++++..
T Consensus         3 avHIGn-G~~vt~tHva~-------~~~~v-----~--g~~----f~~~~--~~ge~~~v~~~~~--~~p~~~ig~g~   57 (105)
T PF03510_consen    3 AVHIGN-GRYVTVTHVAK-------SSDSV-----D--GQP----FKIVK--TDGELCWVQSPLV--HLPAAQIGTGK   57 (105)
T ss_pred             eEEeCC-CEEEEEEEEec-------cCceE-----c--CcC----cEEEE--eccCEEEEECCCC--CCCeeEeccCC
Confidence            456764 99999999999       44332     1  332    13333  4559999999743  47888887543


No 20 
>KOG3627 consensus Trypsin [Amino acid transport and metabolism]
Probab=79.20  E-value=21  Score=29.99  Aligned_cols=86  Identities=22%  Similarity=0.327  Sum_probs=47.6

Q ss_pred             eEEEEEEcCCcEEEEccccccccccCCCCcc--eEEEEEecC-------CCC--eeEEeEEEE---EEcC---C-CcEEE
Q 027022          121 TGSGFVWDKFGHIVTNYHVVAKLATDTSGLH--RCKVSLFDA-------KGN--GFYREGKMV---GCDP---A-YDLAV  182 (229)
Q Consensus       121 ~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~--~~~V~~~~~-------~g~--~~~~~A~vv---~~d~---~-~DlAv  182 (229)
                      ...|.+|++ .+|+|++|.+.+       ..  ...|.+...       .+.  ......+++   .++.   . .||||
T Consensus        39 ~Cggsli~~-~~vltaaHC~~~-------~~~~~~~V~~G~~~~~~~~~~~~~~~~~~v~~~i~H~~y~~~~~~~nDial  110 (256)
T KOG3627|consen   39 LCGGSLISP-RWVLTAAHCVKG-------ASASLYTVRLGEHDINLSVSEGEEQLVGDVEKIIVHPNYNPRTLENNDIAL  110 (256)
T ss_pred             eeeeEEeeC-CEEEEChhhCCC-------CCCcceEEEECccccccccccCchhhhceeeEEEECCCCCCCCCCCCCEEE
Confidence            455667855 599999999994       22  455554210       010  111112333   1222   2 79999


Q ss_pred             EEEccC---CCCccceEcCCCCC---CCCCCeEEEEec
Q 027022          183 LKVDVE---GFELKPVVLGTSHD---LRVGQSCFAIGN  214 (229)
Q Consensus       183 Lki~~~---~~~~~~l~lg~s~~---~~~G~~V~aiG~  214 (229)
                      |+++.+   ...+.++.|-....   ...++.+++.|.
T Consensus       111 l~l~~~v~~~~~i~piclp~~~~~~~~~~~~~~~v~GW  148 (256)
T KOG3627|consen  111 LRLSEPVTFSSHIQPICLPSSADPYFPPGGTTCLVSGW  148 (256)
T ss_pred             EEECCCcccCCcccccCCCCCcccCCCCCCCEEEEEeC
Confidence            999853   13355666632332   444588888884


No 21 
>PF09342 DUF1986:  Domain of unknown function (DUF1986);  InterPro: IPR015420 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This domain is found in serine endopeptidases belonging to MEROPS peptidase family S1A (clan PA). It is found in unusual mosaic proteins, which are encoded by the Drosophila nudel gene (see P98159 from SWISSPROT). Nudel is involved in defining embryonic dorsoventral polarity. Three proteases; ndl, gd and snk process easter to create active easter. Active easter defines cell identities along the dorsal-ventral continuum by activating the spz ligand for the Tl receptor in the ventral region of the embryo. Nudel, pipe and windbeutel together trigger the protease cascade within the extraembryonic perivitelline compartment which induces dorsoventral polarity of the Drosophila embryo [].
Probab=72.68  E-value=65  Score=28.13  Aligned_cols=99  Identities=16%  Similarity=0.247  Sum_probs=60.6

Q ss_pred             cccceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEe----EEEEEEc-----CCCcEEEEEEcc
Q 027022          117 KVEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYRE----GKMVGCD-----PAYDLAVLKVDV  187 (229)
Q Consensus       117 ~~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~----A~vv~~d-----~~~DlAvLki~~  187 (229)
                      .+.-..||++|++ .|||++...+.+...   ....+.+.+  |.|+.+.+.    -+|...|     +.+++.||.++.
T Consensus        25 dG~~~CsgvLlD~-~WlLvsssCl~~I~L---~~~Yvsall--G~~Kt~~~v~Gp~EQI~rVD~~~~V~~S~v~LLHL~~   98 (267)
T PF09342_consen   25 DGRYWCSGVLLDP-HWLLVSSSCLRGISL---SHHYVSALL--GGGKTYLSVDGPHEQISRVDCFKDVPESNVLLLHLEQ   98 (267)
T ss_pred             cCeEEEEEEEecc-ceEEEeccccCCccc---ccceEEEEe--cCcceecccCCChheEEEeeeeeeccccceeeeeecC
Confidence            4567899999998 799999999996532   123444444  445533210    1333333     577999999985


Q ss_pred             CCCC----ccceEcCC-CCCCCCCCeEEEEecCCCCCCcee
Q 027022          188 EGFE----LKPVVLGT-SHDLRVGQSCFAIGNPYGFEDTLT  223 (229)
Q Consensus       188 ~~~~----~~~l~lg~-s~~~~~G~~V~aiG~P~G~~~svt  223 (229)
                      + .+    ..|+-+-+ +.+....+..+++|.-- .+.+.|
T Consensus        99 ~-~~fTr~VlP~flp~~~~~~~~~~~CVAVg~d~-~g~~kt  137 (267)
T PF09342_consen   99 P-ANFTRYVLPTFLPETSNENESDDECVAVGHDD-TGRIKT  137 (267)
T ss_pred             c-ccceeeecccccccccCCCCCCCceEEEEccc-CCceee
Confidence            4 22    22222322 35667777999999754 333333


No 22 
>PF01455 HupF_HypC:  HupF/HypC family;  InterPro: IPR001109 The large subunit of [NiFe]-hydrogenase, as well as other nickel metalloenzymes, is synthesised as a precursor devoid of the metalloenzyme active site. This precursor then undergoes a complex post-translational maturation process that requires a number of accessory proteins. The hydrogenase expression/formation proteins (HupF/HypC) form a family of small proteins that are hydrogenase precursor-specific chaperones required for this maturation process []. They are believed to keep the hydrogenase precursor in a conformation accessible for metal incorporation [, ].; PDB: 3D3R_A 2Z1C_C 2OT2_A.
Probab=64.19  E-value=29  Score=23.91  Aligned_cols=43  Identities=23%  Similarity=0.297  Sum_probs=30.6

Q ss_pred             EeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEE
Q 027022          167 REGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAI  212 (229)
Q Consensus       167 ~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~ai  212 (229)
                      .+++++..+.....|++.+..   ....+.+.=-.++++||+|++-
T Consensus         5 iP~~Vv~v~~~~~~A~v~~~G---~~~~V~~~lv~~v~~Gd~VLVH   47 (68)
T PF01455_consen    5 IPGRVVEVDEDGGMAVVDFGG---VRREVSLALVPDVKVGDYVLVH   47 (68)
T ss_dssp             EEEEEEEEETTTTEEEEEETT---EEEEEEGTTCTSB-TT-EEEEE
T ss_pred             ccEEEEEEeCCCCEEEEEcCC---cEEEEEEEEeCCCCCCCEEEEe
Confidence            578999998889999998863   3455555445569999999873


No 23 
>COG0298 HypC Hydrogenase maturation factor [Posttranslational modification, protein turnover, chaperones]
Probab=64.03  E-value=14  Score=26.51  Aligned_cols=46  Identities=24%  Similarity=0.376  Sum_probs=31.0

Q ss_pred             EeEEEEEEcCCCcEEEEEEccCCCCccceEcCCC-CCCCCCCeEEE-EecC
Q 027022          167 REGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTS-HDLRVGQSCFA-IGNP  215 (229)
Q Consensus       167 ~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s-~~~~~G~~V~a-iG~P  215 (229)
                      .+++++..|.+.++|++.+-.-.   .-+.+.=- ++++.||+|++ +||-
T Consensus         5 iPgqI~~I~~~~~~A~Vd~gGvk---reV~l~Lv~~~v~~GdyVLVHvGfA   52 (82)
T COG0298           5 IPGQIVEIDDNNHLAIVDVGGVK---REVNLDLVGEEVKVGDYVLVHVGFA   52 (82)
T ss_pred             cccEEEEEeCCCceEEEEeccEe---EEEEeeeecCccccCCEEEEEeeEE
Confidence            46789999998889999885321   22222212 28999999986 5553


No 24 
>cd01735 LSm12_N LSm12 belongs to a family of Sm-like proteins that associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet that associates with other Sm proteins to form hexameric and heptameric ring structures.   In addition to the N-terminal Sm-like domain, LSm12 has a novel methyltransferase domain.
Probab=62.11  E-value=32  Score=23.30  Aligned_cols=33  Identities=15%  Similarity=0.139  Sum_probs=26.8

Q ss_pred             cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022          150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD  186 (229)
Q Consensus       150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~  186 (229)
                      +..+.++...++    .+.++|+.+|....+.+|+-.
T Consensus         6 Gs~V~~kTc~g~----~ieGEV~afD~~tk~lIlk~~   38 (61)
T cd01735           6 GSQVSCRTCFEQ----RLQGEVVAFDYPSKMLILKCP   38 (61)
T ss_pred             ccEEEEEecCCc----eEEEEEEEecCCCcEEEEECc
Confidence            346677776655    889999999999999999854


No 25 
>PRK13922 rod shape-determining protein MreC; Provisional
Probab=58.68  E-value=1.2e+02  Score=26.23  Aligned_cols=51  Identities=24%  Similarity=0.167  Sum_probs=31.2

Q ss_pred             CcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEc
Q 027022          178 YDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTF  228 (229)
Q Consensus       178 ~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVS  228 (229)
                      .+-++++=+.....+..--+....++++||.|+.=|.-.-+...+..|.|+
T Consensus       190 ~~~gi~~G~g~~~~l~l~~i~~~~~i~~GD~VvTSGl~g~fP~Gi~VG~V~  240 (276)
T PRK13922        190 GIRGILSGNGSGDNLKLEFIPRSADIKVGDLVVTSGLGGIFPAGLPVGKVT  240 (276)
T ss_pred             CceEEEEecCCCCceEEEecCCCCCCCCCCEEEECCCCCcCCCCCEEEEEE
Confidence            345666665321122333333456799999999999754456667777765


No 26 
>PF01732 DUF31:  Putative peptidase (DUF31);  InterPro: IPR022382  This domain has no known function. It is found in various hypothetical proteins and putative lipoproteins from mycoplasmas. 
Probab=58.28  E-value=7  Score=35.66  Aligned_cols=24  Identities=29%  Similarity=0.451  Sum_probs=19.6

Q ss_pred             ccceEEEEEEcC----Cc------EEEEcccccc
Q 027022          118 VEGTGSGFVWDK----FG------HIVTNYHVVA  141 (229)
Q Consensus       118 ~~~~GSGfiI~~----~G------~IlTn~HVv~  141 (229)
                      ....|||.|+|-    ++      ||.||.||+.
T Consensus        34 ~~~~GT~WIlDy~~~~~~~~p~k~y~ATNlHVa~   67 (374)
T PF01732_consen   34 SSVSGTGWILDYKKPEDNKYPTKWYFATNLHVAS   67 (374)
T ss_pred             ccCcceEEEEEEeccCCCCCCeEEEEEechhhhc
Confidence            346899999972    23      9999999999


No 27 
>COG5640 Secreted trypsin-like serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=54.94  E-value=25  Score=32.33  Aligned_cols=19  Identities=16%  Similarity=0.059  Sum_probs=14.1

Q ss_pred             EEEEcCCcEEEEcccccccc
Q 027022          124 GFVWDKFGHIVTNYHVVAKL  143 (229)
Q Consensus       124 GfiI~~~G~IlTn~HVv~~~  143 (229)
                      |-++..+ ||||++|.+.+.
T Consensus        65 gs~l~~R-YvLTAAHC~~~~   83 (413)
T COG5640          65 GSKLGGR-YVLTAAHCADAS   83 (413)
T ss_pred             cceecce-EEeeehhhccCC
Confidence            3445554 999999999954


No 28 
>PRK10672 rare lipoprotein A; Provisional
Probab=54.21  E-value=52  Score=30.19  Aligned_cols=29  Identities=24%  Similarity=0.155  Sum_probs=19.4

Q ss_pred             hccccccccceEEEEEEcCCcEEEEcccccc
Q 027022          111 VDGEYAKVEGTGSGFVWDKFGHIVTNYHVVA  141 (229)
Q Consensus       111 ~~~~~~~~~~~GSGfiI~~~G~IlTn~HVv~  141 (229)
                      |.+.......+.+|=+++.  +-+|++|---
T Consensus        85 wYg~~f~G~~TA~Ge~~~~--~~~tAAH~tL  113 (361)
T PRK10672         85 IYDAEAGSNLTASGERFDP--NALTAAHPTL  113 (361)
T ss_pred             EeCCccCCCcCcCceeecC--CcCeeeccCC
Confidence            3344445566778888876  5789999544


No 29 
>PF02601 Exonuc_VII_L:  Exonuclease VII, large subunit;  InterPro: IPR020579 Exonuclease VII 3.1.11.6 from EC is composed of two nonidentical subunits; one large subunit and 4 small ones []. Exonuclease VII catalyses exonucleolytic cleavage in either 5'-3' or 3'-5' direction to yield 5'-phosphomononucleotides. The large subunit also contains the OB-fold domains (IPR004365 from INTERPRO) that bind to nucleic acids at the N terminus.  This entry represents Exonuclease VII, large subunit, C-terminal. ; GO: 0008855 exodeoxyribonuclease VII activity
Probab=49.29  E-value=21  Score=31.64  Aligned_cols=34  Identities=24%  Similarity=0.305  Sum_probs=29.3

Q ss_pred             ceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecC
Q 027022          120 GTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDA  160 (229)
Q Consensus       120 ~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~  160 (229)
                      ..|=.++-+++|.+||+..-++       .++.+.|.+.||
T Consensus       280 ~RGYaiv~~~~g~vI~s~~~l~-------~gd~i~i~l~DG  313 (319)
T PF02601_consen  280 KRGYAIVRDKDGKVITSVKQLK-------PGDEIEIRLADG  313 (319)
T ss_pred             hCceEEEECCCCCEECCHHHCC-------CCCEEEEEEcce
Confidence            4466677778999999999999       889999999985


No 30 
>COG3338 Cah Carbonic anhydrase [Inorganic ion transport and metabolism]
Probab=48.17  E-value=32  Score=29.66  Aligned_cols=25  Identities=40%  Similarity=0.446  Sum_probs=23.1

Q ss_pred             CCCeeEEeEEEEEEcCCCcEEEEEE
Q 027022          161 KGNGFYREGKMVGCDPAYDLAVLKV  185 (229)
Q Consensus       161 ~g~~~~~~A~vv~~d~~~DlAvLki  185 (229)
                      +|+.+.+.|..|..|++.+||||-+
T Consensus       124 ~Gk~~pmEaHFVHkd~~g~L~Vl~v  148 (250)
T COG3338         124 DGKSFPMEAHFVHKDAKGTLAVLAV  148 (250)
T ss_pred             ccccccceeeeeecCCCCCEEEEEE
Confidence            6888889999999999999999987


No 31 
>PF15436 PGBA_N:  Plasminogen-binding protein pgbA N-terminal
Probab=46.05  E-value=1.7e+02  Score=24.99  Aligned_cols=74  Identities=16%  Similarity=0.097  Sum_probs=42.3

Q ss_pred             cCCcEEEE--ccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCC-CCccceEcCCCCCCC
Q 027022          128 DKFGHIVT--NYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEG-FELKPVVLGTSHDLR  204 (229)
Q Consensus       128 ~~~G~IlT--n~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~-~~~~~l~lg~s~~~~  204 (229)
                      +.++.++|  +.+..-       +...+.+.-.+.+..  ..-|+.+-..-+...|.+|+.... ..-..++... -.++
T Consensus        12 d~~~~~i~~~~~~l~v-------G~SGiV~h~~~~~~~--~IiA~a~V~~~~~g~A~~kf~~fd~L~Q~aLP~p~-~~pk   81 (218)
T PF15436_consen   12 DDNNKIITFDAPDLKV-------GESGIVVHKFDKDHS--SIIARAVVISKKNGVAKAKFSVFDSLKQDALPTPK-MVPK   81 (218)
T ss_pred             ecCCCEEEecCCcccc-------CCceEEEEEecCCcc--eeeeEEEEEEecCCeeEEEEeehhhhhhhcCCCCc-cccC
Confidence            44444555  444444       555666665543333  455666555557889999985321 1234444432 3689


Q ss_pred             CCCeEEE
Q 027022          205 VGQSCFA  211 (229)
Q Consensus       205 ~G~~V~a  211 (229)
                      .||.|+.
T Consensus        82 ~GD~vil   88 (218)
T PF15436_consen   82 KGDEVIL   88 (218)
T ss_pred             CCCEEEE
Confidence            9999873


No 32 
>PRK10413 hydrogenase 2 accessory protein HypG; Provisional
Probab=43.89  E-value=50  Score=23.66  Aligned_cols=45  Identities=13%  Similarity=0.112  Sum_probs=26.4

Q ss_pred             EeEEEEEEcCCC-cEEEEEEccCCCCccceEcCCC-CCCCCCCeEEE
Q 027022          167 REGKMVGCDPAY-DLAVLKVDVEGFELKPVVLGTS-HDLRVGQSCFA  211 (229)
Q Consensus       167 ~~A~vv~~d~~~-DlAvLki~~~~~~~~~l~lg~s-~~~~~G~~V~a  211 (229)
                      .+++++..+.+. .+|++.+.....+....=+++. .++++||+|++
T Consensus         5 iP~kVi~i~~~~~~~A~vd~~Gv~r~V~l~Lv~~~~~~~~vGDyVLV   51 (82)
T PRK10413          5 VPGQVLAVGEDIHQLAQVEVCGIKRDVNIALICEGNPADLLGQWVLV   51 (82)
T ss_pred             cceEEEEECCCCCcEEEEEcCCeEEEEEeeeeccCCcccccCCEEEE
Confidence            467888887653 6787777532212221112222 25789999987


No 33 
>PRK14864 putative biofilm stress and motility protein A; Provisional
Probab=41.47  E-value=1.5e+02  Score=22.22  Aligned_cols=15  Identities=20%  Similarity=0.191  Sum_probs=6.9

Q ss_pred             chhHHHHHHHHHHHH
Q 027022           30 RRSSIGFGSSVILSS   44 (229)
Q Consensus        30 ~~~~~~~~~~~~~~a   44 (229)
                      +|+++.+++++++++
T Consensus         5 mk~~~~l~~~l~LS~   19 (104)
T PRK14864          5 MRRFASLLLTLLLSA   19 (104)
T ss_pred             HHHHHHHHHHHHHhh
Confidence            455544444444433


No 34 
>cd00600 Sm_like The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=41.40  E-value=88  Score=20.24  Aligned_cols=33  Identities=27%  Similarity=0.349  Sum_probs=26.6

Q ss_pred             cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022          150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD  186 (229)
Q Consensus       150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~  186 (229)
                      ...+.|.+.++.    .|.+.+.+.|+...+.+-...
T Consensus         6 g~~V~V~l~~g~----~~~G~L~~~D~~~Ni~L~~~~   38 (63)
T cd00600           6 GKTVRVELKDGR----VLEGVLVAFDKYMNLVLDDVE   38 (63)
T ss_pred             CCEEEEEECCCc----EEEEEEEEECCCCCEEECCEE
Confidence            357888887754    889999999999988876664


No 35 
>PF08605 Rad9_Rad53_bind:  Fungal Rad9-like Rad53-binding;  InterPro: IPR013914  In Saccharomyces cerevisiae (Baker s yeast), the Rad9 is a key adaptor protein in DNA damage checkpoint pathways. DNA damage induces Rad9 phosphorylation, and Rad53 specifically associates with this region of Rad9, when phosphorylated, via the Rad53 IPR000253 from INTERPRO domain []. There is no clear higher eukaryotic ortholog to Rad9. 
Probab=41.31  E-value=43  Score=26.21  Aligned_cols=47  Identities=23%  Similarity=0.367  Sum_probs=33.8

Q ss_pred             EEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEe
Q 027022          166 YREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIG  213 (229)
Q Consensus       166 ~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG  213 (229)
                      .|+|++++.+...+-.+++++....+.+.-.+ ..-++++||.|-.=|
T Consensus        24 yYPa~~~~~~~~~~~~~V~Fedg~~~i~~~dv-~~LDlRIGD~Vkv~~   70 (131)
T PF08605_consen   24 YYPATCVGSGVDRDRSLVRFEDGTYEIKNEDV-KYLDLRIGDTVKVDG   70 (131)
T ss_pred             EeeEEEEeecCCCCeEEEEEecCceEeCcccE-eeeeeecCCEEEECC
Confidence            78999999998888999999854322322222 123689999887766


No 36 
>TIGR00074 hypC_hupF hydrogenase assembly chaperone HypC/HupF. An additional proposed function is to shuttle the iron atom that has been liganded at the HypC/HypD complex to the precursor of the large hydrogenase (HycE) subunit. PubMed:12441107.
Probab=37.16  E-value=80  Score=22.26  Aligned_cols=40  Identities=20%  Similarity=0.307  Sum_probs=25.9

Q ss_pred             EeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEE
Q 027022          167 REGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFA  211 (229)
Q Consensus       167 ~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~a  211 (229)
                      .+++++..+.  +.|++.+...   ...+.+.=-.++++||+|++
T Consensus         5 iP~~V~~i~~--~~A~v~~~G~---~~~v~l~lv~~~~vGD~VLV   44 (76)
T TIGR00074         5 IPGQVVEIDE--NIALVEFCGI---KRDVSLDLVGEVKVGDYVLV   44 (76)
T ss_pred             cceEEEEEcC--CEEEEEcCCe---EEEEEEEeeCCCCCCCEEEE
Confidence            4678887766  4688877532   23333333357999999986


No 37 
>PF09465 LBR_tudor:  Lamin-B receptor of TUDOR domain;  InterPro: IPR019023  The Lamin-B receptor is a chromatin and lamin binding protein in the inner nuclear membrane. It is one of the integral inner nuclear envelope membrane proteins responsible for targeting nuclear membranes to chromatin, being a downstream effector of Ran, a small Ras-like nuclear GTPase which regulates NE assembly. Lamin-B receptor interacts with importin beta, a Ran-binding protein, thereby directly contributing to the fusion of membrane vesicles and the formation of the nuclear envelope []. ; PDB: 2L8D_A 2DIG_A.
Probab=37.10  E-value=1.2e+02  Score=20.03  Aligned_cols=36  Identities=19%  Similarity=0.075  Sum_probs=27.8

Q ss_pred             CcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEcc
Q 027022          149 GLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDV  187 (229)
Q Consensus       149 ~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~  187 (229)
                      .++.+.+..++.+  . .|.++|..+|...++.-++.+.
T Consensus         8 ~Ge~V~~rWP~s~--l-YYe~kV~~~d~~~~~y~V~Y~D   43 (55)
T PF09465_consen    8 IGEVVMVRWPGSS--L-YYEGKVLSYDSKSDRYTVLYED   43 (55)
T ss_dssp             SS-EEEEE-TTTS----EEEEEEEEEETTTTEEEEEETT
T ss_pred             CCCEEEEECCCCC--c-EEEEEEEEecccCceEEEEEcC
Confidence            5678888887643  2 5799999999999999999974


No 38 
>PF00548 Peptidase_C3:  3C cysteine protease (picornain 3C);  InterPro: IPR000199 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad [].  This signature defines cysteine peptidases belong to MEROPS peptidase family C3 (picornain, clan PA(C)), subfamilies C3A and C3B. The protein fold of this peptidase domain for members of this family resembles that of the serine peptidase, chymotrypsin [], the type example for clan PA. Picornaviral proteins are expressed as a single polyprotein which is cleaved by the viral C3 cysteine protease. The poliovirus polyprotein is selectively cleaved between the Gln-|-Gly bond. In other picornavirus reactions Glu may be substituted for Gln, and Ser or Thr for Gly. ; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis; PDB: 3SJO_E 2H6M_A 1QA7_C 1HAV_B 2HAL_A 2H9H_A 3QZQ_B 3QZR_A 3R0F_B 3SJ9_A ....
Probab=36.23  E-value=2.3e+02  Score=22.86  Aligned_cols=56  Identities=18%  Similarity=0.081  Sum_probs=34.7

Q ss_pred             ccceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcC---CCcEEEEEEcc
Q 027022          118 VEGTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDP---AYDLAVLKVDV  187 (229)
Q Consensus       118 ~~~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~---~~DlAvLki~~  187 (229)
                      ....++++-|-+ .++|-+.| -.       ..+.+.+   +  |+.+.....+...|.   ..||++++++.
T Consensus        23 g~~t~l~~gi~~-~~~lvp~H-~~-------~~~~i~i---~--g~~~~~~d~~~lv~~~~~~~Dl~~v~l~~   81 (172)
T PF00548_consen   23 GEFTMLALGIYD-RYFLVPTH-EE-------PEDTIYI---D--GVEYKVDDSVVLVDRDGVDTDLTLVKLPR   81 (172)
T ss_dssp             EEEEEEEEEEEB-TEEEEEGG-GG-------GCSEEEE---T--TEEEEEEEEEEEEETTSSEEEEEEEEEES
T ss_pred             ceEEEecceEee-eEEEEECc-CC-------CcEEEEE---C--CEEEEeeeeEEEecCCCcceeEEEEEccC
Confidence            456788888875 58999999 22       3334433   1  444333444434454   45999999964


No 39 
>PF03761 DUF316:  Domain of unknown function (DUF316) ;  InterPro: IPR005514 This is a family of uncharacterised proteins from Caenorhabditis elegans.
Probab=34.44  E-value=45  Score=28.69  Aligned_cols=40  Identities=18%  Similarity=0.326  Sum_probs=27.6

Q ss_pred             cCCCcEEEEEEccC-CCCccceEcCCC-CCCCCCCeEEEEec
Q 027022          175 DPAYDLAVLKVDVE-GFELKPVVLGTS-HDLRVGQSCFAIGN  214 (229)
Q Consensus       175 d~~~DlAvLki~~~-~~~~~~l~lg~s-~~~~~G~~V~aiG~  214 (229)
                      ....++.||.++.+ .....++=|.++ ..+..|+.+.+.|+
T Consensus       158 ~~~~~~mIlEl~~~~~~~~~~~Cl~~~~~~~~~~~~~~~yg~  199 (282)
T PF03761_consen  158 NRPYSPMILELEEDFSKNVSPPCLADSSTNWEKGDEVDVYGF  199 (282)
T ss_pred             ccccceEEEEEcccccccCCCEEeCCCccccccCceEEEeec
Confidence            34568999999854 134455555554 45889999998888


No 40 
>cd01722 Sm_F The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit F is capable of forming both homo- and hetero-heptamer ring structures.  To form the hetero-heptamer, Sm subunit F initially binds subunits E and G to form a trimer which then assembles onto snRNA along with the D3/B and D1/D2 heterodimers.
Probab=33.41  E-value=1.1e+02  Score=20.74  Aligned_cols=33  Identities=18%  Similarity=0.145  Sum_probs=26.6

Q ss_pred             cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022          150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD  186 (229)
Q Consensus       150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~  186 (229)
                      .+.+.|.+.++.    .+..++.++|....|.+=...
T Consensus        11 g~~V~V~Lk~g~----~~~G~L~~~D~~mNi~L~~~~   43 (68)
T cd01722          11 GKPVIVKLKWGM----EYKGTLVSVDSYMNLQLANTE   43 (68)
T ss_pred             CCEEEEEECCCc----EEEEEEEEECCCEEEEEeeEE
Confidence            367888997754    889999999999988876553


No 41 
>cd01726 LSm6 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm6 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=32.34  E-value=1.2e+02  Score=20.35  Aligned_cols=33  Identities=15%  Similarity=0.086  Sum_probs=26.7

Q ss_pred             cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022          150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD  186 (229)
Q Consensus       150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~  186 (229)
                      .+.+.|.+.++.    .|..++.++|+...|-+=...
T Consensus        10 ~~~V~V~Lk~g~----~~~G~L~~~D~~mNlvL~~~~   42 (67)
T cd01726          10 GRPVVVKLNSGV----DYRGILACLDGYMNIALEQTE   42 (67)
T ss_pred             CCeEEEEECCCC----EEEEEEEEEccceeeEEeeEE
Confidence            467889998754    889999999999988876654


No 42 
>cd01728 LSm1 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm1 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=29.85  E-value=1.9e+02  Score=20.03  Aligned_cols=57  Identities=18%  Similarity=0.164  Sum_probs=35.8

Q ss_pred             ceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccC---CCCccceEcCCCCCCCCCCeEEEEe
Q 027022          151 HRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVE---GFELKPVVLGTSHDLRVGQSCFAIG  213 (229)
Q Consensus       151 ~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~---~~~~~~l~lg~s~~~~~G~~V~aiG  213 (229)
                      +++.|.+.++  +  .+.+.+.++|+..-|.+=.....   ........+|.  -+-.|+.|+.+|
T Consensus        13 k~v~V~l~~g--r--~~~G~L~~fD~~~NlvL~d~~E~~~~~~~~~~~~lG~--~viRG~~V~~ig   72 (74)
T cd01728          13 KKVVVLLRDG--R--KLIGILRSFDQFANLVLQDTVERIYVGDKYGDIPRGI--FIIRGENVVLLG   72 (74)
T ss_pred             CEEEEEEcCC--e--EEEEEEEEECCcccEEecceEEEEecCCccceeEeeE--EEEECCEEEEEE
Confidence            6788888774  4  78999999999987776544211   11112222221  255577888777


No 43 
>PRK00737 small nuclear ribonucleoprotein; Provisional
Probab=29.70  E-value=1.9e+02  Score=19.78  Aligned_cols=33  Identities=18%  Similarity=0.175  Sum_probs=26.7

Q ss_pred             cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022          150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD  186 (229)
Q Consensus       150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~  186 (229)
                      .+.+.|.+.++  +  .|.+++.++|+...+-+-...
T Consensus        14 ~k~V~V~lk~g--~--~~~G~L~~~D~~mNlvL~d~~   46 (72)
T PRK00737         14 NSPVLVRLKGG--R--EFRGELQGYDIHMNLVLDNAE   46 (72)
T ss_pred             CCEEEEEECCC--C--EEEEEEEEEcccceeEEeeEE
Confidence            35688888774  4  789999999999988887764


No 44 
>TIGR00219 mreC rod shape-determining protein MreC. MreC (murein formation C) is involved in the rod shape determination in E. coli, and more generally in cell shape determination of bacteria whether or not they are rod-shaped. Cells defective in MreC are round. Species with MreC include many of the Proteobacteria, Gram-positives, and spirochetes.
Probab=28.98  E-value=4e+02  Score=23.39  Aligned_cols=52  Identities=17%  Similarity=0.142  Sum_probs=32.0

Q ss_pred             CCcEEEEEEccCCC--CccceEcCCCCCCCCCCeEEEEecCCCCCCceeEeEEc
Q 027022          177 AYDLAVLKVDVEGF--ELKPVVLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTF  228 (229)
Q Consensus       177 ~~DlAvLki~~~~~--~~~~l~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVS  228 (229)
                      ..+-++++=...+.  .+....+....++++||.|+.=|.-.-+...+..|.|+
T Consensus       188 t~~~gi~~G~~~g~~~~l~l~~~~~~~~v~~GD~VvTSGlgg~fP~Gl~VG~V~  241 (283)
T TIGR00219       188 SDFRGLIEGNGYGKTLEMNLVNRPAEKDIKKGDLIVTSGLGGRFPEGYPIGVVT  241 (283)
T ss_pred             CCceEEEEecCCCCCcEEEEEECCCCCCCCCCCEEEECCCCCcCCCCCEEEEEE
Confidence            34557777542111  12223344466899999999988765566667777664


No 45 
>PF02122 Peptidase_S39:  Peptidase S39;  InterPro: IPR000382 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. ORF2 of Potato leafroll virus (PLrV) encodes a polyprotein which is translated following a -1 frameshift. The polyprotein has a putative linear arrangement of membrane achor-VPg-peptidase-polmerase domains. The serine peptidase domain which is found in this group of sequences belongs to MEROPS peptidase family S39 (clan PA(S)). It is likely that the peptidase domain is involved in the cleavage of the polyprotein []. The nucleotide sequence for the RNA of PLrV has been determined [, ]. The sequence contains six large open reading frames (ORFs). The 5' coding region encodes two polypeptides of 28K and 70K, which overlap in different reading frames; it is suggested that the third ORF in the 5' block is translated by frameshift readthrough near the end of the 70K protein, yielding a 118K polypeptide []. Segments of the predicted amino acid sequences of these ORFs resemble those of known viral RNA polymerases, ATP-binding proteins and viral genome-linked proteins. The nucleotide sequence of the genomic RNA of Beet western yellows virus (BWYV) has been determined []. The sequence contains six long ORFs. A cluster of three of these ORFs, including the coat protein cistron, display extensive amino acid sequence similarity to corresponding ORFs of a second luteovirus: Barley yellow dwarf virus [].; GO: 0004252 serine-type endopeptidase activity, 0022415 viral reproductive process, 0016021 integral to membrane; PDB: 1ZYO_A.
Probab=28.84  E-value=88  Score=26.32  Aligned_cols=60  Identities=17%  Similarity=0.063  Sum_probs=24.0

Q ss_pred             ccceEEEEE-EcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022          118 VEGTGSGFV-WDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD  186 (229)
Q Consensus       118 ~~~~GSGfi-I~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~  186 (229)
                      ..+.++.+- ++.+-.++|++||..+       ..... .+.++ .+.-.-.-+.+..+...|++||++.
T Consensus        28 hvGya~cv~l~~g~~~L~ta~Hv~~~-------~~~~~-~~k~g-~kipl~~f~~~~~~~~~D~~il~~P   88 (203)
T PF02122_consen   28 HVGYATCVRLFDGEDALLTARHVWSR-------PSKVT-SLKTG-EKIPLAEFTDLLESRIADFVILRGP   88 (203)
T ss_dssp             ------EEEE----EEEEE-HHHHTS-------SS----EEETT-EEEE--S-EEEEE-TTT-EEEEE--
T ss_pred             ccccceEEECcCCccceecccccCCC-------cccee-EcCCC-CcccchhChhhhCCCccCEEEEecC
Confidence            344555533 2333489999999993       22222 22232 1111122344556889999999996


No 46 
>PF08758 Cadherin_pro:  Cadherin prodomain like;  InterPro: IPR014868 Cadherins are a group of proteins that mediate calcium dependent cell-cell adhesion. They are activated through cleavage of a prosequence in the late Golgi. This protein corresponds to the folded region of the prosequence, and is termed the prodomain. The prodomain shows structural resemblance to the cadherin domain, but lacks all the features known to be important for cadherin-cadherin interactions []. ; GO: 0007155 cell adhesion, 0016021 integral to membrane; PDB: 1OP4_A.
Probab=28.73  E-value=2.3e+02  Score=20.48  Aligned_cols=40  Identities=13%  Similarity=0.104  Sum_probs=26.0

Q ss_pred             ceEEEEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCe
Q 027022          120 GTGSGFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNG  164 (229)
Q Consensus       120 ~~GSGfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~  164 (229)
                      +.=+=|-|.+||-|.|..++..-.     +.....|...|.+++.
T Consensus        43 ssDpdF~V~~DGsVy~~r~v~l~~-----~~~~F~V~a~D~~~~~   82 (90)
T PF08758_consen   43 SSDPDFRVLEDGSVYAKRPVQLSS-----EQRSFTVHAWDSQTQE   82 (90)
T ss_dssp             ---SEEEEETTTEEEEES--S-SS-----S-EEEEEEEEETTTTE
T ss_pred             cCCCCEEEcCCCeEEEeeeEecCC-----CceEEEEEEECCCCCe
Confidence            334478899999999988888732     4457888888888774


No 47 
>PRK09507 cspE cold shock protein CspE; Reviewed
Probab=27.96  E-value=1.5e+02  Score=20.19  Aligned_cols=46  Identities=9%  Similarity=0.051  Sum_probs=30.3

Q ss_pred             EEeEEEEEEcCCCcEEEEEEccCCCC----ccceEcCCCCCCCCCCeEEE
Q 027022          166 YREGKMVGCDPAYDLAVLKVDVEGFE----LKPVVLGTSHDLRVGQSCFA  211 (229)
Q Consensus       166 ~~~A~vv~~d~~~DlAvLki~~~~~~----~~~l~lg~s~~~~~G~~V~a  211 (229)
                      .+..+|..+|.+.+...++.+....+    +..+.-.....++.||.|--
T Consensus         3 ~~~G~Vk~f~~~kGyGFI~~~~g~~dvfvH~s~l~~~g~~~l~~G~~V~f   52 (69)
T PRK09507          3 KIKGNVKWFNESKGFGFITPEDGSKDVFVHFSAIQTNGFKTLAEGQRVEF   52 (69)
T ss_pred             ccceEEEEEeCCCCcEEEecCCCCeeEEEEeecccccCCCCCCCCCEEEE
Confidence            35678888899999998888754322    23333222356899998854


No 48 
>TIGR00237 xseA exodeoxyribonuclease VII, large subunit. This family consist of exodeoxyribonuclease VII, large subunit XseA which catalyses exonucleolytic cleavage in either the 5'-3' or 3'-5' direction to yield 5'-phosphomononucleotides. Exonuclease VII consists of one large subunit and four small subunits.
Probab=26.95  E-value=64  Score=30.21  Aligned_cols=29  Identities=17%  Similarity=0.190  Sum_probs=25.0

Q ss_pred             EEEcCCcEEEEccccccccccCCCCcceEEEEEecC
Q 027022          125 FVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDA  160 (229)
Q Consensus       125 fiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~  160 (229)
                      ++-+++|.|+++.+-+.       ..+.+.+.+.||
T Consensus       398 i~~~~~g~~v~s~~~l~-------~gd~l~i~~~dG  426 (432)
T TIGR00237       398 IALNEKGKAIKSVKQVD-------RGDRLTTKLKDG  426 (432)
T ss_pred             EEEecCCCEecCHHHCC-------CCCEEEEEECCe
Confidence            44467899999999999       789999999985


No 49 
>PRK10943 cold shock-like protein CspC; Provisional
Probab=25.88  E-value=1.5e+02  Score=20.14  Aligned_cols=45  Identities=9%  Similarity=0.036  Sum_probs=29.4

Q ss_pred             EeEEEEEEcCCCcEEEEEEccCCCC----ccceEcCCCCCCCCCCeEEE
Q 027022          167 REGKMVGCDPAYDLAVLKVDVEGFE----LKPVVLGTSHDLRVGQSCFA  211 (229)
Q Consensus       167 ~~A~vv~~d~~~DlAvLki~~~~~~----~~~l~lg~s~~~~~G~~V~a  211 (229)
                      +..+|..+|.+.+...|+.+..+.+    ...+.-..-..+..||.|--
T Consensus         4 ~~G~Vk~f~~~kGfGFI~~~~g~~dvFvH~s~l~~~g~~~l~~G~~V~f   52 (69)
T PRK10943          4 IKGQVKWFNESKGFGFITPADGSKDVFVHFSAIQGNGFKTLAEGQNVEF   52 (69)
T ss_pred             cceEEEEEeCCCCcEEEecCCCCeeEEEEhhHccccCCCCCCCCCEEEE
Confidence            4678888888888888888643222    34443322356889998753


No 50 
>COG5510 Predicted small secreted protein [Function unknown]
Probab=25.56  E-value=77  Score=19.97  Aligned_cols=22  Identities=27%  Similarity=0.389  Sum_probs=13.8

Q ss_pred             chhHHHHHHHHHHHHHhhhccC
Q 027022           30 RRSSIGFGSSVILSSFLVNFCS   51 (229)
Q Consensus        30 ~~~~~~~~~~~~~~a~l~~~~~   51 (229)
                      +++.+.+.+++++++.++.+|.
T Consensus         2 mk~t~l~i~~vll~s~llaaCN   23 (44)
T COG5510           2 MKKTILLIALVLLASTLLAACN   23 (44)
T ss_pred             chHHHHHHHHHHHHHHHHHHhh
Confidence            4455555666666677767774


No 51 
>cd01720 Sm_D2 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit D2 heterodimerizes with subunit D1 and three such heterodimers form a hexameric ring structure with alternating D1 and D2 subunits. The D1 - D2 heterodimer also assembles into a heptameric ring containing D2, D3, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=25.47  E-value=1.7e+02  Score=21.10  Aligned_cols=34  Identities=12%  Similarity=0.139  Sum_probs=27.3

Q ss_pred             CcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022          149 GLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD  186 (229)
Q Consensus       149 ~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~  186 (229)
                      ..+.+.|.+.++.    .+.+++.++|...-|.|=..+
T Consensus        13 ~~~~V~V~lr~~r----~~~G~L~~fD~hmNlvL~d~~   46 (87)
T cd01720          13 NNTQVLINCRNNK----KLLGRVKAFDRHCNMVLENVK   46 (87)
T ss_pred             CCCEEEEEEcCCC----EEEEEEEEecCccEEEEcceE
Confidence            3468889997754    789999999999988876554


No 52 
>cd01731 archaeal_Sm1 The archaeal sm1 proteins: The Sm proteins are conserved in all three domains of life and are always associated with U-rich RNA sequences. They function to mediate RNA-RNA interactions and RNA biogenesis.  All Sm proteins contain a common sequence motif in two segments, Sm1 and Sm2, separated by a short variable linker. Eukaryotic Sm proteins form part of specific small nuclear ribonucleoproteins (snRNPs) that are involved in the processing of pre-mRNAs to mature mRNAs, and are a major component of the eukaryotic spliceosome. Most snRNPs consist of seven Sm proteins (B/B', D1, D2, D3, E, F and G) arranged in a ring on a uridine-rich sequence (Sm site), plus a small nuclear RNA (snRNA) (either U1, U2, U5 or U4/6). Since archaebacteria do not have any splicing apparatus, Sm proteins of archaebacteria may play a more general role. Archaeal Lsm proteins are likely to represent the ancestral Sm domain.
Probab=24.62  E-value=2.2e+02  Score=19.01  Aligned_cols=33  Identities=15%  Similarity=0.187  Sum_probs=27.1

Q ss_pred             cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022          150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD  186 (229)
Q Consensus       150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~  186 (229)
                      .+++.|.+.++  +  .+.+++.++|+...|.+-...
T Consensus        10 ~~~V~V~l~~g--~--~~~G~L~~~D~~mNlvL~~~~   42 (68)
T cd01731          10 NKPVLVKLKGG--K--EVRGRLKSYDQHMNLVLEDAE   42 (68)
T ss_pred             CCEEEEEECCC--C--EEEEEEEEECCcceEEEeeEE
Confidence            36788888774  4  789999999999999887764


No 53 
>PRK09890 cold shock protein CspG; Provisional
Probab=24.43  E-value=2.1e+02  Score=19.49  Aligned_cols=45  Identities=9%  Similarity=0.017  Sum_probs=29.4

Q ss_pred             EeEEEEEEcCCCcEEEEEEccCCCC----ccceEcCCCCCCCCCCeEEE
Q 027022          167 REGKMVGCDPAYDLAVLKVDVEGFE----LKPVVLGTSHDLRVGQSCFA  211 (229)
Q Consensus       167 ~~A~vv~~d~~~DlAvLki~~~~~~----~~~l~lg~s~~~~~G~~V~a  211 (229)
                      +..+|..+|.+.+...|+.+..+.+    ...+.-.....++.||.|--
T Consensus         5 ~~G~Vk~f~~~kGfGFI~~~~g~~dvFvH~s~l~~~~~~~l~~G~~V~f   53 (70)
T PRK09890          5 MTGLVKWFNADKGFGFITPDDGSKDVFVHFTAIQSNEFRTLNENQKVEF   53 (70)
T ss_pred             ceEEEEEEECCCCcEEEecCCCCceEEEEEeeeccCCCCCCCCCCEEEE
Confidence            3578888888888888888743222    33333333357899998854


No 54 
>cd01717 Sm_B The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit B heterodimerizes with subunit D3 and three such heterodimers form a hexameric ring structure with alternating B and D3 subunits.  The D3 - B heterodimer also assembles into a heptameric ring containing D1, D2, E, F, and G subunits.  Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=24.01  E-value=1.9e+02  Score=20.03  Aligned_cols=33  Identities=21%  Similarity=0.296  Sum_probs=26.1

Q ss_pred             cceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022          150 LHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD  186 (229)
Q Consensus       150 ~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~  186 (229)
                      .+.+.|.+.++  +  .+.+.+.++|....|.|=...
T Consensus        10 ~~~V~V~l~dg--R--~~~G~L~~~D~~~NlVL~~~~   42 (79)
T cd01717          10 NYRLRVTLQDG--R--QFVGQFLAFDKHMNLVLSDCE   42 (79)
T ss_pred             CCEEEEEECCC--c--EEEEEEEEEcCccCEEcCCEE
Confidence            36788999874  4  789999999999988765553


No 55 
>cd05701 S1_Rrp5_repeat_hs10 S1_Rrp5_repeat_hs10: Rrp5 is a trans-acting factor important for biogenesis of both the 40S and 60S eukaryotic ribosomal subunits. Rrp5 has two distinct regions, an N-terminal region containing tandemly repeated S1 RNA-binding domains (12 S1 repeats in Saccharomyces cerevisiae Rrp5 and 14 S1 repeats in Homo sapiens Rrp5) and a C-terminal region containing tetratricopeptide repeat (TPR) motifs thought to be involved in protein-protein interactions. Mutational studies have shown that each region represents a specific functional domain. Deletions within the S1-containing region inhibit pre-rRNA processing at either site A3 or A2, whereas deletions within the TPR region confer an inability to support cleavage of A0-A2. This CD includes H. sapiens S1 repeat 10 (hs10). Rrp5 is found in eukaryotes but not in prokaryotes or archaea.
Probab=23.85  E-value=1.6e+02  Score=20.29  Aligned_cols=39  Identities=26%  Similarity=0.198  Sum_probs=21.4

Q ss_pred             EEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEE
Q 027022          171 MVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAI  212 (229)
Q Consensus       171 vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~ai  212 (229)
                      |+.-+...||+.+.+..   .+--...-+++++++|+.+.+.
T Consensus        17 vvSL~~t~~L~a~p~~s---HLNdtfrf~seklkvG~~l~v~   55 (69)
T cd05701          17 IVSLATTGDLAAFPTRS---HLNDTFRFDSEKLSVGQCLDVT   55 (69)
T ss_pred             EEEeeccccEEEEEchh---hccccccccceeeeccceEEEE
Confidence            33444455555555532   1222223357889999988764


No 56 
>cd01730 LSm3 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm3 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=23.60  E-value=1.7e+02  Score=20.53  Aligned_cols=31  Identities=19%  Similarity=0.240  Sum_probs=24.6

Q ss_pred             ceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEE
Q 027022          151 HRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKV  185 (229)
Q Consensus       151 ~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki  185 (229)
                      +.+.|.+.++  +  .+.+++.++|....|.|=..
T Consensus        12 k~V~V~l~~g--r--~~~G~L~~fD~~mNlvL~d~   42 (82)
T cd01730          12 ERVYVKLRGD--R--ELRGRLHAYDQHLNMILGDV   42 (82)
T ss_pred             CEEEEEECCC--C--EEEEEEEEEccceEEeccce
Confidence            5788888774  4  78999999999988876433


No 57 
>PRK10081 entericidin B membrane lipoprotein; Provisional
Probab=23.49  E-value=1.5e+02  Score=19.04  Aligned_cols=22  Identities=23%  Similarity=0.297  Sum_probs=11.7

Q ss_pred             chhHHHHHHHHHHHHHhhhccC
Q 027022           30 RRSSIGFGSSVILSSFLVNFCS   51 (229)
Q Consensus        30 ~~~~~~~~~~~~~~a~l~~~~~   51 (229)
                      ++|.+.+.++++++++++.+|.
T Consensus         2 mKk~i~~i~~~l~~~~~l~~Cn   23 (48)
T PRK10081          2 VKKTIAAIFSVLVLSTVLTACN   23 (48)
T ss_pred             hHHHHHHHHHHHHHHHHHhhhh
Confidence            3455555455555455447774


No 58 
>PRK00286 xseA exodeoxyribonuclease VII large subunit; Reviewed
Probab=22.81  E-value=1.3e+02  Score=28.05  Aligned_cols=30  Identities=20%  Similarity=0.275  Sum_probs=25.1

Q ss_pred             EEEEcCCcEEEEccccccccccCCCCcceEEEEEecC
Q 027022          124 GFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDA  160 (229)
Q Consensus       124 GfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~  160 (229)
                      .++-+++|.++|+.+-++       ..+.+.+.+.||
T Consensus       402 a~v~~~~g~~i~s~~~~~-------~~d~i~i~~~dG  431 (438)
T PRK00286        402 AIVRDEDGKVIRSAKQLK-------PGDRLTIRLADG  431 (438)
T ss_pred             EEEEeCCCCEeccHHHCC-------CCCEEEEEECCe
Confidence            344456799999999999       789999999985


No 59 
>PRK10354 RNA chaperone/anti-terminator; Provisional
Probab=22.31  E-value=2.5e+02  Score=19.08  Aligned_cols=43  Identities=12%  Similarity=0.094  Sum_probs=28.1

Q ss_pred             eEEEEEEcCCCcEEEEEEccCCCC----ccceEcCCCCCCCCCCeEE
Q 027022          168 EGKMVGCDPAYDLAVLKVDVEGFE----LKPVVLGTSHDLRVGQSCF  210 (229)
Q Consensus       168 ~A~vv~~d~~~DlAvLki~~~~~~----~~~l~lg~s~~~~~G~~V~  210 (229)
                      ..+|..+|.+.+...|+.+....+    ...+.-.....++.||.|-
T Consensus         6 ~G~Vk~f~~~kGfGFI~~~~g~~dvfvH~s~l~~~g~~~l~~G~~V~   52 (70)
T PRK10354          6 TGIVKWFNADKGFGFITPDDGSKDVFVHFSAIQNDGYKSLDEGQKVS   52 (70)
T ss_pred             eEEEEEEeCCCCcEEEecCCCCccEEEEEeeccccCCCCCCCCCEEE
Confidence            577888888888888887643222    3333322235689999885


No 60 
>TIGR03497 FliI_clade2 flagellar protein export ATPase FliI. Members of this protein family are the FliI protein of bacterial flagellum systems. This protein acts to drive protein export for flagellar biosynthesis. The most closely related family is the YscN family of bacterial type III secretion systems. This model represents one (of three) segment of the FliI family tree. These have been modeled separately in order to exclude the type III secretion ATPases more effectively.
Probab=22.28  E-value=4.9e+02  Score=24.29  Aligned_cols=38  Identities=24%  Similarity=0.324  Sum_probs=20.8

Q ss_pred             EEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCCCCCCeEEEEecCC
Q 027022          166 YREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDLRVGQSCFAIGNPY  216 (229)
Q Consensus       166 ~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~~~G~~V~aiG~P~  216 (229)
                      ...++|++.+  .|.+++..           +++...++.|+.|...|.|+
T Consensus        33 ~~~~eVi~~~--~~~v~l~~-----------~~~t~gl~~G~~V~~tg~~~   70 (413)
T TIGR03497        33 PVLAEVVGFK--EENVLLMP-----------LGEVEGIGPGSLVIATGRPL   70 (413)
T ss_pred             eEEEEEEEEc--CCeEEEEE-----------ccCccCCCCCCEEEEcCCee
Confidence            3467777777  33344443           34444555666666555544


No 61 
>COG1792 MreC Cell shape-determining protein [Cell envelope biogenesis, outer membrane]
Probab=21.92  E-value=5.5e+02  Score=22.57  Aligned_cols=34  Identities=21%  Similarity=0.205  Sum_probs=26.4

Q ss_pred             EcCCCCCCCCCCeEEEEecCCCCCCceeEeEEcC
Q 027022          196 VLGTSHDLRVGQSCFAIGNPYGFEDTLTTGVTFQ  229 (229)
Q Consensus       196 ~lg~s~~~~~G~~V~aiG~P~G~~~svt~GiVSa  229 (229)
                      .+-...++++||.|.+-|.-.-+...+..|.|++
T Consensus       206 ~~~~~~~i~~GD~vvTSGlgg~fP~Gl~Vg~V~~  239 (284)
T COG1792         206 YLPPNSDIKEGDLVVTSGLGGVFPAGLPVGEVSS  239 (284)
T ss_pred             eccCCCCccCCCEEEecCCCCcCCCCcEEEEEEE
Confidence            3445678999999999888766777788887763


No 62 
>PTZ00138 small nuclear ribonucleoprotein; Provisional
Probab=21.72  E-value=2.4e+02  Score=20.49  Aligned_cols=36  Identities=22%  Similarity=0.371  Sum_probs=28.0

Q ss_pred             CcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEc
Q 027022          149 GLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVD  186 (229)
Q Consensus       149 ~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~  186 (229)
                      ....+.|.+.++.++  .+...+.++|.-..|.+=...
T Consensus        25 ~~~~V~i~l~~~~~r--~~~G~L~gfD~~mNlVL~d~~   60 (89)
T PTZ00138         25 EKTRVQIWLYDHPNL--RIEGKILGFDEYMNMVLDDAE   60 (89)
T ss_pred             CCcEEEEEEEeCCCc--EEEEEEEEEcccceEEEccEE
Confidence            456788888886555  789999999999988776553


No 63 
>PRK08927 fliI flagellum-specific ATP synthase; Validated
Probab=21.72  E-value=5.7e+02  Score=24.18  Aligned_cols=51  Identities=22%  Similarity=0.035  Sum_probs=27.5

Q ss_pred             EEEEEEcCCcEEEEcccc---ccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEE
Q 027022          122 GSGFVWDKFGHIVTNYHV---VAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKV  185 (229)
Q Consensus       122 GSGfiI~~~G~IlTn~HV---v~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki  185 (229)
                      -.|-|..-.|.++...-.   +.       -.+-+.|..  .+|+  ...++|++.+.+.  ++|..
T Consensus        17 ~~g~v~~i~g~~i~v~g~~~~~~-------~ge~~~i~~--~~~~--~~~~eVv~~~~~~--~~l~~   70 (442)
T PRK08927         17 IYGRVVAVRGLLVEVAGPIHALS-------VGARIVVET--RGGR--PVPCEVVGFRGDR--ALLMP   70 (442)
T ss_pred             eeeEEEEEEccEEEEEecCCCCC-------cCCEEEEEc--CCCC--EEEEEEEEEcCCe--EEEEE
Confidence            345555555666665554   22       334455532  2243  3578999888774  44444


No 64 
>PRK15464 cold shock-like protein CspH; Provisional
Probab=21.60  E-value=2.5e+02  Score=19.22  Aligned_cols=44  Identities=11%  Similarity=-0.034  Sum_probs=29.2

Q ss_pred             EeEEEEEEcCCCcEEEEEEccCCCC----ccceEcCCCCCCCCCCeEE
Q 027022          167 REGKMVGCDPAYDLAVLKVDVEGFE----LKPVVLGTSHDLRVGQSCF  210 (229)
Q Consensus       167 ~~A~vv~~d~~~DlAvLki~~~~~~----~~~l~lg~s~~~~~G~~V~  210 (229)
                      +..+|..+|.+.....++.+....+    +..+.-.....++.||.|-
T Consensus         5 ~~G~Vk~fn~~KGfGFI~~~~g~~DvFvH~s~l~~~g~~~l~~G~~V~   52 (70)
T PRK15464          5 MTGIVKTFDRKSGKGFIIPSDGRKEVQVHISAFTPRDAEVLIPGLRVE   52 (70)
T ss_pred             ceEEEEEEECCCCeEEEccCCCCccEEEEehhehhcCCCCCCCCCEEE
Confidence            4678888888888888887653322    3444322334699999874


No 65 
>PF10844 DUF2577:  Protein of unknown function (DUF2577);  InterPro: IPR022555 This family of proteins has no known function
Probab=21.53  E-value=2.5e+02  Score=20.48  Aligned_cols=56  Identities=14%  Similarity=0.143  Sum_probs=34.2

Q ss_pred             EEEEcCCcEEEEccccccccccCCCCcceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEEccCCCCccceEcCCCCCC
Q 027022          124 GFVWDKFGHIVTNYHVVAKLATDTSGLHRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKVDVEGFELKPVVLGTSHDL  203 (229)
Q Consensus       124 GfiI~~~G~IlTn~HVv~~~~~~~~~~~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki~~~~~~~~~l~lg~s~~~  203 (229)
                      +++++++-+++++ |+-+       ....+.+.....+..    ..                         +.+  .+.+
T Consensus        37 ~liL~~~~L~i~~-~l~~-------~~~~~~~~~~~~~~~----~~-------------------------i~~--~~~L   77 (100)
T PF10844_consen   37 KLILDKDFLIIPE-LLKD-------YTRDITIEHNSETDN----IT-------------------------ITF--TDGL   77 (100)
T ss_pred             eEEEchHHEEeeh-hccc-------eEEEEEEeccccccc----ee-------------------------EEE--ecCC
Confidence            4888887788888 7766       444454444332110    11                         344  3478


Q ss_pred             CCCCeEEEEecCCCC
Q 027022          204 RVGQSCFAIGNPYGF  218 (229)
Q Consensus       204 ~~G~~V~aiG~P~G~  218 (229)
                      ++||.|+.+-.-.|.
T Consensus        78 k~GD~V~ll~~~~gQ   92 (100)
T PF10844_consen   78 KVGDKVLLLRVQGGQ   92 (100)
T ss_pred             cCCCEEEEEEecCCC
Confidence            999999998765553


No 66 
>PF04083 Abhydro_lipase:  Partial alpha/beta-hydrolase lipase region;  InterPro: IPR006693 The alpha/beta hydrolase fold is common to several hydrolytic enzymes of widely differing phylogenetic origin and catalytic function. The core of each enzyme is similar: an alpha/beta sheet, not barrel, of eight beta-sheets connected by alpha-helices []. This entry represents the N-terminal part of an alpha/beta hydrolase domain found in a number of lipases.; GO: 0006629 lipid metabolic process; PDB: 1K8Q_B 1HLG_B.
Probab=21.22  E-value=2.1e+02  Score=19.20  Aligned_cols=19  Identities=21%  Similarity=-0.011  Sum_probs=13.9

Q ss_pred             EEEcCCcEEEEcccccccc
Q 027022          125 FVWDKFGHIVTNYHVVAKL  143 (229)
Q Consensus       125 fiI~~~G~IlTn~HVv~~~  143 (229)
                      .|.-+||||||-.++..+.
T Consensus        16 ~V~T~DGYiL~l~RIp~~~   34 (63)
T PF04083_consen   16 EVTTEDGYILTLHRIPPGK   34 (63)
T ss_dssp             EEE-TTSEEEEEEEE-SBT
T ss_pred             EEEeCCCcEEEEEEccCCC
Confidence            3566899999999999864


No 67 
>PRK10781 rcsF outer membrane lipoprotein; Reviewed
Probab=21.11  E-value=1.1e+02  Score=24.14  Aligned_cols=15  Identities=7%  Similarity=0.180  Sum_probs=7.7

Q ss_pred             eEEEEEEcCCcEEEEc
Q 027022          121 TGSGFVWDKFGHIVTN  136 (229)
Q Consensus       121 ~GSGfiI~~~G~IlTn  136 (229)
                      .|.|+||.+ +.++.+
T Consensus       100 gaN~Vvl~~-C~~~~~  114 (133)
T PRK10781        100 KANAVLLHS-CEITSG  114 (133)
T ss_pred             CCCEEEEEE-eeccCC
Confidence            355666654 444443


No 68 
>cd06168 LSm9 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm9 proteins have a single Sm-like domain structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=20.98  E-value=2.7e+02  Score=19.35  Aligned_cols=31  Identities=10%  Similarity=0.172  Sum_probs=25.3

Q ss_pred             ceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEE
Q 027022          151 HRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKV  185 (229)
Q Consensus       151 ~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki  185 (229)
                      +.+.|.+.|+  +  .+.+.+..+|....|-+=..
T Consensus        11 ~~v~V~l~dg--R--~~~G~l~~~D~~~NivL~~~   41 (75)
T cd06168          11 RTMRIHMTDG--R--TLVGVFLCTDRDCNIILGSA   41 (75)
T ss_pred             CeEEEEEcCC--e--EEEEEEEEEcCCCcEEecCc
Confidence            6788999884  4  78999999999988876544


No 69 
>PRK15463 cold shock-like protein CspF; Provisional
Probab=20.96  E-value=2.4e+02  Score=19.31  Aligned_cols=45  Identities=11%  Similarity=0.080  Sum_probs=29.9

Q ss_pred             EeEEEEEEcCCCcEEEEEEccCCCC----ccceEcCCCCCCCCCCeEEE
Q 027022          167 REGKMVGCDPAYDLAVLKVDVEGFE----LKPVVLGTSHDLRVGQSCFA  211 (229)
Q Consensus       167 ~~A~vv~~d~~~DlAvLki~~~~~~----~~~l~lg~s~~~~~G~~V~a  211 (229)
                      +.++|..+|.+.....|..+....+    +..+.-.....|+.||.|--
T Consensus         5 ~~G~Vk~fn~~kGfGFI~~~~g~~DvFvH~sal~~~g~~~l~~G~~V~f   53 (70)
T PRK15463          5 MTGIVKTFDGKSGKGLITPSDGRKDVQVHISALNLRDAEELTTGLRVEF   53 (70)
T ss_pred             ceEEEEEEeCCCceEEEecCCCCccEEEEehhhhhcCCCCCCCCCEEEE
Confidence            4678888888888888888654322    33443322457999998753


No 70 
>cd01732 LSm5 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm4 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=20.64  E-value=2.3e+02  Score=19.69  Aligned_cols=31  Identities=16%  Similarity=0.227  Sum_probs=24.8

Q ss_pred             ceEEEEEecCCCCeeEEeEEEEEEcCCCcEEEEEE
Q 027022          151 HRCKVSLFDAKGNGFYREGKMVGCDPAYDLAVLKV  185 (229)
Q Consensus       151 ~~~~V~~~~~~g~~~~~~A~vv~~d~~~DlAvLki  185 (229)
                      +++.|.+.+  |+  .+.+++.++|+..-|.+=..
T Consensus        14 ~~V~V~l~~--gr--~~~G~L~g~D~~mNlvL~da   44 (76)
T cd01732          14 SRIWIVMKS--DK--EFVGTLLGFDDYVNMVLEDV   44 (76)
T ss_pred             CEEEEEECC--Ce--EEEEEEEEeccceEEEEccE
Confidence            678888877  44  78999999999998876544


No 71 
>PF10518 TAT_signal:  TAT (twin-arginine translocation) pathway signal sequence;  InterPro: IPR019546 The twin-arginine translocation (Tat) pathway serves the role of transporting folded proteins across energy-transducing membranes []. Homologues of the genes that encode the transport apparatus occur in archaea, bacteria, chloroplasts, and plant mitochondria []. In bacteria, the Tat pathway catalyses the export of proteins from the cytoplasm across the inner/cytoplasmic membrane. In chloroplasts, the Tat components are found in the thylakoid membrane and direct the import of proteins from the stroma. The Tat pathway acts separately from the general secretory (Sec) pathway, which transports proteins in an unfolded state []. It is generally accepted that the primary role of the Tat system is to translocate fully folded proteins across membranes. An example of proteins that need to be exported in their 3D conformation are redox proteins that have acquired complex multi-atom cofactors in the bacterial cytoplasm (or the chloroplast stroma or mitochondrial matrix). They include hydrogenases, formate dehydrogenases, nitrate reductases, trimethylamine N-oxide (TMAO) reductases and dimethyl sulphoxide (DMSO) reductases [, ]. The Tat system can also export whole heteroligomeric complexes in which some proteins have no Tat signal. This is the case of the DMSO reductase or formate dehydrogenase complexes. But there are also other cases where the physiological rationale for targeting a protein to the Tat signal is less obvious. Indeed, there are examples of homologous proteins that are in some cases targeted to the Tat pathway and in other cases to the Sec apparatus. Some examples are: copper nitrite reductases, flavin domains of flavocytochrome c and N-acetylmuramoyl-L-alanine amidases []. In halophilic archaea such as Halobacterium almost all secreted proteins appear to be Tat targeted. It has been proposed to be a response to the difficulties these organisms would otherwise face in successfully folding proteins extracellularly at high ionic strength []. The Tat signal peptide consists of three motifs: the positively charged N-terminal motif, the hydrophobic region and the C-terminal region that generally ends with a consensus short motif (A-x-A) specifying cleavage by signal peptidase. Sequence analysis revealed that signal peptides capable of targeting the Tat protein contain the consensus sequence [ST]-R-R-x-F-L-K. The nearly invariant twin-arginine gave rise to the pathway's name. In addition the h-region of Tat signal peptides is typically less hydrophobic than that of Sec-specific signal peptides [, ]. 
Probab=20.57  E-value=1.5e+02  Score=16.26  Aligned_cols=19  Identities=21%  Similarity=0.324  Sum_probs=11.6

Q ss_pred             ccchhHHHHHHHHHHHHHh
Q 027022           28 ITRRSSIGFGSSVILSSFL   46 (229)
Q Consensus        28 ~~~~~~~~~~~~~~~~a~l   46 (229)
                      +.||.++..++++.+++.+
T Consensus         2 ~sRR~fLk~~~a~~a~~~~   20 (26)
T PF10518_consen    2 LSRRQFLKGGAAAAAAAAL   20 (26)
T ss_pred             CcHHHHHHHHHHHHHHHHh
Confidence            3566666666666665555


No 72 
>TIGR00638 Mop molybdenum-pterin binding domain. This model describes a multigene family of molybdenum-pterin binding proteins of about 70 amino acids in Clostridium pasteurianum, as a tandemly-repeated domain C-terminal to an unrelated domain in ModE, a molybdate transport gene repressor of E. coli, and in single or tandemly paired domains in several related proteins.
Probab=20.22  E-value=2.5e+02  Score=18.17  Aligned_cols=46  Identities=22%  Similarity=0.289  Sum_probs=27.2

Q ss_pred             EEeEEEEEEcCCCcEEEEEEccCCC-CccceEcCC----CCCCCCCCeEEEE
Q 027022          166 YREGKMVGCDPAYDLAVLKVDVEGF-ELKPVVLGT----SHDLRVGQSCFAI  212 (229)
Q Consensus       166 ~~~A~vv~~d~~~DlAvLki~~~~~-~~~~l~lg~----s~~~~~G~~V~ai  212 (229)
                      .+.++|.......+.+-+.++..+. .+. ..+..    .-.+++|++|++.
T Consensus         8 ~l~g~I~~i~~~g~~~~v~l~~~~~~~l~-a~i~~~~~~~l~l~~G~~v~~~   58 (69)
T TIGR00638         8 QLKGKVVAIEDGDVNAEVDLLLGGGTKLT-AVITLESVAELGLKPGKEVYAV   58 (69)
T ss_pred             EEEEEEEEEEECCCeEEEEEEECCCCEEE-EEecHHHHhhCCCCCCCEEEEE
Confidence            5677777776666677666664332 111 11211    2257899999875


No 73 
>PRK13684 Ycf48-like protein; Provisional
Probab=20.11  E-value=2e+02  Score=25.65  Aligned_cols=8  Identities=38%  Similarity=0.397  Sum_probs=3.9

Q ss_pred             cEEEEccc
Q 027022          131 GHIVTNYH  138 (229)
Q Consensus       131 G~IlTn~H  138 (229)
                      |+++....
T Consensus       102 ~~~~G~~g  109 (334)
T PRK13684        102 GWIVGQPS  109 (334)
T ss_pred             EEEeCCCc
Confidence            45555444


Done!