Query         017471
Match_columns 371
No_of_seqs    416 out of 3016
Neff          7.8 
Searched_HMMs 46136
Date          Fri Mar 29 08:51:48 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/017471.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/017471hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PRK10139 serine endoprotease;  100.0 2.5E-48 5.4E-53  390.1  29.1  309    5-342    82-414 (455)
  2 TIGR02037 degP_htrA_DO peripla 100.0 3.1E-47 6.7E-52  381.6  29.6  308   10-342    55-386 (428)
  3 PRK10942 serine endoprotease;  100.0 8.1E-46 1.7E-50  373.6  27.7  300   10-342   108-432 (473)
  4 TIGR02038 protease_degS peripl 100.0 7.2E-44 1.6E-48  347.9  26.1  248   12-278    77-349 (351)
  5 PRK10898 serine endoprotease;  100.0 2.6E-43 5.6E-48  343.8  27.0  250   11-279    76-351 (353)
  6 COG0265 DegQ Trypsin-like seri 100.0 2.3E-34 4.9E-39  281.2  23.0  234   25-277   103-340 (347)
  7 KOG1320 Serine protease [Postt  99.9 2.9E-26 6.3E-31  225.8  11.0  320    3-340    77-420 (473)
  8 KOG1421 Predicted signaling-as  99.9 4.7E-24   1E-28  212.4  16.5  303    4-333    75-418 (955)
  9 KOG1320 Serine protease [Postt  99.9 1.3E-21 2.7E-26  193.1  17.5  236   28-277   211-468 (473)
 10 KOG1421 Predicted signaling-as  99.8   3E-18 6.6E-23  171.3  20.7  271   27-333   585-877 (955)
 11 PF13180 PDZ_2:  PDZ domain; PD  99.6   3E-15 6.4E-20  115.8   7.8   81  173-275     1-82  (82)
 12 TIGR03279 cyano_FeS_chp putati  99.5 2.4E-15 5.2E-20  147.7   1.4  102  202-363     2-107 (433)
 13 cd00987 PDZ_serine_protease PD  99.4 6.8E-13 1.5E-17  103.8   9.0   88  173-272     1-89  (90)
 14 cd00986 PDZ_LON_protease PDZ d  99.3 9.9E-12 2.1E-16   95.2   8.4   72  197-278     7-78  (79)
 15 cd00991 PDZ_archaeal_metallopr  99.3 1.1E-11 2.4E-16   95.1   8.0   69  196-274     8-77  (79)
 16 cd00990 PDZ_glycyl_aminopeptid  99.3 2.3E-11 5.1E-16   93.1   8.3   77  173-276     1-78  (80)
 17 TIGR01713 typeII_sec_gspC gene  99.2 5.4E-11 1.2E-15  111.4  10.4  101  153-275   158-259 (259)
 18 TIGR02037 degP_htrA_DO peripla  99.1 1.8E-10   4E-15  115.8   9.1   90  172-272   337-427 (428)
 19 PRK10779 zinc metallopeptidase  99.1 9.9E-11 2.2E-15  118.4   6.9  117  200-342   128-245 (449)
 20 cd00989 PDZ_metalloprotease PD  99.1 2.8E-10 6.1E-15   86.7   6.6   66  198-274    12-78  (79)
 21 cd00988 PDZ_CTP_protease PDZ d  99.0 9.3E-10   2E-14   85.1   8.5   68  197-275    12-83  (85)
 22 PF13365 Trypsin_2:  Trypsin-li  98.8 1.4E-08 2.9E-13   83.0   8.3   87   18-134    31-120 (120)
 23 cd00136 PDZ PDZ domain, also c  98.8   8E-09 1.7E-13   76.8   6.1   65  174-262     2-69  (70)
 24 TIGR00054 RIP metalloprotease   98.7 2.3E-08   5E-13  100.4   7.0   69  198-277   203-272 (420)
 25 smart00228 PDZ Domain present   98.7 6.4E-08 1.4E-12   74.2   7.3   73  173-266    12-85  (85)
 26 PRK10779 zinc metallopeptidase  98.6 5.4E-08 1.2E-12   98.6   7.2   68  199-277   222-290 (449)
 27 TIGR00225 prc C-terminal pepti  98.6 1.2E-07 2.5E-12   92.5   8.0   71  198-279    62-135 (334)
 28 TIGR00054 RIP metalloprotease   98.6 6.1E-08 1.3E-12   97.3   5.4   66  197-274   127-193 (420)
 29 PF00595 PDZ:  PDZ domain (Also  98.5 2.3E-07 4.9E-12   71.2   5.7   72  172-263     9-81  (81)
 30 PRK10139 serine endoprotease;   98.5   3E-07 6.5E-12   93.1   7.3   64  198-273   390-454 (455)
 31 PLN00049 carboxyl-terminal pro  98.5 5.9E-07 1.3E-11   89.3   9.3   69  198-275   102-171 (389)
 32 TIGR02860 spore_IV_B stage IV   98.4 5.1E-07 1.1E-11   88.9   6.4   69  197-276   104-181 (402)
 33 PRK10942 serine endoprotease;   98.4 6.3E-07 1.4E-11   91.2   6.6   64  198-273   408-472 (473)
 34 cd00992 PDZ_signaling PDZ doma  98.4 8.7E-07 1.9E-11   67.7   5.7   52  173-235    12-66  (82)
 35 PF14685 Tricorn_PDZ:  Tricorn   98.3 4.3E-06 9.3E-11   65.2   9.4   65  197-272    11-87  (88)
 36 COG0793 Prc Periplasmic protea  98.3 1.5E-06 3.2E-11   86.7   8.2   83  172-278    99-184 (406)
 37 COG3480 SdrC Predicted secrete  98.2   4E-06 8.6E-11   78.9   7.4   72  197-278   129-201 (342)
 38 PRK09681 putative type II secr  98.0 9.4E-06   2E-10   76.2   6.8   68  198-275   204-275 (276)
 39 PF04495 GRASP55_65:  GRASP55/6  98.0 1.1E-05 2.4E-10   68.3   5.0   86  172-276    25-114 (138)
 40 PF00089 Trypsin:  Trypsin;  In  97.9 6.4E-05 1.4E-09   67.4   9.5  120   40-159    86-219 (220)
 41 PRK11186 carboxy-terminal prot  97.8 5.7E-05 1.2E-09   79.5   8.5   71  198-274   255-332 (667)
 42 COG3975 Predicted protease wit  97.8 4.3E-05 9.4E-10   76.5   7.0   86  175-280   439-527 (558)
 43 KOG3129 26S proteasome regulat  97.7 7.7E-05 1.7E-09   66.3   6.5   73  199-279   140-213 (231)
 44 KOG3553 Tax interaction protei  97.4 0.00017 3.6E-09   56.6   3.4   35  197-231    58-93  (124)
 45 PF12812 PDZ_1:  PDZ-like domai  97.3 0.00033 7.2E-09   53.4   4.5   60  173-236     9-69  (78)
 46 COG3031 PulC Type II secretory  97.3 0.00019 4.1E-09   65.1   3.4   66  199-274   208-274 (275)
 47 COG3591 V8-like Glu-specific e  97.0  0.0036 7.7E-08   58.1   8.5   87   67-163   159-249 (251)
 48 PF00863 Peptidase_C4:  Peptida  96.9   0.015 3.2E-07   53.6  11.6  112   40-163    81-196 (235)
 49 cd00190 Tryp_SPc Trypsin-like   96.3    0.03 6.6E-07   50.2   9.9  100   40-139    88-208 (232)
 50 PF10459 Peptidase_S46:  Peptid  96.2  0.0035 7.7E-08   66.5   3.7   56  108-163   623-686 (698)
 51 KOG3209 WW domain-containing p  96.0  0.0066 1.4E-07   62.9   4.3   55  202-266   782-838 (984)
 52 KOG3580 Tight junction protein  96.0  0.0043 9.2E-08   63.0   2.7   59  198-264   429-488 (1027)
 53 smart00020 Tryp_SPc Trypsin-li  95.9   0.042 9.1E-07   49.5   8.5  100   40-139    88-208 (229)
 54 KOG3550 Receptor targeting pro  95.5   0.028 6.1E-07   47.6   5.3   37  197-233   114-152 (207)
 55 KOG3532 Predicted protein kina  95.4    0.02 4.4E-07   59.2   4.9   50  174-236   387-437 (1051)
 56 PF08192 Peptidase_S64:  Peptid  95.3    0.17 3.6E-06   52.8  11.2  119   38-163   540-688 (695)
 57 PF00949 Peptidase_S7:  Peptida  95.1   0.025 5.5E-07   47.4   3.9   33  108-140    87-119 (132)
 58 PF00947 Pico_P2A:  Picornaviru  94.9    0.16 3.5E-06   42.0   7.8   97   34-138    13-109 (127)
 59 KOG3834 Golgi reassembly stack  94.8    0.05 1.1E-06   53.6   5.6  132  197-362    14-151 (462)
 60 KOG3209 WW domain-containing p  94.7   0.044 9.4E-07   57.0   5.0   58  198-266   923-982 (984)
 61 PF00548 Peptidase_C3:  3C cyst  94.3    0.67 1.4E-05   40.8  11.0  103   33-138    61-170 (172)
 62 KOG3605 Beta amyloid precursor  94.0   0.078 1.7E-06   54.7   5.1  102  118-231   680-790 (829)
 63 KOG3542 cAMP-regulated guanine  93.9    0.05 1.1E-06   56.3   3.5   37  197-233   561-598 (1283)
 64 KOG3651 Protein kinase C, alph  93.8   0.089 1.9E-06   49.6   4.6   39  198-236    30-70  (429)
 65 KOG1892 Actin filament-binding  93.8   0.082 1.8E-06   56.7   4.8   64  194-267   956-1021(1629)
 66 COG0750 Predicted membrane-ass  93.7    0.11 2.3E-06   51.3   5.5   57  202-269   133-194 (375)
 67 KOG2921 Intramembrane metallop  93.6   0.051 1.1E-06   53.0   2.7   40  196-235   218-259 (484)
 68 KOG3580 Tight junction protein  93.4   0.092   2E-06   53.7   4.4   61  197-268    39-100 (1027)
 69 KOG3552 FERM domain protein FR  92.5    0.12 2.7E-06   55.2   3.9   57  198-264    75-131 (1298)
 70 PF00944 Peptidase_S3:  Alphavi  91.4    0.19 4.2E-06   41.9   3.1   26  114-139   102-127 (158)
 71 KOG3606 Cell polarity protein   91.1    0.25 5.5E-06   45.9   3.8   81  147-232   146-230 (358)
 72 KOG3571 Dishevelled 3 and rela  90.9    0.27 5.8E-06   49.5   4.1   37  197-233   276-314 (626)
 73 KOG3551 Syntrophins (type beta  88.9    0.33 7.1E-06   47.4   2.8   72  173-266    96-172 (506)
 74 KOG3549 Syntrophins (type gamm  88.3     0.5 1.1E-05   45.5   3.6   55  199-263    81-137 (505)
 75 PF05579 Peptidase_S32:  Equine  86.4    0.51 1.1E-05   44.0   2.5   23  116-138   206-228 (297)
 76 KOG0609 Calcium/calmodulin-dep  86.0     1.3 2.9E-05   45.1   5.4   57  199-265   147-205 (542)
 77 PF02907 Peptidase_S29:  Hepati  84.3    0.83 1.8E-05   38.2   2.5   25  116-140   106-130 (148)
 78 cd00987 PDZ_serine_protease PD  83.7     1.3 2.9E-05   33.6   3.4   47  296-343     2-49  (90)
 79 KOG3605 Beta amyloid precursor  80.5     1.7 3.6E-05   45.4   3.5   61  203-271   678-740 (829)
 80 PF01732 DUF31:  Putative pepti  76.6     1.8   4E-05   42.8   2.5   26  112-137   349-374 (374)
 81 KOG0606 Microtubule-associated  76.5     2.9 6.3E-05   46.3   4.1   34  200-233   660-694 (1205)
 82 PF05580 Peptidase_S55:  SpoIVB  76.1     2.4 5.1E-05   38.5   2.8   39  114-155   176-214 (218)
 83 PF03761 DUF316:  Domain of unk  73.4      31 0.00068   32.3  10.0   89   40-138   160-254 (282)
 84 COG1625 Fe-S oxidoreductase, r  72.3     2.8   6E-05   41.7   2.5   35  201-235     4-40  (414)
 85 KOG3834 Golgi reassembly stack  68.3     3.9 8.5E-05   40.7   2.5   65  201-276   112-178 (462)
 86 KOG3938 RGS-GAIP interacting p  64.2     5.5 0.00012   37.3   2.5   67  190-264   138-209 (334)
 87 PF03510 Peptidase_C24:  2C end  64.0      38 0.00083   27.2   7.0   52   11-68      3-59  (105)
 88 PF12812 PDZ_1:  PDZ-like domai  63.6      15 0.00031   27.9   4.4   39  291-332     5-44  (78)
 89 PF13180 PDZ_2:  PDZ domain; PD  60.3     6.5 0.00014   29.5   2.0   37  296-342     2-38  (82)
 90 cd01735 LSm12_N LSm12 belongs   55.1      38 0.00082   24.5   5.0   34   17-50      6-39  (61)
 91 TIGR02038 protease_degS peripl  55.0     9.9 0.00021   37.3   2.8   48  296-344   256-304 (351)
 92 cd01726 LSm6 The eukaryotic Sm  51.7      29 0.00062   25.3   4.1   33   17-49     10-42  (67)
 93 PF02122 Peptidase_S39:  Peptid  50.0      21 0.00045   32.3   3.7   46  108-154   137-182 (203)
 94 TIGR03000 plancto_dom_1 Planct  49.8      54  0.0012   24.7   5.3   50  217-275    10-63  (75)
 95 cd00600 Sm_like The eukaryotic  48.2      42  0.0009   23.6   4.5   33   19-51      8-40  (63)
 96 cd01722 Sm_F The eukaryotic Sm  47.7      33 0.00072   25.0   3.9   33   17-49     11-43  (68)
 97 PRK00737 small nuclear ribonuc  47.2      43 0.00093   24.8   4.5   34   17-50     14-47  (72)
 98 cd01717 Sm_B The eukaryotic Sm  47.0      38 0.00082   25.5   4.3   31   20-50     13-43  (79)
 99 cd01730 LSm3 The eukaryotic Sm  46.3      35 0.00076   25.9   4.0   29   20-48     14-42  (82)
100 KOG3553 Tax interaction protei  45.0      32  0.0007   27.4   3.6   27  316-342    57-83  (124)
101 TIGR02860 spore_IV_B stage IV   44.7      20 0.00043   35.9   3.0   37  114-153   356-392 (402)
102 PRK10898 serine endoprotease;   43.4      22 0.00049   34.9   3.2   48  296-344   257-305 (353)
103 cd01732 LSm5 The eukaryotic Sm  43.3      45 0.00098   25.1   4.1   31   18-48     14-44  (76)
104 PF11874 DUF3394:  Domain of un  43.3      19 0.00041   32.0   2.4   28  197-224   121-149 (183)
105 COG0298 HypC Hydrogenase matur  42.4      47   0.001   25.3   4.0   47   30-78      5-52  (82)
106 cd01729 LSm7 The eukaryotic Sm  42.3      53  0.0012   24.9   4.5   30   20-49     15-44  (81)
107 cd01720 Sm_D2 The eukaryotic S  42.2      50  0.0011   25.6   4.3   31   19-49     16-46  (87)
108 cd01731 archaeal_Sm1 The archa  41.9      53  0.0011   23.9   4.3   33   18-50     11-43  (68)
109 KOG1738 Membrane-associated gu  41.4      21 0.00045   37.4   2.7   34  200-233   227-262 (638)
110 cd06168 LSm9 The eukaryotic Sm  40.9      60  0.0013   24.3   4.5   31   19-49     12-42  (75)
111 PF00571 CBS:  CBS domain CBS d  39.5      26 0.00057   23.7   2.3   20  118-137    29-48  (57)
112 COG0260 PepB Leucyl aminopepti  39.0      44 0.00096   34.3   4.6   58  189-251   291-348 (485)
113 smart00651 Sm snRNP Sm protein  37.9      70  0.0015   22.8   4.4   33   18-50      9-41  (67)
114 PF01455 HupF_HypC:  HupF/HypC   36.2      59  0.0013   23.9   3.7   41   30-75      5-47  (68)
115 cd01728 LSm1 The eukaryotic Sm  36.1      82  0.0018   23.5   4.5   31   19-49     14-44  (74)
116 cd01721 Sm_D3 The eukaryotic S  34.3      82  0.0018   23.1   4.3   35   16-50      9-43  (70)
117 cd01727 LSm8 The eukaryotic Sm  32.2      92   0.002   23.0   4.3   31   20-50     12-42  (74)
118 KOG3627 Trypsin [Amino acid tr  30.8      33 0.00072   31.3   2.1   99   41-139   106-228 (256)
119 cd01719 Sm_G The eukaryotic Sm  30.5 1.1E+02  0.0024   22.6   4.5   30   20-49     13-42  (72)
120 PRK05015 aminopeptidase B; Pro  29.0      90  0.0019   31.5   4.8   39  190-230   230-268 (424)
121 PF09465 LBR_tudor:  Lamin-B re  28.3   2E+02  0.0043   20.3   5.0   37   15-51      7-44  (55)
122 COG1958 LSM1 Small nuclear rib  27.1 1.2E+02  0.0026   22.7   4.2   32   19-50     19-50  (79)
123 PF11874 DUF3394:  Domain of un  27.0 1.2E+02  0.0026   27.0   4.7   71  247-336    67-140 (183)
124 PF15483 DUF4641:  Domain of un  25.8      45 0.00097   33.3   2.0   21  344-364   417-438 (445)
125 COG2524 Predicted transcriptio  24.5 2.1E+02  0.0045   27.1   6.0   20  117-137   201-220 (294)
126 cd01723 LSm4 The eukaryotic Sm  24.4 1.7E+02  0.0038   21.7   4.6   33   17-49     11-43  (76)
127 PF12381 Peptidase_C3G:  Tungro  23.0      96  0.0021   28.3   3.4   54  107-163   169-228 (231)
128 PF01423 LSM:  LSM domain ;  In  22.5 1.9E+02  0.0041   20.5   4.4   33   19-51     10-42  (67)
129 KOG2561 Adaptor protein NUB1,   22.3      28  0.0006   35.1  -0.2   23  328-350   195-223 (568)
130 PF10049 DUF2283:  Protein of u  21.8      62  0.0013   22.1   1.6   11  126-136    36-46  (50)
131 COG5233 GRH1 Peripheral Golgi   21.7      47   0.001   32.1   1.2   30  201-230    66-96  (417)
132 PF14827 Cache_3:  Sensory doma  21.7      77  0.0017   25.4   2.4   17  123-139    95-111 (116)
133 TIGR00074 hypC_hupF hydrogenas  21.6 1.6E+02  0.0035   22.1   3.9   41   30-75      5-45  (76)
134 PF15436 PGBA_N:  Plasminogen-b  21.5 3.4E+02  0.0074   24.8   6.7   59   14-74     28-88  (218)
135 cd00991 PDZ_archaeal_metallopr  21.4      55  0.0012   24.2   1.4   27  316-342     8-34  (79)
136 PF11730 DUF3297:  Protein of u  21.3      58  0.0013   23.9   1.4   59  204-274     5-64  (71)
137 PRK09570 rpoH DNA-directed RNA  20.8      59  0.0013   24.7   1.4   19  204-222    39-59  (79)
138 cd00433 Peptidase_M17 Cytosol   20.8 1.1E+02  0.0023   31.5   3.6   28  203-230   290-317 (468)
139 PRK00913 multifunctional amino  20.7 1.1E+02  0.0024   31.5   3.7   27  204-230   305-331 (483)
140 cd01718 Sm_E The eukaryotic Sm  20.6 1.9E+02  0.0041   22.0   4.1   30   20-49     21-52  (79)
141 cd01725 LSm2 The eukaryotic Sm  20.5   2E+02  0.0043   21.8   4.3   33   17-49     11-43  (81)
142 cd01724 Sm_D1 The eukaryotic S  20.5 1.8E+02   0.004   22.5   4.2   35   16-50     10-44  (90)

No 1  
>PRK10139 serine endoprotease; Provisional
Probab=100.00  E-value=2.5e-48  Score=390.08  Aligned_cols=309  Identities=21%  Similarity=0.272  Sum_probs=252.1

Q ss_pred             eccCCeeeeeccEEEEE---EEecCCCe-----------------EeEEEEEecCCCCEEEEEEecCCCcCCccceecCC
Q 017471            5 WSTTRRLNSRNEALILS---TWLLCSPS-----------------APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGE   64 (371)
Q Consensus         5 ~~~~~~~~~~gsg~vi~---~~~~~~~~-----------------~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~   64 (371)
                      |...+...+.||||+++   ++++++.|                 ++|++++.|+++||||||++...   ++++++|++
T Consensus        82 ~~~~~~~~~~GSG~ii~~~~g~IlTn~HVv~~a~~i~V~~~dg~~~~a~vvg~D~~~DlAvlkv~~~~---~l~~~~lg~  158 (455)
T PRK10139         82 DQPAQPFEGLGSGVIIDAAKGYVLTNNHVINQAQKISIQLNDGREFDAKLIGSDDQSDIALLQIQNPS---KLTQIAIAD  158 (455)
T ss_pred             ccccccccceEEEEEEECCCCEEEeChHHhCCCCEEEEEECCCCEEEEEEEEEcCCCCEEEEEecCCC---CCceeEecC
Confidence            44444556789999996   57766665                 79999999999999999998653   789999998


Q ss_pred             CCC--CCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccC
Q 017471           65 LPA--LQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE  142 (371)
Q Consensus        65 s~~--lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~  142 (371)
                      +..  +||+|+++|||++... +++.|+||+.++..... .....+||+|++||+|||||||+|.+||||||+++.+..+
T Consensus       159 s~~~~~G~~V~aiG~P~g~~~-tvt~GivS~~~r~~~~~-~~~~~~iqtda~in~GnSGGpl~n~~G~vIGi~~~~~~~~  236 (455)
T PRK10139        159 SDKLRVGDFAVAVGNPFGLGQ-TATSGIISALGRSGLNL-EGLENFIQTDASINRGNSGGALLNLNGELIGINTAILAPG  236 (455)
T ss_pred             ccccCCCCEEEEEecCCCCCC-ceEEEEEccccccccCC-CCcceEEEECCccCCCCCcceEECCCCeEEEEEEEEEcCC
Confidence            764  6899999999999776 89999999998743222 1234689999999999999999999999999999987543


Q ss_pred             -CccceeccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCE
Q 017471          143 -DVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDI  220 (371)
Q Consensus       143 -~~~~~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDv  220 (371)
                       +..+++||||++.+++++++|.++|++. ++|||+.++++ +++.++.+|++ ...|++|.+|.++|||++ |||+||+
T Consensus       237 ~~~~gigfaIP~~~~~~v~~~l~~~g~v~-r~~LGv~~~~l-~~~~~~~lgl~-~~~Gv~V~~V~~~SpA~~AGL~~GDv  313 (455)
T PRK10139        237 GGSVGIGFAIPSNMARTLAQQLIDFGEIK-RGLLGIKGTEM-SADIAKAFNLD-VQRGAFVSEVLPNSGSAKAGVKAGDI  313 (455)
T ss_pred             CCccceEEEEEhHHHHHHHHHHhhcCccc-ccceeEEEEEC-CHHHHHhcCCC-CCCceEEEEECCCChHHHCCCCCCCE
Confidence             4578999999999999999999999999 99999999999 89999999997 467999999999999999 9999999


Q ss_pred             EEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccccccCCCCCCCCCCCceeeccEE
Q 017471          221 ILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFV  300 (371)
Q Consensus       221 Il~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~~~~~~~~~~~~~~~Gl~  300 (371)
                      |++|||++|.++.++.          ..+....+|++++++|.|+|+.+++++++...+......... .+   .+.|+.
T Consensus       314 Il~InG~~V~s~~dl~----------~~l~~~~~g~~v~l~V~R~G~~~~l~v~~~~~~~~~~~~~~~-~~---~~~g~~  379 (455)
T PRK10139        314 ITSLNGKPLNSFAELR----------SRIATTEPGTKVKLGLLRNGKPLEVEVTLDTSTSSSASAEMI-TP---ALQGAT  379 (455)
T ss_pred             EEEECCEECCCHHHHH----------HHHHhcCCCCEEEEEEEECCEEEEEEEEECCCCCcccccccc-cc---cccccE
Confidence            9999999999999875          567666789999999999999999999985433221111100 01   134655


Q ss_pred             EechHHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhH
Q 017471          301 FSRCLYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLL  342 (371)
Q Consensus       301 ~~~l~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~  342 (371)
                      +++.     +  ......|+++..+.++|||+.--.+.+|+|
T Consensus       380 l~~~-----~--~~~~~~Gv~V~~V~~~spA~~aGL~~GD~I  414 (455)
T PRK10139        380 LSDG-----Q--LKDGTKGIKIDEVVKGSPAAQAGLQKDDVI  414 (455)
T ss_pred             eccc-----c--cccCCCceEEEEeCCCChHHHcCCCCCCEE
Confidence            5431     1  122346899999999999998888888775


No 2  
>TIGR02037 degP_htrA_DO periplasmic serine protease, Do/DeqQ family. This family consists of a set proteins various designated DegP, heat shock protein HtrA, and protease DO. The ortholog in Pseudomonas aeruginosa is designated MucD and is found in an operon that controls mucoid phenotype. This family also includes the DegQ (HhoA) paralog in E. coli which can rescue a DegP mutant, but not the smaller DegS paralog, which cannot. Members of this family are located in the periplasm and have separable functions as both protease and chaperone. Members have a trypsin domain and two copies of a PDZ domain. This protein protects bacteria from thermal and other stresses and may be important for the survival of bacterial pathogens.// The chaperone function is dominant at low temperatures, whereas the proteolytic activity is turned on at elevated temperatures.
Probab=100.00  E-value=3.1e-47  Score=381.60  Aligned_cols=308  Identities=22%  Similarity=0.307  Sum_probs=263.1

Q ss_pred             eeeeeccEEEEE--EEecCCCe-----------------EeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCC--CC
Q 017471           10 RLNSRNEALILS--TWLLCSPS-----------------APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELP--AL   68 (371)
Q Consensus        10 ~~~~~gsg~vi~--~~~~~~~~-----------------~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~--~l   68 (371)
                      ...+.||||+++  ++++++.|                 ++|++++.|+.+|||+||++...   ++++++++++.  .+
T Consensus        55 ~~~~~GSGfii~~~G~IlTn~Hvv~~~~~i~V~~~~~~~~~a~vv~~d~~~DlAllkv~~~~---~~~~~~l~~~~~~~~  131 (428)
T TIGR02037        55 KVRGLGSGVIISADGYILTNNHVVDGADEITVTLSDGREFKAKLVGKDPRTDIAVLKIDAKK---NLPVIKLGDSDKLRV  131 (428)
T ss_pred             cccceeeEEEECCCCEEEEcHHHcCCCCeEEEEeCCCCEEEEEEEEecCCCCEEEEEecCCC---CceEEEccCCCCCCC
Confidence            456789999997  56665555                 79999999999999999999763   68999999765  57


Q ss_pred             CCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccC-Cccce
Q 017471           69 QDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-DVENI  147 (371)
Q Consensus        69 gd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~-~~~~~  147 (371)
                      ||+|+++|||++... +++.|+||+..+.. .....+..++|+|+++++|||||||+|.+|+||||+++.+... +..++
T Consensus       132 G~~v~aiG~p~g~~~-~~t~G~vs~~~~~~-~~~~~~~~~i~tda~i~~GnSGGpl~n~~G~viGI~~~~~~~~g~~~g~  209 (428)
T TIGR02037       132 GDWVLAIGNPFGLGQ-TVTSGIVSALGRSG-LGIGDYENFIQTDAAINPGNSGGPLVNLRGEVIGINTAIYSPSGGNVGI  209 (428)
T ss_pred             CCEEEEEECCCcCCC-cEEEEEEEecccCc-cCCCCccceEEECCCCCCCCCCCceECCCCeEEEEEeEEEcCCCCccce
Confidence            899999999999766 89999999988642 1222334589999999999999999999999999999877543 45689


Q ss_pred             eccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECC
Q 017471          148 GYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDG  226 (371)
Q Consensus       148 ~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG  226 (371)
                      +||||++.+++++++|+++|++. ++|||+.++++ +++.++.+|++. ..|++|.+|.++|||++ ||++||+|++|||
T Consensus       210 ~faiP~~~~~~~~~~l~~~g~~~-~~~lGi~~~~~-~~~~~~~lgl~~-~~Gv~V~~V~~~spA~~aGL~~GDvI~~Vng  286 (428)
T TIGR02037       210 GFAIPSNMAKNVVDQLIEGGKVQ-RGWLGVTIQEV-TSDLAKSLGLEK-QRGALVAQVLPGSPAEKAGLKAGDVILSVNG  286 (428)
T ss_pred             EEEEEhHHHHHHHHHHHhcCcCc-CCcCceEeecC-CHHHHHHcCCCC-CCceEEEEccCCCChHHcCCCCCCEEEEECC
Confidence            99999999999999999999998 99999999999 899999999984 57999999999999999 9999999999999


Q ss_pred             EEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccccccCCCCCCCCCCCceeeccEEEech-H
Q 017471          227 IDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRC-L  305 (371)
Q Consensus       227 ~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~~~~~~~~~~~~~~~Gl~~~~l-~  305 (371)
                      ++|.++.++.          ..+....+|+++++++.|+|+.+++++++...+...+       ++...++|+.++++ +
T Consensus       287 ~~i~~~~~~~----------~~l~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~-------~~~~~~lGi~~~~l~~  349 (428)
T TIGR02037       287 KPISSFADLR----------RAIGTLKPGKKVTLGILRKGKEKTITVTLGASPEEQA-------SSSNPFLGLTVANLSP  349 (428)
T ss_pred             EEcCCHHHHH----------HHHHhcCCCCEEEEEEEECCEEEEEEEEECcCCCccc-------cccccccceEEecCCH
Confidence            9999998865          6676777899999999999999999999876543211       12334799999998 7


Q ss_pred             HHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhH
Q 017471          306 YLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLL  342 (371)
Q Consensus       306 ~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~  342 (371)
                      ..++.++++....|++++.+.++|||+..-...+|+|
T Consensus       350 ~~~~~~~l~~~~~Gv~V~~V~~~SpA~~aGL~~GDvI  386 (428)
T TIGR02037       350 EIRKELRLKGDVKGVVVTKVVSGSPAARAGLQPGDVI  386 (428)
T ss_pred             HHHHHcCCCcCcCceEEEEeCCCCHHHHcCCCCCCEE
Confidence            7777888876668999999999999998877777765


No 3  
>PRK10942 serine endoprotease; Provisional
Probab=100.00  E-value=8.1e-46  Score=373.56  Aligned_cols=300  Identities=20%  Similarity=0.267  Sum_probs=247.5

Q ss_pred             eeeeeccEEEEE---EEecCCCe-----------------EeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCC--C
Q 017471           10 RLNSRNEALILS---TWLLCSPS-----------------APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELP--A   67 (371)
Q Consensus        10 ~~~~~gsg~vi~---~~~~~~~~-----------------~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~--~   67 (371)
                      ...+.||||+++   ++++++.|                 ++|++++.|+.+||||||++...   ++++++++++.  +
T Consensus       108 ~~~~~GSG~ii~~~~G~IlTn~HVv~~a~~i~V~~~dg~~~~a~vv~~D~~~DlAvlki~~~~---~l~~~~lg~s~~l~  184 (473)
T PRK10942        108 KFMALGSGVIIDADKGYVVTNNHVVDNATKIKVQLSDGRKFDAKVVGKDPRSDIALIQLQNPK---NLTAIKMADSDALR  184 (473)
T ss_pred             cccceEEEEEEECCCCEEEeChhhcCCCCEEEEEECCCCEEEEEEEEecCCCCEEEEEecCCC---CCceeEecCccccC
Confidence            346789999997   46665555                 79999999999999999998643   78999999876  4


Q ss_pred             CCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccC-Cccc
Q 017471           68 LQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-DVEN  146 (371)
Q Consensus        68 lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~-~~~~  146 (371)
                      +||+|+++|||++... +++.|+||+..+..... ..+..+||+|+++++|||||||+|.+||||||+++.+..+ +..+
T Consensus       185 ~G~~V~aiG~P~g~~~-tvt~GiVs~~~r~~~~~-~~~~~~iqtda~i~~GnSGGpL~n~~GeviGI~t~~~~~~g~~~g  262 (473)
T PRK10942        185 VGDYTVAIGNPYGLGE-TVTSGIVSALGRSGLNV-ENYENFIQTDAAINRGNSGGALVNLNGELIGINTAILAPDGGNIG  262 (473)
T ss_pred             CCCEEEEEcCCCCCCc-ceeEEEEEEeecccCCc-ccccceEEeccccCCCCCcCccCCCCCeEEEEEEEEEcCCCCccc
Confidence            6899999999999766 89999999998642221 1234689999999999999999999999999999887654 3468


Q ss_pred             eeccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEEC
Q 017471          147 IGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFD  225 (371)
Q Consensus       147 ~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vn  225 (371)
                      ++||||++.+++++++|.+.|++. |+|+|+.++++ ++++++.++++ ...|++|.+|.++|||++ |||+||+|++||
T Consensus       263 ~gfaIP~~~~~~v~~~l~~~g~v~-rg~lGv~~~~l-~~~~a~~~~l~-~~~GvlV~~V~~~SpA~~AGL~~GDvIl~In  339 (473)
T PRK10942        263 IGFAIPSNMVKNLTSQMVEYGQVK-RGELGIMGTEL-NSELAKAMKVD-AQRGAFVSQVLPNSSAAKAGIKAGDVITSLN  339 (473)
T ss_pred             EEEEEEHHHHHHHHHHHHhccccc-cceeeeEeeec-CHHHHHhcCCC-CCCceEEEEECCCChHHHcCCCCCCEEEEEC
Confidence            999999999999999999999999 99999999999 88999999998 467999999999999999 999999999999


Q ss_pred             CEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccccccCCCCCCCCCCCceeeccEEEech-
Q 017471          226 GIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRC-  304 (371)
Q Consensus       226 G~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~~~~~~~~~~~~~~~Gl~~~~l-  304 (371)
                      |++|.++.++.          ..+....+|++++++|.|+|+.+++++++...+.....       +...++|+...++ 
T Consensus       340 G~~V~s~~dl~----------~~l~~~~~g~~v~l~v~R~G~~~~v~v~l~~~~~~~~~-------~~~~~lGl~g~~l~  402 (473)
T PRK10942        340 GKPISSFAALR----------AQVGTMPVGSKLTLGLLRDGKPVNVNVELQQSSQNQVD-------SSNIFNGIEGAELS  402 (473)
T ss_pred             CEECCCHHHHH----------HHHHhcCCCCEEEEEEEECCeEEEEEEEeCcCcccccc-------cccccccceeeecc
Confidence            99999999875          67777788999999999999999999998664221111       1112356544333 


Q ss_pred             HHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhH
Q 017471          305 LYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLL  342 (371)
Q Consensus       305 ~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~  342 (371)
                      +.        ....++++..+.++|||+..-...+|+|
T Consensus       403 ~~--------~~~~gvvV~~V~~~S~A~~aGL~~GDvI  432 (473)
T PRK10942        403 NK--------GGDKGVVVDNVKPGTPAAQIGLKKGDVI  432 (473)
T ss_pred             cc--------cCCCCeEEEEeCCCChHHHcCCCCCCEE
Confidence            11        1125899999999999998777777765


No 4  
>TIGR02038 protease_degS periplasmic serine pepetdase DegS. This family consists of the periplasmic serine protease DegS (HhoB), a shorter paralog of protease DO (HtrA, DegP) and DegQ (HhoA). It is found in E. coli and several other Proteobacteria of the gamma subdivision. It contains a trypsin domain and a single copy of PDZ domain (in contrast to DegP with two copies). A critical role of this DegS is to sense stress in the periplasm and partially degrade an inhibitor of sigma(E).
Probab=100.00  E-value=7.2e-44  Score=347.85  Aligned_cols=248  Identities=24%  Similarity=0.380  Sum_probs=215.3

Q ss_pred             eeeccEEEEE--EEecCCCe-----------------EeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCC--CCCC
Q 017471           12 NSRNEALILS--TWLLCSPS-----------------APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELP--ALQD   70 (371)
Q Consensus        12 ~~~gsg~vi~--~~~~~~~~-----------------~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~--~lgd   70 (371)
                      .+.||||+++  ++++++.|                 ++|++++.|+++|||+||++..    ++++++++++.  ++||
T Consensus        77 ~~~GSG~vi~~~G~IlTn~HVV~~~~~i~V~~~dg~~~~a~vv~~d~~~DlAvlkv~~~----~~~~~~l~~s~~~~~G~  152 (351)
T TIGR02038        77 QGLGSGVIMSKEGYILTNYHVIKKADQIVVALQDGRKFEAELVGSDPLTDLAVLKIEGD----NLPTIPVNLDRPPHVGD  152 (351)
T ss_pred             cceEEEEEEeCCeEEEecccEeCCCCEEEEEECCCCEEEEEEEEecCCCCEEEEEecCC----CCceEeccCcCccCCCC
Confidence            4568999997  66766665                 6899999999999999999986    57888998764  6789


Q ss_pred             eEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccC---Cccce
Q 017471           71 AVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE---DVENI  147 (371)
Q Consensus        71 ~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~---~~~~~  147 (371)
                      +|+++|||++... +++.|+||+.++.... ......+||+|+++++|||||||+|.+||||||+++.+...   ...++
T Consensus       153 ~V~aiG~P~~~~~-s~t~GiIs~~~r~~~~-~~~~~~~iqtda~i~~GnSGGpl~n~~G~vIGI~~~~~~~~~~~~~~g~  230 (351)
T TIGR02038       153 VVLAIGNPYNLGQ-TITQGIISATGRNGLS-SVGRQNFIQTDAAINAGNSGGALINTNGELVGINTASFQKGGDEGGEGI  230 (351)
T ss_pred             EEEEEeCCCCCCC-cEEEEEEEeccCcccC-CCCcceEEEECCccCCCCCcceEECCCCeEEEEEeeeecccCCCCccce
Confidence            9999999998766 8999999999874332 22335689999999999999999999999999998766432   23688


Q ss_pred             eccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECC
Q 017471          148 GYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDG  226 (371)
Q Consensus       148 ~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG  226 (371)
                      +|+||++.++++++++.++|++. ++|||+.++++ ++..++.+|++ ...|++|.+|.++|||++ ||++||+|++|||
T Consensus       231 ~faIP~~~~~~vl~~l~~~g~~~-r~~lGv~~~~~-~~~~~~~lgl~-~~~Gv~V~~V~~~spA~~aGL~~GDvI~~Ing  307 (351)
T TIGR02038       231 NFAIPIKLAHKIMGKIIRDGRVI-RGYIGVSGEDI-NSVVAQGLGLP-DLRGIVITGVDPNGPAARAGILVRDVILKYDG  307 (351)
T ss_pred             EEEecHHHHHHHHHHHhhcCccc-ceEeeeEEEEC-CHHHHHhcCCC-ccccceEeecCCCChHHHCCCCCCCEEEEECC
Confidence            99999999999999999999998 99999999999 88889999997 457999999999999999 9999999999999


Q ss_pred             EEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccc
Q 017471          227 IDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATH  278 (371)
Q Consensus       227 ~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~  278 (371)
                      ++|.++.++.          ..+...++|++++++|.|+|+.+++++++.+.
T Consensus       308 ~~V~s~~dl~----------~~l~~~~~g~~v~l~v~R~g~~~~~~v~l~~~  349 (351)
T TIGR02038       308 KDVIGAEELM----------DRIAETRPGSKVMVTVLRQGKQLELPVTIDEK  349 (351)
T ss_pred             EEcCCHHHHH----------HHHHhcCCCCEEEEEEEECCEEEEEEEEecCC
Confidence            9999998865          56766678999999999999999999988654


No 5  
>PRK10898 serine endoprotease; Provisional
Probab=100.00  E-value=2.6e-43  Score=343.81  Aligned_cols=250  Identities=20%  Similarity=0.342  Sum_probs=214.3

Q ss_pred             eeeeccEEEEE--EEecCCCe-----------------EeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCC--CCC
Q 017471           11 LNSRNEALILS--TWLLCSPS-----------------APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELP--ALQ   69 (371)
Q Consensus        11 ~~~~gsg~vi~--~~~~~~~~-----------------~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~--~lg   69 (371)
                      ..+.||||+++  ++++++.|                 ++|++++.|+.+||||||++..    ++++++++++.  .+|
T Consensus        76 ~~~~GSGfvi~~~G~IlTn~HVv~~a~~i~V~~~dg~~~~a~vv~~d~~~DlAvl~v~~~----~l~~~~l~~~~~~~~G  151 (353)
T PRK10898         76 IRTLGSGVIMDQRGYILTNKHVINDADQIIVALQDGRVFEALLVGSDSLTDLAVLKINAT----NLPVIPINPKRVPHIG  151 (353)
T ss_pred             ccceeeEEEEeCCeEEEecccEeCCCCEEEEEeCCCCEEEEEEEEEcCCCCEEEEEEcCC----CCCeeeccCcCcCCCC
Confidence            34679999997  66766666                 7999999999999999999986    57889998764  578


Q ss_pred             CeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccCC----cc
Q 017471           70 DAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHED----VE  145 (371)
Q Consensus        70 d~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~----~~  145 (371)
                      |+|+++|||++... +++.|+||+.++..... .....+||+|+++++|||||||+|.+||||||+++.+...+    ..
T Consensus       152 ~~V~aiG~P~g~~~-~~t~Giis~~~r~~~~~-~~~~~~iqtda~i~~GnSGGPl~n~~G~vvGI~~~~~~~~~~~~~~~  229 (353)
T PRK10898        152 DVVLAIGNPYNLGQ-TITQGIISATGRIGLSP-TGRQNFLQTDASINHGNSGGALVNSLGELMGINTLSFDKSNDGETPE  229 (353)
T ss_pred             CEEEEEeCCCCcCC-CcceeEEEeccccccCC-ccccceEEeccccCCCCCcceEECCCCeEEEEEEEEecccCCCCccc
Confidence            99999999998765 89999999987643222 12246899999999999999999999999999998765332    25


Q ss_pred             ceeccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEE
Q 017471          146 NIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSF  224 (371)
Q Consensus       146 ~~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~v  224 (371)
                      +++||||++.+++++++|+++|++. ++|||+..+++ ++..+..++++ ...|++|.+|.++|||++ ||++||+|++|
T Consensus       230 g~~faIP~~~~~~~~~~l~~~G~~~-~~~lGi~~~~~-~~~~~~~~~~~-~~~Gv~V~~V~~~spA~~aGL~~GDvI~~I  306 (353)
T PRK10898        230 GIGFAIPTQLATKIMDKLIRDGRVI-RGYIGIGGREI-APLHAQGGGID-QLQGIVVNEVSPDGPAAKAGIQVNDLIISV  306 (353)
T ss_pred             ceEEEEchHHHHHHHHHHhhcCccc-ccccceEEEEC-CHHHHHhcCCC-CCCeEEEEEECCCChHHHcCCCCCCEEEEE
Confidence            7899999999999999999999998 89999999998 77777777876 357999999999999999 99999999999


Q ss_pred             CCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEecccc
Q 017471          225 DGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHR  279 (371)
Q Consensus       225 nG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~  279 (371)
                      ||++|.++.++.          +.+....+|++++++|.|+|+.+++++++.+.+
T Consensus       307 ng~~V~s~~~l~----------~~l~~~~~g~~v~l~v~R~g~~~~~~v~l~~~p  351 (353)
T PRK10898        307 NNKPAISALETM----------DQVAEIRPGSVIPVVVMRDDKQLTLQVTIQEYP  351 (353)
T ss_pred             CCEEcCCHHHHH----------HHHHhcCCCCEEEEEEEECCEEEEEEEEeccCC
Confidence            999999988864          566666789999999999999999999887554


No 6  
>COG0265 DegQ Trypsin-like serine proteases, typically periplasmic, contain C-terminal PDZ domain [Posttranslational modification, protein turnover, chaperones]
Probab=100.00  E-value=2.3e-34  Score=281.25  Aligned_cols=234  Identities=26%  Similarity=0.402  Sum_probs=205.2

Q ss_pred             cCCCeEeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCCC--CCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCC
Q 017471           25 LCSPSAPSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELPA--LQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHG  102 (371)
Q Consensus        25 ~~~~~~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~~--lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~  102 (371)
                      .+++.++|++++.|+..|+|++|++...   .++.+.++++..  +||+++++|+|++... +++.|+||..++......
T Consensus       103 ~dg~~~~a~~vg~d~~~dlavlki~~~~---~~~~~~~~~s~~l~vg~~v~aiGnp~g~~~-tvt~Givs~~~r~~v~~~  178 (347)
T COG0265         103 ADGREVPAKLVGKDPISDLAVLKIDGAG---GLPVIALGDSDKLRVGDVVVAIGNPFGLGQ-TVTSGIVSALGRTGVGSA  178 (347)
T ss_pred             CCCCEEEEEEEecCCccCEEEEEeccCC---CCceeeccCCCCcccCCEEEEecCCCCccc-ceeccEEeccccccccCc
Confidence            4666689999999999999999999874   378889998875  5699999999999665 999999999998622222


Q ss_pred             CeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccCC-ccceeccccCcchhHhHhhhhhCCeeecccccceeeeE
Q 017471          103 STELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHED-VENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQK  181 (371)
Q Consensus       103 ~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~-~~~~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~  181 (371)
                      .....+||+|+++|+||||||++|.+|++|||+++.+...+ ..+++|++|++.++.+++++...|++. ++++|+.+.+
T Consensus       179 ~~~~~~IqtdAain~gnsGgpl~n~~g~~iGint~~~~~~~~~~gigfaiP~~~~~~v~~~l~~~G~v~-~~~lgv~~~~  257 (347)
T COG0265         179 GGYVNFIQTDAAINPGNSGGPLVNIDGEVVGINTAIIAPSGGSSGIGFAIPVNLVAPVLDELISKGKVV-RGYLGVIGEP  257 (347)
T ss_pred             ccccchhhcccccCCCCCCCceEcCCCcEEEEEEEEecCCCCcceeEEEecHHHHHHHHHHHHHcCCcc-ccccceEEEE
Confidence            23567899999999999999999999999999998886553 466999999999999999999888888 9999999998


Q ss_pred             cCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEE
Q 017471          182 MENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAV  260 (371)
Q Consensus       182 ~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l  260 (371)
                      + +...+  +|++ ...|++|.+|.+++||++ |+++||+|+++||+++.+..++.          ..+....+|+++.+
T Consensus       258 ~-~~~~~--~g~~-~~~G~~V~~v~~~spa~~agi~~Gdii~~vng~~v~~~~~l~----------~~v~~~~~g~~v~~  323 (347)
T COG0265         258 L-TADIA--LGLP-VAAGAVVLGVLPGSPAAKAGIKAGDIITAVNGKPVASLSDLV----------AAVASNRPGDEVAL  323 (347)
T ss_pred             c-ccccc--cCCC-CCCceEEEecCCCChHHHcCCCCCCEEEEECCEEccCHHHHH----------HHHhccCCCCEEEE
Confidence            8 65555  7877 667999999999999999 99999999999999999998875          67777779999999


Q ss_pred             EEEECCEEEEEEEEecc
Q 017471          261 KVLRDSKILNFNITLAT  277 (371)
Q Consensus       261 ~v~R~g~~~~~~v~l~~  277 (371)
                      ++.|+|+.+++.+++..
T Consensus       324 ~~~r~g~~~~~~v~l~~  340 (347)
T COG0265         324 KLLRGGKERELAVTLGD  340 (347)
T ss_pred             EEEECCEEEEEEEEecC
Confidence            99999999999999976


No 7  
>KOG1320 consensus Serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=99.93  E-value=2.9e-26  Score=225.77  Aligned_cols=320  Identities=34%  Similarity=0.498  Sum_probs=276.7

Q ss_pred             eeeccCCeeeeeccEEEEE-EEecCCCe---------------------EeEEEEEecCCCCEEEEEEecCCCcCCccce
Q 017471            3 ILWSTTRRLNSRNEALILS-TWLLCSPS---------------------APSATLVTADICIYTMLTVEDDEFWEGVLPV   60 (371)
Q Consensus         3 ~~~~~~~~~~~~gsg~vi~-~~~~~~~~---------------------~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~   60 (371)
                      .+|++.++..+.++||.+. ..+++++|                     +.|++...-.++|+|++.++..+||+++.|+
T Consensus        77 ~pw~~~~q~~~~~s~f~i~~~~lltn~~~v~~~~~~~~v~v~~~gs~~k~~~~v~~~~~~cd~Avv~Ie~~~f~~~~~~~  156 (473)
T KOG1320|consen   77 LPWQRTRQFSSGGSGFAIYGKKLLTNAHVVAPNNDHKFVTVKKHGSPRKYKAFVAAVFEECDLAVVYIESEEFWKGMNPF  156 (473)
T ss_pred             CcceeeehhcccccchhhcccceeecCccccccccccccccccCCCchhhhhhHHHhhhcccceEEEEeeccccCCCccc
Confidence            5899999999999999998 44444444                     4677777778899999999999999999999


Q ss_pred             ecCCCCCCCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecc
Q 017471           61 EFGELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLK  140 (371)
Q Consensus        61 ~l~~s~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~  140 (371)
                      ++++.+.+.+-++++|    ++...+|.|.|++.+...|.++...+..+|+|+++++|+||+|.+...+++.|+++..+.
T Consensus       157 e~~~ip~l~~S~~Vv~----gd~i~VTnghV~~~~~~~y~~~~~~l~~vqi~aa~~~~~s~ep~i~g~d~~~gvA~l~ik  232 (473)
T KOG1320|consen  157 ELGDIPSLNGSGFVVG----GDGIIVTNGHVVRVEPRIYAHSSTVLLRVQIDAAIGPGNSGEPVIVGVDKVAGVAFLKIK  232 (473)
T ss_pred             ccCCCcccCccEEEEc----CCcEEEEeeEEEEEEeccccCCCcceeeEEEEEeecCCccCCCeEEccccccceEEEEEe
Confidence            9999999999999998    455699999999999888888888888999999999999999999988999999998874


Q ss_pred             cCCccceeccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCccccCCCCCCE
Q 017471          141 HEDVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESEVLKPSDI  220 (371)
Q Consensus       141 ~~~~~~~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~GL~~GDv  220 (371)
                      ..+  ++.+.+|...+.++.......+.+.+++.++...+.+++.+.++.+.|..+ +|+.+.++.+.++|.+-+++||.
T Consensus       233 ~~~--~i~~~i~~~~~~~~~~G~~~~a~~~~f~~~nt~t~g~vs~~~R~~~~lg~~-~g~~i~~~~qtd~ai~~~nsg~~  309 (473)
T KOG1320|consen  233 TPE--NILYVIPLGVSSHFRTGVEVSAIGNGFGLLNTLTQGMVSGQLRKSFKLGLE-TGVLISKINQTDAAINPGNSGGP  309 (473)
T ss_pred             cCC--cccceeecceeeeecccceeeccccCceeeeeeeecccccccccccccCcc-cceeeeeecccchhhhcccCCCc
Confidence            322  788999999999998888788888889999999999989999999999877 89999999999988888999999


Q ss_pred             EEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccccccCCCCCCCCCCCceeeccEE
Q 017471          221 ILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFV  300 (371)
Q Consensus       221 Il~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~~~~~~~~~~~~~~~Gl~  300 (371)
                      |+++||..|.    +.+++.+|+.|++.+..+.+++++...+.|.+   ++.++++......|.+.+.+.|.|++..|++
T Consensus       310 ll~~DG~~Ig----Vn~~~~~ri~~~~~iSf~~p~d~vl~~v~r~~---e~~~~lr~~~~~~p~~~~~g~~s~~i~~g~v  382 (473)
T KOG1320|consen  310 LLNLDGEVIG----VNTRKVTRIGFSHGISFKIPIDTVLVIVLRLG---EFQISLRPVKPLVPVHQYIGLPSYYIFAGLV  382 (473)
T ss_pred             EEEecCcEee----eeeeeeEEeeccccceeccCchHhhhhhhhhh---hhceeeccccCcccccccCCceeEEEecceE
Confidence            9999999997    55677889999999999999999999999998   5677777778888888999999999999999


Q ss_pred             Eech--HHHHHHcCcccccCceEEEEEecCChHhHHhhhhhh
Q 017471          301 FSRC--LYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSS  340 (371)
Q Consensus       301 ~~~l--~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~  340 (371)
                      |+++  ++...+    ....++++..+-++||+.--..+-||
T Consensus       383 f~~~~~~~~~~~----~~~q~v~is~Vlp~~~~~~~~~~~g~  420 (473)
T KOG1320|consen  383 FVPLTKSYIFPS----GVVQLVLVSQVLPGSINGGYGLKPGD  420 (473)
T ss_pred             EeecCCCccccc----cceeEEEEEEeccCCCcccccccCCC
Confidence            9876  333211    12268999999999998765555544


No 8  
>KOG1421 consensus Predicted signaling-associated protein (contains a PDZ domain) [General function prediction only]
Probab=99.91  E-value=4.7e-24  Score=212.40  Aligned_cols=303  Identities=16%  Similarity=0.241  Sum_probs=240.4

Q ss_pred             eeccCCeeeeeccEEEEE---EEecCCCeE------------------eEEEEEecCCCCEEEEEEecCCC-cCCcccee
Q 017471            4 LWSTTRRLNSRNEALILS---TWLLCSPSA------------------PSATLVTADICIYTMLTVEDDEF-WEGVLPVE   61 (371)
Q Consensus         4 ~~~~~~~~~~~gsg~vi~---~~~~~~~~~------------------~A~vv~~d~~~DlAlLkv~~~~~-~~~l~~~~   61 (371)
                      ++++.-...+.++||+++   +++++++|+                  +--.++.||.||+.+++++++.. ...+..+.
T Consensus        75 ~fdtesag~~~atgfvvd~~~gyiLtnrhvv~pgP~va~avf~n~ee~ei~pvyrDpVhdfGf~r~dps~ir~s~vt~i~  154 (955)
T KOG1421|consen   75 AFDTESAGESEATGFVVDKKLGYILTNRHVVAPGPFVASAVFDNHEEIEIYPVYRDPVHDFGFFRYDPSTIRFSIVTEIC  154 (955)
T ss_pred             ecccccccccceeEEEEecccceEEEeccccCCCCceeEEEecccccCCcccccCCchhhcceeecChhhcceeeeeccc
Confidence            356677788899999999   888999883                  22345778889999999998753 12344555


Q ss_pred             cCCC-CCCCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCC-----CeEEeEEEEEeeccCCCCCCceecCCCcEEEEE
Q 017471           62 FGEL-PALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHG-----STELLGLQIDAAINSGNSGGPAFNDKGKCVGIA  135 (371)
Q Consensus        62 l~~s-~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~-----~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~  135 (371)
                      +... .++|.++.++||..+. ..++-.|.+|++++.....+     .....++|.-+....|.||.|++|.+|..|.++
T Consensus       155 lap~~akvgseirvvgNDagE-klsIlagflSrldr~apdyg~~~yndfnTfy~QaasstsggssgspVv~i~gyAVAl~  233 (955)
T KOG1421|consen  155 LAPELAKVGSEIRVVGNDAGE-KLSILAGFLSRLDRNAPDYGEDTYNDFNTFYIQAASSTSGGSSGSPVVDIPGYAVALN  233 (955)
T ss_pred             cCccccccCCceEEecCCccc-eEEeehhhhhhccCCCccccccccccccceeeeehhcCCCCCCCCceecccceEEeee
Confidence            5543 3688999999998664 45889999999987533222     122346899999999999999999999999998


Q ss_pred             eeecccCCccceeccccCcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCC-----------CCCc-eEEE
Q 017471          136 FQSLKHEDVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKA-----------DQKG-VRIR  203 (371)
Q Consensus       136 ~~~~~~~~~~~~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~-----------~~~g-v~V~  203 (371)
                      ....   .....+|++|.+.+.+.|.-++++.-.. |+.|-+++..- .-+..+.+||+.           ...| ++|.
T Consensus       234 agg~---~ssas~ffLpLdrV~RaL~clq~n~PIt-RGtLqvefl~k-~~de~rrlGL~sE~eqv~r~k~P~~tgmLvV~  308 (955)
T KOG1421|consen  234 AGGS---ISSASDFFLPLDRVVRALRCLQNNTPIT-RGTLQVEFLHK-LFDECRRLGLSSEWEQVVRTKFPERTGMLVVE  308 (955)
T ss_pred             cCCc---ccccccceeeccchhhhhhhhhcCCCcc-cceEEEEEehh-hhHHHHhcCCcHHHHHHHHhcCcccceeEEEE
Confidence            7644   3456689999999999999997666666 89998888776 677888899864           2345 4577


Q ss_pred             EECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccccccCC
Q 017471          204 RVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIP  283 (371)
Q Consensus       204 ~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~  283 (371)
                      .|.++|||++-|++||++++||+.-+.++.++.           .+.....|+.++|+|+|+|++.+++++........|
T Consensus       309 ~vL~~gpa~k~Le~GDillavN~t~l~df~~l~-----------~iLDegvgk~l~LtI~Rggqelel~vtvqdlh~itp  377 (955)
T KOG1421|consen  309 TVLPEGPAEKKLEPGDILLAVNSTCLNDFEALE-----------QILDEGVGKNLELTIQRGGQELELTVTVQDLHGITP  377 (955)
T ss_pred             EeccCCchhhccCCCcEEEEEcceehHHHHHHH-----------HHHhhccCceEEEEEEeCCEEEEEEEEeccccCCCC
Confidence            899999999999999999999999998888764           444556899999999999999999999987766544


Q ss_pred             CCCCCCCCCceeeccEEEech-HHHHHHcCcccccCceEEEEEecCChHhH
Q 017471          284 SHNKGRPPSYYIIAGFVFSRC-LYLISVLSMERIMNMKLRSSFWTSSCIQC  333 (371)
Q Consensus       284 ~~~~~~~~~~~~~~Gl~~~~l-~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~  333 (371)
                      .       ++..|+|.+|+++ +.++..|.++..  |+|++++. +||+.-
T Consensus       378 ~-------R~levcGav~hdlsyq~ar~y~lP~~--GvyVa~~~-gsf~~~  418 (955)
T KOG1421|consen  378 D-------RFLEVCGAVFHDLSYQLARLYALPVE--GVYVASPG-GSFRHR  418 (955)
T ss_pred             c-------eEEEEcceEecCCCHHHHhhcccccC--cEEEccCC-CCcccc
Confidence            4       4666999999999 888888888744  99999998 888643


No 9  
>KOG1320 consensus Serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=99.88  E-value=1.3e-21  Score=193.07  Aligned_cols=236  Identities=21%  Similarity=0.223  Sum_probs=185.6

Q ss_pred             CeEeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCCC--CCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCC--
Q 017471           28 PSAPSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELPA--LQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGS--  103 (371)
Q Consensus        28 ~~~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~~--lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~--  103 (371)
                      ..+.+.+++.|+..|+|+++++.++  ...++++++.+..  .|+++.++|+|++..+ +.+.|+++...|..+..+.  
T Consensus       211 ~s~ep~i~g~d~~~gvA~l~ik~~~--~i~~~i~~~~~~~~~~G~~~~a~~~~f~~~n-t~t~g~vs~~~R~~~~lg~~~  287 (473)
T KOG1320|consen  211 NSGEPVIVGVDKVAGVAFLKIKTPE--NILYVIPLGVSSHFRTGVEVSAIGNGFGLLN-TLTQGMVSGQLRKSFKLGLET  287 (473)
T ss_pred             ccCCCeEEccccccceEEEEEecCC--cccceeecceeeeecccceeeccccCceeee-eeeecccccccccccccCccc
Confidence            6678999999999999999998664  2378888887665  5799999999999998 9999999999887554332  


Q ss_pred             --eEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccCC-ccceeccccCcchhHhHhhhhhCC---eee-----cc
Q 017471          104 --TELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHED-VENIGYVIPTPVIMHFIQDYEKNG---AYT-----GF  172 (371)
Q Consensus       104 --~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~-~~~~~~aiP~~~i~~~l~~L~~~g---~~~-----g~  172 (371)
                        .-.+++|+|+++++||||||++|.+|++||++++...+-+ ...++|++|.+.+..++.+..+..   +..     .+
T Consensus       288 g~~i~~~~qtd~ai~~~nsg~~ll~~DG~~IgVn~~~~~ri~~~~~iSf~~p~d~vl~~v~r~~e~~~~lr~~~~~~p~~  367 (473)
T KOG1320|consen  288 GVLISKINQTDAAINPGNSGGPLLNLDGEVIGVNTRKVTRIGFSHGISFKIPIDTVLVIVLRLGEFQISLRPVKPLVPVH  367 (473)
T ss_pred             ceeeeeecccchhhhcccCCCcEEEecCcEeeeeeeeeEEeeccccceeccCchHhhhhhhhhhhhceeeccccCccccc
Confidence              3356899999999999999999999999999988764322 357899999999988888763211   111     13


Q ss_pred             cccceeeeEcCCHHHH-----HhccCCC-CCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchh
Q 017471          173 PLLGVEWQKMENPDLR-----VAMSMKA-DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGF  245 (371)
Q Consensus       173 ~~lGi~~~~~~~~~~~-----~~~gl~~-~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~  245 (371)
                      .|+|.....+ ++.+.     +.+-.+. ..++++|.+|.|++++.. ++++||+|.+|||++|.+..++.         
T Consensus       368 ~~~g~~s~~i-~~g~vf~~~~~~~~~~~~~~q~v~is~Vlp~~~~~~~~~~~g~~V~~vng~~V~n~~~l~---------  437 (473)
T KOG1320|consen  368 QYIGLPSYYI-FAGLVFVPLTKSYIFPSGVVQLVLVSQVLPGSINGGYGLKPGDQVVKVNGKPVKNLKHLY---------  437 (473)
T ss_pred             ccCCceeEEE-ecceEEeecCCCccccccceeEEEEEEeccCCCcccccccCCCEEEEECCEEeechHHHH---------
Confidence            4666665555 22211     1121221 125899999999999999 99999999999999999999986         


Q ss_pred             hhhhhccCCCCEEEEEEEECCEEEEEEEEecc
Q 017471          246 SYLVSQKYTGDSAAVKVLRDSKILNFNITLAT  277 (371)
Q Consensus       246 ~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~  277 (371)
                       .++....+++++.+..+|+.|..++.+.++.
T Consensus       438 -~~i~~~~~~~~v~vl~~~~~e~~tl~Il~~~  468 (473)
T KOG1320|consen  438 -ELIEECSTEDKVAVLDRRSAEDATLEILPEH  468 (473)
T ss_pred             -HHHHhcCcCceEEEEEecCccceeEEecccc
Confidence             6788777889999999999999998887653


No 10 
>KOG1421 consensus Predicted signaling-associated protein (contains a PDZ domain) [General function prediction only]
Probab=99.80  E-value=3e-18  Score=171.28  Aligned_cols=271  Identities=15%  Similarity=0.138  Sum_probs=205.6

Q ss_pred             CCeEeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCC-CCCCeEEEEEeCCCCCCc--eeeeeEEeeeeeeeccCCC
Q 017471           27 SPSAPSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELP-ALQDAVTVVGYPIGGDTI--SVTSGVVSRIEILSYVHGS  103 (371)
Q Consensus        27 ~~~~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~-~lgd~V~~iG~p~g~~~~--s~t~G~Vs~~~~~~~~~~~  103 (371)
                      .-.++|++.+.|+.+++|.+|++++.    ...+.|.+.. ..||++...|+....+..  ..+...||.........++
T Consensus       585 S~~i~a~~~fL~~t~n~a~~kydp~~----~~~~kl~~~~v~~gD~~~f~g~~~~~r~ltaktsv~dvs~~~~ps~~~pr  660 (955)
T KOG1421|consen  585 SDGIPANVSFLHPTENVASFKYDPAL----EVQLKLTDTTVLRGDECTFEGFTEDLRALTAKTSVTDVSVVIIPSSVMPR  660 (955)
T ss_pred             cccccceeeEecCccceeEeccChhH----hhhhccceeeEecCCceeEecccccchhhcccceeeeeEEEEecCCCCcc
Confidence            33479999999999999999999973    3556666654 457999999998765531  2233334433332222222


Q ss_pred             ---eEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccC-C--ccceeccccCcchhHhHhhhhhCCeeecccccce
Q 017471          104 ---TELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-D--VENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGV  177 (371)
Q Consensus       104 ---~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~-~--~~~~~~aiP~~~i~~~l~~L~~~g~~~g~~~lGi  177 (371)
                         ..++.|.+++.+..++-.|-+.|.+|+|+|+|...+.+. +  ...+.|.+.+..+.+.++.|+.+++.. ...+|+
T Consensus       661 ~r~~n~e~Is~~~nlsT~c~sg~ltdddg~vvalwl~~~ge~~~~kd~~y~~gl~~~~~l~vl~rlk~g~~~r-p~i~~v  739 (955)
T KOG1421|consen  661 FRATNLEVISFMDNLSTSCLSGRLTDDDGEVVALWLSVVGEDVGGKDYTYKYGLSMSYILPVLERLKLGPSAR-PTIAGV  739 (955)
T ss_pred             eeecceEEEEEeccccccccceEEECCCCeEEEEEeeeeccccCCceeEEEeccchHHHHHHHHHHhcCCCCC-ceeecc
Confidence               346789999999999989999999999999998777543 2  234678899999999999998877776 677899


Q ss_pred             eeeEcCCHHHHHhccCCCC------------CCceEEEEECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchh
Q 017471          178 EWQKMENPDLRVAMSMKAD------------QKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGF  245 (371)
Q Consensus       178 ~~~~~~~~~~~~~~gl~~~------------~~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~  245 (371)
                      +|..+ +...++.+|++.+            .+-.+|++|.+.-+-  -|..||+|+++||+-|+...|+.         
T Consensus       740 ef~~i-~laqar~lglp~e~imk~e~es~~~~ql~~ishv~~~~~k--il~~gdiilsvngk~itr~~dl~---------  807 (955)
T KOG1421|consen  740 EFSHI-TLAQARTLGLPSEFIMKSEEESTIPRQLYVISHVRPLLHK--ILGVGDIILSVNGKMITRLSDLH---------  807 (955)
T ss_pred             ceeeE-EeehhhccCCCHHHHhhhhhcCCCcceEEEEEeeccCccc--ccccccEEEEecCeEEeeehhhh---------
Confidence            99999 8888889999852            124668888775543  59999999999999999998874         


Q ss_pred             hhhhhccCCCCEEEEEEEECCEEEEEEEEeccccccCCCCCCCCCCCceeeccEEEech-HHHHHHcCcccccCceEEEE
Q 017471          246 SYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRC-LYLISVLSMERIMNMKLRSS  324 (371)
Q Consensus       246 ~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~~~~~~~~~~~~~~~Gl~~~~l-~~~~~~~~~~~~~~~~~v~~  324 (371)
                       . +.      .+...|+|+|.+++++++.-+..+         ..+..+|+|..+|++ ..+.++.  .+-..|+|+.+
T Consensus       808 -d-~~------eid~~ilrdg~~~~ikipt~p~~e---------t~r~vi~~gailq~ph~av~~q~--edlp~gvyvt~  868 (955)
T KOG1421|consen  808 -D-FE------EIDAVILRDGIEMEIKIPTYPEYE---------TSRAVIWMGAILQPPHSAVFEQV--EDLPEGVYVTS  868 (955)
T ss_pred             -h-hh------hhheeeeecCcEEEEEeccccccc---------cceEEEEEeccccCchHHHHHHH--hccCCceEEee
Confidence             1 21      578899999999999887765432         224567999999998 6666553  34559999999


Q ss_pred             EecCChHhH
Q 017471          325 FWTSSCIQC  333 (371)
Q Consensus       325 ~~~~Sp~~~  333 (371)
                      ..++|||.-
T Consensus       869 rg~gspalq  877 (955)
T KOG1421|consen  869 RGYGSPALQ  877 (955)
T ss_pred             cccCChhHh
Confidence            999999854


No 11 
>PF13180 PDZ_2:  PDZ domain; PDB: 2L97_A 1Y8T_A 2Z9I_A 1LCY_A 2PZD_B 2P3W_A 1VCW_C 1TE0_B 1SOZ_C 1SOT_C ....
Probab=99.60  E-value=3e-15  Score=115.81  Aligned_cols=81  Identities=32%  Similarity=0.525  Sum_probs=70.0

Q ss_pred             cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc
Q 017471          173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ  251 (371)
Q Consensus       173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~  251 (371)
                      ||||+.+....+            ..|++|.+|.++|||++ ||++||+|++|||++|++..++.          +.+..
T Consensus         1 ~~lGv~~~~~~~------------~~g~~V~~V~~~spA~~aGl~~GD~I~~ing~~v~~~~~~~----------~~l~~   58 (82)
T PF13180_consen    1 GGLGVTVQNLSD------------TGGVVVVSVIPGSPAAKAGLQPGDIILAINGKPVNSSEDLV----------NILSK   58 (82)
T ss_dssp             -E-SEEEEECSC------------SSSEEEEEESTTSHHHHTTS-TTEEEEEETTEESSSHHHHH----------HHHHC
T ss_pred             CEECeEEEEccC------------CCeEEEEEeCCCCcHHHCCCCCCcEEEEECCEEcCCHHHHH----------HHHHh
Confidence            589999998821            46999999999999999 99999999999999999988865          67778


Q ss_pred             cCCCCEEEEEEEECCEEEEEEEEe
Q 017471          252 KYTGDSAAVKVLRDSKILNFNITL  275 (371)
Q Consensus       252 ~~~g~~v~l~v~R~g~~~~~~v~l  275 (371)
                      ..+|+++++++.|+|+.+++++++
T Consensus        59 ~~~g~~v~l~v~R~g~~~~~~v~l   82 (82)
T PF13180_consen   59 GKPGDTVTLTVLRDGEELTVEVTL   82 (82)
T ss_dssp             SSTTSEEEEEEEETTEEEEEEEE-
T ss_pred             CCCCCEEEEEEEECCEEEEEEEEC
Confidence            899999999999999999999875


No 12 
>TIGR03279 cyano_FeS_chp putative FeS-containing Cyanobacterial-specific oxidoreductase. Members of this protein family are predicted FeS-containing oxidoreductases of unknown function, apparently restricted to and universal across the Cyanobacteria. The high trusted cutoff score for this model, 700 bits, excludes homologs from other lineages. This exclusion seems justified because a significant number of sequence positions are simultaneously unique to and invariant across the Cyanobacteria, suggesting a specialized, conserved function, perhaps related to photosynthesis. A distantly related protein family, TIGR03278, in universal in and restricted to archaeal methanogens, and may be linked to methanogenesis.
Probab=99.52  E-value=2.4e-15  Score=147.75  Aligned_cols=102  Identities=21%  Similarity=0.272  Sum_probs=80.7

Q ss_pred             EEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEE-ECCEEEEEEEEecccc
Q 017471          202 IRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVL-RDSKILNFNITLATHR  279 (371)
Q Consensus       202 V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~-R~g~~~~~~v~l~~~~  279 (371)
                      |.+|.|+|||++ ||++||+|++|||++|.++.|+.          ..+    .++.++++|. |+|+..+++++....+
T Consensus         2 I~~V~pgSpAe~AGLe~GD~IlsING~~V~Dw~D~~----------~~l----~~e~l~L~V~~rdGe~~~l~Ie~~~de   67 (433)
T TIGR03279         2 ISAVLPGSIAEELGFEPGDALVSINGVAPRDLIDYQ----------FLC----ADEELELEVLDANGESHQIEIEKDLDE   67 (433)
T ss_pred             cCCcCCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHh----cCCcEEEEEEcCCCeEEEEEEecCCCC
Confidence            668999999999 99999999999999999998864          233    2467899997 8998877776653222


Q ss_pred             ccCCCCCCCCCCCceeeccEEEech--HHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhHHHHHHHHHHHhhhhh
Q 017471          280 RLIPSHNKGRPPSYYIIAGFVFSRC--LYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLLWCLRCLWLILILDMR  357 (371)
Q Consensus       280 ~~~~~~~~~~~~~~~~~~Gl~~~~l--~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~~~~~~~~~~~~~~~~  357 (371)
                      +                +|+.|.+.  ...+ +..-.                             |.|||.-|||.+||
T Consensus        68 d----------------lG~~f~~~~~d~~~-~C~N~-----------------------------C~FCFidQlP~gmR  101 (433)
T TIGR03279        68 D----------------LGLEFTTALFDGLI-QCNNR-----------------------------CPFCFIDQQPPGKR  101 (433)
T ss_pred             C----------------CcEEeccccCCccc-ccCCc-----------------------------CceEeccCCCCCCc
Confidence            2                79998654  2232 55555                             99999999999999


Q ss_pred             hHHHHH
Q 017471          358 RLLTLR  363 (371)
Q Consensus       358 ~~~~~~  363 (371)
                      ++||+|
T Consensus       102 ~sLY~K  107 (433)
T TIGR03279       102 ESLYLK  107 (433)
T ss_pred             Ccceec
Confidence            999976


No 13 
>cd00987 PDZ_serine_protease PDZ domain of tryspin-like serine proteases, such as DegP/HtrA, which are oligomeric proteins involved in heat-shock response, chaperone function, and apoptosis. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.43  E-value=6.8e-13  Score=103.81  Aligned_cols=88  Identities=35%  Similarity=0.599  Sum_probs=74.9

Q ss_pred             cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc
Q 017471          173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ  251 (371)
Q Consensus       173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~  251 (371)
                      +|+|+.++++ +++.+..++++ ...|++|.+|.++|||++ ||++||+|++|||+++.++.++.          ..+..
T Consensus         1 ~~~G~~~~~~-~~~~~~~~~~~-~~~g~~V~~v~~~s~a~~~gl~~GD~I~~Ing~~i~~~~~~~----------~~l~~   68 (90)
T cd00987           1 PWLGVTVQDL-TPDLAEELGLK-DTKGVLVASVDPGSPAAKAGLKPGDVILAVNGKPVKSVADLR----------RALAE   68 (90)
T ss_pred             CccceEEeEC-CHHHHHHcCCC-CCCEEEEEEECCCCHHHHcCCCcCCEEEEECCEECCCHHHHH----------HHHHh
Confidence            5899999999 77777767765 457999999999999998 99999999999999999988764          56665


Q ss_pred             cCCCCEEEEEEEECCEEEEEE
Q 017471          252 KYTGDSAAVKVLRDSKILNFN  272 (371)
Q Consensus       252 ~~~g~~v~l~v~R~g~~~~~~  272 (371)
                      ...++.+.+++.|+|+..+++
T Consensus        69 ~~~~~~i~l~v~r~g~~~~~~   89 (90)
T cd00987          69 LKPGDKVTLTVLRGGKELTVT   89 (90)
T ss_pred             cCCCCEEEEEEEECCEEEEee
Confidence            556899999999999876654


No 14 
>cd00986 PDZ_LON_protease PDZ domain of ATP-dependent LON serine proteases. Most PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this bacterial subfamily of protease-associated PDZ domains a C-terminal beta-strand  is thought to form the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.30  E-value=9.9e-12  Score=95.22  Aligned_cols=72  Identities=28%  Similarity=0.366  Sum_probs=63.5

Q ss_pred             CCceEEEEECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEec
Q 017471          197 QKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLA  276 (371)
Q Consensus       197 ~~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~  276 (371)
                      ..|++|.+|.++|||++||++||+|++|||+++.++.++.          ..+....+|+.+.+++.|+|+.+++++++.
T Consensus         7 ~~Gv~V~~V~~~s~A~~gL~~GD~I~~Ing~~v~~~~~~~----------~~l~~~~~~~~v~l~v~r~g~~~~~~v~l~   76 (79)
T cd00986           7 YHGVYVTSVVEGMPAAGKLKAGDHIIAVDGKPFKEAEELI----------DYIQSKKEGDTVKLKVKREEKELPEDLILK   76 (79)
T ss_pred             ecCEEEEEECCCCchhhCCCCCCEEEEECCEECCCHHHHH----------HHHHhCCCCCEEEEEEEECCEEEEEEEEEe
Confidence            4689999999999998799999999999999999988864          566655688999999999999999999987


Q ss_pred             cc
Q 017471          277 TH  278 (371)
Q Consensus       277 ~~  278 (371)
                      ..
T Consensus        77 ~~   78 (79)
T cd00986          77 TF   78 (79)
T ss_pred             cc
Confidence            53


No 15 
>cd00991 PDZ_archaeal_metalloprotease PDZ domain of archaeal zinc metalloprotases, presumably membrane-associated or integral membrane proteases, which may be involved in signalling and regulatory mechanisms. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.29  E-value=1.1e-11  Score=95.11  Aligned_cols=69  Identities=26%  Similarity=0.273  Sum_probs=60.7

Q ss_pred             CCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEE
Q 017471          196 DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT  274 (371)
Q Consensus       196 ~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~  274 (371)
                      ...|++|.+|.++|||++ ||++||+|++|||+++.++.++.          ..+....+|+++.+++.|+|+..+++++
T Consensus         8 ~~~Gv~V~~V~~~spa~~aGL~~GDiI~~Ing~~v~~~~d~~----------~~l~~~~~g~~v~l~v~r~g~~~~~~~~   77 (79)
T cd00991           8 AVAGVVIVGVIVGSPAENAVLHTGDVIYSINGTPITTLEDFM----------EALKPTKPGEVITVTVLPSTTKLTNVST   77 (79)
T ss_pred             cCCcEEEEEECCCChHHhcCCCCCCEEEEECCEEcCCHHHHH----------HHHhcCCCCCEEEEEEEECCEEEEEEEE
Confidence            357999999999999998 99999999999999999998864          5666656789999999999999887765


No 16 
>cd00990 PDZ_glycyl_aminopeptidase PDZ domain associated with archaeal and bacterial M61 glycyl-aminopeptidases. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand is presumed to form the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.26  E-value=2.3e-11  Score=93.11  Aligned_cols=77  Identities=22%  Similarity=0.384  Sum_probs=63.8

Q ss_pred             cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc
Q 017471          173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ  251 (371)
Q Consensus       173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~  251 (371)
                      +|+|+.+..-              ..|++|.+|.++|||++ ||++||+|++|||+++.++.+             .+..
T Consensus         1 ~~~G~~~~~~--------------~~~~~V~~V~~~s~a~~aGl~~GD~I~~Ing~~v~~~~~-------------~l~~   53 (80)
T cd00990           1 PYLGLTLDKE--------------EGLGKVTFVRDDSPADKAGLVAGDELVAVNGWRVDALQD-------------RLKE   53 (80)
T ss_pred             CcccEEEEcc--------------CCcEEEEEECCCChHHHhCCCCCCEEEEECCEEhHHHHH-------------HHHh
Confidence            5788877532              35799999999999999 999999999999999987443             3444


Q ss_pred             cCCCCEEEEEEEECCEEEEEEEEec
Q 017471          252 KYTGDSAAVKVLRDSKILNFNITLA  276 (371)
Q Consensus       252 ~~~g~~v~l~v~R~g~~~~~~v~l~  276 (371)
                      ..+++.+.+++.|+|+..++++++.
T Consensus        54 ~~~~~~v~l~v~r~g~~~~~~v~~~   78 (80)
T cd00990          54 YQAGDPVELTVFRDDRLIEVPLTLA   78 (80)
T ss_pred             cCCCCEEEEEEEECCEEEEEEEEec
Confidence            4578899999999999988888764


No 17 
>TIGR01713 typeII_sec_gspC general secretion pathway protein C. This model represents GspC, protein C of the main terminal branch of the general secretion pathway, also called type II secretion. This system transports folded proteins across the bacterial outer membrane and is widely distributed in Gram-negative pathogens.
Probab=99.22  E-value=5.4e-11  Score=111.38  Aligned_cols=101  Identities=15%  Similarity=0.183  Sum_probs=86.9

Q ss_pred             CcchhHhHhhhhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecC
Q 017471          153 TPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAN  231 (371)
Q Consensus       153 ~~~i~~~l~~L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~  231 (371)
                      ...+.++++++.++++.. ++|+|+..... +          ....|+.|..+.+++||++ |||+||+|++|||+++++
T Consensus       158 ~~~~~~v~~~l~~~g~~~-~~~lgi~p~~~-~----------g~~~G~~v~~v~~~s~a~~aGLr~GDvIv~ING~~i~~  225 (259)
T TIGR01713       158 IVVSRRIIEELTKDPQKM-FDYIRLSPVMK-N----------DKLEGYRLNPGKDPSLFYKSGLQDGDIAVALNGLDLRD  225 (259)
T ss_pred             hhhHHHHHHHHHHCHHhh-hheEeEEEEEe-C----------CceeEEEEEecCCCCHHHHcCCCCCCEEEEECCEEcCC
Confidence            346788899999989888 89999998655 2          1246999999999999999 999999999999999999


Q ss_pred             CCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEe
Q 017471          232 DGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITL  275 (371)
Q Consensus       232 ~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l  275 (371)
                      +.++.          ..+....++++++++|.|+|+.+++.+.+
T Consensus       226 ~~~~~----------~~l~~~~~~~~v~l~V~R~G~~~~i~v~~  259 (259)
T TIGR01713       226 PEQAF----------QALQMLREETNLTLTVERDGQREDIYVRF  259 (259)
T ss_pred             HHHHH----------HHHHhcCCCCeEEEEEEECCEEEEEEEEC
Confidence            98865          67777788899999999999998888764


No 18 
>TIGR02037 degP_htrA_DO periplasmic serine protease, Do/DeqQ family. This family consists of a set proteins various designated DegP, heat shock protein HtrA, and protease DO. The ortholog in Pseudomonas aeruginosa is designated MucD and is found in an operon that controls mucoid phenotype. This family also includes the DegQ (HhoA) paralog in E. coli which can rescue a DegP mutant, but not the smaller DegS paralog, which cannot. Members of this family are located in the periplasm and have separable functions as both protease and chaperone. Members have a trypsin domain and two copies of a PDZ domain. This protein protects bacteria from thermal and other stresses and may be important for the survival of bacterial pathogens.// The chaperone function is dominant at low temperatures, whereas the proteolytic activity is turned on at elevated temperatures.
Probab=99.11  E-value=1.8e-10  Score=115.84  Aligned_cols=90  Identities=24%  Similarity=0.473  Sum_probs=79.5

Q ss_pred             ccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhh
Q 017471          172 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVS  250 (371)
Q Consensus       172 ~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~  250 (371)
                      ..++|+.+.++ +++.++.++++....|++|.+|.++|||++ ||++||+|++|||++|.++.++.          +.+.
T Consensus       337 ~~~lGi~~~~l-~~~~~~~~~l~~~~~Gv~V~~V~~~SpA~~aGL~~GDvI~~Ing~~V~s~~d~~----------~~l~  405 (428)
T TIGR02037       337 NPFLGLTVANL-SPEIRKELRLKGDVKGVVVTKVVSGSPAARAGLQPGDVILSVNQQPVSSVAELR----------KVLD  405 (428)
T ss_pred             ccccceEEecC-CHHHHHHcCCCcCcCceEEEEeCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHH
Confidence            35799999999 888889999986567999999999999999 99999999999999999998865          6776


Q ss_pred             ccCCCCEEEEEEEECCEEEEEE
Q 017471          251 QKYTGDSAAVKVLRDSKILNFN  272 (371)
Q Consensus       251 ~~~~g~~v~l~v~R~g~~~~~~  272 (371)
                      ..++|++++++|.|+|+...+.
T Consensus       406 ~~~~g~~v~l~v~R~g~~~~~~  427 (428)
T TIGR02037       406 RAKKGGRVALLILRGGATIFVT  427 (428)
T ss_pred             hcCCCCEEEEEEEECCEEEEEE
Confidence            6667999999999999987654


No 19 
>PRK10779 zinc metallopeptidase RseP; Provisional
Probab=99.11  E-value=9.9e-11  Score=118.38  Aligned_cols=117  Identities=12%  Similarity=0.072  Sum_probs=86.6

Q ss_pred             eEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEeccc
Q 017471          200 VRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATH  278 (371)
Q Consensus       200 v~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~  278 (371)
                      .+|++|.++|||++ |||+||+|+++||++|++++++.          ..+....+|+++++++.|+|+.+++++++...
T Consensus       128 ~lV~~V~~~SpA~kAGLk~GDvI~~vnG~~V~~~~~l~----------~~v~~~~~g~~v~v~v~R~gk~~~~~v~l~~~  197 (449)
T PRK10779        128 PVVGEIAPNSIAAQAQIAPGTELKAVDGIETPDWDAVR----------LALVSKIGDESTTITVAPFGSDQRRDKTLDLR  197 (449)
T ss_pred             ccccccCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHhhccCCceEEEEEeCCccceEEEEeccc
Confidence            46899999999999 99999999999999999999976          56777778899999999999999999988654


Q ss_pred             cccCCCCCCCCCCCceeeccEEEechHHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhH
Q 017471          279 RRLIPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLL  342 (371)
Q Consensus       279 ~~~~~~~~~~~~~~~~~~~Gl~~~~l~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~  342 (371)
                      +......  ...  .....|+.  +.          .....+.+..+.++|||+-.-.+-+|.|
T Consensus       198 ~~~~~~~--~~~--~~~~lGl~--~~----------~~~~~~vV~~V~~~SpA~~AGL~~GDvI  245 (449)
T PRK10779        198 HWAFEPD--KQD--PVSSLGIR--PR----------GPQIEPVLAEVQPNSAASKAGLQAGDRI  245 (449)
T ss_pred             ccccCcc--ccc--hhhccccc--cc----------CCCcCcEEEeeCCCCHHHHcCCCCCCEE
Confidence            3221100  000  01123432  21          0112468899999999998777777765


No 20 
>cd00989 PDZ_metalloprotease PDZ domain of bacterial and plant zinc metalloprotases, presumably membrane-associated or integral membrane proteases, which may be involved in signalling and regulatory mechanisms. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.08  E-value=2.8e-10  Score=86.72  Aligned_cols=66  Identities=24%  Similarity=0.335  Sum_probs=55.9

Q ss_pred             CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEE
Q 017471          198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT  274 (371)
Q Consensus       198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~  274 (371)
                      ..++|.+|.++|||++ ||++||+|++|||+++.++.++.          ..+.. ..++.+.+++.|+|+..++.++
T Consensus        12 ~~~~V~~v~~~s~a~~~gl~~GD~I~~ing~~i~~~~~~~----------~~l~~-~~~~~~~l~v~r~~~~~~~~l~   78 (79)
T cd00989          12 IEPVIGEVVPGSPAAKAGLKAGDRILAINGQKIKSWEDLV----------DAVQE-NPGKPLTLTVERNGETITLTLT   78 (79)
T ss_pred             cCcEEEeECCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHHHH-CCCceEEEEEEECCEEEEEEec
Confidence            3588999999999998 99999999999999999988864          45544 3478899999999988777664


No 21 
>cd00988 PDZ_CTP_protease PDZ domain of C-terminal processing-, tail-specific-, and tricorn proteases, which function in posttranslational protein processing, maturation, and disassembly or degradation, in Bacteria, Archaea, and plant chloroplasts. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.05  E-value=9.3e-10  Score=85.13  Aligned_cols=68  Identities=24%  Similarity=0.342  Sum_probs=57.1

Q ss_pred             CCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCC--CCccccccccchhhhhhhccCCCCEEEEEEEEC-CEEEEEE
Q 017471          197 QKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAND--GTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD-SKILNFN  272 (371)
Q Consensus       197 ~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~--~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~-g~~~~~~  272 (371)
                      ..+++|..|.++|||++ ||++||+|++|||+++.++  .++.          ..+. ..+|+++.+++.|+ |+..+++
T Consensus        12 ~~~~~V~~v~~~s~a~~~gl~~GD~I~~vng~~i~~~~~~~~~----------~~l~-~~~~~~i~l~v~r~~~~~~~~~   80 (85)
T cd00988          12 DGGLVITSVLPGSPAAKAGIKAGDIIVAIDGEPVDGLSLEDVV----------KLLR-GKAGTKVRLTLKRGDGEPREVT   80 (85)
T ss_pred             CCeEEEEEecCCCCHHHcCCCCCCEEEEECCEEcCCCCHHHHH----------HHhc-CCCCCEEEEEEEcCCCCEEEEE
Confidence            36899999999999999 9999999999999999998  6643          3343 35688999999999 8888877


Q ss_pred             EEe
Q 017471          273 ITL  275 (371)
Q Consensus       273 v~l  275 (371)
                      +++
T Consensus        81 ~~~   83 (85)
T cd00988          81 LTR   83 (85)
T ss_pred             EEE
Confidence            764


No 22 
>PF13365 Trypsin_2:  Trypsin-like peptidase domain; PDB: 1Y8T_A 2Z9I_A 3QO6_A 1L1J_A 1QY6_A 2O8L_A 3OTP_E 2ZLE_I 1KY9_A 3CS0_A ....
Probab=98.83  E-value=1.4e-08  Score=82.98  Aligned_cols=87  Identities=24%  Similarity=0.325  Sum_probs=52.8

Q ss_pred             EEEEEEecCCCeEe--EEEEEecCC-CCEEEEEEecCCCcCCccceecCCCCCCCCeEEEEEeCCCCCCceeeeeEEeee
Q 017471           18 LILSTWLLCSPSAP--SATLVTADI-CIYTMLTVEDDEFWEGVLPVEFGELPALQDAVTVVGYPIGGDTISVTSGVVSRI   94 (371)
Q Consensus        18 ~vi~~~~~~~~~~~--A~vv~~d~~-~DlAlLkv~~~~~~~~l~~~~l~~s~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~   94 (371)
                      ..+.+...++...+  |++++.|+. +|+|||+++                     .....+..      ....+.....
T Consensus        31 ~~~~~~~~~~~~~~~~~~~~~~~~~~~D~All~v~---------------------~~~~~~~~------~~~~~~~~~~   83 (120)
T PF13365_consen   31 SSVEVVFPDGRRVPPVAEVVYFDPDDYDLALLKVD---------------------PWTGVGGG------VRVPGSTSGV   83 (120)
T ss_dssp             SEEEEEETTSCEEETEEEEEEEETT-TTEEEEEES---------------------CEEEEEEE------EEEEEEEEEE
T ss_pred             CEEEEEecCCCEEeeeEEEEEECCccccEEEEEEe---------------------cccceeee------eEeeeecccc
Confidence            34446666777777  999999999 999999999                     00000000      0111111111


Q ss_pred             eeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEE
Q 017471           95 EILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGI  134 (371)
Q Consensus        95 ~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI  134 (371)
                      ...  ........++ +|+++.+|+||||++|.+|+||||
T Consensus        84 ~~~--~~~~~~~~~~-~~~~~~~G~SGgpv~~~~G~vvGi  120 (120)
T PF13365_consen   84 SPT--STNDNRMLYI-TDADTRPGSSGGPVFDSDGRVVGI  120 (120)
T ss_dssp             EEE--EEEETEEEEE-ESSS-STTTTTSEEEETTSEEEEE
T ss_pred             ccc--cCcccceeEe-eecccCCCcEeHhEECCCCEEEeC
Confidence            100  0001111125 899999999999999999999997


No 23 
>cd00136 PDZ PDZ domain, also called DHR (Dlg homologous region) or GLGF (after a conserved sequence motif). Many PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. Heterodimerization through PDZ-PDZ domain interactions adds to the domain's versatility, and PDZ domain-mediated interactions may be modulated dynamically through target phosphorylation. Some PDZ domains play a role in scaffolding supramolecular complexes. PDZ domains are found in diverse signaling proteins in bacteria, archebacteria, and eurkayotes. This CD contains two distinct structural subgroups with either a N- or C-terminal beta-strand forming the peptide-binding groove base. The circular permutation placing the strand on the N-terminus appears to be found in Eumetazoa only, while the C-terminal variant is found in all three kingdoms of life, and seems to co-occur with protease domains. PDZ domains have been named after PSD95(pos
Probab=98.82  E-value=8e-09  Score=76.78  Aligned_cols=65  Identities=28%  Similarity=0.454  Sum_probs=51.8

Q ss_pred             ccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCC--CCccccccccchhhhhhh
Q 017471          174 LLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAND--GTVPFRHGERIGFSYLVS  250 (371)
Q Consensus       174 ~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~--~~l~~~~~~~~~~~~~~~  250 (371)
                      ++|+.+...+             ..+++|.+|.++|||++ ||++||+|++|||+++.++  .++.          +.+.
T Consensus         2 ~~G~~~~~~~-------------~~~~~V~~v~~~s~a~~~gl~~GD~I~~Ing~~v~~~~~~~~~----------~~l~   58 (70)
T cd00136           2 GLGFSIRGGT-------------EGGVVVLSVEPGSPAERAGLQAGDVILAVNGTDVKNLTLEDVA----------ELLK   58 (70)
T ss_pred             CccEEEecCC-------------CCCEEEEEeCCCCHHHHcCCCCCCEEEEECCEECCCCCHHHHH----------HHHh
Confidence            5777776541             14899999999999999 9999999999999999998  5543          4444


Q ss_pred             ccCCCCEEEEEE
Q 017471          251 QKYTGDSAAVKV  262 (371)
Q Consensus       251 ~~~~g~~v~l~v  262 (371)
                      . .+|+++++++
T Consensus        59 ~-~~g~~v~l~v   69 (70)
T cd00136          59 K-EVGEKVTLTV   69 (70)
T ss_pred             h-CCCCeEEEEE
Confidence            4 3488888876


No 24 
>TIGR00054 RIP metalloprotease RseP. A model that detects fragments as well matches a number of members of the PEPTIDASE FAMILY S2C. The region of match appears not to overlap the active site domain.
Probab=98.71  E-value=2.3e-08  Score=100.36  Aligned_cols=69  Identities=25%  Similarity=0.323  Sum_probs=61.0

Q ss_pred             CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEec
Q 017471          198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLA  276 (371)
Q Consensus       198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~  276 (371)
                      .+++|.+|.++|||++ |||+||+|++|||++|++++|+.          +.+.. .+++++++++.|+|+..++++++.
T Consensus       203 ~g~vV~~V~~~SpA~~aGL~~GD~Iv~Vng~~V~s~~dl~----------~~l~~-~~~~~v~l~v~R~g~~~~~~v~~~  271 (420)
T TIGR00054       203 IEPVLSDVTPNSPAEKAGLKEGDYIQSINGEKLRSWTDFV----------SAVKE-NPGKSMDIKVERNGETLSISLTPE  271 (420)
T ss_pred             cCcEEEEECCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHHHh-CCCCceEEEEEECCEEEEEEEEEc
Confidence            4799999999999999 99999999999999999999875          45544 578889999999999999998885


Q ss_pred             c
Q 017471          277 T  277 (371)
Q Consensus       277 ~  277 (371)
                      .
T Consensus       272 ~  272 (420)
T TIGR00054       272 A  272 (420)
T ss_pred             C
Confidence            3


No 25 
>smart00228 PDZ Domain present in PSD-95, Dlg, and ZO-1/2. Also called DHR (Dlg homologous region) or GLGF (relatively well conserved tetrapeptide in these domains). Some PDZs have been shown to bind C-terminal polypeptides; others appear to bind internal (non-C-terminal) polypeptides. Different PDZs possess different binding specificities.
Probab=98.68  E-value=6.4e-08  Score=74.25  Aligned_cols=73  Identities=25%  Similarity=0.340  Sum_probs=55.2

Q ss_pred             cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc
Q 017471          173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ  251 (371)
Q Consensus       173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~  251 (371)
                      ..+|+.+....+           ...|++|..|.++|||++ ||++||+|++|||+++.+..+..          .....
T Consensus        12 ~~~G~~~~~~~~-----------~~~~~~i~~v~~~s~a~~~gl~~GD~I~~In~~~v~~~~~~~----------~~~~~   70 (85)
T smart00228       12 GGLGFSLVGGKD-----------EGGGVVVSSVVPGSPAAKAGLKVGDVILEVNGTSVEGLTHLE----------AVDLL   70 (85)
T ss_pred             CcccEEEECCCC-----------CCCCEEEEEECCCCHHHHcCCCCCCEEEEECCEECCCCCHHH----------HHHHH
Confidence            467888765411           116899999999999999 99999999999999999876643          22222


Q ss_pred             cCCCCEEEEEEEECC
Q 017471          252 KYTGDSAAVKVLRDS  266 (371)
Q Consensus       252 ~~~g~~v~l~v~R~g  266 (371)
                      ...++.+.+++.|++
T Consensus        71 ~~~~~~~~l~i~r~~   85 (85)
T smart00228       71 KKAGGKVTLTVLRGG   85 (85)
T ss_pred             HhCCCeEEEEEEeCC
Confidence            334668999999875


No 26 
>PRK10779 zinc metallopeptidase RseP; Provisional
Probab=98.63  E-value=5.4e-08  Score=98.55  Aligned_cols=68  Identities=24%  Similarity=0.324  Sum_probs=60.3

Q ss_pred             ceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEecc
Q 017471          199 GVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLAT  277 (371)
Q Consensus       199 gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~  277 (371)
                      +++|.+|.++|||++ ||++||+|++|||++|++++|+.          +.+.. .+|+++.+++.|+|+..++++++..
T Consensus       222 ~~vV~~V~~~SpA~~AGL~~GDvIl~Ing~~V~s~~dl~----------~~l~~-~~~~~v~l~v~R~g~~~~~~v~~~~  290 (449)
T PRK10779        222 EPVLAEVQPNSAASKAGLQAGDRIVKVDGQPLTQWQTFV----------TLVRD-NPGKPLALEIERQGSPLSLTLTPDS  290 (449)
T ss_pred             CcEEEeeCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHh-CCCCEEEEEEEECCEEEEEEEEeee
Confidence            588999999999999 99999999999999999999875          45544 5788999999999999999988863


No 27 
>TIGR00225 prc C-terminal peptidase (prc). A C-terminal peptidase with different substrates in different species including processing of D1 protein of the photosystem II reaction center in higher plants and cleavage of a peptide of 11 residues from the precursor form of penicillin-binding protein in E.coli E.coli and H influenza have the most distal branch of the tree and their proteins have an N-terminal 200 amino acids that show no homology to other proteins in the database.
Probab=98.59  E-value=1.2e-07  Score=92.52  Aligned_cols=71  Identities=21%  Similarity=0.307  Sum_probs=58.0

Q ss_pred             CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCC--CccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEE
Q 017471          198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDG--TVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT  274 (371)
Q Consensus       198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~--~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~  274 (371)
                      .+++|.+|.++|||++ ||++||+|++|||++|.++.  ++.          . .....+|+++.+++.|+|+..+++++
T Consensus        62 ~~~~V~~V~~~spA~~aGL~~GD~I~~Ing~~v~~~~~~~~~----------~-~l~~~~g~~v~l~v~R~g~~~~~~v~  130 (334)
T TIGR00225        62 GEIVIVSPFEGSPAEKAGIKPGDKIIKINGKSVAGMSLDDAV----------A-LIRGKKGTKVSLEILRAGKSKPLTFT  130 (334)
T ss_pred             CEEEEEEeCCCChHHHcCCCCCCEEEEECCEECCCCCHHHHH----------H-hccCCCCCEEEEEEEeCCCCceEEEE
Confidence            4799999999999999 99999999999999999873  221          2 22335789999999999988888877


Q ss_pred             ecccc
Q 017471          275 LATHR  279 (371)
Q Consensus       275 l~~~~  279 (371)
                      +....
T Consensus       131 l~~~~  135 (334)
T TIGR00225       131 LKRDR  135 (334)
T ss_pred             EEEEE
Confidence            76543


No 28 
>TIGR00054 RIP metalloprotease RseP. A model that detects fragments as well matches a number of members of the PEPTIDASE FAMILY S2C. The region of match appears not to overlap the active site domain.
Probab=98.57  E-value=6.1e-08  Score=97.30  Aligned_cols=66  Identities=23%  Similarity=0.305  Sum_probs=55.5

Q ss_pred             CCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEE
Q 017471          197 QKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT  274 (371)
Q Consensus       197 ~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~  274 (371)
                      ..|++|.+|.++|||++ |||+||+|+++||+++.++.++.          ..+....  +++.+++.|+++..+++++
T Consensus       127 ~~g~~V~~V~~~SpA~~AGL~~GDvI~~vng~~v~~~~dl~----------~~ia~~~--~~v~~~I~r~g~~~~l~v~  193 (420)
T TIGR00054       127 EVGPVIELLDKNSIALEAGIEPGDEILSVNGNKIPGFKDVR----------QQIADIA--GEPMVEILAERENWTFEVM  193 (420)
T ss_pred             CCCceeeccCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHhhc--ccceEEEEEecCceEeccc
Confidence            36889999999999999 99999999999999999999875          4555444  6789999999887665443


No 29 
>PF00595 PDZ:  PDZ domain (Also known as DHR or GLGF) Coordinates are not yet available;  InterPro: IPR001478 PDZ domains are found in diverse signalling proteins in bacteria, yeasts, plants, insects and vertebrates [, ]. PDZ domains can occur in one or multiple copies and are nearly always found in cytoplasmic proteins. They bind either the carboxyl-terminal sequences of proteins or internal peptide sequences []. In most cases, interaction between a PDZ domain and its target is constitutive, with a binding affinity of 1 to 10 microns. However, agonist-dependent activation of cell surface receptors is sometimes required to promote interaction with a PDZ protein. PDZ domain proteins are frequently associated with the plasma membrane, a compartment where high concentrations of phosphatidylinositol 4,5-bisphosphate (PIP2) are found. Direct interaction between PIP2 and a subset of class II PDZ domains (syntenin, CASK, Tiam-1) has been demonstrated.  PDZ domains consist of 80 to 90 amino acids comprising six beta-strands (beta-A to beta-F) and two alpha-helices, A and B, compactly arranged in a globular structure. Peptide binding of the ligand takes place in an elongated surface groove as an anti-parallel beta-strand interacts with the beta-B strand and the B helix. The structure of PDZ domains allows binding to a free carboxylate group at the end of a peptide through a carboxylate-binding loop between the beta-A and beta-B strands.; GO: 0005515 protein binding; PDB: 3AXA_A 1WF8_A 1QAV_B 1QAU_A 1B8Q_A 1MC7_A 2KAW_A 1I16_A 1VB7_A 1WI4_A ....
Probab=98.50  E-value=2.3e-07  Score=71.18  Aligned_cols=72  Identities=24%  Similarity=0.322  Sum_probs=52.8

Q ss_pred             ccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhh
Q 017471          172 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVS  250 (371)
Q Consensus       172 ~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~  250 (371)
                      ...+|+.+....+ .         ...+++|.+|.++|||++ ||++||.|++|||+++.++...+        ....+.
T Consensus         9 ~~~lG~~l~~~~~-~---------~~~~~~V~~v~~~~~a~~~gl~~GD~Il~INg~~v~~~~~~~--------~~~~l~   70 (81)
T PF00595_consen    9 NGPLGFTLRGGSD-N---------DEKGVFVSSVVPGSPAERAGLKVGDRILEINGQSVRGMSHDE--------VVQLLK   70 (81)
T ss_dssp             TSBSSEEEEEEST-S---------SSEEEEEEEECTTSHHHHHTSSTTEEEEEETTEESTTSBHHH--------HHHHHH
T ss_pred             CCCcCEEEEecCC-C---------CcCCEEEEEEeCCChHHhcccchhhhhheeCCEeCCCCCHHH--------HHHHHH
Confidence            4568999887621 0         025899999999999999 99999999999999999886543        112222


Q ss_pred             ccCCCCEEEEEEE
Q 017471          251 QKYTGDSAAVKVL  263 (371)
Q Consensus       251 ~~~~g~~v~l~v~  263 (371)
                       . .+.+++|+|+
T Consensus        71 -~-~~~~v~L~V~   81 (81)
T PF00595_consen   71 -S-ASNPVTLTVQ   81 (81)
T ss_dssp             -H-STSEEEEEEE
T ss_pred             -C-CCCcEEEEEC
Confidence             2 3448888774


No 30 
>PRK10139 serine endoprotease; Provisional
Probab=98.47  E-value=3e-07  Score=93.11  Aligned_cols=64  Identities=16%  Similarity=0.323  Sum_probs=56.0

Q ss_pred             CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEE
Q 017471          198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNI  273 (371)
Q Consensus       198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v  273 (371)
                      .|++|.+|.++|||++ |||+||+|++|||++|.++.++.          +.+.+ .+ +++.++|.|+|+...+.+
T Consensus       390 ~Gv~V~~V~~~spA~~aGL~~GD~I~~Ing~~v~~~~~~~----------~~l~~-~~-~~v~l~v~R~g~~~~~~~  454 (455)
T PRK10139        390 KGIKIDEVVKGSPAAQAGLQKDDVIIGVNRDRVNSIAEMR----------KVLAA-KP-AIIALQIVRGNESIYLLL  454 (455)
T ss_pred             CceEEEEeCCCChHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHh-CC-CeEEEEEEECCEEEEEEe
Confidence            5899999999999999 99999999999999999999875          56654 33 689999999999877665


No 31 
>PLN00049 carboxyl-terminal processing protease; Provisional
Probab=98.46  E-value=5.9e-07  Score=89.33  Aligned_cols=69  Identities=20%  Similarity=0.302  Sum_probs=54.9

Q ss_pred             CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEe
Q 017471          198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITL  275 (371)
Q Consensus       198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l  275 (371)
                      .|++|..|.++|||++ ||++||+|++|||++|.++....        +...+ ....|+++.++|.|+|+..+++++-
T Consensus       102 ~g~~V~~V~~~SPA~~aGl~~GD~Iv~InG~~v~~~~~~~--------~~~~l-~g~~g~~v~ltv~r~g~~~~~~l~r  171 (389)
T PLN00049        102 AGLVVVAPAPGGPAARAGIRPGDVILAIDGTSTEGLSLYE--------AADRL-QGPEGSSVELTLRRGPETRLVTLTR  171 (389)
T ss_pred             CcEEEEEeCCCChHHHcCCCCCCEEEEECCEECCCCCHHH--------HHHHH-hcCCCCEEEEEEEECCEEEEEEEEe
Confidence            3899999999999999 99999999999999998753211        11233 3457899999999999887776654


No 32 
>TIGR02860 spore_IV_B stage IV sporulation protein B. SpoIVB, the stage IV sporulation protein B of endospore-forming bacteria such as Bacillus subtilis, is a serine proteinase, expressed in the spore (rather than mother cell) compartment, that participates in a proteolytic activation cascade for Sigma-K. It appears to be universal among endospore-forming bacteria and occurs nowhere else.
Probab=98.39  E-value=5.1e-07  Score=88.86  Aligned_cols=69  Identities=25%  Similarity=0.364  Sum_probs=57.3

Q ss_pred             CCceEEEEEC--------CCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCE
Q 017471          197 QKGVRIRRVD--------PTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK  267 (371)
Q Consensus       197 ~~gv~V~~V~--------~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~  267 (371)
                      .+||+|....        .+|||++ |||+||+|++|||++|++++|+.          +.+... .++++.+++.|+|+
T Consensus       104 t~GVlVvg~~~v~~~~g~~~SPAa~AGLq~GDiIvsING~~V~s~~DL~----------~iL~~~-~g~~V~LtV~R~Ge  172 (402)
T TIGR02860       104 TKGVLVVGFSDIETEKGKIHSPGEEAGIQIGDRILKINGEKIKNMDDLA----------NLINKA-GGEKLTLTIERGGK  172 (402)
T ss_pred             cCEEEEEEEEcccccCCCCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHHHhC-CCCeEEEEEEECCE
Confidence            4689886653        2589998 99999999999999999999875          555544 58899999999999


Q ss_pred             EEEEEEEec
Q 017471          268 ILNFNITLA  276 (371)
Q Consensus       268 ~~~~~v~l~  276 (371)
                      ..++++++.
T Consensus       173 ~~tv~V~Pv  181 (402)
T TIGR02860       173 IIETVIKPV  181 (402)
T ss_pred             EEEEEEEEe
Confidence            998888754


No 33 
>PRK10942 serine endoprotease; Provisional
Probab=98.36  E-value=6.3e-07  Score=91.20  Aligned_cols=64  Identities=20%  Similarity=0.351  Sum_probs=55.7

Q ss_pred             CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEE
Q 017471          198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNI  273 (371)
Q Consensus       198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v  273 (371)
                      .|++|.+|.++|||++ ||++||+|++|||++|.++.++.          +.+.. .+ +.+.++|.|+|+.+.+.+
T Consensus       408 ~gvvV~~V~~~S~A~~aGL~~GDvIv~VNg~~V~s~~dl~----------~~l~~-~~-~~v~l~V~R~g~~~~v~~  472 (473)
T PRK10942        408 KGVVVDNVKPGTPAAQIGLKKGDVIIGANQQPVKNIAELR----------KILDS-KP-SVLALNIQRGDSSIYLLM  472 (473)
T ss_pred             CCeEEEEeCCCChHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHh-CC-CeEEEEEEECCEEEEEEe
Confidence            5899999999999999 99999999999999999999875          55554 33 789999999999877654


No 34 
>cd00992 PDZ_signaling PDZ domain found in a variety of Eumetazoan signaling molecules, often in tandem arrangements. May be responsible for specific protein-protein interactions, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of PDZ domains an N-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in proteases.
Probab=98.35  E-value=8.7e-07  Score=67.65  Aligned_cols=52  Identities=23%  Similarity=0.427  Sum_probs=42.1

Q ss_pred             cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEec--CCCCc
Q 017471          173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIA--NDGTV  235 (371)
Q Consensus       173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~--~~~~l  235 (371)
                      ..+|+.+....+           ...|++|.+|.++|||++ ||++||+|++|||+++.  +..++
T Consensus        12 ~~~G~~~~~~~~-----------~~~~~~V~~v~~~s~a~~~gl~~GD~I~~ing~~i~~~~~~~~   66 (82)
T cd00992          12 GGLGFSLRGGKD-----------SGGGIFVSRVEPGGPAERGGLRVGDRILEVNGVSVEGLTHEEA   66 (82)
T ss_pred             CCcCEEEeCccc-----------CCCCeEEEEECCCChHHhCCCCCCCEEEEECCEEcCccCHHHH
Confidence            458888875511           135899999999999999 99999999999999998  44443


No 35 
>PF14685 Tricorn_PDZ:  Tricorn protease PDZ domain; PDB: 1N6F_D 1N6D_C 1N6E_C 1K32_A.
Probab=98.34  E-value=4.3e-06  Score=65.17  Aligned_cols=65  Identities=22%  Similarity=0.330  Sum_probs=44.8

Q ss_pred             CCceEEEEECCC--------Ccccc-C--CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEEC
Q 017471          197 QKGVRIRRVDPT--------APESE-V--LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD  265 (371)
Q Consensus       197 ~~gv~V~~V~~~--------spA~~-G--L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~  265 (371)
                      ..+..|.+|.++        ||..+ |  +++||+|++|||++++...++.           .+...+.|+++.|+|.+.
T Consensus        11 ~~~y~I~~I~~gd~~~~~~~sPL~~pGv~v~~GD~I~aInG~~v~~~~~~~-----------~lL~~~agk~V~Ltv~~~   79 (88)
T PF14685_consen   11 NGGYRIARIYPGDPWNPNARSPLAQPGVDVREGDYILAINGQPVTADANPY-----------RLLEGKAGKQVLLTVNRK   79 (88)
T ss_dssp             TTEEEEEEE-BS-TTSSS-B-GGGGGS----TT-EEEEETTEE-BTTB-HH-----------HHHHTTTTSEEEEEEE-S
T ss_pred             CCEEEEEEEeCCCCCCccccCCccCCCCCCCCCCEEEEECCEECCCCCCHH-----------HHhcccCCCEEEEEEecC
Confidence            367889999875        78877 6  5699999999999999888763           455567899999999997


Q ss_pred             C-EEEEEE
Q 017471          266 S-KILNFN  272 (371)
Q Consensus       266 g-~~~~~~  272 (371)
                      + +.+++.
T Consensus        80 ~~~~R~v~   87 (88)
T PF14685_consen   80 PGGARTVV   87 (88)
T ss_dssp             TT-EEEEE
T ss_pred             CCCceEEE
Confidence            6 455554


No 36 
>COG0793 Prc Periplasmic protease [Cell envelope biogenesis, outer membrane]
Probab=98.33  E-value=1.5e-06  Score=86.73  Aligned_cols=83  Identities=24%  Similarity=0.390  Sum_probs=63.6

Q ss_pred             ccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCC--ccccccccchhhhh
Q 017471          172 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGT--VPFRHGERIGFSYL  248 (371)
Q Consensus       172 ~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~--l~~~~~~~~~~~~~  248 (371)
                      +..+|++++.-             +..++.|.++.+++||++ ||++||+|++|||+++....-  .           -.
T Consensus        99 ~~GiG~~i~~~-------------~~~~~~V~s~~~~~PA~kagi~~GD~I~~IdG~~~~~~~~~~a-----------v~  154 (406)
T COG0793          99 FGGIGIELQME-------------DIGGVKVVSPIDGSPAAKAGIKPGDVIIKIDGKSVGGVSLDEA-----------VK  154 (406)
T ss_pred             ccceeEEEEEe-------------cCCCcEEEecCCCChHHHcCCCCCCEEEEECCEEccCCCHHHH-----------HH
Confidence            56688888754             126899999999999999 999999999999999987642  1           12


Q ss_pred             hhccCCCCEEEEEEEECCEEEEEEEEeccc
Q 017471          249 VSQKYTGDSAAVKVLRDSKILNFNITLATH  278 (371)
Q Consensus       249 ~~~~~~g~~v~l~v~R~g~~~~~~v~l~~~  278 (371)
                      ..+..+|..|+|++.|.+....+.+++.+.
T Consensus       155 ~irG~~Gt~V~L~i~r~~~~k~~~v~l~Re  184 (406)
T COG0793         155 LIRGKPGTKVTLTILRAGGGKPFTVTLTRE  184 (406)
T ss_pred             HhCCCCCCeEEEEEEEcCCCceeEEEEEEE
Confidence            334578999999999985454555555543


No 37 
>COG3480 SdrC Predicted secreted protein containing a PDZ domain [Signal transduction mechanisms]
Probab=98.20  E-value=4e-06  Score=78.87  Aligned_cols=72  Identities=25%  Similarity=0.295  Sum_probs=65.4

Q ss_pred             CCceEEEEECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEE-CCEEEEEEEEe
Q 017471          197 QKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLR-DSKILNFNITL  275 (371)
Q Consensus       197 ~~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R-~g~~~~~~v~l  275 (371)
                      ..||++..+..+||+..-|++||.|++|||+++.+.+++.          ..+...++|++|++++.| +++....++++
T Consensus       129 y~gvyv~~v~~~~~~~gkl~~gD~i~avdg~~f~s~~e~i----------~~v~~~k~Gd~VtI~~~r~~~~~~~~~~tl  198 (342)
T COG3480         129 YAGVYVLSVIDNSPFKGKLEAGDTIIAVDGEPFTSSDELI----------DYVSSKKPGDEVTIDYERHNETPEIVTITL  198 (342)
T ss_pred             EeeEEEEEccCCcchhceeccCCeEEeeCCeecCCHHHHH----------HHHhccCCCCeEEEEEEeccCCCceEEEEE
Confidence            3699999999999998789999999999999999999976          788888999999999997 88888888888


Q ss_pred             ccc
Q 017471          276 ATH  278 (371)
Q Consensus       276 ~~~  278 (371)
                      ...
T Consensus       199 ~~~  201 (342)
T COG3480         199 IKN  201 (342)
T ss_pred             Eee
Confidence            776


No 38 
>PRK09681 putative type II secretion protein GspC; Provisional
Probab=98.05  E-value=9.4e-06  Score=76.16  Aligned_cols=68  Identities=25%  Similarity=0.314  Sum_probs=55.2

Q ss_pred             CceEEEEECCCCcc---cc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEE
Q 017471          198 KGVRIRRVDPTAPE---SE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNI  273 (371)
Q Consensus       198 ~gv~V~~V~~~spA---~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v  273 (371)
                      .|+.=-++.|+..+   .+ |||+||++++|||.++++.++..          .++........++++|+|||+..++.+
T Consensus       204 ~Gl~GYrl~Pgkd~~lF~~~GLq~GDva~sING~dL~D~~qa~----------~l~~~L~~~tei~ltVeRdGq~~~i~i  273 (276)
T PRK09681        204 EGIVGYAVKPGADRSLFDASGFKEGDIAIALNQQDFTDPRAMI----------ALMRQLPSMDSIQLTVLRKGARHDISI  273 (276)
T ss_pred             CCceEEEECCCCcHHHHHHcCCCCCCEEEEeCCeeCCCHHHHH----------HHHHHhccCCeEEEEEEECCEEEEEEE
Confidence            35333467787544   45 99999999999999999888754          677777888999999999999999988


Q ss_pred             Ee
Q 017471          274 TL  275 (371)
Q Consensus       274 ~l  275 (371)
                      .+
T Consensus       274 ~l  275 (276)
T PRK09681        274 AL  275 (276)
T ss_pred             Ec
Confidence            75


No 39 
>PF04495 GRASP55_65:  GRASP55/65 PDZ-like domain ;  InterPro: IPR007583 GRASP55 (Golgi reassembly stacking protein of 55 kDa) and GRASP65 (a 65 kDa) protein are highly homologous. GRASP55 is a component of the Golgi stacking machinery. GRASP65, an N-ethylmaleimide-sensitive membrane protein required for the stacking of Golgi cisternae in a cell-free system [].; PDB: 3RLE_A 4EDJ_A.
Probab=97.96  E-value=1.1e-05  Score=68.34  Aligned_cols=86  Identities=23%  Similarity=0.331  Sum_probs=54.7

Q ss_pred             ccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCC-CCEEEEECCEEecCCCCccccccccchhhhhh
Q 017471          172 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKP-SDIILSFDGIDIANDGTVPFRHGERIGFSYLV  249 (371)
Q Consensus       172 ~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~-GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~  249 (371)
                      .+.||++++--.. +.       ....+..|.+|.|+|||++ ||++ .|.|+.+|+..+++.+++.          +.+
T Consensus        25 ~g~LG~sv~~~~~-~~-------~~~~~~~Vl~V~p~SPA~~AGL~p~~DyIig~~~~~l~~~~~l~----------~~v   86 (138)
T PF04495_consen   25 QGLLGISVRFESF-EG-------AEEEGWHVLRVAPNSPAAKAGLEPFFDYIIGIDGGLLDDEDDLF----------ELV   86 (138)
T ss_dssp             SSSS-EEEEEEE--TT-------GCCCEEEEEEE-TTSHHHHTT--TTTEEEEEETTCE--STCHHH----------HHH
T ss_pred             CCCCcEEEEEecc-cc-------cccceEEEeEecCCCHHHHCCccccccEEEEccceecCCHHHHH----------HHH
Confidence            4678887764411 10       1246899999999999999 9999 6999999999998776653          444


Q ss_pred             hccCCCCEEEEEEEECC--EEEEEEEEec
Q 017471          250 SQKYTGDSAAVKVLRDS--KILNFNITLA  276 (371)
Q Consensus       250 ~~~~~g~~v~l~v~R~g--~~~~~~v~l~  276 (371)
                      . .+.++++.+.|+...  +.+++++++.
T Consensus        87 ~-~~~~~~l~L~Vyns~~~~vR~V~i~P~  114 (138)
T PF04495_consen   87 E-ANENKPLQLYVYNSKTDSVREVTITPS  114 (138)
T ss_dssp             H-HTTTS-EEEEEEETTTTCEEEEEE---
T ss_pred             H-HcCCCcEEEEEEECCCCeEEEEEEEcC
Confidence            4 567889999999754  4455555543


No 40 
>PF00089 Trypsin:  Trypsin;  InterPro: IPR001254 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine proteases belong to the MEROPS peptidase family S1 (chymotrypsin family, clan PA(S))and to peptidase family S6 (Hap serine peptidases). The chymotrypsin family is almost totally confined to animals, although trypsin-like enzymes are found in actinomycetes of the genera Streptomyces and Saccharopolyspora, and in the fungus Fusarium oxysporum []. The enzymes are inherently secreted, being synthesised with a signal peptide that targets them to the secretory pathway. Animal enzymes are either secreted directly, packaged into vesicles for regulated secretion, or are retained in leukocyte granules []. The Hap family, 'Haemophilus adhesion and penetration', are proteins that play a role in the interaction with human epithelial cells. The serine protease activity is localized at the N-terminal domain, whereas the binding domain is in the C-terminal region. ; GO: 0004252 serine-type endopeptidase activity, 0006508 proteolysis; PDB: 1SPJ_A 1A5I_A 2ZGH_A 2ZKS_A 2ZGJ_A 2ZGC_A 2ODP_A 2I6Q_A 2I6S_A 2ODQ_A ....
Probab=97.91  E-value=6.4e-05  Score=67.39  Aligned_cols=120  Identities=16%  Similarity=0.165  Sum_probs=74.1

Q ss_pred             CCCEEEEEEecC-CCcCCccceecCCCC---CCCCeEEEEEeCCCCCCc---eeeeeEEeeeeee---eccCCCeEEeEE
Q 017471           40 ICIYTMLTVEDD-EFWEGVLPVEFGELP---ALQDAVTVVGYPIGGDTI---SVTSGVVSRIEIL---SYVHGSTELLGL  109 (371)
Q Consensus        40 ~~DlAlLkv~~~-~~~~~l~~~~l~~s~---~lgd~V~~iG~p~g~~~~---s~t~G~Vs~~~~~---~~~~~~~~~~~i  109 (371)
                      .+|+|||+++.+ .+.+.+.++.+....   ..++.+.++|++......   ......+.-+...   ...........+
T Consensus        86 ~~DiAll~L~~~~~~~~~~~~~~l~~~~~~~~~~~~~~~~G~~~~~~~~~~~~~~~~~~~~~~~~~c~~~~~~~~~~~~~  165 (220)
T PF00089_consen   86 DNDIALLKLDRPITFGDNIQPICLPSAGSDPNVGTSCIVVGWGRTSDNGYSSNLQSVTVPVVSRKTCRSSYNDNLTPNMI  165 (220)
T ss_dssp             TTSEEEEEESSSSEHBSSBEESBBTSTTHTTTTTSEEEEEESSBSSTTSBTSBEEEEEEEEEEHHHHHHHTTTTSTTTEE
T ss_pred             cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccc
Confidence            479999999987 333466777777632   567999999998753321   3343444333321   111111111234


Q ss_pred             EEEe----eccCCCCCCceecCCCcEEEEEeeecccCCccceeccccCcchhHh
Q 017471          110 QIDA----AINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHF  159 (371)
Q Consensus       110 ~~da----~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~~~~~~~aiP~~~i~~~  159 (371)
                      ..+.    ....|+||||+++.++.++||.+.....+......+..+++...++
T Consensus       166 c~~~~~~~~~~~g~sG~pl~~~~~~lvGI~s~~~~c~~~~~~~v~~~v~~~~~W  219 (220)
T PF00089_consen  166 CAGSSGSGDACQGDSGGPLICNNNYLVGIVSFGENCGSPNYPGVYTRVSSYLDW  219 (220)
T ss_dssp             EEETTSSSBGGTTTTTSEEEETTEEEEEEEEEESSSSBTTSEEEEEEGGGGHHH
T ss_pred             cccccccccccccccccccccceeeecceeeecCCCCCCCcCEEEEEHHHhhcc
Confidence            4444    6789999999999888899999876432222234666777665554


No 41 
>PRK11186 carboxy-terminal protease; Provisional
Probab=97.82  E-value=5.7e-05  Score=79.50  Aligned_cols=71  Identities=15%  Similarity=0.135  Sum_probs=49.5

Q ss_pred             CceEEEEECCCCcccc--CCCCCCEEEEEC--CEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEEC---CEEEE
Q 017471          198 KGVRIRRVDPTAPESE--VLKPSDIILSFD--GIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD---SKILN  270 (371)
Q Consensus       198 ~gv~V~~V~~~spA~~--GL~~GDvIl~vn--G~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~---g~~~~  270 (371)
                      .+++|.+|.|||||++  ||++||+|++||  |.++.+...+.      +.--..+.+..+|.+|.|+|.|+   ++..+
T Consensus       255 ~~~~V~~vipGsPA~ka~gLk~GD~IlaVn~~g~~~~dv~g~~------~~~vv~lirG~~Gt~V~LtV~r~~~~~~~~~  328 (667)
T PRK11186        255 DYTVINSLVAGGPAAKSKKLSVGDKIVGVGQDGKPIVDVIGWR------LDDVVALIKGPKGSKVRLEILPAGKGTKTRI  328 (667)
T ss_pred             CeEEEEEccCCChHHHhCCCCCCCEEEEECCCCCcccccccCC------HHHHHHHhcCCCCCEEEEEEEeCCCCCceEE
Confidence            4689999999999997  899999999999  56554433221      00012334456899999999994   44555


Q ss_pred             EEEE
Q 017471          271 FNIT  274 (371)
Q Consensus       271 ~~v~  274 (371)
                      ++++
T Consensus       329 vtl~  332 (667)
T PRK11186        329 VTLT  332 (667)
T ss_pred             EEEE
Confidence            5544


No 42 
>COG3975 Predicted protease with the C-terminal PDZ domain [General function prediction only]
Probab=97.81  E-value=4.3e-05  Score=76.50  Aligned_cols=86  Identities=21%  Similarity=0.364  Sum_probs=66.4

Q ss_pred             cceeeeEcCCHHHHHhccCCC--CCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc
Q 017471          175 LGVEWQKMENPDLRVAMSMKA--DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ  251 (371)
Q Consensus       175 lGi~~~~~~~~~~~~~~gl~~--~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~  251 (371)
                      .|+.+.+...+  +-.+|++-  +..+.+|+.|.++|||++ ||.+||.|++|||.   +               ..+.+
T Consensus       439 ~gL~~~~~~~~--~~~LGl~v~~~~g~~~i~~V~~~gPA~~AGl~~Gd~ivai~G~---s---------------~~l~~  498 (558)
T COG3975         439 FGLTFTPKPRE--AYYLGLKVKSEGGHEKITFVFPGGPAYKAGLSPGDKIVAINGI---S---------------DQLDR  498 (558)
T ss_pred             cceEEEecCCC--CcccceEecccCCeeEEEecCCCChhHhccCCCccEEEEEcCc---c---------------ccccc
Confidence            46777665222  34566543  345689999999999999 99999999999999   1               13445


Q ss_pred             cCCCCEEEEEEEECCEEEEEEEEeccccc
Q 017471          252 KYTGDSAAVKVLRDSKILNFNITLATHRR  280 (371)
Q Consensus       252 ~~~g~~v~l~v~R~g~~~~~~v~l~~~~~  280 (371)
                      .+.++.+++++.|.|+.+++.+++.....
T Consensus       499 ~~~~d~i~v~~~~~~~L~e~~v~~~~~~~  527 (558)
T COG3975         499 YKVNDKIQVHVFREGRLREFLVKLGGDPT  527 (558)
T ss_pred             cccccceEEEEccCCceEEeecccCCCcc
Confidence            67899999999999999999988876543


No 43 
>KOG3129 consensus 26S proteasome regulatory complex, subunit PSMD9 [Posttranslational modification, protein turnover, chaperones]
Probab=97.72  E-value=7.7e-05  Score=66.26  Aligned_cols=73  Identities=22%  Similarity=0.216  Sum_probs=59.1

Q ss_pred             ceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEecc
Q 017471          199 GVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLAT  277 (371)
Q Consensus       199 gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~~  277 (371)
                      -++|.+|.|+|||+. ||+.||.|+++....-.++..++        =-..+.+...++.+.+++.|.|+...+.++++.
T Consensus       140 Fa~V~sV~~~SPA~~aGl~~gD~il~fGnV~sgn~~~lq--------~i~~~v~~~e~~~v~v~v~R~g~~v~L~ltP~~  211 (231)
T KOG3129|consen  140 FAVVDSVVPGSPADEAGLCVGDEILKFGNVHSGNFLPLQ--------NIAAVVQSNEDQIVSVTVIREGQKVVLSLTPKK  211 (231)
T ss_pred             eEEEeecCCCChhhhhCcccCceEEEecccccccchhHH--------HHHHHHHhccCcceeEEEecCCCEEEEEeCccc
Confidence            468999999999999 99999999999887776666543        012444567889999999999999999988875


Q ss_pred             cc
Q 017471          278 HR  279 (371)
Q Consensus       278 ~~  279 (371)
                      ..
T Consensus       212 W~  213 (231)
T KOG3129|consen  212 WQ  213 (231)
T ss_pred             cc
Confidence            43


No 44 
>KOG3553 consensus Tax interaction protein TIP1 [Cell wall/membrane/envelope biogenesis]
Probab=97.37  E-value=0.00017  Score=56.58  Aligned_cols=35  Identities=31%  Similarity=0.484  Sum_probs=32.2

Q ss_pred             CCceEEEEECCCCcccc-CCCCCCEEEEECCEEecC
Q 017471          197 QKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAN  231 (371)
Q Consensus       197 ~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~  231 (371)
                      ..|++|++|.++|||+. ||+.+|.|+.+||...+-
T Consensus        58 D~GiYvT~V~eGsPA~~AGLrihDKIlQvNG~DfTM   93 (124)
T KOG3553|consen   58 DKGIYVTRVSEGSPAEIAGLRIHDKILQVNGWDFTM   93 (124)
T ss_pred             CccEEEEEeccCChhhhhcceecceEEEecCceeEE
Confidence            57999999999999999 999999999999987653


No 45 
>PF12812 PDZ_1:  PDZ-like domain
Probab=97.32  E-value=0.00033  Score=53.41  Aligned_cols=60  Identities=12%  Similarity=0.057  Sum_probs=51.8

Q ss_pred             cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCcc
Q 017471          173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVP  236 (371)
Q Consensus       173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~  236 (371)
                      -|.|..++++ +-+.++.++++-   |+++.....++++.. |+.+|-+|.+|||+++.+.+++.
T Consensus         9 ~~~Ga~f~~L-s~q~aR~~~~~~---~gv~v~~~~g~~~~~~~i~~g~iI~~Vn~kpt~~Ld~f~   69 (78)
T PF12812_consen    9 EVCGAVFHDL-SYQQARQYGIPV---GGVYVAVSGGSLAFAGGISKGFIITSVNGKPTPDLDDFI   69 (78)
T ss_pred             EEcCeecccC-CHHHHHHhCCCC---CEEEEEecCCChhhhCCCCCCeEEEeECCcCCcCHHHHH
Confidence            4789999999 899999999973   355666788999988 69999999999999999998864


No 46 
>COG3031 PulC Type II secretory pathway, component PulC [Intracellular trafficking and secretion]
Probab=97.30  E-value=0.00019  Score=65.11  Aligned_cols=66  Identities=18%  Similarity=0.223  Sum_probs=52.4

Q ss_pred             ceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEE
Q 017471          199 GVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT  274 (371)
Q Consensus       199 gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~  274 (371)
                      |..+.-..+++..++ |||.||+.+++|+..+++.+++.          .++.....-+.+.++|.|+|++..+.|.
T Consensus       208 Gyr~~pgkd~slF~~sglq~GDIavaiNnldltdp~~m~----------~llq~l~~m~s~qlTv~R~G~rhdInV~  274 (275)
T COG3031         208 GYRFEPGKDGSLFYKSGLQRGDIAVAINNLDLTDPEDMF----------RLLQMLRNMPSLQLTVIRRGKRHDINVR  274 (275)
T ss_pred             EEEecCCCCcchhhhhcCCCcceEEEecCcccCCHHHHH----------HHHHhhhcCcceEEEEEecCccceeeec
Confidence            333333344566677 99999999999999999999864          6676666667899999999999988875


No 47 
>COG3591 V8-like Glu-specific endopeptidase [Amino acid transport and metabolism]
Probab=96.98  E-value=0.0036  Score=58.13  Aligned_cols=87  Identities=28%  Similarity=0.414  Sum_probs=57.0

Q ss_pred             CCCCeEEEEEeCCCCCC---ceeeeeEEeeeeeeeccCCCeEEeEEEEEeeccCCCCCCceecCCCcEEEEEeeecccCC
Q 017471           67 ALQDAVTVVGYPIGGDT---ISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHED  143 (371)
Q Consensus        67 ~lgd~V~~iG~p~g~~~---~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~  143 (371)
                      .+++.+-++|||.+...   .-...+.|-.+..          ..++.|+.+.||+||.|+++.+.++||+.+......+
T Consensus       159 ~~~d~i~v~GYP~dk~~~~~~~e~t~~v~~~~~----------~~l~y~~dT~pG~SGSpv~~~~~~vigv~~~g~~~~~  228 (251)
T COG3591         159 KANDRITVIGYPGDKPNIGTMWESTGKVNSIKG----------NKLFYDADTLPGSSGSPVLISKDEVIGVHYNGPGANG  228 (251)
T ss_pred             ccCceeEEEeccCCCCcceeEeeecceeEEEec----------ceEEEEecccCCCCCCceEecCceEEEEEecCCCccc
Confidence            45688999999987552   2233444444331          2588999999999999999999999999987654333


Q ss_pred             ccceeccc-cCcchhHhHhhh
Q 017471          144 VENIGYVI-PTPVIMHFIQDY  163 (371)
Q Consensus       144 ~~~~~~ai-P~~~i~~~l~~L  163 (371)
                      ....++++ -...++++++++
T Consensus       229 ~~~~n~~vr~t~~~~~~I~~~  249 (251)
T COG3591         229 GSLANNAVRLTPEILNFIQQN  249 (251)
T ss_pred             ccccCcceEecHHHHHHHHHh
Confidence            33333332 233455555443


No 48 
>PF00863 Peptidase_C4:  Peptidase family C4 This family belongs to family C4 of the peptidase classification.;  InterPro: IPR001730 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad [].  Nuclear inclusion A (NIA) proteases from potyviruses are cysteine peptidases belong to the MEROPS peptidase family C4 (NIa protease family, clan PA(C)) [, ].  Potyviruses include plant viruses in which the single-stranded RNA encodes a polyprotein with NIA protease activity, where proteolytic cleavage is specific for Gln+Gly sites. The NIA protease acts on the polyprotein, releasing itself by Gln+Gly cleavage at both the N- and C-termini. It further processes the polyprotein by cleavage at five similar sites in the C-terminal half of the sequence. In addition to its C-terminal protease activity, the NIA protease contains an N-terminal domain that has been implicated in the transcription process []. This peptidase is present in the nuclear inclusion protein of potyviruses.; GO: 0008234 cysteine-type peptidase activity, 0006508 proteolysis; PDB: 3MMG_B 1Q31_B 1LVB_A 1LVM_A.
Probab=96.88  E-value=0.015  Score=53.58  Aligned_cols=112  Identities=19%  Similarity=0.163  Sum_probs=53.3

Q ss_pred             CCCEEEEEEecCCCcCCccceecC---CCCCCCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeecc
Q 017471           40 ICIYTMLTVEDDEFWEGVLPVEFG---ELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAIN  116 (371)
Q Consensus        40 ~~DlAlLkv~~~~~~~~l~~~~l~---~s~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~  116 (371)
                      ..|+.++|...+     +||.+-.   ..|..+|.|..+|..+.....+.+...-|.+..  ...+.    +..+-.+..
T Consensus        81 ~~DiviirmPkD-----fpPf~~kl~FR~P~~~e~v~mVg~~fq~k~~~s~vSesS~i~p--~~~~~----fWkHwIsTk  149 (235)
T PF00863_consen   81 GRDIVIIRMPKD-----FPPFPQKLKFRAPKEGERVCMVGSNFQEKSISSTVSESSWIYP--EENSH----FWKHWISTK  149 (235)
T ss_dssp             CSSEEEEE--TT-----S----S---B----TT-EEEEEEEECSSCCCEEEEEEEEEEEE--ETTTT----EEEE-C---
T ss_pred             CccEEEEeCCcc-----cCCcchhhhccCCCCCCEEEEEEEEEEcCCeeEEECCceEEee--cCCCC----eeEEEecCC
Confidence            589999998764     5554322   346678999999997754332333332222221  11122    344444556


Q ss_pred             CCCCCCceec-CCCcEEEEEeeecccCCccceeccccCcchhHhHhhh
Q 017471          117 SGNSGGPAFN-DKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDY  163 (371)
Q Consensus       117 ~G~SGGPlvn-~~G~VIGI~~~~~~~~~~~~~~~aiP~~~i~~~l~~L  163 (371)
                      .|+-|.|+++ .+|.+|||.+.... ....++.-++|-+....+++..
T Consensus       150 ~G~CG~PlVs~~Dg~IVGiHsl~~~-~~~~N~F~~f~~~f~~~~l~~~  196 (235)
T PF00863_consen  150 DGDCGLPLVSTKDGKIVGIHSLTSN-TSSRNYFTPFPDDFEEFYLENI  196 (235)
T ss_dssp             TT-TT-EEEETTT--EEEEEEEEET-TTSSEEEEE--TTHHHHHCC-C
T ss_pred             CCccCCcEEEcCCCcEEEEEcCccC-CCCeEEEEcCCHHHHHHHhccc
Confidence            7999999998 66999999876542 2344555556556555555443


No 49 
>cd00190 Tryp_SPc Trypsin-like serine protease; Many of these are synthesized as inactive precursor zymogens that are cleaved during limited proteolysis to generate their active forms. Alignment contains also inactive enzymes that have substitutions of the catalytic triad residues.
Probab=96.33  E-value=0.03  Score=50.23  Aligned_cols=100  Identities=17%  Similarity=0.145  Sum_probs=54.4

Q ss_pred             CCCEEEEEEecCC-CcCCccceecCCC---CCCCCeEEEEEeCCCCCC----ceeeeeEEeeeeeee---ccC---C-Ce
Q 017471           40 ICIYTMLTVEDDE-FWEGVLPVEFGEL---PALQDAVTVVGYPIGGDT----ISVTSGVVSRIEILS---YVH---G-ST  104 (371)
Q Consensus        40 ~~DlAlLkv~~~~-~~~~l~~~~l~~s---~~lgd~V~~iG~p~g~~~----~s~t~G~Vs~~~~~~---~~~---~-~~  104 (371)
                      .+|+|||+++.+. +...+.|+.+...   ...++.+.+.|+......    .......+.-+....   ...   . ..
T Consensus        88 ~~DiAll~L~~~~~~~~~v~picl~~~~~~~~~~~~~~~~G~g~~~~~~~~~~~~~~~~~~~~~~~~C~~~~~~~~~~~~  167 (232)
T cd00190          88 DNDIALLKLKRPVTLSDNVRPICLPSSGYNLPAGTTCTVSGWGRTSEGGPLPDVLQEVNVPIVSNAECKRAYSYGGTITD  167 (232)
T ss_pred             cCCEEEEEECCcccCCCcccceECCCccccCCCCCEEEEEeCCcCCCCCCCCceeeEEEeeeECHHHhhhhccCcccCCC
Confidence            4899999999653 2223677777654   245689999998654321    112222222221100   000   0 00


Q ss_pred             EEeEEEE---EeeccCCCCCCceecCC---CcEEEEEeeec
Q 017471          105 ELLGLQI---DAAINSGNSGGPAFNDK---GKCVGIAFQSL  139 (371)
Q Consensus       105 ~~~~i~~---da~i~~G~SGGPlvn~~---G~VIGI~~~~~  139 (371)
                      ..-+...   +...-+|+||||++...   +.++||.+...
T Consensus       168 ~~~C~~~~~~~~~~c~gdsGgpl~~~~~~~~~lvGI~s~g~  208 (232)
T cd00190         168 NMLCAGGLEGGKDACQGDSGGPLVCNDNGRGVLVGIVSWGS  208 (232)
T ss_pred             ceEeeCCCCCCCccccCCCCCcEEEEeCCEEEEEEEEehhh
Confidence            0001111   23345699999999765   89999986543


No 50 
>PF10459 Peptidase_S46:  Peptidase S46;  InterPro: IPR019500 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This entry represents S46 peptidases, where dipeptidyl-peptidase 7 (DPP-7) is the best-characterised member of this family. It is a serine peptidase that is located on the cell surface and is predicted to have two N-terminal transmembrane domains. 
Probab=96.25  E-value=0.0035  Score=66.48  Aligned_cols=56  Identities=23%  Similarity=0.391  Sum_probs=39.3

Q ss_pred             EEEEEeeccCCCCCCceecCCCcEEEEEeeecccC--C----ccce--eccccCcchhHhHhhh
Q 017471          108 GLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE--D----VENI--GYVIPTPVIMHFIQDY  163 (371)
Q Consensus       108 ~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~--~----~~~~--~~aiP~~~i~~~l~~L  163 (371)
                      .+.++..|..||||+|++|.+||+||++|-...++  +    .+..  +..|.+..+..+|+++
T Consensus       623 ~FlstnDitGGNSGSPvlN~~GeLVGl~FDgn~Esl~~D~~fdp~~~R~I~VDiRyvL~~ldkv  686 (698)
T PF10459_consen  623 NFLSTNDITGGNSGSPVLNAKGELVGLAFDGNWESLSGDIAFDPELNRTIHVDIRYVLWALDKV  686 (698)
T ss_pred             EEEeccCcCCCCCCCccCCCCceEEEEeecCchhhcccccccccccceeEEEEHHHHHHHHHHH
Confidence            46788889999999999999999999998654332  1    1222  3445455566666654


No 51 
>KOG3209 consensus WW domain-containing protein [General function prediction only]
Probab=96.04  E-value=0.0066  Score=62.87  Aligned_cols=55  Identities=27%  Similarity=0.344  Sum_probs=42.9

Q ss_pred             EEEECCCCcccc-C-CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECC
Q 017471          202 IRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDS  266 (371)
Q Consensus       202 V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g  266 (371)
                      |.+|.+||||++ | |+.||.|++|||+.|.+...-.        .-.++  +..|-+|+|+|.-..
T Consensus       782 iGrIieGSPAdRCgkLkVGDrilAVNG~sI~~lsHad--------iv~LI--KdaGlsVtLtIip~e  838 (984)
T KOG3209|consen  782 IGRIIEGSPADRCGKLKVGDRILAVNGQSILNLSHAD--------IVSLI--KDAGLSVTLTIIPPE  838 (984)
T ss_pred             ccccccCChhHhhccccccceEEEecCeeeeccCchh--------HHHHH--HhcCceEEEEEcChh
Confidence            778999999999 5 9999999999999998876532        00122  457889999987643


No 52 
>KOG3580 consensus Tight junction proteins [Signal transduction mechanisms]
Probab=96.01  E-value=0.0043  Score=63.05  Aligned_cols=59  Identities=19%  Similarity=0.286  Sum_probs=45.5

Q ss_pred             CceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEE
Q 017471          198 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLR  264 (371)
Q Consensus       198 ~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R  264 (371)
                      -|+.|..|.++|||++ ||+.||.||.||..+..+.---     +.+   ..+....+|+.+++.-.+
T Consensus       429 VGIFVaGvqegspA~~eGlqEGDQIL~VN~vdF~nl~RE-----eAV---lfLL~lPkGEevtilaQ~  488 (1027)
T KOG3580|consen  429 VGIFVAGVQEGSPAEQEGLQEGDQILKVNTVDFRNLVRE-----EAV---LFLLELPKGEEVTILAQS  488 (1027)
T ss_pred             eeEEEeecccCCchhhccccccceeEEeccccchhhhHH-----HHH---HHHhcCCCCcEEeehhhh
Confidence            4899999999999999 9999999999999988765211     111   245557789988886544


No 53 
>smart00020 Tryp_SPc Trypsin-like serine protease. Many of these are synthesised as inactive precursor zymogens that are cleaved during limited proteolysis to generate their active forms. A few, however, are active as single chain molecules, and others are inactive due to substitutions of the catalytic triad residues.
Probab=95.86  E-value=0.042  Score=49.46  Aligned_cols=100  Identities=17%  Similarity=0.137  Sum_probs=54.6

Q ss_pred             CCCEEEEEEecCC-CcCCccceecCCC---CCCCCeEEEEEeCCCCCC-----ceeeeeEEeeeeeeecc---C-----C
Q 017471           40 ICIYTMLTVEDDE-FWEGVLPVEFGEL---PALQDAVTVVGYPIGGDT-----ISVTSGVVSRIEILSYV---H-----G  102 (371)
Q Consensus        40 ~~DlAlLkv~~~~-~~~~l~~~~l~~s---~~lgd~V~~iG~p~g~~~-----~s~t~G~Vs~~~~~~~~---~-----~  102 (371)
                      .+|+|||+++.+. +.+.+.|+.+...   ...++.+.+.|++.....     .......+.-+......   .     .
T Consensus        88 ~~DiAll~L~~~i~~~~~~~pi~l~~~~~~~~~~~~~~~~g~g~~~~~~~~~~~~~~~~~~~~~~~~~C~~~~~~~~~~~  167 (229)
T smart00020       88 DNDIALLKLKSPVTLSDNVRPICLPSSNYNVPAGTTCTVSGWGRTSEGAGSLPDTLQEVNVPIVSNATCRRAYSGGGAIT  167 (229)
T ss_pred             cCCEEEEEECcccCCCCceeeccCCCcccccCCCCEEEEEeCCCCCCCCCcCCCEeeEEEEEEeCHHHhhhhhccccccC
Confidence            5899999998763 2234667766653   345688999998765421     01122222222110000   0     0


Q ss_pred             CeEEeEEE--EEeeccCCCCCCceecCCC--cEEEEEeeec
Q 017471          103 STELLGLQ--IDAAINSGNSGGPAFNDKG--KCVGIAFQSL  139 (371)
Q Consensus       103 ~~~~~~i~--~da~i~~G~SGGPlvn~~G--~VIGI~~~~~  139 (371)
                      ....-...  .....-+|+||||++...+  .++||.+...
T Consensus       168 ~~~~C~~~~~~~~~~c~gdsG~pl~~~~~~~~l~Gi~s~g~  208 (229)
T smart00020      168 DNMLCAGGLEGGKDACQGDSGGPLVCNDGRWVLVGIVSWGS  208 (229)
T ss_pred             CCcEeecCCCCCCcccCCCCCCeeEEECCCEEEEEEEEECC
Confidence            00000001  1234557999999998765  9999987643


No 54 
>KOG3550 consensus Receptor targeting protein Lin-7 [Extracellular structures]
Probab=95.52  E-value=0.028  Score=47.59  Aligned_cols=37  Identities=24%  Similarity=0.424  Sum_probs=33.0

Q ss_pred             CCceEEEEECCCCcccc--CCCCCCEEEEECCEEecCCC
Q 017471          197 QKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDG  233 (371)
Q Consensus       197 ~~gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~V~~~~  233 (371)
                      ..-++|+++.|++-|++  ||+.||.++++||..|....
T Consensus       114 nspiyisriipggvadrhgglkrgdqllsvngvsvege~  152 (207)
T KOG3550|consen  114 NSPIYISRIIPGGVADRHGGLKRGDQLLSVNGVSVEGEH  152 (207)
T ss_pred             CCceEEEeecCCccccccCcccccceeEeecceeecchh
Confidence            45699999999999998  89999999999999997543


No 55 
>KOG3532 consensus Predicted protein kinase [General function prediction only]
Probab=95.43  E-value=0.02  Score=59.15  Aligned_cols=50  Identities=16%  Similarity=0.458  Sum_probs=42.0

Q ss_pred             ccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCCCcc
Q 017471          174 LLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVP  236 (371)
Q Consensus       174 ~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~  236 (371)
                      -+|+.|..-             ...-|.|..|.+++||.+ .+++||++++|||.||++..+..
T Consensus       387 ~ig~vf~~~-------------~~~~v~v~tv~~ns~a~k~~~~~gdvlvai~~~pi~s~~q~~  437 (1051)
T KOG3532|consen  387 PIGLVFDKN-------------TNRAVKVCTVEDNSLADKAAFKPGDVLVAINNVPIRSERQAT  437 (1051)
T ss_pred             ceeEEEecC-------------CceEEEEEEecCCChhhHhcCCCcceEEEecCccchhHHHHH
Confidence            477777643             245688999999999999 99999999999999999987764


No 56 
>PF08192 Peptidase_S64:  Peptidase family S64;  InterPro: IPR012985 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This family of fungal proteins is involved in the processing of membrane bound transcription factor Stp1 [] and belongs to MEROPS petidase family S64 (clan PA). The processing causes the signalling domain of Stp1 to be passed to the nucleus where several permease genes are induced. The permeases are important for uptake of amino acids, and processing of tp1 only occurs in an amino acid-rich environment. This family is predicted to be distantly related to the trypsin family (MEROPS peptidase family S1) and to have a typical trypsin-like catalytic triad [].
Probab=95.33  E-value=0.17  Score=52.76  Aligned_cols=119  Identities=17%  Similarity=0.285  Sum_probs=73.0

Q ss_pred             cCCCCEEEEEEecCC-----CcCCc------cceecCCC--------CCCCCeEEEEEeCCCCCCceeeeeEEeeeeeee
Q 017471           38 ADICIYTMLTVEDDE-----FWEGV------LPVEFGEL--------PALQDAVTVVGYPIGGDTISVTSGVVSRIEILS   98 (371)
Q Consensus        38 d~~~DlAlLkv~~~~-----~~~~l------~~~~l~~s--------~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~   98 (371)
                      ....|+||++++...     +.+++      |.+.+.+.        ...|.+|+=+|+..+     .|.|+|.++....
T Consensus       540 ~~LsD~AIIkV~~~~~~~N~LGddi~f~~~dP~l~f~NlyV~~~~~~~~~G~~VfK~GrTTg-----yT~G~lNg~klvy  614 (695)
T PF08192_consen  540 KRLSDWAIIKVNKERKCQNYLGDDIQFNEPDPTLMFQNLYVREVVSNLVPGMEVFKVGRTTG-----YTTGILNGIKLVY  614 (695)
T ss_pred             ccccceEEEEeCCCceecCCCCccccccCCCccccccccchhhhhhccCCCCeEEEecccCC-----ccceEecceEEEE
Confidence            345799999999653     22222      33344331        123678999988755     5667777765432


Q ss_pred             ccCCCeE-EeEEEEE----eeccCCCCCCceecCCCc------EEEEEeeecccCCccceeccccCcchhHhHhhh
Q 017471           99 YVHGSTE-LLGLQID----AAINSGNSGGPAFNDKGK------CVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDY  163 (371)
Q Consensus        99 ~~~~~~~-~~~i~~d----a~i~~G~SGGPlvn~~G~------VIGI~~~~~~~~~~~~~~~aiP~~~i~~~l~~L  163 (371)
                      ...+... .+++...    +=...|+||.-+++.-++      |+||.++.-  .....++++.|+..|.+-+++.
T Consensus       615 w~dG~i~s~efvV~s~~~~~Fa~~GDSGS~VLtk~~d~~~gLgvvGMlhsyd--ge~kqfglftPi~~il~rl~~v  688 (695)
T PF08192_consen  615 WADGKIQSSEFVVSSDNNPAFASGGDSGSWVLTKLEDNNKGLGVVGMLHSYD--GEQKQFGLFTPINEILDRLEEV  688 (695)
T ss_pred             ecCCCeEEEEEEEecCCCccccCCCCcccEEEecccccccCceeeEEeeecC--CccceeeccCcHHHHHHHHHHh
Confidence            2233222 2344333    223479999999987555      999998743  2455788888887766555443


No 57 
>PF00949 Peptidase_S7:  Peptidase S7, Flavivirus NS3 serine protease ;  InterPro: IPR001850 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This signature identifies serine peptidases belong to MEROPS peptidase family S7 (flavivirin family, clan PA(S)). The protein fold of the peptidase domain for members of this family resembles that of chymotrypsin, the type example for clan PA.  Flaviviruses produce a polyprotein from the ssRNA genome. The N terminus of the NS3 protein (approx. 180 aa) is required for the processing of the polyprotein. NS3 also has conserved homology with NTP-binding proteins and DEAD family of RNA helicase [, , ].; GO: 0003723 RNA binding, 0003724 RNA helicase activity, 0005524 ATP binding; PDB: 2IJO_B 3E90_D 2GGV_B 2FP7_B 2WV9_A 3U1I_B 3U1J_B 2WZQ_A 2WHX_A 3L6P_A ....
Probab=95.14  E-value=0.025  Score=47.42  Aligned_cols=33  Identities=33%  Similarity=0.511  Sum_probs=23.6

Q ss_pred             EEEEEeeccCCCCCCceecCCCcEEEEEeeecc
Q 017471          108 GLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLK  140 (371)
Q Consensus       108 ~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~  140 (371)
                      +...+....+|.||.|++|.+|+++|+......
T Consensus        87 ~~~~~~d~~~GsSGSpi~n~~g~ivGlYg~g~~  119 (132)
T PF00949_consen   87 IGAIDLDFPKGSSGSPIFNQNGEIVGLYGNGVE  119 (132)
T ss_dssp             EEEE---S-TTGTT-EEEETTSCEEEEEEEEEE
T ss_pred             EEeeecccCCCCCCCceEcCCCcEEEEEcccee
Confidence            345566688999999999999999999877653


No 58 
>PF00947 Pico_P2A:  Picornavirus core protein 2A;  InterPro: IPR000081 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad [].  This domain defines cysteine peptidases belong to MEROPS peptidase family C3 (picornain, clan PA(C)), subfamilies 3CA and 3CB. The protein fold of this peptidase domain for members of this family resembles that of the serine peptidase, chymotrypsin [], the type example for clan PA. Picornaviral proteins are expressed as a single polyprotein which is cleaved by the viral 3C cysteine protease []. The poliovirus polyprotein is selectively cleaved between the Gln-|-Gly bond. In other picornavirus reactions Glu may be substituted for Gln, and Ser or Thr for Gly. ; GO: 0008233 peptidase activity, 0006508 proteolysis, 0016032 viral reproduction; PDB: 2HRV_B 1Z8R_A.
Probab=94.86  E-value=0.16  Score=41.96  Aligned_cols=97  Identities=13%  Similarity=-0.004  Sum_probs=53.2

Q ss_pred             EEEecCCCCEEEEEEecCCCcCCccceecCCCCCCCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEe
Q 017471           34 TLVTADICIYTMLTVEDDEFWEGVLPVEFGELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDA  113 (371)
Q Consensus        34 vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da  113 (371)
                      .+..+...||+|.+.+...    ...++.-+   -..-|+--=.-.-.-..+++.-..-.++..++.+.+...+.+.-..
T Consensus        13 ~v~~~~~rDL~V~~t~a~G----~D~I~~C~---Ct~GvYyCks~~k~yPV~~~~~~~~~i~~s~YYP~h~Q~~~l~g~G   85 (127)
T PF00947_consen   13 TVWEDYTRDLLVDRTTAHG----CDTIPRCD---CTTGVYYCKSKNKYYPVTVTGPTWYWIEESEYYPKHYQYNLLIGEG   85 (127)
T ss_dssp             CCEEECCCTEEEEEECCEE----E--BB-------SEEEEEETTTTCEEEEEEEEECEEEE-SBTTB-SEEEECEEEEE-
T ss_pred             ceehhhCCCEEEEecCCCC----CCcccCcc---CCCCEEEeeECCeEeeEEEeccceEEECCccCchhheecCceeecc
Confidence            4677889999999998763    22222221   0011111000000000122222233455556667777777888899


Q ss_pred             eccCCCCCCceecCCCcEEEEEeee
Q 017471          114 AINSGNSGGPAFNDKGKCVGIAFQS  138 (371)
Q Consensus       114 ~i~~G~SGGPlvn~~G~VIGI~~~~  138 (371)
                      +..||..||+|+-.. -||||.++.
T Consensus        86 p~~PGdCGg~L~C~H-GViGi~Tag  109 (127)
T PF00947_consen   86 PAEPGDCGGILRCKH-GVIGIVTAG  109 (127)
T ss_dssp             SSSTT-TCSEEEETT-CEEEEEEEE
T ss_pred             cCCCCCCCceeEeCC-CeEEEEEeC
Confidence            999999999999766 599999884


No 59 
>KOG3834 consensus Golgi reassembly stacking protein GRASP65, contains PDZ domain [Intracellular trafficking, secretion, and vesicular transport]
Probab=94.82  E-value=0.05  Score=53.59  Aligned_cols=132  Identities=17%  Similarity=0.246  Sum_probs=80.1

Q ss_pred             CCceEEEEECCCCcccc-CCCC-CCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCE--EEEEE
Q 017471          197 QKGVRIRRVDPTAPESE-VLKP-SDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK--ILNFN  272 (371)
Q Consensus       197 ~~gv~V~~V~~~spA~~-GL~~-GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~--~~~~~  272 (371)
                      ..|.-|-+|.++|+|.+ ||.+ -|-|++|||..++...|..          +.+.+....+ |+++++--..  .+.++
T Consensus        14 teg~hvlkVqedSpa~~aglepffdFIvSI~g~rL~~dnd~L----------k~llk~~sek-Vkltv~n~kt~~~R~v~   82 (462)
T KOG3834|consen   14 TEGYHVLKVQEDSPAHKAGLEPFFDFIVSINGIRLNKDNDTL----------KALLKANSEK-VKLTVYNSKTQEVRIVE   82 (462)
T ss_pred             ceeEEEEEeecCChHHhcCcchhhhhhheeCcccccCchHHH----------HHHHHhcccc-eEEEEEecccceeEEEE
Confidence            46888999999999999 9888 5899999999999887753          3444433333 9999886432  33333


Q ss_pred             EEeccccccCCCCCCCCCCCceeeccEEEechHHHHHHcCcccc-cCceEEEEEecCChHhHHhhh-hhhhHHHHHHHHH
Q 017471          273 ITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERI-MNMKLRSSFWTSSCIQCHNCQ-MSSLLWCLRCLWL  350 (371)
Q Consensus       273 v~l~~~~~~~~~~~~~~~~~~~~~~Gl~~~~l~~~~~~~~~~~~-~~~~~v~~~~~~Sp~~~~~~~-~~~~~~~~~~~~~  350 (371)
                      |+........             ++|+.+.-.       +...+ ..-.=+-+|.++|||++.... ..|.+=   =-|-
T Consensus        83 I~ps~~wggq-------------llGvsvrFc-------sf~~A~~~vwHvl~V~p~SPaalAgl~~~~DYiv---G~~~  139 (462)
T KOG3834|consen   83 IVPSNNWGGQ-------------LLGVSVRFC-------SFDGAVESVWHVLSVEPNSPAALAGLRPYTDYIV---GIWD  139 (462)
T ss_pred             eccccccccc-------------ccceEEEec-------cCccchhheeeeeecCCCCHHHhcccccccceEe---cchh
Confidence            3332211100             356654211       11111 111225578999999998776 555432   1245


Q ss_pred             HHhhhhhhHHHH
Q 017471          351 ILILDMRRLLTL  362 (371)
Q Consensus       351 ~~~~~~~~~~~~  362 (371)
                      ++-.+.+.+++|
T Consensus       140 ~~~~~~eDl~~l  151 (462)
T KOG3834|consen  140 AVMHEEEDLFTL  151 (462)
T ss_pred             hhccchHHHHHH
Confidence            566666666554


No 60 
>KOG3209 consensus WW domain-containing protein [General function prediction only]
Probab=94.70  E-value=0.044  Score=57.04  Aligned_cols=58  Identities=14%  Similarity=0.161  Sum_probs=46.9

Q ss_pred             CceEEEEECCCCcccc-C-CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECC
Q 017471          198 KGVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDS  266 (371)
Q Consensus       198 ~gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g  266 (371)
                      -+++|-++.+++||.+ | ++.||.|++|||++.++...-           +++.-.+.|....+.++|.|
T Consensus       923 M~LfVLRlAeDGPA~rdGrm~VGDqi~eINGesTkgmtH~-----------rAIelIk~gg~~vll~Lr~g  982 (984)
T KOG3209|consen  923 MDLFVLRLAEDGPAIRDGRMRVGDQITEINGESTKGMTHD-----------RAIELIKQGGRRVLLLLRRG  982 (984)
T ss_pred             cceEEEEeccCCCccccCceeecceEEEecCcccCCCcHH-----------HHHHHHHhCCeEEEEEeccC
Confidence            4699999999999999 6 999999999999999887653           35555556666667777765


No 61 
>PF00548 Peptidase_C3:  3C cysteine protease (picornain 3C);  InterPro: IPR000199 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad [].  This signature defines cysteine peptidases belong to MEROPS peptidase family C3 (picornain, clan PA(C)), subfamilies C3A and C3B. The protein fold of this peptidase domain for members of this family resembles that of the serine peptidase, chymotrypsin [], the type example for clan PA. Picornaviral proteins are expressed as a single polyprotein which is cleaved by the viral C3 cysteine protease. The poliovirus polyprotein is selectively cleaved between the Gln-|-Gly bond. In other picornavirus reactions Glu may be substituted for Gln, and Ser or Thr for Gly. ; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis; PDB: 3SJO_E 2H6M_A 1QA7_C 1HAV_B 2HAL_A 2H9H_A 3QZQ_B 3QZR_A 3R0F_B 3SJ9_A ....
Probab=94.28  E-value=0.67  Score=40.80  Aligned_cols=103  Identities=17%  Similarity=0.281  Sum_probs=63.3

Q ss_pred             EEEEecCC---CCEEEEEEecCCCcCC-ccceecCCCCCCCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeE
Q 017471           33 ATLVTADI---CIYTMLTVEDDEFWEG-VLPVEFGELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLG  108 (371)
Q Consensus        33 ~vv~~d~~---~DlAlLkv~~~~~~~~-l~~~~l~~s~~lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~  108 (371)
                      .+.-+|.+   .|+++++++..+-+.+ .+.++ ...+...+.+.++-++ .........+.|+..+.. ...+......
T Consensus        61 ~~~lv~~~~~~~Dl~~v~l~~~~kfrDIrk~~~-~~~~~~~~~~l~v~~~-~~~~~~~~v~~v~~~~~i-~~~g~~~~~~  137 (172)
T PF00548_consen   61 SVVLVDRDGVDTDLTLVKLPRNPKFRDIRKFFP-ESIPEYPECVLLVNST-KFPRMIVEVGFVTNFGFI-NLSGTTTPRS  137 (172)
T ss_dssp             EEEEEETTSSEEEEEEEEEESSS-B--GGGGSB-SSGGTEEEEEEEEESS-SSTCEEEEEEEEEEEEEE-EETTEEEEEE
T ss_pred             eEEEecCCCcceeEEEEEccCCcccCchhhhhc-cccccCCCcEEEEECC-CCccEEEEEEEEeecCcc-ccCCCEeeEE
Confidence            43445554   6999999976431222 22333 2222444666666544 333335566666666543 3334444567


Q ss_pred             EEEEeeccCCCCCCceec---CCCcEEEEEeee
Q 017471          109 LQIDAAINSGNSGGPAFN---DKGKCVGIAFQS  138 (371)
Q Consensus       109 i~~da~i~~G~SGGPlvn---~~G~VIGI~~~~  138 (371)
                      +..+++-.+|.-||||+.   ..++++||..+.
T Consensus       138 ~~Y~~~t~~G~CG~~l~~~~~~~~~i~GiHvaG  170 (172)
T PF00548_consen  138 LKYKAPTKPGMCGSPLVSRIGGQGKIIGIHVAG  170 (172)
T ss_dssp             EEEESEEETTGTTEEEEESCGGTTEEEEEEEEE
T ss_pred             EEEccCCCCCccCCeEEEeeccCccEEEEEecc
Confidence            899999999999999985   358999998874


No 62 
>KOG3605 consensus Beta amyloid precursor-binding protein [General function prediction only]
Probab=94.03  E-value=0.078  Score=54.71  Aligned_cols=102  Identities=21%  Similarity=0.262  Sum_probs=68.3

Q ss_pred             CCCCCce-----ecCCCcEEEEEeeecccCCccceeccccCcchhHhHhhhhhCCeeeccc---ccceeeeEcCCHHHHH
Q 017471          118 GNSGGPA-----FNDKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDYEKNGAYTGFP---LLGVEWQKMENPDLRV  189 (371)
Q Consensus       118 G~SGGPl-----vn~~G~VIGI~~~~~~~~~~~~~~~aiP~~~i~~~l~~L~~~g~~~g~~---~lGi~~~~~~~~~~~~  189 (371)
                      =++|||.     +|.--+++.||--.+         ..+|.+.+..++..+++.-.+. +.   .--+.-..+..|+.+-
T Consensus       680 mm~~GpAarsgkLnIGDQiiaING~SL---------VGLPLstcQs~Ik~~KnQT~Vk-ltiV~cpPV~~V~I~RPd~ky  749 (829)
T KOG3605|consen  680 MMHGGPAARSGKLNIGDQIMSINGTSL---------VGLPLSTCQSIIKGLKNQTAVK-LNIVSCPPVTTVLIRRPDLRY  749 (829)
T ss_pred             cccCChhhhcCCccccceeEeecCcee---------ccccHHHHHHHHhcccccceEE-EEEecCCCceEEEeecccchh
Confidence            3556664     444456666653222         4489999999888886544332 11   1112222233778888


Q ss_pred             hccCCCCCCceEEEEECCCCcccc-CCCCCCEEEEECCEEecC
Q 017471          190 AMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAN  231 (371)
Q Consensus       190 ~~gl~~~~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~  231 (371)
                      .+|+. -..|+ |-+...|+-|++ |++.|-.|++|||+.|--
T Consensus       750 QLGFS-VQNGi-ICSLlRGGIAERGGVRVGHRIIEINgQSVVA  790 (829)
T KOG3605|consen  750 QLGFS-VQNGI-ICSLLRGGIAERGGVRVGHRIIEINGQSVVA  790 (829)
T ss_pred             hccce-eeCcE-eehhhcccchhccCceeeeeEEEECCceEEe
Confidence            89997 34566 677889999999 999999999999999753


No 63 
>KOG3542 consensus cAMP-regulated guanine nucleotide exchange factor [Signal transduction mechanisms]
Probab=93.91  E-value=0.05  Score=56.29  Aligned_cols=37  Identities=27%  Similarity=0.339  Sum_probs=33.4

Q ss_pred             CCceEEEEECCCCcccc-CCCCCCEEEEECCEEecCCC
Q 017471          197 QKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDG  233 (371)
Q Consensus       197 ~~gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~  233 (371)
                      ..|++|.+|.|++.|+. ||+.||.|++|||+...+..
T Consensus       561 GfgifV~~V~pgskAa~~GlKRgDqilEVNgQnfenis  598 (1283)
T KOG3542|consen  561 GFGIFVAEVFPGSKAAREGLKRGDQILEVNGQNFENIS  598 (1283)
T ss_pred             cceeEEeeecCCchHHHhhhhhhhhhhhccccchhhhh
Confidence            45899999999999999 99999999999999887654


No 64 
>KOG3651 consensus Protein kinase C, alpha binding protein [Signal transduction mechanisms]
Probab=93.78  E-value=0.089  Score=49.57  Aligned_cols=39  Identities=21%  Similarity=0.318  Sum_probs=34.7

Q ss_pred             CceEEEEECCCCcccc-C-CCCCCEEEEECCEEecCCCCcc
Q 017471          198 KGVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVP  236 (371)
Q Consensus       198 ~gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~  236 (371)
                      .-++|.+|..++||++ | ++.||.|++|||..|+....+.
T Consensus        30 PClYiVQvFD~tPAa~dG~i~~GDEi~avNg~svKGktKve   70 (429)
T KOG3651|consen   30 PCLYIVQVFDKTPAAKDGRIRCGDEIVAVNGISVKGKTKVE   70 (429)
T ss_pred             CeEEEEEeccCCchhccCccccCCeeEEecceeecCccHHH
Confidence            3589999999999999 6 9999999999999999876654


No 65 
>KOG1892 consensus Actin filament-binding protein Afadin [Cytoskeleton]
Probab=93.76  E-value=0.082  Score=56.73  Aligned_cols=64  Identities=17%  Similarity=0.212  Sum_probs=49.3

Q ss_pred             CCCCCceEEEEECCCCcccc-C-CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCE
Q 017471          194 KADQKGVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK  267 (371)
Q Consensus       194 ~~~~~gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~  267 (371)
                      -++.-|++|..|.+|++|+. | |+.||.+++|||+..-...+-.          .+-...+.|..|.+.|...|.
T Consensus       956 Gq~klGIYvKsVV~GgaAd~DGRL~aGDQLLsVdG~SLiGisQEr----------AA~lmtrtg~vV~leVaKqgA 1021 (1629)
T KOG1892|consen  956 GQRKLGIYVKSVVEGGAADHDGRLEAGDQLLSVDGHSLIGISQER----------AARLMTRTGNVVHLEVAKQGA 1021 (1629)
T ss_pred             CccccceEEEEeccCCccccccccccCceeeeecCcccccccHHH----------HHHHHhccCCeEEEehhhhhh
Confidence            33456999999999999998 6 9999999999999987766532          122224578899999877553


No 66 
>COG0750 Predicted membrane-associated Zn-dependent proteases 1 [Cell envelope biogenesis, outer membrane]
Probab=93.71  E-value=0.11  Score=51.29  Aligned_cols=57  Identities=28%  Similarity=0.424  Sum_probs=45.3

Q ss_pred             EEEECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCE---EEEEEEE-CCEEE
Q 017471          202 IRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDS---AAVKVLR-DSKIL  269 (371)
Q Consensus       202 V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~---v~l~v~R-~g~~~  269 (371)
                      +.++..+|+|+. |+++||.|+++|++++.+++++.          +.. ....+..   +.+.+.| +++..
T Consensus       133 ~~~v~~~s~a~~a~l~~Gd~iv~~~~~~i~~~~~~~----------~~~-~~~~~~~~~~~~i~~~~~~~~~~  194 (375)
T COG0750         133 VGEVAPKSAAALAGLRPGDRIVAVDGEKVASWDDVR----------RLL-VAAAGDVFNLLTILVIRLDGEAH  194 (375)
T ss_pred             eeecCCCCHHHHcCCCCCCEEEeECCEEccCHHHHH----------HHH-HhccCCcccceEEEEEeccceee
Confidence            337899999999 99999999999999999998864          223 3334555   8899999 77763


No 67 
>KOG2921 consensus Intramembrane metalloprotease (sterol-regulatory element-binding protein (SREBP) protease) [Posttranslational modification, protein turnover, chaperones]
Probab=93.55  E-value=0.051  Score=53.01  Aligned_cols=40  Identities=25%  Similarity=0.341  Sum_probs=36.3

Q ss_pred             CCCceEEEEECCCCcccc--CCCCCCEEEEECCEEecCCCCc
Q 017471          196 DQKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTV  235 (371)
Q Consensus       196 ~~~gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~V~~~~~l  235 (371)
                      ...|+.|++|...||+..  ||++||+|+++||-+|.+.+|.
T Consensus       218 ~g~gV~Vtev~~~Spl~gprGL~vgdvitsldgcpV~~v~dW  259 (484)
T KOG2921|consen  218 HGEGVTVTEVPSVSPLFGPRGLSVGDVITSLDGCPVHKVSDW  259 (484)
T ss_pred             cCceEEEEeccccCCCcCcccCCccceEEecCCcccCCHHHH
Confidence            357999999999999987  9999999999999999988774


No 68 
>KOG3580 consensus Tight junction proteins [Signal transduction mechanisms]
Probab=93.44  E-value=0.092  Score=53.72  Aligned_cols=61  Identities=28%  Similarity=0.464  Sum_probs=47.6

Q ss_pred             CCceEEEEECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc-cCCCCEEEEEEEECCEE
Q 017471          197 QKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ-KYTGDSAAVKVLRDSKI  268 (371)
Q Consensus       197 ~~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~-~~~g~~v~l~v~R~g~~  268 (371)
                      .+.++|++|.||+||+.-||.||.|+.|||.+..+....         |  ++.+ .+.|+...++|.|..+.
T Consensus        39 etSiViSDVlpGGPAeG~LQenDrvvMVNGvsMenv~ha---------F--AvQqLrksgK~A~ItvkRprkv  100 (1027)
T KOG3580|consen   39 ETSIVISDVLPGGPAEGLLQENDRVVMVNGVSMENVLHA---------F--AVQQLRKSGKVAAITVKRPRKV  100 (1027)
T ss_pred             ceeEEEeeccCCCCcccccccCCeEEEEcCcchhhhHHH---------H--HHHHHHhhccceeEEeccccee
Confidence            456999999999999878999999999999999876542         1  2222 34677888999886543


No 69 
>KOG3552 consensus FERM domain protein FRM-8 [General function prediction only]
Probab=92.50  E-value=0.12  Score=55.25  Aligned_cols=57  Identities=26%  Similarity=0.325  Sum_probs=42.3

Q ss_pred             CceEEEEECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEE
Q 017471          198 KGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLR  264 (371)
Q Consensus       198 ~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R  264 (371)
                      .-|+|..|.+|+|+...|+|||.|++|||.+|+...--     +.|   .++.  ...+.|.|+|.+
T Consensus        75 rPviVr~VT~GGps~GKL~PGDQIl~vN~Epv~dapre-----rvI---dlvR--ace~sv~ltV~q  131 (1298)
T KOG3552|consen   75 RPVIVRFVTEGGPSIGKLQPGDQILAVNGEPVKDAPRE-----RVI---DLVR--ACESSVNLTVCQ  131 (1298)
T ss_pred             CceEEEEecCCCCccccccCCCeEEEecCcccccccHH-----HHH---HHHH--HHhhhcceEEec
Confidence            45889999999999878999999999999999765321     111   1222  245678888887


No 70 
>PF00944 Peptidase_S3:  Alphavirus core protein ;  InterPro: IPR000930 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. Togavirin, also known as Sindbis virus core endopeptidase, is a serine protease resident at the N terminus of the p130 polyprotein of togaviruses []. The endopeptidase signature identifies the peptidase as belonging to the MEROPS peptidase family S3 (togavirin family, clan PA(S)). The polyprotein also includes structural proteins for the nucleocapsid core and for the glycoprotein spikes []. Togavirin is only active while part of the polyprotein, cleavage at a Trp-Ser bond resulting in total lack of activity []. Mutagenesis studies have identified the location of the His-Asp-Ser catalytic triad, and X-ray studies have revealed the protein fold to be similar to that of chymotrypsin [, ].; GO: 0004252 serine-type endopeptidase activity, 0006508 proteolysis, 0016020 membrane; PDB: 2YEW_D 1EP5_A 3J0C_F 1EP6_C 1WYK_D 1DYL_A 1VCQ_B 1VCP_B 1LD4_D 1KXA_A ....
Probab=91.44  E-value=0.19  Score=41.88  Aligned_cols=26  Identities=31%  Similarity=0.631  Sum_probs=22.5

Q ss_pred             eccCCCCCCceecCCCcEEEEEeeec
Q 017471          114 AINSGNSGGPAFNDKGKCVGIAFQSL  139 (371)
Q Consensus       114 ~i~~G~SGGPlvn~~G~VIGI~~~~~  139 (371)
                      .-.+|+||-|++|..|+||||+....
T Consensus       102 ~g~~GDSGRpi~DNsGrVVaIVLGG~  127 (158)
T PF00944_consen  102 VGKPGDSGRPIFDNSGRVVAIVLGGA  127 (158)
T ss_dssp             S-STTSTTEEEESTTSBEEEEEEEEE
T ss_pred             CCCCCCCCCccCcCCCCEEEEEecCC
Confidence            45699999999999999999998765


No 71 
>KOG3606 consensus Cell polarity protein PAR6 [Signal transduction mechanisms]
Probab=91.09  E-value=0.25  Score=45.95  Aligned_cols=81  Identities=19%  Similarity=0.273  Sum_probs=52.0

Q ss_pred             eeccccCcchhHhHhh--hhhCCeeecccccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc-C-CCCCCEEE
Q 017471          147 IGYVIPTPVIMHFIQD--YEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-V-LKPSDIIL  222 (371)
Q Consensus       147 ~~~aiP~~~i~~~l~~--L~~~g~~~g~~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~-G-L~~GDvIl  222 (371)
                      .+-.|.++.+-+.-.+  |.++|.-.   .||+...+- +.--...-|+. ...|+.|++..||+-|+. | |..+|.++
T Consensus       146 VSsIIDVDivPEtHRRVRL~khG~ek---PLGFYIRDG-~SVRVtp~Gle-kvpGIFISRlVpGGLAeSTGLLaVnDEVl  220 (358)
T KOG3606|consen  146 VSSIIDVDIVPETHRRVRLHKHGSEK---PLGFYIRDG-TSVRVTPHGLE-KVPGIFISRLVPGGLAESTGLLAVNDEVL  220 (358)
T ss_pred             eceeeeecccchhhhheehhhcCCCC---CceEEEecC-ceEEecccccc-ccCceEEEeecCCccccccceeeecceeE
Confidence            3344445444443333  33444432   477776544 11111224554 467999999999999999 7 78899999


Q ss_pred             EECCEEecCC
Q 017471          223 SFDGIDIAND  232 (371)
Q Consensus       223 ~vnG~~V~~~  232 (371)
                      +|||.+|...
T Consensus       221 EVNGIEVaGK  230 (358)
T KOG3606|consen  221 EVNGIEVAGK  230 (358)
T ss_pred             EEcCEEeccc
Confidence            9999999764


No 72 
>KOG3571 consensus Dishevelled 3 and related proteins [General function prediction only]
Probab=90.91  E-value=0.27  Score=49.51  Aligned_cols=37  Identities=16%  Similarity=0.331  Sum_probs=32.0

Q ss_pred             CCceEEEEECCCCcccc-C-CCCCCEEEEECCEEecCCC
Q 017471          197 QKGVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDG  233 (371)
Q Consensus       197 ~~gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~  233 (371)
                      ..|++|.+|.++++-+. | +.+||.||.||.....+..
T Consensus       276 DggIYVgsImkgGAVA~DGRIe~GDMiLQVNevsFENmS  314 (626)
T KOG3571|consen  276 DGGIYVGSIMKGGAVALDGRIEPGDMILQVNEVSFENMS  314 (626)
T ss_pred             CCceEEeeeccCceeeccCccCccceEEEeeecchhhcC
Confidence            46899999999998877 6 9999999999998776654


No 73 
>KOG3551 consensus Syntrophins (type beta) [Extracellular structures]
Probab=88.86  E-value=0.33  Score=47.39  Aligned_cols=72  Identities=21%  Similarity=0.264  Sum_probs=49.4

Q ss_pred             cccceeeeEcCCHHHHHhccCCCCCCceEEEEECCCCcccc--CCCCCCEEEEECCEEecCCCCccccccccchhhhhhh
Q 017471          173 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVS  250 (371)
Q Consensus       173 ~~lGi~~~~~~~~~~~~~~gl~~~~~gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~  250 (371)
                      +-|||+++--.           ++.--++|+.+.++-.|.+  -|..||.|++|||....+...-           ..+.
T Consensus        96 gGLGISIKGGr-----------eNkMPIlISKIFkGlAADQt~aL~~gDaIlSVNG~dL~~AtHd-----------eAVq  153 (506)
T KOG3551|consen   96 GGLGISIKGGR-----------ENKMPILISKIFKGLAADQTGALFLGDAILSVNGEDLRDATHD-----------EAVQ  153 (506)
T ss_pred             CcceEEeecCc-----------ccCCceehhHhccccccccccceeeccEEEEecchhhhhcchH-----------HHHH
Confidence            55788776431           1223589999999999998  5999999999999998766432           1222


Q ss_pred             c-cCCCCEEEEEE--EECC
Q 017471          251 Q-KYTGDSAAVKV--LRDS  266 (371)
Q Consensus       251 ~-~~~g~~v~l~v--~R~g  266 (371)
                      . ++.|+.|.+.|  .|+-
T Consensus       154 aLKraGkeV~levKy~REv  172 (506)
T KOG3551|consen  154 ALKRAGKEVLLEVKYMREV  172 (506)
T ss_pred             HHHhhCceeeeeeeeehhc
Confidence            2 45687766544  4543


No 74 
>KOG3549 consensus Syntrophins (type gamma) [Extracellular structures]
Probab=88.32  E-value=0.5  Score=45.54  Aligned_cols=55  Identities=22%  Similarity=0.266  Sum_probs=42.2

Q ss_pred             ceEEEEECCCCcccc-C-CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEE
Q 017471          199 GVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVL  263 (371)
Q Consensus       199 gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~  263 (371)
                      -++|+.+.++-.|+. | |-.||-|+.|||..|+.-..-     +.+   +.+  .+.|+.|+++|.
T Consensus        81 PvviSkI~kdQaAd~tG~LFvGDAilqvNGi~v~~c~He-----evV---~iL--RNAGdeVtlTV~  137 (505)
T KOG3549|consen   81 PVVISKIYKDQAADITGQLFVGDAILQVNGIYVTACPHE-----EVV---NIL--RNAGDEVTLTVK  137 (505)
T ss_pred             cEEeehhhhhhhhhhcCceEeeeeeEEeccEEeecCChH-----HHH---HHH--HhcCCEEEEEeH
Confidence            488999999999988 6 889999999999999865431     111   122  457999998885


No 75 
>PF05579 Peptidase_S32:  Equine arteritis virus serine endopeptidase S32;  InterPro: IPR008760 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to MEROPS peptidase family S32 (clan PA(S)). The type example is equine arteritis virus serine endopeptidase (equine arteritis virus), which is involved in processing of nidovirus polyproteins [].; GO: 0004252 serine-type endopeptidase activity, 0016032 viral reproduction, 0019082 viral protein processing; PDB: 3FAN_A 3FAO_A 1MBM_A.
Probab=86.42  E-value=0.51  Score=44.02  Aligned_cols=23  Identities=30%  Similarity=0.680  Sum_probs=18.6

Q ss_pred             cCCCCCCceecCCCcEEEEEeee
Q 017471          116 NSGNSGGPAFNDKGKCVGIAFQS  138 (371)
Q Consensus       116 ~~G~SGGPlvn~~G~VIGI~~~~  138 (371)
                      ++|+||+|++..+|.+||+.+.+
T Consensus       206 ~~GDSGSPVVt~dg~liGVHTGS  228 (297)
T PF05579_consen  206 GPGDSGSPVVTEDGDLIGVHTGS  228 (297)
T ss_dssp             -GGCTT-EEEETTC-EEEEEEEE
T ss_pred             CCCCCCCccCcCCCCEEEEEecC
Confidence            58999999999999999999864


No 76 
>KOG0609 consensus Calcium/calmodulin-dependent serine protein kinase/membrane-associated guanylate kinase [Signal transduction mechanisms]
Probab=86.02  E-value=1.3  Score=45.10  Aligned_cols=57  Identities=25%  Similarity=0.336  Sum_probs=42.3

Q ss_pred             ceEEEEECCCCcccc-C-CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEEC
Q 017471          199 GVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD  265 (371)
Q Consensus       199 gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~  265 (371)
                      -++|+++..|+.+++ | |+.||.|+++||..+.+..--+        +..++.... | .+++.+.-.
T Consensus       147 ~~~vARI~~GG~~~r~glL~~GD~i~EvNGi~v~~~~~~e--------~q~~l~~~~-G-~itfkiiP~  205 (542)
T KOG0609|consen  147 KVVVARIMHGGMADRQGLLHVGDEILEVNGISVANKSPEE--------LQELLRNSR-G-SITFKIIPS  205 (542)
T ss_pred             ccEEeeeccCCcchhccceeeccchheecCeecccCCHHH--------HHHHHHhCC-C-cEEEEEccc
Confidence            589999999999998 6 9999999999999998763211        224444443 5 577777544


No 77 
>PF02907 Peptidase_S29:  Hepatitis C virus NS3 protease;  InterPro: IPR004109 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This signature identifies the Hepatitis C virus NS3 protein as a serine protease which belongs to MEROPS peptidase family S29 (hepacivirin family, clan PA(S)), which has a trypsin-like fold. The non-structural (NS) protein NS3 is one of the NS proteins involved in replication of the HCV genome. The NS2 proteinase (IPR002518 from INTERPRO), a zinc-dependent enzyme, performs a single proteolytic cut to release the N terminus of NS3. The action of NS3 proteinase (NS3P), which resides in the N-terminal one-third of the NS3 protein, then yields all remaining non-structural proteins. The C-terminal two-thirds of the NS3 protein contain a helicase. The functional relationship between the proteinase and helicase domains is unknown. NS3 has a structural zinc-binding site and requires cofactor NS4. It has been suggested that the NS3 serine protease of hepatitus C is involved in cell transformation and that the ability to transform requires an active enzyme [].; GO: 0008236 serine-type peptidase activity, 0006508 proteolysis, 0019087 transformation of host cell by virus; PDB: 2QV1_B 3LOX_C 2OBQ_C 2OC1_C 2OC0_A 3LON_A 3KNX_A 2O8M_A 2OBO_A 2OC8_A ....
Probab=84.26  E-value=0.83  Score=38.22  Aligned_cols=25  Identities=32%  Similarity=0.547  Sum_probs=19.6

Q ss_pred             cCCCCCCceecCCCcEEEEEeeecc
Q 017471          116 NSGNSGGPAFNDKGKCVGIAFQSLK  140 (371)
Q Consensus       116 ~~G~SGGPlvn~~G~VIGI~~~~~~  140 (371)
                      -.|.||||++-.+|.+|||-.+...
T Consensus       106 lkGSSGgPiLC~~GH~vG~f~aa~~  130 (148)
T PF02907_consen  106 LKGSSGGPILCPSGHAVGMFRAAVC  130 (148)
T ss_dssp             HTT-TT-EEEETTSEEEEEEEEEEE
T ss_pred             EecCCCCcccCCCCCEEEEEEEEEE
Confidence            3799999999999999999766553


No 78 
>cd00987 PDZ_serine_protease PDZ domain of tryspin-like serine proteases, such as DegP/HtrA, which are oligomeric proteins involved in heat-shock response, chaperone function, and apoptosis. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=83.66  E-value=1.3  Score=33.59  Aligned_cols=47  Identities=13%  Similarity=-0.014  Sum_probs=37.2

Q ss_pred             eccEEEech-HHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhHH
Q 017471          296 IAGFVFSRC-LYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLLW  343 (371)
Q Consensus       296 ~~Gl~~~~l-~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~~  343 (371)
                      |+|+.++++ +..+..+++ ....|+.+..+.++||++......+|+|.
T Consensus         2 ~~G~~~~~~~~~~~~~~~~-~~~~g~~V~~v~~~s~a~~~gl~~GD~I~   49 (90)
T cd00987           2 WLGVTVQDLTPDLAEELGL-KDTKGVLVASVDPGSPAAKAGLKPGDVIL   49 (90)
T ss_pred             ccceEEeECCHHHHHHcCC-CCCCEEEEEEECCCCHHHHcCCCcCCEEE
Confidence            689999988 666655554 34578999999999999987777788763


No 79 
>KOG3605 consensus Beta amyloid precursor-binding protein [General function prediction only]
Probab=80.46  E-value=1.7  Score=45.37  Aligned_cols=61  Identities=10%  Similarity=0.149  Sum_probs=39.4

Q ss_pred             EEECCCCcccc-C-CCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEE
Q 017471          203 RRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNF  271 (371)
Q Consensus       203 ~~V~~~spA~~-G-L~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~  271 (371)
                      .....++||++ | |-.||.|++|||..+-..---.        -..++...+.-..|+++|.+=--..++
T Consensus       678 Anmm~~GpAarsgkLnIGDQiiaING~SLVGLPLst--------cQs~Ik~~KnQT~VkltiV~cpPV~~V  740 (829)
T KOG3605|consen  678 ANMMHGGPAARSGKLNIGDQIMSINGTSLVGLPLST--------CQSIIKGLKNQTAVKLNIVSCPPVTTV  740 (829)
T ss_pred             HhcccCChhhhcCCccccceeEeecCceeccccHHH--------HHHHHhcccccceEEEEEecCCCceEE
Confidence            36677999999 5 9999999999998875432110        012333344445688888775444333


No 80 
>PF01732 DUF31:  Putative peptidase (DUF31);  InterPro: IPR022382  This domain has no known function. It is found in various hypothetical proteins and putative lipoproteins from mycoplasmas. 
Probab=76.63  E-value=1.8  Score=42.83  Aligned_cols=26  Identities=31%  Similarity=0.543  Sum_probs=22.3

Q ss_pred             EeeccCCCCCCceecCCCcEEEEEee
Q 017471          112 DAAINSGNSGGPAFNDKGKCVGIAFQ  137 (371)
Q Consensus       112 da~i~~G~SGGPlvn~~G~VIGI~~~  137 (371)
                      +..+..|.||+.|+|.+|++|||.++
T Consensus       349 ~~~l~gGaSGS~V~n~~~~lvGIy~g  374 (374)
T PF01732_consen  349 NYSLGGGASGSMVINQNNELVGIYFG  374 (374)
T ss_pred             ccCCCCCCCcCeEECCCCCEEEEeCC
Confidence            34667899999999999999999753


No 81 
>KOG0606 consensus Microtubule-associated serine/threonine kinase and related proteins [Signal transduction mechanisms; General function prediction only]
Probab=76.50  E-value=2.9  Score=46.30  Aligned_cols=34  Identities=21%  Similarity=0.258  Sum_probs=30.2

Q ss_pred             eEEEEECCCCcccc-CCCCCCEEEEECCEEecCCC
Q 017471          200 VRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDG  233 (371)
Q Consensus       200 v~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~  233 (371)
                      =+|..|.++|||.. |+++||.|+.+||+++....
T Consensus       660 h~v~sv~egsPA~~agls~~DlIthvnge~v~gl~  694 (1205)
T KOG0606|consen  660 HSVGSVEEGSPAFEAGLSAGDLITHVNGEPVHGLV  694 (1205)
T ss_pred             eeeeeecCCCCccccCCCccceeEeccCcccchhh
Confidence            45788999999988 99999999999999997654


No 82 
>PF05580 Peptidase_S55:  SpoIVB peptidase S55;  InterPro: IPR008763 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to the MEROPS peptidase family S55 (SpoIVB peptidase family, clan PA(S)). The protein SpoIVB plays a key role in signalling in the final sigma-K checkpoint of Bacillus subtilis [, ].
Probab=76.05  E-value=2.4  Score=38.52  Aligned_cols=39  Identities=28%  Similarity=0.390  Sum_probs=29.5

Q ss_pred             eccCCCCCCceecCCCcEEEEEeeecccCCccceeccccCcc
Q 017471          114 AINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPV  155 (371)
Q Consensus       114 ~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~~~~~~~aiP~~~  155 (371)
                      -|-.|+||+|++- +|++||=++..+.  +.+..+|.+|++.
T Consensus       176 GIvqGMSGSPI~q-dGKLiGAVthvf~--~dp~~Gygi~ie~  214 (218)
T PF05580_consen  176 GIVQGMSGSPIIQ-DGKLIGAVTHVFV--NDPTKGYGIFIEW  214 (218)
T ss_pred             CEEecccCCCEEE-CCEEEEEEEEEEe--cCCCceeeecHHH
Confidence            4668999999985 8999998776553  4466777787643


No 83 
>PF03761 DUF316:  Domain of unknown function (DUF316) ;  InterPro: IPR005514 This is a family of uncharacterised proteins from Caenorhabditis elegans.
Probab=73.42  E-value=31  Score=32.30  Aligned_cols=89  Identities=19%  Similarity=0.281  Sum_probs=53.0

Q ss_pred             CCCEEEEEEecCCCcCCccceecCCCCC---CCCeEEEEEeCCCCCCceeeeeEEeeeeeeeccCCCeEEeEEEEEeecc
Q 017471           40 ICIYTMLTVEDDEFWEGVLPVEFGELPA---LQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAIN  116 (371)
Q Consensus        40 ~~DlAlLkv~~~~~~~~l~~~~l~~s~~---lgd~V~~iG~p~g~~~~s~t~G~Vs~~~~~~~~~~~~~~~~i~~da~i~  116 (371)
                      .++++||+++.+ +.....|+=+++++.   .++.+.+.|+....   ......+.-....   .   ....+..+....
T Consensus       160 ~~~~mIlEl~~~-~~~~~~~~Cl~~~~~~~~~~~~~~~yg~~~~~---~~~~~~~~i~~~~---~---~~~~~~~~~~~~  229 (282)
T PF03761_consen  160 PYSPMILELEED-FSKNVSPPCLADSSTNWEKGDEVDVYGFNSTG---KLKHRKLKITNCT---K---CAYSICTKQYSC  229 (282)
T ss_pred             ccceEEEEEccc-ccccCCCEEeCCCccccccCceEEEeecCCCC---eEEEEEEEEEEee---c---cceeEecccccC
Confidence            489999999987 333677777877653   35888899882221   1222222211110   0   112345555666


Q ss_pred             CCCCCCcee---cCCCcEEEEEeee
Q 017471          117 SGNSGGPAF---NDKGKCVGIAFQS  138 (371)
Q Consensus       117 ~G~SGGPlv---n~~G~VIGI~~~~  138 (371)
                      .|++|||++   |.+-.|||+.+..
T Consensus       230 ~~d~Gg~lv~~~~gr~tlIGv~~~~  254 (282)
T PF03761_consen  230 KGDRGGPLVKNINGRWTLIGVGASG  254 (282)
T ss_pred             CCCccCeEEEEECCCEEEEEEEccC
Confidence            899999997   4445688886543


No 84 
>COG1625 Fe-S oxidoreductase, related to NifB/MoaA family [Energy production and conversion]
Probab=72.31  E-value=2.8  Score=41.69  Aligned_cols=35  Identities=14%  Similarity=0.194  Sum_probs=30.2

Q ss_pred             EEEEECCCCcccc-CCCCCCEEEEEC-CEEecCCCCc
Q 017471          201 RIRRVDPTAPESE-VLKPSDIILSFD-GIDIANDGTV  235 (371)
Q Consensus       201 ~V~~V~~~spA~~-GL~~GDvIl~vn-G~~V~~~~~l  235 (371)
                      .+..+.+++.++. |+.+||.+.+|| |.+.++-.+.
T Consensus         4 ~i~~v~~~~~~d~~Gfe~~~~l~~Vn~~~~~~~c~~~   40 (414)
T COG1625           4 KISKVGGISGADCDGFEEGDYLLKVNPGFGCKDCIPY   40 (414)
T ss_pred             ceeeccCCCcccccCccccceeeecCCCCCCCcCCCc
Confidence            4678889999999 999999999999 8888776654


No 85 
>KOG3834 consensus Golgi reassembly stacking protein GRASP65, contains PDZ domain [Intracellular trafficking, secretion, and vesicular transport]
Probab=68.29  E-value=3.9  Score=40.71  Aligned_cols=65  Identities=14%  Similarity=0.190  Sum_probs=45.9

Q ss_pred             EEEEECCCCcccc-CCC-CCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEEec
Q 017471          201 RIRRVDPTAPESE-VLK-PSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLA  276 (371)
Q Consensus       201 ~V~~V~~~spA~~-GL~-~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~l~  276 (371)
                      -|-+|.++|||+. ||+ -+|.|+.+-.......+|+.           .+...+.++.+++.|+--.....-+|++.
T Consensus       112 Hvl~V~p~SPaalAgl~~~~DYivG~~~~~~~~~eDl~-----------~lIeshe~kpLklyVYN~D~d~~ReVti~  178 (462)
T KOG3834|consen  112 HVLSVEPNSPAALAGLRPYTDYIVGIWDAVMHEEEDLF-----------TLIESHEGKPLKLYVYNHDTDSCREVTIT  178 (462)
T ss_pred             eeeecCCCCHHHhcccccccceEecchhhhccchHHHH-----------HHHHhccCCCcceeEeecCCCccceEEee
Confidence            3668999999999 999 68999999545555566653           34445678899999987554443444443


No 86 
>KOG3938 consensus RGS-GAIP interacting protein GIPC, contains PDZ domain [Signal transduction mechanisms; Intracellular trafficking, secretion, and vesicular transport]
Probab=64.24  E-value=5.5  Score=37.25  Aligned_cols=67  Identities=13%  Similarity=0.280  Sum_probs=49.5

Q ss_pred             hccCCCCCCc---eEEEEECCCCcccc--CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEE
Q 017471          190 AMSMKADQKG---VRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLR  264 (371)
Q Consensus       190 ~~gl~~~~~g---v~V~~V~~~spA~~--GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R  264 (371)
                      .+|+.-...|   ..|..+.++|--.+  -++.||.|-+|||+.|-.+..++        ..+.+.....|++.++.+.-
T Consensus       138 alGlTITDNG~GyAFIKrIkegsvidri~~i~VGd~IEaiNge~ivG~RHYe--------VArmLKel~rge~ftlrLie  209 (334)
T KOG3938|consen  138 ALGLTITDNGAGYAFIKRIKEGSVIDRIEAICVGDHIEAINGESIVGKRHYE--------VARMLKELPRGETFTLRLIE  209 (334)
T ss_pred             ccceEEeeCCcceeeeEeecCCchhhhhhheeHHhHHHhhcCccccchhHHH--------HHHHHHhcccCCeeEEEeec
Confidence            4555433333   56889999998887  79999999999999998887654        34566666778877776554


No 87 
>PF03510 Peptidase_C24:  2C endopeptidase (C24) cysteine protease family;  InterPro: IPR000317 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad [].  The two signatures that defines this group of calivirus polyproteins identify a cysteine peptidase signature that belongs to MEROPS peptidase family C24 (clan PA(C)). Caliciviruses are positive-stranded ssRNA viruses that cause gastroenteritis. The calicivirus genome contains two open reading frames, ORF1 and ORF2. ORF2 encodes a structural protein []; while ORF1 encodes a non-structural polypeptide, which has RNA helicase, cysteine protease and RNA polymerase activity. The regions of the polyprotein in which these activities lie are similar to proteins produced by the picornaviruses. Two different families of caliciviruses can be distinguished on the basis of sequence similarity, namely those classified as small round structured viruses (SRSVs) and those classed as non-SRSVs. Calicivirus proteases from the non-SRSV group, which are members of the PA protease clan, constitute family C24 of the cysteine proteases (proteases from SRSVs belong to the C37 family). As mentioned above, the protease activity resides within a polyprotein. The enzyme cleaves the polyprotein at sites N-terminal to itself, liberating the polyprotein helicase.; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis
Probab=63.96  E-value=38  Score=27.24  Aligned_cols=52  Identities=8%  Similarity=-0.049  Sum_probs=35.5

Q ss_pred             eeeeccEEEEE-EEecCCCeE----eEEEEEecCCCCEEEEEEecCCCcCCccceecCCCCCC
Q 017471           11 LNSRNEALILS-TWLLCSPSA----PSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELPAL   68 (371)
Q Consensus        11 ~~~~gsg~vi~-~~~~~~~~~----~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~~l   68 (371)
                      .+..|.|..++ -|+......    +-+++.  ..-|+|+++.+..    .+|..++++++.+
T Consensus         3 avHIGnG~~vt~tHva~~~~~v~g~~f~~~~--~~ge~~~v~~~~~----~~p~~~ig~g~Pv   59 (105)
T PF03510_consen    3 AVHIGNGRYVTVTHVAKSSDSVDGQPFKIVK--TDGELCWVQSPLV----HLPAAQIGTGKPV   59 (105)
T ss_pred             eEEeCCCEEEEEEEEeccCceEcCcCcEEEE--eccCEEEEECCCC----CCCeeEeccCCCE
Confidence            35568888888 777776643    223333  4569999999887    4688888875543


No 88 
>PF12812 PDZ_1:  PDZ-like domain
Probab=63.57  E-value=15  Score=27.89  Aligned_cols=39  Identities=10%  Similarity=0.012  Sum_probs=29.2

Q ss_pred             CCceeeccEEEech-HHHHHHcCcccccCceEEEEEecCChHh
Q 017471          291 PSYYIIAGFVFSRC-LYLISVLSMERIMNMKLRSSFWTSSCIQ  332 (371)
Q Consensus       291 ~~~~~~~Gl~~~~l-~~~~~~~~~~~~~~~~~v~~~~~~Sp~~  332 (371)
                      .++..|+|.+|+++ .....+++++   .++++.+...+||+.
T Consensus         5 ~r~v~~~Ga~f~~Ls~q~aR~~~~~---~~gv~v~~~~g~~~~   44 (78)
T PF12812_consen    5 SRFVEVCGAVFHDLSYQQARQYGIP---VGGVYVAVSGGSLAF   44 (78)
T ss_pred             CEEEEEcCeecccCCHHHHHHhCCC---CCEEEEEecCCChhh
Confidence            35677999999998 6667788766   336666778888865


No 89 
>PF13180 PDZ_2:  PDZ domain; PDB: 2L97_A 1Y8T_A 2Z9I_A 1LCY_A 2PZD_B 2P3W_A 1VCW_C 1TE0_B 1SOZ_C 1SOT_C ....
Probab=60.33  E-value=6.5  Score=29.53  Aligned_cols=37  Identities=11%  Similarity=-0.081  Sum_probs=27.6

Q ss_pred             eccEEEechHHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhH
Q 017471          296 IAGFVFSRCLYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLL  342 (371)
Q Consensus       296 ~~Gl~~~~l~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~  342 (371)
                      ++|+.++..          ....+++|.++.++|||+-.-.+.+|+|
T Consensus         2 ~lGv~~~~~----------~~~~g~~V~~V~~~spA~~aGl~~GD~I   38 (82)
T PF13180_consen    2 GLGVTVQNL----------SDTGGVVVVSVIPGSPAAKAGLQPGDII   38 (82)
T ss_dssp             E-SEEEEEC----------SCSSSEEEEEESTTSHHHHTTS-TTEEE
T ss_pred             EECeEEEEc----------cCCCeEEEEEeCCCCcHHHCCCCCCcEE
Confidence            678877543          1156899999999999999988888875


No 90 
>cd01735 LSm12_N LSm12 belongs to a family of Sm-like proteins that associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet that associates with other Sm proteins to form hexameric and heptameric ring structures.   In addition to the N-terminal Sm-like domain, LSm12 has a novel methyltransferase domain.
Probab=55.13  E-value=38  Score=24.46  Aligned_cols=34  Identities=6%  Similarity=-0.088  Sum_probs=27.4

Q ss_pred             EEEEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471           17 ALILSTWLLCSPSAPSATLVTADICIYTMLTVED   50 (371)
Q Consensus        17 g~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~   50 (371)
                      |-.++...-.|.++.++|+++|+...+.+||-.+
T Consensus         6 Gs~V~~kTc~g~~ieGEV~afD~~tk~lIlk~~s   39 (61)
T cd01735           6 GSQVSCRTCFEQRLQGEVVAFDYPSKMLILKCPS   39 (61)
T ss_pred             ccEEEEEecCCceEEEEEEEecCCCcEEEEECcc
Confidence            3445555556999999999999999999999655


No 91 
>TIGR02038 protease_degS periplasmic serine pepetdase DegS. This family consists of the periplasmic serine protease DegS (HhoB), a shorter paralog of protease DO (HtrA, DegP) and DegQ (HhoA). It is found in E. coli and several other Proteobacteria of the gamma subdivision. It contains a trypsin domain and a single copy of PDZ domain (in contrast to DegP with two copies). A critical role of this DegS is to sense stress in the periplasm and partially degrade an inhibitor of sigma(E).
Probab=54.96  E-value=9.9  Score=37.32  Aligned_cols=48  Identities=4%  Similarity=-0.083  Sum_probs=39.8

Q ss_pred             eccEEEech-HHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhHHH
Q 017471          296 IAGFVFSRC-LYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLLWC  344 (371)
Q Consensus       296 ~~Gl~~~~l-~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~~~  344 (371)
                      |+|+.++++ +...+.++++. ..|++|..+.++||++-.-...+|+|..
T Consensus       256 ~lGv~~~~~~~~~~~~lgl~~-~~Gv~V~~V~~~spA~~aGL~~GDvI~~  304 (351)
T TIGR02038       256 YIGVSGEDINSVVAQGLGLPD-LRGIVITGVDPNGPAARAGILVRDVILK  304 (351)
T ss_pred             EeeeEEEECCHHHHHhcCCCc-cccceEeecCCCChHHHCCCCCCCEEEE
Confidence            899999888 77777888864 4799999999999999877777887753


No 92 
>cd01726 LSm6 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm6 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=51.69  E-value=29  Score=25.26  Aligned_cols=33  Identities=6%  Similarity=-0.150  Sum_probs=28.3

Q ss_pred             EEEEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471           17 ALILSTWLLCSPSAPSATLVTADICIYTMLTVE   49 (371)
Q Consensus        17 g~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~   49 (371)
                      |-.+.+.+.+|+++.+++.++|+..|+.+=...
T Consensus        10 ~~~V~V~Lk~g~~~~G~L~~~D~~mNlvL~~~~   42 (67)
T cd01726          10 GRPVVVKLNSGVDYRGILACLDGYMNIALEQTE   42 (67)
T ss_pred             CCeEEEEECCCCEEEEEEEEEccceeeEEeeEE
Confidence            446778999999999999999999999886654


No 93 
>PF02122 Peptidase_S39:  Peptidase S39;  InterPro: IPR000382 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. ORF2 of Potato leafroll virus (PLrV) encodes a polyprotein which is translated following a -1 frameshift. The polyprotein has a putative linear arrangement of membrane achor-VPg-peptidase-polmerase domains. The serine peptidase domain which is found in this group of sequences belongs to MEROPS peptidase family S39 (clan PA(S)). It is likely that the peptidase domain is involved in the cleavage of the polyprotein []. The nucleotide sequence for the RNA of PLrV has been determined [, ]. The sequence contains six large open reading frames (ORFs). The 5' coding region encodes two polypeptides of 28K and 70K, which overlap in different reading frames; it is suggested that the third ORF in the 5' block is translated by frameshift readthrough near the end of the 70K protein, yielding a 118K polypeptide []. Segments of the predicted amino acid sequences of these ORFs resemble those of known viral RNA polymerases, ATP-binding proteins and viral genome-linked proteins. The nucleotide sequence of the genomic RNA of Beet western yellows virus (BWYV) has been determined []. The sequence contains six long ORFs. A cluster of three of these ORFs, including the coat protein cistron, display extensive amino acid sequence similarity to corresponding ORFs of a second luteovirus: Barley yellow dwarf virus [].; GO: 0004252 serine-type endopeptidase activity, 0022415 viral reproductive process, 0016021 integral to membrane; PDB: 1ZYO_A.
Probab=49.98  E-value=21  Score=32.35  Aligned_cols=46  Identities=26%  Similarity=0.337  Sum_probs=18.1

Q ss_pred             EEEEEeeccCCCCCCceecCCCcEEEEEeeecccCCccceeccccCc
Q 017471          108 GLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTP  154 (371)
Q Consensus       108 ~i~~da~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~~~~~~~aiP~~  154 (371)
                      +..+-+.-.+|.||-|.++.+ +++|+.....+....++.++.-|+.
T Consensus       137 ~~~vls~T~~G~SGtp~y~g~-~vvGvH~G~~~~~~~~n~n~~spip  182 (203)
T PF02122_consen  137 FASVLSNTSPGWSGTPYYSGK-NVVGVHTGSPSGSNRENNNRMSPIP  182 (203)
T ss_dssp             EEEE-----TT-TT-EEE-SS--EEEEEEEE----------------
T ss_pred             CCceEcCCCCCCCCCCeEECC-CceEeecCccccccccccccccccc
Confidence            456667778999999999998 9999988753333445666655554


No 94 
>TIGR03000 plancto_dom_1 Planctomycetes uncharacterized domain TIGR03000. Domains described by this model are found, so far, only in the Planctomycetes (Pirellula sp. strain 1 and Gemmata obscuriglobus), in up to six proteins per genome, and may be duplicated within a protein. The function is unknown.
Probab=49.83  E-value=54  Score=24.69  Aligned_cols=50  Identities=28%  Similarity=0.408  Sum_probs=33.8

Q ss_pred             CCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCC----EEEEEEEECCEEEEEEEEe
Q 017471          217 PSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGD----SAAVKVLRDSKILNFNITL  275 (371)
Q Consensus       217 ~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~----~v~l~v~R~g~~~~~~v~l  275 (371)
                      |-|-.+.+||++.++.+..+-         ..-....+|.    ++..++.|||+..+.+-++
T Consensus        10 PadAkl~v~G~~t~~~G~~R~---------F~T~~L~~G~~y~Y~v~a~~~~dG~~~t~~~~V   63 (75)
T TIGR03000        10 PADAKLKVDGKETNGTGTVRT---------FTTPPLEAGKEYEYTVTAEYDRDGRILTRTRTV   63 (75)
T ss_pred             CCCCEEEECCeEcccCccEEE---------EECCCCCCCCEEEEEEEEEEecCCcEEEEEEEE
Confidence            468899999999999887640         1112234565    4677888999876655444


No 95 
>cd00600 Sm_like The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=48.21  E-value=42  Score=23.59  Aligned_cols=33  Identities=9%  Similarity=-0.044  Sum_probs=28.2

Q ss_pred             EEEEEecCCCeEeEEEEEecCCCCEEEEEEecC
Q 017471           19 ILSTWLLCSPSAPSATLVTADICIYTMLTVEDD   51 (371)
Q Consensus        19 vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~~   51 (371)
                      .+.+.+.+++.+.|++.++|+..|+.+-.....
T Consensus         8 ~V~V~l~~g~~~~G~L~~~D~~~Ni~L~~~~~~   40 (63)
T cd00600           8 TVRVELKDGRVLEGVLVAFDKYMNLVLDDVEET   40 (63)
T ss_pred             EEEEEECCCcEEEEEEEEECCCCCEEECCEEEE
Confidence            456888999999999999999999988776543


No 96 
>cd01722 Sm_F The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit F is capable of forming both homo- and hetero-heptamer ring structures.  To form the hetero-heptamer, Sm subunit F initially binds subunits E and G to form a trimer which then assembles onto snRNA along with the D3/B and D1/D2 heterodimers.
Probab=47.66  E-value=33  Score=25.02  Aligned_cols=33  Identities=6%  Similarity=-0.092  Sum_probs=28.0

Q ss_pred             EEEEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471           17 ALILSTWLLCSPSAPSATLVTADICIYTMLTVE   49 (371)
Q Consensus        17 g~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~   49 (371)
                      |-.+.+.+.+|+.+.+++.++|...|+.+=...
T Consensus        11 g~~V~V~Lk~g~~~~G~L~~~D~~mNi~L~~~~   43 (68)
T cd01722          11 GKPVIVKLKWGMEYKGTLVSVDSYMNLQLANTE   43 (68)
T ss_pred             CCEEEEEECCCcEEEEEEEEECCCEEEEEeeEE
Confidence            345678999999999999999999999886554


No 97 
>PRK00737 small nuclear ribonucleoprotein; Provisional
Probab=47.20  E-value=43  Score=24.75  Aligned_cols=34  Identities=6%  Similarity=-0.217  Sum_probs=28.8

Q ss_pred             EEEEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471           17 ALILSTWLLCSPSAPSATLVTADICIYTMLTVED   50 (371)
Q Consensus        17 g~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~   50 (371)
                      +-.+.+.+.+|+++.+++.++|+..|+.+=....
T Consensus        14 ~k~V~V~lk~g~~~~G~L~~~D~~mNlvL~d~~e   47 (72)
T PRK00737         14 NSPVLVRLKGGREFRGELQGYDIHMNLVLDNAEE   47 (72)
T ss_pred             CCEEEEEECCCCEEEEEEEEEcccceeEEeeEEE
Confidence            3356688999999999999999999998877654


No 98 
>cd01717 Sm_B The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit B heterodimerizes with subunit D3 and three such heterodimers form a hexameric ring structure with alternating B and D3 subunits.  The D3 - B heterodimer also assembles into a heptameric ring containing D1, D2, E, F, and G subunits.  Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=47.04  E-value=38  Score=25.50  Aligned_cols=31  Identities=13%  Similarity=-0.037  Sum_probs=26.2

Q ss_pred             EEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471           20 LSTWLLCSPSAPSATLVTADICIYTMLTVED   50 (371)
Q Consensus        20 i~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~   50 (371)
                      +.+++-+|+.+.|++.++|+..|+.|=....
T Consensus        13 V~V~l~dgR~~~G~L~~~D~~~NlVL~~~~E   43 (79)
T cd01717          13 LRVTLQDGRQFVGQFLAFDKHMNLVLSDCEE   43 (79)
T ss_pred             EEEEECCCcEEEEEEEEEcCccCEEcCCEEE
Confidence            4578899999999999999999998765543


No 99 
>cd01730 LSm3 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm3 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=46.29  E-value=35  Score=25.93  Aligned_cols=29  Identities=7%  Similarity=-0.183  Sum_probs=25.1

Q ss_pred             EEEEecCCCeEeEEEEEecCCCCEEEEEE
Q 017471           20 LSTWLLCSPSAPSATLVTADICIYTMLTV   48 (371)
Q Consensus        20 i~~~~~~~~~~~A~vv~~d~~~DlAlLkv   48 (371)
                      +.+.+.+|+.+.|++.++|.+.|+.+=..
T Consensus        14 V~V~l~~gr~~~G~L~~fD~~mNlvL~d~   42 (82)
T cd01730          14 VYVKLRGDRELRGRLHAYDQHLNMILGDV   42 (82)
T ss_pred             EEEEECCCCEEEEEEEEEccceEEeccce
Confidence            45788999999999999999999987544


No 100
>KOG3553 consensus Tax interaction protein TIP1 [Cell wall/membrane/envelope biogenesis]
Probab=45.01  E-value=32  Score=27.40  Aligned_cols=27  Identities=4%  Similarity=-0.121  Sum_probs=22.5

Q ss_pred             ccCceEEEEEecCChHhHHhhhhhhhH
Q 017471          316 IMNMKLRSSFWTSSCIQCHNCQMSSLL  342 (371)
Q Consensus       316 ~~~~~~v~~~~~~Sp~~~~~~~~~~~~  342 (371)
                      ...|+||+.+.+||||+.--+...|-|
T Consensus        57 tD~GiYvT~V~eGsPA~~AGLrihDKI   83 (124)
T KOG3553|consen   57 TDKGIYVTRVSEGSPAEIAGLRIHDKI   83 (124)
T ss_pred             CCccEEEEEeccCChhhhhcceecceE
Confidence            357999999999999999877776643


No 101
>TIGR02860 spore_IV_B stage IV sporulation protein B. SpoIVB, the stage IV sporulation protein B of endospore-forming bacteria such as Bacillus subtilis, is a serine proteinase, expressed in the spore (rather than mother cell) compartment, that participates in a proteolytic activation cascade for Sigma-K. It appears to be universal among endospore-forming bacteria and occurs nowhere else.
Probab=44.68  E-value=20  Score=35.90  Aligned_cols=37  Identities=27%  Similarity=0.447  Sum_probs=26.7

Q ss_pred             eccCCCCCCceecCCCcEEEEEeeecccCCccceeccccC
Q 017471          114 AINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPT  153 (371)
Q Consensus       114 ~i~~G~SGGPlvn~~G~VIGI~~~~~~~~~~~~~~~aiP~  153 (371)
                      -|-.|+||+|++- +|++||=.+.-+-+  .+..+|+|-+
T Consensus       356 GivqGMSGSPi~q-~gkliGAvtHVfvn--dpt~GYGi~i  392 (402)
T TIGR02860       356 GIVQGMSGSPIIQ-NGKVIGAVTHVFVN--DPTSGYGVYI  392 (402)
T ss_pred             CEEecccCCCEEE-CCEEEEEEEEEEec--CCCcceeehH
Confidence            4567999999994 79999987766643  3455566644


No 102
>PRK10898 serine endoprotease; Provisional
Probab=43.41  E-value=22  Score=34.88  Aligned_cols=48  Identities=6%  Similarity=-0.088  Sum_probs=37.3

Q ss_pred             eccEEEech-HHHHHHcCcccccCceEEEEEecCChHhHHhhhhhhhHHH
Q 017471          296 IAGFVFSRC-LYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLLWC  344 (371)
Q Consensus       296 ~~Gl~~~~l-~~~~~~~~~~~~~~~~~v~~~~~~Sp~~~~~~~~~~~~~~  344 (371)
                      |+|+..+++ +.....+++. ...|++|..+.++||++-.-...+|+|..
T Consensus       257 ~lGi~~~~~~~~~~~~~~~~-~~~Gv~V~~V~~~spA~~aGL~~GDvI~~  305 (353)
T PRK10898        257 YIGIGGREIAPLHAQGGGID-QLQGIVVNEVSPDGPAAKAGIQVNDLIIS  305 (353)
T ss_pred             ccceEEEECCHHHHHhcCCC-CCCeEEEEEECCCChHHHcCCCCCCEEEE
Confidence            799988877 5555555554 34899999999999999887888887653


No 103
>cd01732 LSm5 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm4 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=43.28  E-value=45  Score=25.06  Aligned_cols=31  Identities=10%  Similarity=-0.095  Sum_probs=26.2

Q ss_pred             EEEEEEecCCCeEeEEEEEecCCCCEEEEEE
Q 017471           18 LILSTWLLCSPSAPSATLVTADICIYTMLTV   48 (371)
Q Consensus        18 ~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv   48 (371)
                      --+.+++.+++.+.|++.++|...|+.+=..
T Consensus        14 ~~V~V~l~~gr~~~G~L~g~D~~mNlvL~da   44 (76)
T cd01732          14 SRIWIVMKSDKEFVGTLLGFDDYVNMVLEDV   44 (76)
T ss_pred             CEEEEEECCCeEEEEEEEEeccceEEEEccE
Confidence            3455788999999999999999999987554


No 104
>PF11874 DUF3394:  Domain of unknown function (DUF3394);  InterPro: IPR021814  This domain is functionally uncharacterised. This domain is found in bacteria. This presumed domain is about 190 amino acids in length. This domain is found associated with PF06808 from PFAM. 
Probab=43.26  E-value=19  Score=32.00  Aligned_cols=28  Identities=14%  Similarity=0.051  Sum_probs=25.1

Q ss_pred             CCceEEEEECCCCcccc-CCCCCCEEEEE
Q 017471          197 QKGVRIRRVDPTAPESE-VLKPSDIILSF  224 (371)
Q Consensus       197 ~~gv~V~~V~~~spA~~-GL~~GDvIl~v  224 (371)
                      ...+.|..|..+|||++ |+..|+.|+++
T Consensus       121 ~~~~~Vd~v~fgS~A~~~g~d~d~~I~~v  149 (183)
T PF11874_consen  121 GGKVIVDEVEFGSPAEKAGIDFDWEITEV  149 (183)
T ss_pred             CCEEEEEecCCCCHHHHcCCCCCcEEEEE
Confidence            35689999999999999 99999998877


No 105
>COG0298 HypC Hydrogenase maturation factor [Posttranslational modification, protein turnover, chaperones]
Probab=42.41  E-value=47  Score=25.32  Aligned_cols=47  Identities=23%  Similarity=0.220  Sum_probs=32.4

Q ss_pred             EeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCCCCCCeEEE-EEeC
Q 017471           30 APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELPALQDAVTV-VGYP   78 (371)
Q Consensus        30 ~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~~lgd~V~~-iG~p   78 (371)
                      +|++|+.+|...++|++.+-.-.  .+..---+++...+||+|++ +||.
T Consensus         5 iPgqI~~I~~~~~~A~Vd~gGvk--reV~l~Lv~~~v~~GdyVLVHvGfA   52 (82)
T COG0298           5 IPGQIVEIDDNNHLAIVDVGGVK--REVNLDLVGEEVKVGDYVLVHVGFA   52 (82)
T ss_pred             cccEEEEEeCCCceEEEEeccEe--EEEEeeeecCccccCCEEEEEeeEE
Confidence            68999999998889999987642  11222223336688999887 5554


No 106
>cd01729 LSm7 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm7 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=42.28  E-value=53  Score=24.93  Aligned_cols=30  Identities=0%  Similarity=-0.170  Sum_probs=25.5

Q ss_pred             EEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471           20 LSTWLLCSPSAPSATLVTADICIYTMLTVE   49 (371)
Q Consensus        20 i~~~~~~~~~~~A~vv~~d~~~DlAlLkv~   49 (371)
                      +.+.+.+|+++.+++.++|...|+.+=...
T Consensus        15 V~V~l~~gr~~~G~L~~~D~~mNlvL~~~~   44 (81)
T cd01729          15 IRVKFQGGREVTGILKGYDQLLNLVLDDTV   44 (81)
T ss_pred             EEEEECCCcEEEEEEEEEcCcccEEecCEE
Confidence            447788999999999999999999885543


No 107
>cd01720 Sm_D2 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit D2 heterodimerizes with subunit D1 and three such heterodimers form a hexameric ring structure with alternating D1 and D2 subunits. The D1 - D2 heterodimer also assembles into a heptameric ring containing D2, D3, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=42.15  E-value=50  Score=25.57  Aligned_cols=31  Identities=6%  Similarity=-0.090  Sum_probs=26.9

Q ss_pred             EEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471           19 ILSTWLLCSPSAPSATLVTADICIYTMLTVE   49 (371)
Q Consensus        19 vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~   49 (371)
                      .+.+++-+++.+.|++.++|.+.|+.+=...
T Consensus        16 ~V~V~lr~~r~~~G~L~~fD~hmNlvL~d~~   46 (87)
T cd01720          16 QVLINCRNNKKLLGRVKAFDRHCNMVLENVK   46 (87)
T ss_pred             EEEEEEcCCCEEEEEEEEecCccEEEEcceE
Confidence            4568899999999999999999999976654


No 108
>cd01731 archaeal_Sm1 The archaeal sm1 proteins: The Sm proteins are conserved in all three domains of life and are always associated with U-rich RNA sequences. They function to mediate RNA-RNA interactions and RNA biogenesis.  All Sm proteins contain a common sequence motif in two segments, Sm1 and Sm2, separated by a short variable linker. Eukaryotic Sm proteins form part of specific small nuclear ribonucleoproteins (snRNPs) that are involved in the processing of pre-mRNAs to mature mRNAs, and are a major component of the eukaryotic spliceosome. Most snRNPs consist of seven Sm proteins (B/B', D1, D2, D3, E, F and G) arranged in a ring on a uridine-rich sequence (Sm site), plus a small nuclear RNA (snRNA) (either U1, U2, U5 or U4/6). Since archaebacteria do not have any splicing apparatus, Sm proteins of archaebacteria may play a more general role. Archaeal Lsm proteins are likely to represent the ancestral Sm domain.
Probab=41.91  E-value=53  Score=23.85  Aligned_cols=33  Identities=6%  Similarity=-0.146  Sum_probs=28.7

Q ss_pred             EEEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471           18 LILSTWLLCSPSAPSATLVTADICIYTMLTVED   50 (371)
Q Consensus        18 ~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~   50 (371)
                      -.+.+.+.+|+.+.+++.++|+..|+.+-....
T Consensus        11 ~~V~V~l~~g~~~~G~L~~~D~~mNlvL~~~~e   43 (68)
T cd01731          11 KPVLVKLKGGKEVRGRLKSYDQHMNLVLEDAEE   43 (68)
T ss_pred             CEEEEEECCCCEEEEEEEEECCcceEEEeeEEE
Confidence            356688899999999999999999999887754


No 109
>KOG1738 consensus Membrane-associated guanylate kinase-interacting protein/connector enhancer of KSR-like [Nucleotide transport and metabolism]
Probab=41.36  E-value=21  Score=37.39  Aligned_cols=34  Identities=9%  Similarity=0.085  Sum_probs=30.4

Q ss_pred             eEEEEECCCCcccc--CCCCCCEEEEECCEEecCCC
Q 017471          200 VRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDG  233 (371)
Q Consensus       200 v~V~~V~~~spA~~--GL~~GDvIl~vnG~~V~~~~  233 (371)
                      .+|+++.++|||..  -|..||.|+.||++.|-.|.
T Consensus       227 h~~s~~~e~Spad~~~kI~dgdEv~qiN~qtvVgwq  262 (638)
T KOG1738|consen  227 HVTSKIFEQSPADYRQKILDGDEVLQINEQTVVGWQ  262 (638)
T ss_pred             eeccccccCChHHHhhcccCccceeeecccccccch
Confidence            56788999999988  69999999999999988775


No 110
>cd06168 LSm9 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm9 proteins have a single Sm-like domain structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=40.86  E-value=60  Score=24.34  Aligned_cols=31  Identities=10%  Similarity=0.071  Sum_probs=26.6

Q ss_pred             EEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471           19 ILSTWLLCSPSAPSATLVTADICIYTMLTVE   49 (371)
Q Consensus        19 vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~   49 (371)
                      .+.+++.||+.+.+++.++|+..|+.+=...
T Consensus        12 ~v~V~l~dgR~~~G~l~~~D~~~NivL~~~~   42 (75)
T cd06168          12 TMRIHMTDGRTLVGVFLCTDRDCNIILGSAQ   42 (75)
T ss_pred             eEEEEEcCCeEEEEEEEEEcCCCcEEecCcE
Confidence            4568999999999999999999999875554


No 111
>PF00571 CBS:  CBS domain CBS domain web page. Mutations in the CBS domain of Swiss:P35520 lead to homocystinuria.;  InterPro: IPR000644 CBS (cystathionine-beta-synthase) domains are small intracellular modules, mostly found in two or four copies within a protein, that occur in a variety of proteins in bacteria, archaea, and eukaryotes [, ]. Tandem pairs of CBS domains can act as binding domains for adenosine derivatives and may regulate the activity of attached enzymatic or other domains []. In some cases, CBS domains may act as sensors of cellular energy status by being activated by AMP and inhibited by ATP []. In chloride ion channels, the CBS domains have been implicated in intracellular targeting and trafficking, as well as in protein-protein interactions, but results vary with different channels: in the CLC-5 channel, the CBS domain was shown to be required for trafficking [], while in the CLC-1 channel, the CBS domain was shown to be critical for channel function, but not necessary for trafficking []. Recent experiments revealing that CBS domains can bind adenosine-containing ligands such ATP, AMP, or S-adenosylmethionine have led to the hypothesis that CBS domains function as sensors of intracellular metabolites [, ]. Crystallographic studies of CBS domains have shown that pairs of CBS sequences form a globular domain where each CBS unit adopts a beta-alpha-beta-beta-alpha pattern []. Crystal structure of the CBS domains of the AMP-activated protein kinase in complexes with AMP and ATP shows that the phosphate groups of AMP/ATP lie in a surface pocket at the interface of two CBS domains, which is lined with basic residues, many of which are associated with disease-causing mutations [].  In humans, mutations in conserved residues within CBS domains cause a variety of human hereditary diseases, including (with the gene mutated in parentheses): homocystinuria (cystathionine beta-synthase); Wolff-Parkinson-White syndrome (gamma 2 subunit of AMP-activated protein kinase); retinitis pigmentosa (IMP dehydrogenase-1); congenital myotonia, idiopathic generalized epilepsy, hypercalciuric nephrolithiasis, and classic Bartter syndrome (CLC chloride channel family members).; GO: 0005515 protein binding; PDB: 3JTF_A 3TE5_C 3TDH_C 3T4N_C 2QLV_C 3OI8_A 3LV9_A 2QH1_B 1PVM_B 3LQN_A ....
Probab=39.53  E-value=26  Score=23.68  Aligned_cols=20  Identities=40%  Similarity=0.567  Sum_probs=16.3

Q ss_pred             CCCCCceecCCCcEEEEEee
Q 017471          118 GNSGGPAFNDKGKCVGIAFQ  137 (371)
Q Consensus       118 G~SGGPlvn~~G~VIGI~~~  137 (371)
                      +.+.-|++|.+|+++|+.+.
T Consensus        29 ~~~~~~V~d~~~~~~G~is~   48 (57)
T PF00571_consen   29 GISRLPVVDEDGKLVGIISR   48 (57)
T ss_dssp             TSSEEEEESTTSBEEEEEEH
T ss_pred             CCcEEEEEecCCEEEEEEEH
Confidence            45667899999999999764


No 112
>COG0260 PepB Leucyl aminopeptidase [Amino acid transport and metabolism]
Probab=39.03  E-value=44  Score=34.35  Aligned_cols=58  Identities=16%  Similarity=0.251  Sum_probs=34.7

Q ss_pred             HhccCCCCCCceEEEEECCCCccccCCCCCCEEEEECCEEecCCCCccccccccchhhhhhhc
Q 017471          189 VAMSMKADQKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ  251 (371)
Q Consensus       189 ~~~gl~~~~~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~  251 (371)
                      ..++++.+  -+.|....++.|.....||||||++.||+.|.=...=   ..+|+.+.+.+..
T Consensus       291 a~l~l~vn--v~~vl~~~ENm~~g~A~rPGDVits~~GkTVEV~NTD---AEGRLVLADaLtY  348 (485)
T COG0260         291 AELKLPVN--VVGVLPAVENMPSGNAYRPGDVITSMNGKTVEVLNTD---AEGRLVLADALTY  348 (485)
T ss_pred             HHcCCCce--EEEEEeeeccCCCCCCCCCCCeEEecCCcEEEEcccC---ccHHHHHHHHHHH
Confidence            34556532  2334445566666556899999999999987422110   1266666665544


No 113
>smart00651 Sm snRNP Sm proteins. small nuclear ribonucleoprotein particles (snRNPs) involved in pre-mRNA splicing
Probab=37.92  E-value=70  Score=22.80  Aligned_cols=33  Identities=9%  Similarity=-0.125  Sum_probs=27.8

Q ss_pred             EEEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471           18 LILSTWLLCSPSAPSATLVTADICIYTMLTVED   50 (371)
Q Consensus        18 ~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~   50 (371)
                      -.+.+++.+|+.+.|++.++|+..|+-+=....
T Consensus         9 ~~V~V~l~~g~~~~G~L~~~D~~~NlvL~~~~e   41 (67)
T smart00651        9 KRVLVELKNGREYRGTLKGFDQFMNLVLEDVEE   41 (67)
T ss_pred             cEEEEEECCCcEEEEEEEEECccccEEEccEEE
Confidence            356688899999999999999999998866654


No 114
>PF01455 HupF_HypC:  HupF/HypC family;  InterPro: IPR001109 The large subunit of [NiFe]-hydrogenase, as well as other nickel metalloenzymes, is synthesised as a precursor devoid of the metalloenzyme active site. This precursor then undergoes a complex post-translational maturation process that requires a number of accessory proteins. The hydrogenase expression/formation proteins (HupF/HypC) form a family of small proteins that are hydrogenase precursor-specific chaperones required for this maturation process []. They are believed to keep the hydrogenase precursor in a conformation accessible for metal incorporation [, ].; PDB: 3D3R_A 2Z1C_C 2OT2_A.
Probab=36.19  E-value=59  Score=23.89  Aligned_cols=41  Identities=12%  Similarity=-0.006  Sum_probs=29.1

Q ss_pred             EeEEEEEecCCCCEEEEEEecCCCcCCcccee--cCCCCCCCCeEEEE
Q 017471           30 APSATLVTADICIYTMLTVEDDEFWEGVLPVE--FGELPALQDAVTVV   75 (371)
Q Consensus        30 ~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~--l~~s~~lgd~V~~i   75 (371)
                      +|++|+.++.....|++.+....     ..+.  +-+..++||+|++-
T Consensus         5 iP~~Vv~v~~~~~~A~v~~~G~~-----~~V~~~lv~~v~~Gd~VLVH   47 (68)
T PF01455_consen    5 IPGRVVEVDEDGGMAVVDFGGVR-----REVSLALVPDVKVGDYVLVH   47 (68)
T ss_dssp             EEEEEEEEETTTTEEEEEETTEE-----EEEEGTTCTSB-TT-EEEEE
T ss_pred             ccEEEEEEeCCCCEEEEEcCCcE-----EEEEEEEeCCCCCCCEEEEe
Confidence            68999999888999999888642     3333  33446789999886


No 115
>cd01728 LSm1 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm1 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=36.08  E-value=82  Score=23.53  Aligned_cols=31  Identities=3%  Similarity=-0.170  Sum_probs=26.2

Q ss_pred             EEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471           19 ILSTWLLCSPSAPSATLVTADICIYTMLTVE   49 (371)
Q Consensus        19 vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~   49 (371)
                      .+.+.+.+|+.+.|++.++|+..|+.+=...
T Consensus        14 ~v~V~l~~gr~~~G~L~~fD~~~NlvL~d~~   44 (74)
T cd01728          14 KVVVLLRDGRKLIGILRSFDQFANLVLQDTV   44 (74)
T ss_pred             EEEEEEcCCeEEEEEEEEECCcccEEecceE
Confidence            3447888999999999999999999886654


No 116
>cd01721 Sm_D3 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit D3 heterodimerizes with subunit B and three such heterodimers form a hexameric ring structure with alternating B and D3 subunits. The D3 - B heterodimer also assembles into a heptameric ring containing D1, D2, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=34.27  E-value=82  Score=23.10  Aligned_cols=35  Identities=14%  Similarity=0.016  Sum_probs=29.9

Q ss_pred             cEEEEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471           16 EALILSTWLLCSPSAPSATLVTADICIYTMLTVED   50 (371)
Q Consensus        16 sg~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~   50 (371)
                      .|-.+.+.+.+|..+.+++..+|...|+.+-....
T Consensus         9 ~g~~V~VeLk~g~~~~G~L~~~D~~MNl~L~~~~~   43 (70)
T cd01721           9 EGHIVTVELKTGEVYRGKLIEAEDNMNCQLKDVTV   43 (70)
T ss_pred             CCCEEEEEECCCcEEEEEEEEEcCCceeEEEEEEE
Confidence            44566788899999999999999999999988753


No 117
>cd01727 LSm8 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm8 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=32.24  E-value=92  Score=23.05  Aligned_cols=31  Identities=3%  Similarity=-0.166  Sum_probs=26.3

Q ss_pred             EEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471           20 LSTWLLCSPSAPSATLVTADICIYTMLTVED   50 (371)
Q Consensus        20 i~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~   50 (371)
                      +.+.+-+++++.+++.++|+..|+.+=....
T Consensus        12 V~V~l~dgr~~~G~L~~~D~~~NlvL~~~~E   42 (74)
T cd01727          12 VSVITVDGRVIVGTLKGFDQATNLILDDSHE   42 (74)
T ss_pred             EEEEECCCcEEEEEEEEEccccCEEccceEE
Confidence            4477889999999999999999998877543


No 118
>KOG3627 consensus Trypsin [Amino acid transport and metabolism]
Probab=30.83  E-value=33  Score=31.28  Aligned_cols=99  Identities=17%  Similarity=0.148  Sum_probs=49.8

Q ss_pred             CCEEEEEEecC-CCcCCccceecCCCC----CCC-CeEEEEEeCCCC----C-CceeeeeEEeeeeee----eccCC---
Q 017471           41 CIYTMLTVEDD-EFWEGVLPVEFGELP----ALQ-DAVTVVGYPIGG----D-TISVTSGVVSRIEIL----SYVHG---  102 (371)
Q Consensus        41 ~DlAlLkv~~~-~~~~~l~~~~l~~s~----~lg-d~V~~iG~p~g~----~-~~s~t~G~Vs~~~~~----~~~~~---  102 (371)
                      .|+|+|+++.+ .|-+.+.|+.+....    ..+ ..+.+.|+....    . ........+.-+...    .+...   
T Consensus       106 nDiall~l~~~v~~~~~i~piclp~~~~~~~~~~~~~~~v~GWG~~~~~~~~~~~~L~~~~v~i~~~~~C~~~~~~~~~~  185 (256)
T KOG3627|consen  106 NDIALLRLSEPVTFSSHIQPICLPSSADPYFPPGGTTCLVSGWGRTESGGGPLPDTLQEVDVPIISNSECRRAYGGLGTI  185 (256)
T ss_pred             CCEEEEEECCCcccCCcccccCCCCCcccCCCCCCCEEEEEeCCCcCCCCCCCCceeEEEEEeEcChhHhcccccCcccc
Confidence            79999999974 344456666664222    223 677777764321    1 111121122111110    11100   


Q ss_pred             CeEEeEE---EEEeeccCCCCCCceecCC---CcEEEEEeeec
Q 017471          103 STELLGL---QIDAAINSGNSGGPAFNDK---GKCVGIAFQSL  139 (371)
Q Consensus       103 ~~~~~~i---~~da~i~~G~SGGPlvn~~---G~VIGI~~~~~  139 (371)
                      ....-+.   .-....-.|+|||||+-.+   ..++||++...
T Consensus       186 ~~~~~Ca~~~~~~~~~C~GDSGGPLv~~~~~~~~~~GivS~G~  228 (256)
T KOG3627|consen  186 TDTMLCAGGPEGGKDACQGDSGGPLVCEDNGRWVLVGIVSWGS  228 (256)
T ss_pred             CCCEEeeCccCCCCccccCCCCCeEEEeeCCcEEEEEEEEecC
Confidence            0001011   1112234699999998765   69999987754


No 119
>cd01719 Sm_G The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet.  Sm subunit G binds subunits E and F to form a trimer which then assembles onto snRNA along with the D1/D2 and D3/B heterodimers forming a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=30.47  E-value=1.1e+02  Score=22.57  Aligned_cols=30  Identities=10%  Similarity=-0.165  Sum_probs=25.2

Q ss_pred             EEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471           20 LSTWLLCSPSAPSATLVTADICIYTMLTVE   49 (371)
Q Consensus        20 i~~~~~~~~~~~A~vv~~d~~~DlAlLkv~   49 (371)
                      +.+.+-+|+.+.+++.++|...|+.+=...
T Consensus        13 V~V~L~~g~~~~G~L~~~D~~mNlvL~~~~   42 (72)
T cd01719          13 LSLKLNGNRKVSGILRGFDPFMNLVLDDAV   42 (72)
T ss_pred             EEEEECCCeEEEEEEEEEcccccEEeccEE
Confidence            346788999999999999999999886554


No 120
>PRK05015 aminopeptidase B; Provisional
Probab=28.96  E-value=90  Score=31.52  Aligned_cols=39  Identities=13%  Similarity=0.075  Sum_probs=26.6

Q ss_pred             hccCCCCCCceEEEEECCCCccccCCCCCCEEEEECCEEec
Q 017471          190 AMSMKADQKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIA  230 (371)
Q Consensus       190 ~~gl~~~~~gv~V~~V~~~spA~~GL~~GDvIl~vnG~~V~  230 (371)
                      .++++.+  -..|.-..++.+.....+|||||.+-||+.|.
T Consensus       230 ~~~l~~n--V~~il~~aENmisg~A~kpgDVIt~~nGkTVE  268 (424)
T PRK05015        230 TRGLNKR--VKLFLCCAENLISGNAFKLGDIITYRNGKTVE  268 (424)
T ss_pred             hcCCCce--EEEEEEecccCCCCCCCCCCCEEEecCCcEEe
Confidence            3455522  22344455666665678999999999999874


No 121
>PF09465 LBR_tudor:  Lamin-B receptor of TUDOR domain;  InterPro: IPR019023  The Lamin-B receptor is a chromatin and lamin binding protein in the inner nuclear membrane. It is one of the integral inner nuclear envelope membrane proteins responsible for targeting nuclear membranes to chromatin, being a downstream effector of Ran, a small Ras-like nuclear GTPase which regulates NE assembly. Lamin-B receptor interacts with importin beta, a Ran-binding protein, thereby directly contributing to the fusion of membrane vesicles and the formation of the nuclear envelope []. ; PDB: 2L8D_A 2DIG_A.
Probab=28.33  E-value=2e+02  Score=20.35  Aligned_cols=37  Identities=11%  Similarity=-0.101  Sum_probs=28.7

Q ss_pred             ccEEEEEEEecCCCeE-eEEEEEecCCCCEEEEEEecC
Q 017471           15 NEALILSTWLLCSPSA-PSATLVTADICIYTMLTVEDD   51 (371)
Q Consensus        15 gsg~vi~~~~~~~~~~-~A~vv~~d~~~DlAlLkv~~~   51 (371)
                      ..|=++.++.+.+..+ +|+|...|...+++-+++++-
T Consensus         7 ~~Ge~V~~rWP~s~lYYe~kV~~~d~~~~~y~V~Y~DG   44 (55)
T PF09465_consen    7 AIGEVVMVRWPGSSLYYEGKVLSYDSKSDRYTVLYEDG   44 (55)
T ss_dssp             -SS-EEEEE-TTTS-EEEEEEEEEETTTTEEEEEETTS
T ss_pred             cCCCEEEEECCCCCcEEEEEEEEecccCceEEEEEcCC
Confidence            4566778888887775 999999999999999999875


No 122
>COG1958 LSM1 Small nuclear ribonucleoprotein (snRNP) homolog [Transcription]
Probab=27.12  E-value=1.2e+02  Score=22.69  Aligned_cols=32  Identities=9%  Similarity=-0.089  Sum_probs=27.3

Q ss_pred             EEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471           19 ILSTWLLCSPSAPSATLVTADICIYTMLTVED   50 (371)
Q Consensus        19 vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~   50 (371)
                      .+.+.+-+|+.+.|++.++|+..|+.+--+..
T Consensus        19 ~V~V~lk~g~~~~G~L~~~D~~mNlvL~d~~e   50 (79)
T COG1958          19 RVLVKLKNGREYRGTLVGFDQYMNLVLDDVEE   50 (79)
T ss_pred             EEEEEECCCCEEEEEEEEEccceeEEEeceEE
Confidence            44578899999999999999999998876655


No 123
>PF11874 DUF3394:  Domain of unknown function (DUF3394);  InterPro: IPR021814  This domain is functionally uncharacterised. This domain is found in bacteria. This presumed domain is about 190 amino acids in length. This domain is found associated with PF06808 from PFAM. 
Probab=26.97  E-value=1.2e+02  Score=26.99  Aligned_cols=71  Identities=11%  Similarity=-0.010  Sum_probs=40.8

Q ss_pred             hhhhccCCCCEEEEEEEE---CCEEEEEEEEeccccccCCCCCCCCCCCceeeccEEEechHHHHHHcCcccccCceEEE
Q 017471          247 YLVSQKYTGDSAAVKVLR---DSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERIMNMKLRS  323 (371)
Q Consensus       247 ~~~~~~~~g~~v~l~v~R---~g~~~~~~v~l~~~~~~~~~~~~~~~~~~~~~~Gl~~~~l~~~~~~~~~~~~~~~~~v~  323 (371)
                      ..+.+..+|+.+.+.|.+   .|+..+.++.+.-.+.......       ..-.|+.+.            +....+.+.
T Consensus        67 ~~~~~~~~g~~lrl~V~G~~~~G~~~~k~v~lpl~~~~~g~eR-------L~~~GL~l~------------~e~~~~~Vd  127 (183)
T PF11874_consen   67 QVAEQLPPGSSLRLRVEGPDFEGDPVTKTVLLPLGDGADGEER-------LEAAGLTLM------------EEGGKVIVD  127 (183)
T ss_pred             HHHhcCCCCCEEEEEEEccCCCCCceEEEEEEEcCCCCCHHHH-------HHhCCCEEE------------eeCCEEEEE
Confidence            456667789999999988   3555444444433222111000       001455542            223557899


Q ss_pred             EEecCChHhHHhh
Q 017471          324 SFWTSSCIQCHNC  336 (371)
Q Consensus       324 ~~~~~Sp~~~~~~  336 (371)
                      ++..|||++-...
T Consensus       128 ~v~fgS~A~~~g~  140 (183)
T PF11874_consen  128 EVEFGSPAEKAGI  140 (183)
T ss_pred             ecCCCCHHHHcCC
Confidence            9999999876543


No 124
>PF15483 DUF4641:  Domain of unknown function (DUF4641)
Probab=25.82  E-value=45  Score=33.33  Aligned_cols=21  Identities=43%  Similarity=0.605  Sum_probs=17.3

Q ss_pred             HHHHHHHHH-hhhhhhHHHHHH
Q 017471          344 CLRCLWLIL-ILDMRRLLTLRF  364 (371)
Q Consensus       344 ~~~~~~~~~-~~~~~~~~~~~~  364 (371)
                      |+||+|||- |.|.|+-|-.+-
T Consensus       417 CpRC~~LQkEIedLreQLaamq  438 (445)
T PF15483_consen  417 CPRCLVLQKEIEDLREQLAAMQ  438 (445)
T ss_pred             CcccHHHHHHHHHHHHHHHHHH
Confidence            999999985 889998776543


No 125
>COG2524 Predicted transcriptional regulator, contains C-terminal CBS domains [Transcription]
Probab=24.51  E-value=2.1e+02  Score=27.06  Aligned_cols=20  Identities=40%  Similarity=0.625  Sum_probs=17.3

Q ss_pred             CCCCCCceecCCCcEEEEEee
Q 017471          117 SGNSGGPAFNDKGKCVGIAFQ  137 (371)
Q Consensus       117 ~G~SGGPlvn~~G~VIGI~~~  137 (371)
                      .|-.|.|++|.+ +++||.+.
T Consensus       201 ~~i~GaPVvd~d-k~vGiit~  220 (294)
T COG2524         201 KGIRGAPVVDDD-KIVGIITL  220 (294)
T ss_pred             cCccCCceecCC-ceEEEEEH
Confidence            688999999977 99999754


No 126
>cd01723 LSm4 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm4 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=24.40  E-value=1.7e+02  Score=21.69  Aligned_cols=33  Identities=6%  Similarity=-0.158  Sum_probs=28.1

Q ss_pred             EEEEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471           17 ALILSTWLLCSPSAPSATLVTADICIYTMLTVE   49 (371)
Q Consensus        17 g~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~   49 (371)
                      |-.+.+.+.+|..+.+++..+|...|+.+-...
T Consensus        11 g~~V~VeLkng~~~~G~L~~~D~~mNi~L~~~~   43 (76)
T cd01723          11 NHPMLVELKNGETYNGHLVNCDNWMNIHLREVI   43 (76)
T ss_pred             CCEEEEEECCCCEEEEEEEEEcCCCceEEEeEE
Confidence            345667888999999999999999999987764


No 127
>PF12381 Peptidase_C3G:  Tungro spherical virus-type peptidase;  InterPro: IPR024387 This entry represents a rice tungro spherical waikavirus-type peptidase that belongs to MEROPS peptidase family C3G. It is a picornain 3C-type protease, and is responsible for the self-cleavage of the positive single-stranded polyproteins of a number of plant viral genomes. The location of the protease activity of the polyprotein is at the C-terminal end, adjacent and N-terminal to the putative RNA polymerase [, ].
Probab=22.99  E-value=96  Score=28.34  Aligned_cols=54  Identities=24%  Similarity=0.413  Sum_probs=37.1

Q ss_pred             eEEEEEeeccCCCCCCceecCC----CcEEEEEeeecccCCccceeccccCc--chhHhHhhh
Q 017471          107 LGLQIDAAINSGNSGGPAFNDK----GKCVGIAFQSLKHEDVENIGYVIPTP--VIMHFIQDY  163 (371)
Q Consensus       107 ~~i~~da~i~~G~SGGPlvn~~----G~VIGI~~~~~~~~~~~~~~~aiP~~--~i~~~l~~L  163 (371)
                      ..++..++-..|+-|+|++=.+    -+++||..+..   .....+||-++.  .+++.++.|
T Consensus       169 ~gleY~~~t~~GdCGs~i~~~~t~~~RKIvGiHVAG~---~~~~~gYAe~itQEDL~~A~~~l  228 (231)
T PF12381_consen  169 QGLEYQMPTMNGDCGSPIVRNNTQMVRKIVGIHVAGS---ANHAMGYAESITQEDLMRAINKL  228 (231)
T ss_pred             eeeeEECCCcCCCccceeeEcchhhhhhhheeeeccc---ccccceehhhhhHHHHHHHHHhh
Confidence            3577888889999999986322    58999998865   235667777663  344444443


No 128
>PF01423 LSM:  LSM domain ;  InterPro: IPR001163 This family is found in Lsm (like-Sm) proteins and in bacterial Lsm-related Hfq proteins. In each case, the domain adopts a core structure consisting of an open beta-barrel with an SH3-like topology. Lsm (like-Sm) proteins have diverse functions, and are thought to be important modulators of RNA biogenesis and function [, ]. The Sm proteins form part of specific small nuclear ribonucleoproteins (snRNPs) that are involved in the processing of pre-mRNAs to mature mRNAs, and are a major component of the eukaryotic spliceosome. Most snRNPs consist of seven Sm proteins (B/B', D1, D2, D3, E, F and G) arranged in a ring on a uridine-rich sequence (Sm site), plus a small nuclear RNA (snRNA) (either U1, U2, U5 or U4/6) []. All Sm proteins contain a common sequence motif in two segments, Sm1 and Sm2, separated by a short variable linker []. In other snRNPs, certain Sm proteins are replaced with different Lsm proteins, such as with U7 snRNPs, in which the D1 and D2 Sm proteins are replaced with U7-specific Lsm10 and Lsm11 proteins, where Lsm11 plays a role in histone U7-specific RNA processing []. Lsm proteins are also found in archaebacteria, which do not have any splicing apparatus suggesting a more general role for Lsm proteins. The pleiotropic translational regulator Hfq (host factor Q) is a bacterial Lsm-like protein, which modulates the structure of numerous RNA molecules by binding preferentially to A/U-rich sequences in RNA []. Hfq forms an Lsm-like fold, however, unlike the heptameric Sm proteins, Hfq forms a homo-hexameric ring.; PDB: 1D3B_K 2Y9D_D 2Y9A_D 2Y9C_R 3VRI_C 2Y9B_K 3QUI_D 3M4G_H 3INZ_E 1U1S_C ....
Probab=22.54  E-value=1.9e+02  Score=20.48  Aligned_cols=33  Identities=6%  Similarity=-0.040  Sum_probs=28.4

Q ss_pred             EEEEEecCCCeEeEEEEEecCCCCEEEEEEecC
Q 017471           19 ILSTWLLCSPSAPSATLVTADICIYTMLTVEDD   51 (371)
Q Consensus        19 vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~~   51 (371)
                      .+.+.+.+|+.+.|++.++|...|+.+-.....
T Consensus        10 ~V~V~l~~g~~~~G~L~~~D~~~Nl~L~~~~~~   42 (67)
T PF01423_consen   10 RVRVELKNGRTYRGTLVSFDQFMNLVLSDVTET   42 (67)
T ss_dssp             EEEEEETTSEEEEEEEEEEETTEEEEEEEEEEE
T ss_pred             EEEEEEeCCEEEEEEEEEeechheEEeeeEEEE
Confidence            355788899999999999999999998887754


No 129
>KOG2561 consensus Adaptor protein NUB1, contains UBA domain [Posttranslational modification, protein turnover, chaperones; Signal transduction mechanisms]
Probab=22.25  E-value=28  Score=35.15  Aligned_cols=23  Identities=13%  Similarity=0.144  Sum_probs=15.2

Q ss_pred             CChHhHHhhhhh------hhHHHHHHHHH
Q 017471          328 SSCIQCHNCQMS------SLLWCLRCLWL  350 (371)
Q Consensus       328 ~Sp~~~~~~~~~------~~~~~~~~~~~  350 (371)
                      -.+-++++.+=|      ||.||-|||.-
T Consensus       195 ~Cd~klLe~VDNyallnLDIVWCYfrLkn  223 (568)
T KOG2561|consen  195 LCDSKLLELVDNYALLNLDIVWCYFRLKN  223 (568)
T ss_pred             hhhHHHHHhhcchhhhhcchhheehhhcc
Confidence            344556655433      99999998853


No 130
>PF10049 DUF2283:  Protein of unknown function (DUF2283);  InterPro: IPR019270  Members of this family of hypothetical proteins have no known function. 
Probab=21.80  E-value=62  Score=22.09  Aligned_cols=11  Identities=36%  Similarity=0.866  Sum_probs=8.4

Q ss_pred             cCCCcEEEEEe
Q 017471          126 NDKGKCVGIAF  136 (371)
Q Consensus       126 n~~G~VIGI~~  136 (371)
                      |.+|++|||-.
T Consensus        36 d~~G~ivGIEI   46 (50)
T PF10049_consen   36 DEDGRIVGIEI   46 (50)
T ss_pred             CCCCCEEEEEE
Confidence            46799999843


No 131
>COG5233 GRH1 Peripheral Golgi membrane protein [Intracellular trafficking and secretion]
Probab=21.71  E-value=47  Score=32.09  Aligned_cols=30  Identities=23%  Similarity=0.424  Sum_probs=26.6

Q ss_pred             EEEEECCCCcccc-CCCCCCEEEEECCEEec
Q 017471          201 RIRRVDPTAPESE-VLKPSDIILSFDGIDIA  230 (371)
Q Consensus       201 ~V~~V~~~spA~~-GL~~GDvIl~vnG~~V~  230 (371)
                      -+-+|.+.+||++ |.-.||.|+.+|+-++.
T Consensus        66 ~~lrv~~~~~~e~~~~~~~dyilg~n~Dp~~   96 (417)
T COG5233          66 EVLRVNPESPAEKAGMVVGDYILGINEDPLR   96 (417)
T ss_pred             hheeccccChhHhhccccceeEEeecCCcHH
Confidence            4568899999999 99999999999988875


No 132
>PF14827 Cache_3:  Sensory domain of two-component sensor kinase; PDB: 1OJG_A 3BY8_A 1P0Z_I 2V9A_A 2J80_B.
Probab=21.67  E-value=77  Score=25.40  Aligned_cols=17  Identities=24%  Similarity=0.718  Sum_probs=11.7

Q ss_pred             ceecCCCcEEEEEeeec
Q 017471          123 PAFNDKGKCVGIAFQSL  139 (371)
Q Consensus       123 Plvn~~G~VIGI~~~~~  139 (371)
                      |+.|.+|+++|++.-.+
T Consensus        95 PV~d~~g~viG~V~VG~  111 (116)
T PF14827_consen   95 PVYDSDGKVIGVVSVGV  111 (116)
T ss_dssp             EEE-TTS-EEEEEEEEE
T ss_pred             eeECCCCcEEEEEEEEE
Confidence            66788999999986543


No 133
>TIGR00074 hypC_hupF hydrogenase assembly chaperone HypC/HupF. An additional proposed function is to shuttle the iron atom that has been liganded at the HypC/HypD complex to the precursor of the large hydrogenase (HycE) subunit. PubMed:12441107.
Probab=21.59  E-value=1.6e+02  Score=22.14  Aligned_cols=41  Identities=10%  Similarity=0.063  Sum_probs=27.0

Q ss_pred             EeEEEEEecCCCCEEEEEEecCCCcCCccceecCCCCCCCCeEEEE
Q 017471           30 APSATLVTADICIYTMLTVEDDEFWEGVLPVEFGELPALQDAVTVV   75 (371)
Q Consensus        30 ~~A~vv~~d~~~DlAlLkv~~~~~~~~l~~~~l~~s~~lgd~V~~i   75 (371)
                      +|++|+.++.  +.|++.+....   .--.+.+-+..++||+|++-
T Consensus         5 iP~~V~~i~~--~~A~v~~~G~~---~~v~l~lv~~~~vGD~VLVH   45 (76)
T TIGR00074         5 IPGQVVEIDE--NIALVEFCGIK---RDVSLDLVGEVKVGDYVLVH   45 (76)
T ss_pred             cceEEEEEcC--CEEEEEcCCeE---EEEEEEeeCCCCCCCEEEEe
Confidence            6899999876  57888887542   11122333456789998874


No 134
>PF15436 PGBA_N:  Plasminogen-binding protein pgbA N-terminal
Probab=21.52  E-value=3.4e+02  Score=24.84  Aligned_cols=59  Identities=8%  Similarity=0.042  Sum_probs=37.1

Q ss_pred             eccEEEEEEEecCCCeEeEEEEEecCCCCEEEEEEecCCCcC--CccceecCCCCCCCCeEEE
Q 017471           14 RNEALILSTWLLCSPSAPSATLVTADICIYTMLTVEDDEFWE--GVLPVEFGELPALQDAVTV   74 (371)
Q Consensus        14 ~gsg~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~~~~~~--~l~~~~l~~s~~lgd~V~~   74 (371)
                      +.||+|+.-.--+-..+-|+++....+.+.|.+|+.+-+-.+  .+|....  .++.||+|+.
T Consensus        28 G~SGiV~h~~~~~~~~IiA~a~V~~~~~g~A~~kf~~fd~L~Q~aLP~p~~--~pk~GD~vil   88 (218)
T PF15436_consen   28 GESGIVVHKFDKDHSSIIARAVVISKKNGVAKAKFSVFDSLKQDALPTPKM--VPKKGDEVIL   88 (218)
T ss_pred             CCceEEEEEecCCcceeeeEEEEEEecCCeeEEEEeehhhhhhhcCCCCcc--ccCCCCEEEE
Confidence            567777764445666677887777778999999998653111  2222222  3567877664


No 135
>cd00991 PDZ_archaeal_metalloprotease PDZ domain of archaeal zinc metalloprotases, presumably membrane-associated or integral membrane proteases, which may be involved in signalling and regulatory mechanisms. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=21.37  E-value=55  Score=24.23  Aligned_cols=27  Identities=4%  Similarity=-0.089  Sum_probs=20.7

Q ss_pred             ccCceEEEEEecCChHhHHhhhhhhhH
Q 017471          316 IMNMKLRSSFWTSSCIQCHNCQMSSLL  342 (371)
Q Consensus       316 ~~~~~~v~~~~~~Sp~~~~~~~~~~~~  342 (371)
                      ...|+++..+.++||++-.-.+.+|+|
T Consensus         8 ~~~Gv~V~~V~~~spa~~aGL~~GDiI   34 (79)
T cd00991           8 AVAGVVIVGVIVGSPAENAVLHTGDVI   34 (79)
T ss_pred             cCCcEEEEEECCCChHHhcCCCCCCEE
Confidence            346888899999999987767767664


No 136
>PF11730 DUF3297:  Protein of unknown function (DUF3297);  InterPro: IPR021724  This family is expressed in Proteobacteria and Actinobacteria. The function is not known. 
Probab=21.35  E-value=58  Score=23.91  Aligned_cols=59  Identities=17%  Similarity=0.291  Sum_probs=38.6

Q ss_pred             EECCCCcccc-CCCCCCEEEEECCEEecCCCCccccccccchhhhhhhccCCCCEEEEEEEECCEEEEEEEE
Q 017471          204 RVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT  274 (371)
Q Consensus       204 ~V~~~spA~~-GL~~GDvIl~vnG~~V~~~~~l~~~~~~~~~~~~~~~~~~~g~~v~l~v~R~g~~~~~~v~  274 (371)
                      ++.|.||-.. -+-.-|+=+.+||+.=++.+++-+..|        ..+-..|+    ...|.|+.++++++
T Consensus         5 S~~P~Sp~~~~~~l~~~iGIrfng~Er~nVeEYciSEG--------Wvrv~~gk----a~DR~G~Pl~iklk   64 (71)
T PF11730_consen    5 SINPRSPHYDAEVLERGIGIRFNGKERTNVEEYCISEG--------WVRVAAGK----ALDRRGNPLTIKLK   64 (71)
T ss_pred             ccCCCChhhHHHHHhcCcceEECCeEcccceeEeccCC--------EEEeecCc----ccccCCCeeEEEEc
Confidence            4678898877 667778889999999999988752211        11122232    33577777666553


No 137
>PRK09570 rpoH DNA-directed RNA polymerase subunit H; Reviewed
Probab=20.77  E-value=59  Score=24.74  Aligned_cols=19  Identities=26%  Similarity=0.529  Sum_probs=14.0

Q ss_pred             EECCCCcccc--CCCCCCEEE
Q 017471          204 RVDPTAPESE--VLKPSDIIL  222 (371)
Q Consensus       204 ~V~~~spA~~--GL~~GDvIl  222 (371)
                      .+...-|+++  |+++||+|-
T Consensus        39 ~I~~~DPv~r~~g~k~GdVvk   59 (79)
T PRK09570         39 KIKASDPVVKAIGAKPGDVIK   59 (79)
T ss_pred             ceeccChhhhhcCCCCCCEEE
Confidence            3445567766  999999984


No 138
>cd00433 Peptidase_M17 Cytosol aminopeptidase family, N-terminal and catalytic domains.  Family M17 contains zinc- and manganese-dependent exopeptidases ( EC  3.4.11.1), including leucine aminopeptidase. They catalyze removal of amino acids from the N-terminus of a protein and play a key role in protein degradation and in the metabolism of biologically active peptides. They do not contain HEXXH motif (which is used as one of the signature patterns to group the peptidase families) in the metal-binding site. The two associated zinc ions and the active site are entirely enclosed within the C-terminal catalytic domain in leucine aminopeptidase. The enzyme is a hexamer, with the catalytic domains clustered around the three-fold axis, and the two trimers related to one another by a two-fold rotation. The N-terminal domain is structurally similar to the ADP-ribose binding Macro domain. This family includes proteins from bacteria, archaea, animals and plants.
Probab=20.75  E-value=1.1e+02  Score=31.55  Aligned_cols=28  Identities=18%  Similarity=0.299  Sum_probs=21.9

Q ss_pred             EEECCCCccccCCCCCCEEEEECCEEec
Q 017471          203 RRVDPTAPESEVLKPSDIILSFDGIDIA  230 (371)
Q Consensus       203 ~~V~~~spA~~GL~~GDvIl~vnG~~V~  230 (371)
                      .-..+|.+.....+|||||.+.||+.|.
T Consensus       290 ~~~~EN~is~~A~rPgDVi~s~~GkTVE  317 (468)
T cd00433         290 LPLAENMISGNAYRPGDVITSRSGKTVE  317 (468)
T ss_pred             EEeeecCCCCCCCCCCCEeEeCCCcEEE
Confidence            3445666666678999999999999874


No 139
>PRK00913 multifunctional aminopeptidase A; Provisional
Probab=20.67  E-value=1.1e+02  Score=31.55  Aligned_cols=27  Identities=22%  Similarity=0.465  Sum_probs=21.7

Q ss_pred             EECCCCccccCCCCCCEEEEECCEEec
Q 017471          204 RVDPTAPESEVLKPSDIILSFDGIDIA  230 (371)
Q Consensus       204 ~V~~~spA~~GL~~GDvIl~vnG~~V~  230 (371)
                      -..++.|...-.+|||||++.||+.|.
T Consensus       305 ~l~ENm~~~~A~rPgDVi~~~~GkTVE  331 (483)
T PRK00913        305 AACENMPSGNAYRPGDVLTSMSGKTIE  331 (483)
T ss_pred             EeeccCCCCCCCCCCCEEEECCCcEEE
Confidence            345667766678999999999999875


No 140
>cd01718 Sm_E The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet.  Sm subunit E binds subunits F and G to form a trimer which then assembles onto snRNA along with the D1/D2 and D3/B heterodimers forming a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=20.61  E-value=1.9e+02  Score=21.97  Aligned_cols=30  Identities=10%  Similarity=0.152  Sum_probs=24.0

Q ss_pred             EEEEec--CCCeEeEEEEEecCCCCEEEEEEe
Q 017471           20 LSTWLL--CSPSAPSATLVTADICIYTMLTVE   49 (371)
Q Consensus        20 i~~~~~--~~~~~~A~vv~~d~~~DlAlLkv~   49 (371)
                      +.+++.  +++++.+++.++|...|+.+=...
T Consensus        21 V~V~l~~~~g~~~~G~L~gfD~~mNlvL~d~~   52 (79)
T cd01718          21 VQIWLYEQTDLRIEGVIIGFDEYMNLVLDDAE   52 (79)
T ss_pred             EEEEEEeCCCcEEEEEEEEEccceeEEEcCEE
Confidence            345554  899999999999999999876543


No 141
>cd01725 LSm2 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm2 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=20.49  E-value=2e+02  Score=21.77  Aligned_cols=33  Identities=6%  Similarity=-0.092  Sum_probs=28.6

Q ss_pred             EEEEEEEecCCCeEeEEEEEecCCCCEEEEEEe
Q 017471           17 ALILSTWLLCSPSAPSATLVTADICIYTMLTVE   49 (371)
Q Consensus        17 g~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~   49 (371)
                      |-.+.+.+.+|..+.+++..+|...|+-+=.+.
T Consensus        11 g~~V~VeLKng~~~~G~L~~vD~~MNi~L~n~~   43 (81)
T cd01725          11 GKEVTVELKNDLSIRGTLHSVDQYLNIKLTNIS   43 (81)
T ss_pred             CCEEEEEECCCcEEEEEEEEECCCcccEEEEEE
Confidence            445668889999999999999999999888775


No 142
>cd01724 Sm_D1 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit D1 heterodimerizes with subunit D2 and three such heterodimers form a hexameric ring structure with alternating D1 and D2 subunits. The D1 - D2 heterodimer also assembles into a heptameric ring containing DB, D3, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=20.49  E-value=1.8e+02  Score=22.51  Aligned_cols=35  Identities=6%  Similarity=-0.175  Sum_probs=30.1

Q ss_pred             cEEEEEEEecCCCeEeEEEEEecCCCCEEEEEEec
Q 017471           16 EALILSTWLLCSPSAPSATLVTADICIYTMLTVED   50 (371)
Q Consensus        16 sg~vi~~~~~~~~~~~A~vv~~d~~~DlAlLkv~~   50 (371)
                      .|-.+.+.+.+|..+.+++..+|...|+.+-.+..
T Consensus        10 ~g~~V~VeLKng~~~~G~L~~vD~~MNl~L~~a~~   44 (90)
T cd01724          10 TNETVTIELKNGTIVHGTITGVDPSMNTHLKNVKL   44 (90)
T ss_pred             CCCEEEEEECCCCEEEEEEEEEcCceeEEEEEEEE
Confidence            45567788999999999999999999999988753


Done!