Query         009784
Match_columns 526
No_of_seqs    483 out of 3258
Neff          7.3 
Searched_HMMs 46136
Date          Thu Mar 28 17:17:35 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/009784.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/009784hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PRK10139 serine endoprotease;  100.0 2.7E-51 5.9E-56  438.4  38.6  367  111-503    42-431 (455)
  2 TIGR02037 degP_htrA_DO peripla 100.0 1.4E-48   3E-53  417.1  37.6  330  147-502    56-402 (428)
  3 PRK10942 serine endoprotease;  100.0 2.2E-48 4.7E-53  418.0  36.2  329  148-502   110-448 (473)
  4 TIGR02038 protease_degS peripl 100.0 1.6E-47 3.4E-52  398.1  34.5  296  111-433    47-349 (351)
  5 PRK10898 serine endoprotease;  100.0 7.6E-47 1.6E-51  392.8  34.0  297  111-434    47-351 (353)
  6 COG0265 DegQ Trypsin-like seri 100.0 8.5E-37 1.8E-41  317.9  28.5  300  111-433    35-341 (347)
  7 KOG1421 Predicted signaling-as 100.0 1.7E-31 3.6E-36  281.0  20.4  322  112-463    52-395 (955)
  8 KOG1320 Serine protease [Postt 100.0 1.2E-29 2.7E-34  265.5  15.8  372  118-503    56-439 (473)
  9 KOG1320 Serine protease [Postt  99.9 8.3E-25 1.8E-29  229.3  20.3  301  118-432   134-468 (473)
 10 KOG1421 Predicted signaling-as  99.8   2E-17 4.2E-22  175.5  24.2  327  124-492   530-891 (955)
 11 PF13365 Trypsin_2:  Trypsin-li  99.6   6E-15 1.3E-19  128.7  14.1  108  151-289     1-120 (120)
 12 PF13180 PDZ_2:  PDZ domain; PD  99.5   4E-14 8.7E-19  116.6   7.4   81  328-430     1-82  (82)
 13 PF00089 Trypsin:  Trypsin;  In  99.4 1.5E-12 3.2E-17  125.1  14.4  185  131-315     4-220 (220)
 14 cd00190 Tryp_SPc Trypsin-like   99.3   5E-11 1.1E-15  115.2  16.8  161  134-294     7-208 (232)
 15 cd00987 PDZ_serine_protease PD  99.3 1.5E-11 3.3E-16  102.3   8.5   88  328-427     1-89  (90)
 16 smart00020 Tryp_SPc Trypsin-li  99.2 1.5E-10 3.3E-15  112.2  13.7  164  131-294     5-208 (229)
 17 cd00986 PDZ_LON_protease PDZ d  99.2 8.6E-11 1.9E-15   95.9   8.0   72  352-433     7-78  (79)
 18 cd00991 PDZ_archaeal_metallopr  99.1 1.2E-10 2.6E-15   95.2   7.8   69  351-429     8-77  (79)
 19 TIGR01713 typeII_sec_gspC gene  99.1   4E-10 8.6E-15  112.5  11.4  100  309-430   159-259 (259)
 20 cd00990 PDZ_glycyl_aminopeptid  99.1 4.5E-10 9.8E-15   91.5   8.9   77  328-431     1-78  (80)
 21 PRK10779 zinc metallopeptidase  99.0 8.9E-10 1.9E-14  118.9   9.0  132  355-503   128-262 (449)
 22 TIGR02037 degP_htrA_DO peripla  98.9 3.3E-09 7.2E-14  113.9   8.6   90  327-427   337-427 (428)
 23 cd00989 PDZ_metalloprotease PD  98.9 4.1E-09 8.8E-14   85.5   6.8   66  353-429    12-78  (79)
 24 cd00988 PDZ_CTP_protease PDZ d  98.9   1E-08 2.2E-13   84.4   8.3   68  352-430    12-83  (85)
 25 COG3591 V8-like Glu-specific e  98.8 9.2E-08   2E-12   94.1  14.8  161  149-319    64-250 (251)
 26 cd00136 PDZ PDZ domain, also c  98.6 8.4E-08 1.8E-12   75.9   5.8   55  353-418    13-70  (70)
 27 TIGR00054 RIP metalloprotease   98.5 1.7E-07 3.6E-12  100.4   7.6  113  352-501   127-242 (420)
 28 TIGR00054 RIP metalloprotease   98.5 2.4E-07 5.2E-12   99.2   7.2   69  353-432   203-272 (420)
 29 smart00228 PDZ Domain present   98.5 4.8E-07   1E-11   73.8   7.2   59  353-421    26-85  (85)
 30 PRK10779 zinc metallopeptidase  98.4   6E-07 1.3E-11   97.0   7.2   68  354-432   222-290 (449)
 31 TIGR00225 prc C-terminal pepti  98.3 1.3E-06 2.8E-11   90.8   8.3   70  353-433    62-134 (334)
 32 PF00595 PDZ:  PDZ domain (Also  98.3 1.5E-06 3.2E-11   71.0   5.4   72  327-418     9-81  (81)
 33 PRK10139 serine endoprotease;   98.2 2.1E-06 4.6E-11   92.8   6.7   64  353-428   390-454 (455)
 34 TIGR03279 cyano_FeS_chp putati  98.2 1.9E-06 4.2E-11   90.9   6.2   62  357-432     2-65  (433)
 35 KOG3627 Trypsin [Amino acid tr  98.2 9.9E-05 2.2E-09   73.2  17.3  167  128-295    13-229 (256)
 36 PLN00049 carboxyl-terminal pro  98.1 4.9E-06 1.1E-10   88.3   8.0   69  353-430   102-171 (389)
 37 PF14685 Tricorn_PDZ:  Tricorn   98.1 1.6E-05 3.5E-10   66.1   9.1   65  352-427    11-87  (88)
 38 TIGR02860 spore_IV_B stage IV   98.1   4E-06 8.6E-11   88.0   6.3   69  352-431   104-181 (402)
 39 PRK10942 serine endoprotease;   98.1 5.5E-06 1.2E-10   90.0   6.8   64  353-428   408-472 (473)
 40 cd00992 PDZ_signaling PDZ doma  98.0 9.4E-06   2E-10   65.9   5.6   52  328-390    12-66  (82)
 41 PF00863 Peptidase_C4:  Peptida  98.0 0.00061 1.3E-08   66.8  18.2  169  119-317    14-195 (235)
 42 COG3480 SdrC Predicted secrete  97.9 1.4E-05 3.1E-10   80.0   6.1   72  352-433   129-201 (342)
 43 COG0793 Prc Periplasmic protea  97.9   3E-05 6.6E-10   82.5   8.6   84  327-434    99-185 (406)
 44 PF05579 Peptidase_S32:  Equine  97.8 0.00024 5.1E-09   69.8  11.7  115  149-294   112-229 (297)
 45 PRK09681 putative type II secr  97.8 3.9E-05 8.4E-10   76.8   6.0   62  359-430   210-275 (276)
 46 KOG3129 26S proteasome regulat  97.7 6.7E-05 1.5E-09   70.9   6.6   73  354-434   140-213 (231)
 47 PF04495 GRASP55_65:  GRASP55/6  97.5 0.00034 7.4E-09   63.3   7.2   87  327-432    25-115 (138)
 48 PF00548 Peptidase_C3:  3C cyst  97.4  0.0047   1E-07   58.2  14.2  138  147-293    23-170 (172)
 49 PRK11186 carboxy-terminal prot  97.4 0.00068 1.5E-08   76.2   9.9   71  353-429   255-332 (667)
 50 COG3975 Predicted protease wit  97.3 0.00056 1.2E-08   73.1   7.7   85  330-434   439-526 (558)
 51 COG3031 PulC Type II secretory  97.0 0.00067 1.5E-08   65.5   3.7   66  354-429   208-274 (275)
 52 PF03761 DUF316:  Domain of unk  96.8   0.039 8.4E-07   55.8  15.2  107  195-312   160-272 (282)
 53 KOG3553 Tax interaction protei  96.7  0.0018 3.8E-08   54.3   3.5   35  351-385    57-92  (124)
 54 PF08192 Peptidase_S64:  Peptid  96.5    0.03 6.4E-07   61.8  12.5  117  195-318   542-688 (695)
 55 COG5640 Secreted trypsin-like   96.4   0.062 1.3E-06   55.3  13.1   59  149-207    61-135 (413)
 56 PF12812 PDZ_1:  PDZ-like domai  96.3  0.0044 9.6E-08   50.5   3.8   60  328-391     9-69  (78)
 57 PF10459 Peptidase_S46:  Peptid  96.3  0.0048   1E-07   69.8   5.4   55  264-318   624-686 (698)
 58 PF02122 Peptidase_S39:  Peptid  96.2  0.0035 7.5E-08   60.4   3.0  137  159-310    42-183 (203)
 59 KOG3209 WW domain-containing p  95.8   0.031 6.6E-07   61.7   8.5  140  351-502   672-819 (984)
 60 KOG3209 WW domain-containing p  95.8  0.0092   2E-07   65.6   4.5  152  357-518   782-980 (984)
 61 KOG3580 Tight junction protein  95.5   0.011 2.3E-07   63.9   3.5   61  351-419   427-488 (1027)
 62 KOG3580 Tight junction protein  94.7   0.039 8.4E-07   59.7   4.9   74  345-429   212-287 (1027)
 63 PF00949 Peptidase_S7:  Peptida  94.6   0.043 9.3E-07   49.2   4.2   29  266-294    90-118 (132)
 64 KOG3550 Receptor targeting pro  94.0     0.1 2.2E-06   47.2   5.2   36  352-387   114-151 (207)
 65 KOG3834 Golgi reassembly stack  93.3    0.36 7.8E-06   50.8   8.5  114  352-489    14-137 (462)
 66 PF09342 DUF1986:  Domain of un  93.2     1.2 2.6E-05   43.9  11.4   99  135-234    12-131 (267)
 67 KOG3532 Predicted protein kina  93.0     0.1 2.2E-06   57.4   4.2   39  353-391   398-437 (1051)
 68 KOG2921 Intramembrane metallop  91.8    0.12 2.5E-06   53.8   2.7   40  351-390   218-259 (484)
 69 KOG3605 Beta amyloid precursor  91.1    0.37 8.1E-06   53.0   5.8  103  272-386   679-790 (829)
 70 PF00944 Peptidase_S3:  Alphavi  90.8    0.18 3.9E-06   44.8   2.5   27  268-294   101-127 (158)
 71 PF05580 Peptidase_S55:  SpoIVB  90.4     6.6 0.00014   38.1  12.9   41  268-311   175-215 (218)
 72 COG0750 Predicted membrane-ass  90.4    0.47   1E-05   49.9   5.8   57  358-425   134-195 (375)
 73 KOG1892 Actin filament-binding  90.3    0.34 7.4E-06   55.4   4.7   61  351-421   958-1020(1629)
 74 PF02907 Peptidase_S29:  Hepati  90.3    0.37 8.1E-06   42.9   4.0   41  270-311   105-146 (148)
 75 KOG3542 cAMP-regulated guanine  90.2    0.21 4.6E-06   54.9   3.0   42  347-388   556-598 (1283)
 76 KOG3552 FERM domain protein FR  89.4    0.32 6.9E-06   55.5   3.6   57  354-420    76-132 (1298)
 77 KOG3571 Dishevelled 3 and rela  87.5    0.66 1.4E-05   49.8   4.3   38  352-389   276-315 (626)
 78 KOG3551 Syntrophins (type beta  87.3     0.4 8.6E-06   49.8   2.5   56  355-420   112-171 (506)
 79 PF00947 Pico_P2A:  Picornaviru  86.7       2 4.4E-05   38.0   6.2   31  262-293    79-109 (127)
 80 KOG3605 Beta amyloid precursor  84.9     2.3 4.9E-05   47.2   6.8  119  359-505   679-799 (829)
 81 PF10459 Peptidase_S46:  Peptid  84.1    0.53 1.2E-05   53.6   1.8   22  149-170    47-69  (698)
 82 PF02395 Peptidase_S6:  Immunog  81.4     7.2 0.00016   45.1   9.5  161  151-318    67-266 (769)
 83 KOG3606 Cell polarity protein   81.3     1.8 3.9E-05   43.0   4.0   64  318-386   164-229 (358)
 84 PF12812 PDZ_1:  PDZ-like domai  80.6     2.8   6E-05   34.1   4.3   56  446-501     5-69  (78)
 85 KOG3549 Syntrophins (type gamm  80.3     1.8 3.9E-05   44.4   3.7   55  354-418    81-137 (505)
 86 KOG3651 Protein kinase C, alph  79.4     2.7 5.9E-05   42.4   4.6   37  354-390    31-69  (429)
 87 PF03510 Peptidase_C24:  2C end  75.7     9.6 0.00021   32.8   6.3   54  153-220     3-56  (105)
 88 PF01732 DUF31:  Putative pepti  71.0     3.2   7E-05   43.9   2.9   24  269-292   351-374 (374)
 89 KOG0609 Calcium/calmodulin-dep  70.2     6.7 0.00014   42.8   5.0   57  354-420   147-205 (542)
 90 KOG0606 Microtubule-associated  69.6     4.9 0.00011   47.4   4.1   34  355-388   660-694 (1205)
 91 TIGR02860 spore_IV_B stage IV   68.0     3.6 7.9E-05   43.8   2.5   42  268-312   355-396 (402)
 92 KOG1924 RhoA GTPase effector D  54.4      42 0.00091   38.5   7.6   10  197-206   719-728 (1102)
 93 PF05416 Peptidase_C37:  Southa  54.3      95  0.0021   33.3   9.8  135  149-294   379-527 (535)
 94 smart00384 AT_hook DNA binding  46.0      13 0.00028   23.6   1.2   16    5-20      1-16  (26)
 95 PF12381 Peptidase_C3G:  Tungro  45.5      30 0.00065   33.6   4.3   54  263-319   170-229 (231)
 96 PF13180 PDZ_2:  PDZ domain; PD  42.8      63  0.0014   25.7   5.3   50  452-502     3-54  (82)
 97 KOG1924 RhoA GTPase effector D  40.7      80  0.0017   36.4   7.1    9   12-20    502-510 (1102)
 98 KOG3938 RGS-GAIP interacting p  39.7      15 0.00032   36.8   1.2   58  355-420   151-210 (334)
 99 cd01720 Sm_D2 The eukaryotic S  37.1      63  0.0014   26.8   4.4   37  167-204    10-46  (87)
100 cd00600 Sm_like The eukaryotic  34.5 1.1E+02  0.0023   23.0   5.1   33  172-205     7-39  (63)
101 KOG3834 Golgi reassembly stack  31.5      37 0.00081   36.2   2.7   65  356-431   112-180 (462)
102 PF02178 AT_hook:  AT hook moti  31.2      21 0.00045   19.0   0.4   11    5-15      1-11  (13)
103 PF09465 LBR_tudor:  Lamin-B re  30.5 2.1E+02  0.0046   21.7   5.8   38  169-206     7-44  (55)
104 TIGR03000 plancto_dom_1 Planct  30.4 1.5E+02  0.0032   24.0   5.3   49  372-429    10-62  (75)
105 PF00571 CBS:  CBS domain CBS d  29.8      45 0.00099   24.1   2.3   21  272-292    28-48  (57)
106 cd01726 LSm6 The eukaryotic Sm  29.3 1.2E+02  0.0026   23.5   4.7   32  172-204    11-42  (67)
107 PRK00737 small nuclear ribonuc  28.9 1.2E+02  0.0027   23.9   4.8   33  172-205    15-47  (72)
108 cd01731 archaeal_Sm1 The archa  28.8 1.3E+02  0.0028   23.4   4.8   33  172-205    11-43  (68)
109 cd01722 Sm_F The eukaryotic Sm  28.5 1.2E+02  0.0025   23.7   4.5   32  172-204    12-43  (68)
110 cd01730 LSm3 The eukaryotic Sm  27.1 1.1E+02  0.0024   24.9   4.3   31  172-203    12-42  (82)
111 cd01717 Sm_B The eukaryotic Sm  26.2 1.3E+02  0.0029   24.1   4.6   32  172-204    11-42  (79)
112 COG0298 HypC Hydrogenase matur  26.0 1.5E+02  0.0033   24.3   4.7   47  185-234     5-53  (82)
113 cd06168 LSm9 The eukaryotic Sm  25.6 1.6E+02  0.0035   23.6   4.9   32  172-204    11-42  (75)
114 cd01732 LSm5 The eukaryotic Sm  24.7 1.5E+02  0.0032   23.9   4.5   31  172-203    14-44  (76)
115 cd01729 LSm7 The eukaryotic Sm  24.6 1.6E+02  0.0034   24.0   4.8   32  172-204    13-44  (81)
116 cd01735 LSm12_N LSm12 belongs   24.6 2.6E+02  0.0055   21.7   5.6   33  172-205     7-39  (61)
117 PF11874 DUF3394:  Domain of un  24.0      71  0.0015   30.4   2.9   28  352-379   121-149 (183)
118 COG2524 Predicted transcriptio  23.4 2.2E+02  0.0048   28.7   6.2   94  193-292   110-220 (294)
119 cd01719 Sm_G The eukaryotic Sm  22.7 1.9E+02  0.0042   22.9   4.8   32  172-204    11-42  (72)
120 PF00595 PDZ:  PDZ domain (Also  22.5 1.8E+02   0.004   22.8   4.8   52  453-504    13-67  (81)
121 cd01727 LSm8 The eukaryotic Sm  21.9 3.4E+02  0.0073   21.5   6.1   33  172-205    10-42  (74)
122 cd01721 Sm_D3 The eukaryotic S  21.7 2.1E+02  0.0046   22.4   4.9   32  172-204    11-42  (70)
123 smart00651 Sm snRNP Sm protein  21.3 2.2E+02  0.0049   21.6   4.9   33  172-205     9-41  (67)
124 cd01728 LSm1 The eukaryotic Sm  20.8 2.2E+02  0.0047   22.8   4.7   32  172-204    13-44  (74)
125 PF01423 LSM:  LSM domain ;  In  20.7 1.7E+02  0.0037   22.3   4.1   34  172-206     9-42  (67)
126 COG0260 PepB Leucyl aminopepti  20.2 1.3E+02  0.0028   33.1   4.4   46  358-406   303-348 (485)

No 1  
>PRK10139 serine endoprotease; Provisional
Probab=100.00  E-value=2.7e-51  Score=438.40  Aligned_cols=367  Identities=23%  Similarity=0.326  Sum_probs=290.3

Q ss_pred             ChhhhhhhcccCCCeEEEEeeeeCCC-------------CCCccccCCCcceEEEEEEEe--CCEEEecccccCCCCeEE
Q 009784          111 VEPGVARVVPAMDAVVKVFCVHTEPN-------------FSLPWQRKRQYSSSSSGFAIG--GRRVLTNAHSVEHYTQVK  175 (526)
Q Consensus       111 ~~~~v~~~~~~~~SVV~I~~~~~~~~-------------~~~P~~~~~~~~~~GSGfvI~--~g~ILT~aHvV~~~~~i~  175 (526)
                      +.+.++++.|   |||.|.+......             ...||.......+.||||+|+  +||||||+|||.++..+.
T Consensus        42 ~~~~~~~~~p---avV~i~~~~~~~~~~~~~~~~~~~f~~~~~~~~~~~~~~~GSG~ii~~~~g~IlTn~HVv~~a~~i~  118 (455)
T PRK10139         42 LAPMLEKVLP---AVVSVRVEGTASQGQKIPEEFKKFFGDDLPDQPAQPFEGLGSGVIIDAAKGYVLTNNHVINQAQKIS  118 (455)
T ss_pred             HHHHHHHhCC---cEEEEEEEEeecccccCchhHHHhccccCCccccccccceEEEEEEECCCCEEEeChHHhCCCCEEE
Confidence            4455555555   9999987653221             011333333345789999997  589999999999999999


Q ss_pred             EEEcCCCcEEEEEEEEeccCCCeEEEEecccccccCceeeecCCCC--cCCCcEEEEeeCCCCCceeEEEEEEeceeeee
Q 009784          176 LKKRGSDTKYLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGELP--ALQDAVTVVGYPIGGDTISVTSGVVSRIEILS  253 (526)
Q Consensus       176 V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~~--~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~  253 (526)
                      |++. |++.++|++++.|+.+||||||++...   .+++++|+++.  ++|++|+++|||++... +++.|+||+..+..
T Consensus       119 V~~~-dg~~~~a~vvg~D~~~DlAvlkv~~~~---~l~~~~lg~s~~~~~G~~V~aiG~P~g~~~-tvt~GivS~~~r~~  193 (455)
T PRK10139        119 IQLN-DGREFDAKLIGSDDQSDIALLQIQNPS---KLTQIAIADSDKLRVGDFAVAVGNPFGLGQ-TATSGIISALGRSG  193 (455)
T ss_pred             EEEC-CCCEEEEEEEEEcCCCCEEEEEecCCC---CCceeEecCccccCCCCEEEEEecCCCCCC-ceEEEEEccccccc
Confidence            9997 999999999999999999999998643   67899999765  57999999999999776 89999999987643


Q ss_pred             ccCCceeeeEEEEcccCCCCCCCCeeecCCCeEEEEEeeccccC-ccccccccccHHHHHHHHHHHHHcCceeeccccCc
Q 009784          254 YVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-DVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGV  332 (526)
Q Consensus       254 ~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~-~~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi  332 (526)
                      ... .....+||+|+++++|||||||+|.+|+||||+++.+... +..+++||||++.+++++++|+++|++. ++|||+
T Consensus       194 ~~~-~~~~~~iqtda~in~GnSGGpl~n~~G~vIGi~~~~~~~~~~~~gigfaIP~~~~~~v~~~l~~~g~v~-r~~LGv  271 (455)
T PRK10139        194 LNL-EGLENFIQTDASINRGNSGGALLNLNGELIGINTAILAPGGGSVGIGFAIPSNMARTLAQQLIDFGEIK-RGLLGI  271 (455)
T ss_pred             cCC-CCcceEEEECCccCCCCCcceEECCCCeEEEEEEEEEcCCCCccceEEEEEhHHHHHHHHHHhhcCccc-ccceeE
Confidence            221 1234589999999999999999999999999999877543 3578999999999999999999999998 999999


Q ss_pred             eeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCC
Q 009784          333 EWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGD  411 (526)
Q Consensus       333 ~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~  411 (526)
                      .++++ +++.++.+|++ ...|++|.+|.++|||++ |||+||+|++|||++|.++.++.          ..+....+|+
T Consensus       272 ~~~~l-~~~~~~~lgl~-~~~Gv~V~~V~~~SpA~~AGL~~GDvIl~InG~~V~s~~dl~----------~~l~~~~~g~  339 (455)
T PRK10139        272 KGTEM-SADIAKAFNLD-VQRGAFVSEVLPNSGSAKAGVKAGDIITSLNGKPLNSFAELR----------SRIATTEPGT  339 (455)
T ss_pred             EEEEC-CHHHHHhcCCC-CCCceEEEEECCCChHHHCCCCCCCEEEEECCEECCCHHHHH----------HHHHhcCCCC
Confidence            99999 88999999997 467999999999999999 99999999999999999999874          6666667899


Q ss_pred             EEEEEEEECCEEEEEEEEecccccccCCCCCCCCCceEEEeeEEEEecccc--ceeeeeeeecc--hhhhccccccceee
Q 009784          412 SAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLYL--ISVLSMERIMN--MKLRSSFWTSSCIQ  487 (526)
Q Consensus       412 ~v~l~v~R~G~~~~~~v~l~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt~~--~~~~~~~~i~~--~~~~sg~~~~~~~~  487 (526)
                      ++.++|+|+|+.+++++++...+...... ....+   .+.|+.+.+....  ...+.+..+.+  .+.++||+.||.|.
T Consensus       340 ~v~l~V~R~G~~~~l~v~~~~~~~~~~~~-~~~~~---~~~g~~l~~~~~~~~~~Gv~V~~V~~~spA~~aGL~~GD~I~  415 (455)
T PRK10139        340 KVKLGLLRNGKPLEVEVTLDTSTSSSASA-EMITP---ALQGATLSDGQLKDGTKGIKIDEVVKGSPAAQAGLQKDDVII  415 (455)
T ss_pred             EEEEEEEECCEEEEEEEEECCCCCccccc-ccccc---cccccEecccccccCCCceEEEEeCCCChHHHcCCCCCCEEE
Confidence            99999999999999999985443211110 00111   1234444432110  12234445544  45679999999999


Q ss_pred             eecccchhhhHHHHHH
Q 009784          488 CHNCQMSSLLWCLRCL  503 (526)
Q Consensus       488 ~~~~~~~~~~~~~~~~  503 (526)
                      ..|.+..++...|+.+
T Consensus       416 ~Ing~~v~~~~~~~~~  431 (455)
T PRK10139        416 GVNRDRVNSIAEMRKV  431 (455)
T ss_pred             EECCEEcCCHHHHHHH
Confidence            9999998888877654


No 2  
>TIGR02037 degP_htrA_DO periplasmic serine protease, Do/DeqQ family. This family consists of a set proteins various designated DegP, heat shock protein HtrA, and protease DO. The ortholog in Pseudomonas aeruginosa is designated MucD and is found in an operon that controls mucoid phenotype. This family also includes the DegQ (HhoA) paralog in E. coli which can rescue a DegP mutant, but not the smaller DegS paralog, which cannot. Members of this family are located in the periplasm and have separable functions as both protease and chaperone. Members have a trypsin domain and two copies of a PDZ domain. This protein protects bacteria from thermal and other stresses and may be important for the survival of bacterial pathogens.// The chaperone function is dominant at low temperatures, whereas the proteolytic activity is turned on at elevated temperatures.
Probab=100.00  E-value=1.4e-48  Score=417.12  Aligned_cols=330  Identities=25%  Similarity=0.340  Sum_probs=278.5

Q ss_pred             cceEEEEEEEe-CCEEEecccccCCCCeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecccccccCceeeecCCCC--cC
Q 009784          147 YSSSSSGFAIG-GRRVLTNAHSVEHYTQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGELP--AL  223 (526)
Q Consensus       147 ~~~~GSGfvI~-~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~~--~~  223 (526)
                      ..+.||||+|+ +||||||+||+.++..+.|++. +++.++|++++.|+.+||||||++...   .+++++|+++.  +.
T Consensus        56 ~~~~GSGfii~~~G~IlTn~Hvv~~~~~i~V~~~-~~~~~~a~vv~~d~~~DlAllkv~~~~---~~~~~~l~~~~~~~~  131 (428)
T TIGR02037        56 VRGLGSGVIISADGYILTNNHVVDGADEITVTLS-DGREFKAKLVGKDPRTDIAVLKIDAKK---NLPVIKLGDSDKLRV  131 (428)
T ss_pred             ccceeeEEEECCCCEEEEcHHHcCCCCeEEEEeC-CCCEEEEEEEEecCCCCEEEEEecCCC---CceEEEccCCCCCCC
Confidence            45789999999 7899999999999999999998 899999999999999999999998753   68999998654  67


Q ss_pred             CCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecCCCeEEEEEeeccccC-ccccc
Q 009784          224 QDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-DVENI  302 (526)
Q Consensus       224 g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~-~~~~~  302 (526)
                      |++|+++|||++... +++.|+|+...+... ....+..++++|+++++|||||||+|.+|+||||+++.+... +..++
T Consensus       132 G~~v~aiG~p~g~~~-~~t~G~vs~~~~~~~-~~~~~~~~i~tda~i~~GnSGGpl~n~~G~viGI~~~~~~~~g~~~g~  209 (428)
T TIGR02037       132 GDWVLAIGNPFGLGQ-TVTSGIVSALGRSGL-GIGDYENFIQTDAAINPGNSGGPLVNLRGEVIGINTAIYSPSGGNVGI  209 (428)
T ss_pred             CCEEEEEECCCcCCC-cEEEEEEEecccCcc-CCCCccceEEECCCCCCCCCCCceECCCCeEEEEEeEEEcCCCCccce
Confidence            999999999999775 899999998876432 122234579999999999999999999999999998876543 34678


Q ss_pred             cccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECC
Q 009784          303 GYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDG  381 (526)
Q Consensus       303 ~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG  381 (526)
                      +|+||++.+++++++|+++|++. ++|||+.++.+ +++.++.+|++. ..|++|.+|.++|||++ |||+||+|++|||
T Consensus       210 ~faiP~~~~~~~~~~l~~~g~~~-~~~lGi~~~~~-~~~~~~~lgl~~-~~Gv~V~~V~~~spA~~aGL~~GDvI~~Vng  286 (428)
T TIGR02037       210 GFAIPSNMAKNVVDQLIEGGKVQ-RGWLGVTIQEV-TSDLAKSLGLEK-QRGALVAQVLPGSPAEKAGLKAGDVILSVNG  286 (428)
T ss_pred             EEEEEhHHHHHHHHHHHhcCcCc-CCcCceEeecC-CHHHHHHcCCCC-CCceEEEEccCCCChHHcCCCCCCEEEEECC
Confidence            99999999999999999999998 99999999999 889999999974 57999999999999999 9999999999999


Q ss_pred             EEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEecccccccCCCCCCCCCceEEEeeEEEEeccc
Q 009784          382 IDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLY  461 (526)
Q Consensus       382 ~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt~  461 (526)
                      ++|.++.++.          ..+....+|++++++|+|+|+.+++++++...+...+       .+...+.|+.+++++.
T Consensus       287 ~~i~~~~~~~----------~~l~~~~~g~~v~l~v~R~g~~~~~~v~l~~~~~~~~-------~~~~~~lGi~~~~l~~  349 (428)
T TIGR02037       287 KPISSFADLR----------RAIGTLKPGKKVTLGILRKGKEKTITVTLGASPEEQA-------SSSNPFLGLTVANLSP  349 (428)
T ss_pred             EEcCCHHHHH----------HHHHhcCCCCEEEEEEEECCEEEEEEEEECcCCCccc-------cccccccceEEecCCH
Confidence            9999988864          6676777899999999999999999999876543211       1233467889988762


Q ss_pred             cc----------eeeeeeeecc--hhhhccccccceeeeecccchhhhHHHHH
Q 009784          462 LI----------SVLSMERIMN--MKLRSSFWTSSCIQCHNCQMSSLLWCLRC  502 (526)
Q Consensus       462 ~~----------~~~~~~~i~~--~~~~sg~~~~~~~~~~~~~~~~~~~~~~~  502 (526)
                      ..          ..+.+..+.+  .+.++||+.||+|...|.+-..+...++-
T Consensus       350 ~~~~~~~l~~~~~Gv~V~~V~~~SpA~~aGL~~GDvI~~Ing~~V~s~~d~~~  402 (428)
T TIGR02037       350 EIRKELRLKGDVKGVVVTKVVSGSPAARAGLQPGDVILSVNQQPVSSVAELRK  402 (428)
T ss_pred             HHHHHcCCCcCcCceEEEEeCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHHH
Confidence            11          3455555554  34578999999999999988887766553


No 3  
>PRK10942 serine endoprotease; Provisional
Probab=100.00  E-value=2.2e-48  Score=417.96  Aligned_cols=329  Identities=24%  Similarity=0.302  Sum_probs=270.9

Q ss_pred             ceEEEEEEEe--CCEEEecccccCCCCeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecccccccCceeeecCCCC--cC
Q 009784          148 SSSSSGFAIG--GRRVLTNAHSVEHYTQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGELP--AL  223 (526)
Q Consensus       148 ~~~GSGfvI~--~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~~--~~  223 (526)
                      .+.||||+|+  +||||||+|||.++++++|++. |++.++|++++.|+.+||||||++...   .+++++|+++.  ++
T Consensus       110 ~~~GSG~ii~~~~G~IlTn~HVv~~a~~i~V~~~-dg~~~~a~vv~~D~~~DlAvlki~~~~---~l~~~~lg~s~~l~~  185 (473)
T PRK10942        110 MALGSGVIIDADKGYVVTNNHVVDNATKIKVQLS-DGRKFDAKVVGKDPRSDIALIQLQNPK---NLTAIKMADSDALRV  185 (473)
T ss_pred             cceEEEEEEECCCCEEEeChhhcCCCCEEEEEEC-CCCEEEEEEEEecCCCCEEEEEecCCC---CCceeEecCccccCC
Confidence            4689999998  4899999999999999999998 999999999999999999999997543   67899998765  67


Q ss_pred             CCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecCCCeEEEEEeeccccC-ccccc
Q 009784          224 QDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-DVENI  302 (526)
Q Consensus       224 g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~-~~~~~  302 (526)
                      |++|+++|+|++... +++.|+|++..+..... ..+..+||+|+++++|||||||+|.+|+||||+++.+... +..++
T Consensus       186 G~~V~aiG~P~g~~~-tvt~GiVs~~~r~~~~~-~~~~~~iqtda~i~~GnSGGpL~n~~GeviGI~t~~~~~~g~~~g~  263 (473)
T PRK10942        186 GDYTVAIGNPYGLGE-TVTSGIVSALGRSGLNV-ENYENFIQTDAAINRGNSGGALVNLNGELIGINTAILAPDGGNIGI  263 (473)
T ss_pred             CCEEEEEcCCCCCCc-ceeEEEEEEeecccCCc-ccccceEEeccccCCCCCcCccCCCCCeEEEEEEEEEcCCCCcccE
Confidence            999999999998766 89999999887642211 1234579999999999999999999999999999877544 34679


Q ss_pred             cccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECC
Q 009784          303 GYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDG  381 (526)
Q Consensus       303 ~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG  381 (526)
                      +|+||++.+++++++|+++|++. |+|||+.++.+ ++++++.++++ ...|++|.+|.++|||++ |||+||+|++|||
T Consensus       264 gfaIP~~~~~~v~~~l~~~g~v~-rg~lGv~~~~l-~~~~a~~~~l~-~~~GvlV~~V~~~SpA~~AGL~~GDvIl~InG  340 (473)
T PRK10942        264 GFAIPSNMVKNLTSQMVEYGQVK-RGELGIMGTEL-NSELAKAMKVD-AQRGAFVSQVLPNSSAAKAGIKAGDVITSLNG  340 (473)
T ss_pred             EEEEEHHHHHHHHHHHHhccccc-cceeeeEeeec-CHHHHHhcCCC-CCCceEEEEECCCChHHHcCCCCCCEEEEECC
Confidence            99999999999999999999998 99999999999 78899999997 467999999999999999 9999999999999


Q ss_pred             EEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEecccccccCCCCCCCCCceEEEeeEEEEeccc
Q 009784          382 IDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLY  461 (526)
Q Consensus       382 ~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt~  461 (526)
                      ++|.++.++.          ..+....+|+++.++|+|+|+.+++++++...+.....    ....   +.|+....++-
T Consensus       341 ~~V~s~~dl~----------~~l~~~~~g~~v~l~v~R~G~~~~v~v~l~~~~~~~~~----~~~~---~lGl~g~~l~~  403 (473)
T PRK10942        341 KPISSFAALR----------AQVGTMPVGSKLTLGLLRDGKPVNVNVELQQSSQNQVD----SSNI---FNGIEGAELSN  403 (473)
T ss_pred             EECCCHHHHH----------HHHHhcCCCCEEEEEEEECCeEEEEEEEeCcCcccccc----cccc---cccceeeeccc
Confidence            9999999875          66777778999999999999999999988664221110    1111   22333333321


Q ss_pred             --cceeeeeeeecc--hhhhccccccceeeeecccchhhhHHHHH
Q 009784          462 --LISVLSMERIMN--MKLRSSFWTSSCIQCHNCQMSSLLWCLRC  502 (526)
Q Consensus       462 --~~~~~~~~~i~~--~~~~sg~~~~~~~~~~~~~~~~~~~~~~~  502 (526)
                        ....+.+..+.+  .+.++||+.||+|...|.+-..+..+|+-
T Consensus       404 ~~~~~gvvV~~V~~~S~A~~aGL~~GDvIv~VNg~~V~s~~dl~~  448 (473)
T PRK10942        404 KGGDKGVVVDNVKPGTPAAQIGLKKGDVIIGANQQPVKNIAELRK  448 (473)
T ss_pred             ccCCCCeEEEEeCCCChHHHcCCCCCCEEEEECCEEcCCHHHHHH
Confidence              112344545543  44579999999999999999988887765


No 4  
>TIGR02038 protease_degS periplasmic serine pepetdase DegS. This family consists of the periplasmic serine protease DegS (HhoB), a shorter paralog of protease DO (HtrA, DegP) and DegQ (HhoA). It is found in E. coli and several other Proteobacteria of the gamma subdivision. It contains a trypsin domain and a single copy of PDZ domain (in contrast to DegP with two copies). A critical role of this DegS is to sense stress in the periplasm and partially degrade an inhibitor of sigma(E).
Probab=100.00  E-value=1.6e-47  Score=398.05  Aligned_cols=296  Identities=25%  Similarity=0.390  Sum_probs=249.7

Q ss_pred             ChhhhhhhcccCCCeEEEEeeeeCCCCCCccccCCCcceEEEEEEEe-CCEEEecccccCCCCeEEEEEcCCCcEEEEEE
Q 009784          111 VEPGVARVVPAMDAVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIG-GRRVLTNAHSVEHYTQVKLKKRGSDTKYLATV  189 (526)
Q Consensus       111 ~~~~v~~~~~~~~SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~-~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~v  189 (526)
                      +.+.++++.   +|||.|.+.....+.   + ......+.||||+|+ +||||||+|||.++..+.|++. ||+.++|++
T Consensus        47 ~~~~~~~~~---psVV~I~~~~~~~~~---~-~~~~~~~~GSG~vi~~~G~IlTn~HVV~~~~~i~V~~~-dg~~~~a~v  118 (351)
T TIGR02038        47 FNKAVRRAA---PAVVNIYNRSISQNS---L-NQLSIQGLGSGVIMSKEGYILTNYHVIKKADQIVVALQ-DGRKFEAEL  118 (351)
T ss_pred             HHHHHHhcC---CcEEEEEeEeccccc---c-ccccccceEEEEEEeCCeEEEecccEeCCCCEEEEEEC-CCCEEEEEE
Confidence            344455555   599999986543321   1 112345689999999 7899999999999999999997 899999999


Q ss_pred             EEeccCCCeEEEEecccccccCceeeecCCC--CcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEc
Q 009784          190 LAIGTECDIAMLTVEDDEFWEGVLPVEFGEL--PALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQID  267 (526)
Q Consensus       190 v~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~--~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~d  267 (526)
                      ++.|+.+||||||++..    .+++++++++  .+.|++|+++|||.+... +++.|+|+...+..... .....+||+|
T Consensus       119 v~~d~~~DlAvlkv~~~----~~~~~~l~~s~~~~~G~~V~aiG~P~~~~~-s~t~GiIs~~~r~~~~~-~~~~~~iqtd  192 (351)
T TIGR02038       119 VGSDPLTDLAVLKIEGD----NLPTIPVNLDRPPHVGDVVLAIGNPYNLGQ-TITQGIISATGRNGLSS-VGRQNFIQTD  192 (351)
T ss_pred             EEecCCCCEEEEEecCC----CCceEeccCcCccCCCCEEEEEeCCCCCCC-cEEEEEEEeccCcccCC-CCcceEEEEC
Confidence            99999999999999976    3677778754  478999999999998775 89999999987643321 2234689999


Q ss_pred             ccCCCCCCCCeeecCCCeEEEEEeeccccC---ccccccccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHHh
Q 009784          268 AAINSGNSGGPAFNDKGKCVGIAFQSLKHE---DVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRV  344 (526)
Q Consensus       268 a~i~~G~SGGPlvn~~G~VVGI~~~~~~~~---~~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~  344 (526)
                      +.+++|||||||+|.+|+||||+++.+...   ...+++|+||++.+++++++|+++|++. ++|||+.++++ ++..++
T Consensus       193 a~i~~GnSGGpl~n~~G~vIGI~~~~~~~~~~~~~~g~~faIP~~~~~~vl~~l~~~g~~~-r~~lGv~~~~~-~~~~~~  270 (351)
T TIGR02038       193 AAINAGNSGGALINTNGELVGINTASFQKGGDEGGEGINFAIPIKLAHKIMGKIIRDGRVI-RGYIGVSGEDI-NSVVAQ  270 (351)
T ss_pred             CccCCCCCcceEECCCCeEEEEEeeeecccCCCCccceEEEecHHHHHHHHHHHhhcCccc-ceEeeeEEEEC-CHHHHH
Confidence            999999999999999999999998765432   2367899999999999999999999998 99999999998 788888


Q ss_pred             hhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEE
Q 009784          345 AMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKI  423 (526)
Q Consensus       345 ~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~  423 (526)
                      .+|++ ...|++|.+|.++|||++ ||++||+|++|||++|.++.++.          ..+...++|+++.++|+|+|+.
T Consensus       271 ~lgl~-~~~Gv~V~~V~~~spA~~aGL~~GDvI~~Ing~~V~s~~dl~----------~~l~~~~~g~~v~l~v~R~g~~  339 (351)
T TIGR02038       271 GLGLP-DLRGIVITGVDPNGPAARAGILVRDVILKYDGKDVIGAEELM----------DRIAETRPGSKVMVTVLRQGKQ  339 (351)
T ss_pred             hcCCC-ccccceEeecCCCChHHHCCCCCCCEEEEECCEEcCCHHHHH----------HHHHhcCCCCEEEEEEEECCEE
Confidence            99997 357999999999999999 99999999999999999998864          6666667899999999999999


Q ss_pred             EEEEEEeccc
Q 009784          424 LNFNITLATH  433 (526)
Q Consensus       424 ~~~~v~l~~~  433 (526)
                      +++++++.+.
T Consensus       340 ~~~~v~l~~~  349 (351)
T TIGR02038       340 LELPVTIDEK  349 (351)
T ss_pred             EEEEEEecCC
Confidence            9999988654


No 5  
>PRK10898 serine endoprotease; Provisional
Probab=100.00  E-value=7.6e-47  Score=392.79  Aligned_cols=297  Identities=23%  Similarity=0.360  Sum_probs=247.9

Q ss_pred             ChhhhhhhcccCCCeEEEEeeeeCCCCCCccccCCCcceEEEEEEEe-CCEEEecccccCCCCeEEEEEcCCCcEEEEEE
Q 009784          111 VEPGVARVVPAMDAVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIG-GRRVLTNAHSVEHYTQVKLKKRGSDTKYLATV  189 (526)
Q Consensus       111 ~~~~v~~~~~~~~SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~-~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~v  189 (526)
                      ..+.++++.+   |||.|.+.......    .......+.||||+|+ +||||||+|||.++..+.|++. ||+.++|++
T Consensus        47 ~~~~~~~~~p---svV~v~~~~~~~~~----~~~~~~~~~GSGfvi~~~G~IlTn~HVv~~a~~i~V~~~-dg~~~~a~v  118 (353)
T PRK10898         47 YNQAVRRAAP---AVVNVYNRSLNSTS----HNQLEIRTLGSGVIMDQRGYILTNKHVINDADQIIVALQ-DGRVFEALL  118 (353)
T ss_pred             HHHHHHHhCC---cEEEEEeEeccccC----cccccccceeeEEEEeCCeEEEecccEeCCCCEEEEEeC-CCCEEEEEE
Confidence            3445555555   99999986643221    1222344789999999 7899999999999999999997 899999999


Q ss_pred             EEeccCCCeEEEEecccccccCceeeecCCC--CcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEc
Q 009784          190 LAIGTECDIAMLTVEDDEFWEGVLPVEFGEL--PALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQID  267 (526)
Q Consensus       190 v~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~--~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~d  267 (526)
                      ++.|+.+||||||++..    .+++++++++  .+.|++|+++|||.+... +++.|+|+...+..... .....+||+|
T Consensus       119 v~~d~~~DlAvl~v~~~----~l~~~~l~~~~~~~~G~~V~aiG~P~g~~~-~~t~Giis~~~r~~~~~-~~~~~~iqtd  192 (353)
T PRK10898        119 VGSDSLTDLAVLKINAT----NLPVIPINPKRVPHIGDVVLAIGNPYNLGQ-TITQGIISATGRIGLSP-TGRQNFLQTD  192 (353)
T ss_pred             EEEcCCCCEEEEEEcCC----CCCeeeccCcCcCCCCCEEEEEeCCCCcCC-CcceeEEEeccccccCC-ccccceEEec
Confidence            99999999999999875    4677788764  468999999999998765 79999999877643221 1223579999


Q ss_pred             ccCCCCCCCCeeecCCCeEEEEEeeccccCc----cccccccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHH
Q 009784          268 AAINSGNSGGPAFNDKGKCVGIAFQSLKHED----VENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLR  343 (526)
Q Consensus       268 a~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~----~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~  343 (526)
                      +++++|||||||+|.+|+||||+++.+...+    ..+++|+||++.+++++++|+++|++. ++|||+..+.+ ++..+
T Consensus       193 a~i~~GnSGGPl~n~~G~vvGI~~~~~~~~~~~~~~~g~~faIP~~~~~~~~~~l~~~G~~~-~~~lGi~~~~~-~~~~~  270 (353)
T PRK10898        193 ASINHGNSGGALVNSLGELMGINTLSFDKSNDGETPEGIGFAIPTQLATKIMDKLIRDGRVI-RGYIGIGGREI-APLHA  270 (353)
T ss_pred             cccCCCCCcceEECCCCeEEEEEEEEecccCCCCcccceEEEEchHHHHHHHHHHhhcCccc-ccccceEEEEC-CHHHH
Confidence            9999999999999999999999998764322    257899999999999999999999998 99999999988 56666


Q ss_pred             hhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCE
Q 009784          344 VAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK  422 (526)
Q Consensus       344 ~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~  422 (526)
                      ..++++ ...|++|.+|.++|||++ ||++||+|++|||++|.++.++.          ..+....+|++++++|+|+|+
T Consensus       271 ~~~~~~-~~~Gv~V~~V~~~spA~~aGL~~GDvI~~Ing~~V~s~~~l~----------~~l~~~~~g~~v~l~v~R~g~  339 (353)
T PRK10898        271 QGGGID-QLQGIVVNEVSPDGPAAKAGIQVNDLIISVNNKPAISALETM----------DQVAEIRPGSVIPVVVMRDDK  339 (353)
T ss_pred             HhcCCC-CCCeEEEEEECCCChHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHhcCCCCEEEEEEEECCE
Confidence            677775 347999999999999999 99999999999999999988864          666666789999999999999


Q ss_pred             EEEEEEEecccc
Q 009784          423 ILNFNITLATHR  434 (526)
Q Consensus       423 ~~~~~v~l~~~~  434 (526)
                      .+++++++.+.+
T Consensus       340 ~~~~~v~l~~~p  351 (353)
T PRK10898        340 QLTLQVTIQEYP  351 (353)
T ss_pred             EEEEEEEeccCC
Confidence            999999887653


No 6  
>COG0265 DegQ Trypsin-like serine proteases, typically periplasmic, contain C-terminal PDZ domain [Posttranslational modification, protein turnover, chaperones]
Probab=100.00  E-value=8.5e-37  Score=317.92  Aligned_cols=300  Identities=26%  Similarity=0.389  Sum_probs=249.6

Q ss_pred             ChhhhhhhcccCCCeEEEEeeeeCCCCC-CccccCCC-cceEEEEEEEe-CCEEEecccccCCCCeEEEEEcCCCcEEEE
Q 009784          111 VEPGVARVVPAMDAVVKVFCVHTEPNFS-LPWQRKRQ-YSSSSSGFAIG-GRRVLTNAHSVEHYTQVKLKKRGSDTKYLA  187 (526)
Q Consensus       111 ~~~~v~~~~~~~~SVV~I~~~~~~~~~~-~P~~~~~~-~~~~GSGfvI~-~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a  187 (526)
                      +...++++.+   +||.|.......... ++-..... ..+.||||+++ +|||+||.||+.++.++.+.+. +|+.+++
T Consensus        35 ~~~~~~~~~~---~vV~~~~~~~~~~~~~~~~~~~~~~~~~~gSg~i~~~~g~ivTn~hVi~~a~~i~v~l~-dg~~~~a  110 (347)
T COG0265          35 FATAVEKVAP---AVVSIATGLTAKLRSFFPSDPPLRSAEGLGSGFIISSDGYIVTNNHVIAGAEEITVTLA-DGREVPA  110 (347)
T ss_pred             HHHHHHhcCC---cEEEEEeeeeecchhcccCCcccccccccccEEEEcCCeEEEecceecCCcceEEEEeC-CCCEEEE
Confidence            3445555555   999999876544200 00000000 14789999999 9999999999999999999996 9999999


Q ss_pred             EEEEeccCCCeEEEEecccccccCceeeecCCCC--cCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEE
Q 009784          188 TVLAIGTECDIAMLTVEDDEFWEGVLPVEFGELP--ALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQ  265 (526)
Q Consensus       188 ~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~~--~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~  265 (526)
                      ++++.|+..|+|+||++...   .++.+.++++.  .+|++++++|+|++... +++.|+|+...+...........+||
T Consensus       111 ~~vg~d~~~dlavlki~~~~---~~~~~~~~~s~~l~vg~~v~aiGnp~g~~~-tvt~Givs~~~r~~v~~~~~~~~~Iq  186 (347)
T COG0265         111 KLVGKDPISDLAVLKIDGAG---GLPVIALGDSDKLRVGDVVVAIGNPFGLGQ-TVTSGIVSALGRTGVGSAGGYVNFIQ  186 (347)
T ss_pred             EEEecCCccCEEEEEeccCC---CCceeeccCCCCcccCCEEEEecCCCCccc-ceeccEEeccccccccCcccccchhh
Confidence            99999999999999999875   26777888765  46899999999999665 89999999998762222122556899


Q ss_pred             EcccCCCCCCCCeeecCCCeEEEEEeeccccCc-cccccccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHHh
Q 009784          266 IDAAINSGNSGGPAFNDKGKCVGIAFQSLKHED-VENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRV  344 (526)
Q Consensus       266 ~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~-~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~  344 (526)
                      +|+++++||||||++|.+|++|||+++.....+ ..+++|+||++.++.+++++.+.|++. ++++|+.+.++ +.+.+ 
T Consensus       187 tdAain~gnsGgpl~n~~g~~iGint~~~~~~~~~~gigfaiP~~~~~~v~~~l~~~G~v~-~~~lgv~~~~~-~~~~~-  263 (347)
T COG0265         187 TDAAINPGNSGGPLVNIDGEVVGINTAIIAPSGGSSGIGFAIPVNLVAPVLDELISKGKVV-RGYLGVIGEPL-TADIA-  263 (347)
T ss_pred             cccccCCCCCCCceEcCCCcEEEEEEEEecCCCCcceeEEEecHHHHHHHHHHHHHcCCcc-ccccceEEEEc-ccccc-
Confidence            999999999999999999999999999886543 456899999999999999999988887 99999999988 55544 


Q ss_pred             hhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEE
Q 009784          345 AMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKI  423 (526)
Q Consensus       345 ~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~  423 (526)
                       +|++ ...|++|.+|.+++||++ |++.||+|+++||+++.+..++.          ..+....+|+++.++++|+|+.
T Consensus       264 -~g~~-~~~G~~V~~v~~~spa~~agi~~Gdii~~vng~~v~~~~~l~----------~~v~~~~~g~~v~~~~~r~g~~  331 (347)
T COG0265         264 -LGLP-VAAGAVVLGVLPGSPAAKAGIKAGDIITAVNGKPVASLSDLV----------AAVASNRPGDEVALKLLRGGKE  331 (347)
T ss_pred             -cCCC-CCCceEEEecCCCChHHHcCCCCCCEEEEECCEEccCHHHHH----------HHHhccCCCCEEEEEEEECCEE
Confidence             7776 678899999999999999 99999999999999999998865          6677777999999999999999


Q ss_pred             EEEEEEeccc
Q 009784          424 LNFNITLATH  433 (526)
Q Consensus       424 ~~~~v~l~~~  433 (526)
                      +++.+++.+.
T Consensus       332 ~~~~v~l~~~  341 (347)
T COG0265         332 RELAVTLGDR  341 (347)
T ss_pred             EEEEEEecCc
Confidence            9999998773


No 7  
>KOG1421 consensus Predicted signaling-associated protein (contains a PDZ domain) [General function prediction only]
Probab=99.98  E-value=1.7e-31  Score=281.02  Aligned_cols=322  Identities=17%  Similarity=0.264  Sum_probs=266.3

Q ss_pred             hhhhhhhcccCCCeEEEEeeeeCCCCCCccccCCCcceEEEEEEEe--CCEEEecccccCCCCe-EEEEEcCCCcEEEEE
Q 009784          112 EPGVARVVPAMDAVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIG--GRRVLTNAHSVEHYTQ-VKLKKRGSDTKYLAT  188 (526)
Q Consensus       112 ~~~v~~~~~~~~SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~--~g~ILT~aHvV~~~~~-i~V~~~~~g~~~~a~  188 (526)
                      ..|...++.+.+|||.|.+.....     |.......+.++||+++  .||||||+|++..... -.+.+. +..+.+.-
T Consensus        52 e~w~~~ia~VvksvVsI~~S~v~~-----fdtesag~~~atgfvvd~~~gyiLtnrhvv~pgP~va~avf~-n~ee~ei~  125 (955)
T KOG1421|consen   52 EDWRNTIANVVKSVVSIRFSAVRA-----FDTESAGESEATGFVVDKKLGYILTNRHVVAPGPFVASAVFD-NHEEIEIY  125 (955)
T ss_pred             hhhhhhhhhhcccEEEEEehheee-----cccccccccceeEEEEecccceEEEeccccCCCCceeEEEec-ccccCCcc
Confidence            356666677777999999877543     34455667889999999  6999999999985544 455554 77788888


Q ss_pred             EEEeccCCCeEEEEeccccc-ccCceeeecCC-CCcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCc-----eee
Q 009784          189 VLAIGTECDIAMLTVEDDEF-WEGVLPVEFGE-LPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGS-----TEL  261 (526)
Q Consensus       189 vv~~d~~~DlAlLkv~~~~~-~~~~~pl~l~~-~~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~-----~~~  261 (526)
                      .++.|+.||+.+++.+++.+ +..+..+++.. ..++|.++.++|+..+. ..++..|.++++++.....+.     .+.
T Consensus       126 pvyrDpVhdfGf~r~dps~ir~s~vt~i~lap~~akvgseirvvgNDagE-klsIlagflSrldr~apdyg~~~yndfnT  204 (955)
T KOG1421|consen  126 PVYRDPVHDFGFFRYDPSTIRFSIVTEICLAPELAKVGSEIRVVGNDAGE-KLSILAGFLSRLDRNAPDYGEDTYNDFNT  204 (955)
T ss_pred             cccCCchhhcceeecChhhcceeeeeccccCccccccCCceEEecCCccc-eEEeehhhhhhccCCCccccccccccccc
Confidence            89999999999999998754 33456666663 45789999999998774 458999999999876544433     222


Q ss_pred             eEEEEcccCCCCCCCCeeecCCCeEEEEEeeccccCccccccccccHHHHHHHHHHHHHcCceeeccccCceeeeccChh
Q 009784          262 LGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPD  341 (526)
Q Consensus       262 ~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~  341 (526)
                      .++|..+....|.||+|++|.+|..|.++.++.   .....+|++|++.+.+.|..++.+..++ |+.|.++|... ..+
T Consensus       205 fy~QaasstsggssgspVv~i~gyAVAl~agg~---~ssas~ffLpLdrV~RaL~clq~n~PIt-RGtLqvefl~k-~~d  279 (955)
T KOG1421|consen  205 FYIQAASSTSGGSSGSPVVDIPGYAVALNAGGS---ISSASDFFLPLDRVVRALRCLQNNTPIT-RGTLQVEFLHK-LFD  279 (955)
T ss_pred             eeeeehhcCCCCCCCCceecccceEEeeecCCc---ccccccceeeccchhhhhhhhhcCCCcc-cceEEEEEehh-hhH
Confidence            368888899999999999999999999998765   4567799999999999999999877777 99999999988 889


Q ss_pred             HHhhhcCCC-----------Cccce-EEEEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCC
Q 009784          342 LRVAMSMKA-----------DQKGV-RIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYT  409 (526)
Q Consensus       342 ~~~~lgl~~-----------~~~Gv-~V~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~  409 (526)
                      .++.+||+.           ...|+ +|..|.++|||++.|++||++++||+.-+.++..+.          ..++ ...
T Consensus       280 e~rrlGL~sE~eqv~r~k~P~~tgmLvV~~vL~~gpa~k~Le~GDillavN~t~l~df~~l~----------~iLD-egv  348 (955)
T KOG1421|consen  280 ECRRLGLSSEWEQVVRTKFPERTGMLVVETVLPEGPAEKKLEPGDILLAVNSTCLNDFEALE----------QILD-EGV  348 (955)
T ss_pred             HHHhcCCcHHHHHHHHhcCcccceeEEEEEeccCCchhhccCCCcEEEEEcceehHHHHHHH----------HHHh-hcc
Confidence            999999975           23454 567889999999999999999999999999988763          4444 458


Q ss_pred             CCEEEEEEEECCEEEEEEEEecccccccCCCCCCCCCceEEEeeEEEEeccccc
Q 009784          410 GDSAAVKVLRDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLYLI  463 (526)
Q Consensus       410 G~~v~l~v~R~G~~~~~~v~l~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt~~~  463 (526)
                      |+.++|+|+|+|+++++++++...+...|.       ||+.++|++||+++|+.
T Consensus       349 gk~l~LtI~Rggqelel~vtvqdlh~itp~-------R~levcGav~hdlsyq~  395 (955)
T KOG1421|consen  349 GKNLELTIQRGGQELELTVTVQDLHGITPD-------RFLEVCGAVFHDLSYQL  395 (955)
T ss_pred             CceEEEEEEeCCEEEEEEEEeccccCCCCc-------eEEEEcceEecCCCHHH
Confidence            999999999999999999999999988887       99999999999999763


No 8  
>KOG1320 consensus Serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=99.96  E-value=1.2e-29  Score=265.52  Aligned_cols=372  Identities=40%  Similarity=0.557  Sum_probs=328.2

Q ss_pred             hcccCCCeEEEEeeeeCCCCCCccccCCCcceEEEEEEEeCCEEEecccccC---CCCeEEEEEcCCCcEEEEEEEEecc
Q 009784          118 VVPAMDAVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIGGRRVLTNAHSVE---HYTQVKLKKRGSDTKYLATVLAIGT  194 (526)
Q Consensus       118 ~~~~~~SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~~g~ILT~aHvV~---~~~~i~V~~~~~g~~~~a~vv~~d~  194 (526)
                      ......|++.+.+....+.+..||+...+....|+||.+....++||+|++.   +...+.+...+.-+.|.+++...-.
T Consensus        56 ~~~~~~s~~~v~~~~~~~~~~~pw~~~~q~~~~~s~f~i~~~~lltn~~~v~~~~~~~~v~v~~~gs~~k~~~~v~~~~~  135 (473)
T KOG1320|consen   56 VDLALQSVVKVFSVSTEPSSVLPWQRTRQFSSGGSGFAIYGKKLLTNAHVVAPNNDHKFVTVKKHGSPRKYKAFVAAVFE  135 (473)
T ss_pred             ccccccceeEEEeecccccccCcceeeehhcccccchhhcccceeecCccccccccccccccccCCCchhhhhhHHHhhh
Confidence            3445569999999999999999999999888999999999999999999999   6667777766566788999998889


Q ss_pred             CCCeEEEEecccccccCceeeecCCCCcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCC
Q 009784          195 ECDIAMLTVEDDEFWEGVLPVEFGELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGN  274 (526)
Q Consensus       195 ~~DlAlLkv~~~~~~~~~~pl~l~~~~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~  274 (526)
                      +.|+|+|.++..+|+....|+++++.+.+.+.++++|    ++..+++.|.|++.....+..+......+++|+++++|+
T Consensus       136 ~cd~Avv~Ie~~~f~~~~~~~e~~~ip~l~~S~~Vv~----gd~i~VTnghV~~~~~~~y~~~~~~l~~vqi~aa~~~~~  211 (473)
T KOG1320|consen  136 ECDLAVVYIESEEFWKGMNPFELGDIPSLNGSGFVVG----GDGIIVTNGHVVRVEPRIYAHSSTVLLRVQIDAAIGPGN  211 (473)
T ss_pred             cccceEEEEeeccccCCCcccccCCCcccCccEEEEc----CCcEEEEeeEEEEEEeccccCCCcceeeEEEEEeecCCc
Confidence            9999999999999988888999999999999999999    345699999999999888888877788899999999999


Q ss_pred             CCCeeecCCCeEEEEEeeccccCccccccccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHHhhhcCCCCccc
Q 009784          275 SGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKG  354 (526)
Q Consensus       275 SGGPlvn~~G~VVGI~~~~~~~~~~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~~lgl~~~~~G  354 (526)
                      ||+|.+...+++.|+.+...+..+  ++++.||.-.+.+++....+.+.+.++++++...+.+.+.+.++.+.|..+ .|
T Consensus       212 s~ep~i~g~d~~~gvA~l~ik~~~--~i~~~i~~~~~~~~~~G~~~~a~~~~f~~~nt~t~g~vs~~~R~~~~lg~~-~g  288 (473)
T KOG1320|consen  212 SGEPVIVGVDKVAGVAFLKIKTPE--NILYVIPLGVSSHFRTGVEVSAIGNGFGLLNTLTQGMVSGQLRKSFKLGLE-TG  288 (473)
T ss_pred             cCCCeEEccccccceEEEEEecCC--cccceeecceeeeecccceeeccccCceeeeeeeecccccccccccccCcc-cc
Confidence            999999888899999998874322  789999999999999998888988899999999999999999999999877 99


Q ss_pred             eEEEEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEecccc
Q 009784          355 VRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHR  434 (526)
Q Consensus       355 v~V~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~~~  434 (526)
                      +.+.++.+.+.|.+-++.||+|+.+||+.|.    +.++..+|+.|++.+..+.++|++.+.+.|.+   ++++.++...
T Consensus       289 ~~i~~~~qtd~ai~~~nsg~~ll~~DG~~Ig----Vn~~~~~ri~~~~~iSf~~p~d~vl~~v~r~~---e~~~~lr~~~  361 (473)
T KOG1320|consen  289 VLISKINQTDAAINPGNSGGPLLNLDGEVIG----VNTRKVTRIGFSHGISFKIPIDTVLVIVLRLG---EFQISLRPVK  361 (473)
T ss_pred             eeeeeecccchhhhcccCCCcEEEecCcEee----eeeeeeEEeeccccceeccCchHhhhhhhhhh---hhceeecccc
Confidence            9999999999999988999999999999998    55667889999999999999999999999998   6777788888


Q ss_pred             cccCCCCCCCCCceEEEeeEEEEeccc--cce-----eeeeeeecch--hhhccccccceeeeecccchhhhHHHHHH
Q 009784          435 RLIPSHNKGRPPSYYIIAGFVFSRCLY--LIS-----VLSMERIMNM--KLRSSFWTSSCIQCHNCQMSSLLWCLRCL  503 (526)
Q Consensus       435 ~~~p~~~~~~~p~~~i~gG~~f~~lt~--~~~-----~~~~~~i~~~--~~~sg~~~~~~~~~~~~~~~~~~~~~~~~  503 (526)
                      .+.|.+.+...+.|++++|++|++++.  ...     .+.+..+++.  ..+.|+..+|.+...|.+...++-+|+-+
T Consensus       362 ~~~p~~~~~g~~s~~i~~g~vf~~~~~~~~~~~~~~q~v~is~Vlp~~~~~~~~~~~g~~V~~vng~~V~n~~~l~~~  439 (473)
T KOG1320|consen  362 PLVPVHQYIGLPSYYIFAGLVFVPLTKSYIFPSGVVQLVLVSQVLPGSINGGYGLKPGDQVVKVNGKPVKNLKHLYEL  439 (473)
T ss_pred             CcccccccCCceeEEEecceEEeecCCCccccccceeEEEEEEeccCCCcccccccCCCEEEEECCEEeechHHHHHH
Confidence            888889999999999999999999973  222     2445555553  35678889999999999999999988764


No 9  
>KOG1320 consensus Serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=99.93  E-value=8.3e-25  Score=229.33  Aligned_cols=301  Identities=21%  Similarity=0.204  Sum_probs=225.6

Q ss_pred             hcccCCCeEEEEeeeeCCCCCCccccCCCcceEEEEEEEe-CCEEEecccccCCCC-----------eEEEEEcC-CCcE
Q 009784          118 VVPAMDAVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIG-GRRVLTNAHSVEHYT-----------QVKLKKRG-SDTK  184 (526)
Q Consensus       118 ~~~~~~SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~-~g~ILT~aHvV~~~~-----------~i~V~~~~-~g~~  184 (526)
                      ..+...+||.|....--.. ..|+....-....|||||++ +|+++||+||+....           .+.|.... .+..
T Consensus       134 ~~~cd~Avv~Ie~~~f~~~-~~~~e~~~ip~l~~S~~Vv~gd~i~VTnghV~~~~~~~y~~~~~~l~~vqi~aa~~~~~s  212 (473)
T KOG1320|consen  134 FEECDLAVVYIESEEFWKG-MNPFELGDIPSLNGSGFVVGGDGIIVTNGHVVRVEPRIYAHSSTVLLRVQIDAAIGPGNS  212 (473)
T ss_pred             hhcccceEEEEeeccccCC-CcccccCCCcccCccEEEEcCCcEEEEeeEEEEEEeccccCCCcceeeEEEEEeecCCcc
Confidence            3444558999887432111 12455555666789999999 999999999997432           26666652 2488


Q ss_pred             EEEEEEEeccCCCeEEEEecccccccCceeeecCC--CCcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCC----c
Q 009784          185 YLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGE--LPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHG----S  258 (526)
Q Consensus       185 ~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~--~~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~----~  258 (526)
                      +++.+++.|+..|+|+++++.+.  ....+++++-  ....|+++.++|.|++..+ +++.|+++...|..+.-+    .
T Consensus       213 ~ep~i~g~d~~~gvA~l~ik~~~--~i~~~i~~~~~~~~~~G~~~~a~~~~f~~~n-t~t~g~vs~~~R~~~~lg~~~g~  289 (473)
T KOG1320|consen  213 GEPVIVGVDKVAGVAFLKIKTPE--NILYVIPLGVSSHFRTGVEVSAIGNGFGLLN-TLTQGMVSGQLRKSFKLGLETGV  289 (473)
T ss_pred             CCCeEEccccccceEEEEEecCC--cccceeecceeeeecccceeeccccCceeee-eeeecccccccccccccCcccce
Confidence            89999999999999999997553  1356666664  3456899999999999887 799999998877655422    3


Q ss_pred             eeeeEEEEcccCCCCCCCCeeecCCCeEEEEEeeccccC-ccccccccccHHHHHHHHHHHHHcCc---ee-----eccc
Q 009784          259 TELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHE-DVENIGYVIPTPVIMHFIQDYEKNGA---YT-----GFPL  329 (526)
Q Consensus       259 ~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~-~~~~~~~aIP~~~i~~~l~~l~~~g~---~~-----~~~~  329 (526)
                      ....++|+|++++.|+||||++|.+|++||++++..... -..+++|++|.+.++.++.+..+...   ..     .+.|
T Consensus       290 ~i~~~~qtd~ai~~~nsg~~ll~~DG~~IgVn~~~~~ri~~~~~iSf~~p~d~vl~~v~r~~e~~~~lr~~~~~~p~~~~  369 (473)
T KOG1320|consen  290 LISKINQTDAAINPGNSGGPLLNLDGEVIGVNTRKVTRIGFSHGISFKIPIDTVLVIVLRLGEFQISLRPVKPLVPVHQY  369 (473)
T ss_pred             eeeeecccchhhhcccCCCcEEEecCcEeeeeeeeeEEeeccccceeccCchHhhhhhhhhhhhceeeccccCccccccc
Confidence            445689999999999999999999999999998876322 23678999999999988888743221   11     1346


Q ss_pred             cCceeeeccChhH----HhhhcCC-CCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhh
Q 009784          330 LGVEWQKMENPDL----RVAMSMK-ADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYL  403 (526)
Q Consensus       330 LGi~~~~~~~~~~----~~~lgl~-~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~  403 (526)
                      +|.....+...-.    .+.+-.+ ...++++|.+|.+++++.. ++++||+|++|||++|.+..+|.          ++
T Consensus       370 ~g~~s~~i~~g~vf~~~~~~~~~~~~~~q~v~is~Vlp~~~~~~~~~~~g~~V~~vng~~V~n~~~l~----------~~  439 (473)
T KOG1320|consen  370 IGLPSYYIFAGLVFVPLTKSYIFPSGVVQLVLVSQVLPGSINGGYGLKPGDQVVKVNGKPVKNLKHLY----------EL  439 (473)
T ss_pred             CCceeEEEecceEEeecCCCccccccceeEEEEEEeccCCCcccccccCCCEEEEECCEEeechHHHH----------HH
Confidence            6666554421111    1111122 2346899999999999999 99999999999999999999975          78


Q ss_pred             hhhcCCCCEEEEEEEECCEEEEEEEEecc
Q 009784          404 VSQKYTGDSAAVKVLRDSKILNFNITLAT  432 (526)
Q Consensus       404 l~~~~~G~~v~l~v~R~G~~~~~~v~l~~  432 (526)
                      +....++++|.+..+|+.|..++.+....
T Consensus       440 i~~~~~~~~v~vl~~~~~e~~tl~Il~~~  468 (473)
T KOG1320|consen  440 IEECSTEDKVAVLDRRSAEDATLEILPEH  468 (473)
T ss_pred             HHhcCcCceEEEEEecCccceeEEecccc
Confidence            88888889999999999999999887654


No 10 
>KOG1421 consensus Predicted signaling-associated protein (contains a PDZ domain) [General function prediction only]
Probab=99.79  E-value=2e-17  Score=175.49  Aligned_cols=327  Identities=13%  Similarity=0.122  Sum_probs=240.0

Q ss_pred             CeEEEEeeeeCCCCCCccccCCCcceEEEEEEEe--CCEEEecccccC-CCCeEEEEEcCCCcEEEEEEEEeccCCCeEE
Q 009784          124 AVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIG--GRRVLTNAHSVE-HYTQVKLKKRGSDTKYLATVLAIGTECDIAM  200 (526)
Q Consensus       124 SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~--~g~ILT~aHvV~-~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAl  200 (526)
                      +.|.+.......-.++     ......|||.|++  +|++++++.+|. ++.+..|+.. +...++|.+.+.++..++|.
T Consensus       530 ~~~~v~~~~~~~l~g~-----s~~i~kgt~~i~d~~~g~~vvsr~~vp~d~~d~~vt~~-dS~~i~a~~~fL~~t~n~a~  603 (955)
T KOG1421|consen  530 CLVDVEPMMPVNLDGV-----SSDIYKGTALIMDTSKGLGVVSRSVVPSDAKDQRVTEA-DSDGIPANVSFLHPTENVAS  603 (955)
T ss_pred             hhhhheeceeeccccc-----hhhhhcCceEEEEccCCceeEecccCCchhhceEEeec-ccccccceeeEecCccceeE
Confidence            6666666554333221     1123579999999  799999999997 6778999987 78889999999999999999


Q ss_pred             EEecccccccCceeeecCCCC-cCCCcEEEEeeCCCCCceeEEEEEEece-----eeee-ccCCceeeeEEEEcccCCCC
Q 009784          201 LTVEDDEFWEGVLPVEFGELP-ALQDAVTVVGYPIGGDTISVTSGVVSRI-----EILS-YVHGSTELLGLQIDAAINSG  273 (526)
Q Consensus       201 Lkv~~~~~~~~~~pl~l~~~~-~~g~~V~~iG~p~~~~~~sv~~GiVs~~-----~~~~-~~~~~~~~~~i~~da~i~~G  273 (526)
                      +|+++..    ...++|.+.. ..|+++...|+....... .....|..+     .... ......+...|.+++.+.-+
T Consensus       604 ~kydp~~----~~~~kl~~~~v~~gD~~~f~g~~~~~r~l-taktsv~dvs~~~~ps~~~pr~r~~n~e~Is~~~nlsT~  678 (955)
T KOG1421|consen  604 FKYDPAL----EVQLKLTDTTVLRGDECTFEGFTEDLRAL-TAKTSVTDVSVVIIPSSVMPRFRATNLEVISFMDNLSTS  678 (955)
T ss_pred             eccChhH----hhhhccceeeEecCCceeEecccccchhh-cccceeeeeEEEEecCCCCcceeecceEEEEEecccccc
Confidence            9999874    3455565433 568999999998765431 111122221     1111 11223556778888887777


Q ss_pred             CCCCeeecCCCeEEEEEeeccccC-c--cccccccccHHHHHHHHHHHHHcCceeeccccCceeeeccChhHHhhhcCCC
Q 009784          274 NSGGPAFNDKGKCVGIAFQSLKHE-D--VENIGYVIPTPVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKA  350 (526)
Q Consensus       274 ~SGGPlvn~~G~VVGI~~~~~~~~-~--~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~~lgl~~  350 (526)
                      +--|-+.|.+|+|+|++...+.+. +  ...+-|.+.+..++++|++|+.++... ...+|++|..+ +...++.+|++.
T Consensus       679 c~sg~ltdddg~vvalwl~~~ge~~~~kd~~y~~gl~~~~~l~vl~rlk~g~~~r-p~i~~vef~~i-~laqar~lglp~  756 (955)
T KOG1421|consen  679 CLSGRLTDDDGEVVALWLSVVGEDVGGKDYTYKYGLSMSYILPVLERLKLGPSAR-PTIAGVEFSHI-TLAQARTLGLPS  756 (955)
T ss_pred             ccceEEECCCCeEEEEEeeeeccccCCceeEEEeccchHHHHHHHHHHhcCCCCC-ceeeccceeeE-EeehhhccCCCH
Confidence            777789999999999998776543 1  223456788899999999998777765 66789999999 888899999985


Q ss_pred             ------------CccceEEEEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEE
Q 009784          351 ------------DQKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVL  418 (526)
Q Consensus       351 ------------~~~Gv~V~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~  418 (526)
                                  ..+-.+|++|.+..+-  -|..||||+++||+-|+...||.          + +.      .++.+|+
T Consensus       757 e~imk~e~es~~~~ql~~ishv~~~~~k--il~~gdiilsvngk~itr~~dl~----------d-~~------eid~~il  817 (955)
T KOG1421|consen  757 EFIMKSEEESTIPRQLYVISHVRPLLHK--ILGVGDIILSVNGKMITRLSDLH----------D-FE------EIDAVIL  817 (955)
T ss_pred             HHHhhhhhcCCCcceEEEEEeeccCccc--ccccccEEEEecCeEEeeehhhh----------h-hh------hhheeee
Confidence                        2345678888876543  59999999999999999998874          2 21      6899999


Q ss_pred             ECCEEEEEEEEecccccccCCCCCCCCCceEEEeeEEEEecc---------ccceeeeeeeecc-hhhhccccccceeee
Q 009784          419 RDSKILNFNITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCL---------YLISVLSMERIMN-MKLRSSFWTSSCIQC  488 (526)
Q Consensus       419 R~G~~~~~~v~l~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt---------~~~~~~~~~~i~~-~~~~sg~~~~~~~~~  488 (526)
                      |+|..+++++++.+..  ++.       |.++|.|..+|+--         .-++++.+-|-.. .+++ ++.+-..|..
T Consensus       818 rdg~~~~ikipt~p~~--et~-------r~vi~~gailq~ph~av~~q~edlp~gvyvt~rg~gspalq-~l~aa~fita  887 (955)
T KOG1421|consen  818 RDGIEMEIKIPTYPEY--ETS-------RAVIWMGAILQPPHSAVFEQVEDLPEGVYVTSRGYGSPALQ-MLRAAHFITA  887 (955)
T ss_pred             ecCcEEEEEecccccc--ccc-------eEEEEEeccccCchHHHHHHHhccCCceEEeecccCChhHh-hcchheeEEE
Confidence            9999999999877654  333       89999999887632         2356666666555 4564 8888888888


Q ss_pred             eccc
Q 009784          489 HNCQ  492 (526)
Q Consensus       489 ~~~~  492 (526)
                      .|.-
T Consensus       888 vng~  891 (955)
T KOG1421|consen  888 VNGH  891 (955)
T ss_pred             eccc
Confidence            8773


No 11 
>PF13365 Trypsin_2:  Trypsin-like peptidase domain; PDB: 1Y8T_A 2Z9I_A 3QO6_A 1L1J_A 1QY6_A 2O8L_A 3OTP_E 2ZLE_I 1KY9_A 3CS0_A ....
Probab=99.63  E-value=6e-15  Score=128.71  Aligned_cols=108  Identities=33%  Similarity=0.488  Sum_probs=72.7

Q ss_pred             EEEEEEe-CCEEEecccccC--------CCCeEEEEEcCCCcEEE--EEEEEeccC-CCeEEEEecccccccCceeeecC
Q 009784          151 SSGFAIG-GRRVLTNAHSVE--------HYTQVKLKKRGSDTKYL--ATVLAIGTE-CDIAMLTVEDDEFWEGVLPVEFG  218 (526)
Q Consensus       151 GSGfvI~-~g~ILT~aHvV~--------~~~~i~V~~~~~g~~~~--a~vv~~d~~-~DlAlLkv~~~~~~~~~~pl~l~  218 (526)
                      ||||+|+ +|+||||+||+.        ....+.+... ++..+.  +++++.++. +|+|||+++..            
T Consensus         1 GTGf~i~~~g~ilT~~Hvv~~~~~~~~~~~~~~~~~~~-~~~~~~~~~~~~~~~~~~~D~All~v~~~------------   67 (120)
T PF13365_consen    1 GTGFLIGPDGYILTAAHVVEDWNDGKQPDNSSVEVVFP-DGRRVPPVAEVVYFDPDDYDLALLKVDPW------------   67 (120)
T ss_dssp             EEEEEEETTTEEEEEHHHHTCCTT--G-TCSEEEEEET-TSCEEETEEEEEEEETT-TTEEEEEESCE------------
T ss_pred             CEEEEEcCCceEEEchhheecccccccCCCCEEEEEec-CCCEEeeeEEEEEECCccccEEEEEEecc------------
Confidence            8999999 559999999998        4567888887 677777  999999999 99999999910            


Q ss_pred             CCCcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecCCCeEEEE
Q 009784          219 ELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGI  289 (526)
Q Consensus       219 ~~~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI  289 (526)
                               ...+..      ....+..........  .......+ +++.+.+|+|||||||.+|+||||
T Consensus        68 ---------~~~~~~------~~~~~~~~~~~~~~~--~~~~~~~~-~~~~~~~G~SGgpv~~~~G~vvGi  120 (120)
T PF13365_consen   68 ---------TGVGGG------VRVPGSTSGVSPTST--NDNRMLYI-TDADTRPGSSGGPVFDSDGRVVGI  120 (120)
T ss_dssp             ---------EEEEEE------EEEEEEEEEEEEEEE--EETEEEEE-ESSS-STTTTTSEEEETTSEEEEE
T ss_pred             ---------cceeee------eEeeeeccccccccC--cccceeEe-eecccCCCcEeHhEECCCCEEEeC
Confidence                     000000      000000011000000  00111124 899999999999999999999997


No 12 
>PF13180 PDZ_2:  PDZ domain; PDB: 2L97_A 1Y8T_A 2Z9I_A 1LCY_A 2PZD_B 2P3W_A 1VCW_C 1TE0_B 1SOZ_C 1SOT_C ....
Probab=99.50  E-value=4e-14  Score=116.57  Aligned_cols=81  Identities=33%  Similarity=0.535  Sum_probs=69.7

Q ss_pred             cccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhh
Q 009784          328 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ  406 (526)
Q Consensus       328 ~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~  406 (526)
                      ||||+.+....            ...|++|.+|.++|||++ |||+||+|++|||++|.+..++.          ..+..
T Consensus         1 ~~lGv~~~~~~------------~~~g~~V~~V~~~spA~~aGl~~GD~I~~ing~~v~~~~~~~----------~~l~~   58 (82)
T PF13180_consen    1 GGLGVTVQNLS------------DTGGVVVVSVIPGSPAAKAGLQPGDIILAINGKPVNSSEDLV----------NILSK   58 (82)
T ss_dssp             -E-SEEEEECS------------CSSSEEEEEESTTSHHHHTTS-TTEEEEEETTEESSSHHHHH----------HHHHC
T ss_pred             CEECeEEEEcc------------CCCeEEEEEeCCCCcHHHCCCCCCcEEEEECCEEcCCHHHHH----------HHHHh
Confidence            58999999872            246999999999999999 99999999999999999988864          77778


Q ss_pred             cCCCCEEEEEEEECCEEEEEEEEe
Q 009784          407 KYTGDSAAVKVLRDSKILNFNITL  430 (526)
Q Consensus       407 ~~~G~~v~l~v~R~G~~~~~~v~l  430 (526)
                      ..+|++++|+|+|+|+.++++++|
T Consensus        59 ~~~g~~v~l~v~R~g~~~~~~v~l   82 (82)
T PF13180_consen   59 GKPGDTVTLTVLRDGEELTVEVTL   82 (82)
T ss_dssp             SSTTSEEEEEEEETTEEEEEEEE-
T ss_pred             CCCCCEEEEEEEECCEEEEEEEEC
Confidence            889999999999999999999875


No 13 
>PF00089 Trypsin:  Trypsin;  InterPro: IPR001254 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine proteases belong to the MEROPS peptidase family S1 (chymotrypsin family, clan PA(S))and to peptidase family S6 (Hap serine peptidases). The chymotrypsin family is almost totally confined to animals, although trypsin-like enzymes are found in actinomycetes of the genera Streptomyces and Saccharopolyspora, and in the fungus Fusarium oxysporum []. The enzymes are inherently secreted, being synthesised with a signal peptide that targets them to the secretory pathway. Animal enzymes are either secreted directly, packaged into vesicles for regulated secretion, or are retained in leukocyte granules []. The Hap family, 'Haemophilus adhesion and penetration', are proteins that play a role in the interaction with human epithelial cells. The serine protease activity is localized at the N-terminal domain, whereas the binding domain is in the C-terminal region. ; GO: 0004252 serine-type endopeptidase activity, 0006508 proteolysis; PDB: 1SPJ_A 1A5I_A 2ZGH_A 2ZKS_A 2ZGJ_A 2ZGC_A 2ODP_A 2I6Q_A 2I6S_A 2ODQ_A ....
Probab=99.44  E-value=1.5e-12  Score=125.06  Aligned_cols=185  Identities=22%  Similarity=0.269  Sum_probs=118.2

Q ss_pred             eeeCCCCCCccccCCCc---ceEEEEEEEeCCEEEecccccCCCCeEEEEEcC------CC--cEEEEEEEEecc-----
Q 009784          131 VHTEPNFSLPWQRKRQY---SSSSSGFAIGGRRVLTNAHSVEHYTQVKLKKRG------SD--TKYLATVLAIGT-----  194 (526)
Q Consensus       131 ~~~~~~~~~P~~~~~~~---~~~GSGfvI~~g~ILT~aHvV~~~~~i~V~~~~------~g--~~~~a~vv~~d~-----  194 (526)
                      +......++||......   ...|+|++|++.+|||++||+.+..++.+.+..      ++  ..+..+-+..++     
T Consensus         4 g~~~~~~~~p~~v~i~~~~~~~~C~G~li~~~~vLTaahC~~~~~~~~v~~g~~~~~~~~~~~~~~~v~~~~~h~~~~~~   83 (220)
T PF00089_consen    4 GDPASPGEFPWVVSIRYSNGRFFCTGTLISPRWVLTAAHCVDGASDIKVRLGTYSIRNSDGSEQTIKVSKIIIHPKYDPS   83 (220)
T ss_dssp             SEECGTTSSTTEEEEEETTTEEEEEEEEEETTEEEEEGGGHTSGGSEEEEESESBTTSTTTTSEEEEEEEEEEETTSBTT
T ss_pred             CEECCCCCCCeEEEEeeCCCCeeEeEEecccccccccccccccccccccccccccccccccccccccccccccccccccc
Confidence            34455566677654322   468999999999999999999996667665431      22  234444443432     


Q ss_pred             --CCCeEEEEeccc-ccccCceeeecCCCC---cCCCcEEEEeeCCCCCce---eE---EEEEEeceeeeeccCCceeee
Q 009784          195 --ECDIAMLTVEDD-EFWEGVLPVEFGELP---ALQDAVTVVGYPIGGDTI---SV---TSGVVSRIEILSYVHGSTELL  262 (526)
Q Consensus       195 --~~DlAlLkv~~~-~~~~~~~pl~l~~~~---~~g~~V~~iG~p~~~~~~---sv---~~GiVs~~~~~~~~~~~~~~~  262 (526)
                        .+|+|||+++.+ .+...+.++.+....   ..++.+.++||+......   .+   ...+++...+...........
T Consensus        84 ~~~~DiAll~L~~~~~~~~~~~~~~l~~~~~~~~~~~~~~~~G~~~~~~~~~~~~~~~~~~~~~~~~~c~~~~~~~~~~~  163 (220)
T PF00089_consen   84 TYDNDIALLKLDRPITFGDNIQPICLPSAGSDPNVGTSCIVVGWGRTSDNGYSSNLQSVTVPVVSRKTCRSSYNDNLTPN  163 (220)
T ss_dssp             TTTTSEEEEEESSSSEHBSSBEESBBTSTTHTTTTTSEEEEEESSBSSTTSBTSBEEEEEEEEEEHHHHHHHTTTTSTTT
T ss_pred             cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccc
Confidence              479999999987 345578888887632   578899999998753221   23   333333333322111111123


Q ss_pred             EEEEcc----cCCCCCCCCeeecCCCeEEEEEeeccccCccccccccccHHHHHHHH
Q 009784          263 GLQIDA----AINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFI  315 (526)
Q Consensus       263 ~i~~da----~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~~~~~~~aIP~~~i~~~l  315 (526)
                      .+++..    ..+.|+|||||++.++.++||++.+.........++..++....++|
T Consensus       164 ~~c~~~~~~~~~~~g~sG~pl~~~~~~lvGI~s~~~~c~~~~~~~v~~~v~~~~~WI  220 (220)
T PF00089_consen  164 MICAGSSGSGDACQGDSGGPLICNNNYLVGIVSFGENCGSPNYPGVYTRVSSYLDWI  220 (220)
T ss_dssp             EEEEETTSSSBGGTTTTTSEEEETTEEEEEEEEEESSSSBTTSEEEEEEGGGGHHHH
T ss_pred             cccccccccccccccccccccccceeeecceeeecCCCCCCCcCEEEEEHHHhhccC
Confidence            466655    78899999999998777999998874322222346777776655543


No 14 
>cd00190 Tryp_SPc Trypsin-like serine protease; Many of these are synthesized as inactive precursor zymogens that are cleaved during limited proteolysis to generate their active forms. Alignment contains also inactive enzymes that have substitutions of the catalytic triad residues.
Probab=99.32  E-value=5e-11  Score=115.21  Aligned_cols=161  Identities=24%  Similarity=0.236  Sum_probs=98.6

Q ss_pred             CCCCCCccccCCC---cceEEEEEEEeCCEEEecccccCCC--CeEEEEEcCC--------CcEEEEEEEEec-------
Q 009784          134 EPNFSLPWQRKRQ---YSSSSSGFAIGGRRVLTNAHSVEHY--TQVKLKKRGS--------DTKYLATVLAIG-------  193 (526)
Q Consensus       134 ~~~~~~P~~~~~~---~~~~GSGfvI~~g~ILT~aHvV~~~--~~i~V~~~~~--------g~~~~a~vv~~d-------  193 (526)
                      .....+||.....   ....|+|++|++.+|||+|||+.+.  ..+.|.+...        ...+..+-+..+       
T Consensus         7 ~~~~~~Pw~v~i~~~~~~~~C~GtlIs~~~VLTaAhC~~~~~~~~~~v~~g~~~~~~~~~~~~~~~v~~~~~hp~y~~~~   86 (232)
T cd00190           7 AKIGSFPWQVSLQYTGGRHFCGGSLISPRWVLTAAHCVYSSAPSNYTVRLGSHDLSSNEGGGQVIKVKKVIVHPNYNPST   86 (232)
T ss_pred             CCCCCCCCEEEEEccCCcEEEEEEEeeCCEEEECHHhcCCCCCccEEEEeCcccccCCCCceEEEEEEEEEECCCCCCCC
Confidence            3444556655332   3478999999999999999999875  5666665311        122334444444       


Q ss_pred             cCCCeEEEEecccc-cccCceeeecCCC---CcCCCcEEEEeeCCCCCc-------eeEEEEEEeceeeeeccC--Ccee
Q 009784          194 TECDIAMLTVEDDE-FWEGVLPVEFGEL---PALQDAVTVVGYPIGGDT-------ISVTSGVVSRIEILSYVH--GSTE  260 (526)
Q Consensus       194 ~~~DlAlLkv~~~~-~~~~~~pl~l~~~---~~~g~~V~~iG~p~~~~~-------~sv~~GiVs~~~~~~~~~--~~~~  260 (526)
                      ..+|||||+++.+. +...+.|+.|...   ...++.+.+.||......       ......+++...+.....  ....
T Consensus        87 ~~~DiAll~L~~~~~~~~~v~picl~~~~~~~~~~~~~~~~G~g~~~~~~~~~~~~~~~~~~~~~~~~C~~~~~~~~~~~  166 (232)
T cd00190          87 YDNDIALLKLKRPVTLSDNVRPICLPSSGYNLPAGTTCTVSGWGRTSEGGPLPDVLQEVNVPIVSNAECKRAYSYGGTIT  166 (232)
T ss_pred             CcCCEEEEEECCcccCCCcccceECCCccccCCCCCEEEEEeCCcCCCCCCCCceeeEEEeeeECHHHhhhhccCcccCC
Confidence            35899999999763 2334788888754   345789999998765321       112223333322221111  0111


Q ss_pred             eeEEEE-----cccCCCCCCCCeeecCC---CeEEEEEeecc
Q 009784          261 LLGLQI-----DAAINSGNSGGPAFNDK---GKCVGIAFQSL  294 (526)
Q Consensus       261 ~~~i~~-----da~i~~G~SGGPlvn~~---G~VVGI~~~~~  294 (526)
                      ...++.     +...+.|+|||||+...   +.++||.+.+.
T Consensus       167 ~~~~C~~~~~~~~~~c~gdsGgpl~~~~~~~~~lvGI~s~g~  208 (232)
T cd00190         167 DNMLCAGGLEGGKDACQGDSGGPLVCNDNGRGVLVGIVSWGS  208 (232)
T ss_pred             CceEeeCCCCCCCccccCCCCCcEEEEeCCEEEEEEEEehhh
Confidence            122333     33578899999999864   78999998765


No 15 
>cd00987 PDZ_serine_protease PDZ domain of tryspin-like serine proteases, such as DegP/HtrA, which are oligomeric proteins involved in heat-shock response, chaperone function, and apoptosis. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.27  E-value=1.5e-11  Score=102.31  Aligned_cols=88  Identities=35%  Similarity=0.599  Sum_probs=74.0

Q ss_pred             cccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhh
Q 009784          328 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ  406 (526)
Q Consensus       328 ~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~  406 (526)
                      +|+|+.++.+ ++..++.++++ ...|++|.+|.++|||++ ||++||+|++|||++|.++.++.          ..+..
T Consensus         1 ~~~G~~~~~~-~~~~~~~~~~~-~~~g~~V~~v~~~s~a~~~gl~~GD~I~~Ing~~i~~~~~~~----------~~l~~   68 (90)
T cd00987           1 PWLGVTVQDL-TPDLAEELGLK-DTKGVLVASVDPGSPAAKAGLKPGDVILAVNGKPVKSVADLR----------RALAE   68 (90)
T ss_pred             CccceEEeEC-CHHHHHHcCCC-CCCEEEEEEECCCCHHHHcCCCcCCEEEEECCEECCCHHHHH----------HHHHh
Confidence            5899999999 66666666664 457999999999999998 99999999999999999988764          56666


Q ss_pred             cCCCCEEEEEEEECCEEEEEE
Q 009784          407 KYTGDSAAVKVLRDSKILNFN  427 (526)
Q Consensus       407 ~~~G~~v~l~v~R~G~~~~~~  427 (526)
                      ...|+.+.+++.|+|+..+++
T Consensus        69 ~~~~~~i~l~v~r~g~~~~~~   89 (90)
T cd00987          69 LKPGDKVTLTVLRGGKELTVT   89 (90)
T ss_pred             cCCCCEEEEEEEECCEEEEee
Confidence            556899999999999876654


No 16 
>smart00020 Tryp_SPc Trypsin-like serine protease. Many of these are synthesised as inactive precursor zymogens that are cleaved during limited proteolysis to generate their active forms. A few, however, are active as single chain molecules, and others are inactive due to substitutions of the catalytic triad residues.
Probab=99.22  E-value=1.5e-10  Score=112.18  Aligned_cols=164  Identities=24%  Similarity=0.239  Sum_probs=99.9

Q ss_pred             eeeCCCCCCccccCCC---cceEEEEEEEeCCEEEecccccCCCC--eEEEEEcCCC-------cEEEEEEEEec-----
Q 009784          131 VHTEPNFSLPWQRKRQ---YSSSSSGFAIGGRRVLTNAHSVEHYT--QVKLKKRGSD-------TKYLATVLAIG-----  193 (526)
Q Consensus       131 ~~~~~~~~~P~~~~~~---~~~~GSGfvI~~g~ILT~aHvV~~~~--~i~V~~~~~g-------~~~~a~vv~~d-----  193 (526)
                      ........+||.....   ....|+|++|++.+|||+|||+.+..  .+.|.+....       ..+...-+..+     
T Consensus         5 G~~~~~~~~Pw~~~i~~~~~~~~C~GtlIs~~~VLTaahC~~~~~~~~~~v~~g~~~~~~~~~~~~~~v~~~~~~p~~~~   84 (229)
T smart00020        5 GSEANIGSFPWQVSLQYRGGRHFCGGSLISPRWVLTAAHCVYGSDPSNIRVRLGSHDLSSGEEGQVIKVSKVIIHPNYNP   84 (229)
T ss_pred             CCcCCCCCCCcEEEEEEcCCCcEEEEEEecCCEEEECHHHcCCCCCcceEEEeCcccCCCCCCceEEeeEEEEECCCCCC
Confidence            3344455566655322   24579999999999999999998753  6777764221       23344444433     


Q ss_pred             --cCCCeEEEEecccc-cccCceeeecCCC---CcCCCcEEEEeeCCCCCc-----ee---EEEEEEeceeeeeccCC--
Q 009784          194 --TECDIAMLTVEDDE-FWEGVLPVEFGEL---PALQDAVTVVGYPIGGDT-----IS---VTSGVVSRIEILSYVHG--  257 (526)
Q Consensus       194 --~~~DlAlLkv~~~~-~~~~~~pl~l~~~---~~~g~~V~~iG~p~~~~~-----~s---v~~GiVs~~~~~~~~~~--  257 (526)
                        ...|||||+++.+. +...+.|+.+...   ...++.+.+.||+.....     ..   ...-+++...+......  
T Consensus        85 ~~~~~DiAll~L~~~i~~~~~~~pi~l~~~~~~~~~~~~~~~~g~g~~~~~~~~~~~~~~~~~~~~~~~~~C~~~~~~~~  164 (229)
T smart00020       85 STYDNDIALLKLKSPVTLSDNVRPICLPSSNYNVPAGTTCTVSGWGRTSEGAGSLPDTLQEVNVPIVSNATCRRAYSGGG  164 (229)
T ss_pred             CCCcCCEEEEEECcccCCCCceeeccCCCcccccCCCCEEEEEeCCCCCCCCCcCCCEeeEEEEEEeCHHHhhhhhcccc
Confidence              45899999998763 2345788888753   345788999998776430     01   11222222122111100  


Q ss_pred             ceeeeEEEE-----cccCCCCCCCCeeecCCC--eEEEEEeecc
Q 009784          258 STELLGLQI-----DAAINSGNSGGPAFNDKG--KCVGIAFQSL  294 (526)
Q Consensus       258 ~~~~~~i~~-----da~i~~G~SGGPlvn~~G--~VVGI~~~~~  294 (526)
                      ......++.     ....++|+||||++...+  .++||++.+.
T Consensus       165 ~~~~~~~C~~~~~~~~~~c~gdsG~pl~~~~~~~~l~Gi~s~g~  208 (229)
T smart00020      165 AITDNMLCAGGLEGGKDACQGDSGGPLVCNDGRWVLVGIVSWGS  208 (229)
T ss_pred             ccCCCcEeecCCCCCCcccCCCCCCeeEEECCCEEEEEEEEECC
Confidence            011112333     345788999999998654  8999998764


No 17 
>cd00986 PDZ_LON_protease PDZ domain of ATP-dependent LON serine proteases. Most PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this bacterial subfamily of protease-associated PDZ domains a C-terminal beta-strand  is thought to form the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.17  E-value=8.6e-11  Score=95.87  Aligned_cols=72  Identities=28%  Similarity=0.366  Sum_probs=63.7

Q ss_pred             ccceEEEEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEec
Q 009784          352 QKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLA  431 (526)
Q Consensus       352 ~~Gv~V~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~  431 (526)
                      ..|++|.+|.++|||+.||++||+|++|||+++.++.++.          ..+....+|+.+.+++.|+|+..++++++.
T Consensus         7 ~~Gv~V~~V~~~s~A~~gL~~GD~I~~Ing~~v~~~~~~~----------~~l~~~~~~~~v~l~v~r~g~~~~~~v~l~   76 (79)
T cd00986           7 YHGVYVTSVVEGMPAAGKLKAGDHIIAVDGKPFKEAEELI----------DYIQSKKEGDTVKLKVKREEKELPEDLILK   76 (79)
T ss_pred             ecCEEEEEECCCCchhhCCCCCCEEEEECCEECCCHHHHH----------HHHHhCCCCCEEEEEEEECCEEEEEEEEEe
Confidence            4689999999999998899999999999999999988864          566655679999999999999999999987


Q ss_pred             cc
Q 009784          432 TH  433 (526)
Q Consensus       432 ~~  433 (526)
                      ..
T Consensus        77 ~~   78 (79)
T cd00986          77 TF   78 (79)
T ss_pred             cc
Confidence            64


No 18 
>cd00991 PDZ_archaeal_metalloprotease PDZ domain of archaeal zinc metalloprotases, presumably membrane-associated or integral membrane proteases, which may be involved in signalling and regulatory mechanisms. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.15  E-value=1.2e-10  Score=95.22  Aligned_cols=69  Identities=26%  Similarity=0.273  Sum_probs=60.8

Q ss_pred             CccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEE
Q 009784          351 DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT  429 (526)
Q Consensus       351 ~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~  429 (526)
                      ...|++|.+|.++|||++ |||+||+|++|||+++.++.++.          ..+....+|+++.+++.|+|+..+++++
T Consensus         8 ~~~Gv~V~~V~~~spa~~aGL~~GDiI~~Ing~~v~~~~d~~----------~~l~~~~~g~~v~l~v~r~g~~~~~~~~   77 (79)
T cd00991           8 AVAGVVIVGVIVGSPAENAVLHTGDVIYSINGTPITTLEDFM----------EALKPTKPGEVITVTVLPSTTKLTNVST   77 (79)
T ss_pred             cCCcEEEEEECCCChHHhcCCCCCCEEEEECCEEcCCHHHHH----------HHHhcCCCCCEEEEEEEECCEEEEEEEE
Confidence            357999999999999998 99999999999999999988864          5666656799999999999999888775


No 19 
>TIGR01713 typeII_sec_gspC general secretion pathway protein C. This model represents GspC, protein C of the main terminal branch of the general secretion pathway, also called type II secretion. This system transports folded proteins across the bacterial outer membrane and is widely distributed in Gram-negative pathogens.
Probab=99.11  E-value=4e-10  Score=112.52  Aligned_cols=100  Identities=14%  Similarity=0.178  Sum_probs=86.3

Q ss_pred             HHHHHHHHHHHHcCceeeccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCC
Q 009784          309 PVIMHFIQDYEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAND  387 (526)
Q Consensus       309 ~~i~~~l~~l~~~g~~~~~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~  387 (526)
                      ..++++++++.++++.. +.|+|+......           ....|++|..+.++++|++ |||+||+|++|||+++.++
T Consensus       159 ~~~~~v~~~l~~~g~~~-~~~lgi~p~~~~-----------g~~~G~~v~~v~~~s~a~~aGLr~GDvIv~ING~~i~~~  226 (259)
T TIGR01713       159 VVSRRIIEELTKDPQKM-FDYIRLSPVMKN-----------DKLEGYRLNPGKDPSLFYKSGLQDGDIAVALNGLDLRDP  226 (259)
T ss_pred             hhHHHHHHHHHHCHHhh-hheEeEEEEEeC-----------CceeEEEEEecCCCCHHHHcCCCCCCEEEEECCEEcCCH
Confidence            45778899999999888 899999876541           2347999999999999999 9999999999999999999


Q ss_pred             CCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEe
Q 009784          388 GTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITL  430 (526)
Q Consensus       388 ~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l  430 (526)
                      .++.          ..+.....+++++|+|+|+|+.+++++.+
T Consensus       227 ~~~~----------~~l~~~~~~~~v~l~V~R~G~~~~i~v~~  259 (259)
T TIGR01713       227 EQAF----------QALQMLREETNLTLTVERDGQREDIYVRF  259 (259)
T ss_pred             HHHH----------HHHHhcCCCCeEEEEEEECCEEEEEEEEC
Confidence            8864          67777778899999999999998888764


No 20 
>cd00990 PDZ_glycyl_aminopeptidase PDZ domain associated with archaeal and bacterial M61 glycyl-aminopeptidases. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand is presumed to form the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=99.09  E-value=4.5e-10  Score=91.54  Aligned_cols=77  Identities=22%  Similarity=0.384  Sum_probs=63.4

Q ss_pred             cccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhh
Q 009784          328 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ  406 (526)
Q Consensus       328 ~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~  406 (526)
                      +|+|+.+...              ..|++|.+|.++|||++ ||++||+|++|||+++.++.+             .+..
T Consensus         1 ~~~G~~~~~~--------------~~~~~V~~V~~~s~a~~aGl~~GD~I~~Ing~~v~~~~~-------------~l~~   53 (80)
T cd00990           1 PYLGLTLDKE--------------EGLGKVTFVRDDSPADKAGLVAGDELVAVNGWRVDALQD-------------RLKE   53 (80)
T ss_pred             CcccEEEEcc--------------CCcEEEEEECCCChHHHhCCCCCCEEEEECCEEhHHHHH-------------HHHh
Confidence            4778777532              35799999999999999 999999999999999987443             3444


Q ss_pred             cCCCCEEEEEEEECCEEEEEEEEec
Q 009784          407 KYTGDSAAVKVLRDSKILNFNITLA  431 (526)
Q Consensus       407 ~~~G~~v~l~v~R~G~~~~~~v~l~  431 (526)
                      ...|+.+.+++.|+|+..++++++.
T Consensus        54 ~~~~~~v~l~v~r~g~~~~~~v~~~   78 (80)
T cd00990          54 YQAGDPVELTVFRDDRLIEVPLTLA   78 (80)
T ss_pred             cCCCCEEEEEEEECCEEEEEEEEec
Confidence            4578899999999999988888764


No 21 
>PRK10779 zinc metallopeptidase RseP; Provisional
Probab=99.00  E-value=8.9e-10  Score=118.86  Aligned_cols=132  Identities=12%  Similarity=0.080  Sum_probs=94.6

Q ss_pred             eEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEeccc
Q 009784          355 VRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATH  433 (526)
Q Consensus       355 v~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~~  433 (526)
                      .+|.+|.++|||++ |||+||+|++|||++|.+++++.          ..+....+|++++++|+|+|+.+++++++...
T Consensus       128 ~lV~~V~~~SpA~kAGLk~GDvI~~vnG~~V~~~~~l~----------~~v~~~~~g~~v~v~v~R~gk~~~~~v~l~~~  197 (449)
T PRK10779        128 PVVGEIAPNSIAAQAQIAPGTELKAVDGIETPDWDAVR----------LALVSKIGDESTTITVAPFGSDQRRDKTLDLR  197 (449)
T ss_pred             ccccccCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHhhccCCceEEEEEeCCccceEEEEeccc
Confidence            46899999999999 99999999999999999999975          56667778899999999999999998888654


Q ss_pred             ccccCCCCCCCCCceEEEeeEEEEeccccceeeeeeeecc--hhhhccccccceeeeecccchhhhHHHHHH
Q 009784          434 RRLIPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERIMN--MKLRSSFWTSSCIQCHNCQMSSLLWCLRCL  503 (526)
Q Consensus       434 ~~~~p~~~~~~~p~~~i~gG~~f~~lt~~~~~~~~~~i~~--~~~~sg~~~~~~~~~~~~~~~~~~~~~~~~  503 (526)
                      +.......  .. .. ...|  +.+++-.. ...+..+.+  .+.++|++.||+|...|.+-.++...++-+
T Consensus       198 ~~~~~~~~--~~-~~-~~lG--l~~~~~~~-~~vV~~V~~~SpA~~AGL~~GDvIl~Ing~~V~s~~dl~~~  262 (449)
T PRK10779        198 HWAFEPDK--QD-PV-SSLG--IRPRGPQI-EPVLAEVQPNSAASKAGLQAGDRIVKVDGQPLTQWQTFVTL  262 (449)
T ss_pred             ccccCccc--cc-hh-hccc--ccccCCCc-CcEEEeeCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHHHH
Confidence            32111000  00 00 0123  23333211 123444544  456799999999999999888776666543


No 22 
>TIGR02037 degP_htrA_DO periplasmic serine protease, Do/DeqQ family. This family consists of a set proteins various designated DegP, heat shock protein HtrA, and protease DO. The ortholog in Pseudomonas aeruginosa is designated MucD and is found in an operon that controls mucoid phenotype. This family also includes the DegQ (HhoA) paralog in E. coli which can rescue a DegP mutant, but not the smaller DegS paralog, which cannot. Members of this family are located in the periplasm and have separable functions as both protease and chaperone. Members have a trypsin domain and two copies of a PDZ domain. This protein protects bacteria from thermal and other stresses and may be important for the survival of bacterial pathogens.// The chaperone function is dominant at low temperatures, whereas the proteolytic activity is turned on at elevated temperatures.
Probab=98.89  E-value=3.3e-09  Score=113.88  Aligned_cols=90  Identities=24%  Similarity=0.473  Sum_probs=78.6

Q ss_pred             ccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhh
Q 009784          327 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVS  405 (526)
Q Consensus       327 ~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~  405 (526)
                      ..++|+.+..+ ++..++.++++....|++|.+|.++|||++ ||++||+|++|||++|.++.++.          +.+.
T Consensus       337 ~~~lGi~~~~l-~~~~~~~~~l~~~~~Gv~V~~V~~~SpA~~aGL~~GDvI~~Ing~~V~s~~d~~----------~~l~  405 (428)
T TIGR02037       337 NPFLGLTVANL-SPEIRKELRLKGDVKGVVVTKVVSGSPAARAGLQPGDVILSVNQQPVSSVAELR----------KVLD  405 (428)
T ss_pred             ccccceEEecC-CHHHHHHcCCCcCcCceEEEEeCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHH
Confidence            35789999988 788888889886668999999999999999 99999999999999999988864          6676


Q ss_pred             hcCCCCEEEEEEEECCEEEEEE
Q 009784          406 QKYTGDSAAVKVLRDSKILNFN  427 (526)
Q Consensus       406 ~~~~G~~v~l~v~R~G~~~~~~  427 (526)
                      ..+.|+.+.|+|+|+|+...+.
T Consensus       406 ~~~~g~~v~l~v~R~g~~~~~~  427 (428)
T TIGR02037       406 RAKKGGRVALLILRGGATIFVT  427 (428)
T ss_pred             hcCCCCEEEEEEEECCEEEEEE
Confidence            6667999999999999987654


No 23 
>cd00989 PDZ_metalloprotease PDZ domain of bacterial and plant zinc metalloprotases, presumably membrane-associated or integral membrane proteases, which may be involved in signalling and regulatory mechanisms. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=98.88  E-value=4.1e-09  Score=85.49  Aligned_cols=66  Identities=24%  Similarity=0.346  Sum_probs=55.7

Q ss_pred             cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEE
Q 009784          353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT  429 (526)
Q Consensus       353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~  429 (526)
                      ..++|.+|.++|||++ ||++||+|++|||+++.++.++.          ..+... .++.+.+++.|+|+..+++++
T Consensus        12 ~~~~V~~v~~~s~a~~~gl~~GD~I~~ing~~i~~~~~~~----------~~l~~~-~~~~~~l~v~r~~~~~~~~l~   78 (79)
T cd00989          12 IEPVIGEVVPGSPAAKAGLKAGDRILAINGQKIKSWEDLV----------DAVQEN-PGKPLTLTVERNGETITLTLT   78 (79)
T ss_pred             cCcEEEeECCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHHHHC-CCceEEEEEEECCEEEEEEec
Confidence            3488999999999998 99999999999999999988763          455443 478999999999988777664


No 24 
>cd00988 PDZ_CTP_protease PDZ domain of C-terminal processing-, tail-specific-, and tricorn proteases, which function in posttranslational protein processing, maturation, and disassembly or degradation, in Bacteria, Archaea, and plant chloroplasts. May be responsible for substrate recognition and/or binding, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of protease-associated PDZ domains a C-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in Eumetazoan signaling proteins.
Probab=98.85  E-value=1e-08  Score=84.40  Aligned_cols=68  Identities=24%  Similarity=0.343  Sum_probs=57.0

Q ss_pred             ccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCC--CCcccccccchhhhhhhhhcCCCCEEEEEEEEC-CEEEEEE
Q 009784          352 QKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAND--GTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD-SKILNFN  427 (526)
Q Consensus       352 ~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~--~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~-G~~~~~~  427 (526)
                      ..+++|..|.+++||++ ||++||+|++|||+++.++  .++.          ..+.. ..|+.+.+++.|+ |+..+++
T Consensus        12 ~~~~~V~~v~~~s~a~~~gl~~GD~I~~vng~~i~~~~~~~~~----------~~l~~-~~~~~i~l~v~r~~~~~~~~~   80 (85)
T cd00988          12 DGGLVITSVLPGSPAAKAGIKAGDIIVAIDGEPVDGLSLEDVV----------KLLRG-KAGTKVRLTLKRGDGEPREVT   80 (85)
T ss_pred             CCeEEEEEecCCCCHHHcCCCCCCEEEEECCEEcCCCCHHHHH----------HHhcC-CCCCEEEEEEEcCCCCEEEEE
Confidence            36799999999999999 9999999999999999998  5642          44433 4689999999999 8888877


Q ss_pred             EEe
Q 009784          428 ITL  430 (526)
Q Consensus       428 v~l  430 (526)
                      +++
T Consensus        81 ~~~   83 (85)
T cd00988          81 LTR   83 (85)
T ss_pred             EEE
Confidence            654


No 25 
>COG3591 V8-like Glu-specific endopeptidase [Amino acid transport and metabolism]
Probab=98.82  E-value=9.2e-08  Score=94.07  Aligned_cols=161  Identities=22%  Similarity=0.237  Sum_probs=96.2

Q ss_pred             eEEEEEEEeCCEEEecccccCCCC----eEEEEEc---CC-CcEEE--EEEEE-ecc---CCCeEEEEecccccc-----
Q 009784          149 SSSSGFAIGGRRVLTNAHSVEHYT----QVKLKKR---GS-DTKYL--ATVLA-IGT---ECDIAMLTVEDDEFW-----  209 (526)
Q Consensus       149 ~~GSGfvI~~g~ILT~aHvV~~~~----~i~V~~~---~~-g~~~~--a~vv~-~d~---~~DlAlLkv~~~~~~-----  209 (526)
                      ..|++|+|.+..+||++||+....    ++.+..+   ++ +..+.  ..... ...   +.|.+...+.+..+.     
T Consensus        64 ~~~~~~lI~pntvLTa~Hc~~s~~~G~~~~~~~p~g~~~~~~~~~~~~~~~~~~~~g~~~~~d~~~~~v~~~~~~~g~~~  143 (251)
T COG3591          64 LCTAATLIGPNTVLTAGHCIYSPDYGEDDIAAAPPGVNSDGGPFYGITKIEIRVYPGELYKEDGASYDVGEAALESGINI  143 (251)
T ss_pred             ceeeEEEEcCceEEEeeeEEecCCCChhhhhhcCCcccCCCCCCCceeeEEEEecCCceeccCCceeeccHHHhccCCCc
Confidence            455669999999999999997433    2222221   11 21221  11111 122   456666666543321     


Q ss_pred             -cCce--eeecCCCCcCCCcEEEEeeCCCCCc---eeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecCC
Q 009784          210 -EGVL--PVEFGELPALQDAVTVVGYPIGGDT---ISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDK  283 (526)
Q Consensus       210 -~~~~--pl~l~~~~~~g~~V~~iG~p~~~~~---~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~  283 (526)
                       ....  ...+....++++.+.++|||.+...   .-.+.+.|..+..          ..++.++.+.+|+||+|+++.+
T Consensus       144 ~~~~~~~~~~~~~~~~~~d~i~v~GYP~dk~~~~~~~e~t~~v~~~~~----------~~l~y~~dT~pG~SGSpv~~~~  213 (251)
T COG3591         144 GDVVNYLKRNTASEAKANDRITVIGYPGDKPNIGTMWESTGKVNSIKG----------NKLFYDADTLPGSSGSPVLISK  213 (251)
T ss_pred             cccccccccccccccccCceeEEEeccCCCCcceeEeeecceeEEEec----------ceEEEEecccCCCCCCceEecC
Confidence             1111  2223334467788999999987652   1223344433321          2578889999999999999998


Q ss_pred             CeEEEEEeeccccCcccccccc-ccHHHHHHHHHHHH
Q 009784          284 GKCVGIAFQSLKHEDVENIGYV-IPTPVIMHFIQDYE  319 (526)
Q Consensus       284 G~VVGI~~~~~~~~~~~~~~~a-IP~~~i~~~l~~l~  319 (526)
                      .++||+++.+....+....+++ .-...++++++++.
T Consensus       214 ~~vigv~~~g~~~~~~~~~n~~vr~t~~~~~~I~~~~  250 (251)
T COG3591         214 DEVIGVHYNGPGANGGSLANNAVRLTPEILNFIQQNI  250 (251)
T ss_pred             ceEEEEEecCCCcccccccCcceEecHHHHHHHHHhh
Confidence            8999999987754333344433 44566777777653


No 26 
>cd00136 PDZ PDZ domain, also called DHR (Dlg homologous region) or GLGF (after a conserved sequence motif). Many PDZ domains bind C-terminal polypeptides, though binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. Heterodimerization through PDZ-PDZ domain interactions adds to the domain's versatility, and PDZ domain-mediated interactions may be modulated dynamically through target phosphorylation. Some PDZ domains play a role in scaffolding supramolecular complexes. PDZ domains are found in diverse signaling proteins in bacteria, archebacteria, and eurkayotes. This CD contains two distinct structural subgroups with either a N- or C-terminal beta-strand forming the peptide-binding groove base. The circular permutation placing the strand on the N-terminus appears to be found in Eumetazoa only, while the C-terminal variant is found in all three kingdoms of life, and seems to co-occur with protease domains. PDZ domains have been named after PSD95(pos
Probab=98.59  E-value=8.4e-08  Score=75.87  Aligned_cols=55  Identities=29%  Similarity=0.527  Sum_probs=46.2

Q ss_pred             cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCC--CCcccccccchhhhhhhhhcCCCCEEEEEEE
Q 009784          353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAND--GTVPFRHGERIGFSYLVSQKYTGDSAAVKVL  418 (526)
Q Consensus       353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~--~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~  418 (526)
                      .|++|.+|.+++||++ ||++||+|++|||+++.++  .++.          ..+... .|++++|+++
T Consensus        13 ~~~~V~~v~~~s~a~~~gl~~GD~I~~Ing~~v~~~~~~~~~----------~~l~~~-~g~~v~l~v~   70 (70)
T cd00136          13 GGVVVLSVEPGSPAERAGLQAGDVILAVNGTDVKNLTLEDVA----------ELLKKE-VGEKVTLTVR   70 (70)
T ss_pred             CCEEEEEeCCCCHHHHcCCCCCCEEEEECCEECCCCCHHHHH----------HHHhhC-CCCeEEEEEC
Confidence            4899999999999999 9999999999999999998  5543          555554 4889998763


No 27 
>TIGR00054 RIP metalloprotease RseP. A model that detects fragments as well matches a number of members of the PEPTIDASE FAMILY S2C. The region of match appears not to overlap the active site domain.
Probab=98.52  E-value=1.7e-07  Score=100.45  Aligned_cols=113  Identities=18%  Similarity=0.161  Sum_probs=82.6

Q ss_pred             ccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEe
Q 009784          352 QKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITL  430 (526)
Q Consensus       352 ~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l  430 (526)
                      ..|.+|.+|.++|||++ |||+||+|+++||+++.++.++.          ..+....  +++.+++.|+|+..++++++
T Consensus       127 ~~g~~V~~V~~~SpA~~AGL~~GDvI~~vng~~v~~~~dl~----------~~ia~~~--~~v~~~I~r~g~~~~l~v~l  194 (420)
T TIGR00054       127 EVGPVIELLDKNSIALEAGIEPGDEILSVNGNKIPGFKDVR----------QQIADIA--GEPMVEILAERENWTFEVMK  194 (420)
T ss_pred             CCCceeeccCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHhhc--ccceEEEEEecCceEecccc
Confidence            36789999999999999 99999999999999999999875          4455543  68999999999887665543


Q ss_pred             cccccccCCCCCCCCCceEEEeeEEEEeccccceeeeeeeecc--hhhhccccccceeeeecccchhhhHHHH
Q 009784          431 ATHRRLIPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERIMN--MKLRSSFWTSSCIQCHNCQMSSLLWCLR  501 (526)
Q Consensus       431 ~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt~~~~~~~~~~i~~--~~~~sg~~~~~~~~~~~~~~~~~~~~~~  501 (526)
                      .-.+.         .+.                ....+..+.+  .+.++|++.||+|...|.+-.++...++
T Consensus       195 ~~~~~---------~~~----------------~g~vV~~V~~~SpA~~aGL~~GD~Iv~Vng~~V~s~~dl~  242 (420)
T TIGR00054       195 ELIPR---------GPK----------------IEPVLSDVTPNSPAEKAGLKEGDYIQSINGEKLRSWTDFV  242 (420)
T ss_pred             cceec---------CCC----------------cCcEEEEECCCCHHHHcCCCCCCEEEEECCEECCCHHHHH
Confidence            31110         000                0122233333  4557999999999999998877655544


No 28 
>TIGR00054 RIP metalloprotease RseP. A model that detects fragments as well matches a number of members of the PEPTIDASE FAMILY S2C. The region of match appears not to overlap the active site domain.
Probab=98.47  E-value=2.4e-07  Score=99.22  Aligned_cols=69  Identities=25%  Similarity=0.323  Sum_probs=60.8

Q ss_pred             cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEec
Q 009784          353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLA  431 (526)
Q Consensus       353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~  431 (526)
                      .|++|.+|.++|||++ |||+||+|++|||++|.+++|+.          ..+.. .+|+++.++++|+|+..++++++.
T Consensus       203 ~g~vV~~V~~~SpA~~aGL~~GD~Iv~Vng~~V~s~~dl~----------~~l~~-~~~~~v~l~v~R~g~~~~~~v~~~  271 (420)
T TIGR00054       203 IEPVLSDVTPNSPAEKAGLKEGDYIQSINGEKLRSWTDFV----------SAVKE-NPGKSMDIKVERNGETLSISLTPE  271 (420)
T ss_pred             cCcEEEEECCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHHHh-CCCCceEEEEEECCEEEEEEEEEc
Confidence            4799999999999999 99999999999999999999874          45544 578899999999999999988875


Q ss_pred             c
Q 009784          432 T  432 (526)
Q Consensus       432 ~  432 (526)
                      .
T Consensus       272 ~  272 (420)
T TIGR00054       272 A  272 (420)
T ss_pred             C
Confidence            3


No 29 
>smart00228 PDZ Domain present in PSD-95, Dlg, and ZO-1/2. Also called DHR (Dlg homologous region) or GLGF (relatively well conserved tetrapeptide in these domains). Some PDZs have been shown to bind C-terminal polypeptides; others appear to bind internal (non-C-terminal) polypeptides. Different PDZs possess different binding specificities.
Probab=98.46  E-value=4.8e-07  Score=73.84  Aligned_cols=59  Identities=27%  Similarity=0.402  Sum_probs=47.9

Q ss_pred             cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECC
Q 009784          353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDS  421 (526)
Q Consensus       353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G  421 (526)
                      .|++|..|.+++||+. ||++||+|++|||+++.+..+..          ........++.+.+++.|++
T Consensus        26 ~~~~i~~v~~~s~a~~~gl~~GD~I~~In~~~v~~~~~~~----------~~~~~~~~~~~~~l~i~r~~   85 (85)
T smart00228       26 GGVVVSSVVPGSPAAKAGLKVGDVILEVNGTSVEGLTHLE----------AVDLLKKAGGKVTLTVLRGG   85 (85)
T ss_pred             CCEEEEEECCCCHHHHcCCCCCCEEEEECCEECCCCCHHH----------HHHHHHhCCCeEEEEEEeCC
Confidence            6899999999999999 99999999999999999776542          22222334679999999975


No 30 
>PRK10779 zinc metallopeptidase RseP; Provisional
Probab=98.37  E-value=6e-07  Score=97.04  Aligned_cols=68  Identities=24%  Similarity=0.324  Sum_probs=60.1

Q ss_pred             ceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEecc
Q 009784          354 GVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLAT  432 (526)
Q Consensus       354 Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~  432 (526)
                      +++|.+|.++|||++ |||+||+|++|||++|.++.|+.          ..+.. ..|+.+.+++.|+|+..++++++..
T Consensus       222 ~~vV~~V~~~SpA~~AGL~~GDvIl~Ing~~V~s~~dl~----------~~l~~-~~~~~v~l~v~R~g~~~~~~v~~~~  290 (449)
T PRK10779        222 EPVLAEVQPNSAASKAGLQAGDRIVKVDGQPLTQWQTFV----------TLVRD-NPGKPLALEIERQGSPLSLTLTPDS  290 (449)
T ss_pred             CcEEEeeCCCCHHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHh-CCCCEEEEEEEECCEEEEEEEEeee
Confidence            588999999999999 99999999999999999999874          45544 5788999999999999999988753


No 31 
>TIGR00225 prc C-terminal peptidase (prc). A C-terminal peptidase with different substrates in different species including processing of D1 protein of the photosystem II reaction center in higher plants and cleavage of a peptide of 11 residues from the precursor form of penicillin-binding protein in E.coli E.coli and H influenza have the most distal branch of the tree and their proteins have an N-terminal 200 amino acids that show no homology to other proteins in the database.
Probab=98.33  E-value=1.3e-06  Score=90.83  Aligned_cols=70  Identities=21%  Similarity=0.327  Sum_probs=57.5

Q ss_pred             cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCC--CcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEE
Q 009784          353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDG--TVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT  429 (526)
Q Consensus       353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~--~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~  429 (526)
                      .+++|.+|.++|||++ ||++||+|++|||++|.++.  ++          ...+ ....|+++.+++.|+|+..+++++
T Consensus        62 ~~~~V~~V~~~spA~~aGL~~GD~I~~Ing~~v~~~~~~~~----------~~~l-~~~~g~~v~l~v~R~g~~~~~~v~  130 (334)
T TIGR00225        62 GEIVIVSPFEGSPAEKAGIKPGDKIIKINGKSVAGMSLDDA----------VALI-RGKKGTKVSLEILRAGKSKPLTFT  130 (334)
T ss_pred             CEEEEEEeCCCChHHHcCCCCCCEEEEECCEECCCCCHHHH----------HHhc-cCCCCCEEEEEEEeCCCCceEEEE
Confidence            5799999999999999 99999999999999998863  22          1222 335789999999999988888877


Q ss_pred             eccc
Q 009784          430 LATH  433 (526)
Q Consensus       430 l~~~  433 (526)
                      +...
T Consensus       131 l~~~  134 (334)
T TIGR00225       131 LKRD  134 (334)
T ss_pred             EEEE
Confidence            7654


No 32 
>PF00595 PDZ:  PDZ domain (Also known as DHR or GLGF) Coordinates are not yet available;  InterPro: IPR001478 PDZ domains are found in diverse signalling proteins in bacteria, yeasts, plants, insects and vertebrates [, ]. PDZ domains can occur in one or multiple copies and are nearly always found in cytoplasmic proteins. They bind either the carboxyl-terminal sequences of proteins or internal peptide sequences []. In most cases, interaction between a PDZ domain and its target is constitutive, with a binding affinity of 1 to 10 microns. However, agonist-dependent activation of cell surface receptors is sometimes required to promote interaction with a PDZ protein. PDZ domain proteins are frequently associated with the plasma membrane, a compartment where high concentrations of phosphatidylinositol 4,5-bisphosphate (PIP2) are found. Direct interaction between PIP2 and a subset of class II PDZ domains (syntenin, CASK, Tiam-1) has been demonstrated.  PDZ domains consist of 80 to 90 amino acids comprising six beta-strands (beta-A to beta-F) and two alpha-helices, A and B, compactly arranged in a globular structure. Peptide binding of the ligand takes place in an elongated surface groove as an anti-parallel beta-strand interacts with the beta-B strand and the B helix. The structure of PDZ domains allows binding to a free carboxylate group at the end of a peptide through a carboxylate-binding loop between the beta-A and beta-B strands.; GO: 0005515 protein binding; PDB: 3AXA_A 1WF8_A 1QAV_B 1QAU_A 1B8Q_A 1MC7_A 2KAW_A 1I16_A 1VB7_A 1WI4_A ....
Probab=98.26  E-value=1.5e-06  Score=71.04  Aligned_cols=72  Identities=24%  Similarity=0.308  Sum_probs=53.0

Q ss_pred             ccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhh
Q 009784          327 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVS  405 (526)
Q Consensus       327 ~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~  405 (526)
                      ...+|+.+......          ...+++|.+|.++|||++ ||++||.|++|||+.+.++....        ...++.
T Consensus         9 ~~~lG~~l~~~~~~----------~~~~~~V~~v~~~~~a~~~gl~~GD~Il~INg~~v~~~~~~~--------~~~~l~   70 (81)
T PF00595_consen    9 NGPLGFTLRGGSDN----------DEKGVFVSSVVPGSPAERAGLKVGDRILEINGQSVRGMSHDE--------VVQLLK   70 (81)
T ss_dssp             TSBSSEEEEEESTS----------SSEEEEEEEECTTSHHHHHTSSTTEEEEEETTEESTTSBHHH--------HHHHHH
T ss_pred             CCCcCEEEEecCCC----------CcCCEEEEEEeCCChHHhcccchhhhhheeCCEeCCCCCHHH--------HHHHHH
Confidence            45688888866210          126899999999999999 99999999999999999886532        112333


Q ss_pred             hcCCCCEEEEEEE
Q 009784          406 QKYTGDSAAVKVL  418 (526)
Q Consensus       406 ~~~~G~~v~l~v~  418 (526)
                      .  .+.+++|+|+
T Consensus        71 ~--~~~~v~L~V~   81 (81)
T PF00595_consen   71 S--ASNPVTLTVQ   81 (81)
T ss_dssp             H--STSEEEEEEE
T ss_pred             C--CCCcEEEEEC
Confidence            3  3448888874


No 33 
>PRK10139 serine endoprotease; Provisional
Probab=98.20  E-value=2.1e-06  Score=92.79  Aligned_cols=64  Identities=17%  Similarity=0.342  Sum_probs=55.9

Q ss_pred             cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEE
Q 009784          353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNI  428 (526)
Q Consensus       353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v  428 (526)
                      .|++|.+|.++|||++ |||+||+|++|||++|.++.++.          +.+...  .+++.|+|+|+|+.+.+.+
T Consensus       390 ~Gv~V~~V~~~spA~~aGL~~GD~I~~Ing~~v~~~~~~~----------~~l~~~--~~~v~l~v~R~g~~~~~~~  454 (455)
T PRK10139        390 KGIKIDEVVKGSPAAQAGLQKDDVIIGVNRDRVNSIAEMR----------KVLAAK--PAIIALQIVRGNESIYLLL  454 (455)
T ss_pred             CceEEEEeCCCChHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHhC--CCeEEEEEEECCEEEEEEe
Confidence            5899999999999999 99999999999999999999874          556543  3789999999999877654


No 34 
>TIGR03279 cyano_FeS_chp putative FeS-containing Cyanobacterial-specific oxidoreductase. Members of this protein family are predicted FeS-containing oxidoreductases of unknown function, apparently restricted to and universal across the Cyanobacteria. The high trusted cutoff score for this model, 700 bits, excludes homologs from other lineages. This exclusion seems justified because a significant number of sequence positions are simultaneously unique to and invariant across the Cyanobacteria, suggesting a specialized, conserved function, perhaps related to photosynthesis. A distantly related protein family, TIGR03278, in universal in and restricted to archaeal methanogens, and may be linked to methanogenesis.
Probab=98.20  E-value=1.9e-06  Score=90.94  Aligned_cols=62  Identities=19%  Similarity=0.303  Sum_probs=51.7

Q ss_pred             EEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEE-ECCEEEEEEEEecc
Q 009784          357 IRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVL-RDSKILNFNITLAT  432 (526)
Q Consensus       357 V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~-R~G~~~~~~v~l~~  432 (526)
                      |.+|.|+|+|++ ||++||+|++|||++|.++.|+.          ..+    .++.+.++|. |+|+..++++....
T Consensus         2 I~~V~pgSpAe~AGLe~GD~IlsING~~V~Dw~D~~----------~~l----~~e~l~L~V~~rdGe~~~l~Ie~~~   65 (433)
T TIGR03279         2 ISAVLPGSIAEELGFEPGDALVSINGVAPRDLIDYQ----------FLC----ADEELELEVLDANGESHQIEIEKDL   65 (433)
T ss_pred             cCCcCCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHh----cCCcEEEEEEcCCCeEEEEEEecCC
Confidence            678999999999 99999999999999999998863          333    2467999997 89988888877643


No 35 
>KOG3627 consensus Trypsin [Amino acid transport and metabolism]
Probab=98.16  E-value=9.9e-05  Score=73.19  Aligned_cols=167  Identities=22%  Similarity=0.233  Sum_probs=97.6

Q ss_pred             EEeeeeCCCCCCccccCCCc----ceEEEEEEEeCCEEEecccccCCCC--eEEEEEcC--------CC---cEE-EEEE
Q 009784          128 VFCVHTEPNFSLPWQRKRQY----SSSSSGFAIGGRRVLTNAHSVEHYT--QVKLKKRG--------SD---TKY-LATV  189 (526)
Q Consensus       128 I~~~~~~~~~~~P~~~~~~~----~~~GSGfvI~~g~ILT~aHvV~~~~--~i~V~~~~--------~g---~~~-~a~v  189 (526)
                      |.........++||+.....    ...|.|.+|++.||+|++||+.+..  .+.|.+..        .+   ... ..++
T Consensus        13 i~~g~~~~~~~~Pw~~~l~~~~~~~~~Cggsli~~~~vltaaHC~~~~~~~~~~V~~G~~~~~~~~~~~~~~~~~~v~~~   92 (256)
T KOG3627|consen   13 IVGGTEAEPGSFPWQVSLQYGGNGRHLCGGSLISPRWVLTAAHCVKGASASLYTVRLGEHDINLSVSEGEEQLVGDVEKI   92 (256)
T ss_pred             EeCCccCCCCCCCCEEEEEECCCcceeeeeEEeeCCEEEEChhhCCCCCCcceEEEECccccccccccCchhhhceeeEE
Confidence            34444444557788754433    2378888888889999999999865  66666521        11   111 1122


Q ss_pred             EEecc-------C-CCeEEEEeccc-ccccCceeeecCCCC----cCC-CcEEEEeeCCCC----C-c---eeEEEEEEe
Q 009784          190 LAIGT-------E-CDIAMLTVEDD-EFWEGVLPVEFGELP----ALQ-DAVTVVGYPIGG----D-T---ISVTSGVVS  247 (526)
Q Consensus       190 v~~d~-------~-~DlAlLkv~~~-~~~~~~~pl~l~~~~----~~g-~~V~~iG~p~~~----~-~---~sv~~GiVs  247 (526)
                      + .++       . .|||||+++.+ .+.+.+.|+.+....    ..+ ..+++.||....    . .   ..+..-+++
T Consensus        93 i-~H~~y~~~~~~~nDiall~l~~~v~~~~~i~piclp~~~~~~~~~~~~~~~v~GWG~~~~~~~~~~~~L~~~~v~i~~  171 (256)
T KOG3627|consen   93 I-VHPNYNPRTLENNDIALLRLSEPVTFSSHIQPICLPSSADPYFPPGGTTCLVSGWGRTESGGGPLPDTLQEVDVPIIS  171 (256)
T ss_pred             E-ECCCCCCCCCCCCCEEEEEECCCcccCCcccccCCCCCcccCCCCCCCEEEEEeCCCcCCCCCCCCceeEEEEEeEcC
Confidence            2 222       2 79999999975 355678888875322    223 677788875431    1 1   111233333


Q ss_pred             ceeeeeccCCc--eeeeEEEEc-----ccCCCCCCCCeeecCC---CeEEEEEeeccc
Q 009784          248 RIEILSYVHGS--TELLGLQID-----AAINSGNSGGPAFNDK---GKCVGIAFQSLK  295 (526)
Q Consensus       248 ~~~~~~~~~~~--~~~~~i~~d-----a~i~~G~SGGPlvn~~---G~VVGI~~~~~~  295 (526)
                      ...+.......  .....+++.     ...|.|+|||||+..+   ..++||++.+..
T Consensus       172 ~~~C~~~~~~~~~~~~~~~Ca~~~~~~~~~C~GDSGGPLv~~~~~~~~~~GivS~G~~  229 (256)
T KOG3627|consen  172 NSECRRAYGGLGTITDTMLCAGGPEGGKDACQGDSGGPLVCEDNGRWVLVGIVSWGSG  229 (256)
T ss_pred             hhHhcccccCccccCCCEEeeCccCCCCccccCCCCCeEEEeeCCcEEEEEEEEecCC
Confidence            32332221110  111135554     2468899999999764   699999988753


No 36 
>PLN00049 carboxyl-terminal processing protease; Provisional
Probab=98.15  E-value=4.9e-06  Score=88.32  Aligned_cols=69  Identities=20%  Similarity=0.302  Sum_probs=54.9

Q ss_pred             cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEe
Q 009784          353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITL  430 (526)
Q Consensus       353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l  430 (526)
                      .|++|..|.++|||++ ||++||+|++|||++|.++....        +...+ ....|+.+.|+|.|+|+..+++++-
T Consensus       102 ~g~~V~~V~~~SPA~~aGl~~GD~Iv~InG~~v~~~~~~~--------~~~~l-~g~~g~~v~ltv~r~g~~~~~~l~r  171 (389)
T PLN00049        102 AGLVVVAPAPGGPAARAGIRPGDVILAIDGTSTEGLSLYE--------AADRL-QGPEGSSVELTLRRGPETRLVTLTR  171 (389)
T ss_pred             CcEEEEEeCCCChHHHcCCCCCCEEEEECCEECCCCCHHH--------HHHHH-hcCCCCEEEEEEEECCEEEEEEEEe
Confidence            3899999999999999 99999999999999998753210        11333 3457899999999999887776654


No 37 
>PF14685 Tricorn_PDZ:  Tricorn protease PDZ domain; PDB: 1N6F_D 1N6D_C 1N6E_C 1K32_A.
Probab=98.14  E-value=1.6e-05  Score=66.12  Aligned_cols=65  Identities=23%  Similarity=0.385  Sum_probs=43.8

Q ss_pred             ccceEEEEeCCC--------CcccC-C--CCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEEC
Q 009784          352 QKGVRIRRVDPT--------APESE-V--LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD  420 (526)
Q Consensus       352 ~~Gv~V~~V~~~--------spA~~-G--L~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~  420 (526)
                      ..+..|.+|.++        ||..+ |  +++||+|++|||+++....++           +.+...+.|+.|.|+|.+.
T Consensus        11 ~~~y~I~~I~~gd~~~~~~~sPL~~pGv~v~~GD~I~aInG~~v~~~~~~-----------~~lL~~~agk~V~Ltv~~~   79 (88)
T PF14685_consen   11 NGGYRIARIYPGDPWNPNARSPLAQPGVDVREGDYILAINGQPVTADANP-----------YRLLEGKAGKQVLLTVNRK   79 (88)
T ss_dssp             TTEEEEEEE-BS-TTSSS-B-GGGGGS----TT-EEEEETTEE-BTTB-H-----------HHHHHTTTTSEEEEEEE-S
T ss_pred             CCEEEEEEEeCCCCCCccccCCccCCCCCCCCCCEEEEECCEECCCCCCH-----------HHHhcccCCCEEEEEEecC
Confidence            367889999875        67666 5  569999999999999988776           3444556899999999996


Q ss_pred             C-EEEEEE
Q 009784          421 S-KILNFN  427 (526)
Q Consensus       421 G-~~~~~~  427 (526)
                      + +.+++.
T Consensus        80 ~~~~R~v~   87 (88)
T PF14685_consen   80 PGGARTVV   87 (88)
T ss_dssp             TT-EEEEE
T ss_pred             CCCceEEE
Confidence            6 455554


No 38 
>TIGR02860 spore_IV_B stage IV sporulation protein B. SpoIVB, the stage IV sporulation protein B of endospore-forming bacteria such as Bacillus subtilis, is a serine proteinase, expressed in the spore (rather than mother cell) compartment, that participates in a proteolytic activation cascade for Sigma-K. It appears to be universal among endospore-forming bacteria and occurs nowhere else.
Probab=98.11  E-value=4e-06  Score=88.03  Aligned_cols=69  Identities=25%  Similarity=0.364  Sum_probs=57.1

Q ss_pred             ccceEEEEeC--------CCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCE
Q 009784          352 QKGVRIRRVD--------PTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK  422 (526)
Q Consensus       352 ~~Gv~V~~V~--------~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~  422 (526)
                      ..||+|....        .+|||++ |||+||+|++|||++|.++.++.          +.+... .|+++.++|.|+|+
T Consensus       104 t~GVlVvg~~~v~~~~g~~~SPAa~AGLq~GDiIvsING~~V~s~~DL~----------~iL~~~-~g~~V~LtV~R~Ge  172 (402)
T TIGR02860       104 TKGVLVVGFSDIETEKGKIHSPGEEAGIQIGDRILKINGEKIKNMDDLA----------NLINKA-GGEKLTLTIERGGK  172 (402)
T ss_pred             cCEEEEEEEEcccccCCCCCCHHHHcCCCCCCEEEEECCEECCCHHHHH----------HHHHhC-CCCeEEEEEEECCE
Confidence            4688886552        2589988 99999999999999999999874          555554 48999999999999


Q ss_pred             EEEEEEEec
Q 009784          423 ILNFNITLA  431 (526)
Q Consensus       423 ~~~~~v~l~  431 (526)
                      ..++++++.
T Consensus       173 ~~tv~V~Pv  181 (402)
T TIGR02860       173 IIETVIKPV  181 (402)
T ss_pred             EEEEEEEEe
Confidence            998888754


No 39 
>PRK10942 serine endoprotease; Provisional
Probab=98.08  E-value=5.5e-06  Score=90.01  Aligned_cols=64  Identities=20%  Similarity=0.343  Sum_probs=55.5

Q ss_pred             cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEE
Q 009784          353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNI  428 (526)
Q Consensus       353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v  428 (526)
                      .|++|.+|.++|+|++ ||++||+|++|||++|.++.++.          +.+.. . ++.+.|+|+|+|+.+.+.+
T Consensus       408 ~gvvV~~V~~~S~A~~aGL~~GDvIv~VNg~~V~s~~dl~----------~~l~~-~-~~~v~l~V~R~g~~~~v~~  472 (473)
T PRK10942        408 KGVVVDNVKPGTPAAQIGLKKGDVIIGANQQPVKNIAELR----------KILDS-K-PSVLALNIQRGDSSIYLLM  472 (473)
T ss_pred             CCeEEEEeCCCChHHHcCCCCCCEEEEECCEEcCCHHHHH----------HHHHh-C-CCeEEEEEEECCEEEEEEe
Confidence            5899999999999998 99999999999999999999874          55554 3 4789999999999877654


No 40 
>cd00992 PDZ_signaling PDZ domain found in a variety of Eumetazoan signaling molecules, often in tandem arrangements. May be responsible for specific protein-protein interactions, as most PDZ domains bind C-terminal polypeptides, and binding to internal (non-C-terminal) polypeptides and even to lipids has been demonstrated. In this subfamily of PDZ domains an N-terminal beta-strand forms the peptide-binding groove base, a circular permutation with respect to PDZ domains found in proteases.
Probab=98.02  E-value=9.4e-06  Score=65.93  Aligned_cols=52  Identities=23%  Similarity=0.427  Sum_probs=41.3

Q ss_pred             cccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEec--CCCCc
Q 009784          328 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIA--NDGTV  390 (526)
Q Consensus       328 ~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~--~~~~l  390 (526)
                      ..+|+.+....+           ...|++|.+|.++|||++ ||++||+|++|||+++.  +..++
T Consensus        12 ~~~G~~~~~~~~-----------~~~~~~V~~v~~~s~a~~~gl~~GD~I~~ing~~i~~~~~~~~   66 (82)
T cd00992          12 GGLGFSLRGGKD-----------SGGGIFVSRVEPGGPAERGGLRVGDRILEVNGVSVEGLTHEEA   66 (82)
T ss_pred             CCcCEEEeCccc-----------CCCCeEEEEECCCChHHhCCCCCCCEEEEECCEEcCccCHHHH
Confidence            457777765411           135899999999999999 99999999999999998  44443


No 41 
>PF00863 Peptidase_C4:  Peptidase family C4 This family belongs to family C4 of the peptidase classification.;  InterPro: IPR001730 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad [].  Nuclear inclusion A (NIA) proteases from potyviruses are cysteine peptidases belong to the MEROPS peptidase family C4 (NIa protease family, clan PA(C)) [, ].  Potyviruses include plant viruses in which the single-stranded RNA encodes a polyprotein with NIA protease activity, where proteolytic cleavage is specific for Gln+Gly sites. The NIA protease acts on the polyprotein, releasing itself by Gln+Gly cleavage at both the N- and C-termini. It further processes the polyprotein by cleavage at five similar sites in the C-terminal half of the sequence. In addition to its C-terminal protease activity, the NIA protease contains an N-terminal domain that has been implicated in the transcription process []. This peptidase is present in the nuclear inclusion protein of potyviruses.; GO: 0008234 cysteine-type peptidase activity, 0006508 proteolysis; PDB: 3MMG_B 1Q31_B 1LVB_A 1LVM_A.
Probab=97.98  E-value=0.00061  Score=66.77  Aligned_cols=169  Identities=20%  Similarity=0.279  Sum_probs=84.2

Q ss_pred             cccCCCeEEEEeeeeCCCCCCccccCCCcceEEEEEEEe-CCEEEecccccCC-CCeEEEEEcCCCcEEEEE-----EEE
Q 009784          119 VPAMDAVVKVFCVHTEPNFSLPWQRKRQYSSSSSGFAIG-GRRVLTNAHSVEH-YTQVKLKKRGSDTKYLAT-----VLA  191 (526)
Q Consensus       119 ~~~~~SVV~I~~~~~~~~~~~P~~~~~~~~~~GSGfvI~-~g~ILT~aHvV~~-~~~i~V~~~~~g~~~~a~-----vv~  191 (526)
                      .+....|+++......              ...+=+-|. ..+|+||+|.... ...++|... .| .|...     -+.
T Consensus        14 n~Ia~~ic~l~n~s~~--------------~~~~l~gigyG~~iItn~HLf~~nng~L~i~s~-hG-~f~v~nt~~lkv~   77 (235)
T PF00863_consen   14 NPIASNICRLTNESDG--------------GTRSLYGIGYGSYIITNAHLFKRNNGELTIKSQ-HG-EFTVPNTTQLKVH   77 (235)
T ss_dssp             HHHHTTEEEEEEEETT--------------EEEEEEEEEETTEEEEEGGGGSSTTCEEEEEET-TE-EEEECEGGGSEEE
T ss_pred             chhhheEEEEEEEeCC--------------CeEEEEEEeECCEEEEChhhhccCCCeEEEEeC-ce-EEEcCCccccceE
Confidence            3445578888764411              122223333 6699999999964 456777764 33 33321     133


Q ss_pred             eccCCCeEEEEecccccccCceeeecC---CCCcCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcc
Q 009784          192 IGTECDIAMLTVEDDEFWEGVLPVEFG---ELPALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDA  268 (526)
Q Consensus       192 ~d~~~DlAlLkv~~~~~~~~~~pl~l~---~~~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da  268 (526)
                      .=+..||.++|++.+     ++|.+--   ..+..++.|..+|.-+.....+   ..||.........+..   +...-.
T Consensus        78 ~i~~~DiviirmPkD-----fpPf~~kl~FR~P~~~e~v~mVg~~fq~k~~~---s~vSesS~i~p~~~~~---fWkHwI  146 (235)
T PF00863_consen   78 PIEGRDIVIIRMPKD-----FPPFPQKLKFRAPKEGERVCMVGSNFQEKSIS---STVSESSWIYPEENSH---FWKHWI  146 (235)
T ss_dssp             E-TCSSEEEEE--TT-----S----S---B----TT-EEEEEEEECSSCCCE---EEEEEEEEEEEETTTT---EEEE-C
T ss_pred             EeCCccEEEEeCCcc-----cCCcchhhhccCCCCCCEEEEEEEEEEcCCee---EEECCceEEeecCCCC---eeEEEe
Confidence            345799999999874     3554321   3567789999999866543321   2233322211111111   223333


Q ss_pred             cCCCCCCCCeeecC-CCeEEEEEeeccccCccccccccccH--HHHHHHHHH
Q 009784          269 AINSGNSGGPAFND-KGKCVGIAFQSLKHEDVENIGYVIPT--PVIMHFIQD  317 (526)
Q Consensus       269 ~i~~G~SGGPlvn~-~G~VVGI~~~~~~~~~~~~~~~aIP~--~~i~~~l~~  317 (526)
                      ....|+=|+|+++. +|++|||++...   .....+|+.|+  +.+..+++.
T Consensus       147 sTk~G~CG~PlVs~~Dg~IVGiHsl~~---~~~~~N~F~~f~~~f~~~~l~~  195 (235)
T PF00863_consen  147 STKDGDCGLPLVSTKDGKIVGIHSLTS---NTSSRNYFTPFPDDFEEFYLEN  195 (235)
T ss_dssp             ---TT-TT-EEEETTT--EEEEEEEEE---TTTSSEEEEE--TTHHHHHCC-
T ss_pred             cCCCCccCCcEEEcCCCcEEEEEcCcc---CCCCeEEEEcCCHHHHHHHhcc
Confidence            44579999999986 899999999765   23445566554  444444433


No 42 
>COG3480 SdrC Predicted secreted protein containing a PDZ domain [Signal transduction mechanisms]
Probab=97.94  E-value=1.4e-05  Score=79.99  Aligned_cols=72  Identities=25%  Similarity=0.295  Sum_probs=65.7

Q ss_pred             ccceEEEEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEE-CCEEEEEEEEe
Q 009784          352 QKGVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLR-DSKILNFNITL  430 (526)
Q Consensus       352 ~~Gv~V~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R-~G~~~~~~v~l  430 (526)
                      -.|+++..|..++|+...|+.||.|++|||+++.+.+++.          ..+..+++||+|++++.| +++....++++
T Consensus       129 y~gvyv~~v~~~~~~~gkl~~gD~i~avdg~~f~s~~e~i----------~~v~~~k~Gd~VtI~~~r~~~~~~~~~~tl  198 (342)
T COG3480         129 YAGVYVLSVIDNSPFKGKLEAGDTIIAVDGEPFTSSDELI----------DYVSSKKPGDEVTIDYERHNETPEIVTITL  198 (342)
T ss_pred             EeeEEEEEccCCcchhceeccCCeEEeeCCeecCCHHHHH----------HHHhccCCCCeEEEEEEeccCCCceEEEEE
Confidence            4699999999999998899999999999999999999965          888889999999999997 88888888888


Q ss_pred             ccc
Q 009784          431 ATH  433 (526)
Q Consensus       431 ~~~  433 (526)
                      ...
T Consensus       199 ~~~  201 (342)
T COG3480         199 IKN  201 (342)
T ss_pred             Eee
Confidence            776


No 43 
>COG0793 Prc Periplasmic protease [Cell envelope biogenesis, outer membrane]
Probab=97.91  E-value=3e-05  Score=82.52  Aligned_cols=84  Identities=24%  Similarity=0.387  Sum_probs=63.8

Q ss_pred             ccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCC--cccccccchhhhhh
Q 009784          327 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGT--VPFRHGERIGFSYL  403 (526)
Q Consensus       327 ~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~--l~~~~~~~~~~~~~  403 (526)
                      +..+|++++..             +..++.|.++.+++||++ ||++||+|++|||+++....-  .           ..
T Consensus        99 ~~GiG~~i~~~-------------~~~~~~V~s~~~~~PA~kagi~~GD~I~~IdG~~~~~~~~~~a-----------v~  154 (406)
T COG0793          99 FGGIGIELQME-------------DIGGVKVVSPIDGSPAAKAGIKPGDVIIKIDGKSVGGVSLDEA-----------VK  154 (406)
T ss_pred             ccceeEEEEEe-------------cCCCcEEEecCCCChHHHcCCCCCCEEEEECCEEccCCCHHHH-----------HH
Confidence            55677777644             126789999999999999 999999999999999987641  1           12


Q ss_pred             hhhcCCCCEEEEEEEECCEEEEEEEEecccc
Q 009784          404 VSQKYTGDSAAVKVLRDSKILNFNITLATHR  434 (526)
Q Consensus       404 l~~~~~G~~v~l~v~R~G~~~~~~v~l~~~~  434 (526)
                      ..+..+|.+|+|++.|.|....+++++.+..
T Consensus       155 ~irG~~Gt~V~L~i~r~~~~k~~~v~l~Re~  185 (406)
T COG0793         155 LIRGKPGTKVTLTILRAGGGKPFTVTLTREE  185 (406)
T ss_pred             HhCCCCCCeEEEEEEEcCCCceeEEEEEEEE
Confidence            3445689999999999865555666665543


No 44 
>PF05579 Peptidase_S32:  Equine arteritis virus serine endopeptidase S32;  InterPro: IPR008760 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to MEROPS peptidase family S32 (clan PA(S)). The type example is equine arteritis virus serine endopeptidase (equine arteritis virus), which is involved in processing of nidovirus polyproteins [].; GO: 0004252 serine-type endopeptidase activity, 0016032 viral reproduction, 0019082 viral protein processing; PDB: 3FAN_A 3FAO_A 1MBM_A.
Probab=97.80  E-value=0.00024  Score=69.76  Aligned_cols=115  Identities=19%  Similarity=0.226  Sum_probs=63.0

Q ss_pred             eEEEEEEEe-C--CEEEecccccCCCCeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecccccccCceeeecCCCCcCCC
Q 009784          149 SSSSGFAIG-G--RRVLTNAHSVEHYTQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGELPALQD  225 (526)
Q Consensus       149 ~~GSGfvI~-~--g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~~~~g~  225 (526)
                      ..|||=++. +  -.|+|+.||+. .+...|..  .+....   ..++..-|+|.-.++.-.  ...|.++++... .|.
T Consensus       112 s~Gsggvft~~~~~vvvTAtHVlg-~~~a~v~~--~g~~~~---~tF~~~GDfA~~~~~~~~--G~~P~~k~a~~~-~Gr  182 (297)
T PF05579_consen  112 SVGSGGVFTIGGNTVVVTATHVLG-GNTARVSG--VGTRRM---LTFKKNGDFAEADITNWP--GAAPKYKFAQNY-TGR  182 (297)
T ss_dssp             SEEEEEEEECTTEEEEEEEHHHCB-TTEEEEEE--TTEEEE---EEEEEETTEEEEEETTS---S---B--B-TT--SEE
T ss_pred             cccccceEEECCeEEEEEEEEEcC-CCeEEEEe--cceEEE---EEEeccCcEEEEECCCCC--CCCCceeecCCc-ccc
Confidence            355655555 3  38999999998 55666665  343333   335567799999884321  256667776221 111


Q ss_pred             cEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecCCCeEEEEEeecc
Q 009784          226 AVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSL  294 (526)
Q Consensus       226 ~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~  294 (526)
                      .-+   ...    .-+..|.|....+            ++   -..+|+||+|++..+|.+|||++++.
T Consensus       183 AyW---~t~----tGvE~G~ig~~~~------------~~---fT~~GDSGSPVVt~dg~liGVHTGSn  229 (297)
T PF05579_consen  183 AYW---LTS----TGVEPGFIGGGGA------------VC---FTGPGDSGSPVVTEDGDLIGVHTGSN  229 (297)
T ss_dssp             EEE---EET----TEEEEEEEETTEE------------EE---SS-GGCTT-EEEETTC-EEEEEEEEE
T ss_pred             eEE---Ecc----cCcccceecCceE------------EE---EcCCCCCCCccCcCCCCEEEEEecCC
Confidence            100   011    1245566554332            22   23589999999999999999999865


No 45 
>PRK09681 putative type II secretion protein GspC; Provisional
Probab=97.77  E-value=3.9e-05  Score=76.82  Aligned_cols=62  Identities=26%  Similarity=0.330  Sum_probs=51.7

Q ss_pred             EeCCCCcc---cC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEe
Q 009784          359 RVDPTAPE---SE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITL  430 (526)
Q Consensus       359 ~V~~~spA---~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l  430 (526)
                      .+.|+..+   ++ |||+||++++|||..+++.++..          .++.......+++|+|+|||+..++.+.+
T Consensus       210 rl~Pgkd~~lF~~~GLq~GDva~sING~dL~D~~qa~----------~l~~~L~~~tei~ltVeRdGq~~~i~i~l  275 (276)
T PRK09681        210 AVKPGADRSLFDASGFKEGDIAIALNQQDFTDPRAMI----------ALMRQLPSMDSIQLTVLRKGARHDISIAL  275 (276)
T ss_pred             EECCCCcHHHHHHcCCCCCCEEEEeCCeeCCCHHHHH----------HHHHHhccCCeEEEEEEECCEEEEEEEEc
Confidence            55677543   45 99999999999999999887753          66777778899999999999999998875


No 46 
>KOG3129 consensus 26S proteasome regulatory complex, subunit PSMD9 [Posttranslational modification, protein turnover, chaperones]
Probab=97.73  E-value=6.7e-05  Score=70.91  Aligned_cols=73  Identities=22%  Similarity=0.216  Sum_probs=61.2

Q ss_pred             ceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEecc
Q 009784          354 GVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLAT  432 (526)
Q Consensus       354 Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~  432 (526)
                      -++|.+|.|+|||+. ||+.||.|++++...-.++..|+        -...+..+..++.+.++|.|.|+.+.+.++++.
T Consensus       140 Fa~V~sV~~~SPA~~aGl~~gD~il~fGnV~sgn~~~lq--------~i~~~v~~~e~~~v~v~v~R~g~~v~L~ltP~~  211 (231)
T KOG3129|consen  140 FAVVDSVVPGSPADEAGLCVGDEILKFGNVHSGNFLPLQ--------NIAAVVQSNEDQIVSVTVIREGQKVVLSLTPKK  211 (231)
T ss_pred             eEEEeecCCCChhhhhCcccCceEEEecccccccchhHH--------HHHHHHHhccCcceeEEEecCCCEEEEEeCccc
Confidence            578999999999999 99999999999988888777653        113455567899999999999999999998876


Q ss_pred             cc
Q 009784          433 HR  434 (526)
Q Consensus       433 ~~  434 (526)
                      +.
T Consensus       212 W~  213 (231)
T KOG3129|consen  212 WQ  213 (231)
T ss_pred             cc
Confidence            53


No 47 
>PF04495 GRASP55_65:  GRASP55/65 PDZ-like domain ;  InterPro: IPR007583 GRASP55 (Golgi reassembly stacking protein of 55 kDa) and GRASP65 (a 65 kDa) protein are highly homologous. GRASP55 is a component of the Golgi stacking machinery. GRASP65, an N-ethylmaleimide-sensitive membrane protein required for the stacking of Golgi cisternae in a cell-free system [].; PDB: 3RLE_A 4EDJ_A.
Probab=97.47  E-value=0.00034  Score=63.32  Aligned_cols=87  Identities=23%  Similarity=0.325  Sum_probs=55.2

Q ss_pred             ccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCC-CcEEEEECCEEecCCCCcccccccchhhhhhh
Q 009784          327 FPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKP-SDIILSFDGIDIANDGTVPFRHGERIGFSYLV  404 (526)
Q Consensus       327 ~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~-GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l  404 (526)
                      .+.||+.++--. ..-       ....+.-|.+|.|+|||++ ||++ .|.|+.+|+..+.+.++|.          ..+
T Consensus        25 ~g~LG~sv~~~~-~~~-------~~~~~~~Vl~V~p~SPA~~AGL~p~~DyIig~~~~~l~~~~~l~----------~~v   86 (138)
T PF04495_consen   25 QGLLGISVRFES-FEG-------AEEEGWHVLRVAPNSPAAKAGLEPFFDYIIGIDGGLLDDEDDLF----------ELV   86 (138)
T ss_dssp             SSSS-EEEEEEE--TT-------GCCCEEEEEEE-TTSHHHHTT--TTTEEEEEETTCE--STCHHH----------HHH
T ss_pred             CCCCcEEEEEec-ccc-------cccceEEEeEecCCCHHHHCCccccccEEEEccceecCCHHHHH----------HHH
Confidence            466787776441 110       1356899999999999998 9999 6999999998888766542          555


Q ss_pred             hhcCCCCEEEEEEEEC--CEEEEEEEEecc
Q 009784          405 SQKYTGDSAAVKVLRD--SKILNFNITLAT  432 (526)
Q Consensus       405 ~~~~~G~~v~l~v~R~--G~~~~~~v~l~~  432 (526)
                      . .+.++++.|.|+..  .+.+++++++..
T Consensus        87 ~-~~~~~~l~L~Vyns~~~~vR~V~i~P~~  115 (138)
T PF04495_consen   87 E-ANENKPLQLYVYNSKTDSVREVTITPSR  115 (138)
T ss_dssp             H-HTTTS-EEEEEEETTTTCEEEEEE---T
T ss_pred             H-HcCCCcEEEEEEECCCCeEEEEEEEcCC
Confidence            4 45789999999984  455666666553


No 48 
>PF00548 Peptidase_C3:  3C cysteine protease (picornain 3C);  InterPro: IPR000199 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad [].  This signature defines cysteine peptidases belong to MEROPS peptidase family C3 (picornain, clan PA(C)), subfamilies C3A and C3B. The protein fold of this peptidase domain for members of this family resembles that of the serine peptidase, chymotrypsin [], the type example for clan PA. Picornaviral proteins are expressed as a single polyprotein which is cleaved by the viral C3 cysteine protease. The poliovirus polyprotein is selectively cleaved between the Gln-|-Gly bond. In other picornavirus reactions Glu may be substituted for Gln, and Ser or Thr for Gly. ; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis; PDB: 3SJO_E 2H6M_A 1QA7_C 1HAV_B 2HAL_A 2H9H_A 3QZQ_B 3QZR_A 3R0F_B 3SJ9_A ....
Probab=97.39  E-value=0.0047  Score=58.15  Aligned_cols=138  Identities=17%  Similarity=0.269  Sum_probs=83.7

Q ss_pred             cceEEEEEEEeCCEEEecccccCCCCeEEEEEcCCCcEEEE--EEEEecc---CCCeEEEEecccccccCc-eeeecCCC
Q 009784          147 YSSSSSGFAIGGRRVLTNAHSVEHYTQVKLKKRGSDTKYLA--TVLAIGT---ECDIAMLTVEDDEFWEGV-LPVEFGEL  220 (526)
Q Consensus       147 ~~~~GSGfvI~~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a--~vv~~d~---~~DlAlLkv~~~~~~~~~-~pl~l~~~  220 (526)
                      ....++|+.|.+.++|.+.|   .....++.+  +|..++.  .+...+.   ..||++++++...-+.++ +.+.  +.
T Consensus        23 g~~t~l~~gi~~~~~lvp~H---~~~~~~i~i--~g~~~~~~d~~~lv~~~~~~~Dl~~v~l~~~~kfrDIrk~~~--~~   95 (172)
T PF00548_consen   23 GEFTMLALGIYDRYFLVPTH---EEPEDTIYI--DGVEYKVDDSVVLVDRDGVDTDLTLVKLPRNPKFRDIRKFFP--ES   95 (172)
T ss_dssp             EEEEEEEEEEEBTEEEEEGG---GGGCSEEEE--TTEEEEEEEEEEEEETTSSEEEEEEEEEESSS-B--GGGGSB--SS
T ss_pred             ceEEEecceEeeeEEEEECc---CCCcEEEEE--CCEEEEeeeeEEEecCCCcceeEEEEEccCCcccCchhhhhc--cc
Confidence            44678999999999999999   223334444  4555542  3333444   469999999765422222 2222  22


Q ss_pred             C-cCCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecC---CCeEEEEEeec
Q 009784          221 P-ALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFND---KGKCVGIAFQS  293 (526)
Q Consensus       221 ~-~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~---~G~VVGI~~~~  293 (526)
                      . ...+.+.++-.+ ......+..+.++..+.. ...+......+..+++..+|+-||||+..   .++++||+.++
T Consensus        96 ~~~~~~~~l~v~~~-~~~~~~~~v~~v~~~~~i-~~~g~~~~~~~~Y~~~t~~G~CG~~l~~~~~~~~~i~GiHvaG  170 (172)
T PF00548_consen   96 IPEYPECVLLVNST-KFPRMIVEVGFVTNFGFI-NLSGTTTPRSLKYKAPTKPGMCGSPLVSRIGGQGKIIGIHVAG  170 (172)
T ss_dssp             GGTEEEEEEEEESS-SSTCEEEEEEEEEEEEEE-EETTEEEEEEEEEESEEETTGTTEEEEESCGGTTEEEEEEEEE
T ss_pred             cccCCCcEEEEECC-CCccEEEEEEEEeecCcc-ccCCCEeeEEEEEccCCCCCccCCeEEEeeccCccEEEEEecc
Confidence            2 233444444333 333334555556555443 23344444578888999999999999863   58999999885


No 49 
>PRK11186 carboxy-terminal protease; Provisional
Probab=97.39  E-value=0.00068  Score=76.21  Aligned_cols=71  Identities=15%  Similarity=0.135  Sum_probs=49.1

Q ss_pred             cceEEEEeCCCCcccC--CCCCCcEEEEEC--CEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEEC---CEEEE
Q 009784          353 KGVRIRRVDPTAPESE--VLKPSDIILSFD--GIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD---SKILN  425 (526)
Q Consensus       353 ~Gv~V~~V~~~spA~~--GL~~GDvIl~vn--G~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~---G~~~~  425 (526)
                      .+++|.+|.+||||++  ||++||+|++||  |+++.+...+.      +.-...+.....|.+|.|+|.|+   |+..+
T Consensus       255 ~~~~V~~vipGsPA~ka~gLk~GD~IlaVn~~g~~~~dv~g~~------~~~vv~lirG~~Gt~V~LtV~r~~~~~~~~~  328 (667)
T PRK11186        255 DYTVINSLVAGGPAAKSKKLSVGDKIVGVGQDGKPIVDVIGWR------LDDVVALIKGPKGSKVRLEILPAGKGTKTRI  328 (667)
T ss_pred             CeEEEEEccCCChHHHhCCCCCCCEEEEECCCCCcccccccCC------HHHHHHHhcCCCCCEEEEEEEeCCCCCceEE
Confidence            4688999999999997  899999999999  55554432221      11112233455799999999994   45555


Q ss_pred             EEEE
Q 009784          426 FNIT  429 (526)
Q Consensus       426 ~~v~  429 (526)
                      ++++
T Consensus       329 vtl~  332 (667)
T PRK11186        329 VTLT  332 (667)
T ss_pred             EEEE
Confidence            5543


No 50 
>COG3975 Predicted protease with the C-terminal PDZ domain [General function prediction only]
Probab=97.32  E-value=0.00056  Score=73.09  Aligned_cols=85  Identities=21%  Similarity=0.363  Sum_probs=65.3

Q ss_pred             cCceeeeccChhHHhhhcCCC--CccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhh
Q 009784          330 LGVEWQKMENPDLRVAMSMKA--DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ  406 (526)
Q Consensus       330 LGi~~~~~~~~~~~~~lgl~~--~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~  406 (526)
                      .|+.+.....  ..-++|+.-  +..+.+|..|.++|||++ ||.+||.|++|||.      +            ..+.+
T Consensus       439 ~gL~~~~~~~--~~~~LGl~v~~~~g~~~i~~V~~~gPA~~AGl~~Gd~ivai~G~------s------------~~l~~  498 (558)
T COG3975         439 FGLTFTPKPR--EAYYLGLKVKSEGGHEKITFVFPGGPAYKAGLSPGDKIVAINGI------S------------DQLDR  498 (558)
T ss_pred             cceEEEecCC--CCcccceEecccCCeeEEEecCCCChhHhccCCCccEEEEEcCc------c------------ccccc
Confidence            3555555521  134566543  345688999999999999 99999999999999      1            23556


Q ss_pred             cCCCCEEEEEEEECCEEEEEEEEecccc
Q 009784          407 KYTGDSAAVKVLRDSKILNFNITLATHR  434 (526)
Q Consensus       407 ~~~G~~v~l~v~R~G~~~~~~v~l~~~~  434 (526)
                      .+.++.|++.+.|.|+.+++.+++....
T Consensus       499 ~~~~d~i~v~~~~~~~L~e~~v~~~~~~  526 (558)
T COG3975         499 YKVNDKIQVHVFREGRLREFLVKLGGDP  526 (558)
T ss_pred             cccccceEEEEccCCceEEeecccCCCc
Confidence            6789999999999999999988877654


No 51 
>COG3031 PulC Type II secretory pathway, component PulC [Intracellular trafficking and secretion]
Probab=96.96  E-value=0.00067  Score=65.53  Aligned_cols=66  Identities=18%  Similarity=0.223  Sum_probs=52.1

Q ss_pred             ceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEE
Q 009784          354 GVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNIT  429 (526)
Q Consensus       354 Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~  429 (526)
                      |..+.-..+++-.++ |||.||+.+++|+..+++.+++.          .++.....-+.++++|+|+|+..++.|.
T Consensus       208 Gyr~~pgkd~slF~~sglq~GDIavaiNnldltdp~~m~----------~llq~l~~m~s~qlTv~R~G~rhdInV~  274 (275)
T COG3031         208 GYRFEPGKDGSLFYKSGLQRGDIAVAINNLDLTDPEDMF----------RLLQMLRNMPSLQLTVIRRGKRHDINVR  274 (275)
T ss_pred             EEEecCCCCcchhhhhcCCCcceEEEecCcccCCHHHHH----------HHHHhhhcCcceEEEEEecCccceeeec
Confidence            444444445566677 99999999999999999988853          5666666667899999999999988874


No 52 
>PF03761 DUF316:  Domain of unknown function (DUF316) ;  InterPro: IPR005514 This is a family of uncharacterised proteins from Caenorhabditis elegans.
Probab=96.78  E-value=0.039  Score=55.79  Aligned_cols=107  Identities=17%  Similarity=0.241  Sum_probs=62.6

Q ss_pred             CCCeEEEEecccccccCceeeecCCCCc---CCCcEEEEeeCCCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCC
Q 009784          195 ECDIAMLTVEDDEFWEGVLPVEFGELPA---LQDAVTVVGYPIGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAIN  271 (526)
Q Consensus       195 ~~DlAlLkv~~~~~~~~~~pl~l~~~~~---~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~  271 (526)
                      ..+++||+++.+ +.....|+.|+++..   .++.+.+.|+... ..  +....+.-.....      ....+......+
T Consensus       160 ~~~~mIlEl~~~-~~~~~~~~Cl~~~~~~~~~~~~~~~yg~~~~-~~--~~~~~~~i~~~~~------~~~~~~~~~~~~  229 (282)
T PF03761_consen  160 PYSPMILELEED-FSKNVSPPCLADSSTNWEKGDEVDVYGFNST-GK--LKHRKLKITNCTK------CAYSICTKQYSC  229 (282)
T ss_pred             ccceEEEEEccc-ccccCCCEEeCCCccccccCceEEEeecCCC-Ce--EEEEEEEEEEeec------cceeEecccccC
Confidence            479999999988 334788889987543   4688889998222 11  2222222111100      112355566778


Q ss_pred             CCCCCCeeecC-CC--eEEEEEeeccccCccccccccccHHHHH
Q 009784          272 SGNSGGPAFND-KG--KCVGIAFQSLKHEDVENIGYVIPTPVIM  312 (526)
Q Consensus       272 ~G~SGGPlvn~-~G--~VVGI~~~~~~~~~~~~~~~aIP~~~i~  312 (526)
                      .|++|||++.. +|  .||||.+..... ...+..+++.+...+
T Consensus       230 ~~d~Gg~lv~~~~gr~tlIGv~~~~~~~-~~~~~~~f~~v~~~~  272 (282)
T PF03761_consen  230 KGDRGGPLVKNINGRWTLIGVGASGNYE-CNKNNSYFFNVSWYQ  272 (282)
T ss_pred             CCCccCeEEEEECCCEEEEEEEccCCCc-ccccccEEEEHHHhh
Confidence            99999999832 44  699998764321 111244555554443


No 53 
>KOG3553 consensus Tax interaction protein TIP1 [Cell wall/membrane/envelope biogenesis]
Probab=96.66  E-value=0.0018  Score=54.28  Aligned_cols=35  Identities=31%  Similarity=0.505  Sum_probs=32.3

Q ss_pred             CccceEEEEeCCCCcccC-CCCCCcEEEEECCEEec
Q 009784          351 DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIA  385 (526)
Q Consensus       351 ~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~  385 (526)
                      .+.|++|++|..+|||+. ||+.+|.|+.+||...+
T Consensus        57 tD~GiYvT~V~eGsPA~~AGLrihDKIlQvNG~DfT   92 (124)
T KOG3553|consen   57 TDKGIYVTRVSEGSPAEIAGLRIHDKILQVNGWDFT   92 (124)
T ss_pred             CCccEEEEEeccCChhhhhcceecceEEEecCceeE
Confidence            368999999999999999 99999999999998765


No 54 
>PF08192 Peptidase_S64:  Peptidase family S64;  InterPro: IPR012985 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This family of fungal proteins is involved in the processing of membrane bound transcription factor Stp1 [] and belongs to MEROPS petidase family S64 (clan PA). The processing causes the signalling domain of Stp1 to be passed to the nucleus where several permease genes are induced. The permeases are important for uptake of amino acids, and processing of tp1 only occurs in an amino acid-rich environment. This family is predicted to be distantly related to the trypsin family (MEROPS peptidase family S1) and to have a typical trypsin-like catalytic triad [].
Probab=96.49  E-value=0.03  Score=61.81  Aligned_cols=117  Identities=19%  Similarity=0.288  Sum_probs=74.2

Q ss_pred             CCCeEEEEecccc-----cccCc------eeeecCCC--------CcCCCcEEEEeeCCCCCceeEEEEEEeceeeeecc
Q 009784          195 ECDIAMLTVEDDE-----FWEGV------LPVEFGEL--------PALQDAVTVVGYPIGGDTISVTSGVVSRIEILSYV  255 (526)
Q Consensus       195 ~~DlAlLkv~~~~-----~~~~~------~pl~l~~~--------~~~g~~V~~iG~p~~~~~~sv~~GiVs~~~~~~~~  255 (526)
                      -.|+||++++...     +.+++      +.+.+.+.        ...|.+|+=+|...+     .+.|.|.++....+.
T Consensus       542 LsD~AIIkV~~~~~~~N~LGddi~f~~~dP~l~f~NlyV~~~~~~~~~G~~VfK~GrTTg-----yT~G~lNg~klvyw~  616 (695)
T PF08192_consen  542 LSDWAIIKVNKERKCQNYLGDDIQFNEPDPTLMFQNLYVREVVSNLVPGMEVFKVGRTTG-----YTTGILNGIKLVYWA  616 (695)
T ss_pred             ccceEEEEeCCCceecCCCCccccccCCCccccccccchhhhhhccCCCCeEEEecccCC-----ccceEecceEEEEec
Confidence            3699999998653     12222      23344331        123678999998766     456888877655455


Q ss_pred             CCcee-eeEEEEc----ccCCCCCCCCeeecCCCe------EEEEEeeccccCccccccccccHHHHHHHHHHH
Q 009784          256 HGSTE-LLGLQID----AAINSGNSGGPAFNDKGK------CVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDY  318 (526)
Q Consensus       256 ~~~~~-~~~i~~d----a~i~~G~SGGPlvn~~G~------VVGI~~~~~~~~~~~~~~~aIP~~~i~~~l~~l  318 (526)
                      +|... .+++...    .-...|+||+-|++.-+.      |+||..+.-  .....+|++.|+..|.+-|++.
T Consensus       617 dG~i~s~efvV~s~~~~~Fa~~GDSGS~VLtk~~d~~~gLgvvGMlhsyd--ge~kqfglftPi~~il~rl~~v  688 (695)
T PF08192_consen  617 DGKIQSSEFVVSSDNNPAFASGGDSGSWVLTKLEDNNKGLGVVGMLHSYD--GEQKQFGLFTPINEILDRLEEV  688 (695)
T ss_pred             CCCeEEEEEEEecCCCccccCCCCcccEEEecccccccCceeeEEeeecC--CccceeeccCcHHHHHHHHHHh
Confidence            54422 2233333    233579999999986444      999998743  2345789999998887766554


No 55 
>COG5640 Secreted trypsin-like serine protease [Posttranslational modification, protein turnover, chaperones]
Probab=96.37  E-value=0.062  Score=55.28  Aligned_cols=59  Identities=22%  Similarity=0.304  Sum_probs=36.8

Q ss_pred             eEEEEEEEeCCEEEecccccCCCC-----eEEE--EEcC--CCcEEEEEEEEec-------cCCCeEEEEecccc
Q 009784          149 SSSSGFAIGGRRVLTNAHSVEHYT-----QVKL--KKRG--SDTKYLATVLAIG-------TECDIAMLTVEDDE  207 (526)
Q Consensus       149 ~~GSGfvI~~g~ILT~aHvV~~~~-----~i~V--~~~~--~g~~~~a~vv~~d-------~~~DlAlLkv~~~~  207 (526)
                      ..|-|-++...||||+|||+.+..     .+.|  .+.+  .+.....+.++.+       ...|+|++++....
T Consensus        61 tfCGgs~l~~RYvLTAAHC~~~~s~is~d~~~vv~~l~d~Sq~~rg~vr~i~~~efY~~~n~~ND~Av~~l~~~a  135 (413)
T COG5640          61 TFCGGSKLGGRYVLTAAHCADASSPISSDVNRVVVDLNDSSQAERGHVRTIYVHEFYSPGNLGNDIAVLELARAA  135 (413)
T ss_pred             eEeccceecceEEeeehhhccCCCCccccceEEEecccccccccCcceEEEeeecccccccccCcceeecccccc
Confidence            357788888779999999998654     1222  2221  1222334444433       35799999998754


No 56 
>PF12812 PDZ_1:  PDZ-like domain
Probab=96.31  E-value=0.0044  Score=50.46  Aligned_cols=60  Identities=12%  Similarity=0.047  Sum_probs=50.3

Q ss_pred             cccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcc
Q 009784          328 PLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVP  391 (526)
Q Consensus       328 ~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~  391 (526)
                      -+.|..++++ +-+.++.++++   -|+++.....++++.. |+..|-+|++|||+++.+.+++.
T Consensus         9 ~~~Ga~f~~L-s~q~aR~~~~~---~~gv~v~~~~g~~~~~~~i~~g~iI~~Vn~kpt~~Ld~f~   69 (78)
T PF12812_consen    9 EVCGAVFHDL-SYQQARQYGIP---VGGVYVAVSGGSLAFAGGISKGFIITSVNGKPTPDLDDFI   69 (78)
T ss_pred             EEcCeecccC-CHHHHHHhCCC---CCEEEEEecCCChhhhCCCCCCeEEEeECCcCCcCHHHHH
Confidence            3789999998 88899999988   3355556788899888 69999999999999999988753


No 57 
>PF10459 Peptidase_S46:  Peptidase S46;  InterPro: IPR019500 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This entry represents S46 peptidases, where dipeptidyl-peptidase 7 (DPP-7) is the best-characterised member of this family. It is a serine peptidase that is located on the cell surface and is predicted to have two N-terminal transmembrane domains. 
Probab=96.31  E-value=0.0048  Score=69.80  Aligned_cols=55  Identities=25%  Similarity=0.349  Sum_probs=31.2

Q ss_pred             EEEcccCCCCCCCCeeecCCCeEEEEEeeccccCc--------cccccccccHHHHHHHHHHH
Q 009784          264 LQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHED--------VENIGYVIPTPVIMHFIQDY  318 (526)
Q Consensus       264 i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~--------~~~~~~aIP~~~i~~~l~~l  318 (526)
                      +.++..|..||||||++|.+|++||+++-+.-++-        ..+.+..+.+..|+.+|+++
T Consensus       624 FlstnDitGGNSGSPvlN~~GeLVGl~FDgn~Esl~~D~~fdp~~~R~I~VDiRyvL~~ldkv  686 (698)
T PF10459_consen  624 FLSTNDITGGNSGSPVLNAKGELVGLAFDGNWESLSGDIAFDPELNRTIHVDIRYVLWALDKV  686 (698)
T ss_pred             EEeccCcCCCCCCCccCCCCceEEEEeecCchhhcccccccccccceeEEEEHHHHHHHHHHH
Confidence            44556667777777777777777777764432111        11223345555566666554


No 58 
>PF02122 Peptidase_S39:  Peptidase S39;  InterPro: IPR000382 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. ORF2 of Potato leafroll virus (PLrV) encodes a polyprotein which is translated following a -1 frameshift. The polyprotein has a putative linear arrangement of membrane achor-VPg-peptidase-polmerase domains. The serine peptidase domain which is found in this group of sequences belongs to MEROPS peptidase family S39 (clan PA(S)). It is likely that the peptidase domain is involved in the cleavage of the polyprotein []. The nucleotide sequence for the RNA of PLrV has been determined [, ]. The sequence contains six large open reading frames (ORFs). The 5' coding region encodes two polypeptides of 28K and 70K, which overlap in different reading frames; it is suggested that the third ORF in the 5' block is translated by frameshift readthrough near the end of the 70K protein, yielding a 118K polypeptide []. Segments of the predicted amino acid sequences of these ORFs resemble those of known viral RNA polymerases, ATP-binding proteins and viral genome-linked proteins. The nucleotide sequence of the genomic RNA of Beet western yellows virus (BWYV) has been determined []. The sequence contains six long ORFs. A cluster of three of these ORFs, including the coat protein cistron, display extensive amino acid sequence similarity to corresponding ORFs of a second luteovirus: Barley yellow dwarf virus [].; GO: 0004252 serine-type endopeptidase activity, 0022415 viral reproductive process, 0016021 integral to membrane; PDB: 1ZYO_A.
Probab=96.19  E-value=0.0035  Score=60.38  Aligned_cols=137  Identities=22%  Similarity=0.223  Sum_probs=49.3

Q ss_pred             CEEEecccccCCCCeEEEEEcCCCcEEE---EEEEEeccCCCeEEEEeccccc-ccCceeeecCCCCcCC-CcEEEEeeC
Q 009784          159 RRVLTNAHSVEHYTQVKLKKRGSDTKYL---ATVLAIGTECDIAMLTVEDDEF-WEGVLPVEFGELPALQ-DAVTVVGYP  233 (526)
Q Consensus       159 g~ILT~aHvV~~~~~i~V~~~~~g~~~~---a~vv~~d~~~DlAlLkv~~~~~-~~~~~pl~l~~~~~~g-~~V~~iG~p  233 (526)
                      ..++|++||......+....  +|+.++   -+.+..+...|++||+..+.-. .-.++.+.+.....+. ..+.+.++.
T Consensus        42 ~~L~ta~Hv~~~~~~~~~~k--~g~kipl~~f~~~~~~~~~D~~il~~P~n~~s~Lg~k~~~~~~~~~~~~g~~~~y~~~  119 (203)
T PF02122_consen   42 DALLTARHVWSRPSKVTSLK--TGEKIPLAEFTDLLESRIADFVILRGPPNWESKLGVKAAQLSQNSQLAKGPVSFYGFS  119 (203)
T ss_dssp             EEEEE-HHHHTSSS---EEE--TTEEEE--S-EEEEE-TTT-EEEEE--HHHHHHHT-----B----SEEEEESSTTSEE
T ss_pred             cceecccccCCCccceeEcC--CCCcccchhChhhhCCCccCEEEEecCcCHHHHhCcccccccchhhhCCCCeeeeeec
Confidence            49999999999866665554  455443   3556678899999999984310 1134444443222110 000011110


Q ss_pred             CCCCceeEEEEEEeceeeeeccCCceeeeEEEEcccCCCCCCCCeeecCCCeEEEEEeeccccCccccccccccHHH
Q 009784          234 IGGDTISVTSGVVSRIEILSYVHGSTELLGLQIDAAINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPV  310 (526)
Q Consensus       234 ~~~~~~sv~~GiVs~~~~~~~~~~~~~~~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~~~~~~~aIP~~~  310 (526)
                      .+  ....+..-|.      ...+    .+...-+...+|.||.|+++.+ +++|++.+..+....++.++..|+.-
T Consensus       120 ~~--~~~~~sa~i~------g~~~----~~~~vls~T~~G~SGtp~y~g~-~vvGvH~G~~~~~~~~n~n~~spip~  183 (203)
T PF02122_consen  120 SG--EWPCSSAKIP------GTEG----KFASVLSNTSPGWSGTPYYSGK-NVVGVHTGSPSGSNRENNNRMSPIPP  183 (203)
T ss_dssp             EE--EEEEEE-S----------ST----TEEEE-----TT-TT-EEE-SS--EEEEEEEE-----------------
T ss_pred             CC--CceeccCccc------cccC----cCCceEcCCCCCCCCCCeEECC-CceEeecCcccccccccccccccccc
Confidence            00  0111111111      1111    1345556778999999999987 99999998643345566676655433


No 59 
>KOG3209 consensus WW domain-containing protein [General function prediction only]
Probab=95.85  E-value=0.031  Score=61.68  Aligned_cols=140  Identities=16%  Similarity=0.091  Sum_probs=85.9

Q ss_pred             CccceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEE
Q 009784          351 DQKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNI  428 (526)
Q Consensus       351 ~~~Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v  428 (526)
                      ..+-++|..|.+.+.|++  -|++||-|+.|||.+|.....-.     -+   .++.......-|.|+|.|.-..-.   
T Consensus       672 p~qpi~iG~Iv~lGaAe~DGRL~~gDElv~iDG~pV~GksH~~-----vv---~Lm~~AArnghV~LtVRRkv~~~~---  740 (984)
T KOG3209|consen  672 PGQPIYIGAIVPLGAAEEDGRLREGDELVCIDGIPVEGKSHSE-----VV---DLMEAAARNGHVNLTVRRKVRTGP---  740 (984)
T ss_pred             CCCeeEEeeeeecccccccCcccCCCeEEEecCeeccCccHHH-----HH---HHHHHHHhcCceEEEEeeeeeecc---
Confidence            456699999999999998  49999999999999998765421     11   333333334569999988421110   


Q ss_pred             Eecccc--cccCCCCCCCCCceEEEeeEEEEecccc-ceeeeeeeecchh--hh-ccccccceeeeecccchhhhHHHHH
Q 009784          429 TLATHR--RLIPSHNKGRPPSYYIIAGFVFSRCLYL-ISVLSMERIMNMK--LR-SSFWTSSCIQCHNCQMSSLLWCLRC  502 (526)
Q Consensus       429 ~l~~~~--~~~p~~~~~~~p~~~i~gG~~f~~lt~~-~~~~~~~~i~~~~--~~-sg~~~~~~~~~~~~~~~~~~~~~~~  502 (526)
                       -...+  ...+...++...+.....||-|.-++.+ .-.-.++||++.+  -| .-|++||.|..+|.+-.-.+.+--.
T Consensus       741 -~~rsp~~s~~~~~~yDV~lhR~ENeGFGFVi~sS~~kp~sgiGrIieGSPAdRCgkLkVGDrilAVNG~sI~~lsHadi  819 (984)
T KOG3209|consen  741 -ARRSPRNSAAPSGPYDVVLHRKENEGFGFVIMSSQNKPESGIGRIIEGSPADRCGKLKVGDRILAVNGQSILNLSHADI  819 (984)
T ss_pred             -ccCCcccccCCCCCeeeEEecccCCceeEEEEecccCCCCCccccccCChhHhhccccccceEEEecCeeeeccCchhH
Confidence             01111  0111112222222223467777777633 2222388998854  23 3589999999999987666555443


No 60 
>KOG3209 consensus WW domain-containing protein [General function prediction only]
Probab=95.83  E-value=0.0092  Score=65.59  Aligned_cols=152  Identities=20%  Similarity=0.119  Sum_probs=88.8

Q ss_pred             EEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEE-------
Q 009784          357 IRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFN-------  427 (526)
Q Consensus       357 V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~-------  427 (526)
                      |.+|.++|||+.  .|+.||.|++|||+.|.+...-.        ...++.  ..|-+|+|+|.-..+.-..+       
T Consensus       782 iGrIieGSPAdRCgkLkVGDrilAVNG~sI~~lsHad--------iv~LIK--daGlsVtLtIip~ee~~~~~~~~sa~~  851 (984)
T KOG3209|consen  782 IGRIIEGSPADRCGKLKVGDRILAVNGQSILNLSHAD--------IVSLIK--DAGLSVTLTIIPPEEAGPPTSMTSAEK  851 (984)
T ss_pred             ccccccCChhHhhccccccceEEEecCeeeeccCchh--------HHHHHH--hcCceEEEEEcChhccCCCCCCcchhh
Confidence            778999999999  69999999999999999876531        113333  36899999997644321111       


Q ss_pred             ---EEec----ccccccCC----CCCCCCC----------ceEEEeeEEEEecc--------------ccceeeeeeeec
Q 009784          428 ---ITLA----THRRLIPS----HNKGRPP----------SYYIIAGFVFSRCL--------------YLISVLSMERIM  472 (526)
Q Consensus       428 ---v~l~----~~~~~~p~----~~~~~~p----------~~~i~gG~~f~~lt--------------~~~~~~~~~~i~  472 (526)
                         ++..    ..-.+.+.    .....+|          +.-..+++.-.+|.              ....-|.+-|+.
T Consensus       852 ~s~~t~~~~~~q~~glp~~~~s~~~~~pqpdt~~~~~~~~r~~qn~~~~~VelErG~kGFGFSiRGGreynM~LfVLRlA  931 (984)
T KOG3209|consen  852 QSPFTQNGPYEQQYGLPGPRPSVYEEHPQPDTFQGLSINDRMSQNGDLYTVELERGAKGFGFSIRGGREYNMDLFVLRLA  931 (984)
T ss_pred             cCcccccCCHhHccCCCCCCccccccCCCCccccceeccccccccCCeeEEEeeccccccceEeecccccccceEEEEec
Confidence               1100    00000000    0001111          11111222222222              112233455554


Q ss_pred             c--hhhhcc-ccccceeeeecccchhhhHHHHHHHHHHHHHhhhHHHHH
Q 009784          473 N--MKLRSS-FWTSSCIQCHNCQMSSLLWCLRCLWLILILDMRRLLTLR  518 (526)
Q Consensus       473 ~--~~~~sg-~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~  518 (526)
                      +  .++|.| +++||-|...|.+--...-+-|.+=||----||-||.||
T Consensus       932 eDGPA~rdGrm~VGDqi~eINGesTkgmtH~rAIelIk~gg~~vll~Lr  980 (984)
T KOG3209|consen  932 EDGPAIRDGRMRVGDQITEINGESTKGMTHDRAIELIKQGGRRVLLLLR  980 (984)
T ss_pred             cCCCccccCceeecceEEEecCcccCCCcHHHHHHHHHhCCeEEEEEec
Confidence            4  555665 567999999999988888888988888766666555444


No 61 
>KOG3580 consensus Tight junction proteins [Signal transduction mechanisms]
Probab=95.51  E-value=0.011  Score=63.89  Aligned_cols=61  Identities=18%  Similarity=0.262  Sum_probs=46.8

Q ss_pred             CccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEE
Q 009784          351 DQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLR  419 (526)
Q Consensus       351 ~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R  419 (526)
                      ++-|+.|..|..++||++ |||.||.||+||.++..+.--=     +   ...++....+|+.++|.-.+
T Consensus       427 NDVGIFVaGvqegspA~~eGlqEGDQIL~VN~vdF~nl~RE-----e---AVlfLL~lPkGEevtilaQ~  488 (1027)
T KOG3580|consen  427 NDVGIFVAGVQEGSPAEQEGLQEGDQILKVNTVDFRNLVRE-----E---AVLFLLELPKGEEVTILAQS  488 (1027)
T ss_pred             CceeEEEeecccCCchhhccccccceeEEeccccchhhhHH-----H---HHHHHhcCCCCcEEeehhhh
Confidence            467999999999999999 9999999999999987764210     0   11345566789988886544


No 62 
>KOG3580 consensus Tight junction proteins [Signal transduction mechanisms]
Probab=94.72  E-value=0.039  Score=59.75  Aligned_cols=74  Identities=23%  Similarity=0.279  Sum_probs=52.4

Q ss_pred             hhcCCCCccceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCE
Q 009784          345 AMSMKADQKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK  422 (526)
Q Consensus       345 ~lgl~~~~~Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~  422 (526)
                      .||+. -.+-+.|.++...+-|++  +||.||+||+|||....|..--   +     . ..+..+..| ++.|.|+||.+
T Consensus       212 EyGlr-LgSqIFvKeit~~gLAardgnlqEGDiiLkINGtvteNmSLt---D-----a-r~LIEkS~G-KL~lvVlRD~~  280 (1027)
T KOG3580|consen  212 EYGLR-LGSQIFVKEITRTGLAARDGNLQEGDIILKINGTVTENMSLT---D-----A-RKLIEKSRG-KLQLVVLRDSQ  280 (1027)
T ss_pred             hhccc-ccchhhhhhhcccchhhccCCcccccEEEEECcEeeccccch---h-----H-HHHHHhccC-ceEEEEEecCC
Confidence            45554 234578899988887776  8999999999999988775421   1     1 233444445 69999999987


Q ss_pred             EEEEEEE
Q 009784          423 ILNFNIT  429 (526)
Q Consensus       423 ~~~~~v~  429 (526)
                      ..-++|.
T Consensus       281 qtLiNiP  287 (1027)
T KOG3580|consen  281 QTLINIP  287 (1027)
T ss_pred             ceeeecC
Confidence            7666665


No 63 
>PF00949 Peptidase_S7:  Peptidase S7, Flavivirus NS3 serine protease ;  InterPro: IPR001850 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This signature identifies serine peptidases belong to MEROPS peptidase family S7 (flavivirin family, clan PA(S)). The protein fold of the peptidase domain for members of this family resembles that of chymotrypsin, the type example for clan PA.  Flaviviruses produce a polyprotein from the ssRNA genome. The N terminus of the NS3 protein (approx. 180 aa) is required for the processing of the polyprotein. NS3 also has conserved homology with NTP-binding proteins and DEAD family of RNA helicase [, , ].; GO: 0003723 RNA binding, 0003724 RNA helicase activity, 0005524 ATP binding; PDB: 2IJO_B 3E90_D 2GGV_B 2FP7_B 2WV9_A 3U1I_B 3U1J_B 2WZQ_A 2WHX_A 3L6P_A ....
Probab=94.61  E-value=0.043  Score=49.20  Aligned_cols=29  Identities=38%  Similarity=0.672  Sum_probs=21.6

Q ss_pred             EcccCCCCCCCCeeecCCCeEEEEEeecc
Q 009784          266 IDAAINSGNSGGPAFNDKGKCVGIAFQSL  294 (526)
Q Consensus       266 ~da~i~~G~SGGPlvn~~G~VVGI~~~~~  294 (526)
                      .+..+.+|+||+|+||.+|++|||...+.
T Consensus        90 ~~~d~~~GsSGSpi~n~~g~ivGlYg~g~  118 (132)
T PF00949_consen   90 IDLDFPKGSSGSPIFNQNGEIVGLYGNGV  118 (132)
T ss_dssp             E---S-TTGTT-EEEETTSCEEEEEEEEE
T ss_pred             eecccCCCCCCCceEcCCCcEEEEEccce
Confidence            34447799999999999999999987765


No 64 
>KOG3550 consensus Receptor targeting protein Lin-7 [Extracellular structures]
Probab=94.04  E-value=0.1  Score=47.22  Aligned_cols=36  Identities=25%  Similarity=0.448  Sum_probs=32.4

Q ss_pred             ccceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCC
Q 009784          352 QKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIAND  387 (526)
Q Consensus       352 ~~Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~  387 (526)
                      .+-++|+.+.|++-|+.  ||+-||.+++|||..+...
T Consensus       114 nspiyisriipggvadrhgglkrgdqllsvngvsvege  151 (207)
T KOG3550|consen  114 NSPIYISRIIPGGVADRHGGLKRGDQLLSVNGVSVEGE  151 (207)
T ss_pred             CCceEEEeecCCccccccCcccccceeEeecceeecch
Confidence            45699999999999988  9999999999999998753


No 65 
>KOG3834 consensus Golgi reassembly stacking protein GRASP65, contains PDZ domain [Intracellular trafficking, secretion, and vesicular transport]
Probab=93.27  E-value=0.36  Score=50.84  Aligned_cols=114  Identities=16%  Similarity=0.238  Sum_probs=73.6

Q ss_pred             ccceEEEEeCCCCcccC-CCCC-CcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECC--EEEEEE
Q 009784          352 QKGVRIRRVDPTAPESE-VLKP-SDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDS--KILNFN  427 (526)
Q Consensus       352 ~~Gv~V~~V~~~spA~~-GL~~-GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G--~~~~~~  427 (526)
                      ..|.-|-+|..+|++++ ||++ -|.|++|||..++...|..          +.+.+.+.. +|+++|+-..  ..++++
T Consensus        14 teg~hvlkVqedSpa~~aglepffdFIvSI~g~rL~~dnd~L----------k~llk~~se-kVkltv~n~kt~~~R~v~   82 (462)
T KOG3834|consen   14 TEGYHVLKVQEDSPAHKAGLEPFFDFIVSINGIRLNKDNDTL----------KALLKANSE-KVKLTVYNSKTQEVRIVE   82 (462)
T ss_pred             ceeEEEEEeecCChHHhcCcchhhhhhheeCcccccCchHHH----------HHHHHhccc-ceEEEEEecccceeEEEE
Confidence            45788999999999999 9988 5899999999999877642          333333333 3999987643  334444


Q ss_pred             EEecccccccCCCCCCCCCceEEEeeEEEEeccccceeeeeeeecc-----hhhhcccc-ccceeeee
Q 009784          428 ITLATHRRLIPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERIMN-----MKLRSSFW-TSSCIQCH  489 (526)
Q Consensus       428 v~l~~~~~~~p~~~~~~~p~~~i~gG~~f~~lt~~~~~~~~~~i~~-----~~~~sg~~-~~~~~~~~  489 (526)
                      |+......        .  +   +-|....-.++...++.+|-|+.     .+..+||. ++|-|.-.
T Consensus        83 I~ps~~wg--------g--q---llGvsvrFcsf~~A~~~vwHvl~V~p~SPaalAgl~~~~DYivG~  137 (462)
T KOG3834|consen   83 IVPSNNWG--------G--Q---LLGVSVRFCSFDGAVESVWHVLSVEPNSPAALAGLRPYTDYIVGI  137 (462)
T ss_pred             eccccccc--------c--c---ccceEEEeccCccchhheeeeeecCCCCHHHhcccccccceEecc
Confidence            44333211        0  1   23555555555556666666554     34459999 78877765


No 66 
>PF09342 DUF1986:  Domain of unknown function (DUF1986);  InterPro: IPR015420 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This domain is found in serine endopeptidases belonging to MEROPS peptidase family S1A (clan PA). It is found in unusual mosaic proteins, which are encoded by the Drosophila nudel gene (see P98159 from SWISSPROT). Nudel is involved in defining embryonic dorsoventral polarity. Three proteases; ndl, gd and snk process easter to create active easter. Active easter defines cell identities along the dorsal-ventral continuum by activating the spz ligand for the Tl receptor in the ventral region of the embryo. Nudel, pipe and windbeutel together trigger the protease cascade within the extraembryonic perivitelline compartment which induces dorsoventral polarity of the Drosophila embryo [].
Probab=93.19  E-value=1.2  Score=43.90  Aligned_cols=99  Identities=19%  Similarity=0.277  Sum_probs=69.7

Q ss_pred             CCCCCccccCC--CcceEEEEEEEeCCEEEecccccCCC----CeEEEEEcCCCcEEE------EEEEEec-----cCCC
Q 009784          135 PNFSLPWQRKR--QYSSSSSGFAIGGRRVLTNAHSVEHY----TQVKLKKRGSDTKYL------ATVLAIG-----TECD  197 (526)
Q Consensus       135 ~~~~~P~~~~~--~~~~~GSGfvI~~g~ILT~aHvV~~~----~~i~V~~~~~g~~~~------a~vv~~d-----~~~D  197 (526)
                      ..+..||....  .+...|+|++|+..|||++..|+.+-    ..+.+.+. .++.+.      -++..+|     ++.+
T Consensus        12 e~y~WPWlA~IYvdG~~~CsgvLlD~~WlLvsssCl~~I~L~~~YvsallG-~~Kt~~~v~Gp~EQI~rVD~~~~V~~S~   90 (267)
T PF09342_consen   12 EDYHWPWLADIYVDGRYWCSGVLLDPHWLLVSSSCLRGISLSHHYVSALLG-GGKTYLSVDGPHEQISRVDCFKDVPESN   90 (267)
T ss_pred             ccccCcceeeEEEcCeEEEEEEEeccceEEEeccccCCcccccceEEEEec-CcceecccCCChheEEEeeeeeeccccc
Confidence            35667887743  45568999999999999999999863    45677774 565443      2444444     5789


Q ss_pred             eEEEEecccc-cccCceeeecCCC---CcCCCcEEEEeeCC
Q 009784          198 IAMLTVEDDE-FWEGVLPVEFGEL---PALQDAVTVVGYPI  234 (526)
Q Consensus       198 lAlLkv~~~~-~~~~~~pl~l~~~---~~~g~~V~~iG~p~  234 (526)
                      ++||.++.+. |...+.|+-+.+.   ....+.++++|...
T Consensus        91 v~LLHL~~~~~fTr~VlP~flp~~~~~~~~~~~CVAVg~d~  131 (267)
T PF09342_consen   91 VLLLHLEQPANFTRYVLPTFLPETSNENESDDECVAVGHDD  131 (267)
T ss_pred             eeeeeecCcccceeeecccccccccCCCCCCCceEEEEccc
Confidence            9999998764 4455677656541   12346899999877


No 67 
>KOG3532 consensus Predicted protein kinase [General function prediction only]
Probab=92.99  E-value=0.1  Score=57.41  Aligned_cols=39  Identities=15%  Similarity=0.474  Sum_probs=35.2

Q ss_pred             cceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcc
Q 009784          353 KGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVP  391 (526)
Q Consensus       353 ~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~  391 (526)
                      .-|.|-.|.++++|.+ .|++|||+++|||.+|.+..+..
T Consensus       398 ~~v~v~tv~~ns~a~k~~~~~gdvlvai~~~pi~s~~q~~  437 (1051)
T KOG3532|consen  398 RAVKVCTVEDNSLADKAAFKPGDVLVAINNVPIRSERQAT  437 (1051)
T ss_pred             eEEEEEEecCCChhhHhcCCCcceEEEecCccchhHHHHH
Confidence            4577889999999999 99999999999999999987753


No 68 
>KOG2921 consensus Intramembrane metalloprotease (sterol-regulatory element-binding protein (SREBP) protease) [Posttranslational modification, protein turnover, chaperones]
Probab=91.76  E-value=0.12  Score=53.76  Aligned_cols=40  Identities=25%  Similarity=0.341  Sum_probs=36.8

Q ss_pred             CccceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCc
Q 009784          351 DQKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTV  390 (526)
Q Consensus       351 ~~~Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l  390 (526)
                      +..|+.|++|...||+..  ||++||+|+++||-+|++.+|.
T Consensus       218 ~g~gV~Vtev~~~Spl~gprGL~vgdvitsldgcpV~~v~dW  259 (484)
T KOG2921|consen  218 HGEGVTVTEVPSVSPLFGPRGLSVGDVITSLDGCPVHKVSDW  259 (484)
T ss_pred             cCceEEEEeccccCCCcCcccCCccceEEecCCcccCCHHHH
Confidence            467999999999999987  9999999999999999998775


No 69 
>KOG3605 consensus Beta amyloid precursor-binding protein [General function prediction only]
Probab=91.12  E-value=0.37  Score=53.03  Aligned_cols=103  Identities=20%  Similarity=0.263  Sum_probs=69.9

Q ss_pred             CCCCCCeee-----cCCCeEEEEEeeccccCccccccccccHHHHHHHHHHHHHcCceeeccc---cCceeeeccChhHH
Q 009784          272 SGNSGGPAF-----NDKGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDYEKNGAYTGFPL---LGVEWQKMENPDLR  343 (526)
Q Consensus       272 ~G~SGGPlv-----n~~G~VVGI~~~~~~~~~~~~~~~aIP~~~i~~~l~~l~~~g~~~~~~~---LGi~~~~~~~~~~~  343 (526)
                      .=++|||.-     |.-.+++.|+-..+         ..+|.+..+.+++.++..-.+. +-.   --+.-..+..|+.+
T Consensus       679 nmm~~GpAarsgkLnIGDQiiaING~SL---------VGLPLstcQs~Ik~~KnQT~Vk-ltiV~cpPV~~V~I~RPd~k  748 (829)
T KOG3605|consen  679 NMMHGGPAARSGKLNIGDQIMSINGTSL---------VGLPLSTCQSIIKGLKNQTAVK-LNIVSCPPVTTVLIRRPDLR  748 (829)
T ss_pred             hcccCChhhhcCCccccceeEeecCcee---------ccccHHHHHHHHhcccccceEE-EEEecCCCceEEEeecccch
Confidence            456777753     44445666552221         3489999999999886533332 111   11222233478888


Q ss_pred             hhhcCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecC
Q 009784          344 VAMSMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIAN  386 (526)
Q Consensus       344 ~~lgl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~  386 (526)
                      ..||++ .+.|+ |=.+..++-|++ |++.|-.|++|||+.|--
T Consensus       749 yQLGFS-VQNGi-ICSLlRGGIAERGGVRVGHRIIEINgQSVVA  790 (829)
T KOG3605|consen  749 YQLGFS-VQNGI-ICSLLRGGIAERGGVRVGHRIIEINGQSVVA  790 (829)
T ss_pred             hhccce-eeCcE-eehhhcccchhccCceeeeeEEEECCceEEe
Confidence            889997 67787 456889999999 999999999999998854


No 70 
>PF00944 Peptidase_S3:  Alphavirus core protein ;  InterPro: IPR000930 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. Togavirin, also known as Sindbis virus core endopeptidase, is a serine protease resident at the N terminus of the p130 polyprotein of togaviruses []. The endopeptidase signature identifies the peptidase as belonging to the MEROPS peptidase family S3 (togavirin family, clan PA(S)). The polyprotein also includes structural proteins for the nucleocapsid core and for the glycoprotein spikes []. Togavirin is only active while part of the polyprotein, cleavage at a Trp-Ser bond resulting in total lack of activity []. Mutagenesis studies have identified the location of the His-Asp-Ser catalytic triad, and X-ray studies have revealed the protein fold to be similar to that of chymotrypsin [, ].; GO: 0004252 serine-type endopeptidase activity, 0006508 proteolysis, 0016020 membrane; PDB: 2YEW_D 1EP5_A 3J0C_F 1EP6_C 1WYK_D 1DYL_A 1VCQ_B 1VCP_B 1LD4_D 1KXA_A ....
Probab=90.80  E-value=0.18  Score=44.82  Aligned_cols=27  Identities=30%  Similarity=0.614  Sum_probs=23.3

Q ss_pred             ccCCCCCCCCeeecCCCeEEEEEeecc
Q 009784          268 AAINSGNSGGPAFNDKGKCVGIAFQSL  294 (526)
Q Consensus       268 a~i~~G~SGGPlvn~~G~VVGI~~~~~  294 (526)
                      ..-.+|+||-|++|..|+||||+.++.
T Consensus       101 g~g~~GDSGRpi~DNsGrVVaIVLGG~  127 (158)
T PF00944_consen  101 GVGKPGDSGRPIFDNSGRVVAIVLGGA  127 (158)
T ss_dssp             TS-STTSTTEEEESTTSBEEEEEEEEE
T ss_pred             CCCCCCCCCCccCcCCCCEEEEEecCC
Confidence            345689999999999999999998865


No 71 
>PF05580 Peptidase_S55:  SpoIVB peptidase S55;  InterPro: IPR008763 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to the MEROPS peptidase family S55 (SpoIVB peptidase family, clan PA(S)). The protein SpoIVB plays a key role in signalling in the final sigma-K checkpoint of Bacillus subtilis [, ].
Probab=90.44  E-value=6.6  Score=38.13  Aligned_cols=41  Identities=27%  Similarity=0.394  Sum_probs=32.4

Q ss_pred             ccCCCCCCCCeeecCCCeEEEEEeeccccCccccccccccHHHH
Q 009784          268 AAINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPVI  311 (526)
Q Consensus       268 a~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~~~~~~~aIP~~~i  311 (526)
                      ..+-.|+||+|++- +|++||-++..+  -+....||.++++..
T Consensus       175 GGIvqGMSGSPI~q-dGKLiGAVthvf--~~dp~~Gygi~ie~M  215 (218)
T PF05580_consen  175 GGIVQGMSGSPIIQ-DGKLIGAVTHVF--VNDPTKGYGIFIEWM  215 (218)
T ss_pred             CCEEecccCCCEEE-CCEEEEEEEEEE--ecCCCceeeecHHHH
Confidence            35678999999985 899999987766  345778899987653


No 72 
>COG0750 Predicted membrane-associated Zn-dependent proteases 1 [Cell envelope biogenesis, outer membrane]
Probab=90.42  E-value=0.47  Score=49.93  Aligned_cols=57  Identities=28%  Similarity=0.407  Sum_probs=45.2

Q ss_pred             EEeCCCCcccC-CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCE---EEEEEEE-CCEEEE
Q 009784          358 RRVDPTAPESE-VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDS---AAVKVLR-DSKILN  425 (526)
Q Consensus       358 ~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~---v~l~v~R-~G~~~~  425 (526)
                      .++..++++.. |+++||.|+++|++++.++.++.          ..+. ...|..   +.+.+.| +++...
T Consensus       134 ~~v~~~s~a~~a~l~~Gd~iv~~~~~~i~~~~~~~----------~~~~-~~~~~~~~~~~i~~~~~~~~~~~  195 (375)
T COG0750         134 GEVAPKSAAALAGLRPGDRIVAVDGEKVASWDDVR----------RLLV-AAAGDVFNLLTILVIRLDGEAHA  195 (375)
T ss_pred             eecCCCCHHHHcCCCCCCEEEeECCEEccCHHHHH----------HHHH-hccCCcccceEEEEEeccceeee
Confidence            37889999999 99999999999999999998863          3333 334555   8999999 777743


No 73 
>KOG1892 consensus Actin filament-binding protein Afadin [Cytoskeleton]
Probab=90.31  E-value=0.34  Score=55.36  Aligned_cols=61  Identities=20%  Similarity=0.299  Sum_probs=47.2

Q ss_pred             CccceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECC
Q 009784          351 DQKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDS  421 (526)
Q Consensus       351 ~~~Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G  421 (526)
                      +.-|++|..|.+|++|+.  -|+.||.+++|||+.+-...+=.        ..+++  ...|..|.|+|-..|
T Consensus       958 ~klGIYvKsVV~GgaAd~DGRL~aGDQLLsVdG~SLiGisQEr--------AA~lm--trtg~vV~leVaKqg 1020 (1629)
T KOG1892|consen  958 RKLGIYVKSVVEGGAADHDGRLEAGDQLLSVDGHSLIGISQER--------AARLM--TRTGNVVHLEVAKQG 1020 (1629)
T ss_pred             cccceEEEEeccCCccccccccccCceeeeecCcccccccHHH--------HHHHH--hccCCeEEEehhhhh
Confidence            456999999999999987  59999999999999887665521        11222  346889999987655


No 74 
>PF02907 Peptidase_S29:  Hepatitis C virus NS3 protease;  InterPro: IPR004109 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This signature identifies the Hepatitis C virus NS3 protein as a serine protease which belongs to MEROPS peptidase family S29 (hepacivirin family, clan PA(S)), which has a trypsin-like fold. The non-structural (NS) protein NS3 is one of the NS proteins involved in replication of the HCV genome. The NS2 proteinase (IPR002518 from INTERPRO), a zinc-dependent enzyme, performs a single proteolytic cut to release the N terminus of NS3. The action of NS3 proteinase (NS3P), which resides in the N-terminal one-third of the NS3 protein, then yields all remaining non-structural proteins. The C-terminal two-thirds of the NS3 protein contain a helicase. The functional relationship between the proteinase and helicase domains is unknown. NS3 has a structural zinc-binding site and requires cofactor NS4. It has been suggested that the NS3 serine protease of hepatitus C is involved in cell transformation and that the ability to transform requires an active enzyme [].; GO: 0008236 serine-type peptidase activity, 0006508 proteolysis, 0019087 transformation of host cell by virus; PDB: 2QV1_B 3LOX_C 2OBQ_C 2OC1_C 2OC0_A 3LON_A 3KNX_A 2O8M_A 2OBO_A 2OC8_A ....
Probab=90.28  E-value=0.37  Score=42.93  Aligned_cols=41  Identities=29%  Similarity=0.528  Sum_probs=25.8

Q ss_pred             CCCCCCCCeeecCCCeEEEEEeeccccCcc-ccccccccHHHH
Q 009784          270 INSGNSGGPAFNDKGKCVGIAFQSLKHEDV-ENIGYVIPTPVI  311 (526)
Q Consensus       270 i~~G~SGGPlvn~~G~VVGI~~~~~~~~~~-~~~~~aIP~~~i  311 (526)
                      ...|+||||++..+|.+|||..+.....+. ..+-| +|.+.+
T Consensus       105 ~lkGSSGgPiLC~~GH~vG~f~aa~~trgvak~i~f-~P~e~l  146 (148)
T PF02907_consen  105 DLKGSSGGPILCPSGHAVGMFRAAVCTRGVAKAIDF-IPVETL  146 (148)
T ss_dssp             HHTT-TT-EEEETTSEEEEEEEEEEEETTEEEEEEE-EEHHHH
T ss_pred             EEecCCCCcccCCCCCEEEEEEEEEEcCCceeeEEE-Eeeeec
Confidence            347999999999999999997665432222 23333 376543


No 75 
>KOG3542 consensus cAMP-regulated guanine nucleotide exchange factor [Signal transduction mechanisms]
Probab=90.21  E-value=0.21  Score=54.94  Aligned_cols=42  Identities=24%  Similarity=0.279  Sum_probs=35.6

Q ss_pred             cCCCCccceEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCC
Q 009784          347 SMKADQKGVRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDG  388 (526)
Q Consensus       347 gl~~~~~Gv~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~  388 (526)
                      |=.+...|++|.+|.|++.|+. ||+-||.|++|||+...+..
T Consensus       556 GGsEkGfgifV~~V~pgskAa~~GlKRgDqilEVNgQnfenis  598 (1283)
T KOG3542|consen  556 GGSEKGFGIFVAEVFPGSKAAREGLKRGDQILEVNGQNFENIS  598 (1283)
T ss_pred             cCccccceeEEeeecCCchHHHhhhhhhhhhhhccccchhhhh
Confidence            3344567999999999999998 99999999999999776543


No 76 
>KOG3552 consensus FERM domain protein FRM-8 [General function prediction only]
Probab=89.43  E-value=0.32  Score=55.49  Aligned_cols=57  Identities=28%  Similarity=0.353  Sum_probs=43.0

Q ss_pred             ceEEEEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEEC
Q 009784          354 GVRIRRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD  420 (526)
Q Consensus       354 Gv~V~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~  420 (526)
                      -|+|..|.+|+|+...|++||.|+.|||+++...--      +|  ..+++..  -.+.|.|+|.+-
T Consensus        76 PviVr~VT~GGps~GKL~PGDQIl~vN~Epv~dapr------er--vIdlvRa--ce~sv~ltV~qP  132 (1298)
T KOG3552|consen   76 PVIVRFVTEGGPSIGKLQPGDQILAVNGEPVKDAPR------ER--VIDLVRA--CESSVNLTVCQP  132 (1298)
T ss_pred             ceEEEEecCCCCccccccCCCeEEEecCcccccccH------HH--HHHHHHH--HhhhcceEEecc
Confidence            388999999999999999999999999999975431      11  1133333  356789998884


No 77 
>KOG3571 consensus Dishevelled 3 and related proteins [General function prediction only]
Probab=87.51  E-value=0.66  Score=49.78  Aligned_cols=38  Identities=16%  Similarity=0.338  Sum_probs=32.7

Q ss_pred             ccceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCC
Q 009784          352 QKGVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGT  389 (526)
Q Consensus       352 ~~Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~  389 (526)
                      +.|++|.+|.+++.-+.  -+++||.||.||.....++..
T Consensus       276 DggIYVgsImkgGAVA~DGRIe~GDMiLQVNevsFENmSN  315 (626)
T KOG3571|consen  276 DGGIYVGSIMKGGAVALDGRIEPGDMILQVNEVSFENMSN  315 (626)
T ss_pred             CCceEEeeeccCceeeccCccCccceEEEeeecchhhcCc
Confidence            57999999999998666  599999999999988777653


No 78 
>KOG3551 consensus Syntrophins (type beta) [Extracellular structures]
Probab=87.27  E-value=0.4  Score=49.78  Aligned_cols=56  Identities=20%  Similarity=0.280  Sum_probs=40.7

Q ss_pred             eEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEE--EEEC
Q 009784          355 VRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVK--VLRD  420 (526)
Q Consensus       355 v~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~--v~R~  420 (526)
                      ++|+++-++-.|++  .|..||.|++|||..+.+...-+          ..-.-++.|++|.++  +.|+
T Consensus       112 IlISKIFkGlAADQt~aL~~gDaIlSVNG~dL~~AtHde----------AVqaLKraGkeV~levKy~RE  171 (506)
T KOG3551|consen  112 ILISKIFKGLAADQTGALFLGDAILSVNGEDLRDATHDE----------AVQALKRAGKEVLLEVKYMRE  171 (506)
T ss_pred             eehhHhccccccccccceeeccEEEEecchhhhhcchHH----------HHHHHHhhCceeeeeeeeehh
Confidence            88999999999988  79999999999999887654311          111224468876654  4554


No 79 
>PF00947 Pico_P2A:  Picornavirus core protein 2A;  InterPro: IPR000081 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad [].  This domain defines cysteine peptidases belong to MEROPS peptidase family C3 (picornain, clan PA(C)), subfamilies 3CA and 3CB. The protein fold of this peptidase domain for members of this family resembles that of the serine peptidase, chymotrypsin [], the type example for clan PA. Picornaviral proteins are expressed as a single polyprotein which is cleaved by the viral 3C cysteine protease []. The poliovirus polyprotein is selectively cleaved between the Gln-|-Gly bond. In other picornavirus reactions Glu may be substituted for Gln, and Ser or Thr for Gly. ; GO: 0008233 peptidase activity, 0006508 proteolysis, 0016032 viral reproduction; PDB: 2HRV_B 1Z8R_A.
Probab=86.68  E-value=2  Score=38.02  Aligned_cols=31  Identities=19%  Similarity=0.209  Sum_probs=23.9

Q ss_pred             eEEEEcccCCCCCCCCeeecCCCeEEEEEeec
Q 009784          262 LGLQIDAAINSGNSGGPAFNDKGKCVGIAFQS  293 (526)
Q Consensus       262 ~~i~~da~i~~G~SGGPlvn~~G~VVGI~~~~  293 (526)
                      .++....+..||+-||+|+... -||||++++
T Consensus        79 ~~l~g~Gp~~PGdCGg~L~C~H-GViGi~Tag  109 (127)
T PF00947_consen   79 NLLIGEGPAEPGDCGGILRCKH-GVIGIVTAG  109 (127)
T ss_dssp             CEEEEE-SSSTT-TCSEEEETT-CEEEEEEEE
T ss_pred             CceeecccCCCCCCCceeEeCC-CeEEEEEeC
Confidence            4566667889999999999865 499999985


No 80 
>KOG3605 consensus Beta amyloid precursor-binding protein [General function prediction only]
Probab=84.88  E-value=2.3  Score=47.15  Aligned_cols=119  Identities=15%  Similarity=0.123  Sum_probs=70.1

Q ss_pred             EeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCEEEEEEEEecccccc
Q 009784          359 RVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSKILNFNITLATHRRL  436 (526)
Q Consensus       359 ~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~~~~~~v~l~~~~~~  436 (526)
                      ....++||++  .|-.||.|++|||..+-..---.        .+.++...+.-..|+|+|.+=--..++.|.  + +..
T Consensus       679 nmm~~GpAarsgkLnIGDQiiaING~SLVGLPLst--------cQs~Ik~~KnQT~VkltiV~cpPV~~V~I~--R-Pd~  747 (829)
T KOG3605|consen  679 NMMHGGPAARSGKLNIGDQIMSINGTSLVGLPLST--------CQSIIKGLKNQTAVKLNIVSCPPVTTVLIR--R-PDL  747 (829)
T ss_pred             hcccCChhhhcCCccccceeEeecCceeccccHHH--------HHHHHhcccccceEEEEEecCCCceEEEee--c-ccc
Confidence            5567899998  69999999999998775321111        123455554445688888875444443332  1 111


Q ss_pred             cCCCCCCCCCceEEEeeEEEEeccccceeeeeeeecchhhhccccccceeeeecccchhhhHHHHHHHH
Q 009784          437 IPSHNKGRPPSYYIIAGFVFSRCLYLISVLSMERIMNMKLRSSFWTSSCIQCHNCQMSSLLWCLRCLWL  505 (526)
Q Consensus       437 ~p~~~~~~~p~~~i~gG~~f~~lt~~~~~~~~~~i~~~~~~sg~~~~~~~~~~~~~~~~~~~~~~~~~~  505 (526)
                                +|  -.||.+++=    -...+.|- .-+.|-|+++|-.|-..|.|.+--.-|=|..-|
T Consensus       748 ----------ky--QLGFSVQNG----iICSLlRG-GIAERGGVRVGHRIIEINgQSVVA~pHekIV~l  799 (829)
T KOG3605|consen  748 ----------RY--QLGFSVQNG----IICSLLRG-GIAERGGVRVGHRIIEINGQSVVATPHEKIVQL  799 (829)
T ss_pred             ----------hh--hccceeeCc----Eeehhhcc-cchhccCceeeeeEEEECCceEEeccHHHHHHH
Confidence                      11  224443321    01112221 346689999999999999987766666565443


No 81 
>PF10459 Peptidase_S46:  Peptidase S46;  InterPro: IPR019500 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This entry represents S46 peptidases, where dipeptidyl-peptidase 7 (DPP-7) is the best-characterised member of this family. It is a serine peptidase that is located on the cell surface and is predicted to have two N-terminal transmembrane domains. 
Probab=84.06  E-value=0.53  Score=53.60  Aligned_cols=22  Identities=32%  Similarity=0.253  Sum_probs=19.6

Q ss_pred             eEEEEEEEe-CCEEEecccccCC
Q 009784          149 SSSSGFAIG-GRRVLTNAHSVEH  170 (526)
Q Consensus       149 ~~GSGfvI~-~g~ILT~aHvV~~  170 (526)
                      +.|||-+|+ +|+||||.||+.+
T Consensus        47 gGCSgsfVS~~GLvlTNHHC~~~   69 (698)
T PF10459_consen   47 GGCSGSFVSPDGLVLTNHHCGYG   69 (698)
T ss_pred             CceeEEEEcCCceEEecchhhhh
Confidence            459999999 8999999999864


No 82 
>PF02395 Peptidase_S6:  Immunoglobulin A1 protease Serine protease Prosite pattern;  InterPro: IPR000710 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to the MEROPS peptidase family S6 (clan PA(S)). The type sample being the IgA1-specific serine endopeptidase from Neisseria gonorrhoeae []. These cleave prolyl bonds in the hinge regions of immunoglobulin A heavy chains. Similar specificity is shown by the unrelated family of M26 metalloendopeptidases.; GO: 0004252 serine-type endopeptidase activity, 0006508 proteolysis; PDB: 3SZE_A 3H09_B 3SYJ_A 1WXR_A 3AK5_B.
Probab=81.36  E-value=7.2  Score=45.11  Aligned_cols=161  Identities=23%  Similarity=0.249  Sum_probs=74.5

Q ss_pred             EEEEEEeCCEEEecccccCCCCeEEEEEcCCCcEEEEEEEEecc--CCCeEEEEecccccccCceeeecCCCC----cC-
Q 009784          151 SSGFAIGGRRVLTNAHSVEHYTQVKLKKRGSDTKYLATVLAIGT--ECDIAMLTVEDDEFWEGVLPVEFGELP----AL-  223 (526)
Q Consensus       151 GSGfvI~~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~vv~~d~--~~DlAlLkv~~~~~~~~~~pl~l~~~~----~~-  223 (526)
                      |...+|++.||+|.+|+..+...+..--. +...|  +++..+.  ..|+.+-|++.--  ..+.|++.....    .. 
T Consensus        67 G~aTLigpqYiVSV~HN~~gy~~v~FG~~-g~~~Y--~iV~RNn~~~~Df~~pRLnK~V--TEvaP~~~t~~~~~~~~y~  141 (769)
T PF02395_consen   67 GVATLIGPQYIVSVKHNGKGYNSVSFGNE-GQNTY--KIVDRNNYPSGDFHMPRLNKFV--TEVAPAEMTTAGSDSNTYN  141 (769)
T ss_dssp             SS-EEEETTEEEBETTG-TSCCEECESCS-STCEE--EEEEEEBETTSTEBEEEESS-----SS----BBSSTTSTTGGG
T ss_pred             ceEEEecCCeEEEEEccCCCcCceeeccc-CCceE--EEEEccCCCCcccceeecCceE--EEEeccccccccccccccc
Confidence            67889999999999999855544433221 22333  4444433  3699999998632  234555443321    00 


Q ss_pred             ---C-CcEEEEe-------eCCCCCc-------eeEEEEEEeceeeeeccCCceeee-----EEEEc----ccCCCCCCC
Q 009784          224 ---Q-DAVTVVG-------YPIGGDT-------ISVTSGVVSRIEILSYVHGSTELL-----GLQID----AAINSGNSG  276 (526)
Q Consensus       224 ---g-~~V~~iG-------~p~~~~~-------~sv~~GiVs~~~~~~~~~~~~~~~-----~i~~d----a~i~~G~SG  276 (526)
                         . ...+=+|       +..+...       ...+.|.+.....  +..+.....     ....+    ....+|+||
T Consensus       142 d~~rY~~f~R~GsG~Q~i~~~~g~~~~~~~~ay~yltgGt~~~~~~--~~n~~~~~~~~~~~~~~~~~pL~n~~~~GDSG  219 (769)
T PF02395_consen  142 DKERYPAFVRVGSGTQYIKDRNGNGTTILGGAYNYLTGGTVYNLPG--YGNGSMILSGDLKKFNSYNGPLPNYGSPGDSG  219 (769)
T ss_dssp             HTTTC-EEEEEESSSEEEEECCEEEEEEEEETTSCEEEEEESSEEE--EECTCEEEEESTTTCCCCCSSSBEB--TT-TT
T ss_pred             cchhchheeecCCceEEEEcCCCCeeEEEEeccceecCCccccccc--cccceEEEecccccccccCCccccccccCcCC
Confidence               0 1111122       2221100       0123344333110  001100000     01111    234689999


Q ss_pred             Ceee--cC---CCeEEEEEeeccccCccccccccccHHHHHHHHHHH
Q 009784          277 GPAF--ND---KGKCVGIAFQSLKHEDVENIGYVIPTPVIMHFIQDY  318 (526)
Q Consensus       277 GPlv--n~---~G~VVGI~~~~~~~~~~~~~~~aIP~~~i~~~l~~l  318 (526)
                      +|||  |.   +.-++|+.+....-.+..+....+|.+.+.++.++.
T Consensus       220 SPlF~YD~~~kKWvl~Gv~~~~~~~~g~~~~~~~~~~~f~~~~~~~d  266 (769)
T PF02395_consen  220 SPLFAYDKEKKKWVLVGVLSGGNGYNGKGNWWNVIPPDFINQIKQND  266 (769)
T ss_dssp             -EEEEEETTTTEEEEEEEEEEECCCCHSEEEEEEECHHHHHHHHHHC
T ss_pred             CceEEEEccCCeEEEEEEEccccccCCccceeEEecHHHHHHHHhhh
Confidence            9987  43   345999987765332333445568888887777664


No 83 
>KOG3606 consensus Cell polarity protein PAR6 [Signal transduction mechanisms]
Probab=81.29  E-value=1.8  Score=43.02  Aligned_cols=64  Identities=22%  Similarity=0.314  Sum_probs=44.0

Q ss_pred             HHHcCceeeccccCceeeeccChhHHhhhcCCCCccceEEEEeCCCCcccC-C-CCCCcEEEEECCEEecC
Q 009784          318 YEKNGAYTGFPLLGVEWQKMENPDLRVAMSMKADQKGVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIAN  386 (526)
Q Consensus       318 l~~~g~~~~~~~LGi~~~~~~~~~~~~~lgl~~~~~Gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~I~~  386 (526)
                      |.++|.-.   -||+....-.+.. ..-.|+. +..|+.|.+..|++-|+. | |-..|-+++|||.+|..
T Consensus       164 L~khG~ek---PLGFYIRDG~SVR-Vtp~Gle-kvpGIFISRlVpGGLAeSTGLLaVnDEVlEVNGIEVaG  229 (358)
T KOG3606|consen  164 LHKHGSEK---PLGFYIRDGTSVR-VTPHGLE-KVPGIFISRLVPGGLAESTGLLAVNDEVLEVNGIEVAG  229 (358)
T ss_pred             hhhcCCCC---CceEEEecCceEE-ecccccc-ccCceEEEeecCCccccccceeeecceeEEEcCEEecc
Confidence            44555432   3666665431111 1123554 567999999999999999 6 56799999999999964


No 84 
>PF12812 PDZ_1:  PDZ-like domain
Probab=80.61  E-value=2.8  Score=34.09  Aligned_cols=56  Identities=13%  Similarity=-0.023  Sum_probs=41.6

Q ss_pred             CceEEEeeEEEEeccccc--------eeeeeeeecchhhhcc-ccccceeeeecccchhhhHHHH
Q 009784          446 PSYYIIAGFVFSRCLYLI--------SVLSMERIMNMKLRSS-FWTSSCIQCHNCQMSSLLWCLR  501 (526)
Q Consensus       446 p~~~i~gG~~f~~lt~~~--------~~~~~~~i~~~~~~sg-~~~~~~~~~~~~~~~~~~~~~~  501 (526)
                      -+++.++|++|++|+|+.        +++.+.+-..+...+| +..+-.|...|.+...+++.+-
T Consensus         5 ~r~v~~~Ga~f~~Ls~q~aR~~~~~~~gv~v~~~~g~~~~~~~i~~g~iI~~Vn~kpt~~Ld~f~   69 (78)
T PF12812_consen    5 SRFVEVCGAVFHDLSYQQARQYGIPVGGVYVAVSGGSLAFAGGISKGFIITSVNGKPTPDLDDFI   69 (78)
T ss_pred             CEEEEEcCeecccCCHHHHHHhCCCCCEEEEEecCCChhhhCCCCCCeEEEeECCcCCcCHHHHH
Confidence            389999999999999653        2333433333333455 9999999999999998888764


No 85 
>KOG3549 consensus Syntrophins (type gamma) [Extracellular structures]
Probab=80.31  E-value=1.8  Score=44.43  Aligned_cols=55  Identities=22%  Similarity=0.266  Sum_probs=41.0

Q ss_pred             ceEEEEeCCCCcccC-C-CCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEE
Q 009784          354 GVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVL  418 (526)
Q Consensus       354 Gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~  418 (526)
                      -++|.++..+-.|+. | |-.||-|++|||..|..-..=     +-+   .++  .+.||.|+|+|.
T Consensus        81 PvviSkI~kdQaAd~tG~LFvGDAilqvNGi~v~~c~He-----evV---~iL--RNAGdeVtlTV~  137 (505)
T KOG3549|consen   81 PVVISKIYKDQAADITGQLFVGDAILQVNGIYVTACPHE-----EVV---NIL--RNAGDEVTLTVK  137 (505)
T ss_pred             cEEeehhhhhhhhhhcCceEeeeeeEEeccEEeecCChH-----HHH---HHH--HhcCCEEEEEeH
Confidence            388999998888888 5 789999999999999875321     111   222  347999998874


No 86 
>KOG3651 consensus Protein kinase C, alpha binding protein [Signal transduction mechanisms]
Probab=79.41  E-value=2.7  Score=42.41  Aligned_cols=37  Identities=22%  Similarity=0.356  Sum_probs=32.4

Q ss_pred             ceEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCc
Q 009784          354 GVRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTV  390 (526)
Q Consensus       354 Gv~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l  390 (526)
                      =++|.+|-.++||++  -++.||-|++|||..|.....+
T Consensus        31 ClYiVQvFD~tPAa~dG~i~~GDEi~avNg~svKGktKv   69 (429)
T KOG3651|consen   31 CLYIVQVFDKTPAAKDGRIRCGDEIVAVNGISVKGKTKV   69 (429)
T ss_pred             eEEEEEeccCCchhccCccccCCeeEEecceeecCccHH
Confidence            378999999999998  5999999999999999876554


No 87 
>PF03510 Peptidase_C24:  2C endopeptidase (C24) cysteine protease family;  InterPro: IPR000317 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad [].  The two signatures that defines this group of calivirus polyproteins identify a cysteine peptidase signature that belongs to MEROPS peptidase family C24 (clan PA(C)). Caliciviruses are positive-stranded ssRNA viruses that cause gastroenteritis. The calicivirus genome contains two open reading frames, ORF1 and ORF2. ORF2 encodes a structural protein []; while ORF1 encodes a non-structural polypeptide, which has RNA helicase, cysteine protease and RNA polymerase activity. The regions of the polyprotein in which these activities lie are similar to proteins produced by the picornaviruses. Two different families of caliciviruses can be distinguished on the basis of sequence similarity, namely those classified as small round structured viruses (SRSVs) and those classed as non-SRSVs. Calicivirus proteases from the non-SRSV group, which are members of the PA protease clan, constitute family C24 of the cysteine proteases (proteases from SRSVs belong to the C37 family). As mentioned above, the protease activity resides within a polyprotein. The enzyme cleaves the polyprotein at sites N-terminal to itself, liberating the polyprotein helicase.; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis
Probab=75.75  E-value=9.6  Score=32.80  Aligned_cols=54  Identities=15%  Similarity=0.250  Sum_probs=33.8

Q ss_pred             EEEEeCCEEEecccccCCCCeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecccccccCceeeecCCC
Q 009784          153 GFAIGGRRVLTNAHSVEHYTQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGEL  220 (526)
Q Consensus       153 GfvI~~g~ILT~aHvV~~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~  220 (526)
                      ++-|.+|.++|+.||.+.++.+.      |..+  +++.  ..-|+++++.+...    ++..++++.
T Consensus         3 avHIGnG~~vt~tHva~~~~~v~------g~~f--~~~~--~~ge~~~v~~~~~~----~p~~~ig~g   56 (105)
T PF03510_consen    3 AVHIGNGRYVTVTHVAKSSDSVD------GQPF--KIVK--TDGELCWVQSPLVH----LPAAQIGTG   56 (105)
T ss_pred             eEEeCCCEEEEEEEEeccCceEc------CcCc--EEEE--eccCEEEEECCCCC----CCeeEeccC
Confidence            45566899999999998765431      2221  1222  34599999988753    455566543


No 88 
>PF01732 DUF31:  Putative peptidase (DUF31);  InterPro: IPR022382  This domain has no known function. It is found in various hypothetical proteins and putative lipoproteins from mycoplasmas. 
Probab=71.03  E-value=3.2  Score=43.89  Aligned_cols=24  Identities=33%  Similarity=0.588  Sum_probs=21.2

Q ss_pred             cCCCCCCCCeeecCCCeEEEEEee
Q 009784          269 AINSGNSGGPAFNDKGKCVGIAFQ  292 (526)
Q Consensus       269 ~i~~G~SGGPlvn~~G~VVGI~~~  292 (526)
                      .+..|.||+.|+|.+|++|||.++
T Consensus       351 ~l~gGaSGS~V~n~~~~lvGIy~g  374 (374)
T PF01732_consen  351 SLGGGASGSMVINQNNELVGIYFG  374 (374)
T ss_pred             CCCCCCCcCeEECCCCCEEEEeCC
Confidence            556899999999999999999753


No 89 
>KOG0609 consensus Calcium/calmodulin-dependent serine protein kinase/membrane-associated guanylate kinase [Signal transduction mechanisms]
Probab=70.18  E-value=6.7  Score=42.81  Aligned_cols=57  Identities=25%  Similarity=0.336  Sum_probs=42.2

Q ss_pred             ceEEEEeCCCCcccC-C-CCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEEC
Q 009784          354 GVRIRRVDPTAPESE-V-LKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD  420 (526)
Q Consensus       354 Gv~V~~V~~~spA~~-G-L~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~  420 (526)
                      -++|..+..|+-+++ | |+.||.|++|||..+.+..--+        +..++.... | .|+++|.=.
T Consensus       147 ~~~vARI~~GG~~~r~glL~~GD~i~EvNGi~v~~~~~~e--------~q~~l~~~~-G-~itfkiiP~  205 (542)
T KOG0609|consen  147 KVVVARIMHGGMADRQGLLHVGDEILEVNGISVANKSPEE--------LQELLRNSR-G-SITFKIIPS  205 (542)
T ss_pred             ccEEeeeccCCcchhccceeeccchheecCeecccCCHHH--------HHHHHHhCC-C-cEEEEEccc
Confidence            588999999999988 4 8999999999999998753211        224555543 4 688887543


No 90 
>KOG0606 consensus Microtubule-associated serine/threonine kinase and related proteins [Signal transduction mechanisms; General function prediction only]
Probab=69.61  E-value=4.9  Score=47.40  Aligned_cols=34  Identities=21%  Similarity=0.258  Sum_probs=30.0

Q ss_pred             eEEEEeCCCCcccC-CCCCCcEEEEECCEEecCCC
Q 009784          355 VRIRRVDPTAPESE-VLKPSDIILSFDGIDIANDG  388 (526)
Q Consensus       355 v~V~~V~~~spA~~-GL~~GDvIl~vnG~~I~~~~  388 (526)
                      -.|..|.+++||.. ||++||.|+.+||+++....
T Consensus       660 h~v~sv~egsPA~~agls~~DlIthvnge~v~gl~  694 (1205)
T KOG0606|consen  660 HSVGSVEEGSPAFEAGLSAGDLITHVNGEPVHGLV  694 (1205)
T ss_pred             eeeeeecCCCCccccCCCccceeEeccCcccchhh
Confidence            44788999999988 99999999999999997643


No 91 
>TIGR02860 spore_IV_B stage IV sporulation protein B. SpoIVB, the stage IV sporulation protein B of endospore-forming bacteria such as Bacillus subtilis, is a serine proteinase, expressed in the spore (rather than mother cell) compartment, that participates in a proteolytic activation cascade for Sigma-K. It appears to be universal among endospore-forming bacteria and occurs nowhere else.
Probab=68.01  E-value=3.6  Score=43.79  Aligned_cols=42  Identities=24%  Similarity=0.423  Sum_probs=32.2

Q ss_pred             ccCCCCCCCCeeecCCCeEEEEEeeccccCccccccccccHHHHH
Q 009784          268 AAINSGNSGGPAFNDKGKCVGIAFQSLKHEDVENIGYVIPTPVIM  312 (526)
Q Consensus       268 a~i~~G~SGGPlvn~~G~VVGI~~~~~~~~~~~~~~~aIP~~~i~  312 (526)
                      ..+-.|+||+|++- +|++||=++--+.  +....||.|-++...
T Consensus       355 gGivqGMSGSPi~q-~gkliGAvtHVfv--ndpt~GYGi~ie~Ml  396 (402)
T TIGR02860       355 GGIVQGMSGSPIIQ-NGKVIGAVTHVFV--NDPTSGYGVYIEWML  396 (402)
T ss_pred             CCEEecccCCCEEE-CCEEEEEEEEEEe--cCCCcceeehHHHHH
Confidence            35678999999995 8999998877663  456778888776643


No 92 
>KOG1924 consensus RhoA GTPase effector DIA/Diaphanous [Signal transduction mechanisms; Cytoskeleton]
Probab=54.44  E-value=42  Score=38.52  Aligned_cols=10  Identities=30%  Similarity=0.531  Sum_probs=5.6

Q ss_pred             CeEEEEeccc
Q 009784          197 DIAMLTVEDD  206 (526)
Q Consensus       197 DlAlLkv~~~  206 (526)
                      -++||+++.+
T Consensus       719 k~~ILevne~  728 (1102)
T KOG1924|consen  719 KNVILEVNED  728 (1102)
T ss_pred             HHHHhhccHH
Confidence            4566666544


No 93 
>PF05416 Peptidase_C37:  Southampton virus-type processing peptidase;  InterPro: IPR001665 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad [].  This group of cysteine peptidases belong to the MEROPS peptidase family C37, (clan PA(C)). The type example is calicivirin from Southampton virus, an endopeptidase that cleaves the polyprotein at sites N-terminal to itself, liberating the polyprotein helicase. Southampton virus is a positive-stranded ssRNA virus belonging to the Caliciviruses, which are viruses that cause gastroenteritis. The calicivirus genome contains two open reading frames, ORF1 and ORF2. ORF1 encodes a non-structural polypeptide, which has RNA helicase, cysteine protease and RNA polymerase activity []. The regions of the polyprotein in which these activities lie are similar to proteins produced by the picornaviruses []. ORF2 encodes a structural, capsid protein. Two different families of caliciviruses can be distinguished on the basis of sequence similarity, namely the Norwalk-like viruses or small round structured viruses (SRSVs), and those classed as non-SRSVs.; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis; PDB: 2FYQ_A 2FYR_A 1WQS_D 4ASH_A 2IPH_B.
Probab=54.29  E-value=95  Score=33.29  Aligned_cols=135  Identities=16%  Similarity=0.215  Sum_probs=65.8

Q ss_pred             eEEEEEEEeCCEEEecccccCCC-CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecccccccCceeeecCCCCcCCCcE
Q 009784          149 SSSSGFAIGGRRVLTNAHSVEHY-TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDDEFWEGVLPVEFGELPALQDAV  227 (526)
Q Consensus       149 ~~GSGfvI~~g~ILT~aHvV~~~-~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l~~~~~~g~~V  227 (526)
                      +.|=||-++..+++|+-||+... .++.      |  .+-.-+.++..-+++-+++..+- -.++.-+-|.+...-|.-+
T Consensus       379 GsGWGfWVS~~lfITttHViP~g~~E~F------G--v~i~~i~vh~sGeF~~~rFpk~i-RPDvtgmiLEeGapEGtV~  449 (535)
T PF05416_consen  379 GSGWGFWVSPTLFITTTHVIPPGAKEAF------G--VPISQIQVHKSGEFCRFRFPKPI-RPDVTGMILEEGAPEGTVC  449 (535)
T ss_dssp             TTEEEEESSSSEEEEEGGGS-STTSEET------T--EECGGEEEEEETTEEEEEESS-S-STTS---EE-SS--TT-EE
T ss_pred             CCceeeeecceEEEEeeeecCCcchhhh------C--CChhHeEEeeccceEEEecCCCC-CCCccceeeccCCCCceEE
Confidence            46778999999999999999743 2210      0  01111233344566666666542 1234444443333334222


Q ss_pred             -EEEeeCCCCC-ceeEEEEEEeceeeee-ccCCceeeeEEEE-------cccCCCCCCCCeeecCCC---eEEEEEeecc
Q 009784          228 -TVVGYPIGGD-TISVTSGVVSRIEILS-YVHGSTELLGLQI-------DAAINSGNSGGPAFNDKG---KCVGIAFQSL  294 (526)
Q Consensus       228 -~~iG~p~~~~-~~sv~~GiVs~~~~~~-~~~~~~~~~~i~~-------da~i~~G~SGGPlvn~~G---~VVGI~~~~~  294 (526)
                       +.|-.+.|.. .+.+..|......... ...|  ...++-+       |-.+.||+-|.|-+-..|   -|+|++++..
T Consensus       450 siLiKR~sGEllpLAvRMgt~AsmkIqgr~v~G--Q~GMLLTGaNAK~mDLGT~PGDCGcPYvyKrgNd~VV~GVH~AAt  527 (535)
T PF05416_consen  450 SILIKRPSGELLPLAVRMGTHASMKIQGRTVHG--QMGMLLTGANAKGMDLGTIPGDCGCPYVYKRGNDWVVIGVHAAAT  527 (535)
T ss_dssp             EEEEE-TTSBEEEEEEEEEEEEEEEETTEEEEE--EEEEETTSTT-SSTTTS--TTGTT-EEEEEETTEEEEEEEEEEE-
T ss_pred             EEEEEcCCccchhhhhhhccceeEEEcceeecc--eeeeeeecCCccccccCCCCCCCCCceeeecCCcEEEEEEEehhc
Confidence             3355555532 2456677666543210 0111  1123323       335679999999886655   4999998865


No 94 
>smart00384 AT_hook DNA binding domain with preference for A/T rich regions. Small DNA-binding motif first described in the high mobility group non-histone chromosomal protein HMG-I(Y).
Probab=45.96  E-value=13  Score=23.58  Aligned_cols=16  Identities=50%  Similarity=0.667  Sum_probs=12.6

Q ss_pred             hhccCCCCCCCCcccc
Q 009784            5 KRKRGRKPKIPDAEKT   20 (526)
Q Consensus         5 ~~~~~~~~~~~~~~~~   20 (526)
                      +|||||-+|.+.....
T Consensus         1 kRkRGRPrK~~~~~~~   16 (26)
T smart00384        1 KRKRGRPRKAPKDXXX   16 (26)
T ss_pred             CCCCCCCCCCCCcccc
Confidence            6899999998876543


No 95 
>PF12381 Peptidase_C3G:  Tungro spherical virus-type peptidase;  InterPro: IPR024387 This entry represents a rice tungro spherical waikavirus-type peptidase that belongs to MEROPS peptidase family C3G. It is a picornain 3C-type protease, and is responsible for the self-cleavage of the positive single-stranded polyproteins of a number of plant viral genomes. The location of the protease activity of the polyprotein is at the C-terminal end, adjacent and N-terminal to the putative RNA polymerase [, ].
Probab=45.51  E-value=30  Score=33.63  Aligned_cols=54  Identities=26%  Similarity=0.426  Sum_probs=39.6

Q ss_pred             EEEEcccCCCCCCCCeeecC----CCeEEEEEeeccccCccccccccccH--HHHHHHHHHHH
Q 009784          263 GLQIDAAINSGNSGGPAFND----KGKCVGIAFQSLKHEDVENIGYVIPT--PVIMHFIQDYE  319 (526)
Q Consensus       263 ~i~~da~i~~G~SGGPlvn~----~G~VVGI~~~~~~~~~~~~~~~aIP~--~~i~~~l~~l~  319 (526)
                      .++..+....|+=|||++-.    .-+++||+.++.   .+...+||-++  +.+++.+..|.
T Consensus       170 gleY~~~t~~GdCGs~i~~~~t~~~RKIvGiHVAG~---~~~~~gYAe~itQEDL~~A~~~l~  229 (231)
T PF12381_consen  170 GLEYQMPTMNGDCGSPIVRNNTQMVRKIVGIHVAGS---ANHAMGYAESITQEDLMRAINKLE  229 (231)
T ss_pred             eeeEECCCcCCCccceeeEcchhhhhhhheeeeccc---ccccceehhhhhHHHHHHHHHhhc
Confidence            46677788899999997632    368999999876   34567888554  55777766664


No 96 
>PF13180 PDZ_2:  PDZ domain; PDB: 2L97_A 1Y8T_A 2Z9I_A 1LCY_A 2PZD_B 2P3W_A 1VCW_C 1TE0_B 1SOZ_C 1SOT_C ....
Probab=42.76  E-value=63  Score=25.73  Aligned_cols=50  Identities=8%  Similarity=-0.073  Sum_probs=34.1

Q ss_pred             eeEEEEeccccceeeeeeeecc--hhhhccccccceeeeecccchhhhHHHHH
Q 009784          452 AGFVFSRCLYLISVLSMERIMN--MKLRSSFWTSSCIQCHNCQMSSLLWCLRC  502 (526)
Q Consensus       452 gG~~f~~lt~~~~~~~~~~i~~--~~~~sg~~~~~~~~~~~~~~~~~~~~~~~  502 (526)
                      -|+.|...+...+ +.+..+.+  .+.++|+..||+|...+..-..+...|.-
T Consensus         3 lGv~~~~~~~~~g-~~V~~V~~~spA~~aGl~~GD~I~~ing~~v~~~~~~~~   54 (82)
T PF13180_consen    3 LGVTVQNLSDTGG-VVVVSVIPGSPAAKAGLQPGDIILAINGKPVNSSEDLVN   54 (82)
T ss_dssp             -SEEEEECSCSSS-EEEEEESTTSHHHHTTS-TTEEEEEETTEESSSHHHHHH
T ss_pred             ECeEEEEccCCCe-EEEEEeCCCCcHHHCCCCCCcEEEEECCEEcCCHHHHHH
Confidence            4667776665323 33334444  55689999999999999988888777763


No 97 
>KOG1924 consensus RhoA GTPase effector DIA/Diaphanous [Signal transduction mechanisms; Cytoskeleton]
Probab=40.69  E-value=80  Score=36.37  Aligned_cols=9  Identities=33%  Similarity=0.150  Sum_probs=4.5

Q ss_pred             CCCCCcccc
Q 009784           12 PKIPDAEKT   20 (526)
Q Consensus        12 ~~~~~~~~~   20 (526)
                      +||++-|-+
T Consensus       502 ~Ki~~l~ae  510 (1102)
T KOG1924|consen  502 EKIKLLEAE  510 (1102)
T ss_pred             hhcccCchh
Confidence            555554443


No 98 
>KOG3938 consensus RGS-GAIP interacting protein GIPC, contains PDZ domain [Signal transduction mechanisms; Intracellular trafficking, secretion, and vesicular transport]
Probab=39.74  E-value=15  Score=36.77  Aligned_cols=58  Identities=12%  Similarity=0.228  Sum_probs=43.9

Q ss_pred             eEEEEeCCCCcccC--CCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEEC
Q 009784          355 VRIRRVDPTAPESE--VLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRD  420 (526)
Q Consensus       355 v~V~~V~~~spA~~--GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~  420 (526)
                      ..|..+.++|--..  -++.||.|-+|||+.|-....++        ..+.+.....|++.+|++.--
T Consensus       151 AFIKrIkegsvidri~~i~VGd~IEaiNge~ivG~RHYe--------VArmLKel~rge~ftlrLieP  210 (334)
T KOG3938|consen  151 AFIKRIKEGSVIDRIEAICVGDHIEAINGESIVGKRHYE--------VARMLKELPRGETFTLRLIEP  210 (334)
T ss_pred             eeeEeecCCchhhhhhheeHHhHHHhhcCccccchhHHH--------HHHHHHhcccCCeeEEEeecc
Confidence            56778888887776  79999999999999998766543        225566666788887776543


No 99 
>cd01720 Sm_D2 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit D2 heterodimerizes with subunit D1 and three such heterodimers form a hexameric ring structure with alternating D1 and D2 subunits. The D1 - D2 heterodimer also assembles into a heptameric ring containing D2, D3, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=37.07  E-value=63  Score=26.82  Aligned_cols=37  Identities=30%  Similarity=0.493  Sum_probs=30.5

Q ss_pred             ccCCCCeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784          167 SVEHYTQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE  204 (526)
Q Consensus       167 vV~~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~  204 (526)
                      ++.....+.|.+. +++.+.+++.++|.+.++.|=...
T Consensus        10 ~~~~~~~V~V~lr-~~r~~~G~L~~fD~hmNlvL~d~~   46 (87)
T cd01720          10 AVKNNTQVLINCR-NNKKLLGRVKAFDRHCNMVLENVK   46 (87)
T ss_pred             HHcCCCEEEEEEc-CCCEEEEEEEEecCccEEEEcceE
Confidence            3444578999997 899999999999999999876654


No 100
>cd00600 Sm_like The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=34.50  E-value=1.1e+02  Score=23.04  Aligned_cols=33  Identities=12%  Similarity=0.238  Sum_probs=27.8

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecc
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVED  205 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~  205 (526)
                      ..+.|.+. +|+.+.+.+..+|...++.|-....
T Consensus         7 ~~V~V~l~-~g~~~~G~L~~~D~~~Ni~L~~~~~   39 (63)
T cd00600           7 KTVRVELK-DGRVLEGVLVAFDKYMNLVLDDVEE   39 (63)
T ss_pred             CEEEEEEC-CCcEEEEEEEEECCCCCEEECCEEE
Confidence            46888887 9999999999999998888766543


No 101
>KOG3834 consensus Golgi reassembly stacking protein GRASP65, contains PDZ domain [Intracellular trafficking, secretion, and vesicular transport]
Probab=31.54  E-value=37  Score=36.24  Aligned_cols=65  Identities=17%  Similarity=0.232  Sum_probs=45.3

Q ss_pred             EEEEeCCCCcccC-CCC-CCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCCEEEEEEEECCE--EEEEEEEec
Q 009784          356 RIRRVDPTAPESE-VLK-PSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGDSAAVKVLRDSK--ILNFNITLA  431 (526)
Q Consensus       356 ~V~~V~~~spA~~-GL~-~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~~v~l~v~R~G~--~~~~~v~l~  431 (526)
                      -|-+|.++|||+. ||. -+|-|+-+-.......+||.          .++. .+.++.+++-|+.-..  .++++++..
T Consensus       112 Hvl~V~p~SPaalAgl~~~~DYivG~~~~~~~~~eDl~----------~lIe-she~kpLklyVYN~D~d~~ReVti~pn  180 (462)
T KOG3834|consen  112 HVLSVEPNSPAALAGLRPYTDYIVGIWDAVMHEEEDLF----------TLIE-SHEGKPLKLYVYNHDTDSCREVTITPN  180 (462)
T ss_pred             eeeecCCCCHHHhcccccccceEecchhhhccchHHHH----------HHHH-hccCCCcceeEeecCCCccceEEeecc
Confidence            3668999999999 999 68999998444455566652          4444 4568999999987443  344555543


No 102
>PF02178 AT_hook:  AT hook motif;  InterPro: IPR017956 AT hooks are DNA-binding motifs with a preference for A/T rich regions. These motifs are found in a variety of proteins, including the high mobility group (HMG) proteins [], in DNA-binding proteins from plants [] and in hBRG1 protein, a central ATPase of the human switching/sucrose non-fermenting (SWI/SNF) remodeling complex [].  High mobility group (HMG) proteins are a family of relatively low molecular weight non-histone components in chromatin []. HMG-I and HMG-Y (HMGA) are proteins of about 100 amino acid residues which are produced by the alternative splicing of a single gene. HMG-I/Y proteins bind preferentially to the minor groove of AT-rich regions in double-stranded DNA in a non-sequence specific manner [, ]. It is suggested that these proteins could function in nucleosome phasing and in the 3' end processing of mRNA transcripts. They are also involved in the transcription regulation of genes containing, or in close proximity to, AT-rich regions. ; GO: 0003677 DNA binding; PDB: 2EZE_A 2EZD_A 2EZF_A 2EZG_A.
Probab=31.17  E-value=21  Score=18.98  Aligned_cols=11  Identities=64%  Similarity=0.797  Sum_probs=3.8

Q ss_pred             hhccCCCCCCC
Q 009784            5 KRKRGRKPKIP   15 (526)
Q Consensus         5 ~~~~~~~~~~~   15 (526)
                      +|+|||-+|-.
T Consensus         1 ~r~RGRP~k~~   11 (13)
T PF02178_consen    1 KRKRGRPRKNA   11 (13)
T ss_dssp             S--SS--TT--
T ss_pred             CCcCCCCcccc
Confidence            57899887753


No 103
>PF09465 LBR_tudor:  Lamin-B receptor of TUDOR domain;  InterPro: IPR019023  The Lamin-B receptor is a chromatin and lamin binding protein in the inner nuclear membrane. It is one of the integral inner nuclear envelope membrane proteins responsible for targeting nuclear membranes to chromatin, being a downstream effector of Ran, a small Ras-like nuclear GTPase which regulates NE assembly. Lamin-B receptor interacts with importin beta, a Ran-binding protein, thereby directly contributing to the fusion of membrane vesicles and the formation of the nuclear envelope []. ; PDB: 2L8D_A 2DIG_A.
Probab=30.53  E-value=2.1e+02  Score=21.69  Aligned_cols=38  Identities=24%  Similarity=0.211  Sum_probs=29.5

Q ss_pred             CCCCeEEEEEcCCCcEEEEEEEEeccCCCeEEEEeccc
Q 009784          169 EHYTQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDD  206 (526)
Q Consensus       169 ~~~~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~  206 (526)
                      .....+.++.+++..-|++++..+|...++.-++++..
T Consensus         7 ~~Ge~V~~rWP~s~lYYe~kV~~~d~~~~~y~V~Y~DG   44 (55)
T PF09465_consen    7 AIGEVVMVRWPGSSLYYEGKVLSYDSKSDRYTVLYEDG   44 (55)
T ss_dssp             -SS-EEEEE-TTTS-EEEEEEEEEETTTTEEEEEETTS
T ss_pred             cCCCEEEEECCCCCcEEEEEEEEecccCceEEEEEcCC
Confidence            34567899999777788999999999999999999764


No 104
>TIGR03000 plancto_dom_1 Planctomycetes uncharacterized domain TIGR03000. Domains described by this model are found, so far, only in the Planctomycetes (Pirellula sp. strain 1 and Gemmata obscuriglobus), in up to six proteins per genome, and may be duplicated within a protein. The function is unknown.
Probab=30.44  E-value=1.5e+02  Score=24.04  Aligned_cols=49  Identities=29%  Similarity=0.415  Sum_probs=32.0

Q ss_pred             CCcEEEEECCEEecCCCCcccccccchhhhhhhhhcCCCC----EEEEEEEECCEEEEEEEE
Q 009784          372 PSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQKYTGD----SAAVKVLRDSKILNFNIT  429 (526)
Q Consensus       372 ~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~~~~G~----~v~l~v~R~G~~~~~~v~  429 (526)
                      |-|-.+.+||++..+.+...      .   ..-.....|.    ++..++.|||+..+.+-+
T Consensus        10 PadAkl~v~G~~t~~~G~~R------~---F~T~~L~~G~~y~Y~v~a~~~~dG~~~t~~~~   62 (75)
T TIGR03000        10 PADAKLKVDGKETNGTGTVR------T---FTTPPLEAGKEYEYTVTAEYDRDGRILTRTRT   62 (75)
T ss_pred             CCCCEEEECCeEcccCccEE------E---EECCCCCCCCEEEEEEEEEEecCCcEEEEEEE
Confidence            46788999999999888752      0   1111233455    467778899987665433


No 105
>PF00571 CBS:  CBS domain CBS domain web page. Mutations in the CBS domain of Swiss:P35520 lead to homocystinuria.;  InterPro: IPR000644 CBS (cystathionine-beta-synthase) domains are small intracellular modules, mostly found in two or four copies within a protein, that occur in a variety of proteins in bacteria, archaea, and eukaryotes [, ]. Tandem pairs of CBS domains can act as binding domains for adenosine derivatives and may regulate the activity of attached enzymatic or other domains []. In some cases, CBS domains may act as sensors of cellular energy status by being activated by AMP and inhibited by ATP []. In chloride ion channels, the CBS domains have been implicated in intracellular targeting and trafficking, as well as in protein-protein interactions, but results vary with different channels: in the CLC-5 channel, the CBS domain was shown to be required for trafficking [], while in the CLC-1 channel, the CBS domain was shown to be critical for channel function, but not necessary for trafficking []. Recent experiments revealing that CBS domains can bind adenosine-containing ligands such ATP, AMP, or S-adenosylmethionine have led to the hypothesis that CBS domains function as sensors of intracellular metabolites [, ]. Crystallographic studies of CBS domains have shown that pairs of CBS sequences form a globular domain where each CBS unit adopts a beta-alpha-beta-beta-alpha pattern []. Crystal structure of the CBS domains of the AMP-activated protein kinase in complexes with AMP and ATP shows that the phosphate groups of AMP/ATP lie in a surface pocket at the interface of two CBS domains, which is lined with basic residues, many of which are associated with disease-causing mutations [].  In humans, mutations in conserved residues within CBS domains cause a variety of human hereditary diseases, including (with the gene mutated in parentheses): homocystinuria (cystathionine beta-synthase); Wolff-Parkinson-White syndrome (gamma 2 subunit of AMP-activated protein kinase); retinitis pigmentosa (IMP dehydrogenase-1); congenital myotonia, idiopathic generalized epilepsy, hypercalciuric nephrolithiasis, and classic Bartter syndrome (CLC chloride channel family members).; GO: 0005515 protein binding; PDB: 3JTF_A 3TE5_C 3TDH_C 3T4N_C 2QLV_C 3OI8_A 3LV9_A 2QH1_B 1PVM_B 3LQN_A ....
Probab=29.79  E-value=45  Score=24.12  Aligned_cols=21  Identities=38%  Similarity=0.555  Sum_probs=17.4

Q ss_pred             CCCCCCeeecCCCeEEEEEee
Q 009784          272 SGNSGGPAFNDKGKCVGIAFQ  292 (526)
Q Consensus       272 ~G~SGGPlvn~~G~VVGI~~~  292 (526)
                      .+.+.-|++|.+|+++|+++.
T Consensus        28 ~~~~~~~V~d~~~~~~G~is~   48 (57)
T PF00571_consen   28 NGISRLPVVDEDGKLVGIISR   48 (57)
T ss_dssp             HTSSEEEEESTTSBEEEEEEH
T ss_pred             cCCcEEEEEecCCEEEEEEEH
Confidence            356678999999999999875


No 106
>cd01726 LSm6 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm6 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=29.28  E-value=1.2e+02  Score=23.52  Aligned_cols=32  Identities=22%  Similarity=0.255  Sum_probs=27.4

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE  204 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~  204 (526)
                      ..+.|.+. +|+.|.+++.++|+..++.|=...
T Consensus        11 ~~V~V~Lk-~g~~~~G~L~~~D~~mNlvL~~~~   42 (67)
T cd01726          11 RPVVVKLN-SGVDYRGILACLDGYMNIALEQTE   42 (67)
T ss_pred             CeEEEEEC-CCCEEEEEEEEEccceeeEEeeEE
Confidence            46889997 899999999999999998886553


No 107
>PRK00737 small nuclear ribonucleoprotein; Provisional
Probab=28.94  E-value=1.2e+02  Score=23.87  Aligned_cols=33  Identities=6%  Similarity=0.227  Sum_probs=28.3

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecc
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVED  205 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~  205 (526)
                      ..+.|.+. +|+.+.+++.++|...++.|=....
T Consensus        15 k~V~V~lk-~g~~~~G~L~~~D~~mNlvL~d~~e   47 (72)
T PRK00737         15 SPVLVRLK-GGREFRGELQGYDIHMNLVLDNAEE   47 (72)
T ss_pred             CEEEEEEC-CCCEEEEEEEEEcccceeEEeeEEE
Confidence            46888887 8999999999999999998877643


No 108
>cd01731 archaeal_Sm1 The archaeal sm1 proteins: The Sm proteins are conserved in all three domains of life and are always associated with U-rich RNA sequences. They function to mediate RNA-RNA interactions and RNA biogenesis.  All Sm proteins contain a common sequence motif in two segments, Sm1 and Sm2, separated by a short variable linker. Eukaryotic Sm proteins form part of specific small nuclear ribonucleoproteins (snRNPs) that are involved in the processing of pre-mRNAs to mature mRNAs, and are a major component of the eukaryotic spliceosome. Most snRNPs consist of seven Sm proteins (B/B', D1, D2, D3, E, F and G) arranged in a ring on a uridine-rich sequence (Sm site), plus a small nuclear RNA (snRNA) (either U1, U2, U5 or U4/6). Since archaebacteria do not have any splicing apparatus, Sm proteins of archaebacteria may play a more general role. Archaeal Lsm proteins are likely to represent the ancestral Sm domain.
Probab=28.75  E-value=1.3e+02  Score=23.39  Aligned_cols=33  Identities=9%  Similarity=0.172  Sum_probs=28.7

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecc
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVED  205 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~  205 (526)
                      ..+.|.+. +|+.+.+++.++|...++.|-....
T Consensus        11 ~~V~V~l~-~g~~~~G~L~~~D~~mNlvL~~~~e   43 (68)
T cd01731          11 KPVLVKLK-GGKEVRGRLKSYDQHMNLVLEDAEE   43 (68)
T ss_pred             CEEEEEEC-CCCEEEEEEEEECCcceEEEeeEEE
Confidence            56888897 8999999999999999998877654


No 109
>cd01722 Sm_F The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit F is capable of forming both homo- and hetero-heptamer ring structures.  To form the hetero-heptamer, Sm subunit F initially binds subunits E and G to form a trimer which then assembles onto snRNA along with the D3/B and D1/D2 heterodimers.
Probab=28.45  E-value=1.2e+02  Score=23.73  Aligned_cols=32  Identities=16%  Similarity=0.306  Sum_probs=27.2

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE  204 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~  204 (526)
                      ..+.|.+. +|+.+.+++.++|...++.+=...
T Consensus        12 ~~V~V~Lk-~g~~~~G~L~~~D~~mNi~L~~~~   43 (68)
T cd01722          12 KPVIVKLK-WGMEYKGTLVSVDSYMNLQLANTE   43 (68)
T ss_pred             CEEEEEEC-CCcEEEEEEEEECCCEEEEEeeEE
Confidence            46889997 999999999999999888875553


No 110
>cd01730 LSm3 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm3 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=27.05  E-value=1.1e+02  Score=24.86  Aligned_cols=31  Identities=19%  Similarity=0.225  Sum_probs=26.5

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEe
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTV  203 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv  203 (526)
                      ..+.|.+. +|+.+.+++.++|...+|.|=..
T Consensus        12 k~V~V~l~-~gr~~~G~L~~fD~~mNlvL~d~   42 (82)
T cd01730          12 ERVYVKLR-GDRELRGRLHAYDQHLNMILGDV   42 (82)
T ss_pred             CEEEEEEC-CCCEEEEEEEEEccceEEeccce
Confidence            56888887 89999999999999998887544


No 111
>cd01717 Sm_B The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit B heterodimerizes with subunit D3 and three such heterodimers form a hexameric ring structure with alternating B and D3 subunits.  The D3 - B heterodimer also assembles into a heptameric ring containing D1, D2, E, F, and G subunits.  Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=26.17  E-value=1.3e+02  Score=24.10  Aligned_cols=32  Identities=9%  Similarity=0.327  Sum_probs=27.3

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE  204 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~  204 (526)
                      ..+.|.+. +|+.+.+.+.++|...+|.|=...
T Consensus        11 ~~V~V~l~-dgR~~~G~L~~~D~~~NlVL~~~~   42 (79)
T cd01717          11 YRLRVTLQ-DGRQFVGQFLAFDKHMNLVLSDCE   42 (79)
T ss_pred             CEEEEEEC-CCcEEEEEEEEEcCccCEEcCCEE
Confidence            46888887 999999999999999998876554


No 112
>COG0298 HypC Hydrogenase maturation factor [Posttranslational modification, protein turnover, chaperones]
Probab=26.04  E-value=1.5e+02  Score=24.25  Aligned_cols=47  Identities=21%  Similarity=0.376  Sum_probs=31.8

Q ss_pred             EEEEEEEeccCCCeEEEEecccccccCceeeec-CCCCcCCCcEEE-EeeCC
Q 009784          185 YLATVLAIGTECDIAMLTVEDDEFWEGVLPVEF-GELPALQDAVTV-VGYPI  234 (526)
Q Consensus       185 ~~a~vv~~d~~~DlAlLkv~~~~~~~~~~pl~l-~~~~~~g~~V~~-iG~p~  234 (526)
                      ++++++..+...++|++.+-.-.   .---+.+ ....++|++|.+ +||..
T Consensus         5 iPgqI~~I~~~~~~A~Vd~gGvk---reV~l~Lv~~~v~~GdyVLVHvGfAi   53 (82)
T COG0298           5 IPGQIVEIDDNNHLAIVDVGGVK---REVNLDLVGEEVKVGDYVLVHVGFAM   53 (82)
T ss_pred             cccEEEEEeCCCceEEEEeccEe---EEEEeeeecCccccCCEEEEEeeEEE
Confidence            57889999988889999987643   1112222 236688998876 67653


No 113
>cd06168 LSm9 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm9 proteins have a single Sm-like domain structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=25.59  E-value=1.6e+02  Score=23.58  Aligned_cols=32  Identities=9%  Similarity=0.354  Sum_probs=27.2

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE  204 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~  204 (526)
                      ..+.|.+. ||+.+.+++..+|...+|.|=...
T Consensus        11 ~~v~V~l~-dgR~~~G~l~~~D~~~NivL~~~~   42 (75)
T cd06168          11 RTMRIHMT-DGRTLVGVFLCTDRDCNIILGSAQ   42 (75)
T ss_pred             CeEEEEEc-CCeEEEEEEEEEcCCCcEEecCcE
Confidence            46888997 999999999999999998775554


No 114
>cd01732 LSm5 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm4 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=24.68  E-value=1.5e+02  Score=23.90  Aligned_cols=31  Identities=16%  Similarity=0.388  Sum_probs=26.7

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEe
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTV  203 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv  203 (526)
                      ..+.|.+. +|+.+.+++.++|...++.|=..
T Consensus        14 ~~V~V~l~-~gr~~~G~L~g~D~~mNlvL~da   44 (76)
T cd01732          14 SRIWIVMK-SDKEFVGTLLGFDDYVNMVLEDV   44 (76)
T ss_pred             CEEEEEEC-CCeEEEEEEEEeccceEEEEccE
Confidence            57888887 89999999999999999887554


No 115
>cd01729 LSm7 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm7 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=24.64  E-value=1.6e+02  Score=23.96  Aligned_cols=32  Identities=3%  Similarity=0.074  Sum_probs=26.9

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE  204 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~  204 (526)
                      ..+.|.+. +|+.+.+++.++|...+|.|=...
T Consensus        13 k~V~V~l~-~gr~~~G~L~~~D~~mNlvL~~~~   44 (81)
T cd01729          13 KKIRVKFQ-GGREVTGILKGYDQLLNLVLDDTV   44 (81)
T ss_pred             CeEEEEEC-CCcEEEEEEEEEcCcccEEecCEE
Confidence            46888887 899999999999999988875543


No 116
>cd01735 LSm12_N LSm12 belongs to a family of Sm-like proteins that associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet that associates with other Sm proteins to form hexameric and heptameric ring structures.   In addition to the N-terminal Sm-like domain, LSm12 has a novel methyltransferase domain.
Probab=24.63  E-value=2.6e+02  Score=21.69  Aligned_cols=33  Identities=15%  Similarity=0.234  Sum_probs=27.2

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecc
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVED  205 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~  205 (526)
                      ..+.+... .|..++++|+.+|....+.+||-+.
T Consensus         7 s~V~~kTc-~g~~ieGEV~afD~~tk~lIlk~~s   39 (61)
T cd01735           7 SQVSCRTC-FEQRLQGEVVAFDYPSKMLILKCPS   39 (61)
T ss_pred             cEEEEEec-CCceEEEEEEEecCCCcEEEEECcc
Confidence            44556665 6899999999999999999999665


No 117
>PF11874 DUF3394:  Domain of unknown function (DUF3394);  InterPro: IPR021814  This domain is functionally uncharacterised. This domain is found in bacteria. This presumed domain is about 190 amino acids in length. This domain is found associated with PF06808 from PFAM. 
Probab=24.03  E-value=71  Score=30.37  Aligned_cols=28  Identities=14%  Similarity=0.051  Sum_probs=25.1

Q ss_pred             ccceEEEEeCCCCcccC-CCCCCcEEEEE
Q 009784          352 QKGVRIRRVDPTAPESE-VLKPSDIILSF  379 (526)
Q Consensus       352 ~~Gv~V~~V~~~spA~~-GL~~GDvIl~v  379 (526)
                      ...+.|..|..+|||++ |+.-|+.|+++
T Consensus       121 ~~~~~Vd~v~fgS~A~~~g~d~d~~I~~v  149 (183)
T PF11874_consen  121 GGKVIVDEVEFGSPAEKAGIDFDWEITEV  149 (183)
T ss_pred             CCEEEEEecCCCCHHHHcCCCCCcEEEEE
Confidence            45689999999999999 99999988887


No 118
>COG2524 Predicted transcriptional regulator, contains C-terminal CBS domains [Transcription]
Probab=23.39  E-value=2.2e+02  Score=28.72  Aligned_cols=94  Identities=26%  Similarity=0.321  Sum_probs=47.9

Q ss_pred             ccCCCeEEEEecccccccCceeeecCCCCcCC----CcEEEEeeCCCCCc----eeE-EEEEEeceeeeeccCCceeeeE
Q 009784          193 GTECDIAMLTVEDDEFWEGVLPVEFGELPALQ----DAVTVVGYPIGGDT----ISV-TSGVVSRIEILSYVHGSTELLG  263 (526)
Q Consensus       193 d~~~DlAlLkv~~~~~~~~~~pl~l~~~~~~g----~~V~~iG~p~~~~~----~sv-~~GiVs~~~~~~~~~~~~~~~~  263 (526)
                      ++..--|.+++..+     +..+..+++..+|    .++.+.|-=.+.+.    ..+ ...++|--..............
T Consensus       110 ~p~~c~a~i~v~Gd-----i~~~~~gD~VrVGPtP~~klvv~G~V~g~Dd~~~~ilidi~~m~siPk~~V~~~~s~~~i~  184 (294)
T COG2524         110 NPDACRAVIRVVGD-----IRKINIGDSVRVGPTPVNKLVVEGKVIGRDDTANEILIDISKMVSIPKEKVKNLMSKKLIT  184 (294)
T ss_pred             CCcccceEEEEEec-----cccCCCCCeEEECCcccceEEEEeEEecccccCCeEEEEEeeeeecCcchhhhhccCCceE
Confidence            45556778888764     5666677766665    33555554333221    111 1222221110000001111222


Q ss_pred             EEEccc--------CCCCCCCCeeecCCCeEEEEEee
Q 009784          264 LQIDAA--------INSGNSGGPAFNDKGKCVGIAFQ  292 (526)
Q Consensus       264 i~~da~--------i~~G~SGGPlvn~~G~VVGI~~~  292 (526)
                      +..|++        ...|-.|.|++|.+ ++||+.+.
T Consensus       185 v~~d~tl~eaak~f~~~~i~GaPVvd~d-k~vGiit~  220 (294)
T COG2524         185 VRPDDTLREAAKLFYEKGIRGAPVVDDD-KIVGIITL  220 (294)
T ss_pred             ecCCccHHHHHHHHHHcCccCCceecCC-ceEEEEEH
Confidence            333443        24799999999965 99999875


No 119
>cd01719 Sm_G The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet.  Sm subunit G binds subunits E and F to form a trimer which then assembles onto snRNA along with the D1/D2 and D3/B heterodimers forming a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=22.69  E-value=1.9e+02  Score=22.89  Aligned_cols=32  Identities=9%  Similarity=0.104  Sum_probs=26.8

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE  204 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~  204 (526)
                      ..+.|.+. +|+.+.+++.++|...+|.|=...
T Consensus        11 k~V~V~L~-~g~~~~G~L~~~D~~mNlvL~~~~   42 (72)
T cd01719          11 KKLSLKLN-GNRKVSGILRGFDPFMNLVLDDAV   42 (72)
T ss_pred             CeEEEEEC-CCeEEEEEEEEEcccccEEeccEE
Confidence            46888886 899999999999999888875543


No 120
>PF00595 PDZ:  PDZ domain (Also known as DHR or GLGF) Coordinates are not yet available;  InterPro: IPR001478 PDZ domains are found in diverse signalling proteins in bacteria, yeasts, plants, insects and vertebrates [, ]. PDZ domains can occur in one or multiple copies and are nearly always found in cytoplasmic proteins. They bind either the carboxyl-terminal sequences of proteins or internal peptide sequences []. In most cases, interaction between a PDZ domain and its target is constitutive, with a binding affinity of 1 to 10 microns. However, agonist-dependent activation of cell surface receptors is sometimes required to promote interaction with a PDZ protein. PDZ domain proteins are frequently associated with the plasma membrane, a compartment where high concentrations of phosphatidylinositol 4,5-bisphosphate (PIP2) are found. Direct interaction between PIP2 and a subset of class II PDZ domains (syntenin, CASK, Tiam-1) has been demonstrated.  PDZ domains consist of 80 to 90 amino acids comprising six beta-strands (beta-A to beta-F) and two alpha-helices, A and B, compactly arranged in a globular structure. Peptide binding of the ligand takes place in an elongated surface groove as an anti-parallel beta-strand interacts with the beta-B strand and the B helix. The structure of PDZ domains allows binding to a free carboxylate group at the end of a peptide through a carboxylate-binding loop between the beta-A and beta-B strands.; GO: 0005515 protein binding; PDB: 3AXA_A 1WF8_A 1QAV_B 1QAU_A 1B8Q_A 1MC7_A 2KAW_A 1I16_A 1VB7_A 1WI4_A ....
Probab=22.55  E-value=1.8e+02  Score=22.77  Aligned_cols=52  Identities=12%  Similarity=-0.073  Sum_probs=34.7

Q ss_pred             eEEEEeccccc-eeeeeeeecc--hhhhccccccceeeeecccchhhhHHHHHHH
Q 009784          453 GFVFSRCLYLI-SVLSMERIMN--MKLRSSFWTSSCIQCHNCQMSSLLWCLRCLW  504 (526)
Q Consensus       453 G~~f~~lt~~~-~~~~~~~i~~--~~~~sg~~~~~~~~~~~~~~~~~~~~~~~~~  504 (526)
                      |+.+....... ....+..+.+  .+.++|++.||.|...|.+-.....+..+.=
T Consensus        13 G~~l~~~~~~~~~~~~V~~v~~~~~a~~~gl~~GD~Il~INg~~v~~~~~~~~~~   67 (81)
T PF00595_consen   13 GFTLRGGSDNDEKGVFVSSVVPGSPAERAGLKVGDRILEINGQSVRGMSHDEVVQ   67 (81)
T ss_dssp             SEEEEEESTSSSEEEEEEEECTTSHHHHHTSSTTEEEEEETTEESTTSBHHHHHH
T ss_pred             CEEEEecCCCCcCCEEEEEEeCCChHHhcccchhhhhheeCCEeCCCCCHHHHHH
Confidence            55555544332 3444555555  4567899999999999998887776666543


No 121
>cd01727 LSm8 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm8 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=21.91  E-value=3.4e+02  Score=21.46  Aligned_cols=33  Identities=6%  Similarity=0.104  Sum_probs=27.7

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecc
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVED  205 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~  205 (526)
                      ..+.|.+. +|+.+.+++.++|...++.|=....
T Consensus        10 ~~V~V~l~-dgr~~~G~L~~~D~~~NlvL~~~~E   42 (74)
T cd01727          10 KTVSVITV-DGRVIVGTLKGFDQATNLILDDSHE   42 (74)
T ss_pred             CEEEEEEC-CCcEEEEEEEEEccccCEEccceEE
Confidence            46788886 9999999999999999988876543


No 122
>cd01721 Sm_D3 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. Sm subunit D3 heterodimerizes with subunit B and three such heterodimers form a hexameric ring structure with alternating B and D3 subunits. The D3 - B heterodimer also assembles into a heptameric ring containing D1, D2, E, F, and G subunits. Sm-like proteins exist in archaea as well as prokaryotes which form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=21.74  E-value=2.1e+02  Score=22.41  Aligned_cols=32  Identities=9%  Similarity=0.198  Sum_probs=28.3

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE  204 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~  204 (526)
                      ..+.|.+- +|..|.+++..+|...++.+-.+.
T Consensus        11 ~~V~VeLk-~g~~~~G~L~~~D~~MNl~L~~~~   42 (70)
T cd01721          11 HIVTVELK-TGEVYRGKLIEAEDNMNCQLKDVT   42 (70)
T ss_pred             CEEEEEEC-CCcEEEEEEEEEcCCceeEEEEEE
Confidence            56888887 899999999999999999988774


No 123
>smart00651 Sm snRNP Sm proteins. small nuclear ribonucleoprotein particles (snRNPs) involved in pre-mRNA splicing
Probab=21.30  E-value=2.2e+02  Score=21.56  Aligned_cols=33  Identities=15%  Similarity=0.288  Sum_probs=27.4

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEecc
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVED  205 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~  205 (526)
                      ..+.|.+. +|+.+.+.+..+|...++-|=....
T Consensus         9 ~~V~V~l~-~g~~~~G~L~~~D~~~NlvL~~~~e   41 (67)
T smart00651        9 KRVLVELK-NGREYRGTLKGFDQFMNLVLEDVEE   41 (67)
T ss_pred             cEEEEEEC-CCcEEEEEEEEECccccEEEccEEE
Confidence            46888887 8999999999999998888766543


No 124
>cd01728 LSm1 The eukaryotic Sm and Sm-like (LSm) proteins associate with RNA to form the core domain of the ribonucleoprotein particles involved in a variety of RNA processing events including pre-mRNA splicing, telomere replication, and mRNA degradation.  Members of this family share a highly conserved Sm fold containing an N-terminal helix followed by a strongly bent five-stranded antiparallel beta-sheet. LSm1 is one of at least seven subunits that assemble onto U6 snRNA to form a seven-membered ring structure.  Sm-like proteins exist in archaea as well as prokaryotes that form heptameric and hexameric ring structures similar to those found in eukaryotes.
Probab=20.75  E-value=2.2e+02  Score=22.81  Aligned_cols=32  Identities=9%  Similarity=0.155  Sum_probs=27.2

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEec
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVE  204 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~  204 (526)
                      ..+.|.+. +|+.+.+.+.++|+..++.|=...
T Consensus        13 k~v~V~l~-~gr~~~G~L~~fD~~~NlvL~d~~   44 (74)
T cd01728          13 KKVVVLLR-DGRKLIGILRSFDQFANLVLQDTV   44 (74)
T ss_pred             CEEEEEEc-CCeEEEEEEEEECCcccEEecceE
Confidence            46888887 899999999999999998876554


No 125
>PF01423 LSM:  LSM domain ;  InterPro: IPR001163 This family is found in Lsm (like-Sm) proteins and in bacterial Lsm-related Hfq proteins. In each case, the domain adopts a core structure consisting of an open beta-barrel with an SH3-like topology. Lsm (like-Sm) proteins have diverse functions, and are thought to be important modulators of RNA biogenesis and function [, ]. The Sm proteins form part of specific small nuclear ribonucleoproteins (snRNPs) that are involved in the processing of pre-mRNAs to mature mRNAs, and are a major component of the eukaryotic spliceosome. Most snRNPs consist of seven Sm proteins (B/B', D1, D2, D3, E, F and G) arranged in a ring on a uridine-rich sequence (Sm site), plus a small nuclear RNA (snRNA) (either U1, U2, U5 or U4/6) []. All Sm proteins contain a common sequence motif in two segments, Sm1 and Sm2, separated by a short variable linker []. In other snRNPs, certain Sm proteins are replaced with different Lsm proteins, such as with U7 snRNPs, in which the D1 and D2 Sm proteins are replaced with U7-specific Lsm10 and Lsm11 proteins, where Lsm11 plays a role in histone U7-specific RNA processing []. Lsm proteins are also found in archaebacteria, which do not have any splicing apparatus suggesting a more general role for Lsm proteins. The pleiotropic translational regulator Hfq (host factor Q) is a bacterial Lsm-like protein, which modulates the structure of numerous RNA molecules by binding preferentially to A/U-rich sequences in RNA []. Hfq forms an Lsm-like fold, however, unlike the heptameric Sm proteins, Hfq forms a homo-hexameric ring.; PDB: 1D3B_K 2Y9D_D 2Y9A_D 2Y9C_R 3VRI_C 2Y9B_K 3QUI_D 3M4G_H 3INZ_E 1U1S_C ....
Probab=20.65  E-value=1.7e+02  Score=22.28  Aligned_cols=34  Identities=12%  Similarity=0.331  Sum_probs=29.1

Q ss_pred             CeEEEEEcCCCcEEEEEEEEeccCCCeEEEEeccc
Q 009784          172 TQVKLKKRGSDTKYLATVLAIGTECDIAMLTVEDD  206 (526)
Q Consensus       172 ~~i~V~~~~~g~~~~a~vv~~d~~~DlAlLkv~~~  206 (526)
                      ..+.|.+. +|+.+.+.+..+|...++.|-.....
T Consensus         9 ~~V~V~l~-~g~~~~G~L~~~D~~~Nl~L~~~~~~   42 (67)
T PF01423_consen    9 KRVRVELK-NGRTYRGTLVSFDQFMNLVLSDVTET   42 (67)
T ss_dssp             SEEEEEET-TSEEEEEEEEEEETTEEEEEEEEEEE
T ss_pred             cEEEEEEe-CCEEEEEEEEEeechheEEeeeEEEE
Confidence            56889997 99999999999999988888777653


No 126
>COG0260 PepB Leucyl aminopeptidase [Amino acid transport and metabolism]
Probab=20.22  E-value=1.3e+02  Score=33.10  Aligned_cols=46  Identities=17%  Similarity=0.246  Sum_probs=28.7

Q ss_pred             EEeCCCCcccCCCCCCcEEEEECCEEecCCCCcccccccchhhhhhhhh
Q 009784          358 RRVDPTAPESEVLKPSDIILSFDGIDIANDGTVPFRHGERIGFSYLVSQ  406 (526)
Q Consensus       358 ~~V~~~spA~~GL~~GDvIl~vnG~~I~~~~~l~~~~~~~~~~~~~l~~  406 (526)
                      .-...|.|.....+|||||++.||+.|.=...=   -.+|+.+-+.+..
T Consensus       303 l~~~ENm~~g~A~rPGDVits~~GkTVEV~NTD---AEGRLVLADaLtY  348 (485)
T COG0260         303 LPAVENMPSGNAYRPGDVITSMNGKTVEVLNTD---AEGRLVLADALTY  348 (485)
T ss_pred             EeeeccCCCCCCCCCCCeEEecCCcEEEEcccC---ccHHHHHHHHHHH
Confidence            344455555556799999999999887522110   1266666665543


Done!