Query 022276
Match_columns 300
No_of_seqs 235 out of 1711
Neff 8.5
Searched_HMMs 46136
Date Fri Mar 29 09:22:43 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/022276.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/022276hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 KOG1542 Cysteine proteinase Ca 100.0 4.6E-71 1E-75 485.3 20.1 260 25-295 45-306 (372)
2 PTZ00203 cathepsin L protease; 100.0 1.5E-60 3.2E-65 436.9 28.9 238 48-295 33-280 (348)
3 PTZ00021 falcipain-2; Provisio 100.0 3.6E-58 7.7E-63 433.0 24.3 240 48-298 164-417 (489)
4 PTZ00200 cysteine proteinase; 100.0 4.8E-56 1E-60 417.3 27.5 236 48-298 121-383 (448)
5 KOG1543 Cysteine proteinase Ca 100.0 3.2E-52 7E-57 379.2 24.4 232 57-299 30-265 (325)
6 cd02621 Peptidase_C1A_Cathepsi 100.0 6.6E-39 1.4E-43 282.7 15.6 150 137-297 1-168 (243)
7 cd02698 Peptidase_C1A_Cathepsi 100.0 8.6E-38 1.9E-42 274.7 16.0 149 137-298 1-174 (239)
8 cd02248 Peptidase_C1A Peptidas 100.0 4.5E-37 9.8E-42 265.0 16.5 151 138-298 1-153 (210)
9 cd02620 Peptidase_C1A_Cathepsi 100.0 3.4E-37 7.4E-42 270.4 15.1 150 138-297 1-179 (236)
10 PTZ00364 dipeptidyl-peptidase 100.0 2.2E-36 4.8E-41 288.4 14.8 150 135-295 203-376 (548)
11 PTZ00049 cathepsin C-like prot 100.0 3.1E-36 6.8E-41 290.4 15.2 152 134-296 378-591 (693)
12 PF00112 Peptidase_C1: Papain 100.0 1.7E-34 3.8E-39 249.8 12.0 154 137-298 1-160 (219)
13 smart00645 Pept_C1 Papain fami 100.0 9.3E-32 2E-36 225.3 11.8 114 137-298 1-114 (174)
14 cd02619 Peptidase_C1 C1 Peptid 100.0 3E-30 6.5E-35 223.8 15.1 148 140-294 1-157 (223)
15 KOG1544 Predicted cysteine pro 100.0 9.7E-31 2.1E-35 228.0 2.9 207 83-298 152-389 (470)
16 PTZ00462 Serine-repeat antigen 100.0 2.8E-28 6E-33 241.8 14.1 142 149-298 544-716 (1004)
17 PF08246 Inhibitor_I29: Cathep 99.7 2.1E-16 4.6E-21 108.0 7.0 57 53-109 1-58 (58)
18 smart00848 Inhibitor_I29 Cathe 99.5 4.6E-14 1E-18 96.0 5.4 56 53-108 1-57 (57)
19 COG4870 Cysteine protease [Pos 99.4 2.2E-13 4.7E-18 122.4 2.7 151 136-297 98-260 (372)
20 cd00585 Peptidase_C1B Peptidas 98.7 1.8E-07 3.8E-12 88.5 11.6 83 150-235 55-159 (437)
21 PF03051 Peptidase_C1_2: Pepti 97.7 5.8E-05 1.2E-09 71.7 5.4 83 150-235 56-160 (438)
22 COG3579 PepC Aminopeptidase C 91.4 0.26 5.7E-06 44.8 4.2 84 151-235 59-162 (444)
23 KOG4128 Bleomycin hydrolases a 89.9 0.24 5.2E-06 45.0 2.6 86 149-235 62-169 (457)
24 PF08127 Propeptide_C1: Peptid 88.2 0.34 7.3E-06 30.2 1.6 34 82-117 4-37 (41)
25 PF07172 GRP: Glycine rich pro 78.7 1.7 3.7E-05 32.4 2.2 9 1-9 1-9 (95)
26 PF08139 LPAM_1: Prokaryotic m 65.0 4.7 0.0001 22.2 1.4 15 2-16 8-22 (25)
27 COG5510 Predicted small secret 61.9 8.9 0.00019 24.0 2.3 15 1-15 2-16 (44)
28 PRK10081 entericidin B membran 59.4 10 0.00022 24.4 2.3 13 1-13 2-14 (48)
29 PF10731 Anophelin: Thrombin i 57.3 11 0.00025 25.2 2.4 19 1-19 1-20 (65)
30 PRK10386 curli assembly protei 56.8 22 0.00047 28.1 4.4 19 1-19 1-19 (130)
31 PF11777 DUF3316: Protein of u 54.1 11 0.00024 28.9 2.4 19 1-19 1-19 (114)
32 PF05984 Cytomega_UL20A: Cytom 52.5 14 0.0003 26.6 2.4 21 1-21 1-22 (100)
33 PRK09810 entericidin A; Provis 42.7 25 0.00054 21.8 2.2 9 1-9 2-10 (41)
34 PF13529 Peptidase_C39_2: Pept 40.2 1.6E+02 0.0035 22.2 7.3 20 262-281 88-107 (144)
35 PRK10053 hypothetical protein; 40.0 24 0.00052 27.9 2.3 19 1-19 1-19 (130)
36 PF02402 Lysis_col: Lysis prot 36.7 14 0.00031 23.1 0.4 20 1-20 1-22 (46)
37 PRK10449 heat-inducible protei 35.2 32 0.00068 27.5 2.3 19 1-19 1-19 (140)
38 PF06291 Lambda_Bor: Bor prote 32.5 27 0.00058 26.1 1.4 21 1-21 1-21 (97)
39 PF11106 YjbE: Exopolysacchari 32.1 42 0.00091 23.8 2.2 15 1-15 1-15 (80)
40 PF05543 Peptidase_C47: Stapho 32.1 2.8E+02 0.006 23.1 7.3 53 154-223 18-78 (175)
41 PRK13883 conjugal transfer pro 31.7 32 0.00068 28.0 1.8 19 1-19 1-19 (151)
42 PF12276 DUF3617: Protein of u 30.0 41 0.00088 27.2 2.2 15 1-15 1-15 (162)
43 PF10614 CsgF: Type VIII secre 29.6 96 0.0021 24.9 4.1 31 53-83 42-79 (142)
44 PF11153 DUF2931: Protein of u 28.7 42 0.0009 28.8 2.2 18 1-19 1-18 (216)
45 PF11567 PfUIS3: Plasmodium fa 28.6 35 0.00075 24.8 1.3 29 68-108 18-46 (101)
46 KOG4702 Uncharacterized conser 27.5 1.6E+02 0.0035 20.5 4.3 35 48-83 26-60 (77)
47 COG3462 Predicted membrane pro 25.9 2.1E+02 0.0046 21.8 5.1 22 53-74 91-112 (117)
48 PF04202 Mfp-3: Foot protein 3 25.6 62 0.0013 22.3 2.0 16 1-16 1-16 (71)
49 PRK15346 outer membrane secret 23.9 54 0.0012 32.1 2.2 21 1-21 1-21 (499)
50 PRK11443 lipoprotein; Provisio 23.7 61 0.0013 25.4 2.1 18 1-19 1-18 (124)
51 PF08138 Sex_peptide: Sex pept 23.3 27 0.00059 22.9 0.0 12 1-12 1-12 (56)
52 KOG2735 Phosphatidylserine syn 22.9 58 0.0013 30.7 2.0 21 160-180 374-394 (466)
53 PF10880 DUF2673: Protein of u 22.7 86 0.0019 20.8 2.2 21 1-21 1-21 (65)
54 TIGR00156 conserved hypothetic 22.6 72 0.0016 25.1 2.2 12 1-12 1-12 (126)
55 PF02553 CbiN: Cobalt transpor 22.5 74 0.0016 22.5 2.1 12 1-12 1-12 (74)
56 PRK13835 conjugal transfer pro 21.6 65 0.0014 25.9 1.8 19 1-19 1-19 (145)
57 TIGR01165 cbiN cobalt transpor 21.5 87 0.0019 23.1 2.3 12 1-12 3-14 (91)
58 PF11912 DUF3430: Protein of u 20.9 76 0.0016 26.8 2.3 17 1-17 1-17 (212)
59 PF09403 FadA: Adhesion protei 20.9 87 0.0019 24.6 2.4 21 44-64 27-47 (126)
60 COG4871 Uncharacterized protei 20.3 59 0.0013 26.7 1.3 14 153-166 137-152 (193)
No 1
>KOG1542 consensus Cysteine proteinase Cathepsin F [Posttranslational modification, protein turnover, chaperones]
Probab=100.00 E-value=4.6e-71 Score=485.26 Aligned_cols=260 Identities=55% Similarity=0.924 Sum_probs=237.5
Q ss_pred CCCCCceeeecCCCCCCcchhhccHHHHHHHHHHHhCCccCCHHHHHHHHHHHHHHHHHHHHhcCCCC-CeeeeeccCCC
Q 022276 25 NDDDAMIRQVVPSDGEQSEDHLLNAEHHFSLFKSKFSKTYATQEEHDYRFRVFKANLRRAKRRQLLDP-TAVHGVTKFSD 103 (300)
Q Consensus 25 ~~~~~~i~~~~~~~i~~~~~~l~~~~~~F~~f~~~~~k~Y~s~~E~~~r~~~F~~n~~~I~~~N~~~~-s~~~giN~FsD 103 (300)
..++..|+++.... +.+...++.+++|..|+.+|+|+|.+.+|+.+|+.+|+.|+..+++++..++ +.++|+|+|||
T Consensus 45 ~~~~~~i~~v~~~~--~~~~~~l~~~~~F~~F~~kf~r~Y~s~eE~~~Rl~iF~~N~~~a~~~q~~d~gsA~yGvtqFSD 122 (372)
T KOG1542|consen 45 LGDDLTIRQVVRLQ--DLNPRGLGLEDSFKLFTIKFGRSYASREEHAHRLSIFKHNLLRAERLQENDPGSAEYGVTQFSD 122 (372)
T ss_pred cchhhhhhhhhhhc--ccCCcccchHHHHHHHHHhcCcccCcHHHHHHHHHHHHHHHHHHHHhhhcCccccccCccchhh
Confidence 45788888887532 2345666779999999999999999999999999999999999999999887 99999999999
Q ss_pred CChhhHHhhhcCCCcc-CCCCCCCCCCCCCCCCCCCCceecCCCCCccccccCCCCcchHHHHHHHHHHHHHHHhcCCCc
Q 022276 104 LTPSEFRRQFLGLNRR-LRLPADAQKAPILPTNDLPTDFDWRDHGAVTGVKDQGACGSCWSFSATGALEGAHFLSTGELV 182 (300)
Q Consensus 104 lt~~Ef~~~~~g~~~~-~~~~~~~~~~~~~~~~~lP~s~DwR~~g~v~pvknQg~CgsCwAfa~~~~~e~~~~i~~~~~~ 182 (300)
||++||++++++.+.. .+.+.....++..+...+|++||||++|.||||||||.||||||||+++++|++++|++|+++
T Consensus 123 lT~eEFkk~~l~~~~~~~~~~~~~~~~~~~~~~~lP~~fDWR~kgaVTpVKnQG~CGSCWAFS~tG~vEga~~i~~g~Lv 202 (372)
T KOG1542|consen 123 LTEEEFKKIYLGVKRRGSKLPGDAAEAPIEPGESLPESFDWRDKGAVTPVKNQGMCGSCWAFSTTGAVEGAWAIATGKLV 202 (372)
T ss_pred cCHHHHHHHhhccccccccCccccccCcCCCCCCCCcccchhccCCccccccCCcCcchhhhhhhhhhhhHHHhhcCccc
Confidence 9999999999887763 344444455555677899999999999999999999999999999999999999999999999
Q ss_pred cCChhHHHhhCCCCCCCCCCCCCCCCCCCChHHHHHHHHHhCCcCCCcccccCCCCCCCCCCCCCCceEEEceeEEcChh
Q 022276 183 SLSEQQLVDCDHECDPEESGSCDSGCNGGLMNSAFEYILKAGGVEREKDYPYTGTDGGSCKFDKSKIAAAVSNFSVISSD 262 (300)
Q Consensus 183 ~lS~Q~lidC~~~~~~~~~~~~~~gC~GG~~~~a~~y~~~~~G~~~e~~yPY~~~~~~~C~~~~~~~~~~i~~~~~v~~~ 262 (300)
+||||||+||+. +++||+||.+.+||+|+++.+||+.|++|||++.++..|..++...++.|.+|..++.|
T Consensus 203 sLSEQeLvDCD~---------~d~gC~GGl~~nA~~~~~~~gGL~~E~dYPY~g~~~~~C~~~~~~~~v~I~~f~~l~~n 273 (372)
T KOG1542|consen 203 SLSEQELVDCDS---------CDNGCNGGLMDNAFKYIKKAGGLEKEKDYPYTGKKGNQCHFDKSKIVVSIKDFSMLSNN 273 (372)
T ss_pred ccchhhhhcccC---------cCCcCCCCChhHHHHHHHHhCCccccccCCccccCCCccccchhhceEEEeccEecCCC
Confidence 999999999996 59999999999999999888999999999999999459999999999999999999999
Q ss_pred HHHHHHHHHhcCCeEEEEecCCCCCccCeeEec
Q 022276 263 EDQMAANLVKHGPLAGNVASIELPHISFSFLFT 295 (300)
Q Consensus 263 ~~~i~~al~~~GPv~v~i~a~~f~~Y~~Giy~~ 295 (300)
|++|.+.|+++|||+|+|+|..||+|++||..|
T Consensus 274 E~~ia~wLv~~GPi~vgiNa~~mQ~YrgGV~~P 306 (372)
T KOG1542|consen 274 EDQIAAWLVTFGPLSVGINAKPMQFYRGGVSCP 306 (372)
T ss_pred HHHHHHHHHhcCCeEEEEchHHHHHhcccccCC
Confidence 999999999999999999999999999999998
No 2
>PTZ00203 cathepsin L protease; Provisional
Probab=100.00 E-value=1.5e-60 Score=436.88 Aligned_cols=238 Identities=39% Similarity=0.646 Sum_probs=201.3
Q ss_pred cHHHHHHHHHHHhCCccCCHHHHHHHHHHHHHHHHHHHHhcCCCCCeeeeeccCCCCChhhHHhhhcCCCccCC-CCCCC
Q 022276 48 NAEHHFSLFKSKFSKTYATQEEHDYRFRVFKANLRRAKRRQLLDPTAVHGVTKFSDLTPSEFRRQFLGLNRRLR-LPADA 126 (300)
Q Consensus 48 ~~~~~F~~f~~~~~k~Y~s~~E~~~r~~~F~~n~~~I~~~N~~~~s~~~giN~FsDlt~~Ef~~~~~g~~~~~~-~~~~~ 126 (300)
..+.+|++|+.+|+|.|.+.+|+.+|+.+|++|++.|++||+++.+|++|+|+|+|||++||.+.+++...... .....
T Consensus 33 ~~~~~f~~~~~~~~K~Y~~~~E~~~R~~iF~~N~~~I~~~N~~~~~~~lg~N~FaDlT~eEf~~~~l~~~~~~~~~~~~~ 112 (348)
T PTZ00203 33 PAAALFEEFKRTYQRAYGTLTEEQQRLANFERNLELMREHQARNPHARFGITKFFDLSEAEFAARYLNGAAYFAAAKQHA 112 (348)
T ss_pred HHHHHHHHHHHHhCCCCCChHHHHHHHHHHHHHHHHHHHHhccCCCeEEeccccccCCHHHHHHHhcCCCcccccccccc
Confidence 35567999999999999998899999999999999999999888899999999999999999988764221110 00000
Q ss_pred -CCCCC--CCCCCCCCceecCCCCCccccccCCCCcchHHHHHHHHHHHHHHHhcCCCccCChhHHHhhCCCCCCCCCCC
Q 022276 127 -QKAPI--LPTNDLPTDFDWRDHGAVTGVKDQGACGSCWSFSATGALEGAHFLSTGELVSLSEQQLVDCDHECDPEESGS 203 (300)
Q Consensus 127 -~~~~~--~~~~~lP~s~DwR~~g~v~pvknQg~CgsCwAfa~~~~~e~~~~i~~~~~~~lS~Q~lidC~~~~~~~~~~~ 203 (300)
..... ....++|++||||++|+|+||||||.||||||||+++++|++++|++++.+.||+|||+||+..
T Consensus 113 ~~~~~~~~~~~~~lP~~~DWR~~g~VtpVkdQg~CGSCWAfa~~~aiEs~~~i~~~~~~~LSeQqLvdC~~~-------- 184 (348)
T PTZ00203 113 GQHYRKARADLSAVPDAVDWREKGAVTPVKNQGACGSCWAFSAVGNIESQWAVAGHKLVRLSEQQLVSCDHV-------- 184 (348)
T ss_pred cccccccccccccCCCCCcCCcCCCCCCccccCCCccHHHHhhHHHHHHHHHHhcCCCccCCHHHHHhccCC--------
Confidence 00000 1123689999999999999999999999999999999999999999999999999999999864
Q ss_pred CCCCCCCCChHHHHHHHHHh--CCcCCCcccccCCCCCC---CCCCCCC-CceEEEceeEEcChhHHHHHHHHHhcCCeE
Q 022276 204 CDSGCNGGLMNSAFEYILKA--GGVEREKDYPYTGTDGG---SCKFDKS-KIAAAVSNFSVISSDEDQMAANLVKHGPLA 277 (300)
Q Consensus 204 ~~~gC~GG~~~~a~~y~~~~--~G~~~e~~yPY~~~~~~---~C~~~~~-~~~~~i~~~~~v~~~~~~i~~al~~~GPv~ 277 (300)
+.||+||++..||+|++++ +|+++|++|||++.+ + .|+.... ...+++.+|..++.+++.|+++|+++|||+
T Consensus 185 -~~GC~GG~~~~a~~yi~~~~~ggi~~e~~YPY~~~~-~~~~~C~~~~~~~~~~~i~~~~~i~~~e~~~~~~l~~~GPv~ 262 (348)
T PTZ00203 185 -DNGCGGGLMLQAFEWVLRNMNGTVFTEKSYPYVSGN-GDVPECSNSSELAPGARIDGYVSMESSERVMAAWLAKNGPIS 262 (348)
T ss_pred -CCCCCCCCHHHHHHHHHHhcCCCCCccccCCCccCC-CCCCcCCCCcccccceEecceeecCcCHHHHHHHHHhCCCEE
Confidence 7899999999999999865 679999999999876 4 6875433 235678899888778899999999999999
Q ss_pred EEEecCCCCCccCeeEec
Q 022276 278 GNVASIELPHISFSFLFT 295 (300)
Q Consensus 278 v~i~a~~f~~Y~~Giy~~ 295 (300)
|+|++..|++|++|||..
T Consensus 263 v~i~a~~f~~Y~~GIy~~ 280 (348)
T PTZ00203 263 IAVDASSFMSYHSGVLTS 280 (348)
T ss_pred EEEEhhhhcCccCceeec
Confidence 999998899999999985
No 3
>PTZ00021 falcipain-2; Provisional
Probab=100.00 E-value=3.6e-58 Score=433.05 Aligned_cols=240 Identities=31% Similarity=0.564 Sum_probs=203.3
Q ss_pred cHHHHHHHHHHHhCCccCCHHHHHHHHHHHHHHHHHHHHhcCC-CCCeeeeeccCCCCChhhHHhhhcCCCcc-CCC-CC
Q 022276 48 NAEHHFSLFKSKFSKTYATQEEHDYRFRVFKANLRRAKRRQLL-DPTAVHGVTKFSDLTPSEFRRQFLGLNRR-LRL-PA 124 (300)
Q Consensus 48 ~~~~~F~~f~~~~~k~Y~s~~E~~~r~~~F~~n~~~I~~~N~~-~~s~~~giN~FsDlt~~Ef~~~~~g~~~~-~~~-~~ 124 (300)
+...+|++|+.+|+|+|.+.+|+..|+.+|++|+++|++||+. +.+|++|+|+|+|||.+||++.+++.... ... ..
T Consensus 164 e~~~~F~~wk~ky~K~Y~~~eE~~~R~~iF~~Nl~~Ie~hN~~~~~ty~lgiNqFsDlT~EEF~~~~l~~~~~~~~~~~~ 243 (489)
T PTZ00021 164 ENVNSFYLFIKEHGKKYQTPDEMQQRYLSFVENLAKINAHNNKENVLYKKGMNRFGDLSFEEFKKKYLTLKSFDFKSNGK 243 (489)
T ss_pred HHHHHHHHHHHHhCCcCCCHHHHHHHHHHHHHHHHHHHHhhccCCCCEEEeccccccCCHHHHHHHhccccccccccccc
Confidence 4446799999999999999999999999999999999999975 47999999999999999999988775421 000 00
Q ss_pred --C--CCCC----CCCCC--CCCCCceecCCCCCccccccCCCCcchHHHHHHHHHHHHHHHhcCCCccCChhHHHhhCC
Q 022276 125 --D--AQKA----PILPT--NDLPTDFDWRDHGAVTGVKDQGACGSCWSFSATGALEGAHFLSTGELVSLSEQQLVDCDH 194 (300)
Q Consensus 125 --~--~~~~----~~~~~--~~lP~s~DwR~~g~v~pvknQg~CgsCwAfa~~~~~e~~~~i~~~~~~~lS~Q~lidC~~ 194 (300)
. .... ...+. ...|+++|||+.|.|+||||||.||||||||+++++|++++|++++.+.||+|||+||+.
T Consensus 244 ~~~~~~~~~~~~~~~~~~~~~~~P~s~DWR~~g~VtpVKdQG~CGSCWAFAa~~alEs~~~I~~g~~v~LSeQqLVDCs~ 323 (489)
T PTZ00021 244 KSPRVINYDDVIKKYKPKDATFDHAKYDWRLHNGVTPVKDQKNCGSCWAFSTVGVVESQYAIRKNELVSLSEQELVDCSF 323 (489)
T ss_pred cccccccccccccccccccccCCccccccccCCCCCCcccccccccHHHHHHHHHHHHHHHHHcCCCcccCHHHHhhhcc
Confidence 0 0000 00011 124999999999999999999999999999999999999999999999999999999986
Q ss_pred CCCCCCCCCCCCCCCCCChHHHHHHHHHhCCcCCCcccccCCCCCCCCCCCCCCceEEEceeEEcChhHHHHHHHHHhcC
Q 022276 195 ECDPEESGSCDSGCNGGLMNSAFEYILKAGGVEREKDYPYTGTDGGSCKFDKSKIAAAVSNFSVISSDEDQMAANLVKHG 274 (300)
Q Consensus 195 ~~~~~~~~~~~~gC~GG~~~~a~~y~~~~~G~~~e~~yPY~~~~~~~C~~~~~~~~~~i~~~~~v~~~~~~i~~al~~~G 274 (300)
. +.||.||++..||+|+++++|+++|++|||.+..++.|........++|.+|..++ +++|+++|+.+|
T Consensus 324 ~---------n~GC~GG~~~~Af~yi~~~gGl~tE~~YPY~~~~~~~C~~~~~~~~~~i~~y~~i~--~~~lk~al~~~G 392 (489)
T PTZ00021 324 K---------NNGCYGGLIPNAFEDMIELGGLCSEDDYPYVSDTPELCNIDRCKEKYKIKSYVSIP--EDKFKEAIRFLG 392 (489)
T ss_pred C---------CCCCCCcchHhhhhhhhhccccCcccccCccCCCCCccccccccccceeeeEEEec--HHHHHHHHHhcC
Confidence 4 88999999999999998888999999999998744789876666678899998886 578999999999
Q ss_pred CeEEEEecC-CCCCccCeeEecCCC
Q 022276 275 PLAGNVASI-ELPHISFSFLFTVSS 298 (300)
Q Consensus 275 Pv~v~i~a~-~f~~Y~~Giy~~~~~ 298 (300)
||+|+|++. .|++|++|||.++++
T Consensus 393 PVsv~i~a~~~f~~YkgGIy~~~C~ 417 (489)
T PTZ00021 393 PISVSIAVSDDFAFYKGGIFDGECG 417 (489)
T ss_pred CeEEEEEeecccccCCCCcCCCCCC
Confidence 999999995 699999999987543
No 4
>PTZ00200 cysteine proteinase; Provisional
Probab=100.00 E-value=4.8e-56 Score=417.28 Aligned_cols=236 Identities=31% Similarity=0.529 Sum_probs=194.9
Q ss_pred cHHHHHHHHHHHhCCccCCHHHHHHHHHHHHHHHHHHHHhcCCCCCeeeeeccCCCCChhhHHhhhcCCCccCCCC----
Q 022276 48 NAEHHFSLFKSKFSKTYATQEEHDYRFRVFKANLRRAKRRQLLDPTAVHGVTKFSDLTPSEFRRQFLGLNRRLRLP---- 123 (300)
Q Consensus 48 ~~~~~F~~f~~~~~k~Y~s~~E~~~r~~~F~~n~~~I~~~N~~~~s~~~giN~FsDlt~~Ef~~~~~g~~~~~~~~---- 123 (300)
+...+|++|+++|+|.|.+.+|+.+|+.+|++|++.|++||. +.+|++|+|+|+|||++||.+.+++...+....
T Consensus 121 e~~~~F~~f~~ky~K~Y~~~~E~~~R~~iF~~Nl~~I~~hN~-~~~y~lgiN~FsDlT~eEF~~~~~~~~~~~~~~~~~~ 199 (448)
T PTZ00200 121 EVYLEFEEFNKKYNRKHATHAERLNRFLTFRNNYLEVKSHKG-DEPYSKEINKFSDLTEEEFRKLFPVIKVPPKSNSTSH 199 (448)
T ss_pred HHHHHHHHHHHHhCCcCCCHHHHHHHHHHHHHHHHHHHHhcC-cCCeEEeccccccCCHHHHHHHhccCCCccccccccc
Confidence 555789999999999999999999999999999999999996 568999999999999999998877644211000
Q ss_pred -----CC-CCCCCC-----------C----CCCCCCCceecCCCCCccccccCC-CCcchHHHHHHHHHHHHHHHhcCCC
Q 022276 124 -----AD-AQKAPI-----------L----PTNDLPTDFDWRDHGAVTGVKDQG-ACGSCWSFSATGALEGAHFLSTGEL 181 (300)
Q Consensus 124 -----~~-~~~~~~-----------~----~~~~lP~s~DwR~~g~v~pvknQg-~CgsCwAfa~~~~~e~~~~i~~~~~ 181 (300)
.+ ...... . +...+|++||||+.|.|+|||||| .||||||||+++++|++++|++++.
T Consensus 200 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~P~~~DWR~~g~vtpVkdQG~~CGSCWAFat~~aiEs~~~i~~~~~ 279 (448)
T PTZ00200 200 NNDFKARHVSNPTYLKNLKKAKNTDEDVKDPSKITGEGLDWRRADAVTKVKDQGLNCGSCWAFSSVGSVESLYKIYRDKS 279 (448)
T ss_pred ccccccccccccccccccccccccccccccccccCCCCccCCCCCCCCCcccCCCccchHHHHhHHHHHHHHHHHhcCCC
Confidence 00 000000 0 011369999999999999999999 9999999999999999999999999
Q ss_pred ccCChhHHHhhCCCCCCCCCCCCCCCCCCCChHHHHHHHHHhCCcCCCcccccCCCCCCCCCCCCCCceEEEceeEEcCh
Q 022276 182 VSLSEQQLVDCDHECDPEESGSCDSGCNGGLMNSAFEYILKAGGVEREKDYPYTGTDGGSCKFDKSKIAAAVSNFSVISS 261 (300)
Q Consensus 182 ~~lS~Q~lidC~~~~~~~~~~~~~~gC~GG~~~~a~~y~~~~~G~~~e~~yPY~~~~~~~C~~~~~~~~~~i~~~~~v~~ 261 (300)
+.||+|||+||+.. ++||+||++..||+|++++ |+++|++|||++.. +.|...... .+.|.+|..++
T Consensus 280 ~~LSeQqLvDC~~~---------~~GC~GG~~~~A~~yi~~~-Gi~~e~~YPY~~~~-~~C~~~~~~-~~~i~~y~~~~- 346 (448)
T PTZ00200 280 VDLSEQELVNCDTK---------SQGCSGGYPDTALEYVKNK-GLSSSSDVPYLAKD-GKCVVSSTK-KVYIDSYLVAK- 346 (448)
T ss_pred eecCHHHHhhccCc---------cCCCCCCcHHHHHHHHhhc-CccccccCCCCCCC-CCCcCCCCC-eeEecceEecC-
Confidence 99999999999864 7899999999999999776 89999999999988 899865433 46688887665
Q ss_pred hHHHHHHHHHhcCCeEEEEecC-CCCCccCeeEecCCC
Q 022276 262 DEDQMAANLVKHGPLAGNVASI-ELPHISFSFLFTVSS 298 (300)
Q Consensus 262 ~~~~i~~al~~~GPv~v~i~a~-~f~~Y~~Giy~~~~~ 298 (300)
+.+.++++ +.+|||+|+|++. .|++|++|||.++++
T Consensus 347 ~~~~l~~~-l~~GPV~v~i~~~~~f~~Yk~GIy~~~C~ 383 (448)
T PTZ00200 347 GKDVLNKS-LVISPTVVYIAVSRELLKYKSGVYNGECG 383 (448)
T ss_pred HHHHHHHH-HhcCCEEEEeecccccccCCCCccccccC
Confidence 44555555 4689999999996 599999999987543
No 5
>KOG1543 consensus Cysteine proteinase Cathepsin L [Posttranslational modification, protein turnover, chaperones]
Probab=100.00 E-value=3.2e-52 Score=379.22 Aligned_cols=232 Identities=39% Similarity=0.688 Sum_probs=201.1
Q ss_pred HHHhCCccCCHHHHHHHHHHHHHHHHHHHHhcCC-CCCeeeeeccCCCCChhhHHhhhcCCCccCCCCCCCCCCCCCCCC
Q 022276 57 KSKFSKTYATQEEHDYRFRVFKANLRRAKRRQLL-DPTAVHGVTKFSDLTPSEFRRQFLGLNRRLRLPADAQKAPILPTN 135 (300)
Q Consensus 57 ~~~~~k~Y~s~~E~~~r~~~F~~n~~~I~~~N~~-~~s~~~giN~FsDlt~~Ef~~~~~g~~~~~~~~~~~~~~~~~~~~ 135 (300)
+.+|.+.|.+..|...|+.+|++|++.|..||.. ..+|.+|+|+|+|++.+|+++.+.+.+.+.. ............
T Consensus 30 ~~~~~~~y~~~~~~~~r~~~f~~n~~~~~~~n~~~~~~~~~g~n~~~d~~~ee~~~~~~~~~~~~~--~~~~~~~~~~~~ 107 (325)
T KOG1543|consen 30 LVKFLKRYEDRVEKKARRAIFKENLQKIESHNLKYVLSFLMGVNQFADLTTEEFKRKKTGKKPPEI--KRDKFTEKLDGD 107 (325)
T ss_pred hhhhccccccHHHHHHHHHHHHHHHHHHHhhhhhhceeeeeccccccccchHHHHHhhccccCccc--cccccccccchh
Confidence 5667777777788899999999999999999987 7899999999999999999998887765322 111111122345
Q ss_pred CCCCceecCCCC-CccccccCCCCcchHHHHHHHHHHHHHHHhcC-CCccCChhHHHhhCCCCCCCCCCCCCCCCCCCCh
Q 022276 136 DLPTDFDWRDHG-AVTGVKDQGACGSCWSFSATGALEGAHFLSTG-ELVSLSEQQLVDCDHECDPEESGSCDSGCNGGLM 213 (300)
Q Consensus 136 ~lP~s~DwR~~g-~v~pvknQg~CgsCwAfa~~~~~e~~~~i~~~-~~~~lS~Q~lidC~~~~~~~~~~~~~~gC~GG~~ 213 (300)
++|++||||++| .++||||||.||||||||++++||++++|++| .++.||+|||+||+.. +++||.||.+
T Consensus 108 ~~p~s~DwR~~~~~~~~vkdQg~CgsCWAFaa~~aie~~~~i~~g~~l~sLSeq~lvdC~~~--------~~~GC~GG~~ 179 (325)
T KOG1543|consen 108 DLPDSFDWRDKGAVTPPVKDQGSCGSCWAFAATGALEDRYNIKTGGKLLSLSEQDLVDCCGE--------CGDGCNGGEP 179 (325)
T ss_pred hCCCCccccccCCcCCCcCCCCcCcchHHHHHHHHHHHHHHHHhCCccCccChhhhhhccCC--------CCCCcCCCCH
Confidence 899999999996 55569999999999999999999999999999 8999999999999984 5889999999
Q ss_pred HHHHHHHHHhCCcCCCcccccCCCCCCCCCCCCCCceEEEceeEEcChhHHHHHHHHHhcCCeEEEEecCC-CCCccCee
Q 022276 214 NSAFEYILKAGGVEREKDYPYTGTDGGSCKFDKSKIAAAVSNFSVISSDEDQMAANLVKHGPLAGNVASIE-LPHISFSF 292 (300)
Q Consensus 214 ~~a~~y~~~~~G~~~e~~yPY~~~~~~~C~~~~~~~~~~i~~~~~v~~~~~~i~~al~~~GPv~v~i~a~~-f~~Y~~Gi 292 (300)
..||+|++++||+..+.+|||.+.+ +.|..+.......+.++..++.++++|+++|+++|||+|+|+|.. |++|++||
T Consensus 180 ~~A~~yi~~~G~~t~~~~Ypy~~~~-~~C~~~~~~~~~~~~~~~~~~~~e~~i~~~v~~~GPv~v~~~a~~~F~~Y~~GV 258 (325)
T KOG1543|consen 180 KNAFKYIKKNGGVTECENYPYIGKD-GTCKSNKKDKTVTIKGFYNVPANEEAIAEAVAKNGPVSVAIDAYEDFSLYKGGV 258 (325)
T ss_pred HHHHHHHHHhCCCCCCcCCCCcCCC-CCccCCCccceeEeeeeeecCcCHHHHHHHHHhcCCeEEEEeehhhhhhccCce
Confidence 9999999999555559999999999 899998876778888888888899999999999999999999965 99999999
Q ss_pred EecCCCC
Q 022276 293 LFTVSSP 299 (300)
Q Consensus 293 y~~~~~~ 299 (300)
|.++++.
T Consensus 259 y~~~~~~ 265 (325)
T KOG1543|consen 259 YAEEKGD 265 (325)
T ss_pred EeCCCCC
Confidence 9999775
No 6
>cd02621 Peptidase_C1A_CathepsinC Cathepsin C; also known as Dipeptidyl Peptidase I (DPPI), an atypical papain-like cysteine peptidase with chloride dependency and dipeptidyl aminopeptidase activity, resulting from its tetrameric structure which limits substrate access. Each subunit of the tetramer is composed of three peptides: the heavy and light chains, which together adopts the papain fold and forms the catalytic domain; and the residual propeptide region, which forms a beta barrel and points towards the substrate's N-terminus. The subunit composition is the result of the unique characteristic of procathepsin C maturation involving the cleavage of the catalytic domain and the non-autocatalytic excision of an activation peptide within its propeptide region. By removing N-terminal dipeptide extensions, cathepsin C activates granule serine peptidases (granzymes) involved in cell-mediated apoptosis, inflammation and tissue remodelling. Loss-of-function mutations in cathepsin C are assoc
Probab=100.00 E-value=6.6e-39 Score=282.66 Aligned_cols=150 Identities=27% Similarity=0.546 Sum_probs=131.1
Q ss_pred CCCceecCCCC----CccccccCCCCcchHHHHHHHHHHHHHHHhcCC------CccCChhHHHhhCCCCCCCCCCCCCC
Q 022276 137 LPTDFDWRDHG----AVTGVKDQGACGSCWSFSATGALEGAHFLSTGE------LVSLSEQQLVDCDHECDPEESGSCDS 206 (300)
Q Consensus 137 lP~s~DwR~~g----~v~pvknQg~CgsCwAfa~~~~~e~~~~i~~~~------~~~lS~Q~lidC~~~~~~~~~~~~~~ 206 (300)
||++||||+.+ +|+||||||.||||||||+++++|++++|++++ .+.||+|||+||+.. +.
T Consensus 1 lP~~fDwr~~~~~~~~v~~v~dQg~CGsCwAfa~~~~ies~~~i~~~~~~~~~~~~~lS~q~l~dC~~~---------~~ 71 (243)
T cd02621 1 LPKSFDWGDVNNGFNYVSPVRNQGGCGSCYAFASVYALEARIMIASNKTDPLGQQPILSPQHVLSCSQY---------SQ 71 (243)
T ss_pred CCCcccccccCCCCcccccCCCCCcCccHHHHHHHHHHHHHHHHHhCCCCccccCcccCHHHhhhhcCC---------CC
Confidence 69999999998 999999999999999999999999999998876 689999999999864 78
Q ss_pred CCCCCChHHHHHHHHHhCCcCCCcccccCC-CCCCCCCCCC-CCceEEEceeEEcC-----hhHHHHHHHHHhcCCeEEE
Q 022276 207 GCNGGLMNSAFEYILKAGGVEREKDYPYTG-TDGGSCKFDK-SKIAAAVSNFSVIS-----SDEDQMAANLVKHGPLAGN 279 (300)
Q Consensus 207 gC~GG~~~~a~~y~~~~~G~~~e~~yPY~~-~~~~~C~~~~-~~~~~~i~~~~~v~-----~~~~~i~~al~~~GPv~v~ 279 (300)
||+||++..|++|+++. |+++|++|||++ .. +.|.... ....+++..|..+. .++++||++|+++|||+|+
T Consensus 72 GC~GG~~~~a~~~~~~~-Gi~~e~~yPY~~~~~-~~C~~~~~~~~~~~~~~~~~i~~~~~~~~~~~ik~~i~~~GPv~v~ 149 (243)
T cd02621 72 GCDGGFPFLVGKFAEDF-GIVTEDYFPYTADDD-RPCKASPSECRRYYFSDYNYVGGCYGCTNEDEMKWEIYRNGPIVVA 149 (243)
T ss_pred CCCCCCHHHHHHHHHhc-CcCCCceeCCCCCCC-CCCCCCccccccccccceeEcccccccCCHHHHHHHHHHcCCEEEE
Confidence 99999999999999877 899999999998 55 8898655 33445555555442 3789999999999999999
Q ss_pred EecC-CCCCccCeeEecCC
Q 022276 280 VASI-ELPHISFSFLFTVS 297 (300)
Q Consensus 280 i~a~-~f~~Y~~Giy~~~~ 297 (300)
|++. .|++|++|||..+.
T Consensus 150 ~~~~~~F~~Y~~GIy~~~~ 168 (243)
T cd02621 150 FEVYSDFDFYKEGVYHHTD 168 (243)
T ss_pred EEecccccccCCeEECcCC
Confidence 9995 59999999999863
No 7
>cd02698 Peptidase_C1A_CathepsinX Cathepsin X; the only papain-like lysosomal cysteine peptidase exhibiting carboxymonopeptidase activity. It can also act as a carboxydipeptidase, like cathepsin B, but has been shown to preferentially cleave substrates through a monopeptidyl carboxypeptidase pathway. The propeptide region of cathepsin X, the shortest among papain-like peptidases, is covalently attached to the active site cysteine in the inactive form of the enzyme. Little is known about the biological function of cathepsin X. Some studies point to a role in early tumorigenesis. A more recent study indicates that cathepsin X expression is restricted to immune cells suggesting a role in phagocytosis and the regulation of the immune response.
Probab=100.00 E-value=8.6e-38 Score=274.75 Aligned_cols=149 Identities=30% Similarity=0.547 Sum_probs=128.9
Q ss_pred CCCceecCCCC---CccccccCC---CCcchHHHHHHHHHHHHHHHhcC---CCccCChhHHHhhCCCCCCCCCCCCCCC
Q 022276 137 LPTDFDWRDHG---AVTGVKDQG---ACGSCWSFSATGALEGAHFLSTG---ELVSLSEQQLVDCDHECDPEESGSCDSG 207 (300)
Q Consensus 137 lP~s~DwR~~g---~v~pvknQg---~CgsCwAfa~~~~~e~~~~i~~~---~~~~lS~Q~lidC~~~~~~~~~~~~~~g 207 (300)
||++||||+++ +|+|||||| .||||||||++++||++++|+++ ..+.||+|||+||+. +.|
T Consensus 1 lP~~~Dwr~~~~~~~v~~vk~Qg~~~~CGsCwAfa~~~aies~~~i~~~~~~~~~~lS~Q~lldC~~----------~~g 70 (239)
T cd02698 1 LPKSWDWRNVNGVNYVSPTRNQHIPQYCGSCWAHGSTSALADRINIARKGAWPSVYLSVQVVIDCAG----------GGS 70 (239)
T ss_pred CCCCcccccCCCCcccCccccCCCCCCCCcchHHHhHHHHHHHHHHHHCCCCCCcccCHHHHHhCCC----------CCC
Confidence 69999999988 999999998 89999999999999999999875 357899999999985 679
Q ss_pred CCCCChHHHHHHHHHhCCcCCCcccccCCCCCCCCCCC---------------CCCceEEEceeEEcChhHHHHHHHHHh
Q 022276 208 CNGGLMNSAFEYILKAGGVEREKDYPYTGTDGGSCKFD---------------KSKIAAAVSNFSVISSDEDQMAANLVK 272 (300)
Q Consensus 208 C~GG~~~~a~~y~~~~~G~~~e~~yPY~~~~~~~C~~~---------------~~~~~~~i~~~~~v~~~~~~i~~al~~ 272 (300)
|+||++..|++|++++ |+++|++|||.+.+ +.|... +....+++++|..++ ++++||++|++
T Consensus 71 C~GG~~~~a~~~~~~~-Gl~~e~~yPY~~~~-~~C~~~~~~~~c~~~~~c~~~~~~~~~~i~~~~~~~-~~~~i~~~l~~ 147 (239)
T cd02698 71 CHGGDPGGVYEYAHKH-GIPDETCNPYQAKD-GECNPFNRCGTCNPFGECFAIKNYTLYFVSDYGSVS-GRDKMMAEIYA 147 (239)
T ss_pred ccCcCHHHHHHHHHHc-CcCCCCeeCCcCCC-CCCcCCCCCCCcccCcccccccccceEEeeeceecC-CHHHHHHHHHH
Confidence 9999999999999886 89999999999876 566531 112346777887775 67889999999
Q ss_pred cCCeEEEEecC-CCCCccCeeEecCCC
Q 022276 273 HGPLAGNVASI-ELPHISFSFLFTVSS 298 (300)
Q Consensus 273 ~GPv~v~i~a~-~f~~Y~~Giy~~~~~ 298 (300)
+|||+|+|++. .|++|++|||..+++
T Consensus 148 ~GPV~v~i~~~~~f~~Y~~GIy~~~~~ 174 (239)
T cd02698 148 RGPISCGIMATEALENYTGGVYKEYVQ 174 (239)
T ss_pred cCCEEEEEEecccccccCCeEEccCCC
Confidence 99999999996 599999999987654
No 8
>cd02248 Peptidase_C1A Peptidase C1A subfamily (MEROPS database nomenclature); composed of cysteine peptidases (CPs) similar to papain, including the mammalian CPs (cathepsins B, C, F, H, L, K, O, S, V, X and W). Papain is an endopeptidase with specific substrate preferences, primarily for bulky hydrophobic or aromatic residues at the S2 subsite, a hydrophobic pocket in papain that accommodates the P2 sidechain of the substrate (the second residue away from the scissile bond). Most members of the papain subfamily are endopeptidases. Some exceptions to this rule can be explained by specific details of the catalytic domains like the occluding loop in cathepsin B which confers an additional carboxydipeptidyl activity and the mini-chain of cathepsin H resulting in an N-terminal exopeptidase activity. Papain-like CPs have different functions in various organisms. Plant CPs are used to mobilize storage proteins in seeds. Parasitic CPs act extracellularly to help invade tissues and cells, to h
Probab=100.00 E-value=4.5e-37 Score=265.02 Aligned_cols=151 Identities=47% Similarity=0.886 Sum_probs=138.9
Q ss_pred CCceecCCCCCccccccCCCCcchHHHHHHHHHHHHHHHhcCCCccCChhHHHhhCCCCCCCCCCCCCCCCCCCChHHHH
Q 022276 138 PTDFDWRDHGAVTGVKDQGACGSCWSFSATGALEGAHFLSTGELVSLSEQQLVDCDHECDPEESGSCDSGCNGGLMNSAF 217 (300)
Q Consensus 138 P~s~DwR~~g~v~pvknQg~CgsCwAfa~~~~~e~~~~i~~~~~~~lS~Q~lidC~~~~~~~~~~~~~~gC~GG~~~~a~ 217 (300)
|++||||+.+.++||+|||.||+|||||+++++|++++++++....||+|+|++|... .+.||.||.+..|+
T Consensus 1 P~~~d~r~~~~~~~v~dQg~cgsCwAfa~~~~le~~~~i~~~~~~~lS~q~l~~c~~~--------~~~gC~GG~~~~a~ 72 (210)
T cd02248 1 PESVDWREKGAVTPVKDQGSCGSCWAFSTVGALEGAYAIKTGKLVSLSEQQLVDCSTS--------GNNGCNGGNPDNAF 72 (210)
T ss_pred CCcccCCcCCCCCCCccCCCCcchHHhHHHHHHHHHHHHHcCCCcccCHHHHhccCCC--------CCCCCCCCCHHHhH
Confidence 8899999999999999999999999999999999999999998899999999999863 36899999999999
Q ss_pred HHHHHhCCcCCCcccccCCCCCCCCCCCCCCceEEEceeEEcCh-hHHHHHHHHHhcCCeEEEEecC-CCCCccCeeEec
Q 022276 218 EYILKAGGVEREKDYPYTGTDGGSCKFDKSKIAAAVSNFSVISS-DEDQMAANLVKHGPLAGNVASI-ELPHISFSFLFT 295 (300)
Q Consensus 218 ~y~~~~~G~~~e~~yPY~~~~~~~C~~~~~~~~~~i~~~~~v~~-~~~~i~~al~~~GPv~v~i~a~-~f~~Y~~Giy~~ 295 (300)
+++.+. |+++|++|||.+.. ..|+........+|++|..++. +.++||++|+++|||+|+|.+. .|+.|++|||..
T Consensus 73 ~~~~~~-Gi~~e~~yPY~~~~-~~C~~~~~~~~~~i~~~~~i~~~~~~~ik~~l~~~gPV~~~~~~~~~f~~y~~Giy~~ 150 (210)
T cd02248 73 EYVKNG-GLASESDYPYTGKD-GTCKYNSSKVGAKITGYSNVPPGDEEALKAALANYGPVSVAIDASSSFQFYKGGIYSG 150 (210)
T ss_pred HHHHHC-CcCccccCCccCCC-CCccCCCCcccEEEeeEEEcCCCcHHHHHHHHhhcCCEEEEEecCcccccCCCCceeC
Confidence 988776 89999999999877 8898877667899999999876 5889999999999999999996 599999999998
Q ss_pred CCC
Q 022276 296 VSS 298 (300)
Q Consensus 296 ~~~ 298 (300)
+++
T Consensus 151 ~~~ 153 (210)
T cd02248 151 PCC 153 (210)
T ss_pred CCC
Confidence 765
No 9
>cd02620 Peptidase_C1A_CathepsinB Cathepsin B group; composed of cathepsin B and similar proteins, including tubulointerstitial nephritis antigen (TIN-Ag). Cathepsin B is a lysosomal papain-like cysteine peptidase which is expressed in all tissues and functions primarily as an exopeptidase through its carboxydipeptidyl activity. Together with other cathepsins, it is involved in the degradation of proteins, proenzyme activation, Ag processing, metabolism and apoptosis. Cathepsin B has been implicated in a number of human diseases such as cancer, rheumatoid arthritis, osteoporosis and Alzheimer's disease. The unique carboxydipeptidyl activity of cathepsin B is attributed to the presence of an occluding loop in its active site which favors the binding of the C-termini of substrate proteins. Some members of this group do not possess the occluding loop. TIN-Ag is an extracellular matrix basement protein which was originally identified as a target Ag involved in anti-tubular basement membrane
Probab=100.00 E-value=3.4e-37 Score=270.42 Aligned_cols=150 Identities=29% Similarity=0.533 Sum_probs=125.5
Q ss_pred CCceecCCC--CCc--cccccCCCCcchHHHHHHHHHHHHHHHhcC--CCccCChhHHHhhCCCCCCCCCCCCCCCCCCC
Q 022276 138 PTDFDWRDH--GAV--TGVKDQGACGSCWSFSATGALEGAHFLSTG--ELVSLSEQQLVDCDHECDPEESGSCDSGCNGG 211 (300)
Q Consensus 138 P~s~DwR~~--g~v--~pvknQg~CgsCwAfa~~~~~e~~~~i~~~--~~~~lS~Q~lidC~~~~~~~~~~~~~~gC~GG 211 (300)
|++||||++ +++ +||+|||.||||||||++++||++++|+++ +.+.||+|||+||+.. .+.||+||
T Consensus 1 p~~~DwR~~~~~~~~v~~v~dQg~CGsCwAfa~~~~le~~~~i~~~~~~~~~LS~Q~lidC~~~--------~~~gC~GG 72 (236)
T cd02620 1 PESFDAREKWPNCISIGEIRDQGNCGSCWAFSAVEAFSDRLCIQSNGKENVLLSAQDLLSCCSG--------CGDGCNGG 72 (236)
T ss_pred CCcccchhhCCCCCCccccCCcccchhHHHHHHHHHHhhHHHHhcCCCCccccCHHHHHhhcCC--------CCCCCCCC
Confidence 899999997 554 599999999999999999999999999988 7799999999999863 37899999
Q ss_pred ChHHHHHHHHHhCCcCCCcccccCCCCCC------------------CCCCCCC----CceEEEceeEEcChhHHHHHHH
Q 022276 212 LMNSAFEYILKAGGVEREKDYPYTGTDGG------------------SCKFDKS----KIAAAVSNFSVISSDEDQMAAN 269 (300)
Q Consensus 212 ~~~~a~~y~~~~~G~~~e~~yPY~~~~~~------------------~C~~~~~----~~~~~i~~~~~v~~~~~~i~~a 269 (300)
++..||+|++++ |+++|++|||.+.+ . .|..... ....++..+..+..++++||++
T Consensus 73 ~~~~a~~~i~~~-G~~~e~~yPY~~~~-~~~~~~~~~~~~~~~~~~~~C~~~~~~~~~~~~~~~~~~~~~~~~~~~ik~~ 150 (236)
T cd02620 73 YPDAAWKYLTTT-GVVTGGCQPYTIPP-CGHHPEGPPPCCGTPYCTPKCQDGCEKTYEEDKHKGKSAYSVPSDETDIMKE 150 (236)
T ss_pred CHHHHHHHHHhc-CCCcCCEecCcCCC-CccCCCCCCCCCCCCCCCCCCCcCCccccceeeeeecceeeeCCHHHHHHHH
Confidence 999999999887 89999999998765 2 2432221 1124455565665578999999
Q ss_pred HHhcCCeEEEEec-CCCCCccCeeEecCC
Q 022276 270 LVKHGPLAGNVAS-IELPHISFSFLFTVS 297 (300)
Q Consensus 270 l~~~GPv~v~i~a-~~f~~Y~~Giy~~~~ 297 (300)
|+++|||+|+|++ +.|+.|++|||..++
T Consensus 151 l~~~GPv~v~i~~~~~f~~Y~~Giy~~~~ 179 (236)
T cd02620 151 IMTNGPVQAAFTVYEDFLYYKSGVYQHTS 179 (236)
T ss_pred HHHCCCeEEEEEechhhhhcCCcEEeecC
Confidence 9999999999999 469999999998653
No 10
>PTZ00364 dipeptidyl-peptidase I precursor; Provisional
Probab=100.00 E-value=2.2e-36 Score=288.44 Aligned_cols=150 Identities=21% Similarity=0.424 Sum_probs=128.3
Q ss_pred CCCCCceecCCCC---CccccccCCC---CcchHHHHHHHHHHHHHHHhcC------CCccCChhHHHhhCCCCCCCCCC
Q 022276 135 NDLPTDFDWRDHG---AVTGVKDQGA---CGSCWSFSATGALEGAHFLSTG------ELVSLSEQQLVDCDHECDPEESG 202 (300)
Q Consensus 135 ~~lP~s~DwR~~g---~v~pvknQg~---CgsCwAfa~~~~~e~~~~i~~~------~~~~lS~Q~lidC~~~~~~~~~~ 202 (300)
.++|++||||++| +|+||||||. ||||||||+++++|++++|+++ +.+.||+|||+||+..
T Consensus 203 ~~LP~sfDWR~~gg~~~VtpVrdQg~~~~CGSCWAFAav~alEsr~~I~tn~~~~~g~~~~LS~QqLVDCs~~------- 275 (548)
T PTZ00364 203 DPPPAAWSWGDVGGASFLPAAPPASPGRGCNSSYVEAALAAMMARVMVASNRTDPLGQQTFLSARHVLDCSQY------- 275 (548)
T ss_pred cCCCCccccCcCCCCccCCCCcCCCCCCCCcCHHHHHHHHHHHHHHHHHhCCCcccCcccCcCHHHHhcccCC-------
Confidence 5799999999997 7999999999 9999999999999999999984 4688999999999864
Q ss_pred CCCCCCCCCChHHHHHHHHHhCCcCCCccc--ccCCCCCC---CCCCCCCCceEEEce------eEEcChhHHHHHHHHH
Q 022276 203 SCDSGCNGGLMNSAFEYILKAGGVEREKDY--PYTGTDGG---SCKFDKSKIAAAVSN------FSVISSDEDQMAANLV 271 (300)
Q Consensus 203 ~~~~gC~GG~~~~a~~y~~~~~G~~~e~~y--PY~~~~~~---~C~~~~~~~~~~i~~------~~~v~~~~~~i~~al~ 271 (300)
++||+||++..|++|++++ |+++|++| ||++.+ + .|+.......+.+++ |..+..++++|+++|+
T Consensus 276 --n~GCdGG~p~~A~~yi~~~-GI~tE~dY~~PY~~~d-g~~~~Ck~~~~~~~y~~~~~~~I~gyy~~~~~e~~I~~eI~ 351 (548)
T PTZ00364 276 --GQGCAGGFPEEVGKFAETF-GILTTDSYYIPYDSGD-GVERACKTRRPSRRYYFTNYGPLGGYYGAVTDPDEIIWEIY 351 (548)
T ss_pred --CCCCCCCcHHHHHHHHHhC-CcccccccCCCCCCCC-CCCCCCCCCcccceeeeeeeEEecceeecCCcHHHHHHHHH
Confidence 7899999999999999876 89999999 998765 4 588655444444444 4333447888999999
Q ss_pred hcCCeEEEEecC-CCCCccCeeEec
Q 022276 272 KHGPLAGNVASI-ELPHISFSFLFT 295 (300)
Q Consensus 272 ~~GPv~v~i~a~-~f~~Y~~Giy~~ 295 (300)
++|||+|+|++. +|++|++|||..
T Consensus 352 ~~GPVsVaIda~~df~~YksGiy~g 376 (548)
T PTZ00364 352 RHGPVPASVYANSDWYNCDENSTED 376 (548)
T ss_pred HcCCeEEEEEechHHHhcCCCCccC
Confidence 999999999996 599999999873
No 11
>PTZ00049 cathepsin C-like protein; Provisional
Probab=100.00 E-value=3.1e-36 Score=290.41 Aligned_cols=152 Identities=23% Similarity=0.447 Sum_probs=127.9
Q ss_pred CCCCCCceecCCC----CCccccccCCCCcchHHHHHHHHHHHHHHHhcCCC----------ccCChhHHHhhCCCCCCC
Q 022276 134 TNDLPTDFDWRDH----GAVTGVKDQGACGSCWSFSATGALEGAHFLSTGEL----------VSLSEQQLVDCDHECDPE 199 (300)
Q Consensus 134 ~~~lP~s~DwR~~----g~v~pvknQg~CgsCwAfa~~~~~e~~~~i~~~~~----------~~lS~Q~lidC~~~~~~~ 199 (300)
..+||++||||+. +.++||+|||.||||||||++++||++++|++++. ..||+|+|+||+..
T Consensus 378 ~~~LP~sfDWRd~~~~~~~vtpVkdQG~CGSCWAFAat~alEsR~~Ia~~~~l~~~~~~~~~~~LS~QqLLDCs~~---- 453 (693)
T PTZ00049 378 IDELPKNFTWGDPFNNNTREYDVTNQLLCGSCYIASQMYAFKRRIEIALTKNLDKKYLNNFDDLLSIQTVLSCSFY---- 453 (693)
T ss_pred cccCCCCEecCcCCCCCCcccCCCCCccCcHHHHHHHHHHHHHHHHHHhccccccccccccccCcCHHHhcccCCC----
Confidence 3589999999985 67999999999999999999999999999986431 27999999999864
Q ss_pred CCCCCCCCCCCCChHHHHHHHHHhCCcCCCcccccCCCCCCCCCCCCCC-------------------------------
Q 022276 200 ESGSCDSGCNGGLMNSAFEYILKAGGVEREKDYPYTGTDGGSCKFDKSK------------------------------- 248 (300)
Q Consensus 200 ~~~~~~~gC~GG~~~~a~~y~~~~~G~~~e~~yPY~~~~~~~C~~~~~~------------------------------- 248 (300)
++||+||++..|++|+++. ||++|.+|||++.. +.|+.....
T Consensus 454 -----nqGC~GG~~~~A~kya~~~-GI~tEscYPY~a~~-g~C~~~~~~~~~~~~g~~~~~~~~~~~~~~~~~~~~~~~~ 526 (693)
T PTZ00049 454 -----DQGCNGGFPYLVSKMAKLQ-GIPLDKVFPYTATE-QTCPYQVDQSANSMNGSANLRQINAVFFSSETQSDMHADF 526 (693)
T ss_pred -----CCCcCCCcHHHHHHHHHHC-CCCcCCccCCcCCC-CCCCCCCCCccccccccccccccccccccccccccccccc
Confidence 7899999999999999887 89999999999887 788653211
Q ss_pred --------ceEEEceeEEcC--------hhHHHHHHHHHhcCCeEEEEecC-CCCCccCeeEecC
Q 022276 249 --------IAAAVSNFSVIS--------SDEDQMAANLVKHGPLAGNVASI-ELPHISFSFLFTV 296 (300)
Q Consensus 249 --------~~~~i~~~~~v~--------~~~~~i~~al~~~GPv~v~i~a~-~f~~Y~~Giy~~~ 296 (300)
..+.+++|..++ .++++||++|+++|||+|+|+|. .|++|++|||..+
T Consensus 527 ~~~~~~~~~r~y~k~y~yI~g~y~~~~~~~E~~Im~eI~~~GPVsVsIda~~dF~~YksGVY~~~ 591 (693)
T PTZ00049 527 EAPISSEPARWYAKDYNYIGGCYGCNQCNGEKIMMNEIYRNGPIVASFEASPDFYDYADGVYYVE 591 (693)
T ss_pred cccccccccceeeeeeEEecccccccCCCCHHHHHHHHHhcCCEEEEEEechhhhcCCCccccCc
Confidence 123345565553 26889999999999999999996 5999999999853
No 12
>PF00112 Peptidase_C1: Papain family cysteine protease This is family C1 in the peptidase classification. ; InterPro: IPR000668 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This group of proteins belong to the peptidase family C1, sub-family C1A (papain family, clan CA). It includes proteins classed as non-peptidase homologs. These are have either been shown experimentally to lack peptidase activity or lack one or more of the active site residues. The papain family has a wide variety of activities, including broad-range (papain) and narrow-range endo-peptidases, aminopeptidases, dipeptidyl peptidases and enzymes with both exo- and endo-peptidase activity []. Members of the papain family are widespread, found in baculovirus [], eubacteria, yeast, and practically all protozoa, plants and mammals []. The proteins are typically lysosomal or secreted, and proteolytic cleavage of the propeptide is required for enzyme activation, although bleomycin hydrolase is cytosolic in fungi and mammals []. Papain-like cysteine proteinases are essentially synthesised as inactive proenzymes (zymogens) with N-terminal propeptide regions. The activation process of these enzymes includes the removal of propeptide regions. The propeptide regions serve a variety of functions in vivo and in vitro. The pro-region is required for the proper folding of the newly synthesised enzyme, the inactivation of the peptidase domain and stabilisation of the enzyme against denaturing at neutral to alkaline pH conditions. Amino acid residues within the pro-region mediate their membrane association, and play a role in the transport of the proenzyme to lysosomes. Among the most notable features of propeptides is their ability to inhibit the activity of their cognate enzymes and that certain propeptides exhibit high selectivity for inhibition of the peptidases from which they originate []. The catalytic residues of papain are Cys-25 and His-159, other important residues being Gln-19, which helps form the 'oxyanion hole', and Asn-175, which orientates the imidazole ring of His-159. ; GO: 0008234 cysteine-type peptidase activity, 0006508 proteolysis; PDB: 3MOR_B 3HHI_B 1S4V_A 3F75_A 1MEG_A 1PCI_C 1PPO_A 3HD3_B 1F29_A 1EWL_A ....
Probab=100.00 E-value=1.7e-34 Score=249.81 Aligned_cols=154 Identities=36% Similarity=0.736 Sum_probs=130.7
Q ss_pred CCCceecCCC-CCccccccCCCCcchHHHHHHHHHHHHHHHhc-CCCccCChhHHHhhCCCCCCCCCCCCCCCCCCCChH
Q 022276 137 LPTDFDWRDH-GAVTGVKDQGACGSCWSFSATGALEGAHFLST-GELVSLSEQQLVDCDHECDPEESGSCDSGCNGGLMN 214 (300)
Q Consensus 137 lP~s~DwR~~-g~v~pvknQg~CgsCwAfa~~~~~e~~~~i~~-~~~~~lS~Q~lidC~~~~~~~~~~~~~~gC~GG~~~ 214 (300)
||++||||+. +.++||+|||.||+|||||+++++|++++++. +..+.||+|+|++|... .+.+|+||++.
T Consensus 1 lP~~~D~r~~~~~~~~v~dQg~~gsCwafa~~~~~e~~~~~~~~~~~~~lS~q~l~~~~~~--------~~~~c~gg~~~ 72 (219)
T PF00112_consen 1 LPKSFDWRDKGGRITPVRDQGSCGSCWAFAAAAALESRLAIQNNGKNVDLSEQYLIDCSNK--------YNKGCDGGSPF 72 (219)
T ss_dssp STSSEEGGGTTTCSG---BTTSSBTHHHHHHHHHHHHHHHHHHTSSCEEB-HHHHHHHSTG--------TSSTTBBBEHH
T ss_pred CCCCEecccCCCCcCccccCCcccccccchhccceeccccccccccccccccccccccccc--------cccccccCccc
Confidence 7999999998 58999999999999999999999999999998 78899999999999972 26799999999
Q ss_pred HHHHHHHHhCCcCCCcccccCCCCCCCCCCCCCCc-eEEEceeEEcCh-hHHHHHHHHHhcCCeEEEEecCC--CCCccC
Q 022276 215 SAFEYILKAGGVEREKDYPYTGTDGGSCKFDKSKI-AAAVSNFSVISS-DEDQMAANLVKHGPLAGNVASIE--LPHISF 290 (300)
Q Consensus 215 ~a~~y~~~~~G~~~e~~yPY~~~~~~~C~~~~~~~-~~~i~~~~~v~~-~~~~i~~al~~~GPv~v~i~a~~--f~~Y~~ 290 (300)
.|++|++++.|+++|++|||.+.....|....... ..++.+|..+.. +.++||++|+++|||+++|.+.. |+.|++
T Consensus 73 ~a~~~~~~~~Gi~~e~~~pY~~~~~~~c~~~~~~~~~~~i~~~~~~~~~~~~~ik~~L~~~gpV~~~~~~~~~~f~~~~~ 152 (219)
T PF00112_consen 73 DALKYIKNNNGIVTEEDYPYNGNENPTCKSKKSNSYYVKIKGYGKVKDNDIEDIKKALMKYGPVVASIDVSSEDFQNYKS 152 (219)
T ss_dssp HHHHHHHHHTSBEBTTTS--SSSSSCSSCHSGGGEEEBEESEEEEEESTCHHHHHHHHHHHSSEEEEEEEESHHHHTEES
T ss_pred ccceeecccCcccccccccccccccccccccccccccccccccccccccchhHHHHHHhhCceeeeeeeccccccccccc
Confidence 99999998459999999999986635788764443 478889988876 59999999999999999999855 999999
Q ss_pred eeEecCCC
Q 022276 291 SFLFTVSS 298 (300)
Q Consensus 291 Giy~~~~~ 298 (300)
|||.++.+
T Consensus 153 gi~~~~~~ 160 (219)
T PF00112_consen 153 GIYDPPDC 160 (219)
T ss_dssp SEECSTSS
T ss_pred eeeecccc
Confidence 99999854
No 13
>smart00645 Pept_C1 Papain family cysteine protease.
Probab=99.97 E-value=9.3e-32 Score=225.32 Aligned_cols=114 Identities=56% Similarity=0.972 Sum_probs=104.0
Q ss_pred CCCceecCCCCCccccccCCCCcchHHHHHHHHHHHHHHHhcCCCccCChhHHHhhCCCCCCCCCCCCCCCCCCCChHHH
Q 022276 137 LPTDFDWRDHGAVTGVKDQGACGSCWSFSATGALEGAHFLSTGELVSLSEQQLVDCDHECDPEESGSCDSGCNGGLMNSA 216 (300)
Q Consensus 137 lP~s~DwR~~g~v~pvknQg~CgsCwAfa~~~~~e~~~~i~~~~~~~lS~Q~lidC~~~~~~~~~~~~~~gC~GG~~~~a 216 (300)
||++||||+.++++||||||.||+|||||+++++|+++++++++.+.||+|+|+||... .++||+||.+..|
T Consensus 1 lP~~~D~R~~~~~~~v~dQg~CGsCwAfa~~~~ie~~~~i~~~~~~~lS~q~l~~C~~~--------~~~gC~GG~~~~a 72 (174)
T smart00645 1 LPESFDWRKKGAVTPVKDQGQCGSCWAFSATGALEGRYCIKTGKLVSLSEQQLVDCSTG--------GNNGCNGGLPDNA 72 (174)
T ss_pred CCCcCcccccCCCCccccCcccchHHHHHHHHHHHHHHHHhcCCccccCHHHHhhhcCC--------CCCCCCCcCHHHH
Confidence 69999999999999999999999999999999999999999999999999999999873 2569999999999
Q ss_pred HHHHHHhCCcCCCcccccCCCCCCCCCCCCCCceEEEceeEEcChhHHHHHHHHHhcCCeEEEEecCCCCCccCeeEecC
Q 022276 217 FEYILKAGGVEREKDYPYTGTDGGSCKFDKSKIAAAVSNFSVISSDEDQMAANLVKHGPLAGNVASIELPHISFSFLFTV 296 (300)
Q Consensus 217 ~~y~~~~~G~~~e~~yPY~~~~~~~C~~~~~~~~~~i~~~~~v~~~~~~i~~al~~~GPv~v~i~a~~f~~Y~~Giy~~~ 296 (300)
++|+.+++|+++|++|||.+ ++.+.+.+|++|++|||..+
T Consensus 73 ~~~~~~~~Gi~~e~~~PY~~----------------------------------------~~~~~~~~f~~Y~~Gi~~~~ 112 (174)
T smart00645 73 FEYIKKNGGLETESCYPYTG----------------------------------------SVAIDASDFQFYKSGIYDHP 112 (174)
T ss_pred HHHHHHcCCcccccccCccc----------------------------------------EEEEEcccccCCcCeEECCC
Confidence 99998876899999999954 66777778999999999986
Q ss_pred CC
Q 022276 297 SS 298 (300)
Q Consensus 297 ~~ 298 (300)
++
T Consensus 113 ~~ 114 (174)
T smart00645 113 GC 114 (174)
T ss_pred CC
Confidence 43
No 14
>cd02619 Peptidase_C1 C1 Peptidase family (MEROPS database nomenclature), also referred to as the papain family; composed of two subfamilies of cysteine peptidases (CPs), C1A (papain) and C1B (bleomycin hydrolase). Papain-like enzymes are mostly endopeptidases with some exceptions like cathepsins B, C, H and X, which are exopeptidases. Papain-like CPs have different functions in various organisms. Plant CPs are used to mobilize storage proteins in seeds while mammalian CPs are primarily lysosomal enzymes responsible for protein degradation in the lysosome. Papain-like CPs are synthesized as inactive proenzymes with N-terminal propeptide regions, which are removed upon activation. Bleomycin hydrolase (BH) is a CP that detoxifies bleomycin by hydrolysis of an amide group. It acts as a carboxypeptidase on its C-terminus to convert itself into an aminopeptidase and peptide ligase. BH is found in all tissues in mammals as well as in many other eukaryotes. It forms a hexameric ring barrel str
Probab=99.97 E-value=3e-30 Score=223.75 Aligned_cols=148 Identities=26% Similarity=0.453 Sum_probs=125.6
Q ss_pred ceecCCCCCccccccCCCCcchHHHHHHHHHHHHHHHhcC--CCccCChhHHHhhCCCCCCCCCCCCCCCCCCCChHHHH
Q 022276 140 DFDWRDHGAVTGVKDQGACGSCWSFSATGALEGAHFLSTG--ELVSLSEQQLVDCDHECDPEESGSCDSGCNGGLMNSAF 217 (300)
Q Consensus 140 s~DwR~~g~v~pvknQg~CgsCwAfa~~~~~e~~~~i~~~--~~~~lS~Q~lidC~~~~~~~~~~~~~~gC~GG~~~~a~ 217 (300)
.+|||+.+ ++||||||.||+|||||+++++|++++++++ +.+.||+|+|++|..... .....||.||.+..++
T Consensus 1 ~~d~r~~~-~~~v~dQg~~gsCwafa~~~~les~~~~~~~~~~~~~lS~q~l~~c~~~~~----~~~~~~c~gG~~~~~~ 75 (223)
T cd02619 1 SVDLRPLR-LTPVKNQGSRGSCWAFASAYALESAYRIKGGEDEYVDLSPQYLYICANDEC----LGINGSCDGGGPLSAL 75 (223)
T ss_pred CCcchhcC-CCCcccCCCCcCcHHHHHHHHHHHHHHHhcCCcccccCCHHHHHHhccccc----cccCCCCCCCcHHHHH
Confidence 48999998 9999999999999999999999999999988 789999999999987510 0013799999999999
Q ss_pred H-HHHHhCCcCCCcccccCCCCCCCCCCC----CCCceEEEceeEEcCh-hHHHHHHHHHhcCCeEEEEecCC-CCCccC
Q 022276 218 E-YILKAGGVEREKDYPYTGTDGGSCKFD----KSKIAAAVSNFSVISS-DEDQMAANLVKHGPLAGNVASIE-LPHISF 290 (300)
Q Consensus 218 ~-y~~~~~G~~~e~~yPY~~~~~~~C~~~----~~~~~~~i~~~~~v~~-~~~~i~~al~~~GPv~v~i~a~~-f~~Y~~ 290 (300)
. ++.++ |+++|.+|||.... ..|... ......++..|..+.. +.++||++|+++|||+|+|.+.. |..|++
T Consensus 76 ~~~~~~~-Gi~~e~~~Py~~~~-~~~~~~~~~~~~~~~~~~~~y~~~~~~~~~~ik~aL~~~gPv~~~~~~~~~~~~~~~ 153 (223)
T cd02619 76 LKLVALK-GIPPEEDYPYGAES-DGEEPKSEAALNAAKVKLKDYRRVLKNNIEDIKEALAKGGPVVAGFDVYSGFDRLKE 153 (223)
T ss_pred HHHHHHc-CCCccccCCCCCCC-CCCCCCCccchhhcceeecceeEeCchhHHHHHHHHHHCCCEEEEEEcccchhcccC
Confidence 8 66655 99999999999877 566432 3344678889988876 58999999999999999999864 999999
Q ss_pred eeEe
Q 022276 291 SFLF 294 (300)
Q Consensus 291 Giy~ 294 (300)
|+|.
T Consensus 154 ~~~~ 157 (223)
T cd02619 154 GIIY 157 (223)
T ss_pred cccc
Confidence 9974
No 15
>KOG1544 consensus Predicted cysteine proteinase TIN-ag [General function prediction only]
Probab=99.96 E-value=9.7e-31 Score=227.99 Aligned_cols=207 Identities=21% Similarity=0.321 Sum_probs=158.7
Q ss_pred HHHHhcCCCCCeeee-eccCCCCChhhHHhhhcCCCccCCCCCCCC--CCCCCCCCCCCCceecCCC--CCccccccCCC
Q 022276 83 RAKRRQLLDPTAVHG-VTKFSDLTPSEFRRQFLGLNRRLRLPADAQ--KAPILPTNDLPTDFDWRDH--GAVTGVKDQGA 157 (300)
Q Consensus 83 ~I~~~N~~~~s~~~g-iN~FsDlt~~Ef~~~~~g~~~~~~~~~~~~--~~~~~~~~~lP~s~DwR~~--g~v~pvknQg~ 157 (300)
.|++.|+.+.+++.+ ..+|..+|.+.-.+..+|...+...-.... .+.+.+..+||+.|+.|++ +++.|+-|||+
T Consensus 152 ~iE~in~G~YgW~A~NYSaFWGmtL~DGiKyRLGTL~Ps~sv~nMNEi~~~l~p~~~LPE~F~As~KWp~liH~plDQgn 231 (470)
T KOG1544|consen 152 MIEAINQGNYGWQAGNYSAFWGMTLDDGIKYRLGTLRPSSSVMNMNEIYTVLNPGEVLPEAFEASEKWPNLIHEPLDQGN 231 (470)
T ss_pred HHHHHhcCCccccccchhhhhcccccccceeeecccCchhhhhhHHhHhhccCcccccchhhhhhhcCCccccCccccCC
Confidence 344444444444443 247999999887777777654422211111 1223345799999999998 89999999999
Q ss_pred CcchHHHHHHHHHHHHHHHhcCC--CccCChhHHHhhCCCCCCCCCCCCCCCCCCCChHHHHHHHHHhCCcCCCcccccC
Q 022276 158 CGSCWSFSATGALEGAHFLSTGE--LVSLSEQQLVDCDHECDPEESGSCDSGCNGGLMNSAFEYILKAGGVEREKDYPYT 235 (300)
Q Consensus 158 CgsCwAfa~~~~~e~~~~i~~~~--~~~lS~Q~lidC~~~~~~~~~~~~~~gC~GG~~~~a~~y~~~~~G~~~e~~yPY~ 235 (300)
|++.|||+|+++..++++|++.. ...||+|+|++|... ...||.||..+.||=|+.+. |++...||||.
T Consensus 232 Ca~SWafSTaavasDRiAI~S~GR~t~~LSpQnLlSC~~h--------~q~GC~gG~lDRAWWYlRKr-GvVsdhCYP~~ 302 (470)
T KOG1544|consen 232 CAGSWAFSTAAVASDRVAIHSLGRMTPVLSPQNLLSCDTH--------QQQGCRGGRLDRAWWYLRKR-GVVSDHCYPFS 302 (470)
T ss_pred cccceeeeeehhccceeEEeeccccccccChHHhcchhhh--------hhccCccCcccchheeeecc-ccccccccccc
Confidence 99999999999999999998754 468999999999875 47999999999999999887 89999999997
Q ss_pred CCC---CCCC------------------CCC--CCCceEEEceeEEcChhHHHHHHHHHhcCCeEEEEec-CCCCCccCe
Q 022276 236 GTD---GGSC------------------KFD--KSKIAAAVSNFSVISSDEDQMAANLVKHGPLAGNVAS-IELPHISFS 291 (300)
Q Consensus 236 ~~~---~~~C------------------~~~--~~~~~~~i~~~~~v~~~~~~i~~al~~~GPv~v~i~a-~~f~~Y~~G 291 (300)
+.. ++.| ... ..+.+++++--..|..+|++|+++|+.+|||.+.|.+ ++|.+|++|
T Consensus 303 ~dQ~~~~~~C~m~sR~~grgkRqat~~CPn~~~~Sn~iyq~tPPYrVSSnE~eImkElM~NGPVQA~m~VHEDFF~YkgG 382 (470)
T KOG1544|consen 303 GDQAGPAPPCMMHSRAMGRGKRQATAHCPNSYVNSNDIYQVTPPYRVSSNEKEIMKELMENGPVQALMEVHEDFFLYKGG 382 (470)
T ss_pred CCCCCCCCCceeeccccCcccccccCcCCCcccccCceeeecCCeeccCCHHHHHHHHHhCCChhhhhhhhhhhhhhccc
Confidence 532 1334 322 1234566666556777899999999999999999988 669999999
Q ss_pred eEecCCC
Q 022276 292 FLFTVSS 298 (300)
Q Consensus 292 iy~~~~~ 298 (300)
||..+..
T Consensus 383 iY~H~~~ 389 (470)
T KOG1544|consen 383 IYSHTPV 389 (470)
T ss_pred eeecccc
Confidence 9998765
No 16
>PTZ00462 Serine-repeat antigen protein; Provisional
Probab=99.95 E-value=2.8e-28 Score=241.82 Aligned_cols=142 Identities=14% Similarity=0.249 Sum_probs=113.5
Q ss_pred ccccccCCCCcchHHHHHHHHHHHHHHHhcCCCccCChhHHHhhCCCCCCCCCCCCCCCCCCCChH-HHHHHHHHhCCcC
Q 022276 149 VTGVKDQGACGSCWSFSATGALEGAHFLSTGELVSLSEQQLVDCDHECDPEESGSCDSGCNGGLMN-SAFEYILKAGGVE 227 (300)
Q Consensus 149 v~pvknQg~CgsCwAfa~~~~~e~~~~i~~~~~~~lS~Q~lidC~~~~~~~~~~~~~~gC~GG~~~-~a~~y~~~~~G~~ 227 (300)
..||||||.||+|||||+++++|++++|++++.+.||+|+|+||+.. ..+.||.||... .++.|+.++||++
T Consensus 544 ~i~VKDQG~CGSCWAFASaaaLES~~cIkgg~~v~LSeQqLVDCs~~-------~gn~GC~GG~~~~efl~yI~e~GgLp 616 (1004)
T PTZ00462 544 KIQIEDQGNCAISWIFASKYHLETIKCMKGYEPHAISALYIANCSKG-------EHKDRCDEGSNPLEFLQIIEDNGFLP 616 (1004)
T ss_pred CCCcccCCcchHHHHHHHHHHHHHHHHHhcCCCcccCHHHHHhcccc-------cCCCCCCCCCcHHHHHHHHHHcCCCc
Confidence 57899999999999999999999999999999999999999999864 236899999744 5568988887899
Q ss_pred CCcccccCC--CCCCCCCCCCC------------------CceEEEceeEEcChh---------HHHHHHHHHhcCCeEE
Q 022276 228 REKDYPYTG--TDGGSCKFDKS------------------KIAAAVSNFSVISSD---------EDQMAANLVKHGPLAG 278 (300)
Q Consensus 228 ~e~~yPY~~--~~~~~C~~~~~------------------~~~~~i~~~~~v~~~---------~~~i~~al~~~GPv~v 278 (300)
+|++|||.+ .. +.|+.... ...+.+.+|..+... +++|+++|+.+|||+|
T Consensus 617 tESdYPYt~k~~~-g~Cp~~~~~w~n~~~~~kll~~~~~~~~~i~~kgY~~~~s~~~~~n~d~~i~~IK~eI~~kGPVaV 695 (1004)
T PTZ00462 617 ADSNYLYNYTKVG-EDCPDEEDHWMNLLDHGKILNHNKKEPNSLDGKAYRAYESEHFHDKMDAFIKIIKDEIMNKGSVIA 695 (1004)
T ss_pred ccccCCCccCCCC-CCCCCCcccccccccccccccccccccceeeccceEEecccccccchhhHHHHHHHHHHhcCCEEE
Confidence 999999986 34 67974321 112344566655421 4689999999999999
Q ss_pred EEecCCCCCc-cCeeEecCCC
Q 022276 279 NVASIELPHI-SFSFLFTVSS 298 (300)
Q Consensus 279 ~i~a~~f~~Y-~~Giy~~~~~ 298 (300)
+|++..|++| .+|||....|
T Consensus 696 ~IdAsdf~~Y~~sGIyv~~~C 716 (1004)
T PTZ00462 696 YIKAENVLGYEFNGKKVQNLC 716 (1004)
T ss_pred EEEeehHHhhhcCCccccCCC
Confidence 9999888888 5898776533
No 17
>PF08246 Inhibitor_I29: Cathepsin propeptide inhibitor domain (I29); InterPro: IPR013201 Peptide proteinase inhibitors can be found as single domain proteins or as single or multiple domains within proteins; these are referred to as either simple or compound inhibitors, respectively. In many cases they are synthesised as part of a larger precursor protein, either as a prepropeptide or as an N-terminal domain associated with an inactive peptidase or zymogen. This domain prevents access of the substrate to the active site. Removal of the N-terminal inhibitor domain either by interaction with a second peptidase or by autocatalytic cleavage activates the zymogen. Other inhibitors interact direct with proteinases using a simple noncovalent lock and key mechanism; while yet others use a conformational change-based trapping mechanism that depends on their structural and thermodynamic properties. This entry represents a peptidase inhibitor domain, which belongs to MEROPS peptidase inhibitor family I29. The domain is also found at the N terminus of a variety of peptidase precursors that belong to MEROPS peptidase subfamily C1A; these include cathepsin L, papain, and procaricain (P10056 from SWISSPROT) []. It forms an alpha-helical domain that runs through the substrate-binding site, preventing access. Removal of this region by proteolytic cleavage results in activation of the enzyme. This domain is also found, in one or more copies, in a variety of cysteine peptidase inhibitors such as salarin [].; PDB: 3QT4_A 3QJ3_A 2C0Y_A 2L95_A 1CJL_A 1CS8_A 7PCK_A 1BY8_A 1PCI_A 2O6X_A ....
Probab=99.66 E-value=2.1e-16 Score=107.99 Aligned_cols=57 Identities=44% Similarity=0.736 Sum_probs=50.9
Q ss_pred HHHHHHHhCCccCCHHHHHHHHHHHHHHHHHHHHhc-CCCCCeeeeeccCCCCChhhH
Q 022276 53 FSLFKSKFSKTYATQEEHDYRFRVFKANLRRAKRRQ-LLDPTAVHGVTKFSDLTPSEF 109 (300)
Q Consensus 53 F~~f~~~~~k~Y~s~~E~~~r~~~F~~n~~~I~~~N-~~~~s~~~giN~FsDlt~~Ef 109 (300)
|+.|+++|+|.|.+.+|...|+.+|++|++.|.+|| ..+.+|++|+|+|+|||++||
T Consensus 1 F~~~~~~~~k~Y~~~~e~~~R~~~F~~N~~~I~~~N~~~~~~~~~~~N~fsD~t~eEf 58 (58)
T PF08246_consen 1 FEQFKKKYGKSYKSAEEEARRFAIFKENLRRIEEHNANGNNTYKLGLNQFSDMTPEEF 58 (58)
T ss_dssp HHHHHHHCT---SSHHHHHHHHHHHHHHHHHHHHHHHTTSSSEEE-SSTTTTSSHHHH
T ss_pred CHHHHHHcCCCCCCHHHHHHHHHHHHHHHHHHHHHhcCCCCCeEEeCccccCcChhhC
Confidence 899999999999999999999999999999999999 667899999999999999997
No 18
>smart00848 Inhibitor_I29 Cathepsin propeptide inhibitor domain (I29). This domain is found at the N-terminus of some C1 peptidases such as Cathepsin L where it acts as a propeptide. There are also a number of proteins that are composed solely of multiple copies of this domain such as the peptidase inhibitor salarin. This family is classified as I29 by MEROPS. Peptide proteinase inhibitors can be found as single domain proteins or as single or multiple domains within proteins; these are referred to as either simple or compound inhibitors, respectively. In many cases they are synthesised as part of a larger precursor protein, either as a prepropeptide or as an N-terminal domain associated with an inactive peptidase or zymogen. This domain prevents access of the substrate to the active site. Removal of the N-terminal inhibitor domain either by interaction with a second peptidase or by autocatalytic cleavage activates the zymogen. Other inhibitors interact direct with proteinases using a s
Probab=99.49 E-value=4.6e-14 Score=95.97 Aligned_cols=56 Identities=34% Similarity=0.605 Sum_probs=52.5
Q ss_pred HHHHHHHhCCccCCHHHHHHHHHHHHHHHHHHHHhcCCC-CCeeeeeccCCCCChhh
Q 022276 53 FSLFKSKFSKTYATQEEHDYRFRVFKANLRRAKRRQLLD-PTAVHGVTKFSDLTPSE 108 (300)
Q Consensus 53 F~~f~~~~~k~Y~s~~E~~~r~~~F~~n~~~I~~~N~~~-~s~~~giN~FsDlt~~E 108 (300)
|..|+.+|+|.|.+.+|...|+.+|++|++.|..||..+ .+|++|+|+|+|||++|
T Consensus 1 f~~~~~~~~k~y~~~~e~~~r~~~f~~n~~~i~~~N~~~~~~~~~~~N~fsDlt~eE 57 (57)
T smart00848 1 FEQWKKKYGKSYSSEEEELRRFEIFKENLKFIEEHNKKNDHSYTLGLNQFADLTNEE 57 (57)
T ss_pred ChHHHHHhCCCCCCHHHHHHHHHHHHHHHHHHHHHHhcCCCCeEecCcccccCCCCC
Confidence 688999999999999999999999999999999999764 78999999999999886
No 19
>COG4870 Cysteine protease [Posttranslational modification, protein turnover, chaperones]
Probab=99.36 E-value=2.2e-13 Score=122.39 Aligned_cols=151 Identities=26% Similarity=0.408 Sum_probs=102.8
Q ss_pred CCCCceecCCCCCccccccCCCCcchHHHHHHHHHHHHHHHhcCCCccCChhHHHhhCCCCCCCCCCCCCCCC-----CC
Q 022276 136 DLPTDFDWRDHGAVTGVKDQGACGSCWSFSATGALEGAHFLSTGELVSLSEQQLVDCDHECDPEESGSCDSGC-----NG 210 (300)
Q Consensus 136 ~lP~s~DwR~~g~v~pvknQg~CgsCwAfa~~~~~e~~~~i~~~~~~~lS~Q~lidC~~~~~~~~~~~~~~gC-----~G 210 (300)
.+|+.||||+.|.|+|||+||.||+||||++++++|+.+.-.. ...+|+-.+..-...+ +..+| +|
T Consensus 98 s~~~~fd~r~~g~vs~v~dQg~~Gscwaf~t~~sles~l~~~~--~w~~s~~nm~~ll~~~-------ye~~fd~~~~d~ 168 (372)
T COG4870 98 SLPSYFDRRDEGKVSPVKDQGSGGSCWAFATTRSLESYLNPES--AWDFSENNMKNLLGVP-------YEKGFDYTSNDG 168 (372)
T ss_pred cchhheeeeccCCcccccccCcccceEeeeehhhhhheecccc--cccccccchhhhcCCC-------ccccCCCccccC
Confidence 5899999999999999999999999999999999999874443 3445554443322221 12222 37
Q ss_pred CChHHHHHHHHHhCCcCCCcccccCCCCCCCCCCCCCCceEEEceeEEcCh-----hHHHHHHHHHhcCCeEEE--EecC
Q 022276 211 GLMNSAFEYILKAGGVEREKDYPYTGTDGGSCKFDKSKIAAAVSNFSVISS-----DEDQMAANLVKHGPLAGN--VASI 283 (300)
Q Consensus 211 G~~~~a~~y~~~~~G~~~e~~yPY~~~~~~~C~~~~~~~~~~i~~~~~v~~-----~~~~i~~al~~~GPv~v~--i~a~ 283 (300)
|....+..|+.++.|.+.|.+-||.... ..|..... ...+++.-..++. ++..|++++..+|-++.. |++.
T Consensus 169 g~~~m~~a~l~e~sgpv~et~d~y~~~s-~~~~~~~p-~~k~~~~~~~i~~~~~~LdnG~i~~~~~~yg~~s~~~~id~~ 246 (372)
T COG4870 169 GNADMSAAYLTEWSGPVYETDDPYSENS-YFSPTNLP-VTKHVQEAQIIPSRKKYLDNGNIKAMFGFYGAVSSSMYIDAT 246 (372)
T ss_pred CccccccccccccCCcchhhcCcccccc-ccCCcCCc-hhhccccceecccchhhhcccchHHHHhhhccccceeEEecc
Confidence 8888888899999999999999998766 55543221 1223333333332 455688888888877644 5665
Q ss_pred CCCCccCeeEecCC
Q 022276 284 ELPHISFSFLFTVS 297 (300)
Q Consensus 284 ~f~~Y~~Giy~~~~ 297 (300)
.+..-.-++|+..+
T Consensus 247 ~~~~~~~~~~~~~s 260 (372)
T COG4870 247 NSLGICIPYPYVDS 260 (372)
T ss_pred cccccccCCCCCCc
Confidence 54445555555544
No 20
>cd00585 Peptidase_C1B Peptidase C1B subfamily (MEROPS database nomenclature); composed of eukaryotic bleomycin hydrolases (BH) and bacterial aminopeptidases C (pepC). The proteins of this subfamily contain a large insert relative to the C1A peptidase (papain) subfamily. BH is a cysteine peptidase that detoxifies bleomycin by hydrolysis of an amide group. It acts as a carboxypeptidase on its C-terminus to convert itself into an aminopeptidase and peptide ligase. BH is found in all tissues in mammals as well as in many other eukaryotes. Bleomycin, a glycopeptide derived from the fungus Streptomyces verticullus, is an effective anticancer drug due to its ability to induce DNA strand breaks. Human BH is the major cause of tumor cell resistance to bleomycin chemotherapy, and is also genetically linked to Alzheimer's disease. In addition to its peptidase activity, the yeast BH (Gal6) binds DNA and acts as a repressor in the Gal4 regulatory system. BH forms a hexameric ring barrel structure w
Probab=98.68 E-value=1.8e-07 Score=88.52 Aligned_cols=83 Identities=22% Similarity=0.303 Sum_probs=63.4
Q ss_pred cccccCCCCcchHHHHHHHHHHHHHHHh-cCCCccCChhHHHh----------------hCCCCCCCC-----CCCCCCC
Q 022276 150 TGVKDQGACGSCWSFSATGALEGAHFLS-TGELVSLSEQQLVD----------------CDHECDPEE-----SGSCDSG 207 (300)
Q Consensus 150 ~pvknQg~CgsCwAfa~~~~~e~~~~i~-~~~~~~lS~Q~lid----------------C~~~~~~~~-----~~~~~~g 207 (300)
.||+||+.-|-||.||+...++..+..+ +.+.+.||+.++.- +... +.+ +-....-
T Consensus 55 ~~vtnQ~~SGrCW~FA~Ln~lr~~~~k~~~~~~felSq~Yl~f~dklEkaN~fle~ii~~~~~--~~~~R~v~~ll~~~~ 132 (437)
T cd00585 55 EPVTNQKSSGRCWLFAALNVLRHQFMKKLNLKEFEFSQSYLFFWDKLEKANYFLENIIETADE--PLDDRLVQFLLANPQ 132 (437)
T ss_pred CCcccCCCCchhHHHHCHHHHHHHHHHHcCCCCEEeCcHHHHHHHHHHHHHHHHHHHHHHhcC--CCccHHHHHHHhCCc
Confidence 3899999999999999999999988764 55689999988864 2110 000 0002445
Q ss_pred CCCCChHHHHHHHHHhCCcCCCcccccC
Q 022276 208 CNGGLMNSAFEYILKAGGVEREKDYPYT 235 (300)
Q Consensus 208 C~GG~~~~a~~y~~~~~G~~~e~~yPY~ 235 (300)
.+||.-..+...+.++ |+++.+.||=+
T Consensus 133 ~DGGqw~m~~~li~KY-GvVPk~~~pet 159 (437)
T cd00585 133 NDGGQWDMLVNLIEKY-GLVPKSVMPES 159 (437)
T ss_pred CCCCchHHHHHHHHHc-CCCcccccCCC
Confidence 6899999999999888 89999999854
No 21
>PF03051 Peptidase_C1_2: Peptidase C1-like family This family is a subfamily of the Prosite entry; InterPro: IPR004134 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This group of proteins belong to MEROPS peptidase family C1, sub-family C1B (bleomycin hydrolase, clan CA). This family contains prokaryotic and eukaryotic aminopeptidases and bleomycin hydrolases.; GO: 0004197 cysteine-type endopeptidase activity, 0006508 proteolysis; PDB: 3PW3_F 2CB5_A 1CB5_C 2DZZ_A 2E02_A 2E01_A 2E03_A 1A6R_A 1GCB_A 3GCB_A ....
Probab=97.68 E-value=5.8e-05 Score=71.66 Aligned_cols=83 Identities=27% Similarity=0.362 Sum_probs=51.0
Q ss_pred cccccCCCCcchHHHHHHHHHHHHHHHhcC-CCccCChhHHH----------------hhCCCCCCCCC-----CCCCCC
Q 022276 150 TGVKDQGACGSCWSFSATGALEGAHFLSTG-ELVSLSEQQLV----------------DCDHECDPEES-----GSCDSG 207 (300)
Q Consensus 150 ~pvknQg~CgsCwAfa~~~~~e~~~~i~~~-~~~~lS~Q~li----------------dC~~~~~~~~~-----~~~~~g 207 (300)
.||.||..-|-||.||+..+++..+..+.+ +.+.||+-.|. ++... +.+. -.....
T Consensus 56 ~~vtnQk~SGRCW~FA~lN~lR~~~~kk~~l~~felSq~Yl~F~DKlEKaN~fLe~ii~~~~~--~~d~R~v~~ll~~~~ 133 (438)
T PF03051_consen 56 GPVTNQKSSGRCWLFAALNVLRHEIMKKLNLKDFELSQNYLFFWDKLEKANYFLENIIDTADE--PLDDRLVRFLLKNPV 133 (438)
T ss_dssp -S--B--BSSTHHHHHHHHHHHHHHHHHCT-SS--B-HHHHHHHHHHHHHHHHHHHHHHCCTS---TTSHHHHHHHHSTT
T ss_pred CCCCCCCCCCCcchhhchHHHHHHHHHHcCCCceEeechHHHHHHHHHHHHHHHHHHHHHhcC--CcchHHHHHHHhcCC
Confidence 499999999999999999999999887765 67999999875 33211 0000 001234
Q ss_pred CCCCChHHHHHHHHHhCCcCCCcccccC
Q 022276 208 CNGGLMNSAFEYILKAGGVEREKDYPYT 235 (300)
Q Consensus 208 C~GG~~~~a~~y~~~~~G~~~e~~yPY~ 235 (300)
.+||.-..+.+-+.++ |+|+.+.||=+
T Consensus 134 ~DGGqw~~~~nli~KY-GvVPk~~mpet 160 (438)
T PF03051_consen 134 SDGGQWDMVVNLIKKY-GVVPKSVMPET 160 (438)
T ss_dssp -S-B-HHHHHHHHHHH----BGGGSTTG
T ss_pred CCCCchHHHHHHHHHc-CcCcHhhCCCC
Confidence 6899999998888888 89999999975
No 22
>COG3579 PepC Aminopeptidase C [Amino acid transport and metabolism]
Probab=91.43 E-value=0.26 Score=44.82 Aligned_cols=84 Identities=23% Similarity=0.244 Sum_probs=50.2
Q ss_pred ccccCCCCcchHHHHHHHHHHHHHHHhcC-CCccCChhHHHhhCCCCC-------------CCC------CCCCCCCCCC
Q 022276 151 GVKDQGACGSCWSFSATGALEGAHFLSTG-ELVSLSEQQLVDCDHECD-------------PEE------SGSCDSGCNG 210 (300)
Q Consensus 151 pvknQg~CgsCwAfa~~~~~e~~~~i~~~-~~~~lS~Q~lidC~~~~~-------------~~~------~~~~~~gC~G 210 (300)
||-||...|-||-||+...+---+.-+-+ +.+.||..++.--+.... ... +--...--+|
T Consensus 59 ~vtNQk~SGRCWmFAAlNtfRhk~~~el~le~fElSQaytfFwDKlEKaN~FleqIi~tadq~ldsRlv~~LL~~PqqDG 138 (444)
T COG3579 59 KVTNQKQSGRCWMFAALNTFRHKLISELKLEDFELSQAYTFFWDKLEKANWFLEQIIETADQELDSRLVSFLLATPQQDG 138 (444)
T ss_pred ccccccccceehHHHHHHHHHHHHHHhcCcceeehhhHHHHHHHHHHHhhHHHHHHHhhcccchHHHHHHHHHcCccccC
Confidence 89999999999999998886544322222 346677666543221100 000 0000112267
Q ss_pred CChHHHHHHHHHhCCcCCCcccccC
Q 022276 211 GLMNSAFEYILKAGGVEREKDYPYT 235 (300)
Q Consensus 211 G~~~~a~~y~~~~~G~~~e~~yPY~ 235 (300)
|--..-..-+.++ |+++-++||=+
T Consensus 139 GQwdM~v~l~eKY-GvVpK~~ypes 162 (444)
T COG3579 139 GQWDMFVSLFEKY-GVVPKSVYPES 162 (444)
T ss_pred chHHHHHHHHHHh-CCCchhhcccc
Confidence 7666556666666 89999999975
No 23
>KOG4128 consensus Bleomycin hydrolases and aminopeptidases of cysteine protease family [Amino acid transport and metabolism]
Probab=89.94 E-value=0.24 Score=44.95 Aligned_cols=86 Identities=22% Similarity=0.276 Sum_probs=54.9
Q ss_pred ccccccCCCCcchHHHHHHHHHHHHHHHhcC-CCccCChhHHHh----------------hCCCCCCCCCC-----CCCC
Q 022276 149 VTGVKDQGACGSCWSFSATGALEGAHFLSTG-ELVSLSEQQLVD----------------CDHECDPEESG-----SCDS 206 (300)
Q Consensus 149 v~pvknQg~CgsCwAfa~~~~~e~~~~i~~~-~~~~lS~Q~lid----------------C~~~~~~~~~~-----~~~~ 206 (300)
-+||-||..-|-||.|+....+---+..+-+ ..+.||..+|+- -...|.|.+.. ..+-
T Consensus 62 ~~pvtnqkssGrcWift~ln~lrl~~~~kLnl~eFElSqayLFFwdKlErcnyFL~~vvd~a~r~ep~DgRlvq~Ll~nP 141 (457)
T KOG4128|consen 62 RQPVTNQKSSGRCWIFTGLNLLRLEMDRKLNLPEFELSQAYLFFWDKLERCNYFLWTVVDLAMRCEPLDGRLVQNLLKNP 141 (457)
T ss_pred CcccccCcCCCceEEEechhHHHHHHHhcCCcchhhhhhHHHHHHHHHHHHHHHHHHHHHHHhhcCCcccHHHHHHHhCC
Confidence 4699999999999999999886543333322 357788887742 22222222100 0112
Q ss_pred CCCCCChHHHHHHHHHhCCcCCCcccccC
Q 022276 207 GCNGGLMNSAFEYILKAGGVEREKDYPYT 235 (300)
Q Consensus 207 gC~GG~~~~a~~y~~~~~G~~~e~~yPY~ 235 (300)
.=+||.-..-.+.++++ |+..-.|||-.
T Consensus 142 ~~DGGqw~MfvNlVkKY-GviPKkcy~~s 169 (457)
T KOG4128|consen 142 VPDGGQWQMFVNLVKKY-GVIPKKCYLHS 169 (457)
T ss_pred CCCCchHHHHHHHHHHh-CCCcHHhcccc
Confidence 22688777777777777 89999999764
No 24
>PF08127 Propeptide_C1: Peptidase family C1 propeptide; InterPro: IPR012599 This domain is found at the N-terminal of cathepsin B and cathepsin B-like peptidases that belong to MEROPS peptidase subfamily C1A. Cathepsin B are lysosomal cysteine proteinases belonging to the papain superfamily and are unique in their ability to act as both an endo- and an exopeptidases. They are synthesized as inactive zymogens. Activation of the peptidases occurs with the removal of the propeptide [, ]. ; GO: 0004197 cysteine-type endopeptidase activity, 0050790 regulation of catalytic activity; PDB: 1MIR_A 1PBH_A 2PBH_A 3PBH_A.
Probab=88.19 E-value=0.34 Score=30.21 Aligned_cols=34 Identities=18% Similarity=0.172 Sum_probs=19.3
Q ss_pred HHHHHhcCCCCCeeeeeccCCCCChhhHHhhhcCCC
Q 022276 82 RRAKRRQLLDPTAVHGVTKFSDLTPSEFRRQFLGLN 117 (300)
Q Consensus 82 ~~I~~~N~~~~s~~~giN~FsDlt~~Ef~~~~~g~~ 117 (300)
+.|+..|..+.+++.|.| |.+.+.+.++. ++|..
T Consensus 4 e~I~~IN~~~~tWkAG~N-F~~~~~~~ik~-LlGv~ 37 (41)
T PF08127_consen 4 EFIDYINSKNTTWKAGRN-FENTSIEYIKR-LLGVL 37 (41)
T ss_dssp HHHHHHHHCT-SEEE-----SSB-HHHHHH-CS-B-
T ss_pred HHHHHHHcCCCcccCCCC-CCCCCHHHHHH-HcCCC
Confidence 356777777889999999 88888887766 45543
No 25
>PF07172 GRP: Glycine rich protein family; InterPro: IPR010800 This family consists of glycine rich proteins. Some of them may be involved in resistance to environmental stress [].
Probab=78.69 E-value=1.7 Score=32.44 Aligned_cols=9 Identities=22% Similarity=0.113 Sum_probs=4.9
Q ss_pred ChhhHHHHH
Q 022276 1 MERLILSSL 9 (300)
Q Consensus 1 m~~~~ll~l 9 (300)
|++..+|+|
T Consensus 1 MaSK~~llL 9 (95)
T PF07172_consen 1 MASKAFLLL 9 (95)
T ss_pred CchhHHHHH
Confidence 776544443
No 26
>PF08139 LPAM_1: Prokaryotic membrane lipoprotein lipid attachment site; InterPro: IPR012640 In prokaryotes, membrane lipoproteins are synthesized with a precursor signal peptide, which is cleaved by a specific lipoprotein signal peptidase (signal peptidase II). The peptidase recognises a conserved sequence and cuts upstream of a cysteine residue to which a glyceride-fatty acid lipid is attached [,]. This lipid attachment site is found in homologues of the VirB proteins of type IV secretion systems (T4SS). Conjugal transfer across the cell envelope of Gram-negative bacteria is mediated by a supramolecular structure termed mating pair formation (Mpf) complex. Collectively, secretion pathways ancestrally related to bacterial conjugation systems are now known as T4SS. T4SS are involved in the delivery of effector molecules to eukaryotic target cells; each of these systems exports distinct DNA or protein substrates to effect a myriad of changes in host cell physiology during infection [].
Probab=64.96 E-value=4.7 Score=22.18 Aligned_cols=15 Identities=20% Similarity=0.598 Sum_probs=6.5
Q ss_pred hhhHHHHHHHHHHHH
Q 022276 2 ERLILSSLLLLLLSS 16 (300)
Q Consensus 2 ~~~~ll~l~~~~~~~ 16 (300)
+|++++++.++.++.
T Consensus 8 Kkil~~l~a~~~Lag 22 (25)
T PF08139_consen 8 KKILFPLLALFMLAG 22 (25)
T ss_pred HHHHHHHHHHHHHhh
Confidence 555444443333443
No 27
>COG5510 Predicted small secreted protein [Function unknown]
Probab=61.90 E-value=8.9 Score=24.00 Aligned_cols=15 Identities=47% Similarity=0.547 Sum_probs=7.9
Q ss_pred ChhhHHHHHHHHHHH
Q 022276 1 MERLILSSLLLLLLS 15 (300)
Q Consensus 1 m~~~~ll~l~~~~~~ 15 (300)
|+|.+++++++++.+
T Consensus 2 mk~t~l~i~~vll~s 16 (44)
T COG5510 2 MKKTILLIALVLLAS 16 (44)
T ss_pred chHHHHHHHHHHHHH
Confidence 677555554444333
No 28
>PRK10081 entericidin B membrane lipoprotein; Provisional
Probab=59.37 E-value=10 Score=24.40 Aligned_cols=13 Identities=15% Similarity=0.355 Sum_probs=7.2
Q ss_pred ChhhHHHHHHHHH
Q 022276 1 MERLILSSLLLLL 13 (300)
Q Consensus 1 m~~~~ll~l~~~~ 13 (300)
|+|.+.+++++++
T Consensus 2 mKk~i~~i~~~l~ 14 (48)
T PRK10081 2 VKKTIAAIFSVLV 14 (48)
T ss_pred hHHHHHHHHHHHH
Confidence 5666666554443
No 29
>PF10731 Anophelin: Thrombin inhibitor from mosquito; InterPro: IPR018932 Members of this family are all inhibitors of thrombin, the peptidase that is at the end of the blood coagulation cascade and which creates the clot by cleaving fibrinogen. The interaction between thrombin and fibrinogen involves two different areas of contact - via the thrombin active site and via a second substrate-binding site known as an exosite. The inhibitor acts by blocking the exosite, rather than by interacting with the active site. The inhibitors are from mosquitoes that feed on human blood and which, by inhibiting thrombin, prevent the blood from clotting and keep it flowing.
Probab=57.35 E-value=11 Score=25.23 Aligned_cols=19 Identities=32% Similarity=0.562 Sum_probs=10.8
Q ss_pred Ch-hhHHHHHHHHHHHHhhh
Q 022276 1 ME-RLILSSLLLLLLSSVLA 19 (300)
Q Consensus 1 m~-~~~ll~l~~~~~~~~~~ 19 (300)
|+ |++++.||++.|++.+-
T Consensus 1 MA~Kl~vialLC~aLva~vQ 20 (65)
T PF10731_consen 1 MASKLIVIALLCVALVAIVQ 20 (65)
T ss_pred CcchhhHHHHHHHHHHHHHh
Confidence 66 46666666665444333
No 30
>PRK10386 curli assembly protein CsgE; Provisional
Probab=56.83 E-value=22 Score=28.09 Aligned_cols=19 Identities=21% Similarity=0.072 Sum_probs=12.9
Q ss_pred ChhhHHHHHHHHHHHHhhh
Q 022276 1 MERLILSSLLLLLLSSVLA 19 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~~~~ 19 (300)
|+|+...+++.+|++++.+
T Consensus 1 ~~r~~~~~l~~~~l~~~~~ 19 (130)
T PRK10386 1 MKRYLRWIVAAELLFAAGN 19 (130)
T ss_pred ChhHHHHHHHHHHHHhCcc
Confidence 8898777766666555554
No 31
>PF11777 DUF3316: Protein of unknown function (DUF3316); InterPro: IPR016879 There is currently no experimental data for members of this group or their homologues, nor do they exhibit features indicative of any function.
Probab=54.15 E-value=11 Score=28.90 Aligned_cols=19 Identities=53% Similarity=0.679 Sum_probs=13.0
Q ss_pred ChhhHHHHHHHHHHHHhhh
Q 022276 1 MERLILSSLLLLLLSSVLA 19 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~~~~ 19 (300)
|++++|+++++++-+.+.|
T Consensus 1 MKk~~ll~~~ll~s~~a~A 19 (114)
T PF11777_consen 1 MKKIILLASLLLLSSSAFA 19 (114)
T ss_pred CchHHHHHHHHHHHHHHhh
Confidence 8888888866655555555
No 32
>PF05984 Cytomega_UL20A: Cytomegalovirus UL20A protein; InterPro: IPR009245 This family consists of several Cytomegalovirus UL20A proteins. UL20A is thought to be a glycoprotein [].
Probab=52.54 E-value=14 Score=26.64 Aligned_cols=21 Identities=38% Similarity=0.439 Sum_probs=12.3
Q ss_pred Chhh-HHHHHHHHHHHHhhhcc
Q 022276 1 MERL-ILSSLLLLLLSSVLASA 21 (300)
Q Consensus 1 m~~~-~ll~l~~~~~~~~~~~~ 21 (300)
|+|. .+|.||.+-|++++|+.
T Consensus 1 MaRRlwiLslLAVtLtVALAAP 22 (100)
T PF05984_consen 1 MARRLWILSLLAVTLTVALAAP 22 (100)
T ss_pred CchhhHHHHHHHHHHHHHhhcc
Confidence 7765 45556666555555543
No 33
>PRK09810 entericidin A; Provisional
Probab=42.74 E-value=25 Score=21.84 Aligned_cols=9 Identities=56% Similarity=0.774 Sum_probs=5.2
Q ss_pred ChhhHHHHH
Q 022276 1 MERLILSSL 9 (300)
Q Consensus 1 m~~~~ll~l 9 (300)
|+|++++++
T Consensus 2 Mkk~~~l~~ 10 (41)
T PRK09810 2 MKRLIVLVL 10 (41)
T ss_pred hHHHHHHHH
Confidence 666655554
No 34
>PF13529 Peptidase_C39_2: Peptidase_C39 like family; PDB: 3ERV_A.
Probab=40.24 E-value=1.6e+02 Score=22.21 Aligned_cols=20 Identities=15% Similarity=0.192 Sum_probs=15.0
Q ss_pred hHHHHHHHHHhcCCeEEEEe
Q 022276 262 DEDQMAANLVKHGPLAGNVA 281 (300)
Q Consensus 262 ~~~~i~~al~~~GPv~v~i~ 281 (300)
+.+.|+++|....||.+.+.
T Consensus 88 ~~~~i~~~i~~G~Pvi~~~~ 107 (144)
T PF13529_consen 88 SFDDIKQEIDAGRPVIVSVN 107 (144)
T ss_dssp -HHHHHHHHHTT--EEEEEE
T ss_pred cHHHHHHHHHCCCcEEEEEE
Confidence 57889999988779999996
No 35
>PRK10053 hypothetical protein; Provisional
Probab=40.03 E-value=24 Score=27.89 Aligned_cols=19 Identities=21% Similarity=0.265 Sum_probs=11.9
Q ss_pred ChhhHHHHHHHHHHHHhhh
Q 022276 1 MERLILSSLLLLLLSSVLA 19 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~~~~ 19 (300)
|++.+|+++++++.++++|
T Consensus 1 MKK~~~~~~~~~~s~~~~A 19 (130)
T PRK10053 1 MKLQAIALASFLVMPYALA 19 (130)
T ss_pred CcHHHHHHHHHHHHHHHHH
Confidence 8887776666555444444
No 36
>PF02402 Lysis_col: Lysis protein; InterPro: IPR003059 The DNA sequence of the entire colicin E2 operon has been determined []. The operon comprises the colicin activity gene (ceaB), the colicin immunity gene (ceiB) and the lysis gene (celB), which is essential for colicin release from producing cells []. A putative LexA binding site is located upstream from ceaB, and a rho-independent terminator structure is located downstream from celB []. Comparison of the amino acid sequences of colicin E2 and cloacin DF13 reveal extensive similarity. These colicins have different modes of action and recognise different cell surface receptors; the two major regions of heterology at the C terminus, and in the C-terminal end of the central region are thought to correspond to the catalytic and receptor-recognition domains, respectively []. Sequence similarities between colicins E2, A and E1 [] are less striking. The colicin E2 (pyocin) immunity protein does not share similarity with either the colicin E3 or cloacin DF13 [] immunity proteins. By contrast, the lysis proteins of the ColE2, ColE1 and CloDF13 plasmids are almost identical except in the N-terminal regions, which themselves are similar to lipoprotein signal peptides []. Processing of the ColE2 prolysis protein to the mature form is prevented by globomycin, a specific inhibitor of the lipoprotein signal peptidase []. The mature ColE2 lysis protein is located in the cell envelope [].; GO: 0009405 pathogenesis, 0019835 cytolysis, 0019867 outer membrane
Probab=36.74 E-value=14 Score=23.14 Aligned_cols=20 Identities=35% Similarity=0.672 Sum_probs=10.2
Q ss_pred ChhhHHHHHHHH--HHHHhhhc
Q 022276 1 MERLILSSLLLL--LLSSVLAS 20 (300)
Q Consensus 1 m~~~~ll~l~~~--~~~~~~~~ 20 (300)
|++++++.++++ +++++.++
T Consensus 1 MkKi~~~~i~~~~~~L~aCQaN 22 (46)
T PF02402_consen 1 MKKIIFIGIFLLTMLLAACQAN 22 (46)
T ss_pred CcEEEEeHHHHHHHHHHHhhhc
Confidence 777544443333 45555553
No 37
>PRK10449 heat-inducible protein; Provisional
Probab=35.20 E-value=32 Score=27.46 Aligned_cols=19 Identities=21% Similarity=0.434 Sum_probs=12.8
Q ss_pred ChhhHHHHHHHHHHHHhhh
Q 022276 1 MERLILSSLLLLLLSSVLA 19 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~~~~ 19 (300)
|+|+++++++.+++++|.+
T Consensus 1 mk~~~~~~~~~~~l~~C~~ 19 (140)
T PRK10449 1 MKKVVALVALSLLMAGCVS 19 (140)
T ss_pred ChhHHHHHHHHHHHHHhcC
Confidence 8888777666666655555
No 38
>PF06291 Lambda_Bor: Bor protein; InterPro: IPR010438 This family consists of several Bacteriophage lambda Bor and Escherichia coli Iss proteins. Expression of bor significantly increases the survival of the E. coli host cell in animal serum. This property is a well known bacterial virulence determinant indeed, bor and its adjacent sequences are highly homologous to the iss serum resistance locus of the plasmid ColV2-K94, which confers virulence in animals. It has been suggested that lysogeny may generally have a role in bacterial survival in animal hosts, and perhaps in pathogenesis [].
Probab=32.49 E-value=27 Score=26.12 Aligned_cols=21 Identities=33% Similarity=0.586 Sum_probs=15.2
Q ss_pred ChhhHHHHHHHHHHHHhhhcc
Q 022276 1 MERLILSSLLLLLLSSVLASA 21 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~~~~~~ 21 (300)
|+++++...+.++++.++...
T Consensus 1 mKk~ll~~~lallLtgCatqt 21 (97)
T PF06291_consen 1 MKKLLLAAALALLLTGCATQT 21 (97)
T ss_pred CcHHHHHHHHHHHHcccceeE
Confidence 888888777777776666533
No 39
>PF11106 YjbE: Exopolysaccharide production protein YjbE
Probab=32.11 E-value=42 Score=23.79 Aligned_cols=15 Identities=27% Similarity=0.552 Sum_probs=9.2
Q ss_pred ChhhHHHHHHHHHHH
Q 022276 1 MERLILSSLLLLLLS 15 (300)
Q Consensus 1 m~~~~ll~l~~~~~~ 15 (300)
|+|.+++++.++.+.
T Consensus 1 MKK~~~~~~~i~~l~ 15 (80)
T PF11106_consen 1 MKKIIYGLFAILALA 15 (80)
T ss_pred ChhHHHHHHHHHHHH
Confidence 888877555444333
No 40
>PF05543 Peptidase_C47: Staphopain peptidase C47; InterPro: IPR008750 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This group of cysteine peptidases belong to the peptidase family C47 (staphopain family, clan CA). The type example are the staphopains, which are one of four major families of proteinases secreted by the Gram-positive Staphylococcus aureus. These staphylococcal cysteine proteases are secreted as preproenzymes that are proteolytically cleaved to generate the mature enzyme [, , ].; GO: 0008234 cysteine-type peptidase activity, 0006508 proteolysis; PDB: 1X9Y_D 1Y4H_B 1PXV_B 1CV8_A.
Probab=32.09 E-value=2.8e+02 Score=23.15 Aligned_cols=53 Identities=19% Similarity=0.151 Sum_probs=31.7
Q ss_pred cCCCCcchHHHHHHHHHHHHH--------HHhcCCCccCChhHHHhhCCCCCCCCCCCCCCCCCCCChHHHHHHHHHh
Q 022276 154 DQGACGSCWSFSATGALEGAH--------FLSTGELVSLSEQQLVDCDHECDPEESGSCDSGCNGGLMNSAFEYILKA 223 (300)
Q Consensus 154 nQg~CgsCwAfa~~~~~e~~~--------~i~~~~~~~lS~Q~lidC~~~~~~~~~~~~~~gC~GG~~~~a~~y~~~~ 223 (300)
.||.-+-|-+|+.+++|-... .|-+.--..+|+++|.+++. .+.+.++|.+..
T Consensus 18 tQg~~pWCa~Ya~aailN~~~~~~~~~A~~iMr~~yPn~s~~~l~~~~~-----------------~~~~~i~y~ks~ 78 (175)
T PF05543_consen 18 TQGYNPWCAGYAMAAILNATTNTKIYNAKDIMRYLYPNVSEEQLKFTSL-----------------TPNQMIKYAKSQ 78 (175)
T ss_dssp --SSSS-HHHHHHHHHHHHHCT-S---HHHHHHHHSTTS-CCCHHH--B------------------HHHHHHHHHHT
T ss_pred ccCcCcHHHHHHHHHHHHhhhCcCcCCHHHHHHHHCCCCCHHHHhhcCC-----------------CHHHHHHHHHHc
Confidence 488889999999999876542 11111235788888887764 367888887665
No 41
>PRK13883 conjugal transfer protein TrbH; Provisional
Probab=31.69 E-value=32 Score=27.98 Aligned_cols=19 Identities=32% Similarity=0.529 Sum_probs=14.0
Q ss_pred ChhhHHHHHHHHHHHHhhh
Q 022276 1 MERLILSSLLLLLLSSVLA 19 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~~~~ 19 (300)
|.|++++.+|+++|+.|+.
T Consensus 1 Mrk~l~~~~l~l~LaGCAt 19 (151)
T PRK13883 1 MRKIVLLALLALALGGCAT 19 (151)
T ss_pred ChhHHHHHHHHHHHhcccC
Confidence 8888888877776666664
No 42
>PF12276 DUF3617: Protein of unknown function (DUF3617); InterPro: IPR022061 This family of proteins is found in bacteria. Proteins in this family are typically between 155 and 179 amino acids in length. There is a single completely conserved residue C that may be functionally important.
Probab=30.05 E-value=41 Score=27.17 Aligned_cols=15 Identities=47% Similarity=0.634 Sum_probs=9.0
Q ss_pred ChhhHHHHHHHHHHH
Q 022276 1 MERLILSSLLLLLLS 15 (300)
Q Consensus 1 m~~~~ll~l~~~~~~ 15 (300)
|+|.+++++++++++
T Consensus 1 M~~~~~~~~~~~~~~ 15 (162)
T PF12276_consen 1 MKRRLLLALALALLA 15 (162)
T ss_pred CchHHHHHHHHHHHH
Confidence 777766665554443
No 43
>PF10614 CsgF: Type VIII secretion system (T8SS), CsgF protein; InterPro: IPR018893 Fimbriae are cell-surface protein polymers, of e.g. Escherichia coli and Salmonella spp, that mediate interactions important for host and environmental persistence, development of biofilms, motility, colonisation and invasion of cells, and conjugation. Four general assembly pathways for different fimbriae have been proposed, one of which is extracellular nucleation-precipitation (ENP), that differs from the others in that fibre-growth occurs extracellularly. Thin aggregative fimbriae (Tafi) are the only fimbriae dependent on the ENP pathway. Tafi were first identified in Salmonella spp. and the controlling operon termed agf; however subsequent isolation of the homologous operon in E. coli led to its being called csg. Tafi are known as curli because, in the absence of extracellular polysaccharides, their morphology appears curled; however, when expressed with such polysaccharides their morphology appears as a tangled amorphous matrix []. CsgF is one of three putative curli assembly factors appearing to act as a nucleator protein. Unlike eukaryotic amyloid formation, curli biogenesis is a productive pathway requiring a specific assembly machinery [].
Probab=29.63 E-value=96 Score=24.93 Aligned_cols=31 Identities=10% Similarity=0.169 Sum_probs=16.6
Q ss_pred HHHHHHHhCCccCCHHHH-------HHHHHHHHHHHHH
Q 022276 53 FSLFKSKFSKTYATQEEH-------DYRFRVFKANLRR 83 (300)
Q Consensus 53 F~~f~~~~~k~Y~s~~E~-------~~r~~~F~~n~~~ 83 (300)
|..=.+.=+..|+++... ......|.+++++
T Consensus 42 ~LL~~A~AQN~~~dp~~~~~~~~~~~S~l~~F~~sLqs 79 (142)
T PF10614_consen 42 WLLSSAQAQNDFKDPSAEDDFSTSSLSALDRFTQSLQS 79 (142)
T ss_pred HHhhhhhhcCCcCCCccccccccCCCCHHHHHHHHHHH
Confidence 444444445666655443 2236677777763
No 44
>PF11153 DUF2931: Protein of unknown function (DUF2931); InterPro: IPR021326 Some members in this family of proteins are annotated as lipoproteins however this cannot be confirmed. Currently, there is no known function.
Probab=28.74 E-value=42 Score=28.79 Aligned_cols=18 Identities=39% Similarity=0.630 Sum_probs=10.4
Q ss_pred ChhhHHHHHHHHHHHHhhh
Q 022276 1 MERLILSSLLLLLLSSVLA 19 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~~~~ 19 (300)
|+++++|+| +|++++|.+
T Consensus 1 mk~i~~l~l-~lll~~C~~ 18 (216)
T PF11153_consen 1 MKKILLLLL-LLLLTGCST 18 (216)
T ss_pred ChHHHHHHH-HHHHHhhcC
Confidence 777776663 444445444
No 45
>PF11567 PfUIS3: Plasmodium falciparum UIS3 membrane protein; InterPro: IPR021626 UIS3 is a membrane protein essential for sporozoite development in infected hepatocytes. This family is 130-229 of the Plasmodium falciparum UIS3 protein which is compact and has an all alpha-helical structure.PfUIS3(130-229) interacts with lipids, phospholipid lysosomes, the human liver fatty acid-binding protein and with the lipid phosphatidylethanolamine. The interaction with liver fatty acid-binding protein provides the parasite with a method to import essential fatty acids/lipids during rapid growth phases of sporozoites []. ; PDB: 2VWA_C.
Probab=28.57 E-value=35 Score=24.84 Aligned_cols=29 Identities=31% Similarity=0.483 Sum_probs=20.1
Q ss_pred HHHHHHHHHHHHHHHHHHHhcCCCCCeeeeeccCCCCChhh
Q 022276 68 EEHDYRFRVFKANLRRAKRRQLLDPTAVHGVTKFSDLTPSE 108 (300)
Q Consensus 68 ~E~~~r~~~F~~n~~~I~~~N~~~~s~~~giN~FsDlt~~E 108 (300)
+--.+||.+|.+|.+...+| +|++|+.+.
T Consensus 18 DvpiKrfN~F~Dn~rla~qh------------HF~~LSn~Q 46 (101)
T PF11567_consen 18 DVPIKRFNIFMDNARLAAQH------------HFSNLSNEQ 46 (101)
T ss_dssp ---HHHHHHHHHHHHHHHHH------------HHHHS-HHH
T ss_pred cccHHHHHHHHHHHHHHHHH------------HHHhcCcHH
Confidence 44568999999999987777 466776654
No 46
>KOG4702 consensus Uncharacterized conserved protein [Function unknown]
Probab=27.53 E-value=1.6e+02 Score=20.54 Aligned_cols=35 Identities=20% Similarity=0.295 Sum_probs=25.9
Q ss_pred cHHHHHHHHHHHhCCccCCHHHHHHHHHHHHHHHHH
Q 022276 48 NAEHHFSLFKSKFSKTYATQEEHDYRFRVFKANLRR 83 (300)
Q Consensus 48 ~~~~~F~~f~~~~~k~Y~s~~E~~~r~~~F~~n~~~ 83 (300)
+-.+-|++|+..|.+.-.+ .|...|..-|++-+++
T Consensus 26 NQpe~Fee~v~~~krel~p-pe~~~~~EE~~~~lRe 60 (77)
T KOG4702|consen 26 NQPEIFEEFVRGYKRELSP-PEATKRKEEYENFLRE 60 (77)
T ss_pred cChHHHHHHHHhccccCCC-hHHHhhHHHHHHHHHH
Confidence 3445699999999887654 5777788777777664
No 47
>COG3462 Predicted membrane protein [Function unknown]
Probab=25.85 E-value=2.1e+02 Score=21.82 Aligned_cols=22 Identities=23% Similarity=0.270 Sum_probs=14.8
Q ss_pred HHHHHHHhCCccCCHHHHHHHH
Q 022276 53 FSLFKSKFSKTYATQEEHDYRF 74 (300)
Q Consensus 53 F~~f~~~~~k~Y~s~~E~~~r~ 74 (300)
-+--+++|-|---|.||+.++.
T Consensus 91 ~eIlkER~AkGEItEEEY~r~~ 112 (117)
T COG3462 91 EEILKERYAKGEITEEEYRRII 112 (117)
T ss_pred HHHHHHHHhcCCCCHHHHHHHH
Confidence 4555678888877777765443
No 48
>PF04202 Mfp-3: Foot protein 3; InterPro: IPR007328 Mytilus foot protein-3 (Mfp-3) is a highly polymorphic protein family located in the byssal adhesive plaques of blue mussels.
Probab=25.65 E-value=62 Score=22.25 Aligned_cols=16 Identities=38% Similarity=0.538 Sum_probs=8.4
Q ss_pred ChhhHHHHHHHHHHHH
Q 022276 1 MERLILSSLLLLLLSS 16 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~ 16 (300)
|+++.+.+||+|.|..
T Consensus 1 mnn~Si~VLlaLvLIg 16 (71)
T PF04202_consen 1 MNNLSIAVLLALVLIG 16 (71)
T ss_pred CCchhHHHHHHHHHHh
Confidence 7776555554443333
No 49
>PRK15346 outer membrane secretin SsaC; Provisional
Probab=23.95 E-value=54 Score=32.12 Aligned_cols=21 Identities=29% Similarity=0.547 Sum_probs=13.1
Q ss_pred ChhhHHHHHHHHHHHHhhhcc
Q 022276 1 MERLILSSLLLLLLSSVLASA 21 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~~~~~~ 21 (300)
|+|+++|++|+||..+..+++
T Consensus 1 ~~~~~~~~~~~~~~~~~~~~~ 21 (499)
T PRK15346 1 MKKLLILIFLFLLNTAKFAAS 21 (499)
T ss_pred CchhHHHHHHHHHhhhhhhcc
Confidence 777766666666665555544
No 50
>PRK11443 lipoprotein; Provisional
Probab=23.68 E-value=61 Score=25.37 Aligned_cols=18 Identities=39% Similarity=0.488 Sum_probs=9.5
Q ss_pred ChhhHHHHHHHHHHHHhhh
Q 022276 1 MERLILSSLLLLLLSSVLA 19 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~~~~ 19 (300)
|+++++++ ++++|+.+++
T Consensus 1 Mk~~~~~~-~~~lLsgCa~ 18 (124)
T PRK11443 1 MKKFIAPL-LALLLSGCQI 18 (124)
T ss_pred ChHHHHHH-HHHHHHhccC
Confidence 76554444 3445555555
No 51
>PF08138 Sex_peptide: Sex peptide (SP) family; InterPro: IPR012608 This family consists of Sex Peptides (SP) that are found in Drosophila. On mating, Drosophila females decreases her remating rate and increases her egg-laying rate due, in part, to the transfer of SP from the male to the female. SP are found in seminal fluids transferred from the male to the female during mating. The male seminal fluid proteins are referred to as accessory gland proteins (Acps). The SP is one of the most interesting Acps and plays an important role in reproduction [].; GO: 0005179 hormone activity, 0046008 regulation of female receptivity, post-mating, 0005576 extracellular region; PDB: 2LAQ_A.
Probab=23.29 E-value=27 Score=22.86 Aligned_cols=12 Identities=42% Similarity=0.517 Sum_probs=0.0
Q ss_pred ChhhHHHHHHHH
Q 022276 1 MERLILSSLLLL 12 (300)
Q Consensus 1 m~~~~ll~l~~~ 12 (300)
|+..++|+++++
T Consensus 1 Mk~p~~llllvl 12 (56)
T PF08138_consen 1 MKTPIFLLLLVL 12 (56)
T ss_dssp ------------
T ss_pred CcchHHHHHHHH
Confidence 666555554444
No 52
>KOG2735 consensus Phosphatidylserine synthase [Lipid transport and metabolism]
Probab=22.92 E-value=58 Score=30.65 Aligned_cols=21 Identities=38% Similarity=0.612 Sum_probs=19.6
Q ss_pred chHHHHHHHHHHHHHHHhcCC
Q 022276 160 SCWSFSATGALEGAHFLSTGE 180 (300)
Q Consensus 160 sCwAfa~~~~~e~~~~i~~~~ 180 (300)
-||.|+++.++|..+|++-|.
T Consensus 374 qcWv~~aI~~~El~IciKfg~ 394 (466)
T KOG2735|consen 374 QCWVFLAICALELLICIKFGS 394 (466)
T ss_pred hHHHHHHHHHHHhhhheeeCC
Confidence 499999999999999999886
No 53
>PF10880 DUF2673: Protein of unknown function (DUF2673); InterPro: IPR024247 This family of proteins with unknown function appears to be restricted to Rickettsiae spp.
Probab=22.66 E-value=86 Score=20.77 Aligned_cols=21 Identities=38% Similarity=0.488 Sum_probs=10.7
Q ss_pred ChhhHHHHHHHHHHHHhhhcc
Q 022276 1 MERLILSSLLLLLLSSVLASA 21 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~~~~~~ 21 (300)
|++++-++|++.|...+.|++
T Consensus 1 mknllkillilafa~pvfass 21 (65)
T PF10880_consen 1 MKNLLKILLILAFASPVFASS 21 (65)
T ss_pred ChhHHHHHHHHHHhhhHhhhc
Confidence 666544444343555555544
No 54
>TIGR00156 conserved hypothetical protein TIGR00156. As of the last revision, this family consists only of two proteins from Escherichia coli and one from the related species Haemophilus influenzae.
Probab=22.56 E-value=72 Score=25.07 Aligned_cols=12 Identities=17% Similarity=-0.022 Sum_probs=8.1
Q ss_pred ChhhHHHHHHHH
Q 022276 1 MERLILSSLLLL 12 (300)
Q Consensus 1 m~~~~ll~l~~~ 12 (300)
|++++++++++|
T Consensus 1 MKK~~~~~~~~l 12 (126)
T TIGR00156 1 MKFQAIVLASAL 12 (126)
T ss_pred CchHHHHHHHHH
Confidence 888777666533
No 55
>PF02553 CbiN: Cobalt transport protein component CbiN; InterPro: IPR003705 The cobalt transport protein CbiN is part of the active cobalt transport system involved in uptake of cobalt in to the cell involved with cobalamin biosynthesis (vitamin B12). It has been suggested that CbiN may function as the periplasmic binding protein component of the active cobalt transport system [].; GO: 0015087 cobalt ion transmembrane transporter activity, 0006824 cobalt ion transport, 0009236 cobalamin biosynthetic process, 0016020 membrane
Probab=22.49 E-value=74 Score=22.51 Aligned_cols=12 Identities=33% Similarity=0.523 Sum_probs=6.8
Q ss_pred ChhhHHHHHHHH
Q 022276 1 MERLILSSLLLL 12 (300)
Q Consensus 1 m~~~~ll~l~~~ 12 (300)
|++++|++++++
T Consensus 1 ~kn~~l~~~vv~ 12 (74)
T PF02553_consen 1 MKNLLLLLLVVA 12 (74)
T ss_pred CceeHHHHHHHH
Confidence 666666555444
No 56
>PRK13835 conjugal transfer protein TrbH; Provisional
Probab=21.57 E-value=65 Score=25.94 Aligned_cols=19 Identities=37% Similarity=0.595 Sum_probs=13.8
Q ss_pred ChhhHHHHHHHHHHHHhhh
Q 022276 1 MERLILSSLLLLLLSSVLA 19 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~~~~ 19 (300)
|.|++++++++++++.|++
T Consensus 1 mrk~~~~~~~al~LaGCaT 19 (145)
T PRK13835 1 LRRLLAACILALLLSGCQT 19 (145)
T ss_pred ChhHHHHHHHHHHHhcccc
Confidence 7888887777767666665
No 57
>TIGR01165 cbiN cobalt transport protein. This model describes the cobalt transporter in bacteria and its equivalents in archaea. It principally functions in the ion uptake mechanism. It is a multisubunit transporter with two integral membrane proteins and two closely associated cytoplasmic subunits. This transporter belongs to the ABC transporter superfamily (ATP stands for ATP Binding Cassette). This superfamily includes two groups, one which catalyze the uptake of small molecules, including ions from the external milieu and the other group which is engaged in the efflux of small molecular weight compounds and ions from within the cell. Energy derived from the hydrolysis of ATP drive the both the process of uptake and efflux.
Probab=21.50 E-value=87 Score=23.08 Aligned_cols=12 Identities=17% Similarity=0.119 Sum_probs=5.5
Q ss_pred ChhhHHHHHHHH
Q 022276 1 MERLILSSLLLL 12 (300)
Q Consensus 1 m~~~~ll~l~~~ 12 (300)
|++.++|+++++
T Consensus 3 ~~~~~~ll~~v~ 14 (91)
T TIGR01165 3 MKKTIWLLAAVA 14 (91)
T ss_pred cchhHHHHHHHH
Confidence 455554444333
No 58
>PF11912 DUF3430: Protein of unknown function (DUF3430); InterPro: IPR021837 This family of proteins are functionally uncharacterised. This protein is found in eukaryotes. Proteins in this family are typically between 209 to 265 amino acids in length.
Probab=20.93 E-value=76 Score=26.76 Aligned_cols=17 Identities=41% Similarity=0.512 Sum_probs=8.6
Q ss_pred ChhhHHHHHHHHHHHHh
Q 022276 1 MERLILSSLLLLLLSSV 17 (300)
Q Consensus 1 m~~~~ll~l~~~~~~~~ 17 (300)
||=+++|+||++++...
T Consensus 1 MKll~~lilli~~~~~~ 17 (212)
T PF11912_consen 1 MKLLISLILLILLIINF 17 (212)
T ss_pred CcHHHHHHHHHHHHHhh
Confidence 77554555444444443
No 59
>PF09403 FadA: Adhesion protein FadA; InterPro: IPR018543 FadA (Fusobacterium adhesin A) is an adhesin which forms two alpha helices. ; PDB: 3ETZ_B 3ETY_A 2GL2_B 3ETX_C 3ETW_A.
Probab=20.91 E-value=87 Score=24.61 Aligned_cols=21 Identities=19% Similarity=0.294 Sum_probs=12.0
Q ss_pred hhhccHHHHHHHHHHHhCCcc
Q 022276 44 DHLLNAEHHFSLFKSKFSKTY 64 (300)
Q Consensus 44 ~~l~~~~~~F~~f~~~~~k~Y 64 (300)
..+.+.+..|+.-..+-+-.|
T Consensus 27 ~~l~~LEae~q~L~~kE~~r~ 47 (126)
T PF09403_consen 27 SELNQLEAEYQQLEQKEEARY 47 (126)
T ss_dssp HHHHHHHHHHHHHHHHHHHHH
T ss_pred HHHHHHHHHHHHHHHHHHHHH
Confidence 446666777766665544333
No 60
>COG4871 Uncharacterized protein conserved in archaea [Function unknown]
Probab=20.25 E-value=59 Score=26.67 Aligned_cols=14 Identities=36% Similarity=1.061 Sum_probs=9.2
Q ss_pred ccCCCCc--chHHHHH
Q 022276 153 KDQGACG--SCWSFSA 166 (300)
Q Consensus 153 knQg~Cg--sCwAfa~ 166 (300)
-|-|.|| +|+|||.
T Consensus 137 tNCg~CGEqtCmaFAi 152 (193)
T COG4871 137 TNCGKCGEQTCMAFAI 152 (193)
T ss_pred CccccchhHHHHHHHH
Confidence 4555665 6899864
Done!