Query 028703
Match_columns 205
No_of_seqs 136 out of 736
Neff 7.7
Searched_HMMs 46136
Date Fri Mar 29 15:40:33 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/028703.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/028703hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 COG1025 Ptr Secreted/periplasm 100.0 7E-35 1.5E-39 273.6 19.9 182 2-191 753-934 (937)
2 KOG0959 N-arginine dibasic con 100.0 6.1E-35 1.3E-39 277.3 16.8 198 2-200 769-969 (974)
3 PRK15101 protease3; Provisiona 100.0 7.4E-29 1.6E-33 241.0 22.2 184 2-194 773-958 (961)
4 TIGR02110 PQQ_syn_pqqF coenzym 99.4 1.7E-13 3.6E-18 129.3 7.7 64 2-65 632-695 (696)
5 PF05193 Peptidase_M16_C: Pept 98.8 2.9E-08 6.3E-13 76.9 10.1 80 5-87 104-184 (184)
6 COG0612 PqqL Predicted Zn-depe 98.7 5.7E-07 1.2E-11 80.9 15.3 151 5-161 281-432 (438)
7 KOG2067 Mitochondrial processi 97.2 0.0047 1E-07 54.8 11.3 127 12-142 306-432 (472)
8 COG1026 Predicted Zn-dependent 96.9 0.022 4.7E-07 55.8 12.9 146 2-159 810-957 (978)
9 PTZ00432 falcilysin; Provision 96.8 0.036 7.9E-07 55.9 14.1 146 2-161 955-1104(1119)
10 PRK15101 protease3; Provisiona 96.5 0.017 3.6E-07 57.1 10.1 134 5-142 312-449 (961)
11 PTZ00432 falcilysin; Provision 96.5 0.053 1.2E-06 54.7 13.5 149 5-160 408-575 (1119)
12 TIGR02110 PQQ_syn_pqqF coenzym 95.9 0.029 6.4E-07 53.8 7.9 83 5-89 261-348 (696)
13 COG0612 PqqL Predicted Zn-depe 95.1 0.26 5.7E-06 44.3 10.6 98 59-161 108-208 (438)
14 KOG0960 Mitochondrial processi 94.8 0.51 1.1E-05 42.3 11.2 116 35-160 334-449 (467)
15 KOG0961 Predicted Zn2+-depende 94.3 0.52 1.1E-05 45.0 10.6 148 3-156 831-984 (1022)
16 KOG0959 N-arginine dibasic con 93.7 2 4.4E-05 42.8 13.7 114 36-161 582-696 (974)
17 COG1025 Ptr Secreted/periplasm 93.3 2.7 5.9E-05 41.5 13.8 118 35-163 574-691 (937)
18 KOG2681 Metal-dependent phosph 86.6 0.72 1.6E-05 41.7 3.4 111 6-128 40-154 (498)
19 PF09568 RE_MjaI: MjaI restric 86.1 2 4.3E-05 34.0 5.3 39 94-142 45-83 (170)
20 KOG0960 Mitochondrial processi 73.9 24 0.00051 32.0 8.3 69 74-142 141-210 (467)
21 cd08305 Pyrin Pyrin: a protein 63.5 21 0.00045 24.1 4.7 67 70-136 3-70 (73)
22 PHA02698 hypothetical protein; 62.2 27 0.00058 23.9 4.9 36 50-85 39-76 (89)
23 COG1026 Predicted Zn-dependent 56.5 1.4E+02 0.0031 30.1 10.7 143 8-159 305-458 (978)
24 PF09851 SHOCT: Short C-termin 55.7 23 0.0005 19.7 3.2 23 66-88 6-30 (31)
25 PF06518 DUF1104: Protein of u 51.2 63 0.0014 23.0 5.7 36 64-99 49-84 (93)
26 PF05193 Peptidase_M16_C: Pept 49.2 18 0.0004 26.9 2.9 31 128-163 1-31 (184)
27 PF08006 DUF1700: Protein of u 49.1 53 0.0012 25.8 5.6 31 62-92 4-34 (181)
28 PF07609 DUF1572: Protein of u 48.7 23 0.00051 27.9 3.4 87 54-140 4-94 (163)
29 PF08621 RPAP1_N: RPAP1-like, 44.1 34 0.00074 21.4 3.0 23 69-91 9-31 (49)
30 PF14270 DUF4358: Domain of un 40.7 1.1E+02 0.0024 21.8 5.8 61 27-88 33-93 (106)
31 cd08317 Death_ank Death domain 38.1 1.3E+02 0.0028 20.5 8.8 74 54-138 7-83 (84)
32 PF05120 GvpG: Gas vesicle pro 35.1 1.5E+02 0.0033 20.4 5.4 37 52-92 28-66 (79)
33 cd00498 Hsp33 Heat shock prote 32.8 3.1E+02 0.0066 23.2 8.6 117 10-140 133-257 (275)
34 PF13333 rve_2: Integrase core 32.2 51 0.0011 20.3 2.5 32 52-83 18-50 (52)
35 PF14203 DUF4319: Domain of un 32.1 70 0.0015 21.2 3.1 25 59-83 35-59 (64)
36 KOG2085 Serine/threonine prote 30.9 1E+02 0.0022 28.1 4.9 44 64-107 319-366 (457)
37 PF11594 Med28: Mediator compl 29.3 1.5E+02 0.0032 21.7 4.7 47 56-109 5-51 (106)
38 PF04485 NblA: Phycobilisome d 27.4 1E+02 0.0022 19.6 3.2 35 66-100 13-47 (53)
39 PF08671 SinI: Anti-repressor 27.2 84 0.0018 17.5 2.5 19 122-140 9-28 (30)
40 PF02758 PYRIN: PAAD/DAPIN/Pyr 26.9 59 0.0013 22.2 2.3 70 68-138 5-81 (83)
41 KOG2019 Metalloendoprotease HM 26.9 6.4E+02 0.014 25.1 11.9 144 4-161 332-497 (998)
42 PF05553 DUF761: Cotton fibre 26.8 1.4E+02 0.0031 17.5 4.0 20 54-73 2-21 (38)
43 COG1078 HD superfamily phospho 26.6 39 0.00084 30.7 1.6 81 6-97 18-101 (421)
44 PF07735 FBA_2: F-box associat 26.2 1E+02 0.0022 19.8 3.3 26 128-155 43-68 (70)
45 cd08803 Death_ank3 Death domai 26.1 2.3E+02 0.0049 19.6 6.0 58 74-138 26-83 (84)
46 cd08304 DD_superfamily The Dea 25.3 1.9E+02 0.0042 19.0 4.5 63 71-136 4-67 (69)
47 PF00675 Peptidase_M16: Insuli 23.9 2.9E+02 0.0064 20.2 10.9 86 17-113 51-138 (149)
48 PF14659 Phage_int_SAM_3: Phag 23.4 64 0.0014 19.6 1.8 17 126-142 42-58 (58)
49 PF06816 NOD: NOTCH protein; 22.2 1.5E+02 0.0033 19.0 3.3 30 41-73 7-36 (57)
50 PF08700 Vps51: Vps51/Vps67; 22.1 2.4E+02 0.0053 18.8 4.7 43 69-111 13-55 (87)
51 PF11385 DUF3189: Protein of u 21.9 2.2E+02 0.0048 21.9 4.8 55 9-68 33-87 (148)
52 PF11460 DUF3007: Protein of u 21.3 1.3E+02 0.0028 21.9 3.2 19 67-85 82-100 (104)
53 KOG2019 Metalloendoprotease HM 21.2 4.6E+02 0.01 26.0 7.5 142 3-161 841-986 (998)
54 PF11264 ThylakoidFormat: Thyl 20.6 4.1E+02 0.0088 22.0 6.4 58 66-140 53-111 (216)
55 PRK03573 transcriptional regul 20.5 3.6E+02 0.0077 19.9 6.3 60 24-88 71-131 (144)
56 cd08805 Death_ank1 Death domai 20.5 3E+02 0.0066 19.0 8.3 58 74-138 26-83 (84)
57 cd08321 Pyrin_ASC-like Pyrin D 20.1 1.4E+02 0.0031 20.5 3.1 24 69-92 5-28 (82)
No 1
>COG1025 Ptr Secreted/periplasmic Zn-dependent peptidases, insulinase-like [Posttranslational modification, protein turnover, chaperones]
Probab=100.00 E-value=7e-35 Score=273.60 Aligned_cols=182 Identities=29% Similarity=0.462 Sum_probs=165.2
Q ss_pred chHHHHHHHHHchHHHHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHhcCCHHHHH
Q 028703 2 NVKLQLLALIAKQPAFHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLYEMTSDQFK 81 (205)
Q Consensus 2 ~a~~~Ll~~ils~~~f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~~ls~eeF~ 81 (205)
.|+..|+.++++.+||++|||||||||+|+|+++.+.+.+|+.|+|||+.++|++|.+||..|++.+...|.+|++++|+
T Consensus 753 ~a~s~Ll~~l~~~~ff~~LRTkeQLGY~Vfs~~~~v~~~~gi~f~vqS~~~~p~~L~~r~~~F~~~~~~~l~~ms~e~Fe 832 (937)
T COG1025 753 SALSSLLGQLIHPWFFDQLRTKEQLGYAVFSGPREVGRTPGIGFLVQSNSKSPSYLLERINAFLETAEPELREMSEEDFE 832 (937)
T ss_pred HHHHHHHHHHHhHHhHHHhhhhhhcceEEEecceeecCccceEEEEeCCCCChHHHHHHHHHHHHHHHHHHHhCCHHHHH
Confidence 47889999999999999999999999999999999999999999999999999999999999999999999999999999
Q ss_pred HHHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHHHHhhhcCCCCccEEEEEEeeCCC
Q 028703 82 NNVNALIDMKLEKHKNLKEESGFYWREISDGILKFDRREVEVAALRQLTQQELIYFFNENIKAGAPRKKTLSVRVYGSLH 161 (205)
Q Consensus 82 ~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~~~~~~~~~~~~~~l~i~v~~~~~ 161 (205)
.+|++|++++.++++|+.+++.|+|..|..|.++|+++++.|+++++||++++++||.+.+ ...++.++++||.|++.
T Consensus 833 ~~k~alin~il~~~~nl~e~a~r~~~~~~~g~~~Fd~~ek~i~~vk~LT~~~l~~f~~~~l--~~~~g~~l~~~i~g~~~ 910 (937)
T COG1025 833 QIKKALINQILQPPQNLAEEASRLWKAFGRGNLDFDHREKKIEAVKTLTKQKLLDFFENAL--SYEQGSKLLSHIRGQNG 910 (937)
T ss_pred HHHHHHHHHHHccCCCHHHHHHHHHHHhccCCCCcCcHHHHHHHHHhcCHHHHHHHHHHhh--cccccceeeeeeecccc
Confidence 9999999999999999999999999999999999999999999999999999999999999 46788999999999655
Q ss_pred CcccccccCCCCCCCccccCCHHhHhccCC
Q 028703 162 APELKEETSESADPHIVHIDDIFSFRRSQP 191 (205)
Q Consensus 162 ~~~~~~~~~~~~~~~~~~i~d~~~fk~~~~ 191 (205)
.+ +....++-..+++..+++..++
T Consensus 911 e~------~~~~~~~~~~~~~~~~~~~~~~ 934 (937)
T COG1025 911 EA------EYAHPEGWTVLENVSALQQTAP 934 (937)
T ss_pred cc------ccccCCceeeehhhhhhccccc
Confidence 22 2333344455666666665554
No 2
>KOG0959 consensus N-arginine dibasic convertase NRD1 and related Zn2+-dependent endopeptidases, insulinase superfamily [Posttranslational modification, protein turnover, chaperones]
Probab=100.00 E-value=6.1e-35 Score=277.28 Aligned_cols=198 Identities=44% Similarity=0.662 Sum_probs=178.8
Q ss_pred chHHHHHHHHHchHHHHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHhcCCHHHHH
Q 028703 2 NVKLQLLALIAKQPAFHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLYEMTSDQFK 81 (205)
Q Consensus 2 ~a~~~Ll~~ils~~~f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~~ls~eeF~ 81 (205)
.|++.|+.+++++|+|++||||+||||+|+++.+...|+.|+.|+|||+ +++++|+.||+.|++.+++.|.+|++++|+
T Consensus 769 ~~~~~L~~~li~ep~Fd~LRTkeqLGYiv~~~~r~~~G~~~~~i~Vqs~-~~~~~le~rIe~fl~~~~~~i~~m~~e~Fe 847 (974)
T KOG0959|consen 769 NAVLGLLEQLIKEPAFDQLRTKEQLGYIVSTGVRLNYGTVGLQITVQSE-KSVDYLEERIESFLETFLEEIVEMSDEEFE 847 (974)
T ss_pred HHHHHHHHHHhccchHHhhhhHHhhCeEeeeeeeeecCcceeEEEEccC-CCchHHHHHHHHHHHHHHHHHHhcchhhhh
Confidence 4789999999999999999999999999999999999999999999999 999999999999999999999999999999
Q ss_pred HHHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHHHHhhhcCCCCccEEEEEEeeCCC
Q 028703 82 NNVNALIDMKLEKHKNLKEESGFYWREISDGILKFDRREVEVAALRQLTQQELIYFFNENIKAGAPRKKTLSVRVYGSLH 161 (205)
Q Consensus 82 ~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~~~~~~~~~~~~~~l~i~v~~~~~ 161 (205)
.++.++|..+.++|+|+.+++.++|.+|..+.|+|+++++.++++++||+++++.||..++...+.++++++|++.|+..
T Consensus 848 ~~~~~lI~~~~ek~~~l~~e~~~~w~ei~~~~y~f~r~~~~v~~l~~i~k~~~i~~f~~~~~~~a~~~~~lsv~~~~~~~ 927 (974)
T KOG0959|consen 848 KHKSGLIASKLEKPKNLSEESSRYWDEIIIGQYNFDRDEKEVEALKKITKEDVINFFDEYIRKGAAKRKKLSVHVHGKQL 927 (974)
T ss_pred hhHHHHHHHHhhcCcchhHHHHHHHHHHHhhhhcchhhHHHHHHHHhhhHHHHHHHHHhhccccchhcceEEEEecCchh
Confidence 99999999999999999999999999999999999999999999999999999999999998888999999999999977
Q ss_pred CcccccccC---CCCCCCccccCCHHhHhccCCCcCCCCCCc
Q 028703 162 APELKEETS---ESADPHIVHIDDIFSFRRSQPLYGSFKGGF 200 (205)
Q Consensus 162 ~~~~~~~~~---~~~~~~~~~i~d~~~fk~~~~~~~~~~~~~ 200 (205)
..+...... .........|+|+..||+..++||......
T Consensus 928 ~~~~~~~~~~~~~~~~~~~~~I~d~~~fk~~~~l~~~~~~~~ 969 (974)
T KOG0959|consen 928 DEEASSEKIKSQSENLLKIKEITDIVAFKRSLPLYPLVKPVI 969 (974)
T ss_pred hhhhhcccchhhhhhcccccchHHHHHhhccccccccccccc
Confidence 433211111 111122334999999999999999877543
No 3
>PRK15101 protease3; Provisional
Probab=99.97 E-value=7.4e-29 Score=240.96 Aligned_cols=184 Identities=27% Similarity=0.408 Sum_probs=166.0
Q ss_pred chHHHHHHHHHchHHHHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHhcCCHHHHH
Q 028703 2 NVKLQLLALIAKQPAFHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLYEMTSDQFK 81 (205)
Q Consensus 2 ~a~~~Ll~~ils~~~f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~~ls~eeF~ 81 (205)
+++..+|++++++++|++||||+||||+|+|+.....++.|+.|+|||+.++|+++..+|+.|+.++...+++||++||+
T Consensus 773 ~v~~~lLg~~~ssrlf~~LRtk~qLgY~V~s~~~~~~~~~~~~~~vqs~~~~~~~l~~~i~~f~~~~~~~l~~lt~eE~~ 852 (961)
T PRK15101 773 SAYSSLLGQIIQPWFYNQLRTEEQLGYAVFAFPMSVGRQWGMGFLLQSNDKQPAYLWQRYQAFFPQAEAKLRAMKPEEFA 852 (961)
T ss_pred HHHHHHHHHHHhHHHHHHHHHHhhhceEEEEEeeccCCeeeEEEEEECCCCCHHHHHHHHHHHHHHHHHHHHhCCHHHHH
Confidence 57889999999999999999999999999999999999999999999999999999999999999987788899999999
Q ss_pred HHHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHHHHh-hhcCCCCccEEEEEEeeCC
Q 028703 82 NNVNALIDMKLEKHKNLKEESGFYWREISDGILKFDRREVEVAALRQLTQQELIYFFNEN-IKAGAPRKKTLSVRVYGSL 160 (205)
Q Consensus 82 ~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~~~~-~~~~~~~~~~l~i~v~~~~ 160 (205)
.+|+++++++..+++|+.+++.++|.+|..+++.|++.++.++.|++||++|+++|++++ + ...+.+++++|.|..
T Consensus 853 ~~k~~l~~~~~~~~~sl~~~a~~~~~~i~~~~~~fd~~~~~~~~i~~vT~edv~~~~~~~~~---~~~~~~~~~~~~~~~ 929 (961)
T PRK15101 853 QYQQALINQLLQAPQTLGEEASRLSKDFDRGNMRFDSRDKIIAQIKLLTPQKLADFFHQAVI---EPQGLAILSQISGSQ 929 (961)
T ss_pred HHHHHHHHHhcCCCCCHHHHHHHHHHHHhcCCCCcChHHHHHHHHHcCCHHHHHHHHHHHhc---CCCCCEEEEEeeccC
Confidence 999999999999999999999999999999999999999999999999999999999998 5 456668999999987
Q ss_pred CCcccccccCCCC-CCCccccCCHHhHhccCCCcC
Q 028703 161 HAPELKEETSESA-DPHIVHIDDIFSFRRSQPLYG 194 (205)
Q Consensus 161 ~~~~~~~~~~~~~-~~~~~~i~d~~~fk~~~~~~~ 194 (205)
+... ... ..+...|+|+..||+.+++..
T Consensus 930 ~~~~------~~~~~~~~~~~~~~~~~~~~~~~~~ 958 (961)
T PRK15101 930 NGKA------DYAHPKGWKTWENVSALQQTLPVME 958 (961)
T ss_pred cccc------ccccccCCeeeCCHHHHhhcCcccc
Confidence 6321 111 122456999999998887643
No 4
>TIGR02110 PQQ_syn_pqqF coenzyme PQQ biosynthesis probable peptidase PqqF. In a subset of species that make coenzyme PQQ (pyrrolo-quinoline-quinone), this probable peptidase is found in the PQQ biosynthesis region and is thought to act as a protease on PqqA (TIGR02107), a probable peptide precursor of the coenzyme. PQQ is required for some glucose dehydrogenases and alcohol dehydrogenases.
Probab=99.44 E-value=1.7e-13 Score=129.27 Aligned_cols=64 Identities=28% Similarity=0.405 Sum_probs=62.8
Q ss_pred chHHHHHHHHHchHHHHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHH
Q 028703 2 NVKLQLLALIAKQPAFHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFL 65 (205)
Q Consensus 2 ~a~~~Ll~~ils~~~f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl 65 (205)
.|.+.||+++++.|||+.||.++||||+|+|+++.+.|..|+.|.||||..++..|..+|+.||
T Consensus 632 ~aa~rlla~l~~~~f~qrlRve~qlGY~v~~~~~~~~~~~gllf~~QSP~~~~~~l~~h~~~fl 695 (696)
T TIGR02110 632 EAAWRLLAQLLEPPFFQRLRVELQLGYVVFCRYRRVADRDGLLFALQSPDASARELLQHIKRFL 695 (696)
T ss_pred HHHHHHHHHHhchhHHHHHHHhhccceEEEEeeEEcCCcceeEEEEeCCCCCHHHHHHHHHHHh
Confidence 5889999999999999999999999999999999999999999999999999999999999998
No 5
>PF05193 Peptidase_M16_C: Peptidase M16 inactive domain; InterPro: IPR007863 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Metalloproteases are the most diverse of the four main types of protease, with more than 50 families identified to date. In these enzymes, a divalent cation, usually zinc, activates the water molecule. The metal ion is held in place by amino acid ligands, usually three in number. The known metal ligands are His, Glu, Asp or Lys and at least one other residue is required for catalysis, which may play an electrophillic role. Of the known metalloproteases, around half contain an HEXXH motif, which has been shown in crystallographic studies to form part of the metal-binding site []. The HEXXH motif is relatively common, but can be more stringently defined for metalloproteases as 'abXHEbbHbc', where 'a' is most often valine or threonine and forms part of the S1' subsite in thermolysin and neprilysin, 'b' is an uncharged residue, and 'c' a hydrophobic residue. Proline is never found in this site, possibly because it would break the helical structure adopted by this motif in metalloproteases []. These metallopeptidases belong to MEROPS peptidase family M16 (clan ME). They include proteins, which are classified as non-peptidase homologues either have been found experimentally to be without peptidase activity, or lack amino acid residues that are believed to be essential for the catalytic activity. The peptidases in this group of sequences include: Insulinase, insulin-degrading enzyme (3.4.24.56 from EC) Mitochondrial processing peptidase alpha subunit, (Alpha-MPP, 3.4.24.64 from EC) Pitrlysin, Protease III precursor (3.4.24.55 from EC) Nardilysin, (3.4.24.61 from EC) Ubiquinol-cytochrome C reductase complex core protein I,mitochondrial precursor (1.10.2.2 from EC) Coenzyme PQQ synthesis protein F (3.4.99 from EC) These proteins do not share many regions of sequence similarity; the most noticeable is in the N-terminal section. This region includes a conserved histidine followed, two residues later by a glutamate and another histidine. In pitrilysin, it has been shown [] that this H-x-x-E-H motif is involved in enzymatic activity; the two histidines bind zinc and the glutamate is necessary for catalytic activity. The mitochondrial processing peptidase consists of two structurally related domains. One is the active peptidase whereas the other, the C-terminal region, is inactive. The two domains hold the substrate like a clamp [].; GO: 0004222 metalloendopeptidase activity, 0008270 zinc ion binding, 0006508 proteolysis; PDB: 1BE3_B 1PP9_B 2A06_B 1SQB_B 1SQP_B 1L0N_B 1SQX_B 1NU1_B 1L0L_B 2FYU_B ....
Probab=98.85 E-value=2.9e-08 Score=76.87 Aligned_cols=80 Identities=19% Similarity=0.207 Sum_probs=62.3
Q ss_pred HHHHHHHHchHHHHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHh-cCCHHHHHHH
Q 028703 5 LQLLALIAKQPAFHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLY-EMTSDQFKNN 83 (205)
Q Consensus 5 ~~Ll~~ils~~~f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~-~ls~eeF~~~ 83 (205)
..+|...+++++|+.||++++|||.|.++.....+...+.+.++. .|..+...++.++..+....+ ++++++|+.+
T Consensus 104 ~~~l~~~~~s~l~~~lr~~~~l~y~v~~~~~~~~~~~~~~i~~~~---~~~~~~~~~~~~~~~l~~l~~~~~s~~el~~~ 180 (184)
T PF05193_consen 104 SSLLGNGMSSRLFQELREKQGLAYSVSASNSSYRDSGLFSISFQV---TPENLDEAIEAILQELKRLREGGISEEELERA 180 (184)
T ss_dssp HHHHHCSTTSHHHHHHHTTTTSESEEEEEEEEESSEEEEEEEEEE---EGGGHHHHHHHHHHHHHHHHHHCS-HHHHHHH
T ss_pred HHHHhcCccchhHHHHHhccccceEEEeeeeccccceEEEEEEEc---CcccHHHHHHHHHHHHHHHHHcCCCHHHHHHH
Confidence 344444445569999999999999999998888877778888874 444777777778777777776 4999999999
Q ss_pred HHHH
Q 028703 84 VNAL 87 (205)
Q Consensus 84 k~~l 87 (205)
|+.|
T Consensus 181 k~~L 184 (184)
T PF05193_consen 181 KNQL 184 (184)
T ss_dssp HHHH
T ss_pred HhcC
Confidence 9876
No 6
>COG0612 PqqL Predicted Zn-dependent peptidases [General function prediction only]
Probab=98.71 E-value=5.7e-07 Score=80.91 Aligned_cols=151 Identities=15% Similarity=0.105 Sum_probs=122.7
Q ss_pred HHHHHHHHchHHHHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHh-cCCHHHHHHH
Q 028703 5 LQLLALIAKQPAFHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLY-EMTSDQFKNN 83 (205)
Q Consensus 5 ~~Ll~~ils~~~f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~-~ls~eeF~~~ 83 (205)
..+++...++++|..+|.+++|-|.|++........+++.+++......++.....|.+-+..+...+. .+++++++..
T Consensus 281 ~~llgg~~~SrLf~~~re~~glay~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~i~~~~~~~~~~~~~~~t~~~~~~~ 360 (438)
T COG0612 281 NGLLGGGFSSRLFQELREKRGLAYSVSSFSDFLSDSGLFSIYAGTAPENPEKTAELVEEILKALKKGLKGPFTEEELDAA 360 (438)
T ss_pred HHHhCCCcchHHHHHHHHhcCceeeeccccccccccCCceEEEEecCCChhhHHHHHHHHHHHHHHHhccCCCHHHHHHH
Confidence 344444566899999999999999999988888888888888888778888889988888888776654 5899999999
Q ss_pred HHHHHHHHhccCcChHHHHHHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHHHHhhhcCCCCccEEEEEEeeCCC
Q 028703 84 VNALIDMKLEKHKNLKEESGFYWREISDGILKFDRREVEVAALRQLTQQELIYFFNENIKAGAPRKKTLSVRVYGSLH 161 (205)
Q Consensus 84 k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~~~~~~~~~~~~~~l~i~v~~~~~ 161 (205)
|..+...+...-.+....+..++.....+ ......+...+.|+.+|.+|+.++.+.++. ..+ .++.+.|+..
T Consensus 361 k~~~~~~~~~~~~s~~~~~~~~~~~~~~~-~~~~~~~~~~~~i~~vt~~dv~~~a~~~~~--~~~---~~~~~~~p~~ 432 (438)
T COG0612 361 KQLLIGLLLLSLDSPSSIAELLGQYLLLG-GSLITLEELLERIEAVTLEDVNAVAKKLLA--PEN---LTIVVLGPEK 432 (438)
T ss_pred HHHHHHHhhhccCCHHHHHHHHHHHHHhc-CCccCHHHHHHHHHhcCHHHHHHHHHHhcC--CCC---cEEEEEcccc
Confidence 99999999999999999999998887763 233555778889999999999999999993 222 5555556643
No 7
>KOG2067 consensus Mitochondrial processing peptidase, alpha subunit [Posttranslational modification, protein turnover, chaperones]
Probab=97.24 E-value=0.0047 Score=54.85 Aligned_cols=127 Identities=9% Similarity=-0.001 Sum_probs=99.9
Q ss_pred HchHHHHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHhcCCHHHHHHHHHHHHHHH
Q 028703 12 AKQPAFHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLYEMTSDQFKNNVNALIDMK 91 (205)
Q Consensus 12 ls~~~f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~~ls~eeF~~~k~~li~~l 91 (205)
|.+++|..+=..-+-=|...+....+.+++-++++.. .+|+.+...++-...++...-...+.+|++++|..|.+.+
T Consensus 306 MySrLY~~vLNry~wv~sctAfnhsy~DtGlfgi~~s---~~P~~a~~aveli~~e~~~~~~~v~~~el~RAK~qlkS~L 382 (472)
T KOG2067|consen 306 MYSRLYLNVLNRYHWVYSCTAFNHSYSDTGLFGIYAS---APPQAANDAVELIAKEMINMAGGVTQEELERAKTQLKSML 382 (472)
T ss_pred hHHHHHHHHHhhhHHHHHhhhhhccccCCceeEEecc---CCHHHHHHHHHHHHHHHHHHhCCCCHHHHHHHHHHHHHHH
Confidence 5567777777777778899999999999988888877 5799999999999999888888899999999999888887
Q ss_pred hccCcChHHHHHHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHHHHhh
Q 028703 92 LEKHKNLKEESGFYWREISDGILKFDRREVEVAALRQLTQQELIYFFNENI 142 (205)
Q Consensus 92 ~~~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~~~~~ 142 (205)
+-.-.|=--.++-.-.+|+... .--..++.++.|+++|.+|+.++....+
T Consensus 383 lMNLESR~V~~EDvGRQVL~~g-~rk~p~e~~~~Ie~lt~~DI~rva~kvl 432 (472)
T KOG2067|consen 383 LMNLESRPVAFEDVGRQVLTTG-ERKPPDEFIKKIEQLTPSDISRVASKVL 432 (472)
T ss_pred HhcccccchhHHHHhHHHHhcc-CcCCHHHHHHHHHhcCHHHHHHHHHHHh
Confidence 6433332222333445566542 2245688999999999999999999988
No 8
>COG1026 Predicted Zn-dependent peptidases, insulinase-like [General function prediction only]
Probab=96.85 E-value=0.022 Score=55.82 Aligned_cols=146 Identities=13% Similarity=0.160 Sum_probs=104.6
Q ss_pred chHHHHHHHHHch-HHHHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHh-cCCHHH
Q 028703 2 NVKLQLLALIAKQ-PAFHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLY-EMTSDQ 79 (205)
Q Consensus 2 ~a~~~Ll~~ils~-~~f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~-~ls~ee 79 (205)
.+.+.+++.++.. +|.+.+|++ +..|-.++......|+.+++.| .+|+ +....+.|.+....... ++++.+
T Consensus 810 ~~~l~vls~~L~~~~lw~~IR~~-GGAYGa~as~~~~~G~f~f~sY-----RDPn-~~kt~~v~~~~v~~l~s~~~~~~d 882 (978)
T COG1026 810 YAALQVLSEYLGSGYLWNKIREK-GGAYGASASIDANRGVFSFASY-----RDPN-ILKTYKVFRKSVKDLASGNFDERD 882 (978)
T ss_pred chHHHHHHHHhccchhHHHHHhh-ccccccccccccCCCeEEEEec-----CCCc-HHHHHHHHHHHHHHHHcCCCCHHH
Confidence 5778888887765 899999986 7888888888888888877766 5554 45667777776666665 899999
Q ss_pred HHHHHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHHHHhhhcCCCCccEEEEEEeeC
Q 028703 80 FKNNVNALIDMKLEKHKNLKEESGFYWREISDGILKFDRREVEVAALRQLTQQELIYFFNENIKAGAPRKKTLSVRVYGS 159 (205)
Q Consensus 80 F~~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~~~~~~~~~~~~~~l~i~v~~~ 159 (205)
.++++-+.++.+-.+-.+-..-+..+......-.. +.++...++|.++|++|+.+..+.++.+ -...-++.+.|.
T Consensus 883 ~~~~ilg~i~~~d~p~sp~~~~~~s~~~~~sg~~~--~~~qa~re~~l~vt~~di~~~~~~yl~~---~~~e~~i~~~~~ 957 (978)
T COG1026 883 LEEAILGIISTLDTPESPASEGSKSFYRDLSGLTD--EERQAFRERLLDVTKEDIKEVMDKYLLN---FSSENSIAVFAG 957 (978)
T ss_pred HHHHHHHhhcccccccCCcceehhhHHHHHhcCCH--HHHHHHHHHHhcCcHHHHHHHHHHHHhc---ccccceEEEEec
Confidence 99999999999876544433333333333333232 5668888999999999999999998842 222334555554
No 9
>PTZ00432 falcilysin; Provisional
Probab=96.75 E-value=0.036 Score=55.92 Aligned_cols=146 Identities=12% Similarity=-0.014 Sum_probs=105.1
Q ss_pred chHHHHHHHHHc-hHHHHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHh---cCCH
Q 028703 2 NVKLQLLALIAK-QPAFHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLY---EMTS 77 (205)
Q Consensus 2 ~a~~~Ll~~ils-~~~f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~---~ls~ 77 (205)
.+.+.++.++|+ .++++.+|.+ +..|-+++.... .|...+.-| .+|. +...++.|-....-... ++++
T Consensus 955 ~~~l~Vl~~~L~~~yLw~~IR~~-GGAYG~~~~~~~-~G~~~f~SY-----RDPn-~~~Tl~~f~~~~~~l~~~~~~~~~ 1026 (1119)
T PTZ00432 955 DGSFQVIVHYLKNSYLWKTVRMS-LGAYGVFADLLY-TGHVIFMSY-----ADPN-FEKTLEVYKEVASALREAAETLTD 1026 (1119)
T ss_pred CHHHHHHHHHHccccchHHHccc-CCccccCCccCC-CCeEEEEEe-----cCCC-HHHHHHHHHHHHHHHHhcCCCCCH
Confidence 356788888875 6899999976 667777755543 233333322 4444 55778888766554333 4999
Q ss_pred HHHHHHHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHHHHhhhcCCCCccEEEEEEe
Q 028703 78 DQFKNNVNALIDMKLEKHKNLKEESGFYWREISDGILKFDRREVEVAALRQLTQQELIYFFNENIKAGAPRKKTLSVRVY 157 (205)
Q Consensus 78 eeF~~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~~~~~~~~~~~~~~l~i~v~ 157 (205)
+++++++-+.++.+- +|.+-..++.+....+.+| ...+.+++..+.|-+.|++|+.++...+... .. .-.++|.
T Consensus 1027 ~~l~~~iig~~~~~D-~p~~p~~~g~~~~~~~l~g-~t~e~rq~~R~~il~~t~edi~~~a~~~~~~--~~--~~~~~v~ 1100 (1119)
T PTZ00432 1027 KDLLRYKIGKISNID-KPLHVDELSKLALLRIIRN-ESDEDRQKFRKDILETTKEDFYRLADLMEKS--KE--WEKVIAV 1100 (1119)
T ss_pred HHHHHHHHHHHhccC-CCCChHHHHHHHHHHHHcC-CCHHHHHHHHHHHHcCCHHHHHHHHHHHHhh--hc--cCeEEEE
Confidence 999999999999864 5788888888888877775 3457789999999999999999999998832 22 3356666
Q ss_pred eCCC
Q 028703 158 GSLH 161 (205)
Q Consensus 158 ~~~~ 161 (205)
|...
T Consensus 1101 g~~~ 1104 (1119)
T PTZ00432 1101 VNSK 1104 (1119)
T ss_pred ECHH
Confidence 7654
No 10
>PRK15101 protease3; Provisional
Probab=96.54 E-value=0.017 Score=57.10 Aligned_cols=134 Identities=7% Similarity=-0.104 Sum_probs=80.9
Q ss_pred HHHHHHHHchHHHHHhhhccccceEEEEEEeee--CCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHH-hcCCHHHHH
Q 028703 5 LQLLALIAKQPAFHQLRTVEQLGYITALLQRND--FGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKL-YEMTSDQFK 81 (205)
Q Consensus 5 ~~Ll~~ils~~~f~~LRTkqQLGYvV~s~~~~~--~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L-~~ls~eeF~ 81 (205)
..+|++-....++..|+ +++|.|.|+++.... .+...+.+.++......+.+...++.++..+.... .++++++|+
T Consensus 312 ~~ll~~~~~g~l~~~L~-~~gla~~v~s~~~~~~~~~~g~f~i~~~~~~~~~~~~~~v~~~i~~~i~~l~~~g~~~~el~ 390 (961)
T PRK15101 312 SYLIGNRSPGTLSDWLQ-KQGLAEGISAGADPMVDRNSGVFAISVSLTDKGLAQRDQVVAAIFSYLNLLREKGIDKSYFD 390 (961)
T ss_pred HHHhcCCCCCcHHHHHH-HcCccceeeeccccccCCCceEEEEEEEcChHHHHhHHHHHHHHHHHHHHHHhcCCcHHHHH
Confidence 34445545556888886 899999999986643 34555667777432222355666666666665544 379999999
Q ss_pred HHHHHHHHHHhccC-cChHHHHHHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHHHHhh
Q 028703 82 NNVNALIDMKLEKH-KNLKEESGFYWREISDGILKFDRREVEVAALRQLTQQELIYFFNENI 142 (205)
Q Consensus 82 ~~k~~li~~l~~~~-~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~~~~~ 142 (205)
..|+.+.....-.. .+-.+.+...-..+ ..+.+..-......++.++.+++.+.++. +
T Consensus 391 ~~k~~~~~~~~~~~~~~~~~~~~~~~~~~--~~~~~~~~l~~~~~~~~~~~~~i~~~~~~-l 449 (961)
T PRK15101 391 ELAHVLDLDFRYPSITRDMDYIEWLADTM--LRVPVEHTLDAPYIADRYDPKAIKARLAE-M 449 (961)
T ss_pred HHHHHHhccccCCCCCChHHHHHHHHHHh--hhCCHHHheeCchhhhcCCHHHHHHHHhh-c
Confidence 99998877653222 11122333332222 12222211123345788999999999887 5
No 11
>PTZ00432 falcilysin; Provisional
Probab=96.52 E-value=0.053 Score=54.73 Aligned_cols=149 Identities=12% Similarity=0.025 Sum_probs=88.1
Q ss_pred HHHHHHHHchHHHHHhhhccccceEE-EEEEeeeCCeeEEEEEEeCCC-CC----hhHHHHHHHHHHHHHHHHH-hcCCH
Q 028703 5 LQLLALIAKQPAFHQLRTVEQLGYIT-ALLQRNDFGIHGVQFIIQSSV-KG----PKYIDLRVESFLQMFESKL-YEMTS 77 (205)
Q Consensus 5 ~~Ll~~ils~~~f~~LRTkqQLGYvV-~s~~~~~~~~~gl~~~VQS~~-~~----~~~l~~~i~~Fl~~~~~~L-~~ls~ 77 (205)
..+|.+-.++|+|..|| +.+|||.| +++.....+.+.+.+.+++.. .. ++.+..-++...+.+.... +++++
T Consensus 408 s~lLggg~sS~L~q~Lr-E~GLa~svv~~~~~~~~~~~~f~I~l~g~~~~~~~~~~~~~~ev~~~I~~~L~~l~~eGi~~ 486 (1119)
T PTZ00432 408 NYLLLGTPESVLYKALI-DSGLGKKVVGSGLDDYFKQSIFSIGLKGIKETNEKRKDKVHYTFEKVVLNALTKVVTEGFNK 486 (1119)
T ss_pred HHHHcCCCccHHHHHHH-hcCCCcCCCcCcccCCCCceEEEEEEEcCChHhccchhhhHHHHHHHHHHHHHHHHHhCCCH
Confidence 34444455899999999 58999996 556555666677777877421 11 1223333333333333333 37999
Q ss_pred HHHHHHHHHHHHHHhccCcC-------hHHHHHHhHHHHhcCCCCcc--ccHHHHHHHh-cC--CHHHHHHHHHHhhhcC
Q 028703 78 DQFKNNVNALIDMKLEKHKN-------LKEESGFYWREISDGILKFD--RREVEVAALR-QL--TQQELIYFFNENIKAG 145 (205)
Q Consensus 78 eeF~~~k~~li~~l~~~~~s-------l~~~~~~~w~~I~~~~~~F~--~~~~~i~~l~-~i--t~~dl~~f~~~~~~~~ 145 (205)
++++..++.+.-.+++...+ +...+...|. .+.--++ .-+..++.|+ ++ +...+.++.+++|.
T Consensus 487 eele~a~~qlef~~rE~~~~~~p~gl~~~~~~~~~~~---~g~dp~~~l~~~~~l~~lr~~~~~~~~y~e~Li~k~ll-- 561 (1119)
T PTZ00432 487 SAVEASLNNIEFVMKELNLGTYPKGLMLIFLMQSRLQ---YGKDPFEILRFEKLLNELKLRIDNESKYLEKLIEKHLL-- 561 (1119)
T ss_pred HHHHHHHHHHHHHhhhccCCCCCcHHHHHHHHHHHHh---cCCCHHHHHhhHHHHHHHHHHHhcccHHHHHHHHHHcc--
Confidence 99999999998888876432 3334444442 2211222 1234455554 23 22468888999883
Q ss_pred CCCccEEEEEEeeCC
Q 028703 146 APRKKTLSVRVYGSL 160 (205)
Q Consensus 146 ~~~~~~l~i~v~~~~ 160 (205)
.|..++.+.+.+..
T Consensus 562 -~N~h~~~v~~~p~~ 575 (1119)
T PTZ00432 562 -NNNHRVTVHLEAVE 575 (1119)
T ss_pred -CCCeeeEEEEecCC
Confidence 44456777776664
No 12
>TIGR02110 PQQ_syn_pqqF coenzyme PQQ biosynthesis probable peptidase PqqF. In a subset of species that make coenzyme PQQ (pyrrolo-quinoline-quinone), this probable peptidase is found in the PQQ biosynthesis region and is thought to act as a protease on PqqA (TIGR02107), a probable peptide precursor of the coenzyme. PQQ is required for some glucose dehydrogenases and alcohol dehydrogenases.
Probab=95.93 E-value=0.029 Score=53.77 Aligned_cols=83 Identities=12% Similarity=0.061 Sum_probs=57.9
Q ss_pred HHHHHHHHchHHHHHhhhccccceEEEEEEeeeCCee--EEEEEEeC---CCCChhHHHHHHHHHHHHHHHHHhcCCHHH
Q 028703 5 LQLLALIAKQPAFHQLRTVEQLGYITALLQRNDFGIH--GVQFIIQS---SVKGPKYIDLRVESFLQMFESKLYEMTSDQ 79 (205)
Q Consensus 5 ~~Ll~~ils~~~f~~LRTkqQLGYvV~s~~~~~~~~~--gl~~~VQS---~~~~~~~l~~~i~~Fl~~~~~~L~~ls~ee 79 (205)
..+|+.-+++.+|.+|| +++|.|.|+++. ...+.. .+.+++.. +..+.+.+...|.+.+..+.+..-..+.+|
T Consensus 261 ~~iLg~g~sSrL~~~LR-e~GLaysV~s~~-~~~~~g~~lf~I~~~lt~~~~~~~~~v~~~i~~~L~~L~~~~~~~~~ee 338 (696)
T TIGR02110 261 CEFLQDEAPGGLLAQLR-ERGLAESVAATW-LYQDAGQALLALEFSARCISAAAAQQIEQLLTQWLGALAEQTWAEQLEH 338 (696)
T ss_pred HHHhCCCcchHHHHHHH-HCCCEEEEEEec-cccCCCCcEEEEEEEEcCCCccCHHHHHHHHHHHHHHHHhcCCCCCHHH
Confidence 34444455678999999 589999999965 344333 45556653 234677788888888877765544788999
Q ss_pred HHHHHHHHHH
Q 028703 80 FKNNVNALID 89 (205)
Q Consensus 80 F~~~k~~li~ 89 (205)
+++.|+.=..
T Consensus 339 l~rlk~~~~~ 348 (696)
T TIGR02110 339 YAQLAQRRFQ 348 (696)
T ss_pred HHHHHHhhhh
Confidence 9999876433
No 13
>COG0612 PqqL Predicted Zn-dependent peptidases [General function prediction only]
Probab=95.07 E-value=0.26 Score=44.26 Aligned_cols=98 Identities=10% Similarity=0.159 Sum_probs=71.8
Q ss_pred HHHHHHHHHHHHHHh--cCCHHHHHHHHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCcccc-HHHHHHHhcCCHHHHH
Q 028703 59 LRVESFLQMFESKLY--EMTSDQFKNNVNALIDMKLEKHKNLKEESGFYWREISDGILKFDRR-EVEVAALRQLTQQELI 135 (205)
Q Consensus 59 ~~i~~Fl~~~~~~L~--~ls~eeF~~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~-~~~i~~l~~it~~dl~ 135 (205)
..++.-+.-+.+.+. .+++++|+.-|..++..+.....+-...+...|.+...++--+.+. --..+.|++||++|+.
T Consensus 108 ~~~~~~l~llad~l~~p~f~~~~~e~Ek~vil~ei~~~~d~p~~~~~~~l~~~~~~~~p~~~~~~G~~e~I~~it~~dl~ 187 (438)
T COG0612 108 DNLDKALDLLADILLNPTFDEEEVEREKGVILEEIRMRQDDPDDLAFERLLEALYGNHPLGRPILGTEESIEAITREDLK 187 (438)
T ss_pred hhhHHHHHHHHHHHhCCCCCHHHHHHHHHHHHHHHHhhccCchHHHHHHHHHHhhccCCCCCCCCCCHHHHHhCCHHHHH
Confidence 333444444445553 4899999999999999999988888888888888877765444442 1245679999999999
Q ss_pred HHHHHhhhcCCCCccEEEEEEeeCCC
Q 028703 136 YFFNENIKAGAPRKKTLSVRVYGSLH 161 (205)
Q Consensus 136 ~f~~~~~~~~~~~~~~l~i~v~~~~~ 161 (205)
+||++++. .. ...|-|.|.-.
T Consensus 188 ~f~~k~Y~--p~---n~~l~vvGdi~ 208 (438)
T COG0612 188 DFYQKWYQ--PD---NMVLVVVGDVD 208 (438)
T ss_pred HHHHHhcC--cC---ceEEEEecCCC
Confidence 99999993 22 36777778754
No 14
>KOG0960 consensus Mitochondrial processing peptidase, beta subunit, and related enzymes (insulinase superfamily) [Posttranslational modification, protein turnover, chaperones]
Probab=94.83 E-value=0.51 Score=42.27 Aligned_cols=116 Identities=16% Similarity=0.138 Sum_probs=84.4
Q ss_pred eeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHhcCCHHHHHHHHHHHHHHHhccCcChHHHHHHhHHHHhcCCC
Q 028703 35 RNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLYEMTSDQFKNNVNALIDMKLEKHKNLKEESGFYWREISDGIL 114 (205)
Q Consensus 35 ~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~~ls~eeF~~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~ 114 (205)
+...|-.|+.|+.- ++..+..-|..-+.++...-...||+|.+..|+.|..++...-..-..-++-.-.+++..+.
T Consensus 334 YkDTGLwG~y~V~~----~~~~iddl~~~vl~eW~rL~~~vteaEV~RAKn~Lkt~Lll~ldgttpi~ediGrqlL~~Gr 409 (467)
T KOG0960|consen 334 YKDTGLWGIYFVTD----NLTMIDDLIHSVLKEWMRLATSVTEAEVERAKNQLKTNLLLSLDGTTPIAEDIGRQLLTYGR 409 (467)
T ss_pred cccccceeEEEEec----ChhhHHHHHHHHHHHHHHHHhhccHHHHHHHHHHHHHHHHHHhcCCCchHHHHHHHHhhcCC
Confidence 33344555555532 77888888888888887655689999999999999999987665555557777777777554
Q ss_pred CccccHHHHHHHhcCCHHHHHHHHHHhhhcCCCCccEEEEEEeeCC
Q 028703 115 KFDRREVEVAALRQLTQQELIYFFNENIKAGAPRKKTLSVRVYGSL 160 (205)
Q Consensus 115 ~F~~~~~~i~~l~~it~~dl~~f~~~~~~~~~~~~~~l~i~v~~~~ 160 (205)
.-.. .+.-+.|.+||.+++.++..+++- -+.+++-..|+.
T Consensus 410 ri~l-~El~~rId~vt~~~Vr~va~k~iy-----d~~iAia~vG~i 449 (467)
T KOG0960|consen 410 RIPL-AELEARIDAVTAKDVREVASKYIY-----DKDIAIAAVGPI 449 (467)
T ss_pred cCCh-HHHHHHHhhccHHHHHHHHHHHhh-----cCCcceeeeccc
Confidence 4444 445567899999999999999982 234667776774
No 15
>KOG0961 consensus Predicted Zn2+-dependent endopeptidase, insulinase superfamily [Posttranslational modification, protein turnover, chaperones]
Probab=94.31 E-value=0.52 Score=45.04 Aligned_cols=148 Identities=11% Similarity=0.108 Sum_probs=101.9
Q ss_pred hHHHHHHHHH---chHHHHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHhcCCHHH
Q 028703 3 VKLQLLALIA---KQPAFHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLYEMTSDQ 79 (205)
Q Consensus 3 a~~~Ll~~il---s~~~f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~~ls~ee 79 (205)
+...|+++.+ ..||..-+|-- +|.|-.......-++..|++++-- .+|..-.++-....+.+..---++++.+
T Consensus 831 ~~~~l~~~YL~~~eGPfW~~IRG~-GLAYGanm~~~~d~~~~~~~iyr~---ad~~kaye~~rdiV~~~vsG~~e~s~~~ 906 (1022)
T KOG0961|consen 831 IPAMLFGQYLSQCEGPFWRAIRGD-GLAYGANMFVKPDRKQITLSIYRC---ADPAKAYERTRDIVRKIVSGSGEISKAE 906 (1022)
T ss_pred hHHHHHHHHHHhcccchhhhhccc-chhccceeEEeccCCEEEEEeecC---CcHHHHHHHHHHHHHHHhcCceeecHHH
Confidence 4567777754 56999999975 899999999999999999999854 4666666777776665544334699999
Q ss_pred HHHHHHHHHHHHhccCcC-hHHHHHHhHH-HHhcCCCCccccHHHHHHHhcCCHHHHHHHHHHhhhcC-CCCccEEEEEE
Q 028703 80 FKNNVNALIDMKLEKHKN-LKEESGFYWR-EISDGILKFDRREVEVAALRQLTQQELIYFFNENIKAG-APRKKTLSVRV 156 (205)
Q Consensus 80 F~~~k~~li~~l~~~~~s-l~~~~~~~w~-~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~~~~~~~~-~~~~~~l~i~v 156 (205)
|+.+|.+.+..+.+...+ +..-++.+-- .+....-+|+ ....+.|.++|++|+++-...++.+. ++++..-+|.+
T Consensus 907 ~egAk~s~~~~~~~~Eng~~~~a~~~~~l~~~~q~~~~fn--~~~leri~nvT~~~~~~~~~~y~~~~Fds~~~va~i~~ 984 (1022)
T KOG0961|consen 907 FEGAKRSTVFEMMKRENGTVSGAAKISILNNFRQTPHPFN--IDLLERIWNVTSEEMVKIGGPYLARLFDSKCFVASIAV 984 (1022)
T ss_pred hccchHHHHHHHHHHhccceechHHHHHHHHHHhcCCccc--HHHHHHHHHhhHHHHHHhcccceehhhcccCceEEEec
Confidence 999999998888776533 3333333332 3444455555 45778899999999988665554332 33433444444
No 16
>KOG0959 consensus N-arginine dibasic convertase NRD1 and related Zn2+-dependent endopeptidases, insulinase superfamily [Posttranslational modification, protein turnover, chaperones]
Probab=93.66 E-value=2 Score=42.80 Aligned_cols=114 Identities=18% Similarity=0.158 Sum_probs=75.0
Q ss_pred eeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHhcCCHHHHHHHHHHHHHHHhc-cCcChHHHHHHhHHHHhcCCC
Q 028703 36 NDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLYEMTSDQFKNNVNALIDMKLE-KHKNLKEESGFYWREISDGIL 114 (205)
Q Consensus 36 ~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~~ls~eeF~~~k~~li~~l~~-~~~sl~~~~~~~w~~I~~~~~ 114 (205)
...+..|+.+.|-+=...-.-+...+.+++..|. ++++.|+..++.+...+.- ...+-...+..+-..+ ....
T Consensus 582 ~~~s~~G~~~~v~Gfnekl~~ll~~~~~~~~~f~-----~~~~rf~iike~~~~~~~n~~~~~p~~~a~~~~~ll-l~~~ 655 (974)
T KOG0959|consen 582 LSSSSKGVELRVSGFNEKLPLLLEKVVQMMANFE-----LDEDRFEIIKELLKRELRNHAFDNPYQLANDYLLLL-LEES 655 (974)
T ss_pred eeecCCceEEEEeccCcccHHHHHHHHHHHHhcc-----ccHHHHHHHHHHHHHHHhhhhhccHHHHHHHHHHHH-hhcc
Confidence 3345577777777622222223333333444333 8889999999988888776 3333444454444444 4466
Q ss_pred CccccHHHHHHHhcCCHHHHHHHHHHhhhcCCCCccEEEEEEeeCCC
Q 028703 115 KFDRREVEVAALRQLTQQELIYFFNENIKAGAPRKKTLSVRVYGSLH 161 (205)
Q Consensus 115 ~F~~~~~~i~~l~~it~~dl~~f~~~~~~~~~~~~~~l~i~v~~~~~ 161 (205)
.|.. +..+++++.+|.+++..|...+. . +.-+...|.|.-.
T Consensus 656 ~W~~-~e~~~al~~~~le~~~~F~~~~~---~--~~~~e~~i~GN~t 696 (974)
T KOG0959|consen 656 IWSK-EELLEALDDVTLEDLESFISEFL---Q--PFHLELLIHGNLT 696 (974)
T ss_pred ccch-HHHHHHhhcccHHHHHHHHHHHh---h--hhheEEEEecCcc
Confidence 7777 66788999999999999999987 2 4468888888854
No 17
>COG1025 Ptr Secreted/periplasmic Zn-dependent peptidases, insulinase-like [Posttranslational modification, protein turnover, chaperones]
Probab=93.30 E-value=2.7 Score=41.52 Aligned_cols=118 Identities=15% Similarity=0.138 Sum_probs=79.1
Q ss_pred eeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHhcCCHHHHHHHHHHHHHHHhccCcChHHHHHHhHHHHhcCCC
Q 028703 35 RNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLYEMTSDQFKNNVNALIDMKLEKHKNLKEESGFYWREISDGIL 114 (205)
Q Consensus 35 ~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~~ls~eeF~~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~ 114 (205)
....+..|+.+.+-+- ++.+..-++.|+..+.. ...+++.|+.+|..+...+.......--+....--.+...-.
T Consensus 574 s~~~~~~Gl~ltisGf---t~~lp~L~~~~l~~l~~--~~~~~~~f~~~K~~~~~~~~~a~~~~p~~~~~~~l~~l~~~~ 648 (937)
T COG1025 574 SLAANSNGLDLTISGF---TQRLPQLLRAFLDGLFS--LPVDEDRFEQAKSQLSEELKNALTGKPYRQALDGLTGLLQVP 648 (937)
T ss_pred EeecCCCceEEEeecc---ccchHHHHHHHHHHHhc--CCCCHHHHHHHHHHHHHHHHhhhhcCCHHHHHHHhhhhhCCC
Confidence 3334458899998864 34455555566644321 234589999999999988876554433333333333444556
Q ss_pred CccccHHHHHHHhcCCHHHHHHHHHHhhhcCCCCccEEEEEEeeCCCCc
Q 028703 115 KFDRREVEVAALRQLTQQELIYFFNENIKAGAPRKKTLSVRVYGSLHAP 163 (205)
Q Consensus 115 ~F~~~~~~i~~l~~it~~dl~~f~~~~~~~~~~~~~~l~i~v~~~~~~~ 163 (205)
.+.. +...++|++++.+++..|...++ +...+.+.|.|.-...
T Consensus 649 ~~s~-~e~~~~l~~v~~~e~~~f~~~l~-----~~~~lE~lv~Gn~~~~ 691 (937)
T COG1025 649 YWSR-EERRNALESVSVEEFAAFRDTLL-----NGVHLEMLVLGNLTEA 691 (937)
T ss_pred CcCH-HHHHHHhhhccHHHHHHHHHHhh-----hccceeeeeeccchHH
Confidence 6666 66788999999999999999988 2336888999987643
No 18
>KOG2681 consensus Metal-dependent phosphohydrolase [Function unknown]
Probab=86.63 E-value=0.72 Score=41.70 Aligned_cols=111 Identities=15% Similarity=0.106 Sum_probs=67.1
Q ss_pred HHHHHHHchHHHHHhhhccccc--eEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHh-cCCHHHHHH
Q 028703 6 QLLALIAKQPAFHQLRTVEQLG--YITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLY-EMTSDQFKN 82 (205)
Q Consensus 6 ~Ll~~ils~~~f~~LRTkqQLG--YvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~-~ls~eeF~~ 82 (205)
.++..+++.+.|++||-.+||| |.|+.+..-.+....+..+ +.+..+.+++..+- -.+ .+|+-+...
T Consensus 40 pli~~lidt~~FqRLr~vkQlGl~~~vyp~A~HsRfeHsLG~~-----~lA~~~v~~L~~~q-----~~El~It~~d~~~ 109 (498)
T KOG2681|consen 40 PLIIKLIDTPLFQRLRHVKQLGLRYLVYPGANHSRFEHSLGTY-----TLAGILVNALNKNQ-----CPELCITEVDLQA 109 (498)
T ss_pred hHHHHHhccHHHHHHHHHHHhCceeeeccCCccchhhhhhhhH-----HHHHHHHHHHhhcC-----CCCCCCCHHHHHH
Confidence 5788999999999999999987 6666655555544444433 34555555555542 122 577777766
Q ss_pred -HHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCccccHHHHHHHhc
Q 028703 83 -NVNALIDMKLEKHKNLKEESGFYWREISDGILKFDRREVEVAALRQ 128 (205)
Q Consensus 83 -~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~ 128 (205)
.+.+|+..+-..|-|-.-+ +-..-.+..+-.|...+..++.++.
T Consensus 110 vqvA~LLHDIGHGPfSHmFe--~~f~~~v~s~~e~~HE~~si~~i~~ 154 (498)
T KOG2681|consen 110 VQVAALLHDIGHGPFSHLFE--GEFTPMVRSGPEFYHEDMSIDMIKK 154 (498)
T ss_pred HHHHHHHhhcCCCchhhhhh--heecccccCCcccchhhhHHHHHHH
Confidence 4689999998888542211 1112222234455555555554444
No 19
>PF09568 RE_MjaI: MjaI restriction endonuclease; InterPro: IPR019068 There are four classes of restriction endonucleases: types I, II,III and IV. All types of enzymes recognise specific short DNA sequences and carry out the endonucleolytic cleavage of DNA to give specific double-stranded fragments with terminal 5'-phosphates. They differ in their recognition sequence, subunit composition, cleavage position, and cofactor requirements [, ], as summarised below: Type I enzymes (3.1.21.3 from EC) cleave at sites remote from recognition site; require both ATP and S-adenosyl-L-methionine to function; multifunctional protein with both restriction and methylase (2.1.1.72 from EC) activities. Type II enzymes (3.1.21.4 from EC) cleave within or at short specific distances from recognition site; most require magnesium; single function (restriction) enzymes independent of methylase. Type III enzymes (3.1.21.5 from EC) cleave at sites a short distance from recognition site; require ATP (but doesn't hydrolyse it); S-adenosyl-L-methionine stimulates reaction but is not required; exists as part of a complex with a modification methylase methylase (2.1.1.72 from EC). Type IV enzymes target methylated DNA. Type II restriction endonucleases (3.1.21.4 from EC) are components of prokaryotic DNA restriction-modification mechanisms that protect the organism against invading foreign DNA. These site-specific deoxyribonucleases catalyse the endonucleolytic cleavage of DNA to give specific double-stranded fragments with terminal 5'-phosphates. Of the 3000 restriction endonucleases that have been characterised, most are homodimeric or tetrameric enzymes that cleave target DNA at sequence-specific sites close to the recognition site. For homodimeric enzymes, the recognition site is usually a palindromic sequence 4-8 bp in length. Most enzymes require magnesium ions as a cofactor for catalysis. Although they can vary in their mode of recognition, many restriction endonucleases share a similar structural core comprising four beta-strands and one alpha-helix, as well as a similar mechanism of cleavage, suggesting a common ancestral origin []. However, there is still considerable diversity amongst restriction endonucleases [, ]. The target site recognition process triggers large conformational changes of the enzyme and the target DNA, leading to the activation of the catalytic centres. Like other DNA binding proteins, restriction enzymes are capable of non-specific DNA binding as well, which is the prerequisite for efficient target site location by facilitated diffusion. Non-specific binding usually does not involve interactions with the bases but only with the DNA backbone []. This entry includes the MjaI (recognises CTAG but cleavage site unknown) restriction endonuclease. ; GO: 0003677 DNA binding, 0009036 Type II site-specific deoxyribonuclease activity, 0009307 DNA restriction-modification system
Probab=86.14 E-value=2 Score=34.04 Aligned_cols=39 Identities=15% Similarity=0.306 Sum_probs=34.5
Q ss_pred cCcChHHHHHHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHHHHhh
Q 028703 94 KHKNLKEESGFYWREISDGILKFDRREVEVAALRQLTQQELIYFFNENI 142 (205)
Q Consensus 94 ~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~~~~~ 142 (205)
.+..+.+...+.|..|.. .++++++||.+|+.+|.++++
T Consensus 45 ~~e~i~~a~~ki~~~i~e----------~~~a~~~it~ed~~~wv~dLv 83 (170)
T PF09568_consen 45 YPEAIEEATDKIYVMITE----------VKEALNKITEEDCINWVKDLV 83 (170)
T ss_pred ChHHHHHHHHHHHHHHHH----------HHHHHHhCCHHHHHHHHHHhe
Confidence 788899999999999854 557799999999999999987
No 20
>KOG0960 consensus Mitochondrial processing peptidase, beta subunit, and related enzymes (insulinase superfamily) [Posttranslational modification, protein turnover, chaperones]
Probab=73.87 E-value=24 Score=31.99 Aligned_cols=69 Identities=10% Similarity=0.083 Sum_probs=55.1
Q ss_pred cCCHHHHHHHHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCcccc-HHHHHHHhcCCHHHHHHHHHHhh
Q 028703 74 EMTSDQFKNNVNALIDMKLEKHKNLKEESGFYWREISDGILKFDRR-EVEVAALRQLTQQELIYFFNENI 142 (205)
Q Consensus 74 ~ls~eeF~~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~-~~~i~~l~~it~~dl~~f~~~~~ 142 (205)
.+++.+++.-+..++..+.+-++++.+..--+-+....++--..+. ---++-|++|+++|+.+|..+++
T Consensus 141 ~L~~s~IerER~vILrEmqevd~~~~eVVfdhLHatafQgtPL~~tilGp~enI~si~r~DL~~yi~thY 210 (467)
T KOG0960|consen 141 KLEESAIERERDVILREMQEVDKNHQEVVFDHLHATAFQGTPLGRTILGPSENIKSISRADLKDYINTHY 210 (467)
T ss_pred ccchhHHHHHHHHHHHHHHHHHhhhhHHHHHHHHHHHhcCCcccccccChhhhhhhhhHHHHHHHHHhcc
Confidence 5999999999999999999999998887777777666544333321 22466799999999999999998
No 21
>cd08305 Pyrin Pyrin: a protein-protein interaction domain. The Pyrin domain (or PYD), also called DAPIN or PAAD, is a subfamily of the Death Domain (DD) superfamily and it functions in several signaling pathways. The Pyrin domain is found at the N-terminus of a variety of proteins and serves as a linker that recruits other domains into signaling complexes. Pyrin-containing proteins include NALPs, ASC (Apoptosis-associated speck-like protein containing a CARD), and the interferon-inducible p200 (IFI-200) family of proteins which includes the human IFI-16, myeloid cell nuclear differentiation antigen (MNDA) and absent in melanoma (AIM) 2. NALPs are members of the NBS-LRR family of proteins possessing a tripartite domain structure including a C-terminal LRR (leucine-rich repeats), a central nucleotide-binding site (NBS) domain or NACHT (for neuronal apoptosis inhibitor protein, CIITA, HET-E and TP1), and an N-terminal protein-protein interaction domain, which is a Pyrin domain in the case
Probab=63.51 E-value=21 Score=24.11 Aligned_cols=67 Identities=10% Similarity=0.080 Sum_probs=39.5
Q ss_pred HHHhcCCHHHHHHHHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCcc-ccHHHHHHHhcCCHHHHHH
Q 028703 70 SKLYEMTSDQFKNNVNALIDMKLEKHKNLKEESGFYWREISDGILKFD-RREVEVAALRQLTQQELIY 136 (205)
Q Consensus 70 ~~L~~ls~eeF~~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~-~~~~~i~~l~~it~~dl~~ 136 (205)
..|++|+++||.+.|.-|...+...+....+.+..--.......|.=+ --+..+..++.|.+.|+-+
T Consensus 3 ~~Le~L~~~efk~FK~~L~~~~~~~~~~~~~~a~~~la~lL~~~y~~~~a~~~t~~i~~~m~~~dlae 70 (73)
T cd08305 3 TGLENITDEELKRFKSLLANDLFLETKAQLEYTRIQIADLMEQKFGAVSALDKLINIFEDMPLRSLAN 70 (73)
T ss_pred HHHHHcCHHHHHHHHHHHHhcCCCCCcccccccHHHHHHHHHHHcChhHHHHHHHHHHHHcChHHHHH
Confidence 468899999999999999987554454444333111111111122111 1245777778888777654
No 22
>PHA02698 hypothetical protein; Provisional
Probab=62.23 E-value=27 Score=23.95 Aligned_cols=36 Identities=14% Similarity=0.266 Sum_probs=28.6
Q ss_pred CCCChhHHHHHHHHHHHHH--HHHHhcCCHHHHHHHHH
Q 028703 50 SVKGPKYIDLRVESFLQMF--ESKLYEMTSDQFKNNVN 85 (205)
Q Consensus 50 ~~~~~~~l~~~i~~Fl~~~--~~~L~~ls~eeF~~~k~ 85 (205)
+.++|+.+...++.||+++ ...+.=+|.||.++...
T Consensus 39 ~~CsPEdMs~mLD~FLediq~ksElqLLsqEEMdELl~ 76 (89)
T PHA02698 39 PQCSPEDMSDMLDNFLEDIQYKSELQLLSQEEMDELLV 76 (89)
T ss_pred ccCCHHHHHHHHHHHHHHHHHHHHHHHhhHHHHHHHHH
Confidence 3589999999999999885 56777788888777543
No 23
>COG1026 Predicted Zn-dependent peptidases, insulinase-like [General function prediction only]
Probab=56.46 E-value=1.4e+02 Score=30.15 Aligned_cols=143 Identities=16% Similarity=0.027 Sum_probs=87.1
Q ss_pred HHHHHchHHHHHhhhccccc-eEEEEEEeeeCCeeEEEEEEeCC-CCChhHHHHHHHHHHHHHHHHHhcCCHHHHHHHHH
Q 028703 8 LALIAKQPAFHQLRTVEQLG-YITALLQRNDFGIHGVQFIIQSS-VKGPKYIDLRVESFLQMFESKLYEMTSDQFKNNVN 85 (205)
Q Consensus 8 l~~ils~~~f~~LRTkqQLG-YvV~s~~~~~~~~~gl~~~VQS~-~~~~~~l~~~i~~Fl~~~~~~L~~ls~eeF~~~k~ 85 (205)
|..--++||.+.| -|-.|| +.|...+....-..-+.+.+|+- .-..+.+.+.+-+-|+++.+ .+++.+..+..+.
T Consensus 305 Ll~~~asPl~~~l-iesglg~~~~~g~~~~~~~~~~f~v~~~gv~~ek~~~~k~lV~~~L~~l~~--~gi~~~~ie~~~~ 381 (978)
T COG1026 305 LLDSAASPLTQAL-IESGLGFADVSGSYDSDLKETIFSVGLKGVSEEKIAKLKNLVLSTLKELVK--NGIDKKLIEAILH 381 (978)
T ss_pred HccCcccHHHHHH-HHcCCCcccccceeccccceeEEEEEecCCCHHHHHHHHHHHHHHHHHHHH--hcCCHHHHHHHHH
Confidence 3333456899999 788999 66665577777777888888873 33455566666555554332 2588888888888
Q ss_pred HHHHHHhccCc-----ChHHHHHHhHHHHhcCCCCcc--ccHHHHHHH-hcCCHHH-HHHHHHHhhhcCCCCccEEEEEE
Q 028703 86 ALIDMKLEKHK-----NLKEESGFYWREISDGILKFD--RREVEVAAL-RQLTQQE-LIYFFNENIKAGAPRKKTLSVRV 156 (205)
Q Consensus 86 ~li~~l~~~~~-----sl~~~~~~~w~~I~~~~~~F~--~~~~~i~~l-~~it~~d-l~~f~~~~~~~~~~~~~~l~i~v 156 (205)
++.=++++-+. .+...+-.-|.. |.--++ +-...+..| +++++.. +.+..++||. .|...+.|.+
T Consensus 382 q~E~s~ke~~s~pfgl~l~~~~~~gw~~---G~dp~~~Lr~~~~~~~Lr~~le~~~~fe~LI~ky~l---~N~h~~~v~~ 455 (978)
T COG1026 382 QLEFSLKEVKSYPFGLGLMFRSLYGWLN---GGDPEDSLRFLDYLQNLREKLEKGPYFEKLIRKYFL---DNPHYVTVIV 455 (978)
T ss_pred HHHHhhhhhcCCCccHHHHHHhcccccc---CCChhhhhhhHHHHHHHHHhhhcChHHHHHHHHHhh---cCCccEEEEE
Confidence 88877777322 133444444543 222222 123344455 4567766 8889999883 3333455555
Q ss_pred eeC
Q 028703 157 YGS 159 (205)
Q Consensus 157 ~~~ 159 (205)
.+.
T Consensus 456 ~Ps 458 (978)
T COG1026 456 LPS 458 (978)
T ss_pred ecC
Confidence 544
No 24
>PF09851 SHOCT: Short C-terminal domain; InterPro: IPR018649 This family of hypothetical prokaryotic proteins has no known function.
Probab=55.65 E-value=23 Score=19.70 Aligned_cols=23 Identities=9% Similarity=0.335 Sum_probs=16.7
Q ss_pred HHHHHHHh--cCCHHHHHHHHHHHH
Q 028703 66 QMFESKLY--EMTSDQFKNNVNALI 88 (205)
Q Consensus 66 ~~~~~~L~--~ls~eeF~~~k~~li 88 (205)
..+..... .+|++||+..|+.++
T Consensus 6 ~~L~~l~~~G~IseeEy~~~k~~ll 30 (31)
T PF09851_consen 6 EKLKELYDKGEISEEEYEQKKARLL 30 (31)
T ss_pred HHHHHHHHcCCCCHHHHHHHHHHHh
Confidence 34444443 599999999998775
No 25
>PF06518 DUF1104: Protein of unknown function (DUF1104); InterPro: IPR009488 This family consists of several hypothetical proteins of unknown function which appear to be found exclusively in Helicobacter pylori.; PDB: 2XRH_A.
Probab=51.20 E-value=63 Score=22.99 Aligned_cols=36 Identities=14% Similarity=0.190 Sum_probs=19.0
Q ss_pred HHHHHHHHHhcCCHHHHHHHHHHHHHHHhccCcChH
Q 028703 64 FLQMFESKLYEMTSDQFKNNVNALIDMKLEKHKNLK 99 (205)
Q Consensus 64 Fl~~~~~~L~~ls~eeF~~~k~~li~~l~~~~~sl~ 99 (205)
|-..+...+..||.++|...+..+...+...-..++
T Consensus 49 ~~~~~~kn~~~ms~~e~~k~~~ev~k~~~~~~~~mS 84 (93)
T PF06518_consen 49 FKEAARKNLSKMSVEERKKRREEVRKALEKRIKKMS 84 (93)
T ss_dssp HHHHHHHHHTTS-HHHHHHHHHHHHHHHHHT----S
T ss_pred HHHHHHHHHHHCCHHHHHHHHHHHHHHHHHHHHhcc
Confidence 333444556677777777777776666665544443
No 26
>PF05193 Peptidase_M16_C: Peptidase M16 inactive domain; InterPro: IPR007863 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Metalloproteases are the most diverse of the four main types of protease, with more than 50 families identified to date. In these enzymes, a divalent cation, usually zinc, activates the water molecule. The metal ion is held in place by amino acid ligands, usually three in number. The known metal ligands are His, Glu, Asp or Lys and at least one other residue is required for catalysis, which may play an electrophillic role. Of the known metalloproteases, around half contain an HEXXH motif, which has been shown in crystallographic studies to form part of the metal-binding site []. The HEXXH motif is relatively common, but can be more stringently defined for metalloproteases as 'abXHEbbHbc', where 'a' is most often valine or threonine and forms part of the S1' subsite in thermolysin and neprilysin, 'b' is an uncharged residue, and 'c' a hydrophobic residue. Proline is never found in this site, possibly because it would break the helical structure adopted by this motif in metalloproteases []. These metallopeptidases belong to MEROPS peptidase family M16 (clan ME). They include proteins, which are classified as non-peptidase homologues either have been found experimentally to be without peptidase activity, or lack amino acid residues that are believed to be essential for the catalytic activity. The peptidases in this group of sequences include: Insulinase, insulin-degrading enzyme (3.4.24.56 from EC) Mitochondrial processing peptidase alpha subunit, (Alpha-MPP, 3.4.24.64 from EC) Pitrlysin, Protease III precursor (3.4.24.55 from EC) Nardilysin, (3.4.24.61 from EC) Ubiquinol-cytochrome C reductase complex core protein I,mitochondrial precursor (1.10.2.2 from EC) Coenzyme PQQ synthesis protein F (3.4.99 from EC) These proteins do not share many regions of sequence similarity; the most noticeable is in the N-terminal section. This region includes a conserved histidine followed, two residues later by a glutamate and another histidine. In pitrilysin, it has been shown [] that this H-x-x-E-H motif is involved in enzymatic activity; the two histidines bind zinc and the glutamate is necessary for catalytic activity. The mitochondrial processing peptidase consists of two structurally related domains. One is the active peptidase whereas the other, the C-terminal region, is inactive. The two domains hold the substrate like a clamp [].; GO: 0004222 metalloendopeptidase activity, 0008270 zinc ion binding, 0006508 proteolysis; PDB: 1BE3_B 1PP9_B 2A06_B 1SQB_B 1SQP_B 1L0N_B 1SQX_B 1NU1_B 1L0L_B 2FYU_B ....
Probab=49.18 E-value=18 Score=26.95 Aligned_cols=31 Identities=16% Similarity=0.486 Sum_probs=22.3
Q ss_pred cCCHHHHHHHHHHhhhcCCCCccEEEEEEeeCCCCc
Q 028703 128 QLTQQELIYFFNENIKAGAPRKKTLSVRVYGSLHAP 163 (205)
Q Consensus 128 ~it~~dl~~f~~~~~~~~~~~~~~l~i~v~~~~~~~ 163 (205)
+||.+++.+|+++++. +. ...+.+.|.-..+
T Consensus 1 ~it~e~l~~f~~~~y~---p~--n~~l~i~Gd~~~~ 31 (184)
T PF05193_consen 1 NITLEDLRAFYKKFYR---PS--NMTLVIVGDIDPD 31 (184)
T ss_dssp C--HHHHHHHHHHHSS---GG--GEEEEEEESSGHH
T ss_pred CCCHHHHHHHHHHhcC---cc--ceEEEEEcCccHH
Confidence 5899999999999993 22 5777888876643
No 27
>PF08006 DUF1700: Protein of unknown function (DUF1700); InterPro: IPR012963 This family contains many hypothetical bacterial proteins and two putative membrane proteins (Q6GFD0 from SWISSPROT and Q6G806 from SWISSPROT).
Probab=49.06 E-value=53 Score=25.78 Aligned_cols=31 Identities=13% Similarity=0.154 Sum_probs=24.3
Q ss_pred HHHHHHHHHHHhcCCHHHHHHHHHHHHHHHh
Q 028703 62 ESFLQMFESKLYEMTSDQFKNNVNALIDMKL 92 (205)
Q Consensus 62 ~~Fl~~~~~~L~~ls~eeF~~~k~~li~~l~ 92 (205)
++||++++..|..|+++|.++..+-+-.-+.
T Consensus 4 ~efL~~L~~~L~~lp~~e~~e~l~~Y~e~f~ 34 (181)
T PF08006_consen 4 NEFLNELEKYLKKLPEEEREEILEYYEEYFD 34 (181)
T ss_pred HHHHHHHHHHHHcCCHHHHHHHHHHHHHHHH
Confidence 5799999999999999998887665554444
No 28
>PF07609 DUF1572: Protein of unknown function (DUF1572); InterPro: IPR011466 This protein represents proteins with unknown function found in several diverse bacteria.
Probab=48.70 E-value=23 Score=27.90 Aligned_cols=87 Identities=13% Similarity=0.121 Sum_probs=46.0
Q ss_pred hhHHHHHHHHHHHH---HHHHHhcCCHHHHHHHHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCccccHHHHHH-HhcC
Q 028703 54 PKYIDLRVESFLQM---FESKLYEMTSDQFKNNVNALIDMKLEKHKNLKEESGFYWREISDGILKFDRREVEVAA-LRQL 129 (205)
Q Consensus 54 ~~~l~~~i~~Fl~~---~~~~L~~ls~eeF~~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~-l~~i 129 (205)
..+|...+..|-.. ....|..++++++--....-.+++.---..|..-+...|..++..+-.=..|++..+- ....
T Consensus 4 ~~yl~~~~~~f~~~k~~~~~~l~ql~de~l~w~~~~~sNSia~l~~HL~GNm~srw~~fl~~dGek~~R~RD~EF~~~~~ 83 (163)
T PF07609_consen 4 EEYLESVIKRFEYYKRLGEKALAQLSDEQLWWRPNEESNSIANLVLHLSGNMNSRWTDFLTTDGEKYWRNRDAEFENKFV 83 (163)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHhCCHHHhhhccCCCcccHHHHHHHhhccHHHHHHHHhCCCCCccCcCcccccccCCC
Confidence 45677777777644 3567788888887654322122211111223445556788776533211222333332 2457
Q ss_pred CHHHHHHHHHH
Q 028703 130 TQQELIYFFNE 140 (205)
Q Consensus 130 t~~dl~~f~~~ 140 (205)
++++|+.-+++
T Consensus 84 sk~eLl~~~~~ 94 (163)
T PF07609_consen 84 SKEELLARWEK 94 (163)
T ss_pred CHHHHHHHHHH
Confidence 78887776655
No 29
>PF08621 RPAP1_N: RPAP1-like, N-terminal; InterPro: IPR013930 Inhibition of RNA polymerase II-associated protein 1 (RPAP1) synthesis in Saccharomyces cerevisiae (Baker's yeast) results in changes in global gene expression that are similar to those caused by the loss of the RNAPII subunit Rpb11 []. This entry represents the N-terminal region of RPAP-1 that is conserved from yeast to humans.
Probab=44.11 E-value=34 Score=21.36 Aligned_cols=23 Identities=17% Similarity=0.420 Sum_probs=19.8
Q ss_pred HHHHhcCCHHHHHHHHHHHHHHH
Q 028703 69 ESKLYEMTSDQFKNNVNALIDMK 91 (205)
Q Consensus 69 ~~~L~~ls~eeF~~~k~~li~~l 91 (205)
...|.+||++|....++-|..++
T Consensus 9 ~~rL~~MS~eEI~~er~eL~~~L 31 (49)
T PF08621_consen 9 EARLASMSPEEIEEEREELLESL 31 (49)
T ss_pred HHHHHhCCHHHHHHHHHHHHHhC
Confidence 45788999999999999888776
No 30
>PF14270 DUF4358: Domain of unknown function (DUF4358)
Probab=40.68 E-value=1.1e+02 Score=21.76 Aligned_cols=61 Identities=13% Similarity=0.107 Sum_probs=41.7
Q ss_pred ceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHhcCCHHHHHHHHHHHH
Q 028703 27 GYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLYEMTSDQFKNNVNALI 88 (205)
Q Consensus 27 GYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~~ls~eeF~~~k~~li 88 (205)
||++........ ...+.++--.+..+.+.+...|+..+....+.....-+++.....++.+
T Consensus 33 ~~~~~~s~~~~~-~~ei~v~k~kd~~~~e~Vk~~l~~r~~~q~~~f~~Y~p~q~~~l~~a~v 93 (106)
T PF14270_consen 33 DYVIYMSMSNMS-ADEIAVFKAKDGKQAEDVKKALEKRLESQKKSFEGYLPEQYELLENAKV 93 (106)
T ss_pred eEEEEeccccCC-ccEEEEEEECCcCcHHHHHHHHHHHHHHHHHHHhccCHHHHHHHhcCEE
Confidence 566666554332 3344433344556799999999999999888888877777777665543
No 31
>cd08317 Death_ank Death domain associated with Ankyrins. Death Domain (DD) associated with Ankyrins. Ankyrins are modular proteins comprising three conserved domains, an N-terminal membrane-binding domain containing ANK repeats, a spectrin-binding domain and a C-terminal DD. Ankyrins function as adaptor proteins and they interact, through ANK repeats, with structurally diverse membrane proteins, including ion channels/pumps, calcium release channels, and cell adhesion molecules. They play critical roles in the proper expression and membrane localization of these proteins. In mammals, this family includes ankyrin-R for restricted (or ANK1), ankyrin-B for broadly expressed (or ANK2) and ankyrin-G for general or giant (or ANK3). They are expressed in different combinations in many tissues and play non-overlapping functions. In general, DDs are protein-protein interaction domains found in a variety of domain architectures. Their common feature is that they form homodimers by self-associati
Probab=38.11 E-value=1.3e+02 Score=20.48 Aligned_cols=74 Identities=14% Similarity=0.132 Sum_probs=47.3
Q ss_pred hhHHHHHHHHHHHHHHHHHhcCCHHHHHHHHHHHHHHHhccCcChHHHHH---HhHHHHhcCCCCccccHHHHHHHhcCC
Q 028703 54 PKYIDLRVESFLQMFESKLYEMTSDQFKNNVNALIDMKLEKHKNLKEESG---FYWREISDGILKFDRREVEVAALRQLT 130 (205)
Q Consensus 54 ~~~l~~~i~~Fl~~~~~~L~~ls~eeF~~~k~~li~~l~~~~~sl~~~~~---~~w~~I~~~~~~F~~~~~~i~~l~~it 130 (205)
-..|...|-.-+..+...| ++++.+.+..+. +.|.++.+++. +.|.+-... . -..+.++.+|+++.
T Consensus 7 l~~ia~~lG~dW~~LAr~L-g~~~~dI~~i~~-------~~~~~~~eq~~~mL~~W~~r~g~--~-at~~~L~~AL~~i~ 75 (84)
T cd08317 7 LADISNLLGSDWPQLAREL-GVSETDIDLIKA-------ENPNSLAQQAQAMLKLWLEREGK--K-ATGNSLEKALKKIG 75 (84)
T ss_pred HHHHHHHHhhHHHHHHHHc-CCCHHHHHHHHH-------HCCCCHHHHHHHHHHHHHHhcCC--c-chHHHHHHHHHHcC
Confidence 3456666666666666555 488888777654 33555555444 455554221 2 33478999999999
Q ss_pred HHHHHHHH
Q 028703 131 QQELIYFF 138 (205)
Q Consensus 131 ~~dl~~f~ 138 (205)
+.|+.+-+
T Consensus 76 r~Di~~~~ 83 (84)
T cd08317 76 RDDIVEKC 83 (84)
T ss_pred hHHHHHHh
Confidence 99998754
No 32
>PF05120 GvpG: Gas vesicle protein G ; InterPro: IPR007804 Gas vesicles are intracellular, protein-coated, and hollow organelles found in cyanobacteria and halophilic archaea. They are permeable to ambient gases by diffusion and provide buoyancy, enabling cells to move upwards in water to access oxygen and/or light. Proteins containing this family are involved in the formation of gas vesicles [].
Probab=35.12 E-value=1.5e+02 Score=20.41 Aligned_cols=37 Identities=19% Similarity=0.331 Sum_probs=26.4
Q ss_pred CChhHHHHHHHHHHHHHHHHH--hcCCHHHHHHHHHHHHHHHh
Q 028703 52 KGPKYIDLRVESFLQMFESKL--YEMTSDQFKNNVNALIDMKL 92 (205)
Q Consensus 52 ~~~~~l~~~i~~Fl~~~~~~L--~~ls~eeF~~~k~~li~~l~ 92 (205)
++|..|...+.+. ...+ .++|+++|+.....|+..+.
T Consensus 28 ~Dp~~i~~~L~~L----~~~~e~GEIseeEf~~~E~eLL~rL~ 66 (79)
T PF05120_consen 28 YDPAAIRRELAEL----QEALEAGEISEEEFERREDELLDRLE 66 (79)
T ss_pred cCHHHHHHHHHHH----HHHHHcCCCCHHHHHHHHHHHHHHHH
Confidence 5666555555443 3333 48999999999999998876
No 33
>cd00498 Hsp33 Heat shock protein 33 (Hsp33): Cytosolic protein that acts as a molecular chaperone under oxidative conditions. In normal (reducing) cytosolic conditions, four conserved Cys residues are coordinated by a Zn ion. Under oxidative stress (such as heat shock), the Cys are reversibly oxidized to disulfide bonds, which causes the chaperone activity to be turned on. Hsp33 is homodimeric in its functional form.
Probab=32.80 E-value=3.1e+02 Score=23.25 Aligned_cols=117 Identities=15% Similarity=0.005 Sum_probs=65.3
Q ss_pred HHHchHHHHHhhhccccceEEEEEEeeeCC---eeEEEEEEeC-CCCChhHHHHHHHHHHHHHHHHHhcCCHHHHHHH-H
Q 028703 10 LIAKQPAFHQLRTVEQLGYITALLQRNDFG---IHGVQFIIQS-SVKGPKYIDLRVESFLQMFESKLYEMTSDQFKNN-V 84 (205)
Q Consensus 10 ~ils~~~f~~LRTkqQLGYvV~s~~~~~~~---~~gl~~~VQS-~~~~~~~l~~~i~~Fl~~~~~~L~~ls~eeF~~~-k 84 (205)
.-+.+-+=+++++-+|+-=.|..+.....+ ...-.++||. |..+ +...+ .++.....+..+++.+.... .
T Consensus 133 g~iaedl~~Yf~qSEQipt~v~l~v~~~~~~~v~~AgG~liQ~LP~~~-e~~~~----~~e~~~~~~~~~~~~~~~~~~~ 207 (275)
T cd00498 133 GEIAEDLEYYFAQSEQLPSAVGLGVLVNPDGTVKAAGGLLLQVLPGAD-EEDID----AWEKVIKLMPTVSALELLGLSP 207 (275)
T ss_pred CcHHHHHHHHHHhccccceEEEEEEeccCCCCeeEEEEEEEEeCcCCC-hhhHH----HHHHHHHhCCCccHHHHcCCCH
Confidence 345667778899999999999888875322 3455567775 4332 22222 23333333444554433321 2
Q ss_pred HHHHHHHhccCcChHHHHHHhHHHHhcCCCCc---cccHHHHHHHhcCCHHHHHHHHHH
Q 028703 85 NALIDMKLEKHKNLKEESGFYWREISDGILKF---DRREVEVAALRQLTQQELIYFFNE 140 (205)
Q Consensus 85 ~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F---~~~~~~i~~l~~it~~dl~~f~~~ 140 (205)
+.++..+..... +.-+......| ..+++..++|..+.++|+.+.+++
T Consensus 208 e~ll~~lf~~~~---------~~i~~~~~v~f~C~CS~er~~~~L~~Lg~~El~~i~~e 257 (275)
T cd00498 208 EELLYRLFHEEE---------VRILEKQPVRFRCDCSRERVAAALLTLGKEELADMIEE 257 (275)
T ss_pred HHHHHHHhCCCC---------ceeccCCCcCeeCCCCHHHHHHHHHhCCHHHHHHHHHc
Confidence 333333322110 01111222233 468999999999999999998754
No 34
>PF13333 rve_2: Integrase core domain
Probab=32.20 E-value=51 Score=20.32 Aligned_cols=32 Identities=9% Similarity=0.430 Sum_probs=26.0
Q ss_pred CChhHHHHHHHHHHHHHH-HHHhcCCHHHHHHH
Q 028703 52 KGPKYIDLRVESFLQMFE-SKLYEMTSDQFKNN 83 (205)
Q Consensus 52 ~~~~~l~~~i~~Fl~~~~-~~L~~ls~eeF~~~ 83 (205)
.+.+++...|.+++..+. ..|..||+.+|+..
T Consensus 18 ~t~eel~~~I~~YI~~yN~~Rl~~lsP~eyr~~ 50 (52)
T PF13333_consen 18 KTREELKQAIDEYIDYYNNERLKGLSPVEYRNQ 50 (52)
T ss_pred chHHHHHHHHHHHHHHhccCCCCCcCHHHHHHh
Confidence 467899999999998874 35789999998764
No 35
>PF14203 DUF4319: Domain of unknown function (DUF4319); PDB: 2L7K_A.
Probab=32.06 E-value=70 Score=21.17 Aligned_cols=25 Identities=20% Similarity=0.189 Sum_probs=13.2
Q ss_pred HHHHHHHHHHHHHHhcCCHHHHHHH
Q 028703 59 LRVESFLQMFESKLYEMTSDQFKNN 83 (205)
Q Consensus 59 ~~i~~Fl~~~~~~L~~ls~eeF~~~ 83 (205)
..+.+........|..||+++|...
T Consensus 35 ~em~eLa~~tl~KL~~mtD~ef~~l 59 (64)
T PF14203_consen 35 SEMRELAESTLRKLDAMTDAEFAEL 59 (64)
T ss_dssp HHHHHHHHHHHHHHTT--HHHHHHH
T ss_pred HHHHHHHHHHHHHHHhcCHHHHHHh
Confidence 3344555555556667777777653
No 36
>KOG2085 consensus Serine/threonine protein phosphatase 2A, regulatory subunit [Signal transduction mechanisms]
Probab=30.93 E-value=1e+02 Score=28.11 Aligned_cols=44 Identities=18% Similarity=0.325 Sum_probs=34.5
Q ss_pred HHHHHHHHHhcCCHHHHHHHHHHHHHHHhccCcC----hHHHHHHhHH
Q 028703 64 FLQMFESKLYEMTSDQFKNNVNALIDMKLEKHKN----LKEESGFYWR 107 (205)
Q Consensus 64 Fl~~~~~~L~~ls~eeF~~~k~~li~~l~~~~~s----l~~~~~~~w~ 107 (205)
||.++++.|+.+.+.||++...-|-.++..--.| ..+++-.+|+
T Consensus 319 FL~ElEEILe~iep~eFqk~~~PLf~qia~c~sS~HFQVAEraL~~wn 366 (457)
T KOG2085|consen 319 FLNELEEILEVIEPSEFQKIMVPLFRQIARCVSSPHFQVAERALYLWN 366 (457)
T ss_pred eHhhHHHHHHhcCHHHHHHHhHHHHHHHHHHcCChhHHHHHHHHHHHh
Confidence 4556778888999999999998888887665444 4678888887
No 37
>PF11594 Med28: Mediator complex subunit 28; InterPro: IPR021640 Mediator is a large complex of up to 33 proteins that is conserved from plants to fungi to humans - the number and representation of individual subunits varying with species [],[]. It is arranged into four different sections, a core, a head, a tail and a kinase-activity part, and the number of subunits within each of these is what varies with species. Overall, Mediator regulates the transcriptional activity of RNA polymerase II but it would appear that each of the four different sections has a slightly different function []. Subunit Med28 of the Mediator may function as a scaffolding protein within Mediator by maintaining the stability of a submodule within the head module, and components of this submodule act together in a gene-regulatory programme to suppress smooth muscle cell differentiation. Thus, mammalian Mediator subunit Med28 functions as a repressor of smooth muscle-cell differentiation, which could have implications for disorders associated with abnormalities in smooth muscle cell growth and differentiation, including atherosclerosis, asthma, hypertension, and smooth muscle tumours [].
Probab=29.28 E-value=1.5e+02 Score=21.75 Aligned_cols=47 Identities=21% Similarity=0.273 Sum_probs=30.2
Q ss_pred HHHHHHHHHHHHHHHHHhcCCHHHHHHHHHHHHHHHhccCcChHHHHHHhHHHH
Q 028703 56 YIDLRVESFLQMFESKLYEMTSDQFKNNVNALIDMKLEKHKNLKEESGFYWREI 109 (205)
Q Consensus 56 ~l~~~i~~Fl~~~~~~L~~ls~eeF~~~k~~li~~l~~~~~sl~~~~~~~w~~I 109 (205)
+++..+..|+.-.++. +-|=-.|..++ +...+...+.++.+.+-.++
T Consensus 5 ~vEq~~~~FlD~aRq~------e~~FlqKr~~L-S~~kpe~~lkEEi~eLK~El 51 (106)
T PF11594_consen 5 YVEQLIQSFLDVARQM------EAFFLQKRFEL-SAYKPEQVLKEEINELKEEL 51 (106)
T ss_pred HHHHHHHHHHHHHHHH------HHHHHHHHHHH-HhcCHHHHHHHHHHHHHHHH
Confidence 5677777776544332 33444455555 66677778888888887776
No 38
>PF04485 NblA: Phycobilisome degradation protein nblA ; InterPro: IPR007574 In the cyanobacterium Synechococcus species PCC 7942 (P35087 from SWISSPROT), nblA triggers degradation of light-harvesting phycobiliproteins in response to deprivation nutrients including nitrogen, phosphorus and sulphur. The mechanism of nblA function is not known, but it has been hypothesised that nblA may act by disrupting phycobilisome structure, activating a protease or tagging phycobiliproteins for proteolysis. Members of this family have also been identified in the chloroplasts of some red algae.; PDB: 3CS5_D 1OJH_L 2QDO_B 2Q8V_A.
Probab=27.42 E-value=1e+02 Score=19.60 Aligned_cols=35 Identities=14% Similarity=0.252 Sum_probs=24.1
Q ss_pred HHHHHHHhcCCHHHHHHHHHHHHHHHhccCcChHH
Q 028703 66 QMFESKLYEMTSDQFKNNVNALIDMKLEKHKNLKE 100 (205)
Q Consensus 66 ~~~~~~L~~ls~eeF~~~k~~li~~l~~~~~sl~~ 100 (205)
..+.+.+..||.|+-.++.-.+..++.-++.-+..
T Consensus 13 ~~~~~qv~~ls~Eqaq~~Lve~~rqmmikeN~~k~ 47 (53)
T PF04485_consen 13 RSFKDQVQKLSREQAQELLVELYRQMMIKENLIKH 47 (53)
T ss_dssp HHHHHHHCTS-HHHHHHHHHHHHHHHHHHHHHHHH
T ss_pred HHHHHHHHHhCHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 45778889999999888877777776655544333
No 39
>PF08671 SinI: Anti-repressor SinI; InterPro: IPR010981 The SinR repressor is part of a group of Sin (sporulation inhibition) proteins in Bacillus subtilis that regulate the commitment to sporulation in response to extreme adversity []. SinR is a tetrameric repressor protein that binds to the promoters of genes essential for entry into sporulation and prevents their transcription. This repression is overcome through the activity of SinI, which disrupts the SinR tetramer through the formation of a SinI-SinR heterodimer, thereby allowing sporulation to proceed. The SinR structure consists of two domains: a dimerisation domain stabilised by a hydrophobic core, and a DNA-binding domain that is identical to domains of the bacteriophage 434 CI and Cro proteins that regulate prophage induction. The dimerisation domain is a four-helical bundle formed from two helices from the C-terminal residues of SinR and two helices from the central residues of SinI. These regions in SinR and SinI are similar in both structure and sequence. The interaction of SinR monomers to form tetramers is weaker than between SinR and SinI, since SinI can effectively disrupt SinR tetramers. This entry represents the dimerisation domain in both SinI and SinR proteins.; GO: 0005488 binding, 0006355 regulation of transcription, DNA-dependent; PDB: 1B0N_A 2YAL_A.
Probab=27.22 E-value=84 Score=17.53 Aligned_cols=19 Identities=21% Similarity=0.200 Sum_probs=11.9
Q ss_pred HHHHHh-cCCHHHHHHHHHH
Q 028703 122 EVAALR-QLTQQELIYFFNE 140 (205)
Q Consensus 122 ~i~~l~-~it~~dl~~f~~~ 140 (205)
..+|.+ .||++|+.+|+..
T Consensus 9 i~eA~~~Gls~eeir~FL~~ 28 (30)
T PF08671_consen 9 IKEAKESGLSKEEIREFLEF 28 (30)
T ss_dssp HHHHHHTT--HHHHHHHHHH
T ss_pred HHHHHHcCCCHHHHHHHHHh
Confidence 344443 6999999999864
No 40
>PF02758 PYRIN: PAAD/DAPIN/Pyrin domain; InterPro: IPR004020 Pyrin domain was identified as putative protein-protein interaction domain at the N-terminal region of several proteins thought to function in apoptotic and inflammatory signalling pathways. Using secondary structure prediction and potential-based fold recognition methods, the PYRIN domain is predicted to be a member of the six-helix bundle death domain-fold superfamily that includes death domains (DDs), death effector domains (DEDs), and caspase recruitment domains (CARDs). Members of the death domain-fold superfamily are well established mediators of protein-protein interactions found in many proteins involved in apoptosis and inflammation, indicating further that the PYRIN domains serve a similar function. Comparison of a circular dichroism spectrum of the PYRIN domain of CARD7/DEFCAP/NAC/NALP1 with spectra of several proteins known to adopt the death domain-fold provides experimental support for the structure prediction [] It is found in interferon-inducible proteins, pyrin and myeloid cell nuclear differentiation antigen.; PDB: 2DO9_A 2YU0_A 2KN6_A 1UCP_A 2L6A_A 2KM6_A 1PN5_A 2DBG_A 3QF2_B 2HM2_Q.
Probab=26.95 E-value=59 Score=22.24 Aligned_cols=70 Identities=16% Similarity=0.103 Sum_probs=35.8
Q ss_pred HHHHHhcCCHHHHHHHHHHHHHHHhccCcChH----HHHHH--hHHHHhcCCCCcc-ccHHHHHHHhcCCHHHHHHHH
Q 028703 68 FESKLYEMTSDQFKNNVNALIDMKLEKHKNLK----EESGF--YWREISDGILKFD-RREVEVAALRQLTQQELIYFF 138 (205)
Q Consensus 68 ~~~~L~~ls~eeF~~~k~~li~~l~~~~~sl~----~~~~~--~w~~I~~~~~~F~-~~~~~i~~l~~it~~dl~~f~ 138 (205)
+...|++|++++|+..|.-|.........++. +.+++ ...-+.. .|.=. .-+..++.++.|...|+.+=.
T Consensus 5 Ll~~Le~L~~~efk~FK~~L~~~~~~~~~~Ip~~~le~ad~~~la~lLv~-~y~~~~A~~vt~~il~~m~~~dLae~l 81 (83)
T PF02758_consen 5 LLWYLEELSEEEFKRFKWLLKEPVKEGFPPIPRGELEKADREDLADLLVQ-HYGEQRAWEVTLKILEKMNRNDLAEKL 81 (83)
T ss_dssp HHHHHHTS-HHHHHHHHHHHHSTSSTTTCSSSHCHHHHSSHHHHHHHHHH-HTCHHHHHHHHHHHHHHTTCHHHHHHH
T ss_pred HHHHHHhCCHHHHHHHHHHhcchhhcCCCCCCHHHHhhCCHHHHHHHHHH-HcCHHHHHHHHHHHHHHcChHHHHHHH
Confidence 45678999999999999988632222222221 11111 1111111 12111 224566677777777765543
No 41
>KOG2019 consensus Metalloendoprotease HMP1 (insulinase superfamily) [General function prediction only; Posttranslational modification, protein turnover, chaperones]
Probab=26.90 E-value=6.4e+02 Score=25.08 Aligned_cols=144 Identities=13% Similarity=0.168 Sum_probs=79.6
Q ss_pred HHHHHHHHH----chHHHHHhhhccccc--eEEEEEEeeeCCeeEEEEEEeCC-CCChhHHHHHHHHHHHHHHHHHhcCC
Q 028703 4 KLQLLALIA----KQPAFHQLRTVEQLG--YITALLQRNDFGIHGVQFIIQSS-VKGPKYIDLRVESFLQMFESKLYEMT 76 (205)
Q Consensus 4 ~~~Ll~~il----s~~~f~~LRTkqQLG--YvV~s~~~~~~~~~gl~~~VQS~-~~~~~~l~~~i~~Fl~~~~~~L~~ls 76 (205)
.+.+|.++| ++|||.-| -+-+|| ..|.+++...--.+-+.+.+|+- ..+.+.+.+-|..-+. ++-
T Consensus 332 aL~~L~~Ll~~gpsSp~yk~L-iESGLGtEfsvnsG~~~~t~~~~fsVGLqGvseediekve~lV~~t~~-------~la 403 (998)
T KOG2019|consen 332 ALKVLSHLLLDGPSSPFYKAL-IESGLGTEFSVNSGYEDTTLQPQFSVGLQGVSEEDIEKVEELVMNTFN-------KLA 403 (998)
T ss_pred HHHHHHHHhcCCCccHHHHHH-HHcCCCcccccCCCCCcccccceeeeeeccccHHHHHHHHHHHHHHHH-------HHH
Confidence 345555554 78999998 677899 88999999988889999999973 2333333333333332 222
Q ss_pred HHHHHH-HHHHHHHHHhc--cCcCh------HHHHHHhHHHHhcCCCCccc--cHHHHHHHh----cCCHHHHHHHHHHh
Q 028703 77 SDQFKN-NVNALIDMKLE--KHKNL------KEESGFYWREISDGILKFDR--REVEVAALR----QLTQQELIYFFNEN 141 (205)
Q Consensus 77 ~eeF~~-~k~~li~~l~~--~~~sl------~~~~~~~w~~I~~~~~~F~~--~~~~i~~l~----~it~~dl~~f~~~~ 141 (205)
++.|++ .+++++.++.- +.+|. ....-..|. +.-=-|+. -+..++.++ .=++.=+....++|
T Consensus 404 e~gfd~drieAil~qiEislk~qst~fGL~L~~~i~~~W~---~d~DPfE~Lk~~~~L~~lk~~l~ek~~~lfq~lIkkY 480 (998)
T KOG2019|consen 404 ETGFDNDRIEAILHQIEISLKHQSTGFGLSLMQSIISKWI---NDMDPFEPLKFEEQLKKLKQRLAEKSKKLFQPLIKKY 480 (998)
T ss_pred HhccchHHHHHHHHHhhhhhhccccchhHHHHHHHhhhhc---cCCCccchhhhhhHHHHHHHHHhhhchhHHHHHHHHH
Confidence 444555 34566665542 22222 222222232 11111332 122333332 22566677888888
Q ss_pred hhcCCCCccEEEEEEeeCCC
Q 028703 142 IKAGAPRKKTLSVRVYGSLH 161 (205)
Q Consensus 142 ~~~~~~~~~~l~i~v~~~~~ 161 (205)
+. .|..++.+.+.+...
T Consensus 481 il---nn~h~~t~smqpd~e 497 (998)
T KOG2019|consen 481 IL---NNPHCFTFSMQPDPE 497 (998)
T ss_pred Hh---cCCceEEEEecCCch
Confidence 83 445577777765543
No 42
>PF05553 DUF761: Cotton fibre expressed protein; InterPro: IPR008480 This family consists of several plant proteins of unknown function. Three of the sequences from Gossypium hirsutum (Upland cotton) in this family are described as G. hirsutum fibre expressed proteins []. The remaining sequences, found in Arabidopsis thaliana, are uncharacterised.
Probab=26.76 E-value=1.4e+02 Score=17.54 Aligned_cols=20 Identities=30% Similarity=0.514 Sum_probs=16.5
Q ss_pred hhHHHHHHHHHHHHHHHHHh
Q 028703 54 PKYIDLRVESFLQMFESKLY 73 (205)
Q Consensus 54 ~~~l~~~i~~Fl~~~~~~L~ 73 (205)
-++|..+-++|+..|.+.|.
T Consensus 2 ~~evd~rAe~FI~~f~~qlr 21 (38)
T PF05553_consen 2 DDEVDRRAEEFIAKFREQLR 21 (38)
T ss_pred chHHHHHHHHHHHHHHHHHH
Confidence 36789999999999987664
No 43
>COG1078 HD superfamily phosphohydrolases [General function prediction only]
Probab=26.62 E-value=39 Score=30.73 Aligned_cols=81 Identities=19% Similarity=0.165 Sum_probs=45.2
Q ss_pred HHHHHHHchHHHHHhhhccccceE--EEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHhcCCHHHH-HH
Q 028703 6 QLLALIAKQPAFHQLRTVEQLGYI--TALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLYEMTSDQF-KN 82 (205)
Q Consensus 6 ~Ll~~ils~~~f~~LRTkqQLGYv--V~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~~ls~eeF-~~ 82 (205)
.++..++.+|-|++||--+|||=. |+-+..-.+-... -..-+|..++-+-+..-.. +..++++- ..
T Consensus 18 ~~i~~LIdT~~FQRLRrIkQLG~a~lvyPgAnHTRFeHS---------LGV~~la~~~~~~l~~~~~--~~~~~~~~~~~ 86 (421)
T COG1078 18 ELILELIDTPEFQRLRRIKQLGLAYLVYPGANHTRFEHS---------LGVYHLARRLLEHLEKNSE--EEIDEEERLLV 86 (421)
T ss_pred HHHHHHhCCHHHHHHHHhhhccceeEecCCCcccccchh---------hHHHHHHHHHHHHHhhccc--cccchHHHHHH
Confidence 367789999999999999999943 3322222211111 2244555554443322111 22333333 23
Q ss_pred HHHHHHHHHhccCcC
Q 028703 83 NVNALIDMKLEKHKN 97 (205)
Q Consensus 83 ~k~~li~~l~~~~~s 97 (205)
...||+-.+-.+|-|
T Consensus 87 ~~AALLHDIGHgPFS 101 (421)
T COG1078 87 RLAALLHDIGHGPFS 101 (421)
T ss_pred HHHHHHHccCCCccc
Confidence 457888888888765
No 44
>PF07735 FBA_2: F-box associated; InterPro: IPR012885 This domain is found is found towards the C terminus of proteins that contain an F-box, IPR001810 from INTERPRO, suggesting that they are effectors linked with ubiquitination.
Probab=26.15 E-value=1e+02 Score=19.82 Aligned_cols=26 Identities=15% Similarity=0.267 Sum_probs=19.3
Q ss_pred cCCHHHHHHHHHHhhhcCCCCccEEEEE
Q 028703 128 QLTQQELIYFFNENIKAGAPRKKTLSVR 155 (205)
Q Consensus 128 ~it~~dl~~f~~~~~~~~~~~~~~l~i~ 155 (205)
.+|.+|+..|++.++ .+.+++.=.++
T Consensus 43 ~~t~~dln~Flk~W~--~G~~~~Le~l~ 68 (70)
T PF07735_consen 43 KFTNEDLNKFLKHWI--NGSNPRLEYLE 68 (70)
T ss_pred CCCHHHHHHHHHHHH--cCCCcCCcEEE
Confidence 589999999999999 45555543333
No 45
>cd08803 Death_ank3 Death domain of Ankyrin-3. Death Domain (DD) of the human protein ankyrin-3 (ANK-3) and related proteins. Ankyrins are modular proteins comprising three conserved domains, an N-terminal membrane-binding domain containing ANK repeats, a spectrin-binding domain and a C-terminal DD. ANK-3, also called anykyrin-G (for general or giant), is found in neurons and at least one splice variant has been shown to be essential for propagation of action potentials as a binding partner to neurofascin and voltage-gated sodium channels. It is required for maintaining axo-dendritic polarity, and may be a genetic risk factor associated with bipolar disorder. ANK-3 may also play roles in other cell types. Mutations affecting ANK-3 pathways for Na channel localization are associated with Brugada syndrome, a potentially fata arrythmia. In general, DDs are protein-protein interaction domains found in a variety of domain architectures. Their common feature is that they form homodimers by se
Probab=26.09 E-value=2.3e+02 Score=19.58 Aligned_cols=58 Identities=10% Similarity=0.156 Sum_probs=37.0
Q ss_pred cCCHHHHHHHHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHH
Q 028703 74 EMTSDQFKNNVNALIDMKLEKHKNLKEESGFYWREISDGILKFDRREVEVAALRQLTQQELIYFF 138 (205)
Q Consensus 74 ~ls~eeF~~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~ 138 (205)
++++.+....+ .+-|.++.+++...-.....+...--.-+.++.+|.+|.+.|+.+..
T Consensus 26 g~s~~dI~~i~-------~e~p~~~~~q~~~lL~~W~~r~g~~At~~~L~~aL~~i~R~DIv~~~ 83 (84)
T cd08803 26 NFSVDEINQIR-------VENPNSLIAQSFMLLKKWVTRDGKNATTDALTSVLTKINRIDIVTLL 83 (84)
T ss_pred CCCHHHHHHHH-------HhCCCCHHHHHHHHHHHHHHhhCCCcHHHHHHHHHHHCCcHHHHHhc
Confidence 57777766664 23467777665544444433332223346799999999999998753
No 46
>cd08304 DD_superfamily The Death Domain Superfamily of protein-protein interaction domains. The Death Domain (DD) superfamily includes the DD, Pyrin, CARD (Caspase activation and recruitment domain) and DED (Death Effector Domain) families. DDs are protein-protein interaction domains found in a variety of domain architectures. Their common feature is that they form homodimers by self-association or heterodimers by associating with other members of the DD superfamily. They serve as adaptors in signaling pathways and can recruit other proteins into signaling complexes. They are prominent components of the programmed cell death (apoptosis) pathway and are found in a number of other signaling pathways including those that impact innate immunity, inflammation, differentiation, and cancer.
Probab=25.29 E-value=1.9e+02 Score=18.98 Aligned_cols=63 Identities=13% Similarity=0.039 Sum_probs=33.6
Q ss_pred HHhcCCHHHHHHHHHHHHHHHhccCcChHHHH-HHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHH
Q 028703 71 KLYEMTSDQFKNNVNALIDMKLEKHKNLKEES-GFYWREISDGILKFDRREVEVAALRQLTQQELIY 136 (205)
Q Consensus 71 ~L~~ls~eeF~~~k~~li~~l~~~~~sl~~~~-~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~ 136 (205)
.++.|++++|+..+..+.+ .-++..+++-. .+-|-.+.-+.|.-.. ......++.+...++.+
T Consensus 4 L~~~L~~~~~~~l~~~l~~--~~~~~~~e~i~~a~~ll~~l~~~~~~a~-~~~~~vL~~~~~~~la~ 67 (69)
T cd08304 4 LCENLTLEVLQQLKTALKS--RIPPDQVEQISAANELLNILESQYNHTL-QLLFALFEDLGLHNLAR 67 (69)
T ss_pred HHHHhhHhHHHHHHHHHHc--cCCHHHHHHhhHHHHHHHHHHHhCcchH-HHHHHHHHHcCCHhHHh
Confidence 3467888899998888886 12222233222 4555555533332222 34555666665555543
No 47
>PF00675 Peptidase_M16: Insulinase (Peptidase family M16) This is family M16 in the peptidase classification. ; InterPro: IPR011765 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Metalloproteases are the most diverse of the four main types of protease, with more than 50 families identified to date. In these enzymes, a divalent cation, usually zinc, activates the water molecule. The metal ion is held in place by amino acid ligands, usually three in number. The known metal ligands are His, Glu, Asp or Lys and at least one other residue is required for catalysis, which may play an electrophillic role. Of the known metalloproteases, around half contain an HEXXH motif, which has been shown in crystallographic studies to form part of the metal-binding site []. The HEXXH motif is relatively common, but can be more stringently defined for metalloproteases as 'abXHEbbHbc', where 'a' is most often valine or threonine and forms part of the S1' subsite in thermolysin and neprilysin, 'b' is an uncharged residue, and 'c' a hydrophobic residue. Proline is never found in this site, possibly because it would break the helical structure adopted by this motif in metalloproteases []. The majority of the sequences in this entry are metallopeptidases and non-peptidase homologs belong to MEROPS peptidase family M16 (clan ME), subfamilies M16A, M16B and M16C; they include: Insulinase, insulin-degrading enzyme (3.4.24.56 from EC) Mitochondrial processing peptidase alpha subunit, (Alpha-MPP, 3.4.24.64 from EC) Pitrlysin, Protease III precursor (3.4.24.55 from EC) Nardilysin, (3.4.24.61 from EC) Ubiquinol-cytochrome C reductase complex core protein I,mitochondrial precursor (1.10.2.2 from EC) Coenzyme PQQ synthesis protein F (3.4.99 from EC) These proteins do not share many regions of sequence similarity; the most noticeable is in the N-terminal section. This region includes a conserved histidine followed, two residues later by a glutamate and another histidine. In pitrilysin, it has been shown [] that this H-x-x-E-H motif is involved in enzymatic activity; the two histidines bind zinc and the glutamate is necessary for catalytic activity. The proteins classified as non-peptidase homologues either have been found experimentally to be without peptidase activity, or lack amino acid residues that are believed to be essential for the catalytic activity. ; GO: 0004222 metalloendopeptidase activity, 0006508 proteolysis; PDB: 3P7L_A 3P7O_A 3TUV_A 3GO9_A 1BE3_B 1PP9_B 2A06_B 1SQB_B 1SQP_B 1L0N_B ....
Probab=23.94 E-value=2.9e+02 Score=20.17 Aligned_cols=86 Identities=13% Similarity=0.083 Sum_probs=56.0
Q ss_pred HHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHh--cCCHHHHHHHHHHHHHHHhcc
Q 028703 17 FHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFESKLY--EMTSDQFKNNVNALIDMKLEK 94 (205)
Q Consensus 17 f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~--~ls~eeF~~~k~~li~~l~~~ 94 (205)
.+-.+.-+++|=.+. .....+...+.+.+.+. + ++.-|.-+.+.+. .+++++|+..|..+...+.+.
T Consensus 51 ~~l~~~l~~~G~~~~--~~t~~d~t~~~~~~~~~--~-------~~~~l~~l~~~~~~P~f~~~~~~~~r~~~~~ei~~~ 119 (149)
T PF00675_consen 51 DELQEELESLGASFN--ASTSRDSTSYSASVLSE--D-------LEKALELLADMLFNPSFDEEEFEREREQILQEIEEI 119 (149)
T ss_dssp HHHHHHHHHTTCEEE--EEEESSEEEEEEEEEGG--G-------HHHHHHHHHHHHHSBGGCHHHHHHHHHHHHHHHHHH
T ss_pred hhhHHHhhhhccccc--eEecccceEEEEEEecc--c-------chhHHHHHHHHHhCCCCCHHHHHHHHHHHHHHHHHH
Confidence 333444566774443 33345555666655543 2 3333444444443 699999999999999999988
Q ss_pred CcChHHHHHHhHHHHhcCC
Q 028703 95 HKNLKEESGFYWREISDGI 113 (205)
Q Consensus 95 ~~sl~~~~~~~w~~I~~~~ 113 (205)
..+-...+...+.....++
T Consensus 120 ~~~~~~~~~~~l~~~~f~~ 138 (149)
T PF00675_consen 120 KENPQELAFEKLHSAAFRG 138 (149)
T ss_dssp TTHHHHHHHHHHHHHHHTT
T ss_pred HCCHHHHHHHHHHHHHhcc
Confidence 8888888888888776543
No 48
>PF14659 Phage_int_SAM_3: Phage integrase, N-terminal SAM-like domain; PDB: 2KD1_A 2KOB_A 2KHQ_A 3LYS_E 2KIW_A 2KKP_A.
Probab=23.44 E-value=64 Score=19.60 Aligned_cols=17 Identities=24% Similarity=0.597 Sum_probs=12.0
Q ss_pred HhcCCHHHHHHHHHHhh
Q 028703 126 LRQLTQQELIYFFNENI 142 (205)
Q Consensus 126 l~~it~~dl~~f~~~~~ 142 (205)
|++||..++.+|+++++
T Consensus 42 i~~It~~~i~~~~~~l~ 58 (58)
T PF14659_consen 42 IKDITPRDIQNFINELL 58 (58)
T ss_dssp GGG--HHHHHHHHHHH-
T ss_pred HHHCCHHHHHHHHHHcC
Confidence 78899999999988753
No 49
>PF06816 NOD: NOTCH protein; InterPro: IPR010660 NOTCH signalling plays a fundamental role during a great number of developmental processes in multicellular animals []. NOD (NOTCH protein domain) represents a region present in many NOTCH proteins and NOTCH homologues in multiple species such as 0, NOTCH2 and NOTCH3, LIN12, SC1 and TAN1. Role of NOD domain remains to be elucidated.; GO: 0030154 cell differentiation, 0016021 integral to membrane; PDB: 2OO4_A 3ETO_A 3I08_A 3L95_X.
Probab=22.16 E-value=1.5e+02 Score=19.01 Aligned_cols=30 Identities=13% Similarity=0.185 Sum_probs=23.6
Q ss_pred eEEEEEEeCCCCChhHHHHHHHHHHHHHHHHHh
Q 028703 41 HGVQFIIQSSVKGPKYIDLRVESFLQMFESKLY 73 (205)
Q Consensus 41 ~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~~L~ 73 (205)
+.+.++|+ -+|+++.+.-..||.++...|.
T Consensus 7 G~lvivvl---~~P~~f~~~~~~FLr~Ls~~Lr 36 (57)
T PF06816_consen 7 GTLVIVVL---MDPEEFRNNSVQFLRELSRVLR 36 (57)
T ss_dssp SEEEEEES---S-HHHHHHTHHHHHHHHHHHCT
T ss_pred eeEEEEEE---eCHHHHHHHHHHHHHHHHHHHe
Confidence 45667777 7899999999999999887764
No 50
>PF08700 Vps51: Vps51/Vps67; InterPro: IPR014812 The VFT tethering complex (also known as GARP complex, Golgi associated retrograde protein complex, Vps53 tethering complex) is a conserved eukaryotic docking complex which is involved in recycling of proteins from endosomes to the late Golgi. Vps51 (also known as Vps67) is a subunit of VFT and interacts with the SNARE Tlg1 [].
Probab=22.13 E-value=2.4e+02 Score=18.82 Aligned_cols=43 Identities=16% Similarity=0.149 Sum_probs=32.8
Q ss_pred HHHHhcCCHHHHHHHHHHHHHHHhccCcChHHHHHHhHHHHhc
Q 028703 69 ESKLYEMTSDQFKNNVNALIDMKLEKHKNLKEESGFYWREISD 111 (205)
Q Consensus 69 ~~~L~~ls~eeF~~~k~~li~~l~~~~~sl~~~~~~~w~~I~~ 111 (205)
...+...+.++.....+.|...+......|......++.+++.
T Consensus 13 ~~~l~~~s~~~i~~~~~~L~~~i~~~~~eLr~~V~~nY~~fI~ 55 (87)
T PF08700_consen 13 KDLLKNSSIKEIRQLENKLRQEIEEKDEELRKLVYENYRDFIE 55 (87)
T ss_pred HHHHhhCCHHHHHHHHHHHHHHHHHHHHHHHHHHHhhHHHHHH
Confidence 3456678888999999999988888888888777776666543
No 51
>PF11385 DUF3189: Protein of unknown function (DUF3189); InterPro: IPR021525 This family of proteins with unknown function appears to be restricted to Firmicutes
Probab=21.87 E-value=2.2e+02 Score=21.93 Aligned_cols=55 Identities=16% Similarity=0.231 Sum_probs=40.9
Q ss_pred HHHHchHHHHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHH
Q 028703 9 ALIAKQPAFHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMF 68 (205)
Q Consensus 9 ~~ils~~~f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~ 68 (205)
..+++-|+|+.+- ++..|...+.+.....+- +++-+-....+-+...|..|++-+
T Consensus 33 ~el~~lp~fd~~~-~~d~G~l~y~G~De~gn~----VY~lG~~~~~~~~~~al~~l~~i~ 87 (148)
T PF11385_consen 33 EELLSLPYFDKLE-KEDIGRLIYMGTDEYGNE----VYILGRKNNGKIVERALKSLLEIL 87 (148)
T ss_pred HHHhCChhhcCCC-cCcCceEEEEEEcCCCCE----EEEEecCChHHHHHHHHHHHHHHh
Confidence 3578889999984 778999999998866554 344444466788888888887554
No 52
>PF11460 DUF3007: Protein of unknown function (DUF3007); InterPro: IPR021562 This is a family of uncharacterised proteins found in bacteria and eukaryotes.
Probab=21.27 E-value=1.3e+02 Score=21.91 Aligned_cols=19 Identities=11% Similarity=0.293 Sum_probs=14.7
Q ss_pred HHHHHHhcCCHHHHHHHHH
Q 028703 67 MFESKLYEMTSDQFKNNVN 85 (205)
Q Consensus 67 ~~~~~L~~ls~eeF~~~k~ 85 (205)
++.+.+++||+||.+....
T Consensus 82 ~lqkRle~l~~eE~~~L~~ 100 (104)
T PF11460_consen 82 ELQKRLEELSPEELEALQA 100 (104)
T ss_pred HHHHHHHhCCHHHHHHHHH
Confidence 3677889999999877644
No 53
>KOG2019 consensus Metalloendoprotease HMP1 (insulinase superfamily) [General function prediction only; Posttranslational modification, protein turnover, chaperones]
Probab=21.18 E-value=4.6e+02 Score=26.00 Aligned_cols=142 Identities=11% Similarity=0.110 Sum_probs=87.6
Q ss_pred hHHHHHHHHHch-HHHHHhhhccccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHH--HHHHHhcCCHHH
Q 028703 3 VKLQLLALIAKQ-PAFHQLRTVEQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQM--FESKLYEMTSDQ 79 (205)
Q Consensus 3 a~~~Ll~~ils~-~~f~~LRTkqQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~--~~~~L~~ls~ee 79 (205)
|-+.+|+.+|.. .+.++.|.+ +=.|--+|.+....|+..+.-| .+|+-|. -++.|=.. |..-+ ..+.++
T Consensus 841 asl~vlS~~lt~k~Lh~evRek-GGAYGgg~s~~sh~GvfSf~SY-----RDpn~lk-tL~~f~~tgd~~~~~-~~~~~d 912 (998)
T KOG2019|consen 841 ASLQVLSKLLTNKWLHDEVREK-GGAYGGGCSYSSHSGVFSFYSY-----RDPNPLK-TLDIFDGTGDFLRGL-DVDQQD 912 (998)
T ss_pred cHHHHHHHHHHHHHHHHHHHHh-cCccCCccccccccceEEEEec-----cCCchhh-HHHhhcchhhhhhcC-Cccccc
Confidence 456677776554 677888877 4468888888888888777655 5555443 34444322 11122 377888
Q ss_pred HHHHHHHHHHHHhccCcChHHHH-HHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHHHHhhhcCCCCccEEEEEEee
Q 028703 80 FKNNVNALIDMKLEKHKNLKEES-GFYWREISDGILKFDRREVEVAALRQLTQQELIYFFNENIKAGAPRKKTLSVRVYG 158 (205)
Q Consensus 80 F~~~k~~li~~l~~~~~sl~~~~-~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~~~~~~~~~~~~~~l~i~v~~ 158 (205)
+.++|-+.++..-.+ +.-.++. .|+ ..|--+ +.+|..-+.|-.++..|+.++.+.++. ...+-..|-|-|
T Consensus 913 ldeAkl~~f~~VDap-~~P~~kG~~~f----l~gvtD-emkQarREqll~vSl~d~~~vae~yl~---~~~~~~~vav~g 983 (998)
T KOG2019|consen 913 LDEAKLGTFGDVDAP-QLPDAKGLLRF----LLGVTD-EMKQARREQLLAVSLKDFKAVAEAYLG---VGDKGVAVAVAG 983 (998)
T ss_pred hhhhhhhhcccccCC-cCCcccchHHH----HhcCCH-HHHHHHHHHHHhhhHHHHHHHHHHHhc---cCCcceEEEeeC
Confidence 888888888775433 2222222 122 222111 456777788999999999999999983 223345555555
Q ss_pred CCC
Q 028703 159 SLH 161 (205)
Q Consensus 159 ~~~ 161 (205)
+..
T Consensus 984 ~E~ 986 (998)
T KOG2019|consen 984 PED 986 (998)
T ss_pred ccC
Confidence 543
No 54
>PF11264 ThylakoidFormat: Thylakoid formation protein; InterPro: IPR017499 Psp29, originally designated sll1414 (P73956 from SWISSPROT) in Synechocystis sp. (strain PCC 6803), is found universally in Cyanobacteria and in Arabidopsis. It was isolated and partially sequenced from purified photosystem II (PS II) in Synechocystis. While its function is unknown, mutant studies show an impairment in photosystem II biogenesis and/or stability, rather than in PS II core function.; GO: 0010027 thylakoid membrane organization, 0015979 photosynthesis, 0009523 photosystem II
Probab=20.65 E-value=4.1e+02 Score=21.95 Aligned_cols=58 Identities=9% Similarity=0.110 Sum_probs=39.8
Q ss_pred HHHHHHHhcC-CHHHHHHHHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHHHH
Q 028703 66 QMFESKLYEM-TSDQFKNNVNALIDMKLEKHKNLKEESGFYWREISDGILKFDRREVEVAALRQLTQQELIYFFNE 140 (205)
Q Consensus 66 ~~~~~~L~~l-s~eeF~~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~~~ 140 (205)
..|...+++. ++++-..+-+++++.+...|..+.+.+.. ..+.++..+.+|+.+|+..
T Consensus 53 t~fd~fm~GY~p~~~~~~If~Alc~a~~~dp~~~r~dA~~-----------------l~~~a~~~s~~~l~~~l~~ 111 (216)
T PF11264_consen 53 TVFDRFMQGYPPEEDKDSIFNALCQALGFDPEQYRQDAEK-----------------LEEWAKGKSIEDLLSWLSQ 111 (216)
T ss_pred HHHHHHhcCCCChhHHHHHHHHHHHHcCCCHHHHHHHHHH-----------------HHHHHHcCCHHHHHHHHhc
Confidence 3344555666 67778888889998888888776666543 3344567777777777754
No 55
>PRK03573 transcriptional regulator SlyA; Provisional
Probab=20.51 E-value=3.6e+02 Score=19.85 Aligned_cols=60 Identities=12% Similarity=0.186 Sum_probs=40.1
Q ss_pred cccceEEEEEEeeeCCeeEEEEEEeCCCCChhHHHHHHHHHHHHHHH-HHhcCCHHHHHHHHHHHH
Q 028703 24 EQLGYITALLQRNDFGIHGVQFIIQSSVKGPKYIDLRVESFLQMFES-KLYEMTSDQFKNNVNALI 88 (205)
Q Consensus 24 qQLGYvV~s~~~~~~~~~gl~~~VQS~~~~~~~l~~~i~~Fl~~~~~-~L~~ls~eeF~~~k~~li 88 (205)
+..||+.-..........-+.++ -.-..+...+......+.. .+..+++++.+.....+.
T Consensus 71 e~~GlV~r~~~~~DrR~~~l~LT-----~~G~~~~~~~~~~~~~~~~~~~~~l~~ee~~~l~~~l~ 131 (144)
T PRK03573 71 EEKGLISRQTCASDRRAKRIKLT-----EKAEPLISEVEAVINKTRAEILHGISAEEIEQLITLIA 131 (144)
T ss_pred HHCCCEeeecCCCCcCeeeeEEC-----hHHHHHHHHHHHHHHHHHHHHHhCCCHHHHHHHHHHHH
Confidence 56688887766555555555444 2345566677777777655 457999999888776544
No 56
>cd08805 Death_ank1 Death domain of Ankyrin-1. Death Domain (DD) of the human protein ankyrin-1 (ANK-1) and related proteins. Ankyrins are modular proteins comprising three conserved domains, an N-terminal membrane-binding domain containing ANK repeats, a spectrin-binding domain and a C-terminal DD. ANK-1, also called ankyrin-R (for restricted), is found in brain, muscle, and erythrocytes and is thought to function in linking integral membrane proteins to the underlying cytoskeleton. It plays a critical nonredundant role in erythroid development and is associated with hereditary spherocytosis (HS), a common disorder of the red cell membrane. The small alternatively-spliced variant, sANK-1, found in striated muscle and concentrated in the sarcoplasmic reticulum (SR) binds obscurin and titin, which facilitates the anchoring of the network SR to the contractile apparatus. In general, DDs are protein-protein interaction domains found in a variety of domain architectures. Their common featur
Probab=20.46 E-value=3e+02 Score=19.00 Aligned_cols=58 Identities=12% Similarity=0.078 Sum_probs=39.2
Q ss_pred cCCHHHHHHHHHHHHHHHhccCcChHHHHHHhHHHHhcCCCCccccHHHHHHHhcCCHHHHHHHH
Q 028703 74 EMTSDQFKNNVNALIDMKLEKHKNLKEESGFYWREISDGILKFDRREVEVAALRQLTQQELIYFF 138 (205)
Q Consensus 74 ~ls~eeF~~~k~~li~~l~~~~~sl~~~~~~~w~~I~~~~~~F~~~~~~i~~l~~it~~dl~~f~ 138 (205)
++|+.+....+. +.|.|+.+++...-.....+....-..+.++.+|+++.+.|+....
T Consensus 26 ~vs~~dI~~I~~-------e~p~~l~~Q~~~~L~~W~~r~g~~At~~~L~~AL~~i~R~div~~~ 83 (84)
T cd08805 26 QFSVEDINRIRV-------ENPNSLLEQSTALLNLWVDREGENAKMSPLYPALYSIDRLTIVNML 83 (84)
T ss_pred CCCHHHHHHHHH-------hCCCCHHHHHHHHHHHHHHhcCccchHHHHHHHHHHCChHHHHHhh
Confidence 578877777653 4466677766655444444433434557799999999999998754
No 57
>cd08321 Pyrin_ASC-like Pyrin Death Domain found in ASC. Pyrin Death Domain found in ASC (Apoptosis-associated speck-like protein containing a CARD) and similar proteins. ASC is an adaptor molecule that functions in the assembly of the 'inflammasome', a multiprotein platform, which is responsible for caspase-1 activation and regulation of IL-1beta maturation. ASC contains two domains from the Death Domain (DD) superfamily, an N-terminal pyrin-like domain and a C-terminal Caspase activation and recruitment domain (CARD). Through these 2 domains, ASC serves as an adaptor for inflammasome integrity and oligomerizes to form supramolecular assemblies. Other members of this subfamily are associated with ATPase domains and their function remains unknown. In general, Pyrin is a subfamily of the DD superfamily and functions in several signaling pathways. DDs are protein-protein interaction domains found in a variety of domain architectures. Their common feature is that they form homodimers by se
Probab=20.13 E-value=1.4e+02 Score=20.45 Aligned_cols=24 Identities=25% Similarity=0.279 Sum_probs=19.8
Q ss_pred HHHHhcCCHHHHHHHHHHHHHHHh
Q 028703 69 ESKLYEMTSDQFKNNVNALIDMKL 92 (205)
Q Consensus 69 ~~~L~~ls~eeF~~~k~~li~~l~ 92 (205)
...|++|+++||.+.|.-|.+...
T Consensus 5 l~~Le~L~~~ElkkFK~~L~~~~~ 28 (82)
T cd08321 5 LDALEDLEEDELKKFKWKLRDIPL 28 (82)
T ss_pred HHHHHHhCHHHHHHHHHHHhhhhh
Confidence 456889999999999998887643
Done!