Query 006344
Match_columns 649
No_of_seqs 157 out of 195
Neff 6.1
Searched_HMMs 46136
Date Thu Mar 28 21:55:18 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/006344.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/006344hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF13320 DUF4091: Domain of un 99.9 4.1E-25 8.8E-30 183.1 7.6 68 533-606 1-68 (68)
2 PF10633 NPCBM_assoc: NPCBM-as 95.1 0.1 2.2E-06 44.2 7.7 32 154-185 45-76 (78)
3 PF01229 Glyco_hydro_39: Glyco 84.5 3.3 7.2E-05 47.2 8.5 109 313-438 81-202 (486)
4 COG1470 Predicted membrane pro 83.8 7.7 0.00017 43.7 10.4 37 150-186 324-360 (513)
5 COG1470 Predicted membrane pro 79.7 16 0.00034 41.3 11.0 33 154-186 437-469 (513)
6 PF06030 DUF916: Bacterial pro 76.4 66 0.0014 29.9 13.3 107 63-186 5-120 (121)
7 PF15418 DUF4625: Domain of un 70.8 4.8 0.0001 38.1 3.7 91 76-188 28-119 (132)
8 PF14352 DUF4402: Domain of un 63.3 7.6 0.00016 36.1 3.4 33 154-186 94-128 (130)
9 PF02221 E1_DerP2_DerF2: ML do 62.2 10 0.00022 34.7 4.1 35 152-186 85-119 (134)
10 PF01835 A2M_N: MG2 domain; I 57.9 48 0.001 28.8 7.4 26 160-185 61-86 (99)
11 PF02449 Glyco_hydro_42: Beta- 54.7 1.5E+02 0.0032 32.5 12.1 87 494-600 286-373 (374)
12 PF06280 DUF1034: Fn3-like dom 53.7 12 0.00025 33.8 2.8 36 151-186 62-100 (112)
13 PF00150 Cellulase: Cellulase 52.8 49 0.0011 33.7 7.7 102 312-439 59-172 (281)
14 PF13731 WxL: WxL domain surfa 52.0 43 0.00093 33.9 6.9 79 106-185 105-210 (215)
15 smart00633 Glyco_10 Glycosyl h 51.5 50 0.0011 34.2 7.5 98 311-438 11-124 (254)
16 COG5520 O-Glycosyl hydrolase [ 50.7 1.7E+02 0.0036 32.5 11.2 145 382-544 151-310 (433)
17 cd00917 PG-PI_TP The phosphati 42.9 44 0.00094 30.8 4.9 34 152-186 76-109 (122)
18 PLN00180 NDF6 (NDH-dependent f 36.5 46 0.00099 32.3 3.9 87 514-620 90-177 (180)
19 PF09087 Cyc-maltodext_N: Cycl 34.3 2.2E+02 0.0048 25.2 7.6 20 163-183 51-70 (88)
20 KOG1579 Homocysteine S-methylt 34.0 2.5E+02 0.0053 30.6 9.3 122 367-506 132-259 (317)
21 smart00737 ML Domain involved 27.1 1.2E+02 0.0026 27.3 5.0 35 152-186 71-105 (118)
22 PF13204 DUF4038: Protein of u 25.7 8.2E+02 0.018 25.9 12.1 198 305-546 78-286 (289)
23 PF04234 CopC: CopC domain; I 25.6 64 0.0014 28.4 2.9 26 165-191 61-86 (97)
24 PF09099 Qn_am_d_aIII: Quinohe 24.9 69 0.0015 27.9 2.8 21 162-182 49-69 (81)
25 TIGR03769 P_ac_wall_RPT actino 22.9 96 0.0021 23.4 2.9 14 173-186 10-23 (41)
26 PF00868 Transglut_N: Transglu 21.8 6.3E+02 0.014 23.2 9.4 31 155-185 87-117 (118)
27 PRK09778 putative antitoxin of 21.4 1.8E+02 0.0038 26.2 4.6 26 585-610 43-68 (97)
28 PRK10301 hypothetical protein; 21.2 1.2E+02 0.0026 28.2 3.9 25 165-190 88-112 (124)
29 PF09608 Alph_Pro_TM: Putative 21.1 2E+02 0.0044 29.9 5.9 38 151-192 147-184 (236)
30 PF08428 Rib: Rib/alpha-like r 20.3 1.3E+02 0.0027 24.9 3.4 33 154-190 20-52 (65)
No 1
>PF13320 DUF4091: Domain of unknown function (DUF4091)
Probab=99.91 E-value=4.1e-25 Score=183.14 Aligned_cols=68 Identities=47% Similarity=0.862 Sum_probs=63.7
Q ss_pred cCCCEEEEeecccccCCCCCCccccccCCCCCCceEEEccCCcCCCCCCceechhHHHHHHHHHHHHHHHHHHh
Q 006344 533 EGGTGFLYWGANCYEKATVPSAEIRFRRGLPPGDGVLFYPGEVFSSSRQPVASLRLERILSGLQDIEYLNLYAS 606 (649)
Q Consensus 533 ~g~~GfL~W~~n~w~~~~~P~~d~~~~~~~~~GDg~LVYPG~~~~~~~~Pv~SiRle~lReGieDye~L~lL~~ 606 (649)
||++|||||+||+|+++ |+.+++++. |++||++|||||++ .++|++|||||+||+||||||||++|++
T Consensus 1 y~~~G~L~W~~~~w~~d--P~~d~~~~~-~~~GD~~lvYPg~~---~~~p~~SiRle~lr~G~qD~e~l~~l~~ 68 (68)
T PF13320_consen 1 YGFDGFLRWAYNFWNED--PWEDTRFRG-FPAGDGFLVYPGED---TGGPVSSIRLEVLREGIQDYEYLRLLEK 68 (68)
T ss_pred CCCCeEEEecccccccC--cccccCcCc-CCCCCeEEEecCCC---CCCcccCHHHHHHHHHHHHHHHHHHHhC
Confidence 68999999999999876 999999997 99999999999982 3899999999999999999999999985
No 2
>PF10633 NPCBM_assoc: NPCBM-associated, NEW3 domain of alpha-galactosidase; InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=95.13 E-value=0.1 Score=44.19 Aligned_cols=32 Identities=31% Similarity=0.543 Sum_probs=25.9
Q ss_pred eecCCCeeEEEEEEEcCCCCCCceeEEEEEEE
Q 006344 154 SLIPGETTAVWVSIDAPYAQPPGLYEGEIIIT 185 (649)
Q Consensus 154 ~v~~g~~q~lWv~v~VP~~a~pG~Y~G~i~v~ 185 (649)
.|++|+.+.+=++|.+|+++.||.|+.+++++
T Consensus 45 ~l~pG~s~~~~~~V~vp~~a~~G~y~v~~~a~ 76 (78)
T PF10633_consen 45 SLPPGESVTVTFTVTVPADAAPGTYTVTVTAR 76 (78)
T ss_dssp -B-TTSEEEEEEEEEE-TT--SEEEEEEEEEE
T ss_pred cCCCCCEEEEEEEEECCCCCCCceEEEEEEEE
Confidence 68899999999999999999999999999986
No 3
>PF01229 Glyco_hydro_39: Glycosyl hydrolases family 39; InterPro: IPR000514 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. Glycoside hydrolase family 39 GH39 from CAZY comprises enzymes with several known activities; alpha-L-iduronidase (3.2.1.76 from EC); beta-xylosidase (3.2.1.37 from EC). The most highly conserved regions in these enzymes are located in their N-terminal sections. These contain a glutamic acid residue which, on the basis of similarities with other families of glycosyl hydrolases [], probably acts as the proton donor in their catalytic mechanism.; GO: 0004553 hydrolase activity, hydrolyzing O-glycosyl compounds, 0005975 carbohydrate metabolic process; PDB: 2BS9_D 2BFG_E 1W91_B 1UHV_D 1PX8_A.
Probab=84.55 E-value=3.3 Score=47.19 Aligned_cols=109 Identities=24% Similarity=0.348 Sum_probs=62.7
Q ss_pred H-HHHHHHHHHHHHhCCcCccccccCCcceeeeccCCCCCCCCCCccccCccccccccccccCCCCCchhHHHHHHHHHH
Q 006344 313 W-YEALDQHFKWLLQYRISPFFCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYVRKEIE 391 (649)
Q Consensus 313 ~-f~~ldr~~~~~~~~ris~lf~~Wg~~~~i~~y~~pw~~~~~k~~~~f~d~~~~~Y~~~~~~~l~~~d~~~~~L~~~~~ 391 (649)
| |+.+|+-++++++++|.|++ ..|- ...+. +. +.. ..|. |.....|. ..-+.|.+++++|++
T Consensus 81 Ynf~~lD~i~D~l~~~g~~P~v-el~f--~p~~~-----~~-~~~-~~~~------~~~~~~pp-~~~~~W~~lv~~~~~ 143 (486)
T PF01229_consen 81 YNFTYLDQILDFLLENGLKPFV-ELGF--MPMAL-----AS-GYQ-TVFW------YKGNISPP-KDYEKWRDLVRAFAR 143 (486)
T ss_dssp E--HHHHHHHHHHHHCT-EEEE-EE-S--B-GGG-----BS-S---EETT------TTEE-S-B-S-HHHHHHHHHHHHH
T ss_pred CChHHHHHHHHHHHHcCCEEEE-EEEe--chhhh-----cC-CCC-cccc------ccCCcCCc-ccHHHHHHHHHHHHH
Confidence 5 99999999999999999853 1110 00000 00 000 0010 11011111 122469999999999
Q ss_pred HHHhc-cc--ccceeeeecCCCCCc---------cchHHHHHHHHHHHHhCCCCeEEEe
Q 006344 392 LLRTK-AH--WKKAYFYLWDEPLNM---------EHYSSVRNMASELHAYAPDARVLTT 438 (649)
Q Consensus 392 hL~~k-G~--~~~~y~~i~DEP~~~---------~~~~~~r~~~~~ir~~~P~~kil~t 438 (649)
|+..+ |. ....+|=++.||... +=++.|+..++.||++.|++||-..
T Consensus 144 h~~~RYG~~ev~~W~fEiWNEPd~~~f~~~~~~~ey~~ly~~~~~~iK~~~p~~~vGGp 202 (486)
T PF01229_consen 144 HYIDRYGIEEVSTWYFEIWNEPDLKDFWWDGTPEEYFELYDATARAIKAVDPELKVGGP 202 (486)
T ss_dssp HHHHHHHHHHHTTSEEEESS-TTSTTTSGGG-HHHHHHHHHHHHHHHHHH-TTSEEEEE
T ss_pred HHHhhcCCccccceeEEeCcCCCcccccCCCCHHHHHHHHHHHHHHHHHhCCCCcccCc
Confidence 99754 43 335577789999532 2345678899999999999998764
No 4
>COG1470 Predicted membrane protein [Function unknown]
Probab=83.83 E-value=7.7 Score=43.75 Aligned_cols=37 Identities=30% Similarity=0.476 Sum_probs=34.6
Q ss_pred CcceeecCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 006344 150 VCQISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (649)
Q Consensus 150 ~~~~~v~~g~~q~lWv~v~VP~~a~pG~Y~G~i~v~~ 186 (649)
...+.+.+|+...+-++|+.|++|.||.|..+|+++.
T Consensus 324 vt~vkL~~gE~kdvtleV~ps~na~pG~Ynv~I~A~s 360 (513)
T COG1470 324 VTSVKLKPGEEKDVTLEVYPSLNATPGTYNVTITASS 360 (513)
T ss_pred EEEEEecCCCceEEEEEEecCCCCCCCceeEEEEEec
Confidence 3578899999999999999999999999999999985
No 5
>COG1470 Predicted membrane protein [Function unknown]
Probab=79.66 E-value=16 Score=41.32 Aligned_cols=33 Identities=36% Similarity=0.457 Sum_probs=30.5
Q ss_pred eecCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 006344 154 SLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (649)
Q Consensus 154 ~v~~g~~q~lWv~v~VP~~a~pG~Y~G~i~v~~ 186 (649)
.+.||+...+=++|.||++|.+|.|..+|+.++
T Consensus 437 sL~pge~~tV~ltI~vP~~a~aGdY~i~i~~ks 469 (513)
T COG1470 437 SLEPGESKTVSLTITVPEDAGAGDYRITITAKS 469 (513)
T ss_pred ccCCCCcceEEEEEEcCCCCCCCcEEEEEEEee
Confidence 468899999999999999999999999999985
No 6
>PF06030 DUF916: Bacterial protein of unknown function (DUF916); InterPro: IPR010317 This family consists of putative cell surface proteins, from Firmicutes, of unknown function.
Probab=76.35 E-value=66 Score=29.88 Aligned_cols=107 Identities=17% Similarity=0.231 Sum_probs=68.1
Q ss_pred ccCCCCCCCC-CCceeEEeecCceeEEEEEEecCcccCCCCCCcceEEEEcc-ccCCCCCccccccceEEEEeeccCC--
Q 006344 63 NVGPQEMPRP-LEPINLLAARNERESVQIALRPKVSWSSSSTAGVVQVQCSD-LCSASGDRLVVGQSLMLRRVVPMLG-- 138 (649)
Q Consensus 63 kV~p~~~p~~-~~~~~l~aarnE~~sfQi~l~~~~~~~~~~~~~~V~v~~sd-L~s~~G~~~i~~~~i~~~~v~~vpg-- 138 (649)
-|.|+..-.. ...+.|.+.-|+...+|+-+.-. ....-.|.|++.+ .++.+|. +.|.+-
T Consensus 5 p~~p~~Q~~~~~~YFdL~~~P~q~~~l~v~i~N~-----s~~~~tv~v~~~~A~Tn~nG~------------I~Y~~~~~ 67 (121)
T PF06030_consen 5 PVLPENQIDKNVSYFDLKVKPGQKQTLEVRITNN-----SDKEITVKVSANTATTNDNGV------------IDYSQNNP 67 (121)
T ss_pred ecCCccccCCCCCeEEEEeCCCCEEEEEEEEEeC-----CCCCEEEEEEEeeeEecCCEE------------EEECCCCc
Confidence 4566665433 57899999999999999999753 1223344444332 1222221 122211
Q ss_pred CCCc--cccCC---CCCcceeecCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 006344 139 VPDA--LVPLD---LPVCQISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (649)
Q Consensus 139 ~PD~--L~P~~---~~~~~~~v~~g~~q~lWv~v~VP~~a~pG~Y~G~i~v~~ 186 (649)
-.|. -.++. .....++|+|++.+-+=++|.+|+..-.|+.-|.|.|+.
T Consensus 68 ~~d~sl~~~~~~~v~~~~~Vtl~~~~sk~V~~~i~~P~~~f~G~ilGGi~~~e 120 (121)
T PF06030_consen 68 KKDKSLKYPFSDLVKIPKEVTLPPNESKTVTFTIKMPKKAFDGIILGGIYFSE 120 (121)
T ss_pred ccCcccCcchHHhccCCcEEEECCCCEEEEEEEEEcCCCCcCCEEEeeEEEEe
Confidence 1121 11211 011349999999999999999999999999999999984
No 7
>PF15418 DUF4625: Domain of unknown function (DUF4625)
Probab=70.84 E-value=4.8 Score=38.06 Aligned_cols=91 Identities=18% Similarity=0.189 Sum_probs=52.3
Q ss_pred eeEEeecCceeEEEEEEecCcccCCCCCCcceEEEEccccCCCCCccccccceEEEEeeccCCCCCccccCCCCCcceee
Q 006344 76 INLLAARNERESVQIALRPKVSWSSSSTAGVVQVQCSDLCSASGDRLVVGQSLMLRRVVPMLGVPDALVPLDLPVCQISL 155 (649)
Q Consensus 76 ~~l~aarnE~~sfQi~l~~~~~~~~~~~~~~V~v~~sdL~s~~G~~~i~~~~i~~~~v~~vpg~PD~L~P~~~~~~~~~v 155 (649)
-.-.+-||+...|..-+.+. ..++.++|++- +. .+.-..+.. -.+...|+.- .....+
T Consensus 28 ~~~~~~~G~~ihfe~~i~d~------~~i~si~VeIH---~n-fd~H~h~~~-----------~~~~~~~~~~-~~~~~~ 85 (132)
T PF15418_consen 28 NCKVATRGDDIHFEADISDN------SAIKSIKVEIH---NN-FDHHTHSTE-----------AGECEKPWVF-EQDYDI 85 (132)
T ss_pred CCeEEecCCcEEEEEEEEcc------cceeEEEEEEe---cC-cCccccccc-----------ccccccCcEE-EEEEcc
Confidence 34456899999999999875 46888888872 10 000000000 0000111110 001122
Q ss_pred cCC-CeeEEEEEEEcCCCCCCceeEEEEEEEecc
Q 006344 156 IPG-ETTAVWVSIDAPYAQPPGLYEGEIIITSKA 188 (649)
Q Consensus 156 ~~g-~~q~lWv~v~VP~~a~pG~Y~G~i~v~~~~ 188 (649)
..| ...-+=..|.||++++||.|...|+|+.++
T Consensus 86 ~~g~~~~~~h~~i~IPa~a~~G~YH~~i~VtD~~ 119 (132)
T PF15418_consen 86 YGGKKNYDFHEHIDIPADAPAGDYHFMITVTDAA 119 (132)
T ss_pred cCCcccEeEEEeeeCCCCCCCcceEEEEEEEECC
Confidence 222 245677899999999999999999999633
No 8
>PF14352 DUF4402: Domain of unknown function (DUF4402)
Probab=63.29 E-value=7.6 Score=36.07 Aligned_cols=33 Identities=27% Similarity=0.506 Sum_probs=25.9
Q ss_pred eecCCCeeEEEE--EEEcCCCCCCceeEEEEEEEe
Q 006344 154 SLIPGETTAVWV--SIDAPYAQPPGLYEGEIIITS 186 (649)
Q Consensus 154 ~v~~g~~q~lWv--~v~VP~~a~pG~Y~G~i~v~~ 186 (649)
.+..+....+.| ++.|+.++++|.|+|+++|+.
T Consensus 94 ~~~~~g~~~~~VGGtL~v~~~~~~G~YsGt~~VtV 128 (130)
T PF14352_consen 94 TLDTGGSATFNVGGTLNVPANQAAGTYSGTFTVTV 128 (130)
T ss_pred EecCCCcEEEEEEEEEEcCCCCCCeEEEEEEEEEE
Confidence 334455566666 589999999999999999984
No 9
>PF02221 E1_DerP2_DerF2: ML domain; InterPro: IPR003172 The MD-2-related lipid-recognition (ML) domain is implicated in lipid recognition, particularly in the recognition of pathogen related products. It has an immunoglobulin-like beta-sandwich fold similar to that of E-set Ig domains. This domain is present in the following proteins: Epididymal secretory protein E1 (also known as Niemann-Pick C2 protein), which is known to bind cholesterol. Niemann-Pick disease type C2 is a fatal hereditary disease characterised by accumulation of low-density lipoprotein-derived cholesterol in lysosomes []. House-dust mite allergen proteins such as Der f 2 from Dermatophagoides farinae and Der p 2 from Dermatophagoides pteronyssinus []. ; PDB: 2AG9_B 1G13_B 2AG2_B 2AG4_A 1TJJ_C 1PU5_C 1PUB_A 2AF9_A 3T6Q_D 3M7O_B ....
Probab=62.24 E-value=10 Score=34.74 Aligned_cols=35 Identities=26% Similarity=0.338 Sum_probs=32.6
Q ss_pred ceeecCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 006344 152 QISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (649)
Q Consensus 152 ~~~v~~g~~q~lWv~v~VP~~a~pG~Y~G~i~v~~ 186 (649)
...+.+|+....-+++.||...++|.|+++++++.
T Consensus 85 ~CPi~~G~~~~~~~~~~i~~~~p~~~~~i~~~l~d 119 (134)
T PF02221_consen 85 SCPIKAGEYYTYTYTIPIPKIYPPGKYTIQWKLTD 119 (134)
T ss_dssp TSTBTTTEEEEEEEEEEESTTSSSEEEEEEEEEEE
T ss_pred cCccCCCcEEEEEEEEEcccceeeEEEEEEEEEEe
Confidence 56788999999999999999999999999999996
No 10
>PF01835 A2M_N: MG2 domain; InterPro: IPR002890 The proteinase-binding alpha-macroglobulins (A2M) [] are large glycoproteins found in the plasma of vertebrates, in the hemolymph of some invertebrates and in reptilian and avian egg white. A2M-like proteins are able to inhibit all four classes of proteinases by a 'trapping' mechanism. They have a peptide stretch, called the 'bait region', which contains specific cleavage sites for different proteinases. When a proteinase cleaves the bait region, a conformational change is induced in the protein, thus trapping the proteinase. The entrapped enzyme remains active against low molecular weight substrates, whilst its activity toward larger substrates is greatly reduced, due to steric hindrance. Following cleavage in the bait region, a thiol ester bond, formed between the side chains of a cysteine and a glutamine, is cleaved and mediates the covalent binding of the A2M-like protein to the proteinase. This family includes the N-terminal region of the alpha-2-macroglobulin family. The inhibitor domains belong to MEROPS inhibitor family I39.; GO: 0004866 endopeptidase inhibitor activity; PDB: 2B39_B 3KLS_B 3PRX_C 3KM9_B 3PVM_C 3CU7_A 4E0S_A 4A5W_A 4ACQ_C 2P9R_B ....
Probab=57.93 E-value=48 Score=28.81 Aligned_cols=26 Identities=19% Similarity=0.180 Sum_probs=17.3
Q ss_pred eeEEEEEEEcCCCCCCceeEEEEEEE
Q 006344 160 TTAVWVSIDAPYAQPPGLYEGEIIIT 185 (649)
Q Consensus 160 ~q~lWv~v~VP~~a~pG~Y~G~i~v~ 185 (649)
.-.+-.++.+|+++..|.|+.++...
T Consensus 61 ~G~~~~~~~lp~~~~~G~y~i~~~~~ 86 (99)
T PF01835_consen 61 NGIFSGSFQLPDDAPLGTYTIRVKTD 86 (99)
T ss_dssp TTEEEEEEE--SS---EEEEEEEEET
T ss_pred CCEEEEEEECCCCCCCEeEEEEEEEc
Confidence 33567789999999999999999885
No 11
>PF02449 Glyco_hydro_42: Beta-galactosidase; InterPro: IPR013529 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. This group of beta-galactosidase enzymes (3.2.1.23 from EC) belong to the glycosyl hydrolase 42 family GH42 from CAZY. The enzyme catalyses the hydrolysis of terminal, non-reducing terminal beta-D-galactosidase residues.; GO: 0004565 beta-galactosidase activity, 0005975 carbohydrate metabolic process, 0009341 beta-galactosidase complex; PDB: 1KWK_A 1KWG_A 3U7V_A.
Probab=54.67 E-value=1.5e+02 Score=32.52 Aligned_cols=87 Identities=17% Similarity=0.326 Sum_probs=41.1
Q ss_pred CCCceeEEEe-cCCCCCCCCCcccCCchhHHHHHHHHHHHcCCCEEEEeecccccCCCCCCccccccCCCCCCceEEEcc
Q 006344 494 ENGEEWWTYV-CMGPSDPHPNWHLGMRGSQHRAVMWRVWKEGGTGFLYWGANCYEKATVPSAEIRFRRGLPPGDGVLFYP 572 (649)
Q Consensus 494 ~~G~~~W~Y~-C~~p~~~~pN~fid~p~~~~R~~gW~~~k~g~~GfL~W~~n~w~~~~~P~~d~~~~~~~~~GDg~LVYP 572 (649)
+.|++.|.=- +.++ ..+...-..-.+-+.|...|++..+|.+|.++|.+...... .-.|..+.-
T Consensus 286 ~~~kpf~v~E~~~g~-~~~~~~~~~~~pg~~~~~~~~~~A~Ga~~i~~~~wr~~~~g-----~E~~~~g~~--------- 350 (374)
T PF02449_consen 286 AKGKPFWVMEQQPGP-VNWRPYNRPPRPGELRLWSWQAIAHGADGILFWQWRQSRFG-----AEQFHGGLV--------- 350 (374)
T ss_dssp TTT--EEEEEE--S---SSSSS-----TTHHHHHHHHHHHTT-S-EEEC-SB--SSS-----TTTTS--SB---------
T ss_pred cCCCceEeecCCCCC-CCCccCCCCCCCCHHHHHHHHHHHHhCCeeEeeeccCCCCC-----chhhhcccC---------
Confidence 4789888652 3322 12322233344568899999999999999999998665322 111111111
Q ss_pred CCcCCCCCCceechhHHHHHHHHHHHHH
Q 006344 573 GEVFSSSRQPVASLRLERILSGLQDIEY 600 (649)
Q Consensus 573 G~~~~~~~~Pv~SiRle~lReGieDye~ 600 (649)
..++..++.|++-+.+-.++++.
T Consensus 351 -----~~dg~~~~~~~~e~~~~~~~l~~ 373 (374)
T PF02449_consen 351 -----DHDGREPTRRYREVAQLGRELKK 373 (374)
T ss_dssp ------TTS--B-HHHHHHHHHHHHHHT
T ss_pred -----CccCCCCCcHHHHHHHHHHHHhc
Confidence 12334677888877777666553
No 12
>PF06280 DUF1034: Fn3-like domain (DUF1034); InterPro: IPR010435 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This domain of unknown function is present in bacterial and plant peptidases belonging to MEROPS peptidase family S8 (subfamily S8A subtilisin, clan SB). It is C-terminal to and adjacent to the S8 peptidase domain and can be found in conjunction with the PA (Protease associated) domain (IPR003137 from INTERPRO) and additionally in Gram-positive bacteria with the surface protein anchor domain (IPR001899 from INTERPRO).; GO: 0004252 serine-type endopeptidase activity, 0005618 cell wall, 0016020 membrane; PDB: 3EIF_A 1XF1_B.
Probab=53.73 E-value=12 Score=33.82 Aligned_cols=36 Identities=28% Similarity=0.426 Sum_probs=30.7
Q ss_pred cceeecCCCeeEEEEEEEcCCCCCC---ceeEEEEEEEe
Q 006344 151 CQISLIPGETTAVWVSIDAPYAQPP---GLYEGEIIITS 186 (649)
Q Consensus 151 ~~~~v~~g~~q~lWv~v~VP~~a~p---G~Y~G~i~v~~ 186 (649)
..++|+||+.+.+=|+|++|++..+ ..|+|-|.++.
T Consensus 62 ~~vTV~ag~s~~v~vti~~p~~~~~~~~~~~eG~I~~~~ 100 (112)
T PF06280_consen 62 DTVTVPAGQSKTVTVTITPPSGLDASNGPFYEGFITFKS 100 (112)
T ss_dssp EEEEE-TTEEEEEEEEEE--GGGHHTT-EEEEEEEEEES
T ss_pred CeEEECCCCEEEEEEEEEehhcCCcccCCEEEEEEEEEc
Confidence 6799999999999999999998887 99999999995
No 13
>PF00150 Cellulase: Cellulase (glycosyl hydrolase family 5); InterPro: IPR001547 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. Glycoside hydrolase family 5 GH5 from CAZY comprises enzymes with several known activities; endoglucanase (3.2.1.4 from EC); beta-mannanase (3.2.1.78 from EC); exo-1,3-glucanase (3.2.1.58 from EC); endo-1,6-glucanase (3.2.1.75 from EC); xylanase (3.2.1.8 from EC); endoglycoceramidase (3.2.1.123 from EC). The microbial degradation of cellulose and xylans requires several types of enzymes. Fungi and bacteria produces a spectrum of cellulolytic enzymes (cellulases) and xylanases which, on the basis of sequence similarities, can be classified into families. One of these families is known as the cellulase family A [] or as the glycosyl hydrolases family 5 []. One of the conserved regions in this family contains a conserved glutamic acid residue which is potentially involved [] in the catalytic mechanism.; GO: 0004553 hydrolase activity, hydrolyzing O-glycosyl compounds, 0005975 carbohydrate metabolic process; PDB: 3NDY_A 3NDZ_B 1LF1_A 1TVP_B 1TVN_A 3AYR_A 3AYS_A 1QI0_A 1W3K_A 1OCQ_A ....
Probab=52.83 E-value=49 Score=33.72 Aligned_cols=102 Identities=14% Similarity=0.185 Sum_probs=58.5
Q ss_pred hHHHHHHHHHHHHHhCCcCccccccCCcceeeeccCCCCCCCCCCccccCccccccccccccCCCCCchhHHHHHHHHHH
Q 006344 312 EWYEALDQHFKWLLQYRISPFFCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYVRKEIE 391 (649)
Q Consensus 312 ~~f~~ldr~~~~~~~~ris~lf~~Wg~~~~i~~y~~pw~~~~~k~~~~f~d~~~~~Y~~~~~~~l~~~d~~~~~L~~~~~ 391 (649)
.+++.||+-++++.+++|.-+++-... ..|... ... +..... ..+..+++++.++.
T Consensus 59 ~~~~~ld~~v~~a~~~gi~vild~h~~--------~~w~~~-~~~--~~~~~~-------------~~~~~~~~~~~la~ 114 (281)
T PF00150_consen 59 TYLARLDRIVDAAQAYGIYVILDLHNA--------PGWANG-GDG--YGNNDT-------------AQAWFKSFWRALAK 114 (281)
T ss_dssp HHHHHHHHHHHHHHHTT-EEEEEEEES--------TTCSSS-TST--TTTHHH-------------HHHHHHHHHHHHHH
T ss_pred HHHHHHHHHHHHHHhCCCeEEEEeccC--------cccccc-ccc--cccchh-------------hHHHHHhhhhhhcc
Confidence 458999999999999999744321110 112000 000 000000 00113456677777
Q ss_pred HHHhcccccceeeeecCCCCCccc------------hHHHHHHHHHHHHhCCCCeEEEee
Q 006344 392 LLRTKAHWKKAYFYLWDEPLNMEH------------YSSVRNMASELHAYAPDARVLTTY 439 (649)
Q Consensus 392 hL~~kG~~~~~y~~i~DEP~~~~~------------~~~~r~~~~~ir~~~P~~kil~t~ 439 (649)
+++... ..+.+-|+.||..... .+.++++.+.||+..|+..|+...
T Consensus 115 ~y~~~~--~v~~~el~NEP~~~~~~~~w~~~~~~~~~~~~~~~~~~Ir~~~~~~~i~~~~ 172 (281)
T PF00150_consen 115 RYKDNP--PVVGWELWNEPNGGNDDANWNAQNPADWQDWYQRAIDAIRAADPNHLIIVGG 172 (281)
T ss_dssp HHTTTT--TTEEEESSSSGCSTTSTTTTSHHHTHHHHHHHHHHHHHHHHTTSSSEEEEEE
T ss_pred ccCCCC--cEEEEEecCCccccCCccccccccchhhhhHHHHHHHHHHhcCCcceeecCC
Confidence 776332 3556668999964211 356789999999999998877764
No 14
>PF13731 WxL: WxL domain surface cell wall-binding
Probab=52.00 E-value=43 Score=33.92 Aligned_cols=79 Identities=24% Similarity=0.394 Sum_probs=45.6
Q ss_pred ceEEEEccccCCCCCccccccceEEEEeeccC---C--CCCc------cccCCCCCcceeecCCCeeEEE----------
Q 006344 106 VVQVQCSDLCSASGDRLVVGQSLMLRRVVPML---G--VPDA------LVPLDLPVCQISLIPGETTAVW---------- 164 (649)
Q Consensus 106 ~V~v~~sdL~s~~G~~~i~~~~i~~~~v~~vp---g--~PD~------L~P~~~~~~~~~v~~g~~q~lW---------- 164 (649)
.|+|++++|++.+|.. +.+..|.+....... . -|-. |.+-......++-..++-+..|
T Consensus 105 ~L~v~~s~F~~~~~~~-L~ga~l~~~~~~~~~~~~~~~~~~~~~~~~~l~~~~~~~~v~~A~~~~g~G~~~~~~~~~~~~ 183 (215)
T PF13731_consen 105 TLTVKLSPFTNADGDT-LPGATLTFNNGKVQSTANNTNTPTTVSSNITLTPGGQAQTVMSAAKGQGQGTWSYSFGDQDAT 183 (215)
T ss_pred EEEEEeccccccCCcC-cccceEEecCceeEeecccccCCcccccceEeccCCcceeeEeecccccceEEEEEeCCcccc
Confidence 6788888999887665 344455554433221 0 0111 1111110011222345555555
Q ss_pred ----EEEEcCCCCC--CceeEEEEEEE
Q 006344 165 ----VSIDAPYAQP--PGLYEGEIIIT 185 (649)
Q Consensus 165 ----v~v~VP~~a~--pG~Y~G~i~v~ 185 (649)
|.+.||..+. +|.|+++|+=+
T Consensus 184 ~~~~v~L~VP~~~~~~ag~Yt~tlTWt 210 (215)
T PF13731_consen 184 ADTGVSLSVPANTAKQAGTYTATLTWT 210 (215)
T ss_pred cccceEEEeCCCCcccCCcEEEEEEEE
Confidence 8999999998 69999999876
No 15
>smart00633 Glyco_10 Glycosyl hydrolase family 10.
Probab=51.49 E-value=50 Score=34.19 Aligned_cols=98 Identities=12% Similarity=0.182 Sum_probs=58.6
Q ss_pred hhH-HHHHHHHHHHHHhCCcCc--cccccCCcceeeeccCCCCCCCCCCccccCccccccccccccCCCCCchhHHHHHH
Q 006344 311 DEW-YEALDQHFKWLLQYRISP--FFCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYVR 387 (649)
Q Consensus 311 ~~~-f~~ldr~~~~~~~~ris~--lf~~Wg~~~~i~~y~~pw~~~~~k~~~~f~d~~~~~Y~~~~~~~l~~~d~~~~~L~ 387 (649)
+.| |+..|+.++|+.+++|.- +..-|+.. ...| +.....++ -.+.+.+|+.
T Consensus 11 G~~n~~~~D~~~~~a~~~gi~v~gH~l~W~~~------~P~W----------~~~~~~~~----------~~~~~~~~i~ 64 (254)
T smart00633 11 GQFNFSGADAIVNFAKENGIKVRGHTLVWHSQ------TPDW----------VFNLSKET----------LLARLENHIK 64 (254)
T ss_pred CccChHHHHHHHHHHHHCCCEEEEEEEeeccc------CCHh----------hhcCCHHH----------HHHHHHHHHH
Confidence 344 899999999999999871 11234321 1111 11000000 0123566777
Q ss_pred HHHHHHHhcccccceeeeecCCCCCcc-------c----h--HHHHHHHHHHHHhCCCCeEEEe
Q 006344 388 KEIELLRTKAHWKKAYFYLWDEPLNME-------H----Y--SSVRNMASELHAYAPDARVLTT 438 (649)
Q Consensus 388 ~~~~hL~~kG~~~~~y~~i~DEP~~~~-------~----~--~~~r~~~~~ir~~~P~~kil~t 438 (649)
+.++|++.+ -..+ -+..||.+.. . + +.++.+.+.+|++.|++|++..
T Consensus 65 ~v~~ry~g~---i~~w-dV~NE~~~~~~~~~~~~~w~~~~G~~~i~~af~~ar~~~P~a~l~~N 124 (254)
T smart00633 65 TVVGRYKGK---IYAW-DVVNEALHDNGSGLRRSVWYQILGEDYIEKAFRYAREADPDAKLFYN 124 (254)
T ss_pred HHHHHhCCc---ceEE-EEeeecccCCCcccccchHHHhcChHHHHHHHHHHHHhCCCCEEEEe
Confidence 777777644 1112 3678885321 1 2 6688999999999999999886
No 16
>COG5520 O-Glycosyl hydrolase [Cell envelope biogenesis, outer membrane]
Probab=50.70 E-value=1.7e+02 Score=32.53 Aligned_cols=145 Identities=17% Similarity=0.275 Sum_probs=76.2
Q ss_pred HHHHHHHHHHHHHhcccccceeeeecCCCCCccchHH----HHHHHHHHHHhC----CCCeEEEeeccCCCCCCCCCCCc
Q 006344 382 AKDYVRKEIELLRTKAHWKKAYFYLWDEPLNMEHYSS----VRNMASELHAYA----PDARVLTTYYCGPSDAPLGPTPF 453 (649)
Q Consensus 382 ~~~~L~~~~~hL~~kG~~~~~y~~i~DEP~~~~~~~~----~r~~~~~ir~~~----P~~kil~t~~~~p~d~~~~~~~~ 453 (649)
+-+||..|+..++..|.-..+ +.+-.||.-.-.++. ..+..+++++++ ..+||+.- +++
T Consensus 151 yA~~l~~fv~~m~~nGvnlya-lSVQNEPd~~p~~d~~~wtpQe~~rF~~qyl~si~~~~rV~~p------------es~ 217 (433)
T COG5520 151 YADYLNDFVLEMKNNGVNLYA-LSVQNEPDYAPTYDWCWWTPQEELRFMRQYLASINAEMRVIIP------------ESF 217 (433)
T ss_pred HHHHHHHHHHHHHhCCCceeE-EeeccCCcccCCCCcccccHHHHHHHHHHhhhhhccccEEecc------------hhc
Confidence 567899999999999886654 357799954433443 234555666653 34677662 222
Q ss_pred ccccc--cccccCCc----cccccccccccCCchhhhHHHHhhcccCCCceeEEEecCCCCCCCCCcccCCchhHH-HHH
Q 006344 454 ESFVK--VPKFLRPH----TQIYCTSEWVLGNREDLVKDIVTELQPENGEEWWTYVCMGPSDPHPNWHLGMRGSQH-RAV 526 (649)
Q Consensus 454 e~~~~--~p~~~~~~----idi~c~~~wv~~~~~~~~~~~~~~~r~~~G~~~W~Y~C~~p~~~~pN~fid~p~~~~-R~~ 526 (649)
....+ -|.+-+|. ++|. .-||-.++-.++.....+ +...||.+|+=-|..+. .=||.-.- ..... --+
T Consensus 218 ~~~~~~~dp~lnDp~a~a~~~il-g~H~Ygg~v~~~p~~lak--~~~~gKdlwmte~y~~e-sd~~s~dr-~~~~~~~hi 292 (433)
T COG5520 218 KDLPNMSDPILNDPKALANMDIL-GTHLYGGQVSDQPYPLAK--QKPAGKDLWMTECYPPE-SDPNSADR-EALHVALHI 292 (433)
T ss_pred ccccccccccccCHhHhccccee-EeeecccccccchhhHhh--CCCcCCceEEeecccCC-CCCCcchH-HHHHHHHHH
Confidence 11111 12222222 2222 123333443443333332 44569999987787543 23332211 11111 113
Q ss_pred HHHHHHcCCCEEEEeecc
Q 006344 527 MWRVWKEGGTGFLYWGAN 544 (649)
Q Consensus 527 gW~~~k~g~~GfL~W~~n 544 (649)
.--..+-|+.||+.|..-
T Consensus 293 ~~gm~~gg~~ayv~W~i~ 310 (433)
T COG5520 293 HIGMTEGGFQAYVWWNIR 310 (433)
T ss_pred HhhccccCccEEEEEEEe
Confidence 444567789999999853
No 17
>cd00917 PG-PI_TP The phosphatidylinositol/phosphatidylglycerol transfer protein (PG/PI-TP) has been shown to bind phosphatidylglycerol and phosphatidylinositol, but the biological significance of this is still obscure. These proteins belong to the ML domain family.
Probab=42.94 E-value=44 Score=30.80 Aligned_cols=34 Identities=24% Similarity=0.365 Sum_probs=30.4
Q ss_pred ceeecCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 006344 152 QISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (649)
Q Consensus 152 ~~~v~~g~~q~lWv~v~VP~~a~pG~Y~G~i~v~~ 186 (649)
...+.+|+.. +=.++.||...++|.|+++.++..
T Consensus 76 ~CPi~~G~~~-~~~~~~ip~~~P~g~y~v~~~l~d 109 (122)
T cd00917 76 SCPIEPGDKF-LTKLVDLPGEIPPGKYTVSARAYT 109 (122)
T ss_pred cCCcCCCcEE-EEEEeeCCCCCCCceEEEEEEEEC
Confidence 5677889977 888899999999999999999995
No 18
>PLN00180 NDF6 (NDH-dependent flow 6); Provisional
Probab=36.54 E-value=46 Score=32.33 Aligned_cols=87 Identities=21% Similarity=0.339 Sum_probs=56.2
Q ss_pred cccCCchhHHHHHHHHHHHcCCCEEEEeecccccCCCCCCccc-cccCCCCCCceEEEccCCcCCCCCCceechhHHHHH
Q 006344 514 WHLGMRGSQHRAVMWRVWKEGGTGFLYWGANCYEKATVPSAEI-RFRRGLPPGDGVLFYPGEVFSSSRQPVASLRLERIL 592 (649)
Q Consensus 514 ~fid~p~~~~R~~gW~~~k~g~~GfL~W~~n~w~~~~~P~~d~-~~~~~~~~GDg~LVYPG~~~~~~~~Pv~SiRle~lR 592 (649)
|++....+.+-....+.+ --||..|+-+..|||-|. .|+++-+.|-+.-||--. ....+|-|-|+||
T Consensus 90 wHLSD~aiKnVYtfY~mF-------T~WG~~fFgSmKDPfYDSe~YRgdGGDGT~hW~Yd~Q-----Ed~E~sAReeL~R 157 (180)
T PLN00180 90 WHLSDAAIKNVYTFYIMF-------TCWGCLFFGSMKDPFYDSEEYRGDGGDGTGHWVYERQ-----EDIEESARAELWR 157 (180)
T ss_pred hhccHHHHhHHHHHHHHH-------HHHHHhheeccCCcccchHHhcccCCCCceeeEeehH-----HHHHHHHHHHHHH
Confidence 456666666655444433 358877777666798766 466655667777888664 3467899999999
Q ss_pred HHHHHHHHHHHHHhhcCchHHHHHHHHh
Q 006344 593 SGLQDIEYLNLYASRYGRDEGLALLEKT 620 (649)
Q Consensus 593 eGieDye~L~lL~~~~~~~~a~all~~~ 620 (649)
| |+|..++++.|. ++-||+.
T Consensus 158 E-----ELiEEIEQkVGG---LRELEEa 177 (180)
T PLN00180 158 E-----ELIEEIEQKVGG---LRELEEA 177 (180)
T ss_pred H-----HHHHHHHHHhhh---HHHHHHh
Confidence 8 556666665443 3344443
No 19
>PF09087 Cyc-maltodext_N: Cyclomaltodextrinase, N-terminal; InterPro: IPR015171 This domain is found at the N terminus of cyclomaltodextrinase. The domain assumes a beta-sandwich structure composed of the eight antiparallel beta-strands. A ten residue linker is also present at the C-terminal end, which connects the N-terminal domain to a distal domain in the protein. This domain participates in oligomerisation of the protein, wherein the N-terminal domain of one subunit contacts the active centre of the other subunit, and is also required for binding of cyclodextrin to substrate []. ; PDB: 3EDK_B 3EDD_A 3EDJ_B 3EDE_A 1H3G_B 3EDF_B.
Probab=34.35 E-value=2.2e+02 Score=25.19 Aligned_cols=20 Identities=20% Similarity=0.436 Sum_probs=13.3
Q ss_pred EEEEEEcCCCCCCceeEEEEE
Q 006344 163 VWVSIDAPYAQPPGLYEGEII 183 (649)
Q Consensus 163 lWv~v~VP~~a~pG~Y~G~i~ 183 (649)
|-|+++|- +|+||+++..++
T Consensus 51 LFv~L~i~-~akpg~~~i~~~ 70 (88)
T PF09087_consen 51 LFVYLDIS-DAKPGTFTINFK 70 (88)
T ss_dssp EEEEEEE--T--SEEEEEEEE
T ss_pred EEEEEecC-CCCCcEEEEEEE
Confidence 56777777 999999887766
No 20
>KOG1579 consensus Homocysteine S-methyltransferase [Amino acid transport and metabolism]
Probab=33.95 E-value=2.5e+02 Score=30.60 Aligned_cols=122 Identities=13% Similarity=0.170 Sum_probs=72.9
Q ss_pred cccccccCCCCCchhHHHHHHHHHHHHHhcccccc-eeeeecCCCCCccchHHHHHHHHHHHHhCCCCeEEEeeccCCCC
Q 006344 367 AYAVPYSPVLSSNDGAKDYVRKEIELLRTKAHWKK-AYFYLWDEPLNMEHYSSVRNMASELHAYAPDARVLTTYYCGPSD 445 (649)
Q Consensus 367 ~Y~~~~~~~l~~~d~~~~~L~~~~~hL~~kG~~~~-~y~~i~DEP~~~~~~~~~r~~~~~ir~~~P~~kil~t~~~~p~d 445 (649)
+|+-.|....+. +.+++|.+.-++-+-++| .|. ++=.| | +...-+++.+++++..|+.++..+++|.++-
T Consensus 132 eytg~Y~~~~~~-~el~~~~k~qle~~~~~g-vD~L~fETi---p----~~~EA~a~l~~l~~~~~~~p~~is~t~~d~g 202 (317)
T KOG1579|consen 132 EYTGIYGDNVEF-EELYDFFKQQLEVFLEAG-VDLLAFETI---P----NVAEAKAALELLQELGPSKPFWISFTIKDEG 202 (317)
T ss_pred ccccccccccCH-HHHHHHHHHHHHHHHhCC-CCEEEEeec---C----CHHHHHHHHHHHHhcCCCCcEEEEEEecCCC
Confidence 455554443332 347888888888888998 443 23234 4 3456678889999999999999999998765
Q ss_pred CCCCCCCcccccc----cccccCCccccccccccccCCchhhhHHHHhhcc-cCCCceeEEEecCC
Q 006344 446 APLGPTPFESFVK----VPKFLRPHTQIYCTSEWVLGNREDLVKDIVTELQ-PENGEEWWTYVCMG 506 (649)
Q Consensus 446 ~~~~~~~~e~~~~----~p~~~~~~idi~c~~~wv~~~~~~~~~~~~~~~r-~~~G~~~W~Y~C~~ 506 (649)
.......++.++. -+++ ..|.++|... .. ....+.++. .-....+-.|..-+
T Consensus 203 ~l~~G~t~e~~~~~~~~~~~~--~~IGvNC~~~---~~----~~~~~~~L~~~~~~~~llvYPNsG 259 (317)
T KOG1579|consen 203 RLRSGETGEEAAQLLKDGINL--LGIGVNCVSP---NF----VEPLLKELMAKLTKIPLLVYPNSG 259 (317)
T ss_pred cccCCCcHHHHHHHhccCCce--EEEEeccCCc---hh----ccHHHHHHhhccCCCeEEEecCCC
Confidence 5555555555432 1111 2467788752 22 223333332 23455666665543
No 21
>smart00737 ML Domain involved in innate immunity and lipid metabolism. ML (MD-2-related lipid-recognition) is a novel domain identified in MD-1, MD-2, GM2A, Npc2 and multiple proteins of unknown function in plants, animals and fungi. These single-domain proteins were predicted to form a beta-rich fold containing multiple strands, and to mediate diverse biological functions through interacting with specific lipids.
Probab=27.09 E-value=1.2e+02 Score=27.31 Aligned_cols=35 Identities=29% Similarity=0.361 Sum_probs=29.4
Q ss_pred ceeecCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 006344 152 QISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (649)
Q Consensus 152 ~~~v~~g~~q~lWv~v~VP~~a~pG~Y~G~i~v~~ 186 (649)
...+.+|+..-.=.++.||...++|.|+++++++.
T Consensus 71 ~CPl~~G~~~~~~~~~~v~~~~P~~~~~v~~~l~d 105 (118)
T smart00737 71 KCPIEKGETVNYTNSLTVPGIFPPGKYTVKWELTD 105 (118)
T ss_pred CCCCCCCeeEEEEEeeEccccCCCeEEEEEEEEEc
Confidence 46678888765557789999999999999999995
No 22
>PF13204 DUF4038: Protein of unknown function (DUF4038); PDB: 3KZS_D.
Probab=25.69 E-value=8.2e+02 Score=25.93 Aligned_cols=198 Identities=13% Similarity=0.140 Sum_probs=86.2
Q ss_pred ccccCChhHHHHHHHHHHHHHhCCcCcc-ccccCCcceeeeccCCCCCCCCCCccccCccccccccccccCCCCCchhHH
Q 006344 305 GVRHGSDEWYEALDQHFKWLLQYRISPF-FCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAK 383 (649)
Q Consensus 305 ~v~~~~~~~f~~ldr~~~~~~~~ris~l-f~~Wg~~~~i~~y~~pw~~~~~k~~~~f~d~~~~~Y~~~~~~~l~~~d~~~ 383 (649)
.+..-...||+.+|+-++.+.+++|-.. ..-||.+.. ..-| +.. +-+-+.+..+
T Consensus 78 d~~~~N~~YF~~~d~~i~~a~~~Gi~~~lv~~wg~~~~----~~~W----g~~-----------------~~~m~~e~~~ 132 (289)
T PF13204_consen 78 DFTRPNPAYFDHLDRRIEKANELGIEAALVPFWGCPYV----PGTW----GFG-----------------PNIMPPENAE 132 (289)
T ss_dssp --TT----HHHHHHHHHHHHHHTT-EEEEESS-HHHHH--------------------------------TTSS-HHHHH
T ss_pred CCCCCCHHHHHHHHHHHHHHHHCCCeEEEEEEECCccc----cccc----ccc-----------------ccCCCHHHHH
Confidence 3344457789999999999999999842 233432210 0011 000 0011233467
Q ss_pred HHHHHHHHHHHhcccccceeeeecCCCCCccchHHHHHHHHHHHHhCCCCeEEEeeccCCCCCCCCCCCccccccccccc
Q 006344 384 DYVRKEIELLRTKAHWKKAYFYLWDEPLNMEHYSSVRNMASELHAYAPDARVLTTYYCGPSDAPLGPTPFESFVKVPKFL 463 (649)
Q Consensus 384 ~~L~~~~~hL~~kG~~~~~y~~i~DEP~~~~~~~~~r~~~~~ir~~~P~~kil~t~~~~p~d~~~~~~~~e~~~~~p~~~ 463 (649)
.|++=+++.+++.. .-+++..-|.-......+.++++++.|++..|.- .++--.|+. .+..+.|.+.
T Consensus 133 ~Y~~yv~~Ry~~~~--NviW~l~gd~~~~~~~~~~w~~~~~~i~~~dp~~-L~T~H~~~~------~~~~~~~~~~---- 199 (289)
T PF13204_consen 133 RYGRYVVARYGAYP--NVIWILGGDYFDTEKTRADWDAMARGIKENDPYQ-LITIHPCGR------TSSPDWFHDE---- 199 (289)
T ss_dssp HHHHHHHHHHTT-S--SEEEEEESSS--TTSSHHHHHHHHHHHHHH--SS--EEEEE-BT------EBTHHHHTT-----
T ss_pred HHHHHHHHHHhcCC--CCEEEecCccCCCCcCHHHHHHHHHHHHhhCCCC-cEEEeCCCC------CCcchhhcCC----
Confidence 78888888877651 1234434455123567788999999999999965 333322221 0111112222
Q ss_pred CCccccccccccccCCc--hhhhHHHH---hhcccCCCceeEEEecCCCCCCCCCccc----CCchhHHHHHHHHHHHcC
Q 006344 464 RPHTQIYCTSEWVLGNR--EDLVKDIV---TELQPENGEEWWTYVCMGPSDPHPNWHL----GMRGSQHRAVMWRVWKEG 534 (649)
Q Consensus 464 ~~~idi~c~~~wv~~~~--~~~~~~~~---~~~r~~~G~~~W~Y~C~~p~~~~pN~fi----d~p~~~~R~~gW~~~k~g 534 (649)
+.+|..+.- .++. ....-..+ ...+....|++..=-||. ...|...- .....+.|--.|.+.--|
T Consensus 200 -~Wldf~~~Q---sgh~~~~~~~~~~~~~~~~~~~~p~KPvin~Ep~Y--Eg~~~~~~~~~~~~~~~dvrr~aw~svlaG 273 (289)
T PF13204_consen 200 -PWLDFNMYQ---SGHNRYDQDNWYYLPEEFDYRRKPVKPVINGEPCY--EGIPYSRWGYNGRFSAEDVRRRAWWSVLAG 273 (289)
T ss_dssp -TT--SEEEB-----S--TT--THHHH--HHHHTSSS---EEESS-----BT-BTTSS-TS-B--HHHHHHHHHHHHHCT
T ss_pred -CcceEEEee---cCCCcccchHHHHHhhhhhhhhCCCCCEEcCcccc--cCCCCCcCcccCCCCHHHHHHHHHHHHhcC
Confidence 234433221 1221 11111111 222345677765334553 11222111 244567777799999999
Q ss_pred C-CEEEEeecccc
Q 006344 535 G-TGFLYWGANCY 546 (649)
Q Consensus 535 ~-~GfL~W~~n~w 546 (649)
. -|+-|.+-.-|
T Consensus 274 a~aG~tYG~~~iW 286 (289)
T PF13204_consen 274 AYAGHTYGAHGIW 286 (289)
T ss_dssp --SEEEE-BHHHH
T ss_pred CCccccCCCCCcc
Confidence 9 99998875555
No 23
>PF04234 CopC: CopC domain; InterPro: IPR007348 CopC is a bacterial blue copper protein that binds 1 atom of copper per protein molecule. Along with CopA, CopC mediates copper resistance by sequestration of copper in the periplasm [].; GO: 0005507 copper ion binding, 0046688 response to copper ion, 0042597 periplasmic space; PDB: 1IX2_B 1LYQ_A 2C9P_C 2C9R_A 2C9Q_A 1M42_A 1OT4_A 1NM4_A.
Probab=25.63 E-value=64 Score=28.42 Aligned_cols=26 Identities=31% Similarity=0.517 Sum_probs=19.5
Q ss_pred EEEEcCCCCCCceeEEEEEEEeccCcc
Q 006344 165 VSIDAPYAQPPGLYEGEIIITSKADTE 191 (649)
Q Consensus 165 v~v~VP~~a~pG~Y~G~i~v~~~~~g~ 191 (649)
+.+.+|..-++|.|+..-+|.+ +||-
T Consensus 61 ~~~~l~~~l~~G~YtV~wrvvs-~DGH 86 (97)
T PF04234_consen 61 LTVPLPPPLPPGTYTVSWRVVS-ADGH 86 (97)
T ss_dssp EEEEESS---SEEEEEEEEEEE-TTSC
T ss_pred EEEECCCCCCCceEEEEEEEEe-cCCC
Confidence 5788899999999999999986 6664
No 24
>PF09099 Qn_am_d_aIII: Quinohemoprotein amine dehydrogenase, alpha subunit domain III; InterPro: IPR015183 This domain is predominantly found in the prokaryotic protein quinohemoprotein amine dehydrogenase, adopting an immunoglobulin-like beta-sandwich fold, with seven strands arranged into two beta sheets; the fold is possibly related to the immunoglobulin and/or fibronectin type III superfamilies. The precise function of this domain has not, as yet, been defined []. ; PDB: 1JMZ_A 1JMX_A 1PBY_A 1JJU_A.
Probab=24.92 E-value=69 Score=27.86 Aligned_cols=21 Identities=24% Similarity=0.370 Sum_probs=18.8
Q ss_pred EEEEEEEcCCCCCCceeEEEE
Q 006344 162 AVWVSIDAPYAQPPGLYEGEI 182 (649)
Q Consensus 162 ~lWv~v~VP~~a~pG~Y~G~i 182 (649)
.++++|.+.++++||.|+..+
T Consensus 49 ~v~v~V~~aa~a~~G~~~v~v 69 (81)
T PF09099_consen 49 EVVVRVKAAADAAPGIRTVRV 69 (81)
T ss_dssp CEEEEEEEECTSSSEEEEEEE
T ss_pred EEEEEEEEcCCCCCccEEEEe
Confidence 589999999999999998665
No 25
>TIGR03769 P_ac_wall_RPT actinobacterial surface-anchored protein domain. This model describes a repeat domain that one to three times in Actinobacterial proteins, some of which have LPXTG-type sortase recognition motifs for covalent attachment to the Gram-positive cell wall. Where it occurs with duplication in an LPXTG-anchored protein, it tends to be adjacent to the substrate-binding protein of the gene trio of an ABC transporter system, where that substrate-binding protein has a single copy of this same domain. This arrangement suggests a substrate-binding relay system, with the LPXTG protein acting as a substrate receptor.
Probab=22.85 E-value=96 Score=23.40 Aligned_cols=14 Identities=29% Similarity=0.465 Sum_probs=12.3
Q ss_pred CCCceeEEEEEEEe
Q 006344 173 QPPGLYEGEIIITS 186 (649)
Q Consensus 173 a~pG~Y~G~i~v~~ 186 (649)
.+||.|+.+++.+.
T Consensus 10 T~PG~Y~l~~~a~~ 23 (41)
T TIGR03769 10 TKPGTYTLTVQATA 23 (41)
T ss_pred CCCeEEEEEEEEEE
Confidence 58999999999973
No 26
>PF00868 Transglut_N: Transglutaminase family; InterPro: IPR001102 Synonym(s): Protein-glutamine gamma-glutamyltransferase, Fibrinoligase, TGase Protein-glutamine gamma-glutamyltransferases (2.3.2.13 from EC) (TGase) are calcium-dependent enzymes that catalyse the cross-linking of proteins by promoting the formation of isopeptide bonds between the gamma-carboxyl group of a glutamine in one polypeptide chain and the epsilon-amino group of a lysine in a second polypeptide chain. TGases also catalyse the conjugation of polyamines to proteins [, ]. Transglutaminases are widely distributed in various organs, tissues and body fluids. The best known transglutaminase is blood coagulation factor XIII, a plasma tetrameric protein composed of two catalytic A subunits and two non-catalytic B subunits. Factor XIII is responsible for cross-linking fibrin chains, thus stabilising the fibrin clot. There are commonly three domains: N-terminal, middle (IPR013808 from INTERPRO) and C-terminal (IPR013807 from INTERPRO). This entry represents the N-terminal domain found in transglutaminases.; GO: 0018149 peptide cross-linking; PDB: 1L9N_B 1NUF_A 1NUD_A 1NUG_B 1L9M_A 1KV3_C 3S3S_A 2Q3Z_A 3LY6_A 3S3P_A ....
Probab=21.78 E-value=6.3e+02 Score=23.17 Aligned_cols=31 Identities=23% Similarity=0.337 Sum_probs=20.9
Q ss_pred ecCCCeeEEEEEEEcCCCCCCceeEEEEEEE
Q 006344 155 LIPGETTAVWVSIDAPYAQPPGLYEGEIIIT 185 (649)
Q Consensus 155 v~~g~~q~lWv~v~VP~~a~pG~Y~G~i~v~ 185 (649)
+...+-..+=|.|.+|++|.-|.|+-+|.++
T Consensus 87 v~~~~~~~~tv~V~spa~A~VG~y~l~v~~~ 117 (118)
T PF00868_consen 87 VESQDGNSVTVSVTSPANAPVGRYKLSVETK 117 (118)
T ss_dssp EEEEETTEEEEEEE--TTS--EEEEEEEEEE
T ss_pred EEecCCCEEEEEEECCCCCceEEEEEEEEEe
Confidence 3334444578899999999999999999886
No 27
>PRK09778 putative antitoxin of the YafO-YafN toxin-antitoxin system; Provisional
Probab=21.43 E-value=1.8e+02 Score=26.22 Aligned_cols=26 Identities=15% Similarity=0.204 Sum_probs=22.4
Q ss_pred chhHHHHHHHHHHHHHHHHHHhhcCc
Q 006344 585 SLRLERILSGLQDIEYLNLYASRYGR 610 (649)
Q Consensus 585 SiRle~lReGieDye~L~lL~~~~~~ 610 (649)
-=-+|.|.|-++|+|+.++.+++...
T Consensus 43 a~~yE~m~e~LeD~eL~~l~~~R~~~ 68 (97)
T PRK09778 43 ASAFEALMDMLAEQEEKKPIKARFRP 68 (97)
T ss_pred HHHHHHHHHHHHhHHHHHHHHHHcCc
Confidence 33579999999999999999998665
No 28
>PRK10301 hypothetical protein; Provisional
Probab=21.21 E-value=1.2e+02 Score=28.23 Aligned_cols=25 Identities=20% Similarity=0.303 Sum_probs=20.9
Q ss_pred EEEEcCCCCCCceeEEEEEEEeccCc
Q 006344 165 VSIDAPYAQPPGLYEGEIIITSKADT 190 (649)
Q Consensus 165 v~v~VP~~a~pG~Y~G~i~v~~~~~g 190 (649)
+.+.+|..-++|+|+.+-+|.+ +||
T Consensus 88 ~~v~l~~~L~~G~YtV~Wrvvs-~DG 112 (124)
T PRK10301 88 LIVPLADSLKPGTYTVDWHVVS-VDG 112 (124)
T ss_pred EEEECCCCCCCccEEEEEEEEe-cCC
Confidence 4677778889999999999986 565
No 29
>PF09608 Alph_Pro_TM: Putative transmembrane protein (Alph_Pro_TM); InterPro: IPR019088 This entry consists of predicted transmembrane proteins of about 270 amino acids. They are found predominantly, though not exclusively, in alphaproteobacteria, generally only once in each genome.
Probab=21.10 E-value=2e+02 Score=29.87 Aligned_cols=38 Identities=24% Similarity=0.432 Sum_probs=30.7
Q ss_pred cceeecCCCeeEEEEEEEcCCCCCCceeEEEEEEEeccCccc
Q 006344 151 CQISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITSKADTEL 192 (649)
Q Consensus 151 ~~~~v~~g~~q~lWv~v~VP~~a~pG~Y~G~i~v~~~~~g~~ 192 (649)
..+.+..+ +-+..+|.+|++.++|.|+.++-+.. +|++
T Consensus 147 ~~V~~~~~--~lFra~i~LPanvp~G~Y~v~v~l~r--dG~v 184 (236)
T PF09608_consen 147 GGVQFLEG--TLFRARIPLPANVPPGDYTVRVYLFR--DGQV 184 (236)
T ss_pred CeEEEcCC--CeEEEEeEcCCCCCcceEEEEEEEEE--CCEE
Confidence 45666544 47889999999999999999999985 6665
No 30
>PF08428 Rib: Rib/alpha-like repeat; InterPro: IPR012706 This entry represents a region of about 79 amino acids found tandemly repeated up to fourteen times within the proteins that contain it. The repeats lack cysteines and are highly conserved, even at the DNA level, within and between proteins []. Proteins containing these repeats include the Rib and alpha surface antigens of group B Streptococcus, Esp of Enterococcus faecalis (Streptococcus faecalis), and related proteins of Lactobacillus. Most members of this protein family also have the cell wall anchor motif, LPXTG, shared by many staphyloccal and streptococcal surface antigens. These repeats are thought to define protective epitopes and may play a role in generating phenotypic and genotypic variation [].
Probab=20.28 E-value=1.3e+02 Score=24.87 Aligned_cols=33 Identities=30% Similarity=0.537 Sum_probs=24.0
Q ss_pred eecCCCeeEEEEEEEcCCCCCCceeEEEEEEEeccCc
Q 006344 154 SLIPGETTAVWVSIDAPYAQPPGLYEGEIIITSKADT 190 (649)
Q Consensus 154 ~v~~g~~q~lWv~v~VP~~a~pG~Y~G~i~v~~~~~g 190 (649)
+++.|. .--|.+ .|...++|.|+++|+|+- .||
T Consensus 20 ~lP~gt-~~~w~~--~pdt~~~G~~~~~V~Vty-pDg 52 (65)
T PF08428_consen 20 NLPAGT-TYSWKD--KPDTSKPGTKTGKVKVTY-PDG 52 (65)
T ss_pred cCCCCc-ceeecc--CCccccCccEEEEEEEEc-CCC
Confidence 344443 246666 899999999999999995 344
Done!