Query 007391
Match_columns 605
No_of_seqs 148 out of 190
Neff 6.2
Searched_HMMs 46136
Date Thu Mar 28 22:57:28 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/007391.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/007391hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF13320 DUF4091: Domain of un 99.9 1.2E-24 2.6E-29 178.5 7.1 66 533-604 1-66 (68)
2 PF10633 NPCBM_assoc: NPCBM-as 94.6 0.17 3.8E-06 42.3 7.8 32 154-185 45-76 (78)
3 PF01229 Glyco_hydro_39: Glyco 88.3 1.2 2.6E-05 50.3 7.3 109 313-438 81-202 (486)
4 COG1470 Predicted membrane pro 88.1 5.4 0.00012 44.4 11.8 36 151-186 325-360 (513)
5 COG1470 Predicted membrane pro 84.5 12 0.00025 41.9 12.0 33 154-186 437-469 (513)
6 smart00633 Glyco_10 Glycosyl h 74.1 12 0.00026 38.3 7.9 104 306-439 6-125 (254)
7 PF06030 DUF916: Bacterial pro 73.5 65 0.0014 29.5 11.7 106 62-185 4-119 (121)
8 PF15418 DUF4625: Domain of un 71.8 5.2 0.00011 37.4 4.1 85 78-186 30-117 (132)
9 PF00150 Cellulase: Cellulase 67.9 28 0.00061 35.2 8.9 101 311-439 58-172 (281)
10 PF02221 E1_DerP2_DerF2: ML do 64.8 15 0.00033 33.3 5.7 35 152-186 85-119 (134)
11 PF14352 DUF4402: Domain of un 61.2 9.4 0.0002 35.0 3.6 34 153-186 93-128 (130)
12 PF13731 WxL: WxL domain surfa 60.5 30 0.00065 34.6 7.4 79 106-185 105-210 (215)
13 cd00917 PG-PI_TP The phosphati 55.2 21 0.00047 32.5 4.8 34 152-186 76-109 (122)
14 PLN00180 NDF6 (NDH-dependent f 48.2 18 0.0004 34.5 3.2 71 513-595 89-160 (180)
15 PF06280 DUF1034: Fn3-like dom 47.2 19 0.0004 32.1 3.1 36 151-186 62-100 (112)
16 PF01835 A2M_N: MG2 domain; I 46.7 83 0.0018 27.0 7.0 26 160-185 61-86 (99)
17 PF09608 Alph_Pro_TM: Putative 46.1 33 0.00071 35.3 5.0 54 151-221 147-200 (236)
18 COG5520 O-Glycosyl hydrolase [ 44.5 1.8E+02 0.0039 31.9 10.2 147 381-544 150-310 (433)
19 PF02449 Glyco_hydro_42: Beta- 44.4 1.1E+02 0.0024 33.1 9.1 86 494-599 286-372 (374)
20 TIGR02186 alph_Pro_TM conserve 43.0 31 0.00068 36.0 4.3 41 152-196 173-213 (261)
21 PF08428 Rib: Rib/alpha-like r 41.4 49 0.0011 26.9 4.4 33 154-190 20-52 (65)
22 smart00737 ML Domain involved 36.0 69 0.0015 28.5 5.0 35 152-186 71-105 (118)
23 PF13204 DUF4038: Protein of u 29.2 4.1E+02 0.0089 27.9 10.1 201 305-546 78-286 (289)
24 PF12245 Big_3_2: Bacterial Ig 25.0 1.5E+02 0.0033 23.5 4.7 46 162-223 10-55 (60)
25 PF09099 Qn_am_d_aIII: Quinohe 24.9 71 0.0015 27.4 2.8 22 161-182 48-69 (81)
26 TIGR03769 P_ac_wall_RPT actino 24.5 88 0.0019 23.3 2.9 13 173-185 10-22 (41)
27 PF00868 Transglut_N: Transglu 24.0 80 0.0017 28.7 3.2 33 153-185 85-117 (118)
28 PF04234 CopC: CopC domain; I 23.6 1.2E+02 0.0025 26.4 4.1 26 165-191 61-86 (97)
29 COG3693 XynA Beta-1,4-xylanase 20.9 2.1E+02 0.0046 31.0 5.9 106 307-439 73-193 (345)
30 PF14734 DUF4469: Domain of un 20.2 1E+02 0.0022 27.6 3.0 25 162-186 63-87 (102)
No 1
>PF13320 DUF4091: Domain of unknown function (DUF4091)
Probab=99.91 E-value=1.2e-24 Score=178.51 Aligned_cols=66 Identities=42% Similarity=0.698 Sum_probs=61.8
Q ss_pred cCCcEEEEeeccccCCCCCCcccccccCCCCCCccEEEccCCCCCCCCCcccchhHHHHHHHHhHHHHHhhh
Q 007391 533 EGGTGFLYWGANCYEKATVPSAEIRFRRGLPPGDGVLFYPGEVFSSSRQPVASLRLERILSGLQVRWICYYL 604 (605)
Q Consensus 533 ~g~~GfL~W~~n~w~~~~dP~~d~~f~~~~~~GDg~LVYPG~~~~~~~~PvsSiRle~lreGiqDye~~~~l 604 (605)
||++|||||+||+| ..||+.+++|+. |++||++|||||++ .++|++|||||+||+||||||+|++|
T Consensus 1 y~~~G~L~W~~~~w--~~dP~~d~~~~~-~~~GD~~lvYPg~~---~~~p~~SiRle~lr~G~qD~e~l~~l 66 (68)
T PF13320_consen 1 YGFDGFLRWAYNFW--NEDPWEDTRFRG-FPAGDGFLVYPGED---TGGPVSSIRLEVLREGIQDYEYLRLL 66 (68)
T ss_pred CCCCeEEEeccccc--ccCcccccCcCc-CCCCCeEEEecCCC---CCCcccCHHHHHHHHHHHHHHHHHHH
Confidence 68999999999999 569999999986 99999999999983 38999999999999999999999998
No 2
>PF10633 NPCBM_assoc: NPCBM-associated, NEW3 domain of alpha-galactosidase; InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=94.64 E-value=0.17 Score=42.33 Aligned_cols=32 Identities=31% Similarity=0.543 Sum_probs=25.8
Q ss_pred eecCCCeeEEEEEEEcCCCCCCceeEEEEEEE
Q 007391 154 SLIPGETTAVWVSIDAPYAQPPGLYEGEIIIT 185 (605)
Q Consensus 154 ~l~~~~~q~vWV~v~VP~~a~pG~Y~g~v~V~ 185 (605)
.|++|+.+.+=++|.+|++++||.|+.+++++
T Consensus 45 ~l~pG~s~~~~~~V~vp~~a~~G~y~v~~~a~ 76 (78)
T PF10633_consen 45 SLPPGESVTVTFTVTVPADAAPGTYTVTVTAR 76 (78)
T ss_dssp -B-TTSEEEEEEEEEE-TT--SEEEEEEEEEE
T ss_pred cCCCCCEEEEEEEEECCCCCCCceEEEEEEEE
Confidence 68899999999999999999999999999986
No 3
>PF01229 Glyco_hydro_39: Glycosyl hydrolases family 39; InterPro: IPR000514 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. Glycoside hydrolase family 39 GH39 from CAZY comprises enzymes with several known activities; alpha-L-iduronidase (3.2.1.76 from EC); beta-xylosidase (3.2.1.37 from EC). The most highly conserved regions in these enzymes are located in their N-terminal sections. These contain a glutamic acid residue which, on the basis of similarities with other families of glycosyl hydrolases [], probably acts as the proton donor in their catalytic mechanism.; GO: 0004553 hydrolase activity, hydrolyzing O-glycosyl compounds, 0005975 carbohydrate metabolic process; PDB: 2BS9_D 2BFG_E 1W91_B 1UHV_D 1PX8_A.
Probab=88.29 E-value=1.2 Score=50.29 Aligned_cols=109 Identities=24% Similarity=0.318 Sum_probs=62.1
Q ss_pred H-HHHHHHHHHHHHhCCcCccccccCCcceeeeecCCCCCCCCCCccccCCcccccccccCCCCCCChhHHHHHHHHHHH
Q 007391 313 W-YEALDQHFKWLLQYRISPFFCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYVRKEIE 391 (605)
Q Consensus 313 ~-~~~ldrw~~~~~~~~is~~f~~wg~~~~i~~y~~pw~~~~~~~~~yf~~~~~~~Y~~~~~~~~~g~~~~~~~L~~~~~ 391 (605)
| |+.+|+-+++++++||.|+. ..| |+..+-+. +.. ..|. |.....|.. .-+.|.+++++|++
T Consensus 81 Ynf~~lD~i~D~l~~~g~~P~v-el~-------f~p~~~~~-~~~-~~~~------~~~~~~pp~-~~~~W~~lv~~~~~ 143 (486)
T PF01229_consen 81 YNFTYLDQILDFLLENGLKPFV-ELG-------FMPMALAS-GYQ-TVFW------YKGNISPPK-DYEKWRDLVRAFAR 143 (486)
T ss_dssp E--HHHHHHHHHHHHCT-EEEE-EE--------SB-GGGBS-S---EETT------TTEE-S-BS--HHHHHHHHHHHHH
T ss_pred CChHHHHHHHHHHHHcCCEEEE-EEE-------echhhhcC-CCC-cccc------ccCCcCCcc-cHHHHHHHHHHHHH
Confidence 7 99999999999999999942 111 11000000 000 0111 010011111 22479999999999
Q ss_pred HHHH-cCc--cceeeeeecCCCCCc---------cchHHHHHHHHHHHHhCCCCcEEEe
Q 007391 392 LLRT-KAH--WKKAYFYLWDEPLNM---------EHYSSVRNMASELHAYAPDARVLTT 438 (605)
Q Consensus 392 hL~~-kGw--~~~~y~y~~DEP~~~---------~~~~~~~~~~~~ir~~~P~~ki~~t 438 (605)
|+.+ .|. ....+|=++.||... +=++.|+.+++.||++.|++||-..
T Consensus 144 h~~~RYG~~ev~~W~fEiWNEPd~~~f~~~~~~~ey~~ly~~~~~~iK~~~p~~~vGGp 202 (486)
T PF01229_consen 144 HYIDRYGIEEVSTWYFEIWNEPDLKDFWWDGTPEEYFELYDATARAIKAVDPELKVGGP 202 (486)
T ss_dssp HHHHHHHHHHHTTSEEEESS-TTSTTTSGGG-HHHHHHHHHHHHHHHHHH-TTSEEEEE
T ss_pred HHHhhcCCccccceeEEeCcCCCcccccCCCCHHHHHHHHHHHHHHHHHhCCCCcccCc
Confidence 9975 443 233466679999532 2244678889999999999998764
No 4
>COG1470 Predicted membrane protein [Function unknown]
Probab=88.12 E-value=5.4 Score=44.44 Aligned_cols=36 Identities=28% Similarity=0.458 Sum_probs=34.3
Q ss_pred ceeeecCCCeeEEEEEEEcCCCCCCceeEEEEEEEE
Q 007391 151 CQISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (605)
Q Consensus 151 ~~~~l~~~~~q~vWV~v~VP~~a~pG~Y~g~v~V~~ 186 (605)
..+.+.||+...|-++|+.|++|.||.|..+|+++.
T Consensus 325 t~vkL~~gE~kdvtleV~ps~na~pG~Ynv~I~A~s 360 (513)
T COG1470 325 TSVKLKPGEEKDVTLEVYPSLNATPGTYNVTITASS 360 (513)
T ss_pred EEEEecCCCceEEEEEEecCCCCCCCceeEEEEEec
Confidence 689999999999999999999999999999999984
No 5
>COG1470 Predicted membrane protein [Function unknown]
Probab=84.46 E-value=12 Score=41.92 Aligned_cols=33 Identities=36% Similarity=0.457 Sum_probs=30.0
Q ss_pred eecCCCeeEEEEEEEcCCCCCCceeEEEEEEEE
Q 007391 154 SLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (605)
Q Consensus 154 ~l~~~~~q~vWV~v~VP~~a~pG~Y~g~v~V~~ 186 (605)
.|+||+.-.|=++|.||++|.+|.|..+|+.++
T Consensus 437 sL~pge~~tV~ltI~vP~~a~aGdY~i~i~~ks 469 (513)
T COG1470 437 SLEPGESKTVSLTITVPEDAGAGDYRITITAKS 469 (513)
T ss_pred ccCCCCcceEEEEEEcCCCCCCCcEEEEEEEee
Confidence 367888889999999999999999999999994
No 6
>smart00633 Glyco_10 Glycosyl hydrolase family 10.
Probab=74.10 E-value=12 Score=38.33 Aligned_cols=104 Identities=12% Similarity=0.143 Sum_probs=63.0
Q ss_pred cccCChhH-HHHHHHHHHHHHhCCcCcc--ccccCCcceeeeecCCCCCCCCCCccccCCcccccccccCCCCCCChhHH
Q 007391 306 VRHGSDEW-YEALDQHFKWLLQYRISPF--FCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGA 382 (605)
Q Consensus 306 v~~~~~~~-~~~ldrw~~~~~~~~is~~--f~~wg~~~~i~~y~~pw~~~~~~~~~yf~~~~~~~Y~~~~~~~~~g~~~~ 382 (605)
++...+.| |+..|+.++++.+++|.-. .+-|+.. ...|- . ..+ .++- .+.+
T Consensus 6 ~ep~~G~~n~~~~D~~~~~a~~~gi~v~gH~l~W~~~------~P~W~-~--------~~~-~~~~----------~~~~ 59 (254)
T smart00633 6 TEPSRGQFNFSGADAIVNFAKENGIKVRGHTLVWHSQ------TPDWV-F--------NLS-KETL----------LARL 59 (254)
T ss_pred ccCCCCccChHHHHHHHHHHHHCCCEEEEEEEeeccc------CCHhh-h--------cCC-HHHH----------HHHH
Confidence 44555667 9999999999999999731 1233321 11220 0 000 0010 1256
Q ss_pred HHHHHHHHHHHHHcCccceeeeeecCCCCCcc-------c----h--HHHHHHHHHHHHhCCCCcEEEee
Q 007391 383 KDYVRKEIELLRTKAHWKKAYFYLWDEPLNME-------H----Y--SSVRNMASELHAYAPDARVLTTY 439 (605)
Q Consensus 383 ~~~L~~~~~hL~~kGw~~~~y~y~~DEP~~~~-------~----~--~~~~~~~~~ir~~~P~~ki~~t~ 439 (605)
.+|+.+.++|.+.+. ..+ -+..||.+.. . + +.++.+.+.+|+++|+.|++.-.
T Consensus 60 ~~~i~~v~~ry~g~i---~~w-dV~NE~~~~~~~~~~~~~w~~~~G~~~i~~af~~ar~~~P~a~l~~Nd 125 (254)
T smart00633 60 ENHIKTVVGRYKGKI---YAW-DVVNEALHDNGSGLRRSVWYQILGEDYIEKAFRYAREADPDAKLFYND 125 (254)
T ss_pred HHHHHHHHHHhCCcc---eEE-EEeeecccCCCcccccchHHHhcChHHHHHHHHHHHHhCCCCEEEEec
Confidence 777888888776542 123 3578875321 1 2 66789999999999999999853
No 7
>PF06030 DUF916: Bacterial protein of unknown function (DUF916); InterPro: IPR010317 This family consists of putative cell surface proteins, from Firmicutes, of unknown function.
Probab=73.52 E-value=65 Score=29.54 Aligned_cols=106 Identities=18% Similarity=0.288 Sum_probs=68.7
Q ss_pred cccCCCCCCCC-CCceeEEeecCceEEEEEEEecCcccCCCCCCcceEEEEccc-CCCCCCcccccCceEEEEEeecCC-
Q 007391 62 ANVGPQEMPRP-LEPINLLAARNERESVQIALRPKVSWSSSSTAGVVQVQCSDL-CSASGDRLVVGQSLMLRRVVPMLG- 138 (605)
Q Consensus 62 ~kV~~~~~p~~-~~~~~L~a~RgE~~sfQivl~~~~~~~~~~~~~~V~v~~sdl-~s~~G~~~i~~~~i~~~~v~~VpG- 138 (605)
.-|.|+..-.. ..-+.|....|+...+|+-+.-. +.....|+|++.+- ++.+| .+.|.+-
T Consensus 4 ~p~~p~~Q~~~~~~YFdL~~~P~q~~~l~v~i~N~-----s~~~~tv~v~~~~A~Tn~nG------------~I~Y~~~~ 66 (121)
T PF06030_consen 4 TPVLPENQIDKNVSYFDLKVKPGQKQTLEVRITNN-----SDKEITVKVSANTATTNDNG------------VIDYSQNN 66 (121)
T ss_pred eecCCccccCCCCCeEEEEeCCCCEEEEEEEEEeC-----CCCCEEEEEEEeeeEecCCE------------EEEECCCC
Confidence 34566655433 46799999999999999999763 23333444443322 12222 1233221
Q ss_pred -CCC--ccccC----CCCCceeeecCCCeeEEEEEEEcCCCCCCceeEEEEEEE
Q 007391 139 -VPD--ALVPL----DLPVCQISLIPGETTAVWVSIDAPYAQPPGLYEGEIIIT 185 (605)
Q Consensus 139 -~PD--~L~P~----~~~~~~~~l~~~~~q~vWV~v~VP~~a~pG~Y~g~v~V~ 185 (605)
-.| +-.++ ..+ ..+.|+|++.+-|=++|.+|+..-.|..-|-|.|+
T Consensus 67 ~~~d~sl~~~~~~~v~~~-~~Vtl~~~~sk~V~~~i~~P~~~f~G~ilGGi~~~ 119 (121)
T PF06030_consen 67 PKKDKSLKYPFSDLVKIP-KEVTLPPNESKTVTFTIKMPKKAFDGIILGGIYFS 119 (121)
T ss_pred cccCcccCcchHHhccCC-cEEEECCCCEEEEEEEEEcCCCCcCCEEEeeEEEE
Confidence 111 11122 111 34999999999999999999999999999999998
No 8
>PF15418 DUF4625: Domain of unknown function (DUF4625)
Probab=71.81 E-value=5.2 Score=37.40 Aligned_cols=85 Identities=18% Similarity=0.246 Sum_probs=51.2
Q ss_pred EEeecCceEEEEEEEecCcccCCCCCCcceEEEEcccCC--CCCCcccccCceEEEEEeecCCCCCccccCCCCCceeee
Q 007391 78 LLAARNERESVQIALRPKVSWSSSSTAGVVQVQCSDLCS--ASGDRLVVGQSLMLRRVVPMLGVPDALVPLDLPVCQISL 155 (605)
Q Consensus 78 L~a~RgE~~sfQivl~~~~~~~~~~~~~~V~v~~sdl~s--~~G~~~i~~~~i~~~~v~~VpG~PD~L~P~~~~~~~~~l 155 (605)
-.+-||+.+.|..-+.+ ...++.++|.+-.=.. ..+. -.+.+ ..|+.-. ..+.+
T Consensus 30 ~~~~~G~~ihfe~~i~d------~~~i~si~VeIH~nfd~H~h~~--~~~~~---------------~~~~~~~-~~~~~ 85 (132)
T PF15418_consen 30 KVATRGDDIHFEADISD------NSAIKSIKVEIHNNFDHHTHST--EAGEC---------------EKPWVFE-QDYDI 85 (132)
T ss_pred eEEecCCcEEEEEEEEc------ccceeEEEEEEecCcCcccccc--ccccc---------------ccCcEEE-EEEcc
Confidence 45679999999999988 4568888888621000 0000 00000 1111100 11222
Q ss_pred cCC-CeeEEEEEEEcCCCCCCceeEEEEEEEE
Q 007391 156 IPG-ETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (605)
Q Consensus 156 ~~~-~~q~vWV~v~VP~~a~pG~Y~g~v~V~~ 186 (605)
..| ...-+=..|.||++++||.|...|+|+.
T Consensus 86 ~~g~~~~~~h~~i~IPa~a~~G~YH~~i~VtD 117 (132)
T PF15418_consen 86 YGGKKNYDFHEHIDIPADAPAGDYHFMITVTD 117 (132)
T ss_pred cCCcccEeEEEeeeCCCCCCCcceEEEEEEEE
Confidence 222 3456678999999999999999999996
No 9
>PF00150 Cellulase: Cellulase (glycosyl hydrolase family 5); InterPro: IPR001547 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. Glycoside hydrolase family 5 GH5 from CAZY comprises enzymes with several known activities; endoglucanase (3.2.1.4 from EC); beta-mannanase (3.2.1.78 from EC); exo-1,3-glucanase (3.2.1.58 from EC); endo-1,6-glucanase (3.2.1.75 from EC); xylanase (3.2.1.8 from EC); endoglycoceramidase (3.2.1.123 from EC). The microbial degradation of cellulose and xylans requires several types of enzymes. Fungi and bacteria produces a spectrum of cellulolytic enzymes (cellulases) and xylanases which, on the basis of sequence similarities, can be classified into families. One of these families is known as the cellulase family A [] or as the glycosyl hydrolases family 5 []. One of the conserved regions in this family contains a conserved glutamic acid residue which is potentially involved [] in the catalytic mechanism.; GO: 0004553 hydrolase activity, hydrolyzing O-glycosyl compounds, 0005975 carbohydrate metabolic process; PDB: 3NDY_A 3NDZ_B 1LF1_A 1TVP_B 1TVN_A 3AYR_A 3AYS_A 1QI0_A 1W3K_A 1OCQ_A ....
Probab=67.92 E-value=28 Score=35.17 Aligned_cols=101 Identities=13% Similarity=0.206 Sum_probs=61.1
Q ss_pred hhHHHHHHHHHHHHHhCCcCccccccCCcceeeeec--CCCCCCCCCCccccCCcccccccccCCCCCCChhHHHHHHHH
Q 007391 311 DEWYEALDQHFKWLLQYRISPFFCRWGESMRVLTYT--CPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYVRK 388 (605)
Q Consensus 311 ~~~~~~ldrw~~~~~~~~is~~f~~wg~~~~i~~y~--~pw~~~~~~~~~yf~~~~~~~Y~~~~~~~~~g~~~~~~~L~~ 388 (605)
..+++.||+-++++.+++|.-++ +.- ..|... ... +. .... ..+..+.+++.
T Consensus 58 ~~~~~~ld~~v~~a~~~gi~vil----------d~h~~~~w~~~-~~~---~~--~~~~----------~~~~~~~~~~~ 111 (281)
T PF00150_consen 58 ETYLARLDRIVDAAQAYGIYVIL----------DLHNAPGWANG-GDG---YG--NNDT----------AQAWFKSFWRA 111 (281)
T ss_dssp HHHHHHHHHHHHHHHHTT-EEEE----------EEEESTTCSSS-TST---TT--THHH----------HHHHHHHHHHH
T ss_pred HHHHHHHHHHHHHHHhCCCeEEE----------EeccCcccccc-ccc---cc--cchh----------hHHHHHhhhhh
Confidence 35689999999999999997532 221 123100 000 00 0000 01245567777
Q ss_pred HHHHHHHcCccceeeeeecCCCCCccc------------hHHHHHHHHHHHHhCCCCcEEEee
Q 007391 389 EIELLRTKAHWKKAYFYLWDEPLNMEH------------YSSVRNMASELHAYAPDARVLTTY 439 (605)
Q Consensus 389 ~~~hL~~kGw~~~~y~y~~DEP~~~~~------------~~~~~~~~~~ir~~~P~~ki~~t~ 439 (605)
++++++...- .+.+-++.||..... .+.++++++.||++.|+..|+...
T Consensus 112 la~~y~~~~~--v~~~el~NEP~~~~~~~~w~~~~~~~~~~~~~~~~~~Ir~~~~~~~i~~~~ 172 (281)
T PF00150_consen 112 LAKRYKDNPP--VVGWELWNEPNGGNDDANWNAQNPADWQDWYQRAIDAIRAADPNHLIIVGG 172 (281)
T ss_dssp HHHHHTTTTT--TEEEESSSSGCSTTSTTTTSHHHTHHHHHHHHHHHHHHHHTTSSSEEEEEE
T ss_pred hccccCCCCc--EEEEEecCCccccCCccccccccchhhhhHHHHHHHHHHhcCCcceeecCC
Confidence 8888764332 455668999964211 356788999999999998888765
No 10
>PF02221 E1_DerP2_DerF2: ML domain; InterPro: IPR003172 The MD-2-related lipid-recognition (ML) domain is implicated in lipid recognition, particularly in the recognition of pathogen related products. It has an immunoglobulin-like beta-sandwich fold similar to that of E-set Ig domains. This domain is present in the following proteins: Epididymal secretory protein E1 (also known as Niemann-Pick C2 protein), which is known to bind cholesterol. Niemann-Pick disease type C2 is a fatal hereditary disease characterised by accumulation of low-density lipoprotein-derived cholesterol in lysosomes []. House-dust mite allergen proteins such as Der f 2 from Dermatophagoides farinae and Der p 2 from Dermatophagoides pteronyssinus []. ; PDB: 2AG9_B 1G13_B 2AG2_B 2AG4_A 1TJJ_C 1PU5_C 1PUB_A 2AF9_A 3T6Q_D 3M7O_B ....
Probab=64.81 E-value=15 Score=33.28 Aligned_cols=35 Identities=26% Similarity=0.338 Sum_probs=32.7
Q ss_pred eeeecCCCeeEEEEEEEcCCCCCCceeEEEEEEEE
Q 007391 152 QISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (605)
Q Consensus 152 ~~~l~~~~~q~vWV~v~VP~~a~pG~Y~g~v~V~~ 186 (605)
...+.+|+....-+++.||...++|.|+++++++.
T Consensus 85 ~CPi~~G~~~~~~~~~~i~~~~p~~~~~i~~~l~d 119 (134)
T PF02221_consen 85 SCPIKAGEYYTYTYTIPIPKIYPPGKYTIQWKLTD 119 (134)
T ss_dssp TSTBTTTEEEEEEEEEEESTTSSSEEEEEEEEEEE
T ss_pred cCccCCCcEEEEEEEEEcccceeeEEEEEEEEEEe
Confidence 57789999999999999999999999999999995
No 11
>PF14352 DUF4402: Domain of unknown function (DUF4402)
Probab=61.18 E-value=9.4 Score=35.05 Aligned_cols=34 Identities=26% Similarity=0.485 Sum_probs=26.3
Q ss_pred eeecCCCeeEEEE--EEEcCCCCCCceeEEEEEEEE
Q 007391 153 ISLIPGETTAVWV--SIDAPYAQPPGLYEGEIIITS 186 (605)
Q Consensus 153 ~~l~~~~~q~vWV--~v~VP~~a~pG~Y~g~v~V~~ 186 (605)
..+..+....+.| ++.|+.++++|.|+|+++|++
T Consensus 93 ~~~~~~g~~~~~VGGtL~v~~~~~~G~YsGt~~VtV 128 (130)
T PF14352_consen 93 TTLDTGGSATFNVGGTLNVPANQAAGTYSGTFTVTV 128 (130)
T ss_pred eEecCCCcEEEEEEEEEEcCCCCCCeEEEEEEEEEE
Confidence 3344455666666 579999999999999999984
No 12
>PF13731 WxL: WxL domain surface cell wall-binding
Probab=60.54 E-value=30 Score=34.63 Aligned_cols=79 Identities=24% Similarity=0.390 Sum_probs=47.5
Q ss_pred ceEEEEcccCCCCCCcccccCceEEEEEeecC--C---CCC------ccccCCCCCceeeecCCCeeEEE----------
Q 007391 106 VVQVQCSDLCSASGDRLVVGQSLMLRRVVPML--G---VPD------ALVPLDLPVCQISLIPGETTAVW---------- 164 (605)
Q Consensus 106 ~V~v~~sdl~s~~G~~~i~~~~i~~~~v~~Vp--G---~PD------~L~P~~~~~~~~~l~~~~~q~vW---------- 164 (605)
.|+|+.++|++.+|. .+.+..+.+....... + -|- .|.+.......+.-.+++-+..|
T Consensus 105 ~L~v~~s~F~~~~~~-~L~ga~l~~~~~~~~~~~~~~~~~~~~~~~~~l~~~~~~~~v~~A~~~~g~G~~~~~~~~~~~~ 183 (215)
T PF13731_consen 105 TLTVKLSPFTNADGD-TLPGATLTFNNGKVQSTANNTNTPTTVSSNITLTPGGQAQTVMSAAKGQGQGTWSYSFGDQDAT 183 (215)
T ss_pred EEEEEeccccccCCc-CcccceEEecCceeEeecccccCCcccccceEeccCCcceeeEeecccccceEEEEEeCCcccc
Confidence 688889999998765 3455556554433322 0 011 12222111112223355555666
Q ss_pred ----EEEEcCCCCC--CceeEEEEEEE
Q 007391 165 ----VSIDAPYAQP--PGLYEGEIIIT 185 (605)
Q Consensus 165 ----V~v~VP~~a~--pG~Y~g~v~V~ 185 (605)
|.+.||..+. +|.|+++|+=+
T Consensus 184 ~~~~v~L~VP~~~~~~ag~Yt~tlTWt 210 (215)
T PF13731_consen 184 ADTGVSLSVPANTAKQAGTYTATLTWT 210 (215)
T ss_pred cccceEEEeCCCCcccCCcEEEEEEEE
Confidence 7899999998 69999999977
No 13
>cd00917 PG-PI_TP The phosphatidylinositol/phosphatidylglycerol transfer protein (PG/PI-TP) has been shown to bind phosphatidylglycerol and phosphatidylinositol, but the biological significance of this is still obscure. These proteins belong to the ML domain family.
Probab=55.15 E-value=21 Score=32.47 Aligned_cols=34 Identities=24% Similarity=0.365 Sum_probs=30.5
Q ss_pred eeeecCCCeeEEEEEEEcCCCCCCceeEEEEEEEE
Q 007391 152 QISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (605)
Q Consensus 152 ~~~l~~~~~q~vWV~v~VP~~a~pG~Y~g~v~V~~ 186 (605)
...+.+|+.. +=.++.||...++|.|+++.+++.
T Consensus 76 ~CPi~~G~~~-~~~~~~ip~~~P~g~y~v~~~l~d 109 (122)
T cd00917 76 SCPIEPGDKF-LTKLVDLPGEIPPGKYTVSARAYT 109 (122)
T ss_pred cCCcCCCcEE-EEEEeeCCCCCCCceEEEEEEEEC
Confidence 5778899987 888899999999999999999994
No 14
>PLN00180 NDF6 (NDH-dependent flow 6); Provisional
Probab=48.16 E-value=18 Score=34.55 Aligned_cols=71 Identities=20% Similarity=0.317 Sum_probs=51.9
Q ss_pred CccccCchhhHHHHHHHHHHcCCcEEEEeeccccCCCCCCcccc-cccCCCCCCccEEEccCCCCCCCCCcccchhHHHH
Q 007391 513 NWHLGMRGSQHRAVMWRVWKEGGTGFLYWGANCYEKATVPSAEI-RFRRGLPPGDGVLFYPGEVFSSSRQPVASLRLERI 591 (605)
Q Consensus 513 N~fid~p~~~~R~lgW~~~k~g~~GfL~W~~n~w~~~~dP~~d~-~f~~~~~~GDg~LVYPG~~~~~~~~PvsSiRle~l 591 (605)
=|+++...+.+-.+..+.+ --||..|+.+-.|||-|+ .||++-+.|-+.-||--. ....+|-|-|.|
T Consensus 89 IwHLSD~aiKnVYtfY~mF-------T~WG~~fFgSmKDPfYDSe~YRgdGGDGT~hW~Yd~Q-----Ed~E~sAReeL~ 156 (180)
T PLN00180 89 IWHLSDAAIKNVYTFYIMF-------TCWGCLFFGSMKDPFYDSEEYRGDGGDGTGHWVYERQ-----EDIEESARAELW 156 (180)
T ss_pred hhhccHHHHhHHHHHHHHH-------HHHHHhheeccCCcccchHHhcccCCCCceeeEeehH-----HHHHHHHHHHHH
Confidence 3566677777755444443 358888887779998876 477776777788899774 337899999999
Q ss_pred HHHH
Q 007391 592 LSGL 595 (605)
Q Consensus 592 reGi 595 (605)
||-+
T Consensus 157 REEL 160 (180)
T PLN00180 157 REEL 160 (180)
T ss_pred HHHH
Confidence 9853
No 15
>PF06280 DUF1034: Fn3-like domain (DUF1034); InterPro: IPR010435 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This domain of unknown function is present in bacterial and plant peptidases belonging to MEROPS peptidase family S8 (subfamily S8A subtilisin, clan SB). It is C-terminal to and adjacent to the S8 peptidase domain and can be found in conjunction with the PA (Protease associated) domain (IPR003137 from INTERPRO) and additionally in Gram-positive bacteria with the surface protein anchor domain (IPR001899 from INTERPRO).; GO: 0004252 serine-type endopeptidase activity, 0005618 cell wall, 0016020 membrane; PDB: 3EIF_A 1XF1_B.
Probab=47.24 E-value=19 Score=32.13 Aligned_cols=36 Identities=28% Similarity=0.426 Sum_probs=30.4
Q ss_pred ceeeecCCCeeEEEEEEEcCCCCCC---ceeEEEEEEEE
Q 007391 151 CQISLIPGETTAVWVSIDAPYAQPP---GLYEGEIIITS 186 (605)
Q Consensus 151 ~~~~l~~~~~q~vWV~v~VP~~a~p---G~Y~g~v~V~~ 186 (605)
..+.|+||+.+-|=|+|.+|++..+ ..|.|-|.++.
T Consensus 62 ~~vTV~ag~s~~v~vti~~p~~~~~~~~~~~eG~I~~~~ 100 (112)
T PF06280_consen 62 DTVTVPAGQSKTVTVTITPPSGLDASNGPFYEGFITFKS 100 (112)
T ss_dssp EEEEE-TTEEEEEEEEEE--GGGHHTT-EEEEEEEEEES
T ss_pred CeEEECCCCEEEEEEEEEehhcCCcccCCEEEEEEEEEc
Confidence 5899999999999999999998776 89999999994
No 16
>PF01835 A2M_N: MG2 domain; InterPro: IPR002890 The proteinase-binding alpha-macroglobulins (A2M) [] are large glycoproteins found in the plasma of vertebrates, in the hemolymph of some invertebrates and in reptilian and avian egg white. A2M-like proteins are able to inhibit all four classes of proteinases by a 'trapping' mechanism. They have a peptide stretch, called the 'bait region', which contains specific cleavage sites for different proteinases. When a proteinase cleaves the bait region, a conformational change is induced in the protein, thus trapping the proteinase. The entrapped enzyme remains active against low molecular weight substrates, whilst its activity toward larger substrates is greatly reduced, due to steric hindrance. Following cleavage in the bait region, a thiol ester bond, formed between the side chains of a cysteine and a glutamine, is cleaved and mediates the covalent binding of the A2M-like protein to the proteinase. This family includes the N-terminal region of the alpha-2-macroglobulin family. The inhibitor domains belong to MEROPS inhibitor family I39.; GO: 0004866 endopeptidase inhibitor activity; PDB: 2B39_B 3KLS_B 3PRX_C 3KM9_B 3PVM_C 3CU7_A 4E0S_A 4A5W_A 4ACQ_C 2P9R_B ....
Probab=46.74 E-value=83 Score=26.95 Aligned_cols=26 Identities=19% Similarity=0.180 Sum_probs=17.3
Q ss_pred eeEEEEEEEcCCCCCCceeEEEEEEE
Q 007391 160 TTAVWVSIDAPYAQPPGLYEGEIIIT 185 (605)
Q Consensus 160 ~q~vWV~v~VP~~a~pG~Y~g~v~V~ 185 (605)
.-.+-.++.+|+++..|.|+.++...
T Consensus 61 ~G~~~~~~~lp~~~~~G~y~i~~~~~ 86 (99)
T PF01835_consen 61 NGIFSGSFQLPDDAPLGTYTIRVKTD 86 (99)
T ss_dssp TTEEEEEEE--SS---EEEEEEEEET
T ss_pred CCEEEEEEECCCCCCCEeEEEEEEEc
Confidence 34567789999999999999999885
No 17
>PF09608 Alph_Pro_TM: Putative transmembrane protein (Alph_Pro_TM); InterPro: IPR019088 This entry consists of predicted transmembrane proteins of about 270 amino acids. They are found predominantly, though not exclusively, in alphaproteobacteria, generally only once in each genome.
Probab=46.13 E-value=33 Score=35.28 Aligned_cols=54 Identities=19% Similarity=0.212 Sum_probs=39.7
Q ss_pred ceeeecCCCeeEEEEEEEcCCCCCCceeEEEEEEEEcCCCccccccccccchhhhccccceeeeeccCCCC
Q 007391 151 CQISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITSKADTELSSQCLGKGEKHRLFMELRNCLDNVEPIEG 221 (605)
Q Consensus 151 ~~~~l~~~~~q~vWV~v~VP~~a~pG~Y~g~v~V~~~~~g~~~~~~~~~~~~~~~~~~l~~~l~V~~~~lp 221 (605)
..+.+..+ +-..-+|++|++.++|.|+.++-+. .+|+++++. ...|+|...=+.
T Consensus 147 ~~V~~~~~--~lFra~i~LPanvp~G~Y~v~v~l~--rdG~vv~~~-------------~~~l~V~KvG~e 200 (236)
T PF09608_consen 147 GGVQFLEG--TLFRARIPLPANVPPGDYTVRVYLF--RDGQVVASQ-------------ETPLRVRKVGFE 200 (236)
T ss_pred CeEEEcCC--CeEEEEeEcCCCCCcceEEEEEEEE--ECCEEEEEE-------------eeEEEEEEccHH
Confidence 35665444 4778999999999999999999998 578776643 556666544443
No 18
>COG5520 O-Glycosyl hydrolase [Cell envelope biogenesis, outer membrane]
Probab=44.45 E-value=1.8e+02 Score=31.90 Aligned_cols=147 Identities=16% Similarity=0.259 Sum_probs=75.3
Q ss_pred HHHHHHHHHHHHHHHcCccceeeeeecCCCCCccchHH----HHHHHHHHHHh----CCCCcEEEeeccCCCCCCCCCCC
Q 007391 381 GAKDYVRKEIELLRTKAHWKKAYFYLWDEPLNMEHYSS----VRNMASELHAY----APDARVLTTYYCGPSDAPLGPTP 452 (605)
Q Consensus 381 ~~~~~L~~~~~hL~~kGw~~~~y~y~~DEP~~~~~~~~----~~~~~~~ir~~----~P~~ki~~t~~~~p~~~~~~~~~ 452 (605)
.+.+||..|+...+..|.-..+ +.+-.||.-...++. .++..++++++ ....||+.-. +
T Consensus 150 ~yA~~l~~fv~~m~~nGvnlya-lSVQNEPd~~p~~d~~~wtpQe~~rF~~qyl~si~~~~rV~~pe------------s 216 (433)
T COG5520 150 DYADYLNDFVLEMKNNGVNLYA-LSVQNEPDYAPTYDWCWWTPQEELRFMRQYLASINAEMRVIIPE------------S 216 (433)
T ss_pred HHHHHHHHHHHHHHhCCCceeE-EeeccCCcccCCCCcccccHHHHHHHHHHhhhhhccccEEecch------------h
Confidence 4778999999999999974443 356799953323332 33455666664 3346777632 1
Q ss_pred cccc--ccccccCCCc----cccccccccccCcchhhhHHHHhhcccCCCceeEEEeeCCCCCCCCCccccCchhhHHHH
Q 007391 453 FESF--VKVPKFLRPH----TQIYCTSEWVLGNREDLVKDIVTELQPENGEEWWTYVCMGPSDPHPNWHLGMRGSQHRAV 526 (605)
Q Consensus 453 ~e~~--v~vp~~~~~~----idi~c~~~wv~g~~~~~~~~~~~~~~~~~G~~~W~Y~C~~p~~~~pN~fid~p~~~~R~l 526 (605)
+... ..-|++.+|+ ++|-- -.|-.++-.+++.-. ++ +...||.+|+=-|..+.. =||.-.-.-.--.--+
T Consensus 217 ~~~~~~~~dp~lnDp~a~a~~~ilg-~H~Ygg~v~~~p~~l-ak-~~~~gKdlwmte~y~~es-d~~s~dr~~~~~~~hi 292 (433)
T COG5520 217 FKDLPNMSDPILNDPKALANMDILG-THLYGGQVSDQPYPL-AK-QKPAGKDLWMTECYPPES-DPNSADREALHVALHI 292 (433)
T ss_pred cccccccccccccCHhHhcccceeE-eeecccccccchhhH-hh-CCCcCCceEEeecccCCC-CCCcchHHHHHHHHHH
Confidence 1111 1124444442 22211 011112222222222 22 236699999777777653 2443221000000113
Q ss_pred HHHHHHcCCcEEEEeecc
Q 007391 527 MWRVWKEGGTGFLYWGAN 544 (605)
Q Consensus 527 gW~~~k~g~~GfL~W~~n 544 (605)
.--..+-|+.||+.|..-
T Consensus 293 ~~gm~~gg~~ayv~W~i~ 310 (433)
T COG5520 293 HIGMTEGGFQAYVWWNIR 310 (433)
T ss_pred HhhccccCccEEEEEEEe
Confidence 334567789999999864
No 19
>PF02449 Glyco_hydro_42: Beta-galactosidase; InterPro: IPR013529 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. This group of beta-galactosidase enzymes (3.2.1.23 from EC) belong to the glycosyl hydrolase 42 family GH42 from CAZY. The enzyme catalyses the hydrolysis of terminal, non-reducing terminal beta-D-galactosidase residues.; GO: 0004565 beta-galactosidase activity, 0005975 carbohydrate metabolic process, 0009341 beta-galactosidase complex; PDB: 1KWK_A 1KWG_A 3U7V_A.
Probab=44.41 E-value=1.1e+02 Score=33.10 Aligned_cols=86 Identities=17% Similarity=0.311 Sum_probs=39.7
Q ss_pred CCCceeEEEe-eCCCCCCCCCccccCchhhHHHHHHHHHHcCCcEEEEeeccccCCCCCCcccccccCCCCCCccEEEcc
Q 007391 494 ENGEEWWTYV-CMGPSDPHPNWHLGMRGSQHRAVMWRVWKEGGTGFLYWGANCYEKATVPSAEIRFRRGLPPGDGVLFYP 572 (605)
Q Consensus 494 ~~G~~~W~Y~-C~~p~~~~pN~fid~p~~~~R~lgW~~~k~g~~GfL~W~~n~w~~~~dP~~d~~f~~~~~~GDg~LVYP 572 (605)
+.|++.|.=- +.++. .+...-..-.+-+.|...|++..+|.+|.++|.+...... .+. |.++. +
T Consensus 286 ~~~kpf~v~E~~~g~~-~~~~~~~~~~pg~~~~~~~~~~A~Ga~~i~~~~wr~~~~g---~E~--~~~g~-------~-- 350 (374)
T PF02449_consen 286 AKGKPFWVMEQQPGPV-NWRPYNRPPRPGELRLWSWQAIAHGADGILFWQWRQSRFG---AEQ--FHGGL-------V-- 350 (374)
T ss_dssp TTT--EEEEEE--S---SSSSS-----TTHHHHHHHHHHHTT-S-EEEC-SB--SSS---TTT--TS--S-------B--
T ss_pred cCCCceEeecCCCCCC-CCccCCCCCCCCHHHHHHHHHHHHhCCeeEeeeccCCCCC---chh--hhccc-------C--
Confidence 5788888532 22221 1222222334467899999999999999999997654211 111 11111 1
Q ss_pred CCCCCCCCCcccchhHHHHHHHHhHHH
Q 007391 573 GEVFSSSRQPVASLRLERILSGLQVRW 599 (605)
Q Consensus 573 G~~~~~~~~PvsSiRle~lreGiqDye 599 (605)
+ .++...+.|++-+.+-.++.+
T Consensus 351 ~-----~dg~~~~~~~~e~~~~~~~l~ 372 (374)
T PF02449_consen 351 D-----HDGREPTRRYREVAQLGRELK 372 (374)
T ss_dssp ------TTS--B-HHHHHHHHHHHHHH
T ss_pred C-----ccCCCCCcHHHHHHHHHHHHh
Confidence 1 223357788888887777654
No 20
>TIGR02186 alph_Pro_TM conserved hypothetical protein. This family consists of predicted transmembrane proteins of about 270 amino acids. Members are found, so far, only among the Alphaproteobacteria and only once in each genome.
Probab=42.99 E-value=31 Score=35.97 Aligned_cols=41 Identities=12% Similarity=0.169 Sum_probs=31.7
Q ss_pred eeeecCCCeeEEEEEEEcCCCCCCceeEEEEEEEEcCCCcccccc
Q 007391 152 QISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITSKADTELSSQC 196 (605)
Q Consensus 152 ~~~l~~~~~q~vWV~v~VP~~a~pG~Y~g~v~V~~~~~g~~~~~~ 196 (605)
.+.+..++ =.--+|.+|++.+.|.|+.++-+. .+|+++++.
T Consensus 173 gV~~~~~~--LFra~i~LPAnvp~G~Y~v~v~L~--r~G~vv~~~ 213 (261)
T TIGR02186 173 GVTFISPT--LFRATLRLPANVPNGTHEVRAYLF--RGGVFIART 213 (261)
T ss_pred eEEEcCCc--eEEEeeecCCCCCCceEEEEEEEE--eCCEEEEEE
Confidence 45554433 367789999999999999999998 588877653
No 21
>PF08428 Rib: Rib/alpha-like repeat; InterPro: IPR012706 This entry represents a region of about 79 amino acids found tandemly repeated up to fourteen times within the proteins that contain it. The repeats lack cysteines and are highly conserved, even at the DNA level, within and between proteins []. Proteins containing these repeats include the Rib and alpha surface antigens of group B Streptococcus, Esp of Enterococcus faecalis (Streptococcus faecalis), and related proteins of Lactobacillus. Most members of this protein family also have the cell wall anchor motif, LPXTG, shared by many staphyloccal and streptococcal surface antigens. These repeats are thought to define protective epitopes and may play a role in generating phenotypic and genotypic variation [].
Probab=41.45 E-value=49 Score=26.95 Aligned_cols=33 Identities=30% Similarity=0.537 Sum_probs=24.0
Q ss_pred eecCCCeeEEEEEEEcCCCCCCceeEEEEEEEEcCCC
Q 007391 154 SLIPGETTAVWVSIDAPYAQPPGLYEGEIIITSKADT 190 (605)
Q Consensus 154 ~l~~~~~q~vWV~v~VP~~a~pG~Y~g~v~V~~~~~g 190 (605)
+++.|. .--|.+ .|...++|.|+++|+|+. .+|
T Consensus 20 ~lP~gt-~~~w~~--~pdt~~~G~~~~~V~Vty-pDg 52 (65)
T PF08428_consen 20 NLPAGT-TYSWKD--KPDTSKPGTKTGKVKVTY-PDG 52 (65)
T ss_pred cCCCCc-ceeecc--CCccccCccEEEEEEEEc-CCC
Confidence 344443 346666 899999999999999996 344
No 22
>smart00737 ML Domain involved in innate immunity and lipid metabolism. ML (MD-2-related lipid-recognition) is a novel domain identified in MD-1, MD-2, GM2A, Npc2 and multiple proteins of unknown function in plants, animals and fungi. These single-domain proteins were predicted to form a beta-rich fold containing multiple strands, and to mediate diverse biological functions through interacting with specific lipids.
Probab=36.05 E-value=69 Score=28.53 Aligned_cols=35 Identities=29% Similarity=0.361 Sum_probs=29.3
Q ss_pred eeeecCCCeeEEEEEEEcCCCCCCceeEEEEEEEE
Q 007391 152 QISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (605)
Q Consensus 152 ~~~l~~~~~q~vWV~v~VP~~a~pG~Y~g~v~V~~ 186 (605)
...+.+|+..-.=.++.||...++|.|+++++++.
T Consensus 71 ~CPl~~G~~~~~~~~~~v~~~~P~~~~~v~~~l~d 105 (118)
T smart00737 71 KCPIEKGETVNYTNSLTVPGIFPPGKYTVKWELTD 105 (118)
T ss_pred CCCCCCCeeEEEEEeeEccccCCCeEEEEEEEEEc
Confidence 56788888655556779999999999999999994
No 23
>PF13204 DUF4038: Protein of unknown function (DUF4038); PDB: 3KZS_D.
Probab=29.20 E-value=4.1e+02 Score=27.89 Aligned_cols=201 Identities=13% Similarity=0.115 Sum_probs=88.3
Q ss_pred ccccCChhHHHHHHHHHHHHHhCCcCc-cccccCCcceeeeecCCCCCCCCCCccccCCcccccccccCCCCCCChhHHH
Q 007391 305 GVRHGSDEWYEALDQHFKWLLQYRISP-FFCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAK 383 (605)
Q Consensus 305 ~v~~~~~~~~~~ldrw~~~~~~~~is~-~f~~wg~~~~i~~y~~pw~~~~~~~~~yf~~~~~~~Y~~~~~~~~~g~~~~~ 383 (605)
+..+-+..||+.+|+-++.+.++||-. +..-||.+.. .--| +. + +-+=+.+..+
T Consensus 78 d~~~~N~~YF~~~d~~i~~a~~~Gi~~~lv~~wg~~~~----~~~W----g~-----~------------~~~m~~e~~~ 132 (289)
T PF13204_consen 78 DFTRPNPAYFDHLDRRIEKANELGIEAALVPFWGCPYV----PGTW----GF-----G------------PNIMPPENAE 132 (289)
T ss_dssp --TT----HHHHHHHHHHHHHHTT-EEEEESS-HHHHH--------------------------------TTSS-HHHHH
T ss_pred CCCCCCHHHHHHHHHHHHHHHHCCCeEEEEEEECCccc----cccc----cc-----c------------ccCCCHHHHH
Confidence 344455678999999999999999986 2223432110 0011 00 0 0011223567
Q ss_pred HHHHHHHHHHHHcCccceeeeeecCCCCCccchHHHHHHHHHHHHhCCCCcEEEeec-cCCCCCCCCCCCcccccccccc
Q 007391 384 DYVRKEIELLRTKAHWKKAYFYLWDEPLNMEHYSSVRNMASELHAYAPDARVLTTYY-CGPSDAPLGPTPFESFVKVPKF 462 (605)
Q Consensus 384 ~~L~~~~~hL~~kGw~~~~y~y~~DEP~~~~~~~~~~~~~~~ir~~~P~~ki~~t~~-~~p~~~~~~~~~~e~~v~vp~~ 462 (605)
.|++=+++.+++.- ..+++..-|.-......+.++++++.|++.+|.- +.|+. ++ ..+..+.+-+.
T Consensus 133 ~Y~~yv~~Ry~~~~--NviW~l~gd~~~~~~~~~~w~~~~~~i~~~dp~~--L~T~H~~~------~~~~~~~~~~~--- 199 (289)
T PF13204_consen 133 RYGRYVVARYGAYP--NVIWILGGDYFDTEKTRADWDAMARGIKENDPYQ--LITIHPCG------RTSSPDWFHDE--- 199 (289)
T ss_dssp HHHHHHHHHHTT-S--SEEEEEESSS--TTSSHHHHHHHHHHHHHH--SS---EEEEE-B------TEBTHHHHTT----
T ss_pred HHHHHHHHHHhcCC--CCEEEecCccCCCCcCHHHHHHHHHHHHhhCCCC--cEEEeCCC------CCCcchhhcCC---
Confidence 77887777777552 2334444555123567888899999999999966 44442 21 10111111111
Q ss_pred CCCccccccccccccCcchhhhHHH--HhhcccCCCceeEEEeeCCCCCCCCCccc---cCchhhHHHHHHHHHHcCC-c
Q 007391 463 LRPHTQIYCTSEWVLGNREDLVKDI--VTELQPENGEEWWTYVCMGPSDPHPNWHL---GMRGSQHRAVMWRVWKEGG-T 536 (605)
Q Consensus 463 ~~~~idi~c~~~wv~g~~~~~~~~~--~~~~~~~~G~~~W~Y~C~~p~~~~pN~fi---d~p~~~~R~lgW~~~k~g~-~ 536 (605)
+-+|..+.-........+..... ....+....|++.-=-||...-+ .+..- .....+-|--.|.+..-|. -
T Consensus 200 --~Wldf~~~Qsgh~~~~~~~~~~~~~~~~~~~~p~KPvin~Ep~YEg~~-~~~~~~~~~~~~~dvrr~aw~svlaGa~a 276 (289)
T PF13204_consen 200 --PWLDFNMYQSGHNRYDQDNWYYLPEEFDYRRKPVKPVINGEPCYEGIP-YSRWGYNGRFSAEDVRRRAWWSVLAGAYA 276 (289)
T ss_dssp --TT--SEEEB--S--TT--THHHH--HHHHTSSS---EEESS---BT-B-TTSS-TS-B--HHHHHHHHHHHHHCT--S
T ss_pred --CcceEEEeecCCCcccchHHHHHhhhhhhhhCCCCCEEcCcccccCCC-CCcCcccCCCCHHHHHHHHHHHHhcCCCc
Confidence 22444433221110011111111 12324467777754346664321 11221 2355677788999999999 9
Q ss_pred EEEEeecccc
Q 007391 537 GFLYWGANCY 546 (605)
Q Consensus 537 GfL~W~~n~w 546 (605)
|+-|.+-.-|
T Consensus 277 G~tYG~~~iW 286 (289)
T PF13204_consen 277 GHTYGAHGIW 286 (289)
T ss_dssp EEEE-BHHHH
T ss_pred cccCCCCCcc
Confidence 9999876656
No 24
>PF12245 Big_3_2: Bacterial Ig-like domain (group 3); InterPro: IPR022038 This family of proteins is found in bacteria. They have two conserved sequence motifs: AGN and GMT.
Probab=25.01 E-value=1.5e+02 Score=23.54 Aligned_cols=46 Identities=22% Similarity=0.289 Sum_probs=32.9
Q ss_pred EEEEEEEcCCCCCCceeEEEEEEEEcCCCccccccccccchhhhccccceeeeeccCCCCCC
Q 007391 162 AVWVSIDAPYAQPPGLYEGEIIITSKADTELSSQCLGKGEKHRLFMELRNCLDNVEPIEGKP 223 (605)
Q Consensus 162 ~vWV~v~VP~~a~pG~Y~g~v~V~~~~~g~~~~~~~~~~~~~~~~~~l~~~l~V~~~~lp~~ 223 (605)
.+|.. -+|.+...|.|+.+++++.+++.. .+.....-+.+...|.|
T Consensus 10 ~~~~~-~~P~~~~dg~yt~~v~a~D~AGN~---------------~~~~~~~~i~d~~~p~p 55 (60)
T PF12245_consen 10 GVWST-VIPENDADGEYTLTVTATDKAGNT---------------SSSTTQIVIVDNTAPAP 55 (60)
T ss_pred cceec-cccCccCCccEEEEEEEEECCCCE---------------EEeeeEEEEEcCCCCCc
Confidence 44543 369988899999999999655432 24466777778887766
No 25
>PF09099 Qn_am_d_aIII: Quinohemoprotein amine dehydrogenase, alpha subunit domain III; InterPro: IPR015183 This domain is predominantly found in the prokaryotic protein quinohemoprotein amine dehydrogenase, adopting an immunoglobulin-like beta-sandwich fold, with seven strands arranged into two beta sheets; the fold is possibly related to the immunoglobulin and/or fibronectin type III superfamilies. The precise function of this domain has not, as yet, been defined []. ; PDB: 1JMZ_A 1JMX_A 1PBY_A 1JJU_A.
Probab=24.90 E-value=71 Score=27.43 Aligned_cols=22 Identities=23% Similarity=0.353 Sum_probs=19.2
Q ss_pred eEEEEEEEcCCCCCCceeEEEE
Q 007391 161 TAVWVSIDAPYAQPPGLYEGEI 182 (605)
Q Consensus 161 q~vWV~v~VP~~a~pG~Y~g~v 182 (605)
--|+++|.+.++++||.|+..+
T Consensus 48 ~~v~v~V~~aa~a~~G~~~v~v 69 (81)
T PF09099_consen 48 DEVVVRVKAAADAAPGIRTVRV 69 (81)
T ss_dssp TCEEEEEEEECTSSSEEEEEEE
T ss_pred CEEEEEEEEcCCCCCccEEEEe
Confidence 3589999999999999998665
No 26
>TIGR03769 P_ac_wall_RPT actinobacterial surface-anchored protein domain. This model describes a repeat domain that one to three times in Actinobacterial proteins, some of which have LPXTG-type sortase recognition motifs for covalent attachment to the Gram-positive cell wall. Where it occurs with duplication in an LPXTG-anchored protein, it tends to be adjacent to the substrate-binding protein of the gene trio of an ABC transporter system, where that substrate-binding protein has a single copy of this same domain. This arrangement suggests a substrate-binding relay system, with the LPXTG protein acting as a substrate receptor.
Probab=24.55 E-value=88 Score=23.29 Aligned_cols=13 Identities=31% Similarity=0.473 Sum_probs=11.9
Q ss_pred CCCceeEEEEEEE
Q 007391 173 QPPGLYEGEIIIT 185 (605)
Q Consensus 173 a~pG~Y~g~v~V~ 185 (605)
.+||.|+.+++.+
T Consensus 10 T~PG~Y~l~~~a~ 22 (41)
T TIGR03769 10 TKPGTYTLTVQAT 22 (41)
T ss_pred CCCeEEEEEEEEE
Confidence 5899999999997
No 27
>PF00868 Transglut_N: Transglutaminase family; InterPro: IPR001102 Synonym(s): Protein-glutamine gamma-glutamyltransferase, Fibrinoligase, TGase Protein-glutamine gamma-glutamyltransferases (2.3.2.13 from EC) (TGase) are calcium-dependent enzymes that catalyse the cross-linking of proteins by promoting the formation of isopeptide bonds between the gamma-carboxyl group of a glutamine in one polypeptide chain and the epsilon-amino group of a lysine in a second polypeptide chain. TGases also catalyse the conjugation of polyamines to proteins [, ]. Transglutaminases are widely distributed in various organs, tissues and body fluids. The best known transglutaminase is blood coagulation factor XIII, a plasma tetrameric protein composed of two catalytic A subunits and two non-catalytic B subunits. Factor XIII is responsible for cross-linking fibrin chains, thus stabilising the fibrin clot. There are commonly three domains: N-terminal, middle (IPR013808 from INTERPRO) and C-terminal (IPR013807 from INTERPRO). This entry represents the N-terminal domain found in transglutaminases.; GO: 0018149 peptide cross-linking; PDB: 1L9N_B 1NUF_A 1NUD_A 1NUG_B 1L9M_A 1KV3_C 3S3S_A 2Q3Z_A 3LY6_A 3S3P_A ....
Probab=24.00 E-value=80 Score=28.75 Aligned_cols=33 Identities=21% Similarity=0.307 Sum_probs=22.0
Q ss_pred eeecCCCeeEEEEEEEcCCCCCCceeEEEEEEE
Q 007391 153 ISLIPGETTAVWVSIDAPYAQPPGLYEGEIIIT 185 (605)
Q Consensus 153 ~~l~~~~~q~vWV~v~VP~~a~pG~Y~g~v~V~ 185 (605)
..+...+-..+=|.|.+|++|.-|.|+.+|.++
T Consensus 85 a~v~~~~~~~~tv~V~spa~A~VG~y~l~v~~~ 117 (118)
T PF00868_consen 85 ARVESQDGNSVTVSVTSPANAPVGRYKLSVETK 117 (118)
T ss_dssp EEEEEEETTEEEEEEE--TTS--EEEEEEEEEE
T ss_pred EEEEecCCCEEEEEEECCCCCceEEEEEEEEEe
Confidence 333444445588899999999999999999886
No 28
>PF04234 CopC: CopC domain; InterPro: IPR007348 CopC is a bacterial blue copper protein that binds 1 atom of copper per protein molecule. Along with CopA, CopC mediates copper resistance by sequestration of copper in the periplasm [].; GO: 0005507 copper ion binding, 0046688 response to copper ion, 0042597 periplasmic space; PDB: 1IX2_B 1LYQ_A 2C9P_C 2C9R_A 2C9Q_A 1M42_A 1OT4_A 1NM4_A.
Probab=23.57 E-value=1.2e+02 Score=26.44 Aligned_cols=26 Identities=31% Similarity=0.517 Sum_probs=19.4
Q ss_pred EEEEcCCCCCCceeEEEEEEEEcCCCc
Q 007391 165 VSIDAPYAQPPGLYEGEIIITSKADTE 191 (605)
Q Consensus 165 V~v~VP~~a~pG~Y~g~v~V~~~~~g~ 191 (605)
+.+.+|..-++|.|+..-+|-+ +||-
T Consensus 61 ~~~~l~~~l~~G~YtV~wrvvs-~DGH 86 (97)
T PF04234_consen 61 LTVPLPPPLPPGTYTVSWRVVS-ADGH 86 (97)
T ss_dssp EEEEESS---SEEEEEEEEEEE-TTSC
T ss_pred EEEECCCCCCCceEEEEEEEEe-cCCC
Confidence 5788899899999999999975 6664
No 29
>COG3693 XynA Beta-1,4-xylanase [Carbohydrate transport and metabolism]
Probab=20.92 E-value=2.1e+02 Score=30.96 Aligned_cols=106 Identities=14% Similarity=0.182 Sum_probs=59.6
Q ss_pred ccCChhH-HHHHHHHHHHHHhCCcCccc--cccCCcceeeeecCCCCCCCCCCccccCCcccccccccCCCCCCChh---
Q 007391 307 RHGSDEW-YEALDQHFKWLLQYRISPFF--CRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSND--- 380 (605)
Q Consensus 307 ~~~~~~~-~~~ldrw~~~~~~~~is~~f--~~wg~~~~i~~y~~pw~~~~~~~~~yf~~~~~~~Y~~~~~~~~~g~~--- 380 (605)
....+.| |+.-|+.++|+++|++.-.+ .-|... .-+| ++.+. +++..
T Consensus 73 ~p~~G~f~Fe~AD~ia~FAr~h~m~lhGHtLvW~~q------~P~W---------~~~~e------------~~~~~~~~ 125 (345)
T COG3693 73 EPERGRFNFEAADAIANFARKHNMPLHGHTLVWHSQ------VPDW---------LFGDE------------LSKEALAK 125 (345)
T ss_pred cCCCCccCccchHHHHHHHHHcCCeeccceeeeccc------CCch---------hhccc------------cChHHHHH
Confidence 3345567 89999999999999998322 124321 1233 11111 11111
Q ss_pred HHHHHHHHHHHHHHHc-CccceeeeeecCCCCC--------ccchHHHHHHHHHHHHhCCCCcEEEee
Q 007391 381 GAKDYVRKEIELLRTK-AHWKKAYFYLWDEPLN--------MEHYSSVRNMASELHAYAPDARVLTTY 439 (605)
Q Consensus 381 ~~~~~L~~~~~hL~~k-Gw~~~~y~y~~DEP~~--------~~~~~~~~~~~~~ir~~~P~~ki~~t~ 439 (605)
..+..+...++|.|.. -.||-+-=-+.|+|.. ..-.+.|+.+....|+++|+.|.+.-.
T Consensus 126 ~~e~hI~tV~~rYkg~~~sWDVVNE~vdd~g~~R~s~w~~~~~gpd~I~~aF~~AreadP~AkL~~ND 193 (345)
T COG3693 126 MVEEHIKTVVGRYKGSVASWDVVNEAVDDQGSLRRSAWYDGGTGPDYIKLAFHIAREADPDAKLVIND 193 (345)
T ss_pred HHHHHHHHHHHhccCceeEEEecccccCCCchhhhhhhhccCCccHHHHHHHHHHHhhCCCceEEeec
Confidence 3344455555555431 1344333333455521 123456788999999999999998844
No 30
>PF14734 DUF4469: Domain of unknown function (DUF4469) with IG-like fold
Probab=20.19 E-value=1e+02 Score=27.59 Aligned_cols=25 Identities=16% Similarity=0.092 Sum_probs=21.4
Q ss_pred EEEEEEEcCCCCCCceeEEEEEEEE
Q 007391 162 AVWVSIDAPYAQPPGLYEGEIIITS 186 (605)
Q Consensus 162 ~vWV~v~VP~~a~pG~Y~g~v~V~~ 186 (605)
|==+.+.||++.++|.|+.+|+=..
T Consensus 63 ps~l~~~lPa~L~~G~Y~l~V~Tq~ 87 (102)
T PF14734_consen 63 PSRLIFILPADLAAGEYTLEVRTQY 87 (102)
T ss_pred CcEEEEECcCccCceEEEEEEEEEe
Confidence 4447899999999999999999874
Done!