Query 013386
Match_columns 444
No_of_seqs 131 out of 138
Neff 4.9
Searched_HMMs 46136
Date Fri Mar 29 03:19:59 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/013386.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/013386hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF10633 NPCBM_assoc: NPCBM-as 93.5 0.14 3.1E-06 41.3 4.8 32 154-185 45-76 (78)
2 PF00150 Cellulase: Cellulase 90.1 0.63 1.4E-05 44.5 5.9 104 310-439 57-172 (281)
3 PF15418 DUF4625: Domain of un 89.9 2.4 5.2E-05 38.5 9.0 88 77-189 29-120 (132)
4 PF01229 Glyco_hydro_39: Glyco 89.8 1 2.2E-05 48.4 7.8 108 313-441 82-205 (486)
5 COG1470 Predicted membrane pro 86.5 5.7 0.00012 43.2 10.6 38 150-187 324-361 (513)
6 PF06030 DUF916: Bacterial pro 79.7 44 0.00096 29.7 12.5 108 62-186 4-120 (121)
7 PF14352 DUF4402: Domain of un 74.3 3.6 7.8E-05 36.3 3.5 31 156-186 96-128 (130)
8 PF13731 WxL: WxL domain surfa 73.0 13 0.00028 35.7 7.3 79 106-185 105-210 (215)
9 COG1470 Predicted membrane pro 71.8 5.5 0.00012 43.3 4.8 33 154-186 437-469 (513)
10 smart00633 Glyco_10 Glycosyl h 69.7 20 0.00044 35.0 8.0 98 312-439 13-125 (254)
11 cd00917 PG-PI_TP The phosphati 67.4 9.9 0.00021 33.4 4.7 33 153-186 77-109 (122)
12 PF02221 E1_DerP2_DerF2: ML do 64.2 12 0.00026 32.4 4.6 35 153-187 86-120 (134)
13 PF01835 A2M_N: MG2 domain; I 62.8 33 0.00071 28.3 6.8 28 158-185 59-86 (99)
14 PF10003 DUF2244: Integral mem 60.2 7.5 0.00016 35.3 2.7 53 159-213 88-140 (140)
15 PF06280 DUF1034: Fn3-like dom 59.0 9.8 0.00021 32.6 3.1 37 151-187 62-101 (112)
16 PF13204 DUF4038: Protein of u 52.7 59 0.0013 32.8 7.9 103 309-442 82-188 (289)
17 smart00737 ML Domain involved 46.9 37 0.00081 29.0 4.8 34 153-186 72-105 (118)
18 PF13304 AAA_21: AAA domain; P 37.1 38 0.00082 30.2 3.4 38 403-440 259-297 (303)
19 PF09099 Qn_am_d_aIII: Quinohe 28.2 59 0.0013 27.3 2.9 22 161-182 48-69 (81)
20 PF00868 Transglut_N: Transglu 28.1 68 0.0015 28.3 3.4 27 159-185 91-117 (118)
21 PF12891 Glyco_hydro_44: Glyco 27.6 95 0.0021 31.2 4.7 29 414-442 154-182 (239)
22 PLN02475 5-methyltetrahydropte 25.6 88 0.0019 36.3 4.6 60 380-443 179-247 (766)
23 PF02228 Gag_p19: Major core p 25.1 54 0.0012 27.8 2.1 45 283-327 7-56 (92)
24 PF14734 DUF4469: Domain of un 25.0 74 0.0016 27.8 3.0 23 164-186 65-87 (102)
25 PF09608 Alph_Pro_TM: Putative 23.9 1E+02 0.0022 30.7 4.2 34 151-186 147-180 (236)
26 PF05205 COMPASS-Shg1: COMPASS 22.0 82 0.0018 27.4 2.7 37 387-423 1-37 (106)
27 PF08428 Rib: Rib/alpha-like r 20.5 2.9E+02 0.0063 21.8 5.4 31 153-186 19-49 (65)
28 PRK01254 hypothetical protein; 20.4 1.7E+02 0.0037 33.7 5.4 71 367-442 489-566 (707)
29 PF09153 DUF1938: Domain of un 20.2 82 0.0018 26.9 2.2 24 375-398 30-53 (86)
No 1
>PF10633 NPCBM_assoc: NPCBM-associated, NEW3 domain of alpha-galactosidase; InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=93.45 E-value=0.14 Score=41.29 Aligned_cols=32 Identities=31% Similarity=0.543 Sum_probs=25.7
Q ss_pred eeCCCCeeEEEEEEEcCCCCCCceeEEEEEEE
Q 013386 154 SLIPGETTAVWVSIDAPYAQPPGLYEGEIIIT 185 (444)
Q Consensus 154 ~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt 185 (444)
.|++|+.+.+=++|.+|++++||.|..+++++
T Consensus 45 ~l~pG~s~~~~~~V~vp~~a~~G~y~v~~~a~ 76 (78)
T PF10633_consen 45 SLPPGESVTVTFTVTVPADAAPGTYTVTVTAR 76 (78)
T ss_dssp -B-TTSEEEEEEEEEE-TT--SEEEEEEEEEE
T ss_pred cCCCCCEEEEEEEEECCCCCCCceEEEEEEEE
Confidence 68899999999999999999999999999886
No 2
>PF00150 Cellulase: Cellulase (glycosyl hydrolase family 5); InterPro: IPR001547 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. Glycoside hydrolase family 5 GH5 from CAZY comprises enzymes with several known activities; endoglucanase (3.2.1.4 from EC); beta-mannanase (3.2.1.78 from EC); exo-1,3-glucanase (3.2.1.58 from EC); endo-1,6-glucanase (3.2.1.75 from EC); xylanase (3.2.1.8 from EC); endoglycoceramidase (3.2.1.123 from EC). The microbial degradation of cellulose and xylans requires several types of enzymes. Fungi and bacteria produces a spectrum of cellulolytic enzymes (cellulases) and xylanases which, on the basis of sequence similarities, can be classified into families. One of these families is known as the cellulase family A [] or as the glycosyl hydrolases family 5 []. One of the conserved regions in this family contains a conserved glutamic acid residue which is potentially involved [] in the catalytic mechanism.; GO: 0004553 hydrolase activity, hydrolyzing O-glycosyl compounds, 0005975 carbohydrate metabolic process; PDB: 3NDY_A 3NDZ_B 1LF1_A 1TVP_B 1TVN_A 3AYR_A 3AYS_A 1QI0_A 1W3K_A 1OCQ_A ....
Probab=90.14 E-value=0.63 Score=44.54 Aligned_cols=104 Identities=13% Similarity=0.129 Sum_probs=66.6
Q ss_pred CHHHHHHHHHHHHHHHhCcccCCCcCCCCceEEEeecCCCCCCCCcccccccccccceeecccCCCCCChHHHHHHHHHH
Q 013386 310 SDEWYEALDQHFKWLLQYRISPFFCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYVRKE 389 (444)
Q Consensus 310 s~~~f~~L~~~~~~~~~~riS~~f~~Wg~~mrv~~~~~~w~~d~~~~~~y~~d~~~~~y~vp~~~~~~g~~~~~~~L~~~ 389 (444)
.+..++.|++.+++|.++.|.....-.+. ..|..+.... . ....-.+.++++++.+
T Consensus 57 ~~~~~~~ld~~v~~a~~~gi~vild~h~~--------~~w~~~~~~~---------~-------~~~~~~~~~~~~~~~l 112 (281)
T PF00150_consen 57 DETYLARLDRIVDAAQAYGIYVILDLHNA--------PGWANGGDGY---------G-------NNDTAQAWFKSFWRAL 112 (281)
T ss_dssp THHHHHHHHHHHHHHHHTT-EEEEEEEES--------TTCSSSTSTT---------T-------THHHHHHHHHHHHHHH
T ss_pred cHHHHHHHHHHHHHHHhCCCeEEEEeccC--------cccccccccc---------c-------cchhhHHHHHhhhhhh
Confidence 45788999999999999988754321111 1231111100 0 0000122466688889
Q ss_pred HHHHHhcCchhhhhhhhcCCCCCccc------------HHHHHHHHHHHHhhCCCCcEEEEE
Q 013386 390 IELLRTKAHWKKAYFYLWDEPLNMEH------------YSSVRNMASELHAYAPDARVLTTY 439 (444)
Q Consensus 390 ~~~Lr~kGw~~k~yfyl~DEP~~~e~------------~~~~r~a~~~ir~~~Pd~ril~t~ 439 (444)
++++|.. -....|=|+.||..... .+.++++++.||+..|+..|+...
T Consensus 113 a~~y~~~--~~v~~~el~NEP~~~~~~~~w~~~~~~~~~~~~~~~~~~Ir~~~~~~~i~~~~ 172 (281)
T PF00150_consen 113 AKRYKDN--PPVVGWELWNEPNGGNDDANWNAQNPADWQDWYQRAIDAIRAADPNHLIIVGG 172 (281)
T ss_dssp HHHHTTT--TTTEEEESSSSGCSTTSTTTTSHHHTHHHHHHHHHHHHHHHHTTSSSEEEEEE
T ss_pred ccccCCC--CcEEEEEecCCccccCCccccccccchhhhhHHHHHHHHHHhcCCcceeecCC
Confidence 9998743 34678889999996312 257899999999999998888765
No 3
>PF15418 DUF4625: Domain of unknown function (DUF4625)
Probab=89.85 E-value=2.4 Score=38.46 Aligned_cols=88 Identities=19% Similarity=0.252 Sum_probs=53.2
Q ss_pred EEEeecCceeEEEEEEccCCCcCCCCCCCceEEEeec-c--ccCCCCcccccCceEEEEeeecCCCCcccccCCCCccee
Q 013386 77 NLLAARNERESVQIALRPKVSWSSSSTAGVVQVQCSD-L--CSASGDRLVVGQSLMLRRVVPMLGVPDALVPLDLPVCQI 153 (444)
Q Consensus 77 ~LsAaRGE~vSfQlvl~s~~~~~~~~~~~~V~Vs~sd-L--~s~~g~~~i~~~~I~lr~V~yVlGyPD~LvP~d~~~~~v 153 (444)
.-.+-||+.+.|.+-+... ..++.++|++-. + .+-++ ..++. -.|... .+.+
T Consensus 29 ~~~~~~G~~ihfe~~i~d~------~~i~si~VeIH~nfd~H~h~~---~~~~~---------------~~~~~~-~~~~ 83 (132)
T PF15418_consen 29 CKVATRGDDIHFEADISDN------SAIKSIKVEIHNNFDHHTHST---EAGEC---------------EKPWVF-EQDY 83 (132)
T ss_pred CeEEecCCcEEEEEEEEcc------cceeEEEEEEecCcCcccccc---ccccc---------------ccCcEE-EEEE
Confidence 4567899999999999863 567788877721 1 11010 01000 011110 0112
Q ss_pred eeCCC-CeeEEEEEEEcCCCCCCceeEEEEEEEeccC
Q 013386 154 SLIPG-ETTAVWVSIDAPYAQPPGLYEGEIIITSKAD 189 (444)
Q Consensus 154 ~l~ag-~~q~lWI~V~VP~~a~pG~Y~GtVtVt~~~~ 189 (444)
.+..| .+.-+=..|.||++++||.|.-.|+|+.+++
T Consensus 84 ~~~~g~~~~~~h~~i~IPa~a~~G~YH~~i~VtD~~G 120 (132)
T PF15418_consen 84 DIYGGKKNYDFHEHIDIPADAPAGDYHFMITVTDAAG 120 (132)
T ss_pred cccCCcccEeEEEeeeCCCCCCCcceEEEEEEEECCC
Confidence 22222 3456678999999999999999999997544
No 4
>PF01229 Glyco_hydro_39: Glycosyl hydrolases family 39; InterPro: IPR000514 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. Glycoside hydrolase family 39 GH39 from CAZY comprises enzymes with several known activities; alpha-L-iduronidase (3.2.1.76 from EC); beta-xylosidase (3.2.1.37 from EC). The most highly conserved regions in these enzymes are located in their N-terminal sections. These contain a glutamic acid residue which, on the basis of similarities with other families of glycosyl hydrolases [], probably acts as the proton donor in their catalytic mechanism.; GO: 0004553 hydrolase activity, hydrolyzing O-glycosyl compounds, 0005975 carbohydrate metabolic process; PDB: 2BS9_D 2BFG_E 1W91_B 1UHV_D 1PX8_A.
Probab=89.78 E-value=1 Score=48.44 Aligned_cols=108 Identities=20% Similarity=0.279 Sum_probs=65.4
Q ss_pred HHHHHHHHHHHHHhCcccCC----CcCCCCceEEEeecCCCCCCCCcccccccccccceeecccCCCCCChHHHHHHHHH
Q 013386 313 WYEALDQHFKWLLQYRISPF----FCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYVRK 388 (444)
Q Consensus 313 ~f~~L~~~~~~~~~~riS~~----f~~Wg~~mrv~~~~~~w~~d~~~~~~y~~d~~~~~y~vp~~~~~~g~~~~~~~L~~ 388 (444)
.|..||+-++.+++.+|.|+ |.+=+-.... ...|.| .....| ....+.+.+++++
T Consensus 82 nf~~lD~i~D~l~~~g~~P~vel~f~p~~~~~~~---~~~~~~--------------~~~~~p----p~~~~~W~~lv~~ 140 (486)
T PF01229_consen 82 NFTYLDQILDFLLENGLKPFVELGFMPMALASGY---QTVFWY--------------KGNISP----PKDYEKWRDLVRA 140 (486)
T ss_dssp --HHHHHHHHHHHHCT-EEEEEE-SB-GGGBSS-----EETTT--------------TEE-S-----BS-HHHHHHHHHH
T ss_pred ChHHHHHHHHHHHHcCCEEEEEEEechhhhcCCC---Cccccc--------------cCCcCC----cccHHHHHHHHHH
Confidence 48999999999999999994 5431100000 000000 000011 1234589999999
Q ss_pred HHHHHHhc-Cc--hhhhhhhhcCCCCCc---------ccHHHHHHHHHHHHhhCCCCcEEEEEee
Q 013386 389 EIELLRTK-AH--WKKAYFYLWDEPLNM---------EHYSSVRNMASELHAYAPDARVLTTYYC 441 (444)
Q Consensus 389 ~~~~Lr~k-Gw--~~k~yfyl~DEP~~~---------e~~~~~r~a~~~ir~~~Pd~ril~t~~~ 441 (444)
+++|+..+ |. .++.||=+|.||... +=++.|+++++.||+++|++||--...|
T Consensus 141 ~~~h~~~RYG~~ev~~W~fEiWNEPd~~~f~~~~~~~ey~~ly~~~~~~iK~~~p~~~vGGp~~~ 205 (486)
T PF01229_consen 141 FARHYIDRYGIEEVSTWYFEIWNEPDLKDFWWDGTPEEYFELYDATARAIKAVDPELKVGGPAFA 205 (486)
T ss_dssp HHHHHHHHHHHHHHTTSEEEESS-TTSTTTSGGG-HHHHHHHHHHHHHHHHHH-TTSEEEEEEEE
T ss_pred HHHHHHhhcCCccccceeEEeCcCCCcccccCCCCHHHHHHHHHHHHHHHHHhCCCCcccCcccc
Confidence 99999764 32 445578789999751 2245788999999999999998766555
No 5
>COG1470 Predicted membrane protein [Function unknown]
Probab=86.51 E-value=5.7 Score=43.18 Aligned_cols=38 Identities=29% Similarity=0.464 Sum_probs=35.5
Q ss_pred cceeeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEec
Q 013386 150 VCQISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITSK 187 (444)
Q Consensus 150 ~~~v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~~ 187 (444)
...+.|.||+...+-+.|+.|++|.||.|..+|+++.+
T Consensus 324 vt~vkL~~gE~kdvtleV~ps~na~pG~Ynv~I~A~s~ 361 (513)
T COG1470 324 VTSVKLKPGEEKDVTLEVYPSLNATPGTYNVTITASSS 361 (513)
T ss_pred EEEEEecCCCceEEEEEEecCCCCCCCceeEEEEEecc
Confidence 56789999999999999999999999999999999863
No 6
>PF06030 DUF916: Bacterial protein of unknown function (DUF916); InterPro: IPR010317 This family consists of putative cell surface proteins, from Firmicutes, of unknown function.
Probab=79.68 E-value=44 Score=29.70 Aligned_cols=108 Identities=16% Similarity=0.222 Sum_probs=67.8
Q ss_pred cccCCCCCC-CCCCceEEEeecCceeEEEEEEccCCCcCCCCCCCceEEEeeccc-cCCCCcccccCceEEEEeeecC--
Q 013386 62 ANVGPQEMP-RPLEPINLLAARNERESVQIALRPKVSWSSSSTAGVVQVQCSDLC-SASGDRLVVGQSLMLRRVVPML-- 137 (444)
Q Consensus 62 ~KVfpde~P-~~~~~~~LsAaRGE~vSfQlvl~s~~~~~~~~~~~~V~Vs~sdL~-s~~g~~~i~~~~I~lr~V~yVl-- 137 (444)
.-|.|+..- .....+.|...-|+...+|+.+... ++....|++++.+=. +.+| .+.|..
T Consensus 4 ~p~~p~~Q~~~~~~YFdL~~~P~q~~~l~v~i~N~-----s~~~~tv~v~~~~A~Tn~nG------------~I~Y~~~~ 66 (121)
T PF06030_consen 4 TPVLPENQIDKNVSYFDLKVKPGQKQTLEVRITNN-----SDKEITVKVSANTATTNDNG------------VIDYSQNN 66 (121)
T ss_pred eecCCccccCCCCCeEEEEeCCCCEEEEEEEEEeC-----CCCCEEEEEEEeeeEecCCE------------EEEECCCC
Confidence 345666553 2357899999999999999999762 122223443333221 2222 112211
Q ss_pred -CC-CcccccCC---CCcceeeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 013386 138 -GV-PDALVPLD---LPVCQISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (444)
Q Consensus 138 -Gy-PD~LvP~d---~~~~~v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~ 186 (444)
.. +++-.++. .....+.|+|++.+-+=++|.+|+..-.|..-|-|.|+.
T Consensus 67 ~~~d~sl~~~~~~~v~~~~~Vtl~~~~sk~V~~~i~~P~~~f~G~ilGGi~~~e 120 (121)
T PF06030_consen 67 PKKDKSLKYPFSDLVKIPKEVTLPPNESKTVTFTIKMPKKAFDGIILGGIYFSE 120 (121)
T ss_pred cccCcccCcchHHhccCCcEEEECCCCEEEEEEEEEcCCCCcCCEEEeeEEEEe
Confidence 01 11111221 012349999999999999999999999999999999985
No 7
>PF14352 DUF4402: Domain of unknown function (DUF4402)
Probab=74.34 E-value=3.6 Score=36.30 Aligned_cols=31 Identities=26% Similarity=0.480 Sum_probs=25.0
Q ss_pred CCCCeeEEEE--EEEcCCCCCCceeEEEEEEEe
Q 013386 156 IPGETTAVWV--SIDAPYAQPPGLYEGEIIITS 186 (444)
Q Consensus 156 ~ag~~q~lWI--~V~VP~~a~pG~Y~GtVtVt~ 186 (444)
..+....++| ++.|++++++|.|+|+++|+.
T Consensus 96 ~~~g~~~~~VGGtL~v~~~~~~G~YsGt~~VtV 128 (130)
T PF14352_consen 96 DTGGSATFNVGGTLNVPANQAAGTYSGTFTVTV 128 (130)
T ss_pred cCCCcEEEEEEEEEEcCCCCCCeEEEEEEEEEE
Confidence 3444556666 589999999999999999986
No 8
>PF13731 WxL: WxL domain surface cell wall-binding
Probab=73.04 E-value=13 Score=35.70 Aligned_cols=79 Identities=24% Similarity=0.398 Sum_probs=48.0
Q ss_pred ceEEEeeccccCCCCcccccCceEEEEeeec--CC---CCc------ccccCCCCcceeeeCCCCeeEEE----------
Q 013386 106 VVQVQCSDLCSASGDRLVVGQSLMLRRVVPM--LG---VPD------ALVPLDLPVCQISLIPGETTAVW---------- 164 (444)
Q Consensus 106 ~V~Vs~sdL~s~~g~~~i~~~~I~lr~V~yV--lG---yPD------~LvP~d~~~~~v~l~ag~~q~lW---------- 164 (444)
.|+|+.++|++.+|.. +.+..|.+...... .+ -|- .|.+.......+...+++.+..|
T Consensus 105 ~L~v~~s~F~~~~~~~-L~ga~l~~~~~~~~~~~~~~~~~~~~~~~~~l~~~~~~~~v~~A~~~~g~G~~~~~~~~~~~~ 183 (215)
T PF13731_consen 105 TLTVKLSPFTNADGDT-LPGATLTFNNGKVQSTANNTNTPTTVSSNITLTPGGQAQTVMSAAKGQGQGTWSYSFGDQDAT 183 (215)
T ss_pred EEEEEeccccccCCcC-cccceEEecCceeEeecccccCCcccccceEeccCCcceeeEeecccccceEEEEEeCCcccc
Confidence 4788888999888654 55555655543322 11 111 12222211122333456666666
Q ss_pred ----EEEEcCCCCC--CceeEEEEEEE
Q 013386 165 ----VSIDAPYAQP--PGLYEGEIIIT 185 (444)
Q Consensus 165 ----I~V~VP~~a~--pG~Y~GtVtVt 185 (444)
|.+.||.++. +|.|+++|+=+
T Consensus 184 ~~~~v~L~VP~~~~~~ag~Yt~tlTWt 210 (215)
T PF13731_consen 184 ADTGVSLSVPANTAKQAGTYTATLTWT 210 (215)
T ss_pred cccceEEEeCCCCcccCCcEEEEEEEE
Confidence 8899999998 79999999876
No 9
>COG1470 Predicted membrane protein [Function unknown]
Probab=71.84 E-value=5.5 Score=43.30 Aligned_cols=33 Identities=36% Similarity=0.457 Sum_probs=30.5
Q ss_pred eeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 013386 154 SLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (444)
Q Consensus 154 ~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~ 186 (444)
.+.||+.-.+=++|.||++|.||.|+.+|+.++
T Consensus 437 sL~pge~~tV~ltI~vP~~a~aGdY~i~i~~ks 469 (513)
T COG1470 437 SLEPGESKTVSLTITVPEDAGAGDYRITITAKS 469 (513)
T ss_pred ccCCCCcceEEEEEEcCCCCCCCcEEEEEEEee
Confidence 357899999999999999999999999999987
No 10
>smart00633 Glyco_10 Glycosyl hydrolase family 10.
Probab=69.73 E-value=20 Score=35.04 Aligned_cols=98 Identities=10% Similarity=0.122 Sum_probs=62.6
Q ss_pred HHHHHHHHHHHHHHhCcccCC--CcCCCCceEEEeecCCCCCCCCcccccccccccceeecccCCCCCChHHHHHHHHHH
Q 013386 312 EWYEALDQHFKWLLQYRISPF--FCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYVRKE 389 (444)
Q Consensus 312 ~~f~~L~~~~~~~~~~riS~~--f~~Wg~~mrv~~~~~~w~~d~~~~~~y~~d~~~~~y~vp~~~~~~g~~~~~~~L~~~ 389 (444)
..|+.+++.+++|.++.|.-. .+-|+.+ .| .|+.+.. + -.-..++++|+++.
T Consensus 13 ~n~~~~D~~~~~a~~~gi~v~gH~l~W~~~-------------~P---~W~~~~~---------~-~~~~~~~~~~i~~v 66 (254)
T smart00633 13 FNFSGADAIVNFAKENGIKVRGHTLVWHSQ-------------TP---DWVFNLS---------K-ETLLARLENHIKTV 66 (254)
T ss_pred cChHHHHHHHHHHHHCCCEEEEEEEeeccc-------------CC---HhhhcCC---------H-HHHHHHHHHHHHHH
Confidence 458899999999999987632 1224332 22 2222100 0 00123567788888
Q ss_pred HHHHHhcCchhhhhhhhcCCCCCcc-------cH------HHHHHHHHHHHhhCCCCcEEEEE
Q 013386 390 IELLRTKAHWKKAYFYLWDEPLNME-------HY------SSVRNMASELHAYAPDARVLTTY 439 (444)
Q Consensus 390 ~~~Lr~kGw~~k~yfyl~DEP~~~e-------~~------~~~r~a~~~ir~~~Pd~ril~t~ 439 (444)
+.|++.+.. +.-++.||.+.. .+ +-++.+.+.+|+++|+++++..=
T Consensus 67 ~~ry~g~i~----~wdV~NE~~~~~~~~~~~~~w~~~~G~~~i~~af~~ar~~~P~a~l~~Nd 125 (254)
T smart00633 67 VGRYKGKIY----AWDVVNEALHDNGSGLRRSVWYQILGEDYIEKAFRYAREADPDAKLFYND 125 (254)
T ss_pred HHHhCCcce----EEEEeeecccCCCcccccchHHHhcChHHHHHHHHHHHHhCCCCEEEEec
Confidence 888886622 345788887521 12 57889999999999999998863
No 11
>cd00917 PG-PI_TP The phosphatidylinositol/phosphatidylglycerol transfer protein (PG/PI-TP) has been shown to bind phosphatidylglycerol and phosphatidylinositol, but the biological significance of this is still obscure. These proteins belong to the ML domain family.
Probab=67.42 E-value=9.9 Score=33.42 Aligned_cols=33 Identities=24% Similarity=0.374 Sum_probs=29.0
Q ss_pred eeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 013386 153 ISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (444)
Q Consensus 153 v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~ 186 (444)
=++.+|+.. +=.++.||...++|.|+++.++..
T Consensus 77 CPi~~G~~~-~~~~~~ip~~~P~g~y~v~~~l~d 109 (122)
T cd00917 77 CPIEPGDKF-LTKLVDLPGEIPPGKYTVSARAYT 109 (122)
T ss_pred CCcCCCcEE-EEEEeeCCCCCCCceEEEEEEEEC
Confidence 456788877 788899999999999999999986
No 12
>PF02221 E1_DerP2_DerF2: ML domain; InterPro: IPR003172 The MD-2-related lipid-recognition (ML) domain is implicated in lipid recognition, particularly in the recognition of pathogen related products. It has an immunoglobulin-like beta-sandwich fold similar to that of E-set Ig domains. This domain is present in the following proteins: Epididymal secretory protein E1 (also known as Niemann-Pick C2 protein), which is known to bind cholesterol. Niemann-Pick disease type C2 is a fatal hereditary disease characterised by accumulation of low-density lipoprotein-derived cholesterol in lysosomes []. House-dust mite allergen proteins such as Der f 2 from Dermatophagoides farinae and Der p 2 from Dermatophagoides pteronyssinus []. ; PDB: 2AG9_B 1G13_B 2AG2_B 2AG4_A 1TJJ_C 1PU5_C 1PUB_A 2AF9_A 3T6Q_D 3M7O_B ....
Probab=64.22 E-value=12 Score=32.39 Aligned_cols=35 Identities=26% Similarity=0.350 Sum_probs=31.9
Q ss_pred eeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEec
Q 013386 153 ISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITSK 187 (444)
Q Consensus 153 v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~~ 187 (444)
=++.+|+..-..+++.||...++|.|+++++++..
T Consensus 86 CPi~~G~~~~~~~~~~i~~~~p~~~~~i~~~l~d~ 120 (134)
T PF02221_consen 86 CPIKAGEYYTYTYTIPIPKIYPPGKYTIQWKLTDQ 120 (134)
T ss_dssp STBTTTEEEEEEEEEEESTTSSSEEEEEEEEEEET
T ss_pred CccCCCcEEEEEEEEEcccceeeEEEEEEEEEEeC
Confidence 46789999999999999999999999999999974
No 13
>PF01835 A2M_N: MG2 domain; InterPro: IPR002890 The proteinase-binding alpha-macroglobulins (A2M) [] are large glycoproteins found in the plasma of vertebrates, in the hemolymph of some invertebrates and in reptilian and avian egg white. A2M-like proteins are able to inhibit all four classes of proteinases by a 'trapping' mechanism. They have a peptide stretch, called the 'bait region', which contains specific cleavage sites for different proteinases. When a proteinase cleaves the bait region, a conformational change is induced in the protein, thus trapping the proteinase. The entrapped enzyme remains active against low molecular weight substrates, whilst its activity toward larger substrates is greatly reduced, due to steric hindrance. Following cleavage in the bait region, a thiol ester bond, formed between the side chains of a cysteine and a glutamine, is cleaved and mediates the covalent binding of the A2M-like protein to the proteinase. This family includes the N-terminal region of the alpha-2-macroglobulin family. The inhibitor domains belong to MEROPS inhibitor family I39.; GO: 0004866 endopeptidase inhibitor activity; PDB: 2B39_B 3KLS_B 3PRX_C 3KM9_B 3PVM_C 3CU7_A 4E0S_A 4A5W_A 4ACQ_C 2P9R_B ....
Probab=62.85 E-value=33 Score=28.32 Aligned_cols=28 Identities=21% Similarity=0.215 Sum_probs=18.5
Q ss_pred CCeeEEEEEEEcCCCCCCceeEEEEEEE
Q 013386 158 GETTAVWVSIDAPYAQPPGLYEGEIIIT 185 (444)
Q Consensus 158 g~~q~lWI~V~VP~~a~pG~Y~GtVtVt 185 (444)
...-.+-.++.+|++++.|.|+.++...
T Consensus 59 ~~~G~~~~~~~lp~~~~~G~y~i~~~~~ 86 (99)
T PF01835_consen 59 NENGIFSGSFQLPDDAPLGTYTIRVKTD 86 (99)
T ss_dssp TCTTEEEEEEE--SS---EEEEEEEEET
T ss_pred CCCCEEEEEEECCCCCCCEeEEEEEEEc
Confidence 3444677889999999999999999885
No 14
>PF10003 DUF2244: Integral membrane protein (DUF2244); InterPro: IPR019253 This entry consists of various bacterial putative membrane proteins with no known function.
Probab=60.22 E-value=7.5 Score=35.25 Aligned_cols=53 Identities=23% Similarity=0.324 Sum_probs=38.9
Q ss_pred CeeEEEEEEEcCCCCCCceeEEEEEEEeccCccccccccccchhhhhhhhhhccc
Q 013386 159 ETTAVWVSIDAPYAQPPGLYEGEIIITSKADTELSSQCLGKGEKHRLFMELRNCL 213 (444)
Q Consensus 159 ~~q~lWI~V~VP~~a~pG~Y~GtVtVt~~~~ge~~~~~~~~~~~~~~~~~~~~~l 213 (444)
+..+.|+.|.+..+..+ ..-.|++++++..-.-+.-|+++||..|+.+|+..|
T Consensus 88 ~~~~~w~rv~~~~~~~~--~~~~l~L~~~g~~veiG~fL~~~eR~~la~~L~~aL 140 (140)
T PF10003_consen 88 EFNPYWVRVELEEDPGP--GPPRLTLRSRGREVEIGRFLNPEEREELARELRRAL 140 (140)
T ss_pred EEcCCeEEEEEEcCCCC--CCcEEEEEECCEEEEEccCCCHHHHHHHHHHHHhhC
Confidence 45689999999998887 555666665434112234899999999999998754
No 15
>PF06280 DUF1034: Fn3-like domain (DUF1034); InterPro: IPR010435 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This domain of unknown function is present in bacterial and plant peptidases belonging to MEROPS peptidase family S8 (subfamily S8A subtilisin, clan SB). It is C-terminal to and adjacent to the S8 peptidase domain and can be found in conjunction with the PA (Protease associated) domain (IPR003137 from INTERPRO) and additionally in Gram-positive bacteria with the surface protein anchor domain (IPR001899 from INTERPRO).; GO: 0004252 serine-type endopeptidase activity, 0005618 cell wall, 0016020 membrane; PDB: 3EIF_A 1XF1_B.
Probab=58.96 E-value=9.8 Score=32.62 Aligned_cols=37 Identities=27% Similarity=0.416 Sum_probs=31.2
Q ss_pred ceeeeCCCCeeEEEEEEEcCCCCCC---ceeEEEEEEEec
Q 013386 151 CQISLIPGETTAVWVSIDAPYAQPP---GLYEGEIIITSK 187 (444)
Q Consensus 151 ~~v~l~ag~~q~lWI~V~VP~~a~p---G~Y~GtVtVt~~ 187 (444)
..+.|+||+++-+=|++.+|++..+ ..|.|-|.++..
T Consensus 62 ~~vTV~ag~s~~v~vti~~p~~~~~~~~~~~eG~I~~~~~ 101 (112)
T PF06280_consen 62 DTVTVPAGQSKTVTVTITPPSGLDASNGPFYEGFITFKSS 101 (112)
T ss_dssp EEEEE-TTEEEEEEEEEE--GGGHHTT-EEEEEEEEEESS
T ss_pred CeEEECCCCEEEEEEEEEehhcCCcccCCEEEEEEEEEcC
Confidence 5799999999999999999998887 899999999973
No 16
>PF13204 DUF4038: Protein of unknown function (DUF4038); PDB: 3KZS_D.
Probab=52.71 E-value=59 Score=32.85 Aligned_cols=103 Identities=17% Similarity=0.255 Sum_probs=56.7
Q ss_pred CCHHHHHHHHHHHHHHHhCcccCCC-cCCCCceEEEeec-CCCCCCCCcccccccccccceeecccCCCCCChHHHHHHH
Q 013386 309 GSDEWYEALDQHFKWLLQYRISPFF-CRWGESMRVLTYT-CPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYV 386 (444)
Q Consensus 309 gs~~~f~~L~~~~~~~~~~riS~~f-~~Wg~~mrv~~~~-~~w~~d~~~~~~y~~d~~~~~y~vp~~~~~~g~~~~~~~L 386 (444)
..+++|+.|++-++.+.+++|-+.. .=||.. |+ +.|+.....+ +.+..+.|+
T Consensus 82 ~N~~YF~~~d~~i~~a~~~Gi~~~lv~~wg~~-----~~~~~Wg~~~~~m---------------------~~e~~~~Y~ 135 (289)
T PF13204_consen 82 PNPAYFDHLDRRIEKANELGIEAALVPFWGCP-----YVPGTWGFGPNIM---------------------PPENAERYG 135 (289)
T ss_dssp ----HHHHHHHHHHHHHHTT-EEEEESS-HHH-----HH-------TTSS----------------------HHHHHHHH
T ss_pred CCHHHHHHHHHHHHHHHHCCCeEEEEEEECCc-----cccccccccccCC---------------------CHHHHHHHH
Confidence 3588999999999999999988642 223332 11 2344432222 345788899
Q ss_pred HHHHHHHHhcC--chhhhhhhhcCCCCCcccHHHHHHHHHHHHhhCCCCcEEEEEeec
Q 013386 387 RKEIELLRTKA--HWKKAYFYLWDEPLNMEHYSSVRNMASELHAYAPDARVLTTYYCG 442 (444)
Q Consensus 387 ~~~~~~Lr~kG--w~~k~yfyl~DEP~~~e~~~~~r~a~~~ir~~~Pd~ril~t~~~~ 442 (444)
+=+++++++.. ||..+= |.....++-+.++++++.|++.+|.- .+|-=.||
T Consensus 136 ~yv~~Ry~~~~NviW~l~g----d~~~~~~~~~~w~~~~~~i~~~dp~~-L~T~H~~~ 188 (289)
T PF13204_consen 136 RYVVARYGAYPNVIWILGG----DYFDTEKTRADWDAMARGIKENDPYQ-LITIHPCG 188 (289)
T ss_dssp HHHHHHHTT-SSEEEEEES----SS--TTSSHHHHHHHHHHHHHH--SS--EEEEE-B
T ss_pred HHHHHHHhcCCCCEEEecC----ccCCCCcCHHHHHHHHHHHHhhCCCC-cEEEeCCC
Confidence 99999999984 232211 11011255668999999999999977 55555554
No 17
>smart00737 ML Domain involved in innate immunity and lipid metabolism. ML (MD-2-related lipid-recognition) is a novel domain identified in MD-1, MD-2, GM2A, Npc2 and multiple proteins of unknown function in plants, animals and fungi. These single-domain proteins were predicted to form a beta-rich fold containing multiple strands, and to mediate diverse biological functions through interacting with specific lipids.
Probab=46.87 E-value=37 Score=29.03 Aligned_cols=34 Identities=29% Similarity=0.357 Sum_probs=27.6
Q ss_pred eeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 013386 153 ISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (444)
Q Consensus 153 v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~ 186 (444)
=++.+|+..-.=.++.||...++|.|+++++++.
T Consensus 72 CPl~~G~~~~~~~~~~v~~~~P~~~~~v~~~l~d 105 (118)
T smart00737 72 CPIEKGETVNYTNSLTVPGIFPPGKYTVKWELTD 105 (118)
T ss_pred CCCCCCeeEEEEEeeEccccCCCeEEEEEEEEEc
Confidence 3456787655556779999999999999999986
No 18
>PF13304 AAA_21: AAA domain; PDB: 3QKS_B 1US8_B 1F2U_B 1F2T_B 3QKT_A 1II8_B 3QKR_B 3QKU_A.
Probab=37.07 E-value=38 Score=30.16 Aligned_cols=38 Identities=29% Similarity=0.276 Sum_probs=32.2
Q ss_pred hhhhcCCCCCcccHHHHHHHHHHHHhhCC-CCcEEEEEe
Q 013386 403 YFYLWDEPLNMEHYSSVRNMASELHAYAP-DARVLTTYY 440 (444)
Q Consensus 403 yfyl~DEP~~~e~~~~~r~a~~~ir~~~P-d~ril~t~~ 440 (444)
.+.+.|||..-=|.+..+..++.+++..+ +.+++.|+-
T Consensus 259 ~illiDEpE~~LHp~~q~~l~~~l~~~~~~~~QviitTH 297 (303)
T PF13304_consen 259 SILLIDEPENHLHPSWQRKLIELLKELSKKNIQVIITTH 297 (303)
T ss_dssp SEEEEESSSTTSSHHHHHHHHHHHHHTGGGSSEEEEEES
T ss_pred eEEEecCCcCCCCHHHHHHHHHHHHhhCccCCEEEEeCc
Confidence 44569999987788899999999999988 899998874
No 19
>PF09099 Qn_am_d_aIII: Quinohemoprotein amine dehydrogenase, alpha subunit domain III; InterPro: IPR015183 This domain is predominantly found in the prokaryotic protein quinohemoprotein amine dehydrogenase, adopting an immunoglobulin-like beta-sandwich fold, with seven strands arranged into two beta sheets; the fold is possibly related to the immunoglobulin and/or fibronectin type III superfamilies. The precise function of this domain has not, as yet, been defined []. ; PDB: 1JMZ_A 1JMX_A 1PBY_A 1JJU_A.
Probab=28.19 E-value=59 Score=27.33 Aligned_cols=22 Identities=23% Similarity=0.353 Sum_probs=19.0
Q ss_pred eEEEEEEEcCCCCCCceeEEEE
Q 013386 161 TAVWVSIDAPYAQPPGLYEGEI 182 (444)
Q Consensus 161 q~lWI~V~VP~~a~pG~Y~GtV 182 (444)
-.++++|.+.++++||.|+..+
T Consensus 48 ~~v~v~V~~aa~a~~G~~~v~v 69 (81)
T PF09099_consen 48 DEVVVRVKAAADAAPGIRTVRV 69 (81)
T ss_dssp TCEEEEEEEECTSSSEEEEEEE
T ss_pred CEEEEEEEEcCCCCCccEEEEe
Confidence 3589999999999999997654
No 20
>PF00868 Transglut_N: Transglutaminase family; InterPro: IPR001102 Synonym(s): Protein-glutamine gamma-glutamyltransferase, Fibrinoligase, TGase Protein-glutamine gamma-glutamyltransferases (2.3.2.13 from EC) (TGase) are calcium-dependent enzymes that catalyse the cross-linking of proteins by promoting the formation of isopeptide bonds between the gamma-carboxyl group of a glutamine in one polypeptide chain and the epsilon-amino group of a lysine in a second polypeptide chain. TGases also catalyse the conjugation of polyamines to proteins [, ]. Transglutaminases are widely distributed in various organs, tissues and body fluids. The best known transglutaminase is blood coagulation factor XIII, a plasma tetrameric protein composed of two catalytic A subunits and two non-catalytic B subunits. Factor XIII is responsible for cross-linking fibrin chains, thus stabilising the fibrin clot. There are commonly three domains: N-terminal, middle (IPR013808 from INTERPRO) and C-terminal (IPR013807 from INTERPRO). This entry represents the N-terminal domain found in transglutaminases.; GO: 0018149 peptide cross-linking; PDB: 1L9N_B 1NUF_A 1NUD_A 1NUG_B 1L9M_A 1KV3_C 3S3S_A 2Q3Z_A 3LY6_A 3S3P_A ....
Probab=28.11 E-value=68 Score=28.29 Aligned_cols=27 Identities=26% Similarity=0.405 Sum_probs=19.1
Q ss_pred CeeEEEEEEEcCCCCCCceeEEEEEEE
Q 013386 159 ETTAVWVSIDAPYAQPPGLYEGEIIIT 185 (444)
Q Consensus 159 ~~q~lWI~V~VP~~a~pG~Y~GtVtVt 185 (444)
+...+=|.|.+|++|+-|.|+-.|.++
T Consensus 91 ~~~~~tv~V~spa~A~VG~y~l~v~~~ 117 (118)
T PF00868_consen 91 DGNSVTVSVTSPANAPVGRYKLSVETK 117 (118)
T ss_dssp ETTEEEEEEE--TTS--EEEEEEEEEE
T ss_pred CCCEEEEEEECCCCCceEEEEEEEEEe
Confidence 334577899999999999999999886
No 21
>PF12891 Glyco_hydro_44: Glycoside hydrolase family 44; InterPro: IPR024745 This is a family of putative bacterial glycoside hydrolases.; PDB: 3IK2_A 3ZQ9_A 2YJQ_B 2YKK_A 2YIH_A 2EEX_A 2EQD_A 2E0P_A 2E4T_A 2EO7_A ....
Probab=27.61 E-value=95 Score=31.18 Aligned_cols=29 Identities=28% Similarity=0.311 Sum_probs=21.9
Q ss_pred ccHHHHHHHHHHHHhhCCCCcEEEEEeec
Q 013386 414 EHYSSVRNMASELHAYAPDARVLTTYYCG 442 (444)
Q Consensus 414 e~~~~~r~a~~~ir~~~Pd~ril~t~~~~ 442 (444)
|-.+.+-+.++.||+.+|+++|+-.--||
T Consensus 154 El~~r~i~~AkaiK~~DP~a~v~GP~~wg 182 (239)
T PF12891_consen 154 ELRDRSIEYAKAIKAADPDAKVFGPVEWG 182 (239)
T ss_dssp HHHHHHHHHHHHHHHH-TTSEEEEEEE-S
T ss_pred HHHHHHHHHHHHHHhhCCCCeEeechhhc
Confidence 44456778899999999999999776666
No 22
>PLN02475 5-methyltetrahydropteroyltriglutamate--homocysteine methyltransferase
Probab=25.61 E-value=88 Score=36.26 Aligned_cols=60 Identities=18% Similarity=0.323 Sum_probs=38.1
Q ss_pred HHHHHHHHH---HHHHHHhcCchhhhhhhhcCCCCCc-----ccHHHHHHHHHHHHhhCCCCcEE-EEEeecc
Q 013386 380 DGAKDYVRK---EIELLRTKAHWKKAYFYLWDEPLNM-----EHYSSVRNMASELHAYAPDARVL-TTYYCGI 443 (444)
Q Consensus 380 ~~~~~~L~~---~~~~Lr~kGw~~k~yfyl~DEP~~~-----e~~~~~r~a~~~ir~~~Pd~ril-~t~~~~~ 443 (444)
+.+.+.++. +++.|.+.| -.++- .|||.-. +..+..+.+.+.|.+..|+.+|+ +|||+++
T Consensus 179 ~ll~~L~~~y~~~l~~L~~~G---v~~IQ-iDEP~L~~d~~~~~~~~~~~ay~~l~~~~~~~~i~l~TyFg~~ 247 (766)
T PLN02475 179 SLLDKILPVYKEVIAELKAAG---ASWIQ-FDEPALVMDLESHKLQAFKTAYAELESTLSGLNVLVETYFADV 247 (766)
T ss_pred HHHHHHHHHHHHHHHHHHHCC---CCEEE-EeCchhhcCCCHHHHHHHHHHHHHHHhccCCCeEEEEccCCCC
Confidence 355555554 444555554 33455 7998642 24567788888888888888886 6666653
No 23
>PF02228 Gag_p19: Major core protein p19; InterPro: IPR003139 Retroviral matrix proteins (or major core proteins) are components of envelope-associated capsids, which line the inner surface of virus envelopes and are associated with viral membranes []. Matrix proteins are produced as part of Gag precursor polyproteins. During viral maturation, the Gag polyprotein is cleaved into major structural proteins by the viral protease, yielding the matrix (MA), capsid (CA), nucleocapsid (NC), and some smaller peptides. Gag-derived proteins govern the entire assembly and release of the virus particles, with matrix proteins playing key roles in Gag stability, capsid assembly, transport and budding. Although matrix proteins from different retroviruses appear to perform similar functions and can have similar structural folds, their primary sequences can be very different. This entry represents matrix proteins from delta-retroviruses such as Human T-lymphotropic virus 1 and Human T-cell leukemia virus 2 (HTLV-2), both members of the human oncovirus subclass of retroviruses [, ].; GO: 0005198 structural molecule activity, 0019013 viral nucleocapsid; PDB: 1JVR_A.
Probab=25.12 E-value=54 Score=27.84 Aligned_cols=45 Identities=20% Similarity=0.268 Sum_probs=27.2
Q ss_pred CCCCCCCCcccCCC----hhhHhhhcCCCCC-CHHHHHHHHHHHHHHHhC
Q 013386 283 LPATPSLPAVIGIS----DTVIEDRFGVRHG-SDEWYEALDQHFKWLLQY 327 (444)
Q Consensus 283 LP~tp~l~~~~g~~----~~~i~~~~~v~~g-s~~~f~~L~~~~~~~~~~ 327 (444)
++..|.-..--|++ -.-.|..|.++.| |++.|..|+++.+|.++-
T Consensus 7 ~~~sPip~~PrGls~hhWLNflQaAyRL~PgPS~~DF~qLr~flk~alkT 56 (92)
T PF02228_consen 7 RSASPIPKPPRGLSTHHWLNFLQAAYRLQPGPSSFDFHQLRNFLKLALKT 56 (92)
T ss_dssp SSS--S--SS-SSTHHHHHHHHHHHHHSS---STTTHHHHHHHHHHHHT-
T ss_pred CCCCCCCCCCCCcCHHHHHHHHHHHHhcCCCCCcccHHHHHHHHHHHHcC
Confidence 34444444444663 3445667888888 888999999999999864
No 24
>PF14734 DUF4469: Domain of unknown function (DUF4469) with IG-like fold
Probab=25.03 E-value=74 Score=27.76 Aligned_cols=23 Identities=17% Similarity=0.110 Sum_probs=20.3
Q ss_pred EEEEEcCCCCCCceeEEEEEEEe
Q 013386 164 WVSIDAPYAQPPGLYEGEIIITS 186 (444)
Q Consensus 164 WI~V~VP~~a~pG~Y~GtVtVt~ 186 (444)
=+.+.||++-++|.|+.+|+=+-
T Consensus 65 ~l~~~lPa~L~~G~Y~l~V~Tq~ 87 (102)
T PF14734_consen 65 RLIFILPADLAAGEYTLEVRTQY 87 (102)
T ss_pred EEEEECcCccCceEEEEEEEEEe
Confidence 37899999999999999998875
No 25
>PF09608 Alph_Pro_TM: Putative transmembrane protein (Alph_Pro_TM); InterPro: IPR019088 This entry consists of predicted transmembrane proteins of about 270 amino acids. They are found predominantly, though not exclusively, in alphaproteobacteria, generally only once in each genome.
Probab=23.92 E-value=1e+02 Score=30.67 Aligned_cols=34 Identities=24% Similarity=0.413 Sum_probs=27.9
Q ss_pred ceeeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 013386 151 CQISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (444)
Q Consensus 151 ~~v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~ 186 (444)
..+.+..+ +-..-+|.+|++.++|.|+.++-+-.
T Consensus 147 ~~V~~~~~--~lFra~i~LPanvp~G~Y~v~v~l~r 180 (236)
T PF09608_consen 147 GGVQFLEG--TLFRARIPLPANVPPGDYTVRVYLFR 180 (236)
T ss_pred CeEEEcCC--CeEEEEeEcCCCCCcceEEEEEEEEE
Confidence 44665444 47789999999999999999999986
No 26
>PF05205 COMPASS-Shg1: COMPASS (Complex proteins associated with Set1p) component shg1
Probab=21.99 E-value=82 Score=27.41 Aligned_cols=37 Identities=14% Similarity=0.192 Sum_probs=28.6
Q ss_pred HHHHHHHHhcCchhhhhhhhcCCCCCcccHHHHHHHH
Q 013386 387 RKEIELLRTKAHWKKAYFYLWDEPLNMEHYSSVRNMA 423 (444)
Q Consensus 387 ~~~~~~Lr~kGw~~k~yfyl~DEP~~~e~~~~~r~a~ 423 (444)
+++++++|++|+||+.-=-++++....+.|+.++...
T Consensus 1 ~~Lv~~fKk~G~FD~lRk~~l~~~~~~~~~~~l~~~v 37 (106)
T PF05205_consen 1 KQLVEEFKKQGHFDKLRKECLADFDTSPAYQNLRQRV 37 (106)
T ss_pred ChHHHHHHhCCChHHHHHHHHHhccccHHHHHHHHHH
Confidence 3689999999999988888887776666666665554
No 27
>PF08428 Rib: Rib/alpha-like repeat; InterPro: IPR012706 This entry represents a region of about 79 amino acids found tandemly repeated up to fourteen times within the proteins that contain it. The repeats lack cysteines and are highly conserved, even at the DNA level, within and between proteins []. Proteins containing these repeats include the Rib and alpha surface antigens of group B Streptococcus, Esp of Enterococcus faecalis (Streptococcus faecalis), and related proteins of Lactobacillus. Most members of this protein family also have the cell wall anchor motif, LPXTG, shared by many staphyloccal and streptococcal surface antigens. These repeats are thought to define protective epitopes and may play a role in generating phenotypic and genotypic variation [].
Probab=20.55 E-value=2.9e+02 Score=21.83 Aligned_cols=31 Identities=29% Similarity=0.489 Sum_probs=23.8
Q ss_pred eeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 013386 153 ISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS 186 (444)
Q Consensus 153 v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~ 186 (444)
-++++|+. --|.+ .|....||.|.++|+|+=
T Consensus 19 ~~lP~gt~-~~w~~--~pdt~~~G~~~~~V~Vty 49 (65)
T PF08428_consen 19 DNLPAGTT-YSWKD--KPDTSKPGTKTGKVKVTY 49 (65)
T ss_pred ccCCCCcc-eeecc--CCccccCccEEEEEEEEc
Confidence 34444444 46776 899999999999999995
No 28
>PRK01254 hypothetical protein; Provisional
Probab=20.42 E-value=1.7e+02 Score=33.70 Aligned_cols=71 Identities=17% Similarity=0.228 Sum_probs=51.4
Q ss_pred eeecccCCCCCChHHHHHHHHHHHHHHHhcCchhhhhhhhcCCCCC------cccHHHHHHHHHHHHhhCC-CCcEEEEE
Q 013386 367 AYAVPYSPVLSSNDGAKDYVRKEIELLRTKAHWKKAYFYLWDEPLN------MEHYSSVRNMASELHAYAP-DARVLTTY 439 (444)
Q Consensus 367 ~y~vp~~~~~~g~~~~~~~L~~~~~~Lr~kGw~~k~yfyl~DEP~~------~e~~~~~r~a~~~ir~~~P-d~ril~t~ 439 (444)
+|.+||+..+.. ++|++.++++ +=-|+++.+..|++|+-+. ...++.++++.+.+++..| +.-+.+++
T Consensus 489 ~SgiR~Dl~l~d----~elIeel~~~-hV~g~LkVppEH~Sd~VLk~M~Kp~~~~~e~F~e~f~rirk~~gk~q~Lipyf 563 (707)
T PRK01254 489 ASGVRYDLAVED----PRYVKELVTH-HVGGYLKIAPEHTEEGPLSKMMKPGMGSYDRFKELFDKYSKEAGKEQYLIPYF 563 (707)
T ss_pred EcCCCccccccC----HHHHHHHHHh-CCccccccccccCCHHHHHHhCCCCcccHHHHHHHHHHHHHHCCCCeEEEEeE
Confidence 467888774332 4577777775 6678899999998887322 3567899999999999998 56666665
Q ss_pred eec
Q 013386 440 YCG 442 (444)
Q Consensus 440 ~~~ 442 (444)
..|
T Consensus 564 IvG 566 (707)
T PRK01254 564 ISA 566 (707)
T ss_pred EEE
Confidence 655
No 29
>PF09153 DUF1938: Domain of unknown function (DUF1938); InterPro: IPR015236 This domain, which is predominantly found in the archaeal protein O6-alkylguanine-DNA alkyltransferase, adopts a secondary structure consisting of a three stranded antiparallel beta-sheet and three alpha helices. The exact function has not, as yet, been defined, though it has been postulated that this domain may confer thermostability to the protein []. ; GO: 0005737 cytoplasm; PDB: 1MGT_A.
Probab=20.20 E-value=82 Score=26.93 Aligned_cols=24 Identities=17% Similarity=0.350 Sum_probs=20.0
Q ss_pred CCCChHHHHHHHHHHHHHHHhcCc
Q 013386 375 VLSSNDGAKDYVRKEIELLRTKAH 398 (444)
Q Consensus 375 ~~~g~~~~~~~L~~~~~~Lr~kGw 398 (444)
+++|..++++-++++++||+..|-
T Consensus 30 slDg~efl~eri~~L~~~L~kRgv 53 (86)
T PF09153_consen 30 SLDGEEFLRERISRLIEFLKKRGV 53 (86)
T ss_dssp ESSHHHHHH-HHHHHHHHHHHTT-
T ss_pred EeccHHHHHHHHHHHHHHHHhcCc
Confidence 468889999999999999999975
Done!