Query         013386
Match_columns 444
No_of_seqs    131 out of 138
Neff          4.9 
Searched_HMMs 46136
Date          Fri Mar 29 03:19:59 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/013386.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/013386hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF10633 NPCBM_assoc:  NPCBM-as  93.5    0.14 3.1E-06   41.3   4.8   32  154-185    45-76  (78)
  2 PF00150 Cellulase:  Cellulase   90.1    0.63 1.4E-05   44.5   5.9  104  310-439    57-172 (281)
  3 PF15418 DUF4625:  Domain of un  89.9     2.4 5.2E-05   38.5   9.0   88   77-189    29-120 (132)
  4 PF01229 Glyco_hydro_39:  Glyco  89.8       1 2.2E-05   48.4   7.8  108  313-441    82-205 (486)
  5 COG1470 Predicted membrane pro  86.5     5.7 0.00012   43.2  10.6   38  150-187   324-361 (513)
  6 PF06030 DUF916:  Bacterial pro  79.7      44 0.00096   29.7  12.5  108   62-186     4-120 (121)
  7 PF14352 DUF4402:  Domain of un  74.3     3.6 7.8E-05   36.3   3.5   31  156-186    96-128 (130)
  8 PF13731 WxL:  WxL domain surfa  73.0      13 0.00028   35.7   7.3   79  106-185   105-210 (215)
  9 COG1470 Predicted membrane pro  71.8     5.5 0.00012   43.3   4.8   33  154-186   437-469 (513)
 10 smart00633 Glyco_10 Glycosyl h  69.7      20 0.00044   35.0   8.0   98  312-439    13-125 (254)
 11 cd00917 PG-PI_TP The phosphati  67.4     9.9 0.00021   33.4   4.7   33  153-186    77-109 (122)
 12 PF02221 E1_DerP2_DerF2:  ML do  64.2      12 0.00026   32.4   4.6   35  153-187    86-120 (134)
 13 PF01835 A2M_N:  MG2 domain;  I  62.8      33 0.00071   28.3   6.8   28  158-185    59-86  (99)
 14 PF10003 DUF2244:  Integral mem  60.2     7.5 0.00016   35.3   2.7   53  159-213    88-140 (140)
 15 PF06280 DUF1034:  Fn3-like dom  59.0     9.8 0.00021   32.6   3.1   37  151-187    62-101 (112)
 16 PF13204 DUF4038:  Protein of u  52.7      59  0.0013   32.8   7.9  103  309-442    82-188 (289)
 17 smart00737 ML Domain involved   46.9      37 0.00081   29.0   4.8   34  153-186    72-105 (118)
 18 PF13304 AAA_21:  AAA domain; P  37.1      38 0.00082   30.2   3.4   38  403-440   259-297 (303)
 19 PF09099 Qn_am_d_aIII:  Quinohe  28.2      59  0.0013   27.3   2.9   22  161-182    48-69  (81)
 20 PF00868 Transglut_N:  Transglu  28.1      68  0.0015   28.3   3.4   27  159-185    91-117 (118)
 21 PF12891 Glyco_hydro_44:  Glyco  27.6      95  0.0021   31.2   4.7   29  414-442   154-182 (239)
 22 PLN02475 5-methyltetrahydropte  25.6      88  0.0019   36.3   4.6   60  380-443   179-247 (766)
 23 PF02228 Gag_p19:  Major core p  25.1      54  0.0012   27.8   2.1   45  283-327     7-56  (92)
 24 PF14734 DUF4469:  Domain of un  25.0      74  0.0016   27.8   3.0   23  164-186    65-87  (102)
 25 PF09608 Alph_Pro_TM:  Putative  23.9   1E+02  0.0022   30.7   4.2   34  151-186   147-180 (236)
 26 PF05205 COMPASS-Shg1:  COMPASS  22.0      82  0.0018   27.4   2.7   37  387-423     1-37  (106)
 27 PF08428 Rib:  Rib/alpha-like r  20.5 2.9E+02  0.0063   21.8   5.4   31  153-186    19-49  (65)
 28 PRK01254 hypothetical protein;  20.4 1.7E+02  0.0037   33.7   5.4   71  367-442   489-566 (707)
 29 PF09153 DUF1938:  Domain of un  20.2      82  0.0018   26.9   2.2   24  375-398    30-53  (86)

No 1  
>PF10633 NPCBM_assoc:  NPCBM-associated, NEW3 domain of alpha-galactosidase;  InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=93.45  E-value=0.14  Score=41.29  Aligned_cols=32  Identities=31%  Similarity=0.543  Sum_probs=25.7

Q ss_pred             eeCCCCeeEEEEEEEcCCCCCCceeEEEEEEE
Q 013386          154 SLIPGETTAVWVSIDAPYAQPPGLYEGEIIIT  185 (444)
Q Consensus       154 ~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt  185 (444)
                      .|++|+.+.+=++|.+|++++||.|..+++++
T Consensus        45 ~l~pG~s~~~~~~V~vp~~a~~G~y~v~~~a~   76 (78)
T PF10633_consen   45 SLPPGESVTVTFTVTVPADAAPGTYTVTVTAR   76 (78)
T ss_dssp             -B-TTSEEEEEEEEEE-TT--SEEEEEEEEEE
T ss_pred             cCCCCCEEEEEEEEECCCCCCCceEEEEEEEE
Confidence            68899999999999999999999999999886


No 2  
>PF00150 Cellulase:  Cellulase (glycosyl hydrolase family 5);  InterPro: IPR001547 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. Glycoside hydrolase family 5 GH5 from CAZY comprises enzymes with several known activities; endoglucanase (3.2.1.4 from EC); beta-mannanase (3.2.1.78 from EC); exo-1,3-glucanase (3.2.1.58 from EC); endo-1,6-glucanase (3.2.1.75 from EC); xylanase (3.2.1.8 from EC); endoglycoceramidase (3.2.1.123 from EC). The microbial degradation of cellulose and xylans requires several types of enzymes. Fungi and bacteria produces a spectrum of cellulolytic enzymes (cellulases) and xylanases which, on the basis of sequence similarities, can be classified into families. One of these families is known as the cellulase family A [] or as the glycosyl hydrolases family 5 []. One of the conserved regions in this family contains a conserved glutamic acid residue which is potentially involved [] in the catalytic mechanism.; GO: 0004553 hydrolase activity, hydrolyzing O-glycosyl compounds, 0005975 carbohydrate metabolic process; PDB: 3NDY_A 3NDZ_B 1LF1_A 1TVP_B 1TVN_A 3AYR_A 3AYS_A 1QI0_A 1W3K_A 1OCQ_A ....
Probab=90.14  E-value=0.63  Score=44.54  Aligned_cols=104  Identities=13%  Similarity=0.129  Sum_probs=66.6

Q ss_pred             CHHHHHHHHHHHHHHHhCcccCCCcCCCCceEEEeecCCCCCCCCcccccccccccceeecccCCCCCChHHHHHHHHHH
Q 013386          310 SDEWYEALDQHFKWLLQYRISPFFCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYVRKE  389 (444)
Q Consensus       310 s~~~f~~L~~~~~~~~~~riS~~f~~Wg~~mrv~~~~~~w~~d~~~~~~y~~d~~~~~y~vp~~~~~~g~~~~~~~L~~~  389 (444)
                      .+..++.|++.+++|.++.|.....-.+.        ..|..+....         .       ....-.+.++++++.+
T Consensus        57 ~~~~~~~ld~~v~~a~~~gi~vild~h~~--------~~w~~~~~~~---------~-------~~~~~~~~~~~~~~~l  112 (281)
T PF00150_consen   57 DETYLARLDRIVDAAQAYGIYVILDLHNA--------PGWANGGDGY---------G-------NNDTAQAWFKSFWRAL  112 (281)
T ss_dssp             THHHHHHHHHHHHHHHHTT-EEEEEEEES--------TTCSSSTSTT---------T-------THHHHHHHHHHHHHHH
T ss_pred             cHHHHHHHHHHHHHHHhCCCeEEEEeccC--------cccccccccc---------c-------cchhhHHHHHhhhhhh
Confidence            45788999999999999988754321111        1231111100         0       0000122466688889


Q ss_pred             HHHHHhcCchhhhhhhhcCCCCCccc------------HHHHHHHHHHHHhhCCCCcEEEEE
Q 013386          390 IELLRTKAHWKKAYFYLWDEPLNMEH------------YSSVRNMASELHAYAPDARVLTTY  439 (444)
Q Consensus       390 ~~~Lr~kGw~~k~yfyl~DEP~~~e~------------~~~~r~a~~~ir~~~Pd~ril~t~  439 (444)
                      ++++|..  -....|=|+.||.....            .+.++++++.||+..|+..|+...
T Consensus       113 a~~y~~~--~~v~~~el~NEP~~~~~~~~w~~~~~~~~~~~~~~~~~~Ir~~~~~~~i~~~~  172 (281)
T PF00150_consen  113 AKRYKDN--PPVVGWELWNEPNGGNDDANWNAQNPADWQDWYQRAIDAIRAADPNHLIIVGG  172 (281)
T ss_dssp             HHHHTTT--TTTEEEESSSSGCSTTSTTTTSHHHTHHHHHHHHHHHHHHHHTTSSSEEEEEE
T ss_pred             ccccCCC--CcEEEEEecCCccccCCccccccccchhhhhHHHHHHHHHHhcCCcceeecCC
Confidence            9998743  34678889999996312            257899999999999998888765


No 3  
>PF15418 DUF4625:  Domain of unknown function (DUF4625)
Probab=89.85  E-value=2.4  Score=38.46  Aligned_cols=88  Identities=19%  Similarity=0.252  Sum_probs=53.2

Q ss_pred             EEEeecCceeEEEEEEccCCCcCCCCCCCceEEEeec-c--ccCCCCcccccCceEEEEeeecCCCCcccccCCCCccee
Q 013386           77 NLLAARNERESVQIALRPKVSWSSSSTAGVVQVQCSD-L--CSASGDRLVVGQSLMLRRVVPMLGVPDALVPLDLPVCQI  153 (444)
Q Consensus        77 ~LsAaRGE~vSfQlvl~s~~~~~~~~~~~~V~Vs~sd-L--~s~~g~~~i~~~~I~lr~V~yVlGyPD~LvP~d~~~~~v  153 (444)
                      .-.+-||+.+.|.+-+...      ..++.++|++-. +  .+-++   ..++.               -.|... .+.+
T Consensus        29 ~~~~~~G~~ihfe~~i~d~------~~i~si~VeIH~nfd~H~h~~---~~~~~---------------~~~~~~-~~~~   83 (132)
T PF15418_consen   29 CKVATRGDDIHFEADISDN------SAIKSIKVEIHNNFDHHTHST---EAGEC---------------EKPWVF-EQDY   83 (132)
T ss_pred             CeEEecCCcEEEEEEEEcc------cceeEEEEEEecCcCcccccc---ccccc---------------ccCcEE-EEEE
Confidence            4567899999999999863      567788877721 1  11010   01000               011110 0112


Q ss_pred             eeCCC-CeeEEEEEEEcCCCCCCceeEEEEEEEeccC
Q 013386          154 SLIPG-ETTAVWVSIDAPYAQPPGLYEGEIIITSKAD  189 (444)
Q Consensus       154 ~l~ag-~~q~lWI~V~VP~~a~pG~Y~GtVtVt~~~~  189 (444)
                      .+..| .+.-+=..|.||++++||.|.-.|+|+.+++
T Consensus        84 ~~~~g~~~~~~h~~i~IPa~a~~G~YH~~i~VtD~~G  120 (132)
T PF15418_consen   84 DIYGGKKNYDFHEHIDIPADAPAGDYHFMITVTDAAG  120 (132)
T ss_pred             cccCCcccEeEEEeeeCCCCCCCcceEEEEEEEECCC
Confidence            22222 3456678999999999999999999997544


No 4  
>PF01229 Glyco_hydro_39:  Glycosyl hydrolases family 39;  InterPro: IPR000514 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. Glycoside hydrolase family 39 GH39 from CAZY comprises enzymes with several known activities; alpha-L-iduronidase (3.2.1.76 from EC); beta-xylosidase (3.2.1.37 from EC). The most highly conserved regions in these enzymes are located in their N-terminal sections. These contain a glutamic acid residue which, on the basis of similarities with other families of glycosyl hydrolases [], probably acts as the proton donor in their catalytic mechanism.; GO: 0004553 hydrolase activity, hydrolyzing O-glycosyl compounds, 0005975 carbohydrate metabolic process; PDB: 2BS9_D 2BFG_E 1W91_B 1UHV_D 1PX8_A.
Probab=89.78  E-value=1  Score=48.44  Aligned_cols=108  Identities=20%  Similarity=0.279  Sum_probs=65.4

Q ss_pred             HHHHHHHHHHHHHhCcccCC----CcCCCCceEEEeecCCCCCCCCcccccccccccceeecccCCCCCChHHHHHHHHH
Q 013386          313 WYEALDQHFKWLLQYRISPF----FCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYVRK  388 (444)
Q Consensus       313 ~f~~L~~~~~~~~~~riS~~----f~~Wg~~mrv~~~~~~w~~d~~~~~~y~~d~~~~~y~vp~~~~~~g~~~~~~~L~~  388 (444)
                      .|..||+-++.+++.+|.|+    |.+=+-....   ...|.|              .....|    ....+.+.+++++
T Consensus        82 nf~~lD~i~D~l~~~g~~P~vel~f~p~~~~~~~---~~~~~~--------------~~~~~p----p~~~~~W~~lv~~  140 (486)
T PF01229_consen   82 NFTYLDQILDFLLENGLKPFVELGFMPMALASGY---QTVFWY--------------KGNISP----PKDYEKWRDLVRA  140 (486)
T ss_dssp             --HHHHHHHHHHHHCT-EEEEEE-SB-GGGBSS-----EETTT--------------TEE-S-----BS-HHHHHHHHHH
T ss_pred             ChHHHHHHHHHHHHcCCEEEEEEEechhhhcCCC---Cccccc--------------cCCcCC----cccHHHHHHHHHH
Confidence            48999999999999999994    5431100000   000000              000011    1234589999999


Q ss_pred             HHHHHHhc-Cc--hhhhhhhhcCCCCCc---------ccHHHHHHHHHHHHhhCCCCcEEEEEee
Q 013386          389 EIELLRTK-AH--WKKAYFYLWDEPLNM---------EHYSSVRNMASELHAYAPDARVLTTYYC  441 (444)
Q Consensus       389 ~~~~Lr~k-Gw--~~k~yfyl~DEP~~~---------e~~~~~r~a~~~ir~~~Pd~ril~t~~~  441 (444)
                      +++|+..+ |.  .++.||=+|.||...         +=++.|+++++.||+++|++||--...|
T Consensus       141 ~~~h~~~RYG~~ev~~W~fEiWNEPd~~~f~~~~~~~ey~~ly~~~~~~iK~~~p~~~vGGp~~~  205 (486)
T PF01229_consen  141 FARHYIDRYGIEEVSTWYFEIWNEPDLKDFWWDGTPEEYFELYDATARAIKAVDPELKVGGPAFA  205 (486)
T ss_dssp             HHHHHHHHHHHHHHTTSEEEESS-TTSTTTSGGG-HHHHHHHHHHHHHHHHHH-TTSEEEEEEEE
T ss_pred             HHHHHHhhcCCccccceeEEeCcCCCcccccCCCCHHHHHHHHHHHHHHHHHhCCCCcccCcccc
Confidence            99999764 32  445578789999751         2245788999999999999998766555


No 5  
>COG1470 Predicted membrane protein [Function unknown]
Probab=86.51  E-value=5.7  Score=43.18  Aligned_cols=38  Identities=29%  Similarity=0.464  Sum_probs=35.5

Q ss_pred             cceeeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEec
Q 013386          150 VCQISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITSK  187 (444)
Q Consensus       150 ~~~v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~~  187 (444)
                      ...+.|.||+...+-+.|+.|++|.||.|..+|+++.+
T Consensus       324 vt~vkL~~gE~kdvtleV~ps~na~pG~Ynv~I~A~s~  361 (513)
T COG1470         324 VTSVKLKPGEEKDVTLEVYPSLNATPGTYNVTITASSS  361 (513)
T ss_pred             EEEEEecCCCceEEEEEEecCCCCCCCceeEEEEEecc
Confidence            56789999999999999999999999999999999863


No 6  
>PF06030 DUF916:  Bacterial protein of unknown function (DUF916);  InterPro: IPR010317 This family consists of putative cell surface proteins, from Firmicutes, of unknown function. 
Probab=79.68  E-value=44  Score=29.70  Aligned_cols=108  Identities=16%  Similarity=0.222  Sum_probs=67.8

Q ss_pred             cccCCCCCC-CCCCceEEEeecCceeEEEEEEccCCCcCCCCCCCceEEEeeccc-cCCCCcccccCceEEEEeeecC--
Q 013386           62 ANVGPQEMP-RPLEPINLLAARNERESVQIALRPKVSWSSSSTAGVVQVQCSDLC-SASGDRLVVGQSLMLRRVVPML--  137 (444)
Q Consensus        62 ~KVfpde~P-~~~~~~~LsAaRGE~vSfQlvl~s~~~~~~~~~~~~V~Vs~sdL~-s~~g~~~i~~~~I~lr~V~yVl--  137 (444)
                      .-|.|+..- .....+.|...-|+...+|+.+...     ++....|++++.+=. +.+|            .+.|..  
T Consensus         4 ~p~~p~~Q~~~~~~YFdL~~~P~q~~~l~v~i~N~-----s~~~~tv~v~~~~A~Tn~nG------------~I~Y~~~~   66 (121)
T PF06030_consen    4 TPVLPENQIDKNVSYFDLKVKPGQKQTLEVRITNN-----SDKEITVKVSANTATTNDNG------------VIDYSQNN   66 (121)
T ss_pred             eecCCccccCCCCCeEEEEeCCCCEEEEEEEEEeC-----CCCCEEEEEEEeeeEecCCE------------EEEECCCC
Confidence            345666553 2357899999999999999999762     122223443333221 2222            112211  


Q ss_pred             -CC-CcccccCC---CCcceeeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 013386          138 -GV-PDALVPLD---LPVCQISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS  186 (444)
Q Consensus       138 -Gy-PD~LvP~d---~~~~~v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~  186 (444)
                       .. +++-.++.   .....+.|+|++.+-+=++|.+|+..-.|..-|-|.|+.
T Consensus        67 ~~~d~sl~~~~~~~v~~~~~Vtl~~~~sk~V~~~i~~P~~~f~G~ilGGi~~~e  120 (121)
T PF06030_consen   67 PKKDKSLKYPFSDLVKIPKEVTLPPNESKTVTFTIKMPKKAFDGIILGGIYFSE  120 (121)
T ss_pred             cccCcccCcchHHhccCCcEEEECCCCEEEEEEEEEcCCCCcCCEEEeeEEEEe
Confidence             01 11111221   012349999999999999999999999999999999985


No 7  
>PF14352 DUF4402:  Domain of unknown function (DUF4402)
Probab=74.34  E-value=3.6  Score=36.30  Aligned_cols=31  Identities=26%  Similarity=0.480  Sum_probs=25.0

Q ss_pred             CCCCeeEEEE--EEEcCCCCCCceeEEEEEEEe
Q 013386          156 IPGETTAVWV--SIDAPYAQPPGLYEGEIIITS  186 (444)
Q Consensus       156 ~ag~~q~lWI--~V~VP~~a~pG~Y~GtVtVt~  186 (444)
                      ..+....++|  ++.|++++++|.|+|+++|+.
T Consensus        96 ~~~g~~~~~VGGtL~v~~~~~~G~YsGt~~VtV  128 (130)
T PF14352_consen   96 DTGGSATFNVGGTLNVPANQAAGTYSGTFTVTV  128 (130)
T ss_pred             cCCCcEEEEEEEEEEcCCCCCCeEEEEEEEEEE
Confidence            3444556666  589999999999999999986


No 8  
>PF13731 WxL:  WxL domain surface cell wall-binding
Probab=73.04  E-value=13  Score=35.70  Aligned_cols=79  Identities=24%  Similarity=0.398  Sum_probs=48.0

Q ss_pred             ceEEEeeccccCCCCcccccCceEEEEeeec--CC---CCc------ccccCCCCcceeeeCCCCeeEEE----------
Q 013386          106 VVQVQCSDLCSASGDRLVVGQSLMLRRVVPM--LG---VPD------ALVPLDLPVCQISLIPGETTAVW----------  164 (444)
Q Consensus       106 ~V~Vs~sdL~s~~g~~~i~~~~I~lr~V~yV--lG---yPD------~LvP~d~~~~~v~l~ag~~q~lW----------  164 (444)
                      .|+|+.++|++.+|.. +.+..|.+......  .+   -|-      .|.+.......+...+++.+..|          
T Consensus       105 ~L~v~~s~F~~~~~~~-L~ga~l~~~~~~~~~~~~~~~~~~~~~~~~~l~~~~~~~~v~~A~~~~g~G~~~~~~~~~~~~  183 (215)
T PF13731_consen  105 TLTVKLSPFTNADGDT-LPGATLTFNNGKVQSTANNTNTPTTVSSNITLTPGGQAQTVMSAAKGQGQGTWSYSFGDQDAT  183 (215)
T ss_pred             EEEEEeccccccCCcC-cccceEEecCceeEeecccccCCcccccceEeccCCcceeeEeecccccceEEEEEeCCcccc
Confidence            4788888999888654 55555655543322  11   111      12222211122333456666666          


Q ss_pred             ----EEEEcCCCCC--CceeEEEEEEE
Q 013386          165 ----VSIDAPYAQP--PGLYEGEIIIT  185 (444)
Q Consensus       165 ----I~V~VP~~a~--pG~Y~GtVtVt  185 (444)
                          |.+.||.++.  +|.|+++|+=+
T Consensus       184 ~~~~v~L~VP~~~~~~ag~Yt~tlTWt  210 (215)
T PF13731_consen  184 ADTGVSLSVPANTAKQAGTYTATLTWT  210 (215)
T ss_pred             cccceEEEeCCCCcccCCcEEEEEEEE
Confidence                8899999998  79999999876


No 9  
>COG1470 Predicted membrane protein [Function unknown]
Probab=71.84  E-value=5.5  Score=43.30  Aligned_cols=33  Identities=36%  Similarity=0.457  Sum_probs=30.5

Q ss_pred             eeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 013386          154 SLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS  186 (444)
Q Consensus       154 ~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~  186 (444)
                      .+.||+.-.+=++|.||++|.||.|+.+|+.++
T Consensus       437 sL~pge~~tV~ltI~vP~~a~aGdY~i~i~~ks  469 (513)
T COG1470         437 SLEPGESKTVSLTITVPEDAGAGDYRITITAKS  469 (513)
T ss_pred             ccCCCCcceEEEEEEcCCCCCCCcEEEEEEEee
Confidence            357899999999999999999999999999987


No 10 
>smart00633 Glyco_10 Glycosyl hydrolase family 10.
Probab=69.73  E-value=20  Score=35.04  Aligned_cols=98  Identities=10%  Similarity=0.122  Sum_probs=62.6

Q ss_pred             HHHHHHHHHHHHHHhCcccCC--CcCCCCceEEEeecCCCCCCCCcccccccccccceeecccCCCCCChHHHHHHHHHH
Q 013386          312 EWYEALDQHFKWLLQYRISPF--FCRWGESMRVLTYTCPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYVRKE  389 (444)
Q Consensus       312 ~~f~~L~~~~~~~~~~riS~~--f~~Wg~~mrv~~~~~~w~~d~~~~~~y~~d~~~~~y~vp~~~~~~g~~~~~~~L~~~  389 (444)
                      ..|+.+++.+++|.++.|.-.  .+-|+.+             .|   .|+.+..         + -.-..++++|+++.
T Consensus        13 ~n~~~~D~~~~~a~~~gi~v~gH~l~W~~~-------------~P---~W~~~~~---------~-~~~~~~~~~~i~~v   66 (254)
T smart00633       13 FNFSGADAIVNFAKENGIKVRGHTLVWHSQ-------------TP---DWVFNLS---------K-ETLLARLENHIKTV   66 (254)
T ss_pred             cChHHHHHHHHHHHHCCCEEEEEEEeeccc-------------CC---HhhhcCC---------H-HHHHHHHHHHHHHH
Confidence            458899999999999987632  1224332             22   2222100         0 00123567788888


Q ss_pred             HHHHHhcCchhhhhhhhcCCCCCcc-------cH------HHHHHHHHHHHhhCCCCcEEEEE
Q 013386          390 IELLRTKAHWKKAYFYLWDEPLNME-------HY------SSVRNMASELHAYAPDARVLTTY  439 (444)
Q Consensus       390 ~~~Lr~kGw~~k~yfyl~DEP~~~e-------~~------~~~r~a~~~ir~~~Pd~ril~t~  439 (444)
                      +.|++.+..    +.-++.||.+..       .+      +-++.+.+.+|+++|+++++..=
T Consensus        67 ~~ry~g~i~----~wdV~NE~~~~~~~~~~~~~w~~~~G~~~i~~af~~ar~~~P~a~l~~Nd  125 (254)
T smart00633       67 VGRYKGKIY----AWDVVNEALHDNGSGLRRSVWYQILGEDYIEKAFRYAREADPDAKLFYND  125 (254)
T ss_pred             HHHhCCcce----EEEEeeecccCCCcccccchHHHhcChHHHHHHHHHHHHhCCCCEEEEec
Confidence            888886622    345788887521       12      57889999999999999998863


No 11 
>cd00917 PG-PI_TP The phosphatidylinositol/phosphatidylglycerol transfer protein (PG/PI-TP) has been shown to bind phosphatidylglycerol and phosphatidylinositol, but the biological significance of this is still obscure. These proteins belong to the ML domain family.
Probab=67.42  E-value=9.9  Score=33.42  Aligned_cols=33  Identities=24%  Similarity=0.374  Sum_probs=29.0

Q ss_pred             eeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 013386          153 ISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS  186 (444)
Q Consensus       153 v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~  186 (444)
                      =++.+|+.. +=.++.||...++|.|+++.++..
T Consensus        77 CPi~~G~~~-~~~~~~ip~~~P~g~y~v~~~l~d  109 (122)
T cd00917          77 CPIEPGDKF-LTKLVDLPGEIPPGKYTVSARAYT  109 (122)
T ss_pred             CCcCCCcEE-EEEEeeCCCCCCCceEEEEEEEEC
Confidence            456788877 788899999999999999999986


No 12 
>PF02221 E1_DerP2_DerF2:  ML domain;  InterPro: IPR003172  The MD-2-related lipid-recognition (ML) domain is implicated in lipid recognition, particularly in the recognition of pathogen related products. It has an immunoglobulin-like beta-sandwich fold similar to that of E-set Ig domains. This domain is present in the following proteins:  Epididymal secretory protein E1 (also known as Niemann-Pick C2 protein), which is known to bind cholesterol. Niemann-Pick disease type C2 is a fatal hereditary disease characterised by accumulation of low-density lipoprotein-derived cholesterol in lysosomes [].  House-dust mite allergen proteins such as Der f 2 from Dermatophagoides farinae and Der p 2 from Dermatophagoides pteronyssinus [].  ; PDB: 2AG9_B 1G13_B 2AG2_B 2AG4_A 1TJJ_C 1PU5_C 1PUB_A 2AF9_A 3T6Q_D 3M7O_B ....
Probab=64.22  E-value=12  Score=32.39  Aligned_cols=35  Identities=26%  Similarity=0.350  Sum_probs=31.9

Q ss_pred             eeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEec
Q 013386          153 ISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITSK  187 (444)
Q Consensus       153 v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~~  187 (444)
                      =++.+|+..-..+++.||...++|.|+++++++..
T Consensus        86 CPi~~G~~~~~~~~~~i~~~~p~~~~~i~~~l~d~  120 (134)
T PF02221_consen   86 CPIKAGEYYTYTYTIPIPKIYPPGKYTIQWKLTDQ  120 (134)
T ss_dssp             STBTTTEEEEEEEEEEESTTSSSEEEEEEEEEEET
T ss_pred             CccCCCcEEEEEEEEEcccceeeEEEEEEEEEEeC
Confidence            46789999999999999999999999999999974


No 13 
>PF01835 A2M_N:  MG2 domain;  InterPro: IPR002890 The proteinase-binding alpha-macroglobulins (A2M) [] are large glycoproteins found in the plasma of vertebrates, in the hemolymph of some invertebrates and in reptilian and avian egg white. A2M-like proteins are able to inhibit all four classes of proteinases by a 'trapping' mechanism. They have a peptide stretch, called the 'bait region', which contains specific cleavage sites for different proteinases. When a proteinase cleaves the bait region, a conformational change is induced in the protein, thus trapping the proteinase. The entrapped enzyme remains active against low molecular weight substrates, whilst its activity toward larger substrates is greatly reduced, due to steric hindrance. Following cleavage in the bait region, a thiol ester bond, formed between the side chains of a cysteine and a glutamine, is cleaved and mediates the covalent binding of the A2M-like protein to the proteinase. This family includes the N-terminal region of the alpha-2-macroglobulin family. The inhibitor domains belong to MEROPS inhibitor family I39.; GO: 0004866 endopeptidase inhibitor activity; PDB: 2B39_B 3KLS_B 3PRX_C 3KM9_B 3PVM_C 3CU7_A 4E0S_A 4A5W_A 4ACQ_C 2P9R_B ....
Probab=62.85  E-value=33  Score=28.32  Aligned_cols=28  Identities=21%  Similarity=0.215  Sum_probs=18.5

Q ss_pred             CCeeEEEEEEEcCCCCCCceeEEEEEEE
Q 013386          158 GETTAVWVSIDAPYAQPPGLYEGEIIIT  185 (444)
Q Consensus       158 g~~q~lWI~V~VP~~a~pG~Y~GtVtVt  185 (444)
                      ...-.+-.++.+|++++.|.|+.++...
T Consensus        59 ~~~G~~~~~~~lp~~~~~G~y~i~~~~~   86 (99)
T PF01835_consen   59 NENGIFSGSFQLPDDAPLGTYTIRVKTD   86 (99)
T ss_dssp             TCTTEEEEEEE--SS---EEEEEEEEET
T ss_pred             CCCCEEEEEEECCCCCCCEeEEEEEEEc
Confidence            3444677889999999999999999885


No 14 
>PF10003 DUF2244:  Integral membrane protein (DUF2244);  InterPro: IPR019253  This entry consists of various bacterial putative membrane proteins with no known function. 
Probab=60.22  E-value=7.5  Score=35.25  Aligned_cols=53  Identities=23%  Similarity=0.324  Sum_probs=38.9

Q ss_pred             CeeEEEEEEEcCCCCCCceeEEEEEEEeccCccccccccccchhhhhhhhhhccc
Q 013386          159 ETTAVWVSIDAPYAQPPGLYEGEIIITSKADTELSSQCLGKGEKHRLFMELRNCL  213 (444)
Q Consensus       159 ~~q~lWI~V~VP~~a~pG~Y~GtVtVt~~~~ge~~~~~~~~~~~~~~~~~~~~~l  213 (444)
                      +..+.|+.|.+..+..+  ..-.|++++++..-.-+.-|+++||..|+.+|+..|
T Consensus        88 ~~~~~w~rv~~~~~~~~--~~~~l~L~~~g~~veiG~fL~~~eR~~la~~L~~aL  140 (140)
T PF10003_consen   88 EFNPYWVRVELEEDPGP--GPPRLTLRSRGREVEIGRFLNPEEREELARELRRAL  140 (140)
T ss_pred             EEcCCeEEEEEEcCCCC--CCcEEEEEECCEEEEEccCCCHHHHHHHHHHHHhhC
Confidence            45689999999998887  555666665434112234899999999999998754


No 15 
>PF06280 DUF1034:  Fn3-like domain (DUF1034);  InterPro: IPR010435 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This domain of unknown function is present in bacterial and plant peptidases belonging to MEROPS peptidase family S8 (subfamily S8A subtilisin, clan SB). It is C-terminal to and adjacent to the S8 peptidase domain and can be found in conjunction with the PA (Protease associated) domain (IPR003137 from INTERPRO) and additionally in Gram-positive bacteria with the surface protein anchor domain (IPR001899 from INTERPRO).; GO: 0004252 serine-type endopeptidase activity, 0005618 cell wall, 0016020 membrane; PDB: 3EIF_A 1XF1_B.
Probab=58.96  E-value=9.8  Score=32.62  Aligned_cols=37  Identities=27%  Similarity=0.416  Sum_probs=31.2

Q ss_pred             ceeeeCCCCeeEEEEEEEcCCCCCC---ceeEEEEEEEec
Q 013386          151 CQISLIPGETTAVWVSIDAPYAQPP---GLYEGEIIITSK  187 (444)
Q Consensus       151 ~~v~l~ag~~q~lWI~V~VP~~a~p---G~Y~GtVtVt~~  187 (444)
                      ..+.|+||+++-+=|++.+|++..+   ..|.|-|.++..
T Consensus        62 ~~vTV~ag~s~~v~vti~~p~~~~~~~~~~~eG~I~~~~~  101 (112)
T PF06280_consen   62 DTVTVPAGQSKTVTVTITPPSGLDASNGPFYEGFITFKSS  101 (112)
T ss_dssp             EEEEE-TTEEEEEEEEEE--GGGHHTT-EEEEEEEEEESS
T ss_pred             CeEEECCCCEEEEEEEEEehhcCCcccCCEEEEEEEEEcC
Confidence            5799999999999999999998887   899999999973


No 16 
>PF13204 DUF4038:  Protein of unknown function (DUF4038); PDB: 3KZS_D.
Probab=52.71  E-value=59  Score=32.85  Aligned_cols=103  Identities=17%  Similarity=0.255  Sum_probs=56.7

Q ss_pred             CCHHHHHHHHHHHHHHHhCcccCCC-cCCCCceEEEeec-CCCCCCCCcccccccccccceeecccCCCCCChHHHHHHH
Q 013386          309 GSDEWYEALDQHFKWLLQYRISPFF-CRWGESMRVLTYT-CPWPADHPKSDEYFSDPRLAAYAVPYSPVLSSNDGAKDYV  386 (444)
Q Consensus       309 gs~~~f~~L~~~~~~~~~~riS~~f-~~Wg~~mrv~~~~-~~w~~d~~~~~~y~~d~~~~~y~vp~~~~~~g~~~~~~~L  386 (444)
                      ..+++|+.|++-++.+.+++|-+.. .=||..     |+ +.|+.....+                     +.+..+.|+
T Consensus        82 ~N~~YF~~~d~~i~~a~~~Gi~~~lv~~wg~~-----~~~~~Wg~~~~~m---------------------~~e~~~~Y~  135 (289)
T PF13204_consen   82 PNPAYFDHLDRRIEKANELGIEAALVPFWGCP-----YVPGTWGFGPNIM---------------------PPENAERYG  135 (289)
T ss_dssp             ----HHHHHHHHHHHHHHTT-EEEEESS-HHH-----HH-------TTSS----------------------HHHHHHHH
T ss_pred             CCHHHHHHHHHHHHHHHHCCCeEEEEEEECCc-----cccccccccccCC---------------------CHHHHHHHH
Confidence            3588999999999999999988642 223332     11 2344432222                     345788899


Q ss_pred             HHHHHHHHhcC--chhhhhhhhcCCCCCcccHHHHHHHHHHHHhhCCCCcEEEEEeec
Q 013386          387 RKEIELLRTKA--HWKKAYFYLWDEPLNMEHYSSVRNMASELHAYAPDARVLTTYYCG  442 (444)
Q Consensus       387 ~~~~~~Lr~kG--w~~k~yfyl~DEP~~~e~~~~~r~a~~~ir~~~Pd~ril~t~~~~  442 (444)
                      +=+++++++..  ||..+=    |.....++-+.++++++.|++.+|.- .+|-=.||
T Consensus       136 ~yv~~Ry~~~~NviW~l~g----d~~~~~~~~~~w~~~~~~i~~~dp~~-L~T~H~~~  188 (289)
T PF13204_consen  136 RYVVARYGAYPNVIWILGG----DYFDTEKTRADWDAMARGIKENDPYQ-LITIHPCG  188 (289)
T ss_dssp             HHHHHHHTT-SSEEEEEES----SS--TTSSHHHHHHHHHHHHHH--SS--EEEEE-B
T ss_pred             HHHHHHHhcCCCCEEEecC----ccCCCCcCHHHHHHHHHHHHhhCCCC-cEEEeCCC
Confidence            99999999984  232211    11011255668999999999999977 55555554


No 17 
>smart00737 ML Domain involved in innate immunity and lipid metabolism. ML (MD-2-related lipid-recognition) is a novel domain identified in MD-1, MD-2, GM2A, Npc2 and multiple proteins of unknown function in plants, animals and fungi. These single-domain proteins were predicted to form a beta-rich fold containing multiple strands, and to mediate diverse biological functions through interacting with specific lipids.
Probab=46.87  E-value=37  Score=29.03  Aligned_cols=34  Identities=29%  Similarity=0.357  Sum_probs=27.6

Q ss_pred             eeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 013386          153 ISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS  186 (444)
Q Consensus       153 v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~  186 (444)
                      =++.+|+..-.=.++.||...++|.|+++++++.
T Consensus        72 CPl~~G~~~~~~~~~~v~~~~P~~~~~v~~~l~d  105 (118)
T smart00737       72 CPIEKGETVNYTNSLTVPGIFPPGKYTVKWELTD  105 (118)
T ss_pred             CCCCCCeeEEEEEeeEccccCCCeEEEEEEEEEc
Confidence            3456787655556779999999999999999986


No 18 
>PF13304 AAA_21:  AAA domain; PDB: 3QKS_B 1US8_B 1F2U_B 1F2T_B 3QKT_A 1II8_B 3QKR_B 3QKU_A.
Probab=37.07  E-value=38  Score=30.16  Aligned_cols=38  Identities=29%  Similarity=0.276  Sum_probs=32.2

Q ss_pred             hhhhcCCCCCcccHHHHHHHHHHHHhhCC-CCcEEEEEe
Q 013386          403 YFYLWDEPLNMEHYSSVRNMASELHAYAP-DARVLTTYY  440 (444)
Q Consensus       403 yfyl~DEP~~~e~~~~~r~a~~~ir~~~P-d~ril~t~~  440 (444)
                      .+.+.|||..-=|.+..+..++.+++..+ +.+++.|+-
T Consensus       259 ~illiDEpE~~LHp~~q~~l~~~l~~~~~~~~QviitTH  297 (303)
T PF13304_consen  259 SILLIDEPENHLHPSWQRKLIELLKELSKKNIQVIITTH  297 (303)
T ss_dssp             SEEEEESSSTTSSHHHHHHHHHHHHHTGGGSSEEEEEES
T ss_pred             eEEEecCCcCCCCHHHHHHHHHHHHhhCccCCEEEEeCc
Confidence            44569999987788899999999999988 899998874


No 19 
>PF09099 Qn_am_d_aIII:  Quinohemoprotein amine dehydrogenase, alpha subunit domain III;  InterPro: IPR015183 This domain is predominantly found in the prokaryotic protein quinohemoprotein amine dehydrogenase, adopting an immunoglobulin-like beta-sandwich fold, with seven strands arranged into two beta sheets; the fold is possibly related to the immunoglobulin and/or fibronectin type III superfamilies. The precise function of this domain has not, as yet, been defined []. ; PDB: 1JMZ_A 1JMX_A 1PBY_A 1JJU_A.
Probab=28.19  E-value=59  Score=27.33  Aligned_cols=22  Identities=23%  Similarity=0.353  Sum_probs=19.0

Q ss_pred             eEEEEEEEcCCCCCCceeEEEE
Q 013386          161 TAVWVSIDAPYAQPPGLYEGEI  182 (444)
Q Consensus       161 q~lWI~V~VP~~a~pG~Y~GtV  182 (444)
                      -.++++|.+.++++||.|+..+
T Consensus        48 ~~v~v~V~~aa~a~~G~~~v~v   69 (81)
T PF09099_consen   48 DEVVVRVKAAADAAPGIRTVRV   69 (81)
T ss_dssp             TCEEEEEEEECTSSSEEEEEEE
T ss_pred             CEEEEEEEEcCCCCCccEEEEe
Confidence            3589999999999999997654


No 20 
>PF00868 Transglut_N:  Transglutaminase family;  InterPro: IPR001102 Synonym(s): Protein-glutamine gamma-glutamyltransferase, Fibrinoligase, TGase Protein-glutamine gamma-glutamyltransferases (2.3.2.13 from EC) (TGase) are calcium-dependent enzymes that catalyse the cross-linking of proteins by promoting the formation of isopeptide bonds between the gamma-carboxyl group of a glutamine in one polypeptide chain and the epsilon-amino group of a lysine in a second polypeptide chain. TGases also catalyse the conjugation of polyamines to proteins [, ]. Transglutaminases are widely distributed in various organs, tissues and body fluids. The best known transglutaminase is blood coagulation factor XIII, a plasma tetrameric protein composed of two catalytic A subunits and two non-catalytic B subunits. Factor XIII is responsible for cross-linking fibrin chains, thus stabilising the fibrin clot. There are commonly three domains: N-terminal, middle (IPR013808 from INTERPRO) and C-terminal (IPR013807 from INTERPRO). This entry represents the N-terminal domain found in transglutaminases.; GO: 0018149 peptide cross-linking; PDB: 1L9N_B 1NUF_A 1NUD_A 1NUG_B 1L9M_A 1KV3_C 3S3S_A 2Q3Z_A 3LY6_A 3S3P_A ....
Probab=28.11  E-value=68  Score=28.29  Aligned_cols=27  Identities=26%  Similarity=0.405  Sum_probs=19.1

Q ss_pred             CeeEEEEEEEcCCCCCCceeEEEEEEE
Q 013386          159 ETTAVWVSIDAPYAQPPGLYEGEIIIT  185 (444)
Q Consensus       159 ~~q~lWI~V~VP~~a~pG~Y~GtVtVt  185 (444)
                      +...+=|.|.+|++|+-|.|+-.|.++
T Consensus        91 ~~~~~tv~V~spa~A~VG~y~l~v~~~  117 (118)
T PF00868_consen   91 DGNSVTVSVTSPANAPVGRYKLSVETK  117 (118)
T ss_dssp             ETTEEEEEEE--TTS--EEEEEEEEEE
T ss_pred             CCCEEEEEEECCCCCceEEEEEEEEEe
Confidence            334577899999999999999999886


No 21 
>PF12891 Glyco_hydro_44:  Glycoside hydrolase family 44;  InterPro: IPR024745 This is a family of putative bacterial glycoside hydrolases.; PDB: 3IK2_A 3ZQ9_A 2YJQ_B 2YKK_A 2YIH_A 2EEX_A 2EQD_A 2E0P_A 2E4T_A 2EO7_A ....
Probab=27.61  E-value=95  Score=31.18  Aligned_cols=29  Identities=28%  Similarity=0.311  Sum_probs=21.9

Q ss_pred             ccHHHHHHHHHHHHhhCCCCcEEEEEeec
Q 013386          414 EHYSSVRNMASELHAYAPDARVLTTYYCG  442 (444)
Q Consensus       414 e~~~~~r~a~~~ir~~~Pd~ril~t~~~~  442 (444)
                      |-.+.+-+.++.||+.+|+++|+-.--||
T Consensus       154 El~~r~i~~AkaiK~~DP~a~v~GP~~wg  182 (239)
T PF12891_consen  154 ELRDRSIEYAKAIKAADPDAKVFGPVEWG  182 (239)
T ss_dssp             HHHHHHHHHHHHHHHH-TTSEEEEEEE-S
T ss_pred             HHHHHHHHHHHHHHhhCCCCeEeechhhc
Confidence            44456778899999999999999776666


No 22 
>PLN02475 5-methyltetrahydropteroyltriglutamate--homocysteine methyltransferase
Probab=25.61  E-value=88  Score=36.26  Aligned_cols=60  Identities=18%  Similarity=0.323  Sum_probs=38.1

Q ss_pred             HHHHHHHHH---HHHHHHhcCchhhhhhhhcCCCCCc-----ccHHHHHHHHHHHHhhCCCCcEE-EEEeecc
Q 013386          380 DGAKDYVRK---EIELLRTKAHWKKAYFYLWDEPLNM-----EHYSSVRNMASELHAYAPDARVL-TTYYCGI  443 (444)
Q Consensus       380 ~~~~~~L~~---~~~~Lr~kGw~~k~yfyl~DEP~~~-----e~~~~~r~a~~~ir~~~Pd~ril-~t~~~~~  443 (444)
                      +.+.+.++.   +++.|.+.|   -.++- .|||.-.     +..+..+.+.+.|.+..|+.+|+ +|||+++
T Consensus       179 ~ll~~L~~~y~~~l~~L~~~G---v~~IQ-iDEP~L~~d~~~~~~~~~~~ay~~l~~~~~~~~i~l~TyFg~~  247 (766)
T PLN02475        179 SLLDKILPVYKEVIAELKAAG---ASWIQ-FDEPALVMDLESHKLQAFKTAYAELESTLSGLNVLVETYFADV  247 (766)
T ss_pred             HHHHHHHHHHHHHHHHHHHCC---CCEEE-EeCchhhcCCCHHHHHHHHHHHHHHHhccCCCeEEEEccCCCC
Confidence            355555554   444555554   33455 7998642     24567788888888888888886 6666653


No 23 
>PF02228 Gag_p19:  Major core protein p19;  InterPro: IPR003139 Retroviral matrix proteins (or major core proteins) are components of envelope-associated capsids, which line the inner surface of virus envelopes and are associated with viral membranes []. Matrix proteins are produced as part of Gag precursor polyproteins. During viral maturation, the Gag polyprotein is cleaved into major structural proteins by the viral protease, yielding the matrix (MA), capsid (CA), nucleocapsid (NC), and some smaller peptides. Gag-derived proteins govern the entire assembly and release of the virus particles, with matrix proteins playing key roles in Gag stability, capsid assembly, transport and budding. Although matrix proteins from different retroviruses appear to perform similar functions and can have similar structural folds, their primary sequences can be very different. This entry represents matrix proteins from delta-retroviruses such as Human T-lymphotropic virus 1 and Human T-cell leukemia virus 2 (HTLV-2), both members of the human oncovirus subclass of retroviruses [, ].; GO: 0005198 structural molecule activity, 0019013 viral nucleocapsid; PDB: 1JVR_A.
Probab=25.12  E-value=54  Score=27.84  Aligned_cols=45  Identities=20%  Similarity=0.268  Sum_probs=27.2

Q ss_pred             CCCCCCCCcccCCC----hhhHhhhcCCCCC-CHHHHHHHHHHHHHHHhC
Q 013386          283 LPATPSLPAVIGIS----DTVIEDRFGVRHG-SDEWYEALDQHFKWLLQY  327 (444)
Q Consensus       283 LP~tp~l~~~~g~~----~~~i~~~~~v~~g-s~~~f~~L~~~~~~~~~~  327 (444)
                      ++..|.-..--|++    -.-.|..|.++.| |++.|..|+++.+|.++-
T Consensus         7 ~~~sPip~~PrGls~hhWLNflQaAyRL~PgPS~~DF~qLr~flk~alkT   56 (92)
T PF02228_consen    7 RSASPIPKPPRGLSTHHWLNFLQAAYRLQPGPSSFDFHQLRNFLKLALKT   56 (92)
T ss_dssp             SSS--S--SS-SSTHHHHHHHHHHHHHSS---STTTHHHHHHHHHHHHT-
T ss_pred             CCCCCCCCCCCCcCHHHHHHHHHHHHhcCCCCCcccHHHHHHHHHHHHcC
Confidence            34444444444663    3445667888888 888999999999999864


No 24 
>PF14734 DUF4469:  Domain of unknown function (DUF4469) with IG-like fold
Probab=25.03  E-value=74  Score=27.76  Aligned_cols=23  Identities=17%  Similarity=0.110  Sum_probs=20.3

Q ss_pred             EEEEEcCCCCCCceeEEEEEEEe
Q 013386          164 WVSIDAPYAQPPGLYEGEIIITS  186 (444)
Q Consensus       164 WI~V~VP~~a~pG~Y~GtVtVt~  186 (444)
                      =+.+.||++-++|.|+.+|+=+-
T Consensus        65 ~l~~~lPa~L~~G~Y~l~V~Tq~   87 (102)
T PF14734_consen   65 RLIFILPADLAAGEYTLEVRTQY   87 (102)
T ss_pred             EEEEECcCccCceEEEEEEEEEe
Confidence            37899999999999999998875


No 25 
>PF09608 Alph_Pro_TM:  Putative transmembrane protein (Alph_Pro_TM);  InterPro: IPR019088  This entry consists of predicted transmembrane proteins of about 270 amino acids. They are found predominantly, though not exclusively, in alphaproteobacteria, generally only once in each genome. 
Probab=23.92  E-value=1e+02  Score=30.67  Aligned_cols=34  Identities=24%  Similarity=0.413  Sum_probs=27.9

Q ss_pred             ceeeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 013386          151 CQISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS  186 (444)
Q Consensus       151 ~~v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~  186 (444)
                      ..+.+..+  +-..-+|.+|++.++|.|+.++-+-.
T Consensus       147 ~~V~~~~~--~lFra~i~LPanvp~G~Y~v~v~l~r  180 (236)
T PF09608_consen  147 GGVQFLEG--TLFRARIPLPANVPPGDYTVRVYLFR  180 (236)
T ss_pred             CeEEEcCC--CeEEEEeEcCCCCCcceEEEEEEEEE
Confidence            44665444  47789999999999999999999986


No 26 
>PF05205 COMPASS-Shg1:  COMPASS (Complex proteins associated with Set1p) component shg1
Probab=21.99  E-value=82  Score=27.41  Aligned_cols=37  Identities=14%  Similarity=0.192  Sum_probs=28.6

Q ss_pred             HHHHHHHHhcCchhhhhhhhcCCCCCcccHHHHHHHH
Q 013386          387 RKEIELLRTKAHWKKAYFYLWDEPLNMEHYSSVRNMA  423 (444)
Q Consensus       387 ~~~~~~Lr~kGw~~k~yfyl~DEP~~~e~~~~~r~a~  423 (444)
                      +++++++|++|+||+.-=-++++....+.|+.++...
T Consensus         1 ~~Lv~~fKk~G~FD~lRk~~l~~~~~~~~~~~l~~~v   37 (106)
T PF05205_consen    1 KQLVEEFKKQGHFDKLRKECLADFDTSPAYQNLRQRV   37 (106)
T ss_pred             ChHHHHHHhCCChHHHHHHHHHhccccHHHHHHHHHH
Confidence            3689999999999988888887776666666665554


No 27 
>PF08428 Rib:  Rib/alpha-like repeat;  InterPro: IPR012706 This entry represents a region of about 79 amino acids found tandemly repeated up to fourteen times within the proteins that contain it. The repeats lack cysteines and are highly conserved, even at the DNA level, within and between proteins []. Proteins containing these repeats include the Rib and alpha surface antigens of group B Streptococcus, Esp of Enterococcus faecalis (Streptococcus faecalis), and related proteins of Lactobacillus. Most members of this protein family also have the cell wall anchor motif, LPXTG, shared by many staphyloccal and streptococcal surface antigens. These repeats are thought to define protective epitopes and may play a role in generating phenotypic and genotypic variation [].
Probab=20.55  E-value=2.9e+02  Score=21.83  Aligned_cols=31  Identities=29%  Similarity=0.489  Sum_probs=23.8

Q ss_pred             eeeCCCCeeEEEEEEEcCCCCCCceeEEEEEEEe
Q 013386          153 ISLIPGETTAVWVSIDAPYAQPPGLYEGEIIITS  186 (444)
Q Consensus       153 v~l~ag~~q~lWI~V~VP~~a~pG~Y~GtVtVt~  186 (444)
                      -++++|+. --|.+  .|....||.|.++|+|+=
T Consensus        19 ~~lP~gt~-~~w~~--~pdt~~~G~~~~~V~Vty   49 (65)
T PF08428_consen   19 DNLPAGTT-YSWKD--KPDTSKPGTKTGKVKVTY   49 (65)
T ss_pred             ccCCCCcc-eeecc--CCccccCccEEEEEEEEc
Confidence            34444444 46776  899999999999999995


No 28 
>PRK01254 hypothetical protein; Provisional
Probab=20.42  E-value=1.7e+02  Score=33.70  Aligned_cols=71  Identities=17%  Similarity=0.228  Sum_probs=51.4

Q ss_pred             eeecccCCCCCChHHHHHHHHHHHHHHHhcCchhhhhhhhcCCCCC------cccHHHHHHHHHHHHhhCC-CCcEEEEE
Q 013386          367 AYAVPYSPVLSSNDGAKDYVRKEIELLRTKAHWKKAYFYLWDEPLN------MEHYSSVRNMASELHAYAP-DARVLTTY  439 (444)
Q Consensus       367 ~y~vp~~~~~~g~~~~~~~L~~~~~~Lr~kGw~~k~yfyl~DEP~~------~e~~~~~r~a~~~ir~~~P-d~ril~t~  439 (444)
                      +|.+||+..+..    ++|++.++++ +=-|+++.+..|++|+-+.      ...++.++++.+.+++..| +.-+.+++
T Consensus       489 ~SgiR~Dl~l~d----~elIeel~~~-hV~g~LkVppEH~Sd~VLk~M~Kp~~~~~e~F~e~f~rirk~~gk~q~Lipyf  563 (707)
T PRK01254        489 ASGVRYDLAVED----PRYVKELVTH-HVGGYLKIAPEHTEEGPLSKMMKPGMGSYDRFKELFDKYSKEAGKEQYLIPYF  563 (707)
T ss_pred             EcCCCccccccC----HHHHHHHHHh-CCccccccccccCCHHHHHHhCCCCcccHHHHHHHHHHHHHHCCCCeEEEEeE
Confidence            467888774332    4577777775 6678899999998887322      3567899999999999998 56666665


Q ss_pred             eec
Q 013386          440 YCG  442 (444)
Q Consensus       440 ~~~  442 (444)
                      ..|
T Consensus       564 IvG  566 (707)
T PRK01254        564 ISA  566 (707)
T ss_pred             EEE
Confidence            655


No 29 
>PF09153 DUF1938:  Domain of unknown function (DUF1938);  InterPro: IPR015236 This domain, which is predominantly found in the archaeal protein O6-alkylguanine-DNA alkyltransferase, adopts a secondary structure consisting of a three stranded antiparallel beta-sheet and three alpha helices. The exact function has not, as yet, been defined, though it has been postulated that this domain may confer thermostability to the protein []. ; GO: 0005737 cytoplasm; PDB: 1MGT_A.
Probab=20.20  E-value=82  Score=26.93  Aligned_cols=24  Identities=17%  Similarity=0.350  Sum_probs=20.0

Q ss_pred             CCCChHHHHHHHHHHHHHHHhcCc
Q 013386          375 VLSSNDGAKDYVRKEIELLRTKAH  398 (444)
Q Consensus       375 ~~~g~~~~~~~L~~~~~~Lr~kGw  398 (444)
                      +++|..++++-++++++||+..|-
T Consensus        30 slDg~efl~eri~~L~~~L~kRgv   53 (86)
T PF09153_consen   30 SLDGEEFLRERISRLIEFLKKRGV   53 (86)
T ss_dssp             ESSHHHHHH-HHHHHHHHHHHTT-
T ss_pred             EeccHHHHHHHHHHHHHHHHhcCc
Confidence            468889999999999999999975


Done!