Query         011001
Match_columns 496
No_of_seqs    61 out of 63
Neff          2.9 
Searched_HMMs 46136
Date          Fri Mar 29 06:59:14 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/011001.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/011001hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 KOG3849 GDP-fucose protein O-f  99.9 9.9E-28 2.1E-32  237.5  14.0  265   73-421    27-335 (386)
  2 PF10250 O-FucT:  GDP-fucose pr  99.9   4E-27 8.6E-32  225.9   3.7   74   78-158     1-76  (351)
  3 PF05830 NodZ:  Nodulation prot  96.5   0.065 1.4E-06   55.4  14.5   81  310-392   135-224 (321)
  4 PF03254 XG_FTase:  Xyloglucan   89.7    0.33 7.1E-06   52.6   3.9   38   74-111   110-147 (476)
  5 PF01531 Glyco_transf_11:  Glyc  56.2      19 0.00041   35.9   4.9   57  342-403   162-225 (298)
  6 COG0856 Orotate phosphoribosyl  41.9      17 0.00036   36.1   1.9   26  466-491   144-169 (203)
  7 COG1040 ComFC Predicted amidop  37.1      20 0.00042   35.1   1.6   20  472-491   193-212 (225)
  8 TIGR00201 comF comF family pro  36.2      21 0.00045   33.3   1.6   20  472-491   161-180 (190)
  9 PF13768 VWA_3:  von Willebrand  33.2 1.1E+02  0.0023   26.7   5.4  128  266-437    24-152 (155)
 10 KOG2097 Predicted N6-adenine m  26.6      19 0.00041   38.2  -0.4   41  137-185   171-214 (397)
 11 PF07172 GRP:  Glycine rich pro  26.2      51  0.0011   28.9   2.2   19   24-42      3-21  (95)
 12 PF12967 DUF3855:  Domain of Un  25.5      47   0.001   31.4   1.9   45  387-431    73-137 (158)
 13 PF12404 DUF3663:  Peptidase ;   24.7 1.1E+02  0.0024   26.5   3.9   46  384-432     2-49  (77)
 14 cd01770 p47_UBX p47-like ubiqu  24.2   1E+02  0.0022   25.7   3.5   55  342-397     1-58  (79)
 15 smart00166 UBX Domain present   23.6 1.2E+02  0.0025   24.7   3.7   56  343-400     2-60  (80)
 16 PF07315 DUF1462:  Protein of u  22.1      59  0.0013   29.1   1.8   23  299-321    40-62  (93)
 17 cd01461 vWA_interalpha_trypsin  20.6 2.3E+02   0.005   24.5   5.2   75  363-439    81-157 (171)

No 1  
>KOG3849 consensus GDP-fucose protein O-fucosyltransferase [Posttranslational modification, protein turnover, chaperones]
Probab=99.95  E-value=9.9e-28  Score=237.50  Aligned_cols=265  Identities=23%  Similarity=0.363  Sum_probs=169.1

Q ss_pred             CCCceEEEecCCC-CcchHHHHHHHHHHHHHhcceeecCcccccccccCCCCCCccccccccccccccccHHHHhhccce
Q 011001           73 PDKKFFLYAPHSG-FSNQLGEFKNAILMAGILNRTLIVPPVLDHHAVALGSCPKFRVQSPNQMRISVWHHAIELLRSGRY  151 (496)
Q Consensus        73 ~~ekyl~Y~PhsG-F~NQ~~~f~nAl~lAk~LNRTLivPp~~~hha~p~~s~pK~Rv~~~~~vr~~~~~~v~ell~~~Ry  151 (496)
                      .+++||+|||||| |+||++||+++|+|||+|||||||||||.... |  ++.+      .+|++..||+++++.+|+|+
T Consensus        27 DP~GYl~yCPCMGRFGNQaDhFLGsLAFAKaLnRTL~lPpwiEy~~-p--e~~n------~~vpf~~yF~vepl~~YhRV   97 (386)
T KOG3849|consen   27 DPAGYLLYCPCMGRFGNQADHFLGSLAFAKALNRTLVLPPWIEYKH-P--ETKN------LMVPFEFYFQVEPLAKYHRV   97 (386)
T ss_pred             CCCccEEEccccccccchHHHHHHHHHHHHHhcccccCCcchhccC-C--cccc------cccchhheeecccHhhhhhh
Confidence            4899999999999 99999999999999999999999999994433 1  2234      69999999999999999999


Q ss_pred             eehhhhhcccccccCCcccccccccchhhhccccchhhhhcccCCCcchhhhHHHhhhhhcCCCCCCCC------ceEEe
Q 011001          152 VSMADIIDISSLVSSSMVKVLDFRRFASLWCGLDVDLACLISLNTQPSLLDRLRQCVSMLSGLNGNVDG------CFFAV  225 (496)
Q Consensus       152 Vsm~dfmDls~ia~~~~V~pId~R~f~S~Wcgv~~~~~c~~~l~~~~~~~~~~~~c~slL~~~~g~~~~------cvY~V  225 (496)
                      |+|+|||  +.|+|+  +||...|+.           .|..+- .|-+  .+-..|-+-    .||.=|      .|-+|
T Consensus        98 itm~dFm--~klapt--hwp~~~Rva-----------~c~k~a-~qr~--pdkp~Ch~K----eGNPFGPfWDqfhvsFv  155 (386)
T KOG3849|consen   98 ITMQDFM--KKLAPT--HWPGTPRVA-----------ICDKSA-AQRS--PDKPGCHSK----EGNPFGPFWDQFHVSFV  155 (386)
T ss_pred             eeHHHHH--HHhCcc--cCCCCccee-----------eeehhh-hccC--CCCCCCccc----CCCCCCCchhheEeeee
Confidence            9999999  999999  999999975           233321 1111  111223221    122211      12222


Q ss_pred             cccCccceeeccCCCCCCCCCCCCchHHhhhhhhhhHHHHHHHHHHhhCCCCccCcceEEEeeccccccccCceeeeecc
Q 011001          226 DDDCRTTVWTYQSGDEDGVLDPFQPDEQLKKKKKVSYVRRRRDVYKALGSGSKADSATILAFGTLFTAPYKGSQLYIDIN  305 (496)
Q Consensus       226 ~ddcrtTvwtYq~~~~d~~LdsFq~de~Lk~~Kkisyvrrrrdvyk~lG~gs~a~~a~lLaFGSLFs~~YkGse~~idi~  305 (496)
                      .+.    .        =+.+ .|.. .++.         .|..-.+-+    .+|+.-||||-+-= +||-.       .
T Consensus       156 ~sE----~--------f~~i-~Fd~-~~~~---------~~~kW~~kf----p~eeyPVLAf~gAP-A~FPv-------~  200 (386)
T KOG3849|consen  156 GSE----Y--------FGDI-GFDL-NQMG---------SRKKWLEKF----PSEEYPVLAFSGAP-APFPV-------K  200 (386)
T ss_pred             ccc----c--------cccc-ccch-hhcc---------hHHHHHhhC----CcccCceeeecCCC-CCCcc-------c
Confidence            221    0        0111 2211 1111         122222222    35999999997543 33331       1


Q ss_pred             cCcchHHHHHHHHhcccccchHHHHHhhHHHHHHhcCCCeeEEEEeec-----------ch--h----------hhhhHH
Q 011001          306 AAPRDQRIQSLIENIEFIPFVPEILSAGKKYAFETIKAPFLCAQLRLL-----------DG--Q----------FKNHWK  362 (496)
Q Consensus       306 ~s~~d~~~~sl~~~~~~lpf~p~i~~agk~~a~~~ik~pFlcaqLRll-----------DG--q----------FKnH~~  362 (496)
                      +..  -.+|.      .|.++-++..+||+||+..+..||+++|||-+           ||  +          .++|-.
T Consensus       201 ~e~--~~lQk------Yl~WS~r~~e~~k~fI~a~L~rpfvgiHLRng~DWvraCehikd~~~~hlfASpQClGy~~~~g  272 (386)
T KOG3849|consen  201 GEV--WSLQK------YLRWSSRITEQAKKFISANLARPFVGIHLRNGADWVRACEHIKDTTNRHLFASPQCLGYGHHLG  272 (386)
T ss_pred             ccc--ccHHH------HHHHHHHHHHHHHHHHHHhcCcceeEEEeecCchHHHHHHHhcccCCCccccChhhcccccccc
Confidence            111  12333      35566789999999999999999999999965           21  1          122221


Q ss_pred             ------------HHHHHHHHHHHHhhhcCCCceeEEEecCCCCCCccccccccccc--CCCceEEEEeccccH
Q 011001          363 ------------ATFLRLKEKLDSLRQKGPQPINIFVMTDLPVTNWTGNYLGDLAK--DTDSFKLYFLRKEDE  421 (496)
Q Consensus       363 ------------~Tf~~lk~kLesl~~~~~~pi~iFvMTDLp~~nWt~tyl~dl~~--~~~~ykl~~l~e~d~  421 (496)
                                  +-...+|+++.+++    ..-++||.||      ...|.++|-.  ..-.-++|.|+++|.
T Consensus       273 aLt~e~C~Psk~~I~rqik~~v~si~----dakSVfVAsD------s~hmi~Eln~aL~~~~i~vh~l~pdd~  335 (386)
T KOG3849|consen  273 ALTKEICSPSKQQILRQIKEKVGSIG----DAKSVFVASD------SDHMIDELNEALKPYEIEVHRLEPDDM  335 (386)
T ss_pred             ccchhhhCccHHHHHHHHHHHHhhhc----ccceEEEecc------chhhhHHHHHhhcccceeEEecCcccc
Confidence                        22334555555443    3557999999      3467666642  233567899999874


No 2  
>PF10250 O-FucT:  GDP-fucose protein O-fucosyltransferase;  InterPro: IPR019378  This is a family of conserved proteins representing the enzyme responsible for adding O-fucose to EGF (epidermal growth factor-like) repeats. Six highly conserved cysteines are present as well as a DXD-like motif (ERD), conserved in mammals, Drosophila, and Caenorhabditis elegans. Both features are characteristic of several glycosyltransferase families. The enzyme is a membrane-bound protein released by proteolysis and, as for most glycosyltransferases, is strongly activated by manganese []. ; PDB: 3ZY6_A 3ZY3_A 3ZY5_A 3ZY2_A 3ZY4_A.
Probab=99.93  E-value=4e-27  Score=225.86  Aligned_cols=74  Identities=30%  Similarity=0.510  Sum_probs=52.7

Q ss_pred             EEEecCCC-CcchHHHHHHHHHHHHHhcceeecCcccccccccCCCCCCccccccccccccccccHHHHhhcc-ceeehh
Q 011001           78 FLYAPHSG-FSNQLGEFKNAILMAGILNRTLIVPPVLDHHAVALGSCPKFRVQSPNQMRISVWHHAIELLRSG-RYVSMA  155 (496)
Q Consensus        78 l~Y~PhsG-F~NQ~~~f~nAl~lAk~LNRTLivPp~~~hha~p~~s~pK~Rv~~~~~vr~~~~~~v~ell~~~-RyVsm~  155 (496)
                      |.|+|++| |+||+++|+||+++|++|||||||||+..|.  .|.+-.+     -.+++++.+|.+..+.++. ++|+|+
T Consensus         1 ~~y~p~~GGfnNQr~~~~~a~~~A~~LnRTLVLPp~~~~~--~~~~~~~-----~~~ipf~~~fD~~~l~~~~~~vi~~~   73 (351)
T PF10250_consen    1 LVYDPCMGGFNNQRMGFENAVVFAKALNRTLVLPPFIKHY--HWKDQSK-----QRHIPFSDFFDVEHLRKFLRPVITME   73 (351)
T ss_dssp             EEE---SSSHHHHHHHHHHHHHHHHHHT-EEE--EEEEES--SSS---------EEEEEHHHHB-HHHHTTTS--EE-HH
T ss_pred             CccCCCCCCHHHHHHHHHHHHHHHHHhCCEEEcCCccccc--ccccccc-----ccccChhhhccHHHHHHHhhCceehh
Confidence            68999977 9999999999999999999999999999752  2321111     2578888889999999999 889999


Q ss_pred             hhh
Q 011001          156 DII  158 (496)
Q Consensus       156 dfm  158 (496)
                      +++
T Consensus        74 ef~   76 (351)
T PF10250_consen   74 EFL   76 (351)
T ss_dssp             HHH
T ss_pred             eec
Confidence            998


No 3  
>PF05830 NodZ:  Nodulation protein Z (NodZ);  InterPro: IPR008716 The nodulation genes of Rhizobia are regulated by the nodD gene product in response to host-produced flavonoids and appear to encode enzymes involved in the production of a lipo-chitose signal molecule required for infection and nodule formation. NodZ is required for the addition of a 2-O-methylfucose residue to the terminal reducing N-acetylglucosamine of the nodulation signal. This substitution is essential for the biological activity of this molecule. Mutations in nodZ result in defective nodulation. nodZ represents a unique nodulation gene that is not under the control of NodD and yet is essential for the synthesis of an active nodulation signal [].; GO: 0016758 transferase activity, transferring hexosyl groups, 0009312 oligosaccharide biosynthetic process, 0009877 nodulation; PDB: 3SIX_A 2HLH_A 2HHC_A 3SIW_A 2OCX_A.
Probab=96.55  E-value=0.065  Score=55.43  Aligned_cols=81  Identities=20%  Similarity=0.284  Sum_probs=48.3

Q ss_pred             hHHHHHHHHhcccccchHHHHHhhHHHHHHhcCC-CeeEEEEeecchhh----hhhHH---HHHHHHHHHHHHhh-hcCC
Q 011001          310 DQRIQSLIENIEFIPFVPEILSAGKKYAFETIKA-PFLCAQLRLLDGQF----KNHWK---ATFLRLKEKLDSLR-QKGP  380 (496)
Q Consensus       310 d~~~~sl~~~~~~lpf~p~i~~agk~~a~~~ik~-pFlcaqLRllDGqF----KnH~~---~Tf~~lk~kLesl~-~~~~  380 (496)
                      |+.+.+.|-  +.+--.|+|.+....+.++.+++ +-.|+|+|-|+|+=    ..+|.   .++..++..++.++ +..+
T Consensus       135 ~~~aeR~if--~slkpR~eIqarID~iy~ehf~g~~~IGVHVRhGngeD~~~h~~~~~D~e~~L~~V~~ai~~ak~~~~~  212 (321)
T PF05830_consen  135 DEEAEREIF--SSLKPRPEIQARIDAIYREHFAGYSVIGVHVRHGNGEDIMDHAPYWADEERALRQVCTAIDKAKALAPP  212 (321)
T ss_dssp             -HHHHHHHH--HHS-B-HHHHHHHHHHHHHHTTTSEEEEEEE---------------HHHHHHHHHHHHHHHHHHTS--S
T ss_pred             hhHHHHHHH--HhCCCCHHHHHHHHHHHHHHcCCCceEEEEEeccCCcchhccCccccCchHHHHHHHHHHHHHHhccCC
Confidence            466667666  33667899999999999998865 59999999998842    24454   45666666666664 4455


Q ss_pred             CceeEEEecCCC
Q 011001          381 QPINIFVMTDLP  392 (496)
Q Consensus       381 ~pi~iFvMTDLp  392 (496)
                      .|+.|||.||=+
T Consensus       213 k~~~IFLATDSa  224 (321)
T PF05830_consen  213 KPVRIFLATDSA  224 (321)
T ss_dssp             S-EEEEEEES-H
T ss_pred             CCeeEEEecCcH
Confidence            799999999954


No 4  
>PF03254 XG_FTase:  Xyloglucan fucosyltransferase;  InterPro: IPR004938  Plant cell walls are crucial for development, signal transduction, and disease resistance in plants. Cell walls are made of cellulose, hemicelluloses, and pectins. Xyloglucan (XG), the principal load-bearing hemicellulose of dicotyledonous plants, has a terminal fucosyl residue. This fucosyltransferase adds this residue []. ; GO: 0008107 galactoside 2-alpha-L-fucosyltransferase activity, 0042546 cell wall biogenesis, 0016020 membrane
Probab=89.70  E-value=0.33  Score=52.64  Aligned_cols=38  Identities=34%  Similarity=0.603  Sum_probs=36.2

Q ss_pred             CCceEEEecCCCCcchHHHHHHHHHHHHHhcceeecCc
Q 011001           74 DKKFFLYAPHSGFSNQLGEFKNAILMAGILNRTLIVPP  111 (496)
Q Consensus        74 ~ekyl~Y~PhsGF~NQ~~~f~nAl~lAk~LNRTLivPp  111 (496)
                      +=|||.|.|.+|.||++..+--|++-|-+.||.|+|.+
T Consensus       110 ~CkYvVw~~~~GLGNRmLslaSaFLYAlLT~RVLLV~~  147 (476)
T PF03254_consen  110 ECKYVVWIPYSGLGNRMLSLASAFLYALLTNRVLLVDP  147 (476)
T ss_pred             CCcEEEEecCCchHHHHHHHHHHHHHHHHhCcEEEEec
Confidence            55999999999999999999999999999999999976


No 5  
>PF01531 Glyco_transf_11:  Glycosyl transferase family 11;  InterPro: IPR002516 The biosynthesis of disaccharides, oligosaccharides and polysaccharides involves the action of hundreds of different glycosyltransferases. These enzymes catalyse the transfer of sugar moieties from activated donor molecules to specific acceptor molecules, forming glycosidic bonds. A classification of glycosyltransferases using nucleotide diphospho-sugar, nucleotide monophospho-sugar and sugar phosphates (2.4.1.- from EC) and related proteins into distinct sequence based families has been described []. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. The same three-dimensional fold is expected to occur within each of the families. Because 3-D structures are better conserved than sequences, several of the families defined on the basis of sequence similarities may have similar 3-D structures and therefore form 'clans'. Glycosyltransferase family 11 GT11 from CAZY comprises enzymes with only one known activity; galactoside 2-L-fucosyltransferase (2.4.1.69 from EC).  Some of the proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Galactoside 2-L-fucosyltransferase 1 (2.4.1.69 from EC) and Galactoside 2-L-fucosyltransferase 2 (2.4.1.69 from EC) belong to the Hh blood group system and are associated with H/h and Se/se antigens.; GO: 0008107 galactoside 2-alpha-L-fucosyltransferase activity, 0005975 carbohydrate metabolic process, 0016020 membrane
Probab=56.18  E-value=19  Score=35.92  Aligned_cols=57  Identities=21%  Similarity=0.387  Sum_probs=40.0

Q ss_pred             CCCeeEEEEeecchhhhh-------hHHHHHHHHHHHHHHhhhcCCCceeEEEecCCCCCCcccccccc
Q 011001          342 KAPFLCAQLRLLDGQFKN-------HWKATFLRLKEKLDSLRQKGPQPINIFVMTDLPVTNWTGNYLGD  403 (496)
Q Consensus       342 k~pFlcaqLRllDGqFKn-------H~~~Tf~~lk~kLesl~~~~~~pi~iFvMTDLp~~nWt~tyl~d  403 (496)
                      ...++|+|+|-||  |..       |...+.+-.++.++-++.+-+.| .+||.+|  .-+|.+..+..
T Consensus       162 ~~~~V~VHIRRGD--y~~~~~~~~~~~~~~~~Yy~~Ai~~i~~~~~~~-~f~ifSD--D~~w~k~~l~~  225 (298)
T PF01531_consen  162 NSNSVCVHIRRGD--YVSNGNHNWKHGICDKDYYKKAIEYIREKVKNP-KFFIFSD--DIEWCKENLKF  225 (298)
T ss_pred             CCCeEEEEEEchh--ccccccccccCCCCCHHHHHHHHHHHHHhCCCC-EEEEEcC--CHHHHHHHHhh
Confidence            3579999999998  432       33456677888888887666555 4677777  33588776654


No 6  
>COG0856 Orotate phosphoribosyltransferase homologs [Nucleotide transport and metabolism]
Probab=41.91  E-value=17  Score=36.07  Aligned_cols=26  Identities=38%  Similarity=0.468  Sum_probs=22.1

Q ss_pred             ccccccccCCCchhHHHHHHHHhccc
Q 011001          466 CATVGFVGTAGSTLAESIELMRKFDV  491 (496)
Q Consensus       466 CAslGFvGT~GSTia~~ie~mRk~~~  491 (496)
                      |.-.-=|-|+||||.|-||++|+.+.
T Consensus       144 cvIVDDvittG~Ti~E~Ie~lke~g~  169 (203)
T COG0856         144 CVIVDDVITTGSTIKETIEQLKEEGG  169 (203)
T ss_pred             EEEEecccccChhHHHHHHHHHHcCC
Confidence            55566678999999999999999874


No 7  
>COG1040 ComFC Predicted amidophosphoribosyltransferases [General function prediction only]
Probab=37.14  E-value=20  Score=35.08  Aligned_cols=20  Identities=40%  Similarity=0.534  Sum_probs=18.6

Q ss_pred             ccCCCchhHHHHHHHHhccc
Q 011001          472 VGTAGSTLAESIELMRKFDV  491 (496)
Q Consensus       472 vGT~GSTia~~ie~mRk~~~  491 (496)
                      |-|+|+|+.+.-+.||+.|+
T Consensus       193 V~TTGaTl~~~~~~L~~~Ga  212 (225)
T COG1040         193 VYTTGATLKEAAKLLREAGA  212 (225)
T ss_pred             ccccHHHHHHHHHHHHHcCC
Confidence            67999999999999999985


No 8  
>TIGR00201 comF comF family protein. This protein is found in species that do (Bacillus subtilis, Haemophilus influenzae) or do not (E. coli, Borrelia burgdorferi) have described systems for natural transformation with exogenous DNA. It is involved in competence for transformation in Bacillus subtilis.
Probab=36.21  E-value=21  Score=33.32  Aligned_cols=20  Identities=35%  Similarity=0.434  Sum_probs=18.1

Q ss_pred             ccCCCchhHHHHHHHHhccc
Q 011001          472 VGTAGSTLAESIELMRKFDV  491 (496)
Q Consensus       472 vGT~GSTia~~ie~mRk~~~  491 (496)
                      |-|+|+|+.+..+.+++.|+
T Consensus       161 V~TTGaTl~~~~~~L~~~Ga  180 (190)
T TIGR00201       161 VVTTGATLHEIARLLLELGA  180 (190)
T ss_pred             eeccHHHHHHHHHHHHHcCC
Confidence            56999999999999999875


No 9  
>PF13768 VWA_3:  von Willebrand factor type A domain
Probab=33.18  E-value=1.1e+02  Score=26.67  Aligned_cols=128  Identities=20%  Similarity=0.243  Sum_probs=75.7

Q ss_pred             HHHHHHhhCCCCccCcceEEEeeccccccccCceeeeecccCcchHHHHHHHHhcccccchHHHHHhhHHHHHHhcCCCe
Q 011001          266 RRDVYKALGSGSKADSATILAFGTLFTAPYKGSQLYIDINAAPRDQRIQSLIENIEFIPFVPEILSAGKKYAFETIKAPF  345 (496)
Q Consensus       266 rrdvyk~lG~gs~a~~a~lLaFGSLFs~~YkGse~~idi~~s~~d~~~~sl~~~~~~lpf~p~i~~agk~~a~~~ik~pF  345 (496)
                      -+-+.++|++|.   ..+|++||+-... +.           +            ...|..++-+..+.+++++ +.++ 
T Consensus        24 l~~~l~~L~~~d---~fnii~f~~~~~~-~~-----------~------------~~~~~~~~~~~~a~~~I~~-~~~~-   74 (155)
T PF13768_consen   24 LRAILRSLPPGD---RFNIIAFGSSVRP-LF-----------P------------GLVPATEENRQEALQWIKS-LEAN-   74 (155)
T ss_pred             HHHHHHhCCCCC---EEEEEEeCCEeeE-cc-----------h------------hHHHHhHHHHHHHHHHHHH-hccc-
Confidence            355677899865   7899999984311 11           0            0234445555666666544 2211 


Q ss_pred             eEEEEeecchhhhhhHHHHHHHHHHHHHHhhhcCCCceeEEEecCCCCCCccccc-ccccccCCCceEEEEeccccHHHH
Q 011001          346 LCAQLRLLDGQFKNHWKATFLRLKEKLDSLRQKGPQPINIFVMTDLPVTNWTGNY-LGDLAKDTDSFKLYFLRKEDELLA  424 (496)
Q Consensus       346 lcaqLRllDGqFKnH~~~Tf~~lk~kLesl~~~~~~pi~iFvMTDLp~~nWt~ty-l~dl~~~~~~ykl~~l~e~d~lv~  424 (496)
                              +|.     ...-.+|+..+..+ .....+-+|+++||=.+ .++... +..+.+...++.+|.+.-++..-.
T Consensus        75 --------~G~-----t~l~~aL~~a~~~~-~~~~~~~~IilltDG~~-~~~~~~i~~~v~~~~~~~~i~~~~~g~~~~~  139 (155)
T PF13768_consen   75 --------SGG-----TDLLAALRAALALL-QRPGCVRAIILLTDGQP-VSGEEEILDLVRRARGHIRIFTFGIGSDADA  139 (155)
T ss_pred             --------CCC-----ccHHHHHHHHHHhc-ccCCCccEEEEEEeccC-CCCHHHHHHHHHhcCCCceEEEEEECChhHH
Confidence                    111     12223444444433 34556788899998665 333333 444444456788988888877777


Q ss_pred             HHHHHHHHhccCc
Q 011001          425 QTAQKLATAGHGL  437 (496)
Q Consensus       425 ~ta~kl~~a~hg~  437 (496)
                      +.-++|+.+.+|.
T Consensus       140 ~~L~~LA~~~~G~  152 (155)
T PF13768_consen  140 DFLRELARATGGS  152 (155)
T ss_pred             HHHHHHHHcCCCE
Confidence            8888888888874


No 10 
>KOG2097 consensus Predicted N6-adenine methylase involved in transcription regulation [Transcription]
Probab=26.60  E-value=19  Score=38.21  Aligned_cols=41  Identities=20%  Similarity=0.428  Sum_probs=28.1

Q ss_pred             cccccHHHHhhcc-ceeehhhhhcc--cccccCCcccccccccchhhhcccc
Q 011001          137 SVWHHAIELLRSG-RYVSMADIIDI--SSLVSSSMVKVLDFRRFASLWCGLD  185 (496)
Q Consensus       137 ~~~~~v~ell~~~-RyVsm~dfmDl--s~ia~~~~V~pId~R~f~S~Wcgv~  185 (496)
                      ++|.+...|+... .++.|+||++|  ++|        ++.|.|+.+|||-.
T Consensus       171 eeyv~~~g~~t~n~~fw~~~di~nL~id~i--------aa~psFlFlW~gs~  214 (397)
T KOG2097|consen  171 EEYVRMAGCLTENMQFWTWDDIQNLPIDEI--------AAKPSFLFLWCGSG  214 (397)
T ss_pred             HHHHHhccccccCceEecHHHhhcCchhhh--------ccCCceEEEEecCc
Confidence            3455555655544 77999999954  444        44689999998753


No 11 
>PF07172 GRP:  Glycine rich protein family;  InterPro: IPR010800 This family consists of glycine rich proteins. Some of them may be involved in resistance to environmental stress [].
Probab=26.16  E-value=51  Score=28.90  Aligned_cols=19  Identities=32%  Similarity=0.362  Sum_probs=11.4

Q ss_pred             CcchHHHHHHHHHHHHHHH
Q 011001           24 SPFFILSITIFTFLLLFIA   42 (496)
Q Consensus        24 ~~~~l~~~~~~~~~~~~~~   42 (496)
                      |+.+||+..+|+++||+++
T Consensus         3 SK~~llL~l~LA~lLlisS   21 (95)
T PF07172_consen    3 SKAFLLLGLLLAALLLISS   21 (95)
T ss_pred             hhHHHHHHHHHHHHHHHHh
Confidence            4567777666666554444


No 12 
>PF12967 DUF3855:  Domain of Unknown Function with PDB structure (DUF3855);  InterPro: IPR024482 This domain forms an unusual alpha/beta fold where a six-stranded antiparallel beta-sheet is wrapped around a central alpha-helix, flanked by an additional alpha-helix and a small sub-domain consisting of a single beta-strand and a two-stranded antiparallel beta-sheet []. It shows weak structural similarities to phosphoribosylformylglycinamidine synthases and some thioesterase superfamily members, but its function is unknown.; PDB: 1O22_A.
Probab=25.47  E-value=47  Score=31.37  Aligned_cols=45  Identities=22%  Similarity=0.579  Sum_probs=25.6

Q ss_pred             EecCCCCCCccc---------ccccc-----cccCCCceEEEEe------ccccHHHHHHHHHHH
Q 011001          387 VMTDLPVTNWTG---------NYLGD-----LAKDTDSFKLYFL------RKEDELLAQTAQKLA  431 (496)
Q Consensus       387 vMTDLp~~nWt~---------tyl~d-----l~~~~~~ykl~~l------~e~d~lv~~ta~kl~  431 (496)
                      ..||||-.+||.         +||+|     +.+|...|++|.-      +.+||+|.+--+-..
T Consensus        73 ~a~~lplg~w~~l~nvfvee~~yl~~y~~mki~s~~n~y~~yvpys~vk~knr~e~v~~fmkyff  137 (158)
T PF12967_consen   73 NAVDLPLGDWTDLNNVFVEEISYLDSYDYMKIHSEKNWYKIYVPYSSVKSKNRNEVVEEFMKYFF  137 (158)
T ss_dssp             --TT--SSS-----S-EEEEEEE-EEETTEEEEEETTEEEEEEEGGGSTT--HHHHHHHHHHHHH
T ss_pred             ccccCCcchhHHHHHHHHHhhhhhhccCCeEEeccCcEEEEEeehHHhhhccHHHHHHHHHHHHH
Confidence            469999999985         35554     4568889999963      667888887655443


No 13 
>PF12404 DUF3663:  Peptidase ;  InterPro: IPR008330 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Metalloproteases are the most diverse of the four main types of protease, with more than 50 families identified to date. In these enzymes, a divalent cation, usually zinc, activates the water molecule. The metal ion is held in place by amino acid ligands, usually three in number. The known metal ligands are His, Glu, Asp or Lys and at least one other residue is required for catalysis, which may play an electrophillic role. Of the known metalloproteases, around half contain an HEXXH motif, which has been shown in crystallographic studies to form part of the metal-binding site []. The HEXXH motif is relatively common, but can be more stringently defined for metalloproteases as 'abXHEbbHbc', where 'a' is most often valine or threonine and forms part of the S1' subsite in thermolysin and neprilysin, 'b' is an uncharged residue, and 'c' a hydrophobic residue. Proline is never found in this site, possibly because it would break the helical structure adopted by this motif in metalloproteases []. This family represents the peptidase B group of leucyl aminopeptidases, which are restricted to the gammaproteobacteria. They contain a C-terminal aminopeptidase catalytic domain and an N-terminal domain of unknown function. They are zinc-dependent exopeptidases (3.4.11.1 from EC) and belong to MEROPS peptidase family M17 (leucyl aminopeptidase family, clan MF). They selectively release N-terminal amino acid residues from polypeptides and proteins and are involved in the processing, catabolism and degradation of intracellular proteins [, , ]. Leucyl aminopeptidase forms a homohexamer containing two trimers stacked on top of one another []. Each monomer binds two zinc ions. The zinc-binding and catalytic sites are located within the C-terminal catalytic domain []. The same catalytic aminopeptidase domain is found in the other M17 peptidases IPR011356 from INTERPRO. These two groups of aminopeptidases differ by their N-terminal domains. The N-terminal domain in members of IPR011356 from INTERPRO has been implicated in DNA binding [, ] and it is not associated with members of this family which have a different N-terminal domain and therefore are not expected to bind DNA or be involved in transcriptional regulation. In addition, there are related proteins with the same catalytic domain and unique N-terminal sequences unrelated to any of the two N-terminal domains discussed above. For additional information please see [, , , ]. ; GO: 0004177 aminopeptidase activity, 0008235 metalloexopeptidase activity, 0030145 manganese ion binding, 0005737 cytoplasm
Probab=24.73  E-value=1.1e+02  Score=26.47  Aligned_cols=46  Identities=15%  Similarity=0.320  Sum_probs=35.5

Q ss_pred             eEEEecCCCCCCcccccccccccCCCceEEEEeccccHH--HHHHHHHHHH
Q 011001          384 NIFVMTDLPVTNWTGNYLGDLAKDTDSFKLYFLRKEDEL--LAQTAQKLAT  432 (496)
Q Consensus       384 ~iFvMTDLp~~nWt~tyl~dl~~~~~~ykl~~l~e~d~l--v~~ta~kl~~  432 (496)
                      +|++-++-+.+.|-..  +.|+-+.+-..||- .++|+|  |+++||||-.
T Consensus         2 ~V~LS~~~A~a~WG~~--AllSf~~~ga~IHl-~~~~~l~~IQrAaRkLd~   49 (77)
T PF12404_consen    2 QVTLSQQPAAAHWGEK--ALLSFNEQGATIHL-SEGDDLRAIQRAARKLDG   49 (77)
T ss_pred             eEEeeCCCChhHhCcC--cEEEEcCCCEEEEE-CCCcchHHHHHHHHHHhh
Confidence            5788899999999855  34555667888987 666664  8999999976


No 14 
>cd01770 p47_UBX p47-like ubiquitin domain. p47_UBX  p47 is an adaptor molecule of the cytosolic AAA ATPase p97. The principal role of the p97-p47 complex is to regulate membrane fusion events. Mono-ubiquitin recognition by p47 is crucial for p97-p47-mediated Golgi membrane fusion events.  p47 has carboxy-terminal SEP and UBX domains.  The UBX domain has a beta-grasp fold similar to that of ubiquitin however, UBX lacks the c-terminal double glycine motif and is thus unlikely to be conjugated to other proteins.
Probab=24.20  E-value=1e+02  Score=25.70  Aligned_cols=55  Identities=20%  Similarity=0.236  Sum_probs=37.1

Q ss_pred             CCCeeEEEEeecchhhh-hh--HHHHHHHHHHHHHHhhhcCCCceeEEEecCCCCCCcc
Q 011001          342 KAPFLCAQLRLLDGQFK-NH--WKATFLRLKEKLDSLRQKGPQPINIFVMTDLPVTNWT  397 (496)
Q Consensus       342 k~pFlcaqLRllDGqFK-nH--~~~Tf~~lk~kLesl~~~~~~pi~iFvMTDLp~~nWt  397 (496)
                      ++|---+|+||.||.=. .+  ...|+..|++=+++-.. ++..-+.-+||-.|..+.+
T Consensus         1 ~~p~t~iqiRlpdG~r~~~rF~~~~tv~~l~~~v~~~~~-~~~~~~f~L~t~fP~k~l~   58 (79)
T cd01770           1 LEPTTSIQIRLADGKRLVQKFNSSHRVSDVRDFIVNARP-EFAARPFTLMTAFPVKELS   58 (79)
T ss_pred             CCCeeEEEEECCCCCEEEEEeCCCCcHHHHHHHHHHhCC-CCCCCCEEEecCCCCcccC
Confidence            35667899999999432 22  33788999998886542 2223455668888987665


No 15 
>smart00166 UBX Domain present in ubiquitin-regulatory proteins. Present in FAF1 and Shp1p.
Probab=23.57  E-value=1.2e+02  Score=24.73  Aligned_cols=56  Identities=20%  Similarity=0.210  Sum_probs=38.1

Q ss_pred             CCeeEEEEeecchhhh-h--hHHHHHHHHHHHHHHhhhcCCCceeEEEecCCCCCCccccc
Q 011001          343 APFLCAQLRLLDGQFK-N--HWKATFLRLKEKLDSLRQKGPQPINIFVMTDLPVTNWTGNY  400 (496)
Q Consensus       343 ~pFlcaqLRllDGqFK-n--H~~~Tf~~lk~kLesl~~~~~~pi~iFvMTDLp~~nWt~ty  400 (496)
                      ++..-+|+|+.||.-. .  +...|+..|++-+.+....+..|  .-++|-.|...++..+
T Consensus         2 ~~~~~I~iRlPdG~ri~~~F~~~~tl~~v~~~v~~~~~~~~~~--f~L~t~~Prk~l~~~d   60 (80)
T smart00166        2 SDQCRLQIRLPDGSRLVRRFPSSDTLRTVYEFVSAALTDGNDP--FTLNSPFPRRTFTKDD   60 (80)
T ss_pred             CCeEEEEEEcCCCCEEEEEeCCCCcHHHHHHHHHHcccCCCCC--EEEEeCCCCcCCcccc
Confidence            3556789999999822 1  23488999999996544333344  5678888888776543


No 16 
>PF07315 DUF1462:  Protein of unknown function (DUF1462);  InterPro: IPR009190 There are currently no experimental data for members of this group of bacterial proteins or their homologues. A crystal structure of Q7A6J8 from SWISSPROT revealed a thioredoxin-like fold, its core consisting of three layers alpha/beta/alpha.; PDB: 1XG8_A.
Probab=22.08  E-value=59  Score=29.06  Aligned_cols=23  Identities=35%  Similarity=0.502  Sum_probs=15.7

Q ss_pred             eeeeecccCcchHHHHHHHHhcc
Q 011001          299 QLYIDINAAPRDQRIQSLIENIE  321 (496)
Q Consensus       299 e~~idi~~s~~d~~~~sl~~~~~  321 (496)
                      -.||||++++.+..-+.++++|+
T Consensus        40 ~~YiDi~~p~~~~~~~~~a~~I~   62 (93)
T PF07315_consen   40 FTYIDIENPPENDHDQQFAERIL   62 (93)
T ss_dssp             EEEEETTT----HHHHHHHHHHH
T ss_pred             EEEEecCCCCccHHHHHHHHHHH
Confidence            45999999998877788888876


No 17 
>cd01461 vWA_interalpha_trypsin_inhibitor vWA_interalpha trypsin inhibitor (ITI): ITI is a glycoprotein composed of three polypeptides- two heavy chains and one light chain (bikunin). Bikunin confers the protease-inhibitor function while the heavy chains are involved in rendering stability to the extracellular matrix by binding to hyaluronic acid. The heavy chains carry the VWA domain with a conserved MIDAS motif. Although the exact role of the VWA domains remains unknown, it has been speculated to be involved in mediating protein-protein interactions with the components of the extracellular matrix.
Probab=20.64  E-value=2.3e+02  Score=24.47  Aligned_cols=75  Identities=19%  Similarity=0.211  Sum_probs=44.1

Q ss_pred             HHHHHHHHHHHHhhhcCCCceeEEEecCCCCCCcccccccccccC-C-CceEEEEeccccHHHHHHHHHHHHhccCcee
Q 011001          363 ATFLRLKEKLDSLRQKGPQPINIFVMTDLPVTNWTGNYLGDLAKD-T-DSFKLYFLRKEDELLAQTAQKLATAGHGLRY  439 (496)
Q Consensus       363 ~Tf~~lk~kLesl~~~~~~pi~iFvMTDLp~~nWt~tyl~dl~~~-~-~~ykl~~l~e~d~lv~~ta~kl~~a~hg~r~  439 (496)
                      ....+|+..++.+......+-.|+++||--..+.  ..+.+.++. . ...++|.+.-+++.=....++++++.-|.-+
T Consensus        81 ~l~~al~~a~~~l~~~~~~~~~iillTDG~~~~~--~~~~~~~~~~~~~~i~i~~i~~g~~~~~~~l~~ia~~~gG~~~  157 (171)
T cd01461          81 NMNDALEAALELLNSSPGSVPQIILLTDGEVTNE--SQILKNVREALSGRIRLFTFGIGSDVNTYLLERLAREGRGIAR  157 (171)
T ss_pred             CHHHHHHHHHHhhccCCCCccEEEEEeCCCCCCH--HHHHHHHHHhcCCCceEEEEEeCCccCHHHHHHHHHcCCCeEE
Confidence            4566777777776554556788899999764332  211112211 1 2567777776543334556777777766655


Done!