Query         012348
Match_columns 465
No_of_seqs    133 out of 372
Neff          3.6 
Searched_HMMs 46136
Date          Fri Mar 29 01:41:45 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/012348.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/012348hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF02362 B3:  B3 DNA binding do  99.7 8.2E-17 1.8E-21  132.0  11.1   98  359-460     1-99  (100)
  2 PF03754 DUF313:  Domain of unk  98.2 3.3E-06   7E-11   74.9   7.0   80  353-432    18-114 (114)
  3 PF09217 EcoRII-N:  Restriction  97.7 0.00011 2.3E-09   68.6   7.8   89  356-444     7-110 (156)
  4 smart00249 PHD PHD zinc finger  69.7     2.2 4.8E-05   29.7   0.9   25   77-101    10-34  (47)
  5 PF10844 DUF2577:  Protein of u  66.2      17 0.00037   31.4   5.8   82  354-458    16-99  (100)
  6 PRK03760 hypothetical protein;  39.3      59  0.0013   29.1   4.8   29  417-445    88-116 (117)
  7 PF13248 zf-ribbon_3:  zinc-rib  38.0      15 0.00033   24.6   0.8   15   77-91     12-26  (26)
  8 PF08922 DUF1905:  Domain of un  36.6   1E+02  0.0022   25.6   5.6   79  359-444     1-79  (80)
  9 PF02643 DUF192:  Uncharacteriz  36.2      82  0.0018   27.4   5.2   52  393-444    49-107 (108)
 10 KOG4718 Non-SMC (structural ma  33.4      16 0.00034   36.7   0.3   18   82-99    195-212 (235)
 11 PF04014 Antitoxin-MazE:  Antid  30.5      57  0.0012   24.2   2.8   23  427-449    13-35  (47)
 12 TIGR01643 YD_repeat_2x YD repe  30.2      76  0.0016   22.4   3.3   22  392-413     4-25  (42)
 13 COG1998 RPS31 Ribosomal protei  29.5      27 0.00058   27.8   1.0   13   82-94     20-41  (51)
 14 PF03120 DNA_ligase_OB:  NAD-de  29.4      41 0.00088   28.8   2.1   22  427-448    42-63  (82)
 15 PF09297 zf-NADH-PPase:  NADH p  27.6      31 0.00067   24.0   0.9   25   58-91      5-31  (32)
 16 cd05829 Sortase_E Sortase E (S  27.2 1.2E+02  0.0025   27.6   4.8   39  418-456    49-94  (144)
 17 PF12760 Zn_Tnp_IS1595:  Transp  24.5      41 0.00089   25.1   1.2   26   58-90     20-46  (46)
 18 cd04459 Rho_CSD Rho_CSD: Rho p  23.4      71  0.0015   26.4   2.4   26  428-453    34-60  (68)
 19 TIGR03784 marine_sortase sorta  23.3      84  0.0018   29.9   3.2   81  371-459    51-133 (174)
 20 TIGR00375 conserved hypothetic  22.8      36 0.00079   36.2   0.8   44   54-107   238-282 (374)
 21 PRK09570 rpoH DNA-directed RNA  22.3   1E+02  0.0022   26.4   3.3   32  427-459    44-75  (79)
 22 PF01191 RNA_pol_Rpb5_C:  RNA p  22.3 1.2E+02  0.0026   25.6   3.6   32  427-459    41-72  (74)
 23 PF05593 RHS_repeat:  RHS Repea  22.2 1.4E+02   0.003   21.3   3.4   23  391-413     3-25  (38)

No 1  
>PF02362 B3:  B3 DNA binding domain;  InterPro: IPR003340 Two DNA binding proteins, RAV1 and RAV2 from Arabidopsis thaliana contain two distinct amino acid sequence domains found only in higher plant species. The N-terminal regions of RAV1 and RAV2 are homologous to the AP2 DNA-binding domain (see IPR001471 from INTERPRO) present in a family of transcription factors, while the C-terminal region exhibits homology to the highly conserved C-terminal domain, designated B3, of VP1/ABI3 transcription factors []. The AP2 and B3-like domains of RAV1 bind autonomously to the CAACA and CACCTG motifs, respectively, and together achieve a high affinity and specificity of binding. It has been suggested that the AP2 and B3-like domains of RAV1 are connected by a highly flexible structure enabling the two domains to bind to the CAACA and CACCTG motifs in various spacings and orientations [].; GO: 0003677 DNA binding, 0006355 regulation of transcription, DNA-dependent; PDB: 1WID_A 1YEL_A.
Probab=99.71  E-value=8.2e-17  Score=131.96  Aligned_cols=98  Identities=22%  Similarity=0.325  Sum_probs=71.9

Q ss_pred             EEEecccccCCCCCcEEeehhhhhhcCCCCCCCCCceEEEEeCCCCeEEeEEEEcCCCCCCceeec-CchhHhhhcCCCC
Q 012348          359 FEKMLSASDAGRIGRLVLPKKCAEAYFPPISQPEGLPLKVQDSKGKEWIFQFRFWPNNNSRMYVLE-GVTPCIQNMQLQA  437 (465)
Q Consensus       359 F~KvLT~SDVg~lgRLVIPK~~AEa~FP~L~~~~Gv~L~veD~~Gk~W~Frfrfw~NnsSR~YVLt-GW~~FVrsK~Lqa  437 (465)
                      |.|+|+++|+.+..+|+||++.+++|.  .....++.|.++|..|++|.+++.++  +.++.|+|+ ||.+||++++|++
T Consensus         1 F~K~l~~s~~~~~~~l~iP~~f~~~~~--~~~~~~~~v~l~~~~g~~W~v~~~~~--~~~~~~~l~~GW~~Fv~~n~L~~   76 (100)
T PF02362_consen    1 FFKVLKPSDVSSSCRLIIPKEFAKKHG--GNKRKSREVTLKDPDGRSWPVKLKYR--KNSGRYYLTGGWKKFVRDNGLKE   76 (100)
T ss_dssp             EEEE--TTCCCCTT-EEE-HHHHTTTS----SS--CEEEEEETTTEEEEEEEEEE--CCTTEEEEETTHHHHHHHCT--T
T ss_pred             CEEEEEccCcCCCCEEEeCHHHHHHhC--CCcCCCeEEEEEeCCCCEEEEEEEEE--ccCCeEEECCCHHHHHHHcCCCC
Confidence            899999999998889999999999882  22235789999999999999999987  445568886 9999999999999


Q ss_pred             CCEEEEeecCCCCceEEEEEeee
Q 012348          438 GDIGNQPKSQESYEMFSFMWKVH  460 (465)
Q Consensus       438 GDtVtF~R~~ngg~rf~I~~rr~  460 (465)
                      ||.|+|+...++...+.+.+.++
T Consensus        77 GD~~~F~~~~~~~~~~~v~i~~~   99 (100)
T PF02362_consen   77 GDVCVFELIGNSNFTLKVHIFRK   99 (100)
T ss_dssp             T-EEEEEE-SSSCE-EEEEEE--
T ss_pred             CCEEEEEEecCCCceEEEEEEEC
Confidence            99999999876555567777654


No 2  
>PF03754 DUF313:  Domain of unknown function (DUF313) ;  InterPro: IPR005508 This is a family of proteins from Arabidopsis thaliana (Mouse-ear cress) with uncharacterised function.
Probab=98.22  E-value=3.3e-06  Score=74.89  Aligned_cols=80  Identities=21%  Similarity=0.433  Sum_probs=63.9

Q ss_pred             CcccceEEEecccccCC-CCCcEEeehhhhhh--cCCC-----C-------CCCCCceEEEEeCCCCeEEeEEEEcCC-C
Q 012348          353 SVITPLFEKMLSASDAG-RIGRLVLPKKCAEA--YFPP-----I-------SQPEGLPLKVQDSKGKEWIFQFRFWPN-N  416 (465)
Q Consensus       353 s~~~~LF~KvLT~SDVg-~lgRLVIPK~~AEa--~FP~-----L-------~~~~Gv~L~veD~~Gk~W~Frfrfw~N-n  416 (465)
                      .....+++|+|++|||. ..+||.||......  +|=+     +       ....|+.+.+.|..+++|..+++.|.- +
T Consensus        18 ~d~kli~~K~L~~tDv~~~qsRLsmP~~qi~~~dFLt~eE~~~i~~~~~~~~~~~Gv~V~lvdp~~~~~~m~lkkW~mg~   97 (114)
T PF03754_consen   18 EDPKLIIEKTLFKTDVDPHQSRLSMPFNQIIDNDFLTEEEKRIIKEEKKNNDKKKGVEVILVDPSLRKWTMRLKKWNMGN   97 (114)
T ss_pred             CCCeEEEeeeecccCCCCCCceeeccHHHhcccccCCHHHHHHHHHhhccCcccCCceEEEECCcCcEEEEEEEEecccC
Confidence            45578999999999999 48899999987633  2211     2       235689999999999999999999965 4


Q ss_pred             CCCceeec-CchhHhhh
Q 012348          417 NSRMYVLE-GVTPCIQN  432 (465)
Q Consensus       417 sSR~YVLt-GW~~FVrs  432 (465)
                      ..-.|+|. ||.+.|.+
T Consensus        98 ~~~~YvL~~gWn~VV~~  114 (114)
T PF03754_consen   98 GTSNYVLNSGWNKVVED  114 (114)
T ss_pred             CceEEEEEcChHhhccC
Confidence            45689996 99999863


No 3  
>PF09217 EcoRII-N:  Restriction endonuclease EcoRII, N-terminal;  InterPro: IPR023372 There are four classes of restriction endonucleases: types I, II,III and IV. All types of enzymes recognise specific short DNA sequences and carry out the endonucleolytic cleavage of DNA to give specific double-stranded fragments with terminal 5'-phosphates. They differ in their recognition sequence, subunit composition, cleavage position, and cofactor requirements [, ], as summarised below:   Type I enzymes (3.1.21.3 from EC) cleave at sites remote from recognition site; require both ATP and S-adenosyl-L-methionine to function; multifunctional protein with both restriction and methylase (2.1.1.72 from EC) activities. Type II enzymes (3.1.21.4 from EC) cleave within or at short specific distances from recognition site; most require magnesium; single function (restriction) enzymes independent of methylase. Type III enzymes (3.1.21.5 from EC) cleave at sites a short distance from recognition site; require ATP (but doesn't hydrolyse it); S-adenosyl-L-methionine stimulates reaction but is not required; exists as part of a complex with a modification methylase methylase (2.1.1.72 from EC). Type IV enzymes target methylated DNA.   Type II restriction endonucleases (3.1.21.4 from EC) are components of prokaryotic DNA restriction-modification mechanisms that protect the organism against invading foreign DNA. These site-specific deoxyribonucleases catalyse the endonucleolytic cleavage of DNA to give specific double-stranded fragments with terminal 5'-phosphates. Of the 3000 restriction endonucleases that have been characterised, most are homodimeric or tetrameric enzymes that cleave target DNA at sequence-specific sites close to the recognition site. For homodimeric enzymes, the recognition site is usually a palindromic sequence 4-8 bp in length. Most enzymes require magnesium ions as a cofactor for catalysis. Although they can vary in their mode of recognition, many restriction endonucleases share a similar structural core comprising four beta-strands and one alpha-helix, as well as a similar mechanism of cleavage, suggesting a common ancestral origin []. However, there is still considerable diversity amongst restriction endonucleases [, ]. The target site recognition process triggers large conformational changes of the enzyme and the target DNA, leading to the activation of the catalytic centres. Like other DNA binding proteins, restriction enzymes are capable of non-specific DNA binding as well, which is the prerequisite for efficient target site location by facilitated diffusion. Non-specific binding usually does not involve interactions with the bases but only with the DNA backbone [].  This entry represents the N-terminal effector-binding domain of the type II restriction endonuclease EcoRII, which has a DNA recognition fold, allowing for binding to 5'-CCWGG sequences. It assumes a structure composed of an eight-stranded beta-sheet with the strands in the order of b2, b5, b4, b3, b7, b6, b1 and b8. They are mostly antiparallel to each other except that b3 is parallel to b7. Alternatively, it may also be viewed as consisting of two mini beta-sheets of four antiparallel beta-strands, sheet I from beta-strands b2, b5, b4, b3 and sheet II from strands b7, b6, b1, b8, folded into an open mixed beta-barrel with a novel topology. Sheet I has a simple Greek key motif while sheet II does not [].  The domain represented by this entry is only found in bacterial proteins.; PDB: 3HQF_A 1NA6_A.
Probab=97.73  E-value=0.00011  Score=68.63  Aligned_cols=89  Identities=21%  Similarity=0.301  Sum_probs=56.5

Q ss_pred             cceEEEecccccCCC----CCcEEeehhhhhhcCCCCCCCCC----ceEEEEeCCC--CeEEeEEEEcCC----CCCCce
Q 012348          356 TPLFEKMLSASDAGR----IGRLVLPKKCAEAYFPPISQPEG----LPLKVQDSKG--KEWIFQFRFWPN----NNSRMY  421 (465)
Q Consensus       356 ~~LF~KvLT~SDVg~----lgRLVIPK~~AEa~FP~L~~~~G----v~L~veD~~G--k~W~Frfrfw~N----nsSR~Y  421 (465)
                      ...|.|.||+.|++.    ..+++|||..++.+||.+...++    +.|.+.+..+  ..|.+||+|..|    +.+..|
T Consensus         7 ~~~~~K~LSaNDtGaTGgHQaGiyIpk~~~~~lFp~~~~~~~~Np~~~~~~~~~s~~~~~~~~r~iYYnn~~~~gTRNE~   86 (156)
T PF09217_consen    7 WAIYCKRLSANDTGATGGHQAGIYIPKSAAELLFPSINHTKEENPDIWLKARWQSHFVTDSQVRFIYYNNRLFGGTRNEY   86 (156)
T ss_dssp             EEEEEEE--CCCCTTTSSS--EEEE-HHHHHHH-GGG-SSSSSS-EEEEEEEETTTT---EEEEEEEE-CCCTTSS--EE
T ss_pred             eEEEEEEccCCCCCCcCcccceeEecccHHHHhCCCCCcccccCCceeEEEEECCCCccceeEEEEEEcccccCCCcCce
Confidence            468999999999994    34799999999999988765433    7788888866  678899999955    346779


Q ss_pred             eecCchhHhhhcC-CCCCCEEEEe
Q 012348          422 VLEGVTPCIQNMQ-LQAGDIGNQP  444 (465)
Q Consensus       422 VLtGW~~FVrsK~-LqaGDtVtF~  444 (465)
                      -||+|+....-.+ =.+||.++|-
T Consensus        87 RIT~~G~~~~~~~~~~tGaL~vla  110 (156)
T PF09217_consen   87 RITRFGRGFPLQNPENTGALLVLA  110 (156)
T ss_dssp             EEE---TTSGGG-GGGTT-EEEEE
T ss_pred             EEeeecCCCccCCccccccEEEEE
Confidence            9999986665333 3689987765


No 4  
>smart00249 PHD PHD zinc finger. The plant homeodomain (PHD) finger is a C4HC3 zinc-finger-like motif found in nuclear proteins thought to be involved in epigenetics and chromatin-mediated transcriptional regulation. The PHD finger binds two zinc ions using the so-called 'cross-brace' motif and is thus structurally related to the PF10844 DUF2577:  Protein of unknown function (DUF2577);  InterPro: IPR022555 This family of proteins has no known function
Probab=66.22  E-value=17  Score=31.36  Aligned_cols=82  Identities=12%  Similarity=0.140  Sum_probs=50.1

Q ss_pred             cccceEEEecccccCC-C-CCcEEeehhhhhhcCCCCCCCCCceEEEEeCCCCeEEeEEEEcCCCCCCceeecCchhHhh
Q 012348          354 VITPLFEKMLSASDAG-R-IGRLVLPKKCAEAYFPPISQPEGLPLKVQDSKGKEWIFQFRFWPNNNSRMYVLEGVTPCIQ  431 (465)
Q Consensus       354 ~~~~LF~KvLT~SDVg-~-lgRLVIPK~~AEa~FP~L~~~~Gv~L~veD~~Gk~W~Frfrfw~NnsSR~YVLtGW~~FVr  431 (465)
                      .....|-++++.+-+. + .++++||++..  +++..-......+.+....+..            ..        .+.-
T Consensus        16 p~~i~~G~V~s~~PL~I~i~~~liL~~~~L--~i~~~l~~~~~~~~~~~~~~~~------------~~--------~i~~   73 (100)
T PF10844_consen   16 PVDIVIGTVVSVPPLKIKIDQKLILDKDFL--IIPELLKDYTRDITIEHNSETD------------NI--------TITF   73 (100)
T ss_pred             CceeEEEEEEecccEEEEECCeEEEchHHE--EeehhccceEEEEEEecccccc------------ce--------eEEE
Confidence            3444899999999854 2 34599998754  5565323333444444332110            00        0445


Q ss_pred             hcCCCCCCEEEEeecCCCCceEEEEEe
Q 012348          432 NMQLQAGDIGNQPKSQESYEMFSFMWK  458 (465)
Q Consensus       432 sK~LqaGDtVtF~R~~ngg~rf~I~~r  458 (465)
                      ...|++||.|...|.+.+ -+|.|=-|
T Consensus        74 ~~~Lk~GD~V~ll~~~~g-Q~yiVlDk   99 (100)
T PF10844_consen   74 TDGLKVGDKVLLLRVQGG-QKYIVLDK   99 (100)
T ss_pred             ecCCcCCCEEEEEEecCC-CEEEEEEe
Confidence            578999999999997766 57776443


No 6  
>PRK03760 hypothetical protein; Provisional
Probab=39.26  E-value=59  Score=29.13  Aligned_cols=29  Identities=21%  Similarity=0.323  Sum_probs=21.3

Q ss_pred             CCCceeecCchhHhhhcCCCCCCEEEEee
Q 012348          417 NSRMYVLEGVTPCIQNMQLQAGDIGNQPK  445 (465)
Q Consensus       417 sSR~YVLtGW~~FVrsK~LqaGDtVtF~R  445 (465)
                      ..-.|||+==.-++.+.++++||.|.|-+
T Consensus        88 ~~a~~VLEl~aG~~~~~gi~~Gd~v~~~~  116 (117)
T PRK03760         88 KPARYIIEGPVGKIRVLKVEVGDEIEWID  116 (117)
T ss_pred             ccceEEEEeCCChHHHcCCCCCCEEEEee
Confidence            35669997222335789999999998866


No 7  
>PF13248 zf-ribbon_3:  zinc-ribbon domain
Probab=38.00  E-value=15  Score=24.60  Aligned_cols=15  Identities=20%  Similarity=0.662  Sum_probs=12.5

Q ss_pred             CCCCccccccCCCce
Q 012348           77 NASGWRCCESCGKRV   91 (465)
Q Consensus        77 ~~sGWR~C~~C~Krl   91 (465)
                      .+.++|-|..||.+|
T Consensus        12 ~~~~~~fC~~CG~~L   26 (26)
T PF13248_consen   12 IDPDAKFCPNCGAKL   26 (26)
T ss_pred             CCcccccChhhCCCC
Confidence            477899999999875


No 8  
>PF08922 DUF1905:  Domain of unknown function (DUF1905);  InterPro: IPR015018 This family consist of hypothetical bacterial proteins. ; PDB: 2D9R_A.
Probab=36.61  E-value=1e+02  Score=25.63  Aligned_cols=79  Identities=19%  Similarity=0.257  Sum_probs=39.6

Q ss_pred             EEEecccccCCCCCcEEeehhhhhhcCCCCCCCCCceEEEEeCCCCeEEeEEEEcCCCCCCceeecCchhHhhhcCCCCC
Q 012348          359 FEKMLSASDAGRIGRLVLPKKCAEAYFPPISQPEGLPLKVQDSKGKEWIFQFRFWPNNNSRMYVLEGVTPCIQNMQLQAG  438 (465)
Q Consensus       359 F~KvLT~SDVg~lgRLVIPK~~AEa~FP~L~~~~Gv~L~veD~~Gk~W~Frfrfw~NnsSR~YVLtGW~~FVrsK~LqaG  438 (465)
                      |+..|-+.+-+ -..+.||.+-++++-..  +-..+.+.++ ..|.+|+=  +..+ ...+.|+|-==.+.-++-++.+|
T Consensus         1 F~a~l~~~~~~-~~fv~vP~~v~~~l~~~--~~g~v~V~~t-I~g~~~~~--sl~p-~g~G~~~Lpv~~~vRk~~g~~~G   73 (80)
T PF08922_consen    1 FTATLWKGEGG-WTFVEVPFDVAEELGEG--GWGRVPVRGT-IDGHPWRT--SLFP-MGNGGYILPVKAAVRKAIGKEAG   73 (80)
T ss_dssp             EEEE-EE-TTS--EEEE--S-HHHHH--S----S-EEEEEE-ETTEEEEE--EEEE-SSTT-EEEEE-HHHHHHHT--TT
T ss_pred             CeEEEEecCCc-eEEEEeCHHHHHHhccc--cCCceEEEEE-ECCEEEEE--EEEE-CCCCCEEEEEcHHHHHHcCCCCC
Confidence            45555555443 23577899888876332  1123555554 36666655  5555 23466777423466788899999


Q ss_pred             CEEEEe
Q 012348          439 DIGNQP  444 (465)
Q Consensus       439 DtVtF~  444 (465)
                      |+|.+.
T Consensus        74 d~V~v~   79 (80)
T PF08922_consen   74 DTVEVT   79 (80)
T ss_dssp             SEEEEE
T ss_pred             CEEEEE
Confidence            999863


No 9  
>PF02643 DUF192:  Uncharacterized ACR, COG1430;  InterPro: IPR003795 This entry describes proteins of unknown function.; PDB: 3M7A_B 3PJY_B.
Probab=36.18  E-value=82  Score=27.42  Aligned_cols=52  Identities=21%  Similarity=0.200  Sum_probs=27.6

Q ss_pred             CceEEEEeCCCCeEEeEEEEcCC-------CCCCceeecCchhHhhhcCCCCCCEEEEe
Q 012348          393 GLPLKVQDSKGKEWIFQFRFWPN-------NNSRMYVLEGVTPCIQNMQLQAGDIGNQP  444 (465)
Q Consensus       393 Gv~L~veD~~Gk~W~Frfrfw~N-------nsSR~YVLtGW~~FVrsK~LqaGDtVtF~  444 (465)
                      .+.|.+.|..|+.=....-..|.       ...-.|||+-=..++..+++++||.|.|-
T Consensus        49 pLDi~fld~~g~Vv~i~~~~~P~~~~~~~~~~~a~~vLE~~aG~~~~~~i~~Gd~v~~~  107 (108)
T PF02643_consen   49 PLDIAFLDSDGRVVKIERMVPPWRTYPCPSYKPARYVLELPAGWFEKLGIKVGDRVRIE  107 (108)
T ss_dssp             -EEEEEE-TTSBEEEEEEEE-TT--S-EEECCEECEEEEEETTHHHHHT--TT-EEE--
T ss_pred             eEEEEEECCCCeEEEEEccCCCCccCCCCCCCccCEEEEcCCCchhhcCCCCCCEEEec
Confidence            35666677766654443322111       12357999844556789999999999873


No 10 
>KOG4718 consensus Non-SMC (structural maintenance of chromosomes) element 1 protein (NSE1) [Chromatin structure and dynamics]
Probab=33.42  E-value=16  Score=36.72  Aligned_cols=18  Identities=39%  Similarity=0.829  Sum_probs=15.5

Q ss_pred             cccccCCCceeechhhhh
Q 012348           82 RCCESCGKRVHCGCITSV   99 (465)
Q Consensus        82 R~C~~C~KrlHCGCI~S~   99 (465)
                      +.|.+||-|.|||||.--
T Consensus       195 ~rCg~c~i~~h~~c~qty  212 (235)
T KOG4718|consen  195 IRCGSCNIQYHRGCIQTY  212 (235)
T ss_pred             eccCcccchhhhHHHHHH
Confidence            568899999999999753


No 11 
>PF04014 Antitoxin-MazE:  Antidote-toxin recognition MazE;  InterPro: IPR007159 This domain is found in AbrB from Bacillus subtilis. The product of the abrB gene is an ambiactive repressor and activator of the transcription of genes expressed during the transition state between vegetative growth and the onset of stationary phase and sporulation []. AbrB is thought to interact directly with the transcription initiation regions of genes under its control []. AbrB contains a helix-turn-helix structure, but this domain ends before the helix-turn-helix begins []. The product of the B. subtilis gene spoVT is another member of this family and is also a transcriptional regulator []. DNA-binding activity in this AbrB homologue requires hexamerisation []. Another family member has been isolated from the Sulfolobus solfataricus and has been identified as a homologue of bacterial repressor-like proteins. The Escherichia coli family member SohA or Prl1F appears to be bifunctional and is able to regulate its own expression as well as relieve the export block imposed by high-level synthesis of beta-galactosidase hybrid proteins [].; PDB: 2L66_A 2GLW_A 3TND_D 2W1T_B 2RO5_B 2FY9_A 2RO3_B 1UB4_C 3ZVK_G 1YFB_B ....
Probab=30.51  E-value=57  Score=24.22  Aligned_cols=23  Identities=13%  Similarity=0.156  Sum_probs=19.7

Q ss_pred             hhHhhhcCCCCCCEEEEeecCCC
Q 012348          427 TPCIQNMQLQAGDIGNQPKSQES  449 (465)
Q Consensus       427 ~~FVrsK~LqaGDtVtF~R~~ng  449 (465)
                      .++.+..+|++||.|.|.-..++
T Consensus        13 k~~~~~l~l~~Gd~v~i~~~~~g   35 (47)
T PF04014_consen   13 KEIREKLGLKPGDEVEIEVEGDG   35 (47)
T ss_dssp             HHHHHHTTSSTTTEEEEEEETTS
T ss_pred             HHHHHHcCCCCCCEEEEEEeCCC
Confidence            36778889999999999988775


No 12 
>TIGR01643 YD_repeat_2x YD repeat (two copies). This model describes two tandem copies of a 21-residue extracellular repeat found in Gram-negative, Gram-positive, and animal proteins. The repeat is named for a YD dipeptide, the most strongly conserved motif of the repeat. These repeats appear in general to be involved in binding carbohydrate; the chicken teneurin-1 YD-repeat region has been shown to bind heparin.
Probab=30.24  E-value=76  Score=22.36  Aligned_cols=22  Identities=14%  Similarity=0.149  Sum_probs=18.1

Q ss_pred             CCceEEEEeCCCCeEEeEEEEc
Q 012348          392 EGLPLKVQDSKGKEWIFQFRFW  413 (465)
Q Consensus       392 ~Gv~L~veD~~Gk~W~Frfrfw  413 (465)
                      .|..+.+.|..|..|+|.|--.
T Consensus         4 ~g~l~~~~~p~G~~~~~~YD~~   25 (42)
T TIGR01643         4 AGRLTGSTDADGTTTRYTYDAA   25 (42)
T ss_pred             CCCEEEEECCCCCEEEEEECCC
Confidence            4778899999999999987643


No 13 
>COG1998 RPS31 Ribosomal protein S27AE [Translation, ribosomal structure and biogenesis]
Probab=29.53  E-value=27  Score=27.84  Aligned_cols=13  Identities=54%  Similarity=1.349  Sum_probs=8.7

Q ss_pred             cccccCC---------Cceeec
Q 012348           82 RCCESCG---------KRVHCG   94 (465)
Q Consensus        82 R~C~~C~---------KrlHCG   94 (465)
                      |.|..||         +|+|||
T Consensus        20 ~~CPrCG~gvfmA~H~dR~~CG   41 (51)
T COG1998          20 RFCPRCGPGVFMADHKDRWACG   41 (51)
T ss_pred             ccCCCCCCcchhhhcCceeEec
Confidence            4566666         488887


No 14 
>PF03120 DNA_ligase_OB:  NAD-dependent DNA ligase OB-fold domain;  InterPro: IPR004150 DNA ligases catalyse the crucial step of joining the breaks in duplex DNA during DNA replication, repair and recombination, utilizing either ATP or NAD(+) as a cofactor []. This family is a small domain found after the adenylation domain DNA_ligase_N in NAD+-dependent ligases (IPR001679 from INTERPRO). OB-fold domains generally are involved in nucleic acid binding.; GO: 0003911 DNA ligase (NAD+) activity, 0006260 DNA replication, 0006281 DNA repair; PDB: 2OWO_A 1TAE_A 3UQ8_A 1DGS_A 1V9P_B 3SGI_A.
Probab=29.37  E-value=41  Score=28.81  Aligned_cols=22  Identities=14%  Similarity=0.290  Sum_probs=17.7

Q ss_pred             hhHhhhcCCCCCCEEEEeecCC
Q 012348          427 TPCIQNMQLQAGDIGNQPKSQE  448 (465)
Q Consensus       427 ~~FVrsK~LqaGDtVtF~R~~n  448 (465)
                      .+|+++++|..||.|.++|.-+
T Consensus        42 ~~~i~~~~i~~Gd~V~V~raGd   63 (82)
T PF03120_consen   42 YDYIKELDIRIGDTVLVTRAGD   63 (82)
T ss_dssp             HHHHHHTT-BBT-EEEEEEETT
T ss_pred             HHHHHHcCCCCCCEEEEEECCC
Confidence            6899999999999999999654


No 15 
>PF09297 zf-NADH-PPase:  NADH pyrophosphatase zinc ribbon domain;  InterPro: IPR015376 This domain has a zinc ribbon structure and is often found between two NUDIX domains.; GO: 0016787 hydrolase activity, 0046872 metal ion binding; PDB: 1VK6_A 2GB5_A.
Probab=27.55  E-value=31  Score=23.97  Aligned_cols=25  Identities=32%  Similarity=0.800  Sum_probs=13.3

Q ss_pred             hh-ccccccccccccccccCCCCCc-cccccCCCce
Q 012348           58 LC-VYRSIYEEGRFCDTFHVNASGW-RCCESCGKRV   91 (465)
Q Consensus        58 LC-~CgsayE~~~FCd~FH~~~sGW-R~C~~C~Krl   91 (465)
                      -| +||+.-+         ..+.|| |-|.+||...
T Consensus         5 fC~~CG~~t~---------~~~~g~~r~C~~Cg~~~   31 (32)
T PF09297_consen    5 FCGRCGAPTK---------PAPGGWARRCPSCGHEH   31 (32)
T ss_dssp             B-TTT--BEE---------E-SSSS-EEESSSS-EE
T ss_pred             ccCcCCcccc---------CCCCcCEeECCCCcCEe
Confidence            36 7777543         345677 6799998753


No 16 
>cd05829 Sortase_E Sortase E (SrtE) is a membrane transpeptidase found in gram-positive bacteria that cleaves surface proteins at a cell sorting motif and catalyzes a transpeptidation reaction in which the surface protein substrate is covalently linked to peptidoglycan for display on the bacterial surface. Sortases are grouped into different classes and subfamilies based on sequence, membrane topology, genomic positioning, and cleavage site preference. The function of Sortase E is unknown. In two different sortase families, the N-terminus either functions as both a signal peptide for secretion and a stop-transfer signal for membrane anchoring, or it contains a signal peptide only and the C-terminus serves as a membrane anchor. Most gram-positive bacteria contain more than one sortase and it is thought that the different sortases anchor different surface protein classes. The sortase domain is a modified beta-barrel flanked by two (SrtA) or three (SrtB) short alpha-helices.
Probab=27.16  E-value=1.2e+02  Score=27.64  Aligned_cols=39  Identities=15%  Similarity=0.112  Sum_probs=26.1

Q ss_pred             CCceeec---Cch----hHhhhcCCCCCCEEEEeecCCCCceEEEE
Q 012348          418 SRMYVLE---GVT----PCIQNMQLQAGDIGNQPKSQESYEMFSFM  456 (465)
Q Consensus       418 SR~YVLt---GW~----~FVrsK~LqaGDtVtF~R~~ngg~rf~I~  456 (465)
                      .+.++|.   +|.    .|-+=.+|++||.|.+.........|.+.
T Consensus        49 ~Gn~viaGH~~~~g~~~~F~~L~~l~~GD~I~v~~~~g~~~~Y~V~   94 (144)
T cd05829          49 KGTAVLAGHVDSRGGPAVFFRLGDLRKGDKVEVTRADGQTATFRVD   94 (144)
T ss_pred             CCCEEEEEecCCCCCChhhcchhcCCCCCEEEEEECCCCEEEEEEe
Confidence            4566773   332    38888999999999998854433444443


No 17 
>PF12760 Zn_Tnp_IS1595:  Transposase zinc-ribbon domain;  InterPro: IPR024442 This zinc binding domain is found in a range of transposase proteins such as ISSPO8, ISSOD11, ISRSSP2 etc. It may be a zinc-binding beta ribbon domain that could bind DNA.
Probab=24.46  E-value=41  Score=25.10  Aligned_cols=26  Identities=23%  Similarity=0.645  Sum_probs=17.9

Q ss_pred             hh-ccccccccccccccccCCCCCccccccCCCc
Q 012348           58 LC-VYRSIYEEGRFCDTFHVNASGWRCCESCGKR   90 (465)
Q Consensus        58 LC-~CgsayE~~~FCd~FH~~~sGWR~C~~C~Kr   90 (465)
                      -| +||+.       +.+.....+-..|..|+|+
T Consensus        20 ~CP~Cg~~-------~~~~~~~~~~~~C~~C~~q   46 (46)
T PF12760_consen   20 VCPHCGST-------KHYRLKTRGRYRCKACRKQ   46 (46)
T ss_pred             CCCCCCCe-------eeEEeCCCCeEECCCCCCc
Confidence            48 99985       2233344788899999875


No 18 
>cd04459 Rho_CSD Rho_CSD: Rho protein cold-shock domain (CSD). Rho protein is a transcription termination factor in most bacteria. In bacteria, there are two distinct mechanisms for mRNA transcription termination. In intrinsic termination, RNA polymerase and nascent mRNA are released from DNA template by an mRNA stem loop structure, which resembles the transcription termination mechanism used by eukaryotic pol III. The second mechanism is mediated by Rho factor. Rho factor terminates transcription by using energy from ATP hydrolysis to forcibly dissociate the transcripts from RNA polymerase. Rho protein contains an N-terminal S1-like domain, which binds single-stranded RNA. Rho has a C-terminal ATPase domain which hydrolyzes ATP to provide energy to strip RNA polymerase and mRNA from the DNA template. Rho functions as a homohexamer.
Probab=23.37  E-value=71  Score=26.39  Aligned_cols=26  Identities=19%  Similarity=0.272  Sum_probs=19.0

Q ss_pred             hHhhhcCCCCCCEEEE-eecCCCCceE
Q 012348          428 PCIQNMQLQAGDIGNQ-PKSQESYEMF  453 (465)
Q Consensus       428 ~FVrsK~LqaGDtVtF-~R~~ngg~rf  453 (465)
                      .-||..+|+.||.|.= -|...++++|
T Consensus        34 ~~Irr~~LR~GD~V~G~vr~p~~~ek~   60 (68)
T cd04459          34 SQIRRFNLRTGDTVVGQIRPPKEGERY   60 (68)
T ss_pred             HHHHHhCCCCCCEEEEEEeCCCCCCCc
Confidence            5789999999999874 4544444554


No 19 
>TIGR03784 marine_sortase sortase, marine proteobacterial type. Members of this protein family are sortase enzymes, cysteine transpeptidases involved in protein sorting activities. Members of this family tend to be found in proteobacteria, rather than in Gram-positive bacteria where sortases attach proteins to the Gram-positive cell wall or participate in pilin cross-linking. Many species with this sortase appear to contain a signal target sequence, a protein with a Vault protein inter-alpha-trypsin domain (pfam08487) and a von Willebrand factor type A domain (pfam00092), encoded by an adjacent gene. These sortases are designated subfamily 6 according to Comfort and Clubb (2004).
Probab=23.30  E-value=84  Score=29.90  Aligned_cols=81  Identities=15%  Similarity=0.149  Sum_probs=47.0

Q ss_pred             CCcEEeehhhhhhcCCCCCCCCCceEEEEeCCCCeEEeEEEEcCCCCCCceeec--CchhHhhhcCCCCCCEEEEeecCC
Q 012348          371 IGRLVLPKKCAEAYFPPISQPEGLPLKVQDSKGKEWIFQFRFWPNNNSRMYVLE--GVTPCIQNMQLQAGDIGNQPKSQE  448 (465)
Q Consensus       371 lgRLVIPK~~AEa~FP~L~~~~Gv~L~veD~~Gk~W~Frfrfw~NnsSR~YVLt--GW~~FVrsK~LqaGDtVtF~R~~n  448 (465)
                      .+||.||+--.+  +|-+.+..+..|.+    | .-.+..+-.| +..+.++|.  .-+.|-+=.+|+.||.|.+.....
T Consensus        51 va~L~IP~lg~~--~~V~~G~s~~~L~~----G-~G~~~~t~~P-G~~Gn~VIAGHrdt~F~~L~~L~~GD~I~v~~~~g  122 (174)
T TIGR03784        51 VAKLSAPRLGAS--LYVLAGASGRNLAF----G-PGHMLATAQP-GAQGNSVIAGHRDTHFAFLQELRPGDVIRLQTPDG  122 (174)
T ss_pred             eEEEEEccCCCc--eeEEecCCHHHHhh----e-eEEecCCCCC-CCCCcEEEEeeCCccCCChhhCCCCCEEEEEECCC
Confidence            579999985432  45554444333321    1 1112222233 234677884  334699999999999999987655


Q ss_pred             CCceEEEEEee
Q 012348          449 SYEMFSFMWKV  459 (465)
Q Consensus       449 gg~rf~I~~rr  459 (465)
                      ...+|.+.-.+
T Consensus       123 ~~~~Y~V~~~~  133 (174)
T TIGR03784       123 QWQSYQVTATR  133 (174)
T ss_pred             eEEEEEEeEEE
Confidence            43456665443


No 20 
>TIGR00375 conserved hypothetical protein TIGR00375. The member of this family from Methanococcus jannaschii, MJ0043, is considerably longer and appears to contain an intein N-terminal to the region of homology.
Probab=22.83  E-value=36  Score=36.24  Aligned_cols=44  Identities=23%  Similarity=0.330  Sum_probs=29.5

Q ss_pred             HHHHhh-ccccccccccccccccCCCCCccccccCCCceeechhhhhhhhhhhhc
Q 012348           54 LTLILC-VYRSIYEEGRFCDTFHVNASGWRCCESCGKRVHCGCITSVHAFTLLDA  107 (465)
Q Consensus        54 ~~~~LC-~CgsayE~~~FCd~FH~~~sGWR~C~~C~KrlHCGCI~S~~~~~lLD~  107 (465)
                      |+-+-| +|+.-|+..   |   ..+-||+ |. |||+|-=|  |.--.-||=|.
T Consensus       238 Yh~~~c~~C~~~~~~~---~---~~~~~~~-Cp-CG~~i~~G--V~~Rv~eLad~  282 (374)
T TIGR00375       238 YHQTACEACGEPAVSE---D---AETACAN-CP-CGGRIKKG--VSDRLRELSDQ  282 (374)
T ss_pred             cchhhhcccCCcCCch---h---hhhcCCC-CC-CCCcceec--hHHHHHHHhcC
Confidence            678889 998877643   1   2233788 88 99998766  45555566563


No 21 
>PRK09570 rpoH DNA-directed RNA polymerase subunit H; Reviewed
Probab=22.29  E-value=1e+02  Score=26.37  Aligned_cols=32  Identities=9%  Similarity=0.262  Sum_probs=23.8

Q ss_pred             hhHhhhcCCCCCCEEEEeecCCCCceEEEEEee
Q 012348          427 TPCIQNMQLQAGDIGNQPKSQESYEMFSFMWKV  459 (465)
Q Consensus       427 ~~FVrsK~LqaGDtVtF~R~~ngg~rf~I~~rr  459 (465)
                      -+.++.-+++.||+|-+.|..... .-.+.||.
T Consensus        44 DPv~r~~g~k~GdVvkI~R~S~ta-G~~v~YR~   75 (79)
T PRK09570         44 DPVVKAIGAKPGDVIKIVRKSPTA-GEAVYYRL   75 (79)
T ss_pred             ChhhhhcCCCCCCEEEEEECCCCC-CccEEEEE
Confidence            367888899999999999986542 22566664


No 22 
>PF01191 RNA_pol_Rpb5_C:  RNA polymerase Rpb5, C-terminal domain;  InterPro: IPR000783  Prokaryotes contain a single DNA-dependent RNA polymerase (RNAP; 2.7.7.6 from EC) that is responsible for the transcription of all genes, while eukaryotes have three classes of RNAPs (I-III) that transcribe different sets of genes. Each class of RNA polymerase is an assemblage of ten to twelve different polypeptides. Certain subunits of RNAPs, including RPB5 (POLR2E in mammals), are common to all three eukaryotic polymerases. RPB5 plays a role in the transcription activation process. Eukaryotic RPB5 has a bipartite structure consisting of a unique N-terminal region (IPR005571 from INTERPRO), plus a C-terminal region that is structurally homologous to the prokaryotic RPB5 homologue, subunit H (gene rpoH) [, , , ]. This entry represents prokaryotic subunit H and the C-terminal domain of eukaryotic RPB5, which share a two-layer alpha/beta fold, with a core structure of beta/alpha/beta/alpha/beta(2). ; GO: 0003677 DNA binding, 0003899 DNA-directed RNA polymerase activity, 0006351 transcription, DNA-dependent; PDB: 1EIK_A 2Y0S_Z 1DZF_A 3GTG_E 2VUM_E 3GTP_E 3GTO_E 3S17_E 3S1R_E 1I3Q_E ....
Probab=22.29  E-value=1.2e+02  Score=25.57  Aligned_cols=32  Identities=13%  Similarity=0.225  Sum_probs=23.0

Q ss_pred             hhHhhhcCCCCCCEEEEeecCCCCceEEEEEee
Q 012348          427 TPCIQNMQLQAGDIGNQPKSQESYEMFSFMWKV  459 (465)
Q Consensus       427 ~~FVrsK~LqaGDtVtF~R~~ngg~rf~I~~rr  459 (465)
                      .+.++..+++.||+|-+.|..... .-.+.||.
T Consensus        41 DPv~r~~g~k~GdVvkI~R~S~ta-G~~v~YR~   72 (74)
T PF01191_consen   41 DPVARYLGAKPGDVVKIIRKSETA-GEYVTYRL   72 (74)
T ss_dssp             SHHHHHTT--TTSEEEEEEEETTT-SEEEEEEE
T ss_pred             ChhhhhcCCCCCCEEEEEecCCCC-CCcEEEEE
Confidence            478888999999999999977652 34666663


No 23 
>PF05593 RHS_repeat:  RHS Repeat;  InterPro: IPR006530 These sequences contain two tandem copies of a 21-residue extracellular repeat that is found in Gram-negative, Gram-positive, and animal proteins. The repeat is named for a YD dipeptide, the most strongly conserved motif of the repeat. These repeats appear in general to be involved in binding carbohydrate; the chicken teneurin-1 YD-repeat region has been shown to bind heparin [, , ].
Probab=22.18  E-value=1.4e+02  Score=21.29  Aligned_cols=23  Identities=17%  Similarity=0.268  Sum_probs=18.1

Q ss_pred             CCCceEEEEeCCCCeEEeEEEEc
Q 012348          391 PEGLPLKVQDSKGKEWIFQFRFW  413 (465)
Q Consensus       391 ~~Gv~L~veD~~Gk~W~Frfrfw  413 (465)
                      ..|..+.+.|..|.+|+|.|--.
T Consensus         3 ~~G~l~~~~d~~G~~~~y~YD~~   25 (38)
T PF05593_consen    3 ANGRLTSVTDPDGRTTRYTYDAA   25 (38)
T ss_pred             CCCCEEEEEcCCCCEEEEEECCC
Confidence            35778889999999998776544


Done!