Query         039930
Match_columns 245
No_of_seqs    116 out of 169
Neff          3.7 
Searched_HMMs 46136
Date          Fri Mar 29 03:27:33 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/039930.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/039930hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF06884 DUF1264:  Protein of u 100.0  1E-102  2E-107  667.0  14.3  171   40-212     1-171 (171)
  2 PF04934 Med6:  MED6 mediator s  37.4      12 0.00025   31.6   0.1   36   89-131    58-93  (140)
  3 PF01435 Peptidase_M48:  Peptid  36.0      18  0.0004   30.2   1.1   36   79-114    51-90  (226)
  4 smart00136 LamNT Laminin N-ter  31.8      15 0.00033   33.4  -0.1   79   69-158    33-131 (238)
  5 TIGR00787 dctP tripartite ATP-  31.5      30 0.00064   30.4   1.7   22   94-115   205-226 (257)
  6 PF09022 Staphostatin_A:  Staph  30.9      56  0.0012   27.0   3.0   33   65-100    47-79  (105)
  7 PF01082 Cu2_monooxygen:  Coppe  29.9      55  0.0012   26.7   2.9   73   75-159    16-94  (132)
  8 PF00505 HMG_box:  HMG (high mo  29.8      17 0.00037   25.3  -0.1   16  102-117    36-51  (69)
  9 smart00398 HMG high mobility g  27.8      30 0.00066   23.7   0.9   15  103-117    38-52  (70)
 10 cd01390 HMGB-UBF_HMG-box HMGB-  25.8      35 0.00076   23.4   0.9   15  103-117    37-51  (66)
 11 cd00084 HMG-box High Mobility   25.1      37  0.0008   22.9   1.0   15  103-117    37-51  (66)
 12 COG0518 GuaA GMP synthase - Gl  23.9      32  0.0007   30.3   0.6   52  103-155   121-175 (198)
 13 PF04512 Baculo_PEP_N:  Baculov  23.1      42 0.00092   27.1   1.1   37   92-128    15-56  (97)
 14 PF03480 SBP_bac_7:  Bacterial   22.9      58  0.0013   28.9   2.0   27   94-120   205-231 (286)
 15 PF01221 Dynein_light:  Dynein   22.9      42 0.00091   25.5   1.0   13  144-156    45-57  (89)
 16 PF07277 SapC:  SapC;  InterPro  21.3 1.4E+02  0.0031   26.6   4.2   80   19-118   122-201 (221)
 17 COG1638 DctP TRAP-type C4-dica  20.7      62  0.0013   30.7   1.8   23   94-116   236-258 (332)

No 1  
>PF06884 DUF1264:  Protein of unknown function (DUF1264);  InterPro: IPR010686 This family contains a number of bacterial and eukaryotic proteins of unknown function that are approximately 200 residues long. Some family members are annotated as putative lipoproteins.
Probab=100.00  E-value=1.1e-102  Score=667.04  Aligned_cols=171  Identities=56%  Similarity=1.060  Sum_probs=169.7

Q ss_pred             CCchhhhhhhhceeeeeecCCCCCceeeeeeeeeccCceeeeEeeeCCCCCcceeeeeeeecHHHHhhCChhhhhccccc
Q 039930           40 SLKPIKQMSQHVCTFALYSHDMPRQIETHHYVSRLNQDFLQCAVYDSDDSTARLIGVEYIVSDKIFEALPLEEQKLWHSH  119 (245)
Q Consensus        40 ~~~Pi~~i~~hL~aFH~ya~d~~rqvEAhHYC~~lneD~~QCliYDs~~~~ArLIGIEYiISe~lf~tLP~eEkklWHsH  119 (245)
                      +|+||++||+||||||+|++||+|||||||||+|||+||+||+|||||++|||||||||||||+||+|||+|||||||||
T Consensus         1 nf~Pv~~i~~hl~~fH~y~~d~~rqveahHyC~~vned~~QC~iyDs~~~~ArLIGvEYiISe~lf~tLP~eEkklWHsH   80 (171)
T PF06884_consen    1 NFKPVKQICQHLCAFHFYADDPTRQVEAHHYCSHVNEDFRQCLIYDSNEPNARLIGVEYIISEKLFETLPEEEKKLWHSH   80 (171)
T ss_pred             CCCchHHHHhHhhhhhcccCCCCceeeeeeeeeecCCcceEEEEecCCCCCcceeeeeEEEcHHHHhhCCHHHHhccCCc
Confidence            69999999999999999999999999999999999999999999999999999999999999999999999999999999


Q ss_pred             ceeeccceeecCCCCccccchHHHHHhhccCceEEeeccCCCCCCCCCCCccccccCcCCCCCCChhHHHhhhhhcCCCh
Q 039930          120 AYEVKSGLWVNPRVPEMIGRPELDNMAKTYGKFWCTWQVDRGDRLPLGAPALMMSPQAVNLGMVRPDLVQKRDDKYNIST  199 (245)
Q Consensus       120 ~yEVkSG~Lv~P~vP~~aE~~~M~~l~~tYGKt~HtWq~Drgd~LPlG~P~LMmsft~d~~gq~~~~lv~~RD~r~gi~t  199 (245)
                      +|||+||+|++|+||+++|+++|++|++|||||||||||||||+||||+||||||||+|  ||++++||++||+||||||
T Consensus        81 ~~EVkSG~L~~p~vP~~ae~~~m~~l~~tYGKt~HtWq~Drgd~LPlG~P~LM~sft~d--gq~~~~lv~~RD~r~gv~t  158 (171)
T PF06884_consen   81 VYEVKSGMLVMPGVPEAAEKAEMEKLVKTYGKTWHTWQVDRGDKLPLGPPQLMMSFTRD--GQVDPELVKERDERFGVDT  158 (171)
T ss_pred             ceeeeeeeEecCCCCHHHHHHHHHHHHhhhCCeEEeccCCCCCCCCCCCCeeccccCCc--ccCCHHHHHHHHHhcCCCH
Confidence            99999999999999999999999999999999999999999999999999999999999  9999999999999999999


Q ss_pred             HHHHhhccCCCCC
Q 039930          200 DALKQSRVEIDEP  212 (245)
Q Consensus       200 ~~kr~~R~~i~~p  212 (245)
                      ++||++|++|++|
T Consensus       159 ~~kr~~R~~i~~~  171 (171)
T PF06884_consen  159 EEKRESRADIPEP  171 (171)
T ss_pred             HHHHHHhcccccC
Confidence            9999999999986


No 2  
>PF04934 Med6:  MED6 mediator sub complex component;  InterPro: IPR007018 The Mediator complex is a coactivator involved in the regulated transcription of nearly all RNA polymerase II-dependent genes. Mediator functions as a bridge to convey information from gene-specific regulatory proteins to the basal RNA polymerase II transcription machinery. The Mediator complex, having a compact conformation in its free form, is recruited to promoters by direct interactions with regulatory proteins and serves for the assembly of a functional preinitiation complex with RNA polymerase II and the general transcription factors. On recruitment the Mediator complex unfolds to an extended conformation and partially surrounds RNA polymerase II, specifically interacting with the unphosphorylated form of the C-terminal domain (CTD) of RNA polymerase II. The Mediator complex dissociates from the RNA polymerase II holoenzyme and stays at the promoter when transcriptional elongation begins.  The Mediator complex is composed of at least 31 subunits: MED1, MED4, MED6, MED7, MED8, MED9, MED10, MED11, MED12, MED13, MED13L, MED14, MED15, MED16, MED17, MED18, MED19, MED20, MED21, MED22, MED23, MED24, MED25, MED26, MED27, MED29, MED30, MED31, CCNC, CDK8 and CDC2L6/CDK11.  The subunits form at least three structurally distinct submodules. The head and the middle modules interact directly with RNA polymerase II, whereas the elongated tail module interacts with gene-specific regulatory proteins. Mediator containing the CDK8 module is less active than Mediator lacking this module in supporting transcriptional activation.   The head module contains: MED6, MED8, MED11, SRB4/MED17, SRB5/MED18, ROX3/MED19, SRB2/MED20 and SRB6/MED22.  The middle module contains: MED1, MED4, NUT1/MED5, MED7, CSE2/MED9, NUT2/MED10, SRB7/MED21 and SOH1/MED31. CSE2/MED9 interacts directly with MED4.  The tail module contains: MED2, PGD1/MED3, RGR1/MED14, GAL11/MED15 and SIN4/MED16.  The CDK8 module contains: MED12, MED13, CCNC and CDK8.   Individual preparations of the Mediator complex lacking one or more distinct subunits have been variously termed ARC, CRSP, DRIP, PC2, SMCC and TRAP. Regulation of mRNA synthesis requires intermediary proteins that transduce regulatory signals from upstream transcriptional activator proteins to basal transcription machinery at the core promoter. Three types of intermediary factors that enable the basal transcription machinery to respond to transcriptional activator proteins bound to regulatory DNA sequences have been identified: (i) TAFIIs, which associate with TATA-binding protein (TBP) to form TFIID; (ii) mediator, which associates with RNA polymerase II to form a holo-polymerase; and (iii) coactivators such as human upstream stimulatory activity (USA), mammalian CBP/P300, yeast ADA complex, and HMG proteins. The interaction of these multiprotein complexes with activators and general transcription factors is essential for transcriptional regulation.  This family of proteins represent the transcriptional mediator protein subunit 6 that is required for activation of many RNA polymerase II promoters and which are conserved from yeast to humans []..; GO: 0001104 RNA polymerase II transcription cofactor activity, 0006357 regulation of transcription from RNA polymerase II promoter, 0016592 mediator complex; PDB: 3RJ1_N.
Probab=37.38  E-value=12  Score=31.57  Aligned_cols=36  Identities=22%  Similarity=0.326  Sum_probs=2.2

Q ss_pred             CCcceeeeeeeecHHHHhhCChhhhhcccccceeeccceeecC
Q 039930           89 STARLIGVEYIVSDKIFEALPLEEQKLWHSHAYEVKSGLWVNP  131 (245)
Q Consensus        89 ~~ArLIGIEYiISe~lf~tLP~eEkklWHsH~yEVkSG~Lv~P  131 (245)
                      .-.++.||||+|...       .|--+|..++.+..|+.-+.|
T Consensus        58 ~L~~m~GiEY~l~~~-------~eP~l~vI~Kq~r~~~~~~~~   93 (140)
T PF04934_consen   58 QLRNMKGIEYVLAHV-------QEPGLFVIRKQRRQSPDEVTP   93 (140)
T ss_dssp             TTTS---------------------------------------
T ss_pred             HHhhcCCeEEEEecc-------CCCCEEEEEEeeccCCCcceE
Confidence            346788999988654       455899999999988776665


No 3  
>PF01435 Peptidase_M48:  Peptidase family M48 This is family M48 in the peptidase classification. ;  InterPro: IPR001915 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Metalloproteases are the most diverse of the four main types of protease, with more than 50 families identified to date. In these enzymes, a divalent cation, usually zinc, activates the water molecule. The metal ion is held in place by amino acid ligands, usually three in number. The known metal ligands are His, Glu, Asp or Lys and at least one other residue is required for catalysis, which may play an electrophillic role. Of the known metalloproteases, around half contain an HEXXH motif, which has been shown in crystallographic studies to form part of the metal-binding site []. The HEXXH motif is relatively common, but can be more stringently defined for metalloproteases as 'abXHEbbHbc', where 'a' is most often valine or threonine and forms part of the S1' subsite in thermolysin and neprilysin, 'b' is an uncharged residue, and 'c' a hydrophobic residue. Proline is never found in this site, possibly because it would break the helical structure adopted by this motif in metalloproteases []. This group of metallopeptidases belong to MEROPS peptidase family M48 (Ste24 endopeptidase family, clan M-); members of both subfamily are represented. The members of this set of proteins are mostly described as probable protease htpX homologue (3.4.24 from EC) or CAAX prenyl protease 1, which proteolytically removes the C-terminal three residues of farnesylated proteins. They are integral membrane proteins associated with the endoplasmic reticulum and Golgi, binding one zinc ion per subunit. In Saccharomyces cerevisiae (Baker's yeast) Ste24p is required for the first NH2-terminal proteolytic processing event within the a-factor precursor, which takes place after COOH-terminal CAAX modification is complete. The Ste24p contains multiple predicted membrane spans, a zinc metalloprotease motif (HEXXH), and a COOH-terminal ER retrieval signal (KKXX). The HEXXH protease motif is critical for Ste24p activity, since Ste24p fails to function when conserved residues within this motif are mutated.  The Ste24p homologues occur in a diverse group of organisms, including Escherichia coli, Schizosaccharomyces pombe (Fission yeast), Haemophilus influenzae, and Homo sapiens (Human), which indicates that the gene is highly conserved throughout evolution. Ste24p and the proteins related to it define a subfamily of proteins that are likely to function as intracellular, membrane-associated zinc metalloproteases [].  HtpX is a zinc-dependent endoprotease member of the membrane-localized proteolytic system in E. coli, which participates in the proteolytic quality control of membrane proteins in conjunction with FtsH, a membrane-bound and ATP-dependent protease. Biochemical characterisation revealed that HtpX undergoes self-degradation upon cell disruption or membrane solubilization. It can also degraded casein and cleaves solubilized membrane proteins, for example, SecY []. Expression of HtpX in the plasma membrane is under the control of CpxR, with the metalloproteinase active site of HtpX located on the cytosolic side of the membrane. This suggests a potential role for HtpX in the response to mis-folded proteins [].; GO: 0004222 metalloendopeptidase activity, 0006508 proteolysis, 0016020 membrane; PDB: 3CQB_A 3C37_B.
Probab=36.03  E-value=18  Score=30.23  Aligned_cols=36  Identities=25%  Similarity=0.251  Sum_probs=29.4

Q ss_pred             eeeEeeeCCCCCcceeeeee----eecHHHHhhCChhhhh
Q 039930           79 LQCAVYDSDDSTARLIGVEY----IVSDKIFEALPLEEQK  114 (245)
Q Consensus        79 ~QCliYDs~~~~ArLIGIEY----iISe~lf~tLP~eEkk  114 (245)
                      ..-.|+|++..||-.+|.-.    +|+..|++.|+++|-.
T Consensus        51 ~~v~v~~~~~~NA~~~g~~~~~~I~v~~~ll~~~~~~el~   90 (226)
T PF01435_consen   51 PRVYVIDSPSPNAFATGGGPRKRIVVTSGLLESLSEDELA   90 (226)
T ss_dssp             -EEEEE--SSEEEEEETTTC--EEEEEHHHHHHSSHHHHH
T ss_pred             CeEEEEcCCCCcEEEEccCCCcEEEEeChhhhcccHHHHH
Confidence            45678889999999999987    8999999999999965


No 4  
>smart00136 LamNT Laminin N-terminal domain (domain VI). N-terminal domain of laminins and laminin-related protein such as Unc-6/ netrins.
Probab=31.84  E-value=15  Score=33.45  Aligned_cols=79  Identities=19%  Similarity=0.329  Sum_probs=47.0

Q ss_pred             eeeee--ccCceeeeEeeeCCCCCcceeeeeeeecHHHHhhCChhhhhcccccc-----------------eeeccceee
Q 039930           69 HYVSR--LNQDFLQCAVYDSDDSTARLIGVEYIVSDKIFEALPLEEQKLWHSHA-----------------YEVKSGLWV  129 (245)
Q Consensus        69 HYC~~--lneD~~QCliYDs~~~~ArLIGIEYiISe~lf~tLP~eEkklWHsH~-----------------yEVkSG~Lv  129 (245)
                      .||..  .++.-.+|-+=|+..++ +=-.+||||.-..     +.....|-|-.                 |||--=.|.
T Consensus        33 ~yC~~~~~~~~~~~C~~CDa~~p~-~~Hp~~~l~D~~~-----~~~~TwWQS~~~~~~~~~VtitLdL~k~fevtyi~l~  106 (238)
T smart00136       33 RYCKLVGHTEQGKKCDYCDARNPR-RSHPAENLTDGNN-----PNNPTWWQSEPLSNGPQNVNLTLDLGKEFHVTYVILK  106 (238)
T ss_pred             ceeEeccccCcCCcCCCCCCCCcc-ccCCHHHhhccCC-----CCCCceecCCCcCCCCccEEEEEecCCEEEEEEEEEE
Confidence            46665  24556789888988763 5567888875321     22456777765                 666532222


Q ss_pred             cC-CCCccccchHHHHHhhccCceEEeecc
Q 039930          130 NP-RVPEMIGRPELDNMAKTYGKFWCTWQV  158 (245)
Q Consensus       130 ~P-~vP~~aE~~~M~~l~~tYGKt~HtWq~  158 (245)
                      .- .-|.   ...+|+-  .|||||.-||.
T Consensus       107 F~s~RPa---~~i~erS--d~G~tW~p~qy  131 (238)
T smart00136      107 FCSPRPS---LWILERS--DFGKTWQPWQY  131 (238)
T ss_pred             ecCCCCc---eEEEeec--CCCCCCcEeee
Confidence            21 2332   1222322  79999999998


No 5  
>TIGR00787 dctP tripartite ATP-independent periplasmic transporter solute receptor, DctP family. TRAP-T (Tripartite ATP-independent Periplasmic Transporter) family proteins generally consist of three components, and these systems have so far been found in Gram-negative bacteria, Gram-postive bacteria and archaea. The best characterized example is the DctPQM system of Rhodobacter capsulatus, a C4 dicarboxylate (malate, fumarate, succinate) transporter. This model represents the DctP family, one of at least three major families of extracytoplasmic solute receptor for TRAP family transporters. Other are the SnoM family (see pfam03480) and TAXI (TRAP-associated extracytoplasmic immunogenic) family.
Probab=31.52  E-value=30  Score=30.43  Aligned_cols=22  Identities=23%  Similarity=0.487  Sum_probs=19.2

Q ss_pred             eeeeeeecHHHHhhCChhhhhc
Q 039930           94 IGVEYIVSDKIFEALPLEEQKL  115 (245)
Q Consensus        94 IGIEYiISe~lf~tLP~eEkkl  115 (245)
                      .+.-++||.+.|++||+|.|+.
T Consensus       205 ~~~~~~~n~~~~~~L~~e~q~~  226 (257)
T TIGR00787       205 LGYLVVVNKAFWKSLPPDLQAV  226 (257)
T ss_pred             cceEEEEeHHHHhcCCHHHHHH
Confidence            4456899999999999999985


No 6  
>PF09022 Staphostatin_A:  Staphostatin A;  InterPro: IPR015112 The staphostatin A polypeptide chain folds into a slightly deformed, eight-stranded beta-barrel, with strands beta-4 through beta-8 forming an antiparallel sheet while the N terminus forms a psi-loop motif. Members of this family constitute a class of cysteine protease inhibitors distinct in the fold and the mechanism of action from any known inhibitors of these enzymes []. ; PDB: 1OH1_A.
Probab=30.92  E-value=56  Score=26.98  Aligned_cols=33  Identities=30%  Similarity=0.455  Sum_probs=23.7

Q ss_pred             eeeeeeeeeccCceeeeEeeeCCCCCcceeeeeeee
Q 039930           65 IETHHYVSRLNQDFLQCAVYDSDDSTARLIGVEYIV  100 (245)
Q Consensus        65 vEAhHYC~~lneD~~QCliYDs~~~~ArLIGIEYiI  100 (245)
                      +-..-||+-+|+|-||-++|.-+.++   |=||-+|
T Consensus        47 ~~p~YiC~~in~~~r~iiL~n~~n~s---IvIEi~i   79 (105)
T PF09022_consen   47 IQPDYICKYINTDSRQIILYNKDNSS---IVIEIII   79 (105)
T ss_dssp             TT--EEEEEEETTTTEEEEEETTTS----EEEEEE-
T ss_pred             CCcceeeeeecCCceEEEEEcCCCCc---EEEEEEE
Confidence            34567999999999999999887766   3477655


No 7  
>PF01082 Cu2_monooxygen:  Copper type II ascorbate-dependent monooxygenase, N-terminal domain;  InterPro: IPR000323 Copper type II, ascorbate-dependent monooxygenases [] are a class of enzymes that requires copper as a cofactor and which uses ascorbate as an electron donor. This family contains two related enzymes, Dopamine-beta-monooxygenase (1.14.17.1 from EC) and Peptidyl-glycine alpha-amidating monooxygenase (1.14.17.3 from EC). There are a few regions of sequence similarities between these two enzymes, two of these regions contain clusters of conserved histidine residues which are most probably involved in binding copper.; GO: 0004497 monooxygenase activity, 0005507 copper ion binding, 0016715 oxidoreductase activity, acting on paired donors, with incorporation or reduction of molecular oxygen, reduced ascorbate as one donor, and incorporation of one atom of oxygen, 0055114 oxidation-reduction process; PDB: 1YI9_A 3MLL_A 1SDW_A 3MID_A 1YIP_A 3PHM_A 3MIC_A 3MIB_A 1OPM_A 3MIG_A ....
Probab=29.91  E-value=55  Score=26.74  Aligned_cols=73  Identities=15%  Similarity=0.292  Sum_probs=40.1

Q ss_pred             cCceeeeEeeeCCC--CCcceeeeeeeecHHHHhhCChhhhhcccccceeeccceeecCCCCc--ccc--chHHHHHhhc
Q 039930           75 NQDFLQCAVYDSDD--STARLIGVEYIVSDKIFEALPLEEQKLWHSHAYEVKSGLWVNPRVPE--MIG--RPELDNMAKT  148 (245)
Q Consensus        75 neD~~QCliYDs~~--~~ArLIGIEYiISe~lf~tLP~eEkklWHsH~yEVkSG~Lv~P~vP~--~aE--~~~M~~l~~t  148 (245)
                      ++|...|.+++-+.  .+-.+||+|-+|+..         +-.=|.-.|+-.++  ..+ ++.  ..+  ...+......
T Consensus        16 ~~t~Y~C~~~~lp~~~~~~hIi~~ep~~~~~---------~~VHHmlly~C~~~--~~~-~~~~~~~~C~~~~~~~~~~~   83 (132)
T PF01082_consen   16 QDTTYWCFVFKLPDLTEKHHIIGFEPIITNG---------EVVHHMLLYGCDGD--NDP-LSRESGGECYSANMPGDFGS   83 (132)
T ss_dssp             SSSEEEEEEEE-S-S-S-EEEEEEEEE--T----------TTEEEEEEEEES-----SB-SSSSSSEEG-------GG-S
T ss_pred             CCCeEEEEEEECCcccccceEEEeeeeeccC---------CcEEEEEEEEECCC--Ccc-CcccCCCcccceeccccccc
Confidence            47999999999888  889999999999985         33667777888776  111 111  111  1233333444


Q ss_pred             cCceEEeeccC
Q 039930          149 YGKFWCTWQVD  159 (245)
Q Consensus       149 YGKt~HtWq~D  159 (245)
                      -..++..|-..
T Consensus        84 C~~ii~aWA~G   94 (132)
T PF01082_consen   84 CSTIIAAWAPG   94 (132)
T ss_dssp             B-EEEEEEETT
T ss_pred             ccceEEEEcCC
Confidence            47889999874


No 8  
>PF00505 HMG_box:  HMG (high mobility group) box;  InterPro: IPR000910 High mobility group (HMG or HMGB) proteins are a family of relatively low molecular weight non-histone components in chromatin. HMG1 (also called HMG-T in fish) and HMG2 are two highly related proteins that bind single-stranded DNA preferentially and unwind double-stranded DNA. Although they have no sequence specificity, they have a high affinity for bent or distorted DNA, and bend linear DNA. HMG1 and HMG2 contain two DNA-binding HMG-box domains (A and B) that show structural and functional differences, and have a long acidic C-terminal domain rich in aspartic and glutamic acid residues. The acidic tail modulates the affinity of the tandem HMG boxes in HMG1 and 2 for a variety of DNA targets. HMG1 and 2 appear to play important architectural roles in the assembly of nucleoprotein complexes in a variety of biological processes, for example V(D)J recombination, the initiation of transcription, and DNA repair []. The profile in this entry describing the HMG-domains is much more general than the signature. In addition to the HMG1 and HMG2 proteins, HMG-domains occur in single or multiple copies in the following protein classes; the SOX family of transcription factors; SRY sex determining region Y protein and related proteins []; LEF1 lymphoid enhancer binding factor 1 []; SSRP recombination signal recognition protein; MTF1 mitochondrial transcription factor 1; UBF1/2 nucleolar transcription factors; Abf2 yeast ARS-binding factor []; and Saccharomyces cerevisiae transcription factors Ixr1, Rox1, Nhp6a, Nhp6b and Spp41.; GO: 0003677 DNA binding; PDB: 1I11_A 1J3C_A 1J3D_A 1WZ6_A 1WGF_A 2D7L_A 1GT0_D 3U2B_C 2CRJ_A 2CS1_A ....
Probab=29.80  E-value=17  Score=25.29  Aligned_cols=16  Identities=19%  Similarity=0.428  Sum_probs=13.9

Q ss_pred             HHHHhhCChhhhhccc
Q 039930          102 DKIFEALPLEEQKLWH  117 (245)
Q Consensus       102 e~lf~tLP~eEkklWH  117 (245)
                      -+.|+.|+++||+.|+
T Consensus        36 ~~~W~~l~~~eK~~y~   51 (69)
T PF00505_consen   36 AQMWKNLSEEEKAPYK   51 (69)
T ss_dssp             HHHHHCSHHHHHHHHH
T ss_pred             HHHHhcCCHHHHHHHH
Confidence            4689999999999885


No 9  
>smart00398 HMG high mobility group.
Probab=27.76  E-value=30  Score=23.66  Aligned_cols=15  Identities=20%  Similarity=0.324  Sum_probs=13.2

Q ss_pred             HHHhhCChhhhhccc
Q 039930          103 KIFEALPLEEQKLWH  117 (245)
Q Consensus       103 ~lf~tLP~eEkklWH  117 (245)
                      ..|+.|+++||+.|-
T Consensus        38 ~~W~~l~~~ek~~y~   52 (70)
T smart00398       38 ERWKLLSEEEKAPYE   52 (70)
T ss_pred             HHHHcCCHHHHHHHH
Confidence            679999999999884


No 10 
>cd01390 HMGB-UBF_HMG-box HMGB-UBF_HMG-box, class II and III members of the HMG-box superfamily of DNA-binding proteins. These proteins bind the minor groove of DNA in a non-sequence specific fashion and contain two or more tandem HMG boxes. Class II members include non-histone chromosomal proteins, HMG1 and HMG2, which bind to bent or distorted DNA such as four-way DNA junctions, synthetic DNA cruciforms, kinked cisplatin-modified DNA, DNA bulges, cross-overs in supercoiled DNA, and can cause looping of linear DNA. Class III members include nucleolar and mitochondrial transcription factors, UBF and mtTF1, which bind four-way DNA junctions.
Probab=25.79  E-value=35  Score=23.38  Aligned_cols=15  Identities=27%  Similarity=0.441  Sum_probs=13.1

Q ss_pred             HHHhhCChhhhhccc
Q 039930          103 KIFEALPLEEQKLWH  117 (245)
Q Consensus       103 ~lf~tLP~eEkklWH  117 (245)
                      +.|+.|+++||+.|.
T Consensus        37 ~~W~~ls~~eK~~y~   51 (66)
T cd01390          37 EKWKELSEEEKKKYE   51 (66)
T ss_pred             HHHHhCCHHHHHHHH
Confidence            579999999999874


No 11 
>cd00084 HMG-box High Mobility Group (HMG)-box is found in a variety of eukaryotic chromosomal proteins and transcription factors. HMGs bind to the minor groove of DNA and have been classified by DNA binding preferences. Two phylogenically distinct groups of Class I proteins bind DNA in a sequence specific fashion and contain a single HMG box. One group (SOX-TCF) includes transcription factors, TCF-1, -3, -4; and also SRY and LEF-1, which bind four-way DNA junctions and duplex DNA targets. The second group (MATA) includes fungal mating type gene products MC, MATA1 and Ste11. Class II and III proteins (HMGB-UBF) bind DNA in a non-sequence specific fashion and contain two or more tandem HMG boxes. Class II members include non-histone chromosomal proteins, HMG1 and HMG2, which bind to bent or distorted DNA such as four-way DNA junctions, synthetic DNA cruciforms, kinked cisplatin-modified DNA, DNA bulges, cross-overs in supercoiled DNA, and can cause looping of linear DNA. Class III member
Probab=25.09  E-value=37  Score=22.94  Aligned_cols=15  Identities=27%  Similarity=0.567  Sum_probs=13.0

Q ss_pred             HHHhhCChhhhhccc
Q 039930          103 KIFEALPLEEQKLWH  117 (245)
Q Consensus       103 ~lf~tLP~eEkklWH  117 (245)
                      +.|+.|+++||..|-
T Consensus        37 ~~W~~l~~~~k~~y~   51 (66)
T cd00084          37 EMWKSLSEEEKKKYE   51 (66)
T ss_pred             HHHHhCCHHHHHHHH
Confidence            579999999999884


No 12 
>COG0518 GuaA GMP synthase - Glutamine amidotransferase domain [Nucleotide transport and metabolism]
Probab=23.93  E-value=32  Score=30.30  Aligned_cols=52  Identities=21%  Similarity=0.250  Sum_probs=34.6

Q ss_pred             HHHhhCChhhhhcccccceeec---cceeecCCCCccccchHHHHHhhccCceEEe
Q 039930          103 KIFEALPLEEQKLWHSHAYEVK---SGLWVNPRVPEMIGRPELDNMAKTYGKFWCT  155 (245)
Q Consensus       103 ~lf~tLP~eEkklWHsH~yEVk---SG~Lv~P~vP~~aE~~~M~~l~~tYGKt~Ht  155 (245)
                      .||+-||+.+...||||.-+|.   .|--+... .+...++.|+.=.+.||=-||.
T Consensus       121 ~l~~gl~~~~~~v~~sH~D~v~~lP~g~~vlA~-s~~cp~qa~~~~~~~~gvQFHp  175 (198)
T COG0518         121 PLFAGLPDLFTTVFMSHGDTVVELPEGAVVLAS-SETCPNQAFRYGKRAYGVQFHP  175 (198)
T ss_pred             ccccCCccccCccccchhCccccCCCCCEEEec-CCCChhhheecCCcEEEEeeee
Confidence            4999999888889999988775   33322222 2344566666446777766664


No 13 
>PF04512 Baculo_PEP_N:  Baculovirus polyhedron envelope protein, PEP, N terminus;  InterPro: IPR007600 Polyhedra are large crystalline occlusion bodies containing nucleopolyhedrovirus virions, and surrounded by an electron-dense structure called the polyhedron envelope or polyhedron calyx. The polyhedron envelope (associated) protein PEP is thought to be an integral part of the polyhedron envelope. PEP is concentrated at the surface of polyhedra, and is thought to be important for the proper formation of the periphery of polyhedra. It is thought that PEP may stabilise polyhedra and protect them from fusion or aggregation [].; GO: 0005198 structural molecule activity, 0019028 viral capsid, 0019031 viral envelope
Probab=23.12  E-value=42  Score=27.09  Aligned_cols=37  Identities=22%  Similarity=0.384  Sum_probs=26.5

Q ss_pred             ceeeeeeee-----cHHHHhhCChhhhhcccccceeecccee
Q 039930           92 RLIGVEYIV-----SDKIFEALPLEEQKLWHSHAYEVKSGLW  128 (245)
Q Consensus        92 rLIGIEYiI-----Se~lf~tLP~eEkklWHsH~yEVkSG~L  128 (245)
                      --+|+++|+     +...+++||..|||+|---.-.+-|..+
T Consensus        15 ~WvgaDEil~IL~lp~s~l~~iP~~~kk~w~dl~~~~~s~K~   56 (97)
T PF04512_consen   15 LWVGADEILSILRLPCSALQSIPRSHKKLWKDLEPCVDSNKL   56 (97)
T ss_pred             EEecHHHHHHHhCCCHHHHHHcCHHHHHHHHHhcccCCCcee
Confidence            346777775     4788999999999999655444444433


No 14 
>PF03480 SBP_bac_7:  Bacterial extracellular solute-binding protein, family 7;  InterPro: IPR018389 This family of proteins are involved in binding extracellular solutes for transport across the bacterial cytoplasmic membrane. This family includes a C4-dicarboxylate-binding protein DctP [, ] and the sialic acid-binding protein SiaP. The structure of the SiaP receptor has revealed an overall topology similar to ATP binding cassette ESR (extracytoplasmic solute receptors) proteins []. Upon binding of sialic acid, SiaP undergoes domain closure about a hinge region and kinking of an alpha-helix hinge component [].; GO: 0006810 transport, 0030288 outer membrane-bounded periplasmic space; PDB: 2HZK_C 2HZL_B 2HPG_C 2XWI_A 2XWK_A 2WX9_A 2CEY_A 2WYP_A 3B50_A 2CEX_B ....
Probab=22.95  E-value=58  Score=28.93  Aligned_cols=27  Identities=22%  Similarity=0.321  Sum_probs=20.5

Q ss_pred             eeeeeeecHHHHhhCChhhhhcccccc
Q 039930           94 IGVEYIVSDKIFEALPLEEQKLWHSHA  120 (245)
Q Consensus        94 IGIEYiISe~lf~tLP~eEkklWHsH~  120 (245)
                      -+.-++||.+.|++||+|.|+.=-.-.
T Consensus       205 ~~~~~~~n~~~w~~L~~e~q~~l~~~~  231 (286)
T PF03480_consen  205 SPYAVIMNKDWWDSLPDEDQEALDDAA  231 (286)
T ss_dssp             EEEEEEEEHHHHHHS-HHHHHHHHHHH
T ss_pred             cceEEEEcHHHHhcCCHHHHHHHHHHH
Confidence            346789999999999999998644333


No 15 
>PF01221 Dynein_light:  Dynein light chain type 1 ;  InterPro: IPR001372 Dynein is a multisubunit microtubule-dependent motor enzyme that acts as the force generating protein of eukaryotic cilia and flagella. The cytoplasmic isoform of dynein acts as a motor for the intracellular retrograde motility of vesicles and organelles along microtubules.  Dynein is composed of a number of ATP-binding large subunits (see IPR004273 from INTERPRO), intermediate size subunits and small subunits. Among the small subunits, there is a family of highly conserved proteins which make up this family [, ]. Both type 1 (DLC1) and 2 (DLC2) dynein light chains have a similar two-layer alpha-beta core structure consisting of beta-alpha(2)-beta-X-beta(2) [, ].; GO: 0007017 microtubule-based process, 0005875 microtubule associated complex; PDB: 1F95_A 1F96_A 1F3C_A 3P8M_B 2XQQ_C 1RE6_A 1CMI_A 1PWK_A 1PWJ_A 4DS1_C ....
Probab=22.93  E-value=42  Score=25.53  Aligned_cols=13  Identities=31%  Similarity=0.741  Sum_probs=10.0

Q ss_pred             HHhhccCceEEee
Q 039930          144 NMAKTYGKFWCTW  156 (245)
Q Consensus       144 ~l~~tYGKt~HtW  156 (245)
                      .+-++||++||-.
T Consensus        45 ~lD~~yG~~Wh~I   57 (89)
T PF01221_consen   45 ELDKKYGPTWHCI   57 (89)
T ss_dssp             HHHHHHSS-EEEE
T ss_pred             HHhcccCCceEEE
Confidence            5778999999984


No 16 
>PF07277 SapC:  SapC;  InterPro: IPR010836 This family contains a number of bacterial SapC proteins approximately 250 residues long. In Campylobacter fetus, SapC forms part of a paracrystalline surface layer (S-layer) that confers serum resistance [].
Probab=21.31  E-value=1.4e+02  Score=26.58  Aligned_cols=80  Identities=21%  Similarity=0.238  Sum_probs=61.7

Q ss_pred             CCCCCCchhhhHHHHHHhhhcCCchhhhhhhhceeeeeecCCCCCceeeeeeeeeccCceeeeEeeeCCCCCcceeeeee
Q 039930           19 PPGKPMTMKQHVLDKGAAMMQSLKPIKQMSQHVCTFALYSHDMPRQIETHHYVSRLNQDFLQCAVYDSDDSTARLIGVEY   98 (245)
Q Consensus        19 ~pG~~~~~~t~~Le~ga~~lQ~~~Pi~~i~~hL~aFH~ya~d~~rqvEAhHYC~~lneD~~QCliYDs~~~~ArLIGIEY   98 (245)
                      .-|+++..-.++++.=....+++.-.+++++.|...-....                   .+.-|--.++...+|-|+ |
T Consensus       122 ~~G~~T~~l~~~~~~L~~~~~~~~~T~~f~~~L~~~~Ll~~-------------------~~l~v~~~~g~~~~l~G~-~  181 (221)
T PF07277_consen  122 EDGEPTEYLQQVLNFLQQYQQGRQQTQAFIKALAELGLLEP-------------------WTLTVTLDDGEKHNLNGF-Y  181 (221)
T ss_pred             CCCCcCHHHHHHHHHHHHHHHHHHHHHHHHHHHHHcCCCcc-------------------cEEEEEeCCCCeeeccCc-e
Confidence            56999999999999888888888889999988865433321                   122333367777888887 9


Q ss_pred             eecHHHHhhCChhhhhcccc
Q 039930           99 IVSDKIFEALPLEEQKLWHS  118 (245)
Q Consensus        99 iISe~lf~tLP~eEkklWHs  118 (245)
                      +|+|+-++.|+++.-.-||.
T Consensus       182 ~Vde~kL~~L~de~l~~L~~  201 (221)
T PF07277_consen  182 TVDEEKLNALSDEALLELHK  201 (221)
T ss_pred             EECHHHHhcCCHHHHHHHHH
Confidence            99999999999999777764


No 17 
>COG1638 DctP TRAP-type C4-dicarboxylate transport system, periplasmic component [Carbohydrate transport and metabolism]
Probab=20.65  E-value=62  Score=30.66  Aligned_cols=23  Identities=26%  Similarity=0.563  Sum_probs=18.7

Q ss_pred             eeeeeeecHHHHhhCChhhhhcc
Q 039930           94 IGVEYIVSDKIFEALPLEEQKLW  116 (245)
Q Consensus        94 IGIEYiISe~lf~tLP~eEkklW  116 (245)
                      ...=++||.+.|++||+|+|+.=
T Consensus       236 ~~~~~~~s~~~w~~L~~e~q~il  258 (332)
T COG1638         236 LPLAVLVSKAFWDSLPEEDQTIL  258 (332)
T ss_pred             cceeeEEcHHHHhcCCHHHHHHH
Confidence            34447899999999999998853


Done!