Query         039418
Match_columns 297
No_of_seqs    168 out of 337
Neff          6.7 
Searched_HMMs 46136
Date          Fri Mar 29 09:29:09 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/039418.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/039418hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF10536 PMD:  Plant mobile dom 100.0   2E-29 4.3E-34  241.6   8.2  171  127-297     1-184 (363)
  2 PTZ00199 high mobility group p  99.5 3.3E-14 7.1E-19  111.6   5.9   61   30-90     21-82  (94)
  3 cd01389 MATA_HMG-box MATA_HMG-  99.5 3.8E-14 8.2E-19  107.0   5.7   59   31-90      1-59  (77)
  4 cd01390 HMGB-UBF_HMG-box HMGB-  99.4 1.1E-13 2.5E-18  100.5   5.5   57   32-89      1-57  (66)
  5 cd01388 SOX-TCF_HMG-box SOX-TC  99.4 2.5E-13 5.4E-18  101.2   5.6   58   32-90      2-59  (72)
  6 PF00505 HMG_box:  HMG (high mo  99.4 2.7E-13 5.8E-18   99.4   4.8   58   32-90      1-58  (69)
  7 smart00398 HMG high mobility g  99.4   9E-13 1.9E-17   96.4   5.7   59   31-90      1-59  (70)
  8 PF09011 HMG_box_2:  HMG-box do  99.4 6.3E-13 1.4E-17   99.3   4.9   59   30-89      2-61  (73)
  9 cd00084 HMG-box High Mobility   99.4 1.2E-12 2.7E-17   94.5   5.5   59   32-91      1-59  (66)
 10 KOG0381 HMG box-containing pro  99.2 2.4E-11 5.2E-16   95.0   5.7   60   30-90     21-80  (96)
 11 PF09331 DUF1985:  Domain of un  99.0 4.7E-10   1E-14   94.5   6.7  123  156-279    14-142 (142)
 12 KOG0527 HMG-box transcription   99.0 3.4E-10 7.5E-15  107.1   4.4   60   30-90     61-120 (331)
 13 COG5648 NHP6B Chromatin-associ  99.0 6.6E-10 1.4E-14   97.7   5.3   61   29-90     68-128 (211)
 14 KOG0526 Nucleosome-binding fac  98.2   8E-07 1.7E-11   87.5   3.4   56   30-90    534-589 (615)
 15 KOG3248 Transcription factor T  97.5 7.5E-05 1.6E-09   70.1   3.4   57   33-90    193-249 (421)
 16 KOG0528 HMG-box transcription   97.4 4.6E-05 9.9E-10   74.6   1.2   63   31-94    325-387 (511)
 17 KOG2746 HMG-box transcription   95.5  0.0041 8.8E-08   63.4   0.2   62   32-94    182-245 (683)
 18 KOG4715 SWI/SNF-related matrix  95.4   0.017 3.6E-07   54.3   3.9   61   27-88     59-120 (410)
 19 PF06382 DUF1074:  Protein of u  95.0   0.026 5.7E-07   48.9   3.6   48   35-87     82-129 (183)
 20 PF14887 HMG_box_5:  HMG (high   93.1    0.15 3.3E-06   38.4   4.0   55   32-88      4-58  (85)
 21 PF04690 YABBY:  YABBY protein;  91.7    0.26 5.7E-06   42.8   4.4   44   32-76    122-165 (170)
 22 COG5648 NHP6B Chromatin-associ  90.1    0.16 3.5E-06   45.3   1.6   56   32-88    144-199 (211)
 23 PF03078 ATHILA:  ATHILA ORF-1   76.2      51  0.0011   33.1  12.4  168  101-278    63-262 (458)
 24 PF11304 DUF3106:  Protein of u  65.1       8 0.00017   30.9   3.4   21   67-87     34-54  (107)
 25 PF04769 MAT_Alpha1:  Mating-ty  61.8      17 0.00038   32.4   5.2   44   29-77     41-84  (201)
 26 PF08073 CHDNT:  CHDNT (NUC034)  44.5      25 0.00054   24.9   2.7   39   37-76     14-52  (55)
 27 PF11943 DUF3460:  Protein of u  41.2      57  0.0012   23.5   4.1   38   44-84      8-47  (60)
 28 PF10234 Cluap1:  Clusterin-ass  39.3      15 0.00033   34.2   1.2   32  124-155     1-37  (267)
 29 COG5202 Predicted membrane pro  39.2      19  0.0004   35.1   1.8   27   12-38    189-217 (512)
 30 PF06945 DUF1289:  Protein of u  33.6      21 0.00046   24.5   0.9   20   69-88     28-47  (51)
 31 PF01418 HTH_6:  Helix-turn-hel  31.2      33 0.00071   25.3   1.7   60   65-135     5-67  (77)
 32 PF13875 DUF4202:  Domain of un  28.8      79  0.0017   27.9   3.9   64   11-80     99-169 (185)
 33 PF14513 DAG_kinase_N:  Diacylg  28.5      11 0.00023   31.8  -1.5   72   69-155     3-82  (138)
 34 cd09071 FAR_C C-terminal domai  28.4      55  0.0012   24.5   2.6   21  261-282    70-90  (92)
 35 PF05494 Tol_Tol_Ttg2:  Toluene  27.0      48   0.001   28.1   2.3   30   59-88     39-69  (170)
 36 PF03457 HA:  Helicase associat  26.4      42  0.0009   23.9   1.5   16  113-128    52-67  (68)
 37 PF07970 COPIIcoated_ERV:  Endo  25.8      71  0.0015   28.6   3.2   35   16-51     13-47  (222)
 38 TIGR03481 HpnM hopanoid biosyn  25.6      47   0.001   29.3   2.0   30   59-88     65-95  (198)
 39 cd07321 Extradiol_Dioxygenase_  25.2      69  0.0015   23.9   2.5   32  114-148    34-65  (77)
 40 PRK15117 ABC transporter perip  24.9      47   0.001   29.6   1.9   25   64-88     75-99  (211)
 41 cd02988 Phd_like_VIAF Phosduci  23.8      48   0.001   29.1   1.7   17   22-38      4-20  (192)
 42 PF00701 DHDPS:  Dihydrodipicol  23.2      66  0.0014   29.6   2.6  102   70-179    47-155 (289)
 43 PF12650 DUF3784:  Domain of un  23.2      45 0.00098   25.6   1.3   17   70-86     25-41  (97)
 44 PF03015 Sterile:  Male sterili  22.3      85  0.0018   23.8   2.6   54  227-283    33-91  (94)
 45 smart00271 DnaJ DnaJ molecular  22.1      99  0.0022   20.9   2.7   36   42-77     18-58  (60)
 46 PF06628 Catalase-rel:  Catalas  21.8      50  0.0011   23.9   1.1   23   66-88     12-34  (68)
 47 PF00226 DnaJ:  DnaJ domain;  I  20.1 1.1E+02  0.0023   21.1   2.6   39   42-80     17-60  (64)

No 1  
>PF10536 PMD:  Plant mobile domain;  InterPro: IPR019557  This entry represents a domain found in a variety of transposases []. 
Probab=99.96  E-value=2e-29  Score=241.61  Aligned_cols=171  Identities=19%  Similarity=0.322  Sum_probs=146.2

Q ss_pred             Ccccccccc--cccccHHHHHHHHhccccCcceEEECCeEeecCccchhheeccccCCccccccCChh---HHHHHHhhh
Q 039418          127 GLGSIIDLK--CGRLKRKLCAWLVERIDTARCVLQLNGHELELSPNSFGYIMGVTDGGMPMELQGDSA---EVAAYLDKF  201 (297)
Q Consensus       127 GFg~LL~i~--~~~l~~~L~~wL~~~~d~~t~~~~i~g~~i~iT~~dV~~VLGLP~gG~~v~~~~~~~---~~~~l~~~~  201 (297)
                      |||+|+.|.  ..++++.|+.+|+++|+++|++|++++++++||++||..|+|||+.|.+|....+.+   .++++....
T Consensus         1 ~~g~~~~i~~s~~~~~~~li~al~erW~~et~tF~~~~gEmtiTL~DV~~llGLpi~G~pv~~~~~~~~~~~~~~ll~~~   80 (363)
T PF10536_consen    1 GFGILDAIMASRITIDRSLISALVERWDPETNTFHFPWGEMTITLEDVAMLLGLPIDGRPVTGPLPPDWRDLCEELLGVS   80 (363)
T ss_pred             CchhHhhhhhhcCCCCHHHHHHHHHHhCcccCeeecccccccchhhhhhhccccccccccccCccccchhhHHHHHhccc
Confidence            899999999  899999999999999999999999999999999999999999999999998754332   333333222


Q ss_pred             cc----CCCccchHHHHHHHhcCCCC-CchhhhHHHhhhhcceeCCCCCC-ccCcchhhhhhccccCcccchhHHHHHHH
Q 039418          202 NA----TSRGINIKTMEDILLTSKDA-DNDFKVAFMLFTLCTLLCPPGGV-HISYSFLFTLKDVHSIRNRNWATFCFERL  275 (297)
Q Consensus       202 ~~----~~~~isl~~L~~~ll~~~~~-~d~f~r~Fll~~i~~~L~Ptt~~-~vs~~yl~~l~D~~~I~~ynW~~~Vld~L  275 (297)
                      ..    .+..+.+++|++.+...+++ ++.+.||||++.+|++|||+++. +|+..|++++.|++.+++||||++||++|
T Consensus        81 ~~~~~~~~~~~~~~wl~~~~~~~~~~d~~~~~rAFll~~lg~~lfp~~~~~~v~~~~l~~~~~l~~~~~~~wg~a~La~l  160 (363)
T PF10536_consen   81 PQIKSKKGSSIRLSWLEEFFSNRPEDDEEQYHRAFLLYWLGSFLFPDKSGDYVSPRYLPLAVDLARIKRYAWGSAVLAYL  160 (363)
T ss_pred             ccccccccccchhhheeccccccccchHHHHHHHHHHHhhhceeccCCCcceeeeeEEeeeeccccccccccHHHHHHHH
Confidence            11    23456778888886333333 24799999999999999999998 89999999999999999999999999999


Q ss_pred             HHHHHhhhccC--CceeecceecC
Q 039418          276 MRGITRYKDEK--LAHVGGCLLYL  297 (297)
Q Consensus       276 ~~~l~k~~~~k--~~~i~GCllfL  297 (297)
                      +++|++...+.  ..+++||+.||
T Consensus       161 y~~L~~~~~~~~~~~~~~g~~~ll  184 (363)
T PF10536_consen  161 YRDLCKASRKSASQSNIGGPLWLL  184 (363)
T ss_pred             HHHHHHHhhhcccccccccceeee
Confidence            99999988876  78999999986


No 2  
>PTZ00199 high mobility group protein; Provisional
Probab=99.49  E-value=3.3e-14  Score=111.64  Aligned_cols=61  Identities=18%  Similarity=0.269  Sum_probs=56.9

Q ss_pred             CCCCCCCcchhhhHHHHHHHHHHhCCCCc-chhHHHHHHHhhhcCCChHhhhhhhhhhhcCC
Q 039418           30 SKDRENNHGFISFFAESVRQLKAKDGRAC-ITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGG   90 (297)
Q Consensus        30 ~~~~r~~~~f~~~~~~~~~~~~~~~~~~~-~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~   90 (297)
                      |+||||+||||+|++++|.+++++||+.. .+.+|++++|+.|++||++||++|.++|.+..
T Consensus        21 ~~PKrP~sAY~~F~~~~R~~i~~~~P~~~~~~~evsk~ige~Wk~ls~eeK~~y~~~A~~dk   82 (94)
T PTZ00199         21 NAPKRALSAYMFFAKEKRAEIIAENPELAKDVAAVGKMVGEAWNKLSEEEKAPYEKKAQEDK   82 (94)
T ss_pred             CCCCCCCcHHHHHHHHHHHHHHHHCcCCcccHHHHHHHHHHHHHcCCHHHHHHHHHHHHHHH
Confidence            78999999999999999999999999875 58999999999999999999999999998743


No 3  
>cd01389 MATA_HMG-box MATA_HMG-box, class I member of the HMG-box superfamily of DNA-binding proteins. These proteins contain a single HMG box, and bind the minor groove of DNA in a highly sequence-specific manner. Members include the fungal mating type gene products MC, MATA1 and Ste11.
Probab=99.49  E-value=3.8e-14  Score=106.96  Aligned_cols=59  Identities=22%  Similarity=0.259  Sum_probs=55.7

Q ss_pred             CCCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcCC
Q 039418           31 KDRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGG   90 (297)
Q Consensus        31 ~~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~   90 (297)
                      +||||+||||+|++++|++++++||+ .++.+|++.+|..|+.||++||++|.+.|.+..
T Consensus         1 ~~kRP~naf~lf~~~~r~~~~~~~p~-~~~~eisk~~g~~Wk~ls~eeK~~y~~~A~~~k   59 (77)
T cd01389           1 KIPRPRNAFILYRQDKHAQLKTENPG-LTNNEISRIIGRMWRSESPEVKAYYKELAEEEK   59 (77)
T ss_pred             CCCCCCcHHHHHHHHHHHHHHHHCCC-CCHHHHHHHHHHHHhhCCHHHHHHHHHHHHHHH
Confidence            58999999999999999999999994 689999999999999999999999999998754


No 4  
>cd01390 HMGB-UBF_HMG-box HMGB-UBF_HMG-box, class II and III members of the HMG-box superfamily of DNA-binding proteins. These proteins bind the minor groove of DNA in a non-sequence specific fashion and contain two or more tandem HMG boxes. Class II members include non-histone chromosomal proteins, HMG1 and HMG2, which bind to bent or distorted DNA such as four-way DNA junctions, synthetic DNA cruciforms, kinked cisplatin-modified DNA, DNA bulges, cross-overs in supercoiled DNA, and can cause looping of linear DNA. Class III members include nucleolar and mitochondrial transcription factors, UBF and mtTF1, which bind four-way DNA junctions.
Probab=99.45  E-value=1.1e-13  Score=100.46  Aligned_cols=57  Identities=23%  Similarity=0.332  Sum_probs=54.4

Q ss_pred             CCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcC
Q 039418           32 DRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRG   89 (297)
Q Consensus        32 ~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~   89 (297)
                      ||||+|||++|+.|+|.++++.||+ .++.+|++.+|..|++||++||++|.++|++.
T Consensus         1 Pkrp~saf~~f~~~~r~~~~~~~p~-~~~~~i~~~~~~~W~~ls~~eK~~y~~~a~~~   57 (66)
T cd01390           1 PKRPLSAYFLFSQEQRPKLKKENPD-ASVTEVTKILGEKWKELSEEEKKKYEEKAEKD   57 (66)
T ss_pred             CCCCCcHHHHHHHHHHHHHHHHCcC-CCHHHHHHHHHHHHHhCCHHHHHHHHHHHHHH
Confidence            8999999999999999999999995 68999999999999999999999999999874


No 5  
>cd01388 SOX-TCF_HMG-box SOX-TCF_HMG-box, class I member of the HMG-box superfamily of DNA-binding proteins. These proteins contain a single HMG box, and bind the minor groove of DNA in a highly sequence-specific manner. Members include SRY and its homologs in insects and vertebrates, and transcription factor-like proteins, TCF-1, -3, -4, and LEF-1. They appear to bind the minor groove of the A/T C A A A G/C-motif.
Probab=99.42  E-value=2.5e-13  Score=101.25  Aligned_cols=58  Identities=17%  Similarity=0.178  Sum_probs=54.7

Q ss_pred             CCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcCC
Q 039418           32 DRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGG   90 (297)
Q Consensus        32 ~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~   90 (297)
                      +|||+||||+|+++.|+.++++||+ .++.+|+|.+|+.|+.||++||++|.+.|++..
T Consensus         2 iKrP~naf~~F~~~~r~~~~~~~p~-~~~~eisk~l~~~Wk~ls~~eK~~y~~~a~~~k   59 (72)
T cd01388           2 IKRPMNAFMLFSKRHRRKVLQEYPL-KENRAISKILGDRWKALSNEEKQPYYEEAKKLK   59 (72)
T ss_pred             CCCCCcHHHHHHHHHHHHHHHHCCC-CCHHHHHHHHHHHHHcCCHHHHHHHHHHHHHHH
Confidence            6999999999999999999999995 699999999999999999999999999998744


No 6  
>PF00505 HMG_box:  HMG (high mobility group) box;  InterPro: IPR000910 High mobility group (HMG or HMGB) proteins are a family of relatively low molecular weight non-histone components in chromatin. HMG1 (also called HMG-T in fish) and HMG2 are two highly related proteins that bind single-stranded DNA preferentially and unwind double-stranded DNA. Although they have no sequence specificity, they have a high affinity for bent or distorted DNA, and bend linear DNA. HMG1 and HMG2 contain two DNA-binding HMG-box domains (A and B) that show structural and functional differences, and have a long acidic C-terminal domain rich in aspartic and glutamic acid residues. The acidic tail modulates the affinity of the tandem HMG boxes in HMG1 and 2 for a variety of DNA targets. HMG1 and 2 appear to play important architectural roles in the assembly of nucleoprotein complexes in a variety of biological processes, for example V(D)J recombination, the initiation of transcription, and DNA repair []. The profile in this entry describing the HMG-domains is much more general than the signature. In addition to the HMG1 and HMG2 proteins, HMG-domains occur in single or multiple copies in the following protein classes; the SOX family of transcription factors; SRY sex determining region Y protein and related proteins []; LEF1 lymphoid enhancer binding factor 1 []; SSRP recombination signal recognition protein; MTF1 mitochondrial transcription factor 1; UBF1/2 nucleolar transcription factors; Abf2 yeast ARS-binding factor []; and Saccharomyces cerevisiae transcription factors Ixr1, Rox1, Nhp6a, Nhp6b and Spp41.; GO: 0003677 DNA binding; PDB: 1I11_A 1J3C_A 1J3D_A 1WZ6_A 1WGF_A 2D7L_A 1GT0_D 3U2B_C 2CRJ_A 2CS1_A ....
Probab=99.40  E-value=2.7e-13  Score=99.44  Aligned_cols=58  Identities=26%  Similarity=0.386  Sum_probs=52.5

Q ss_pred             CCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcCC
Q 039418           32 DRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGG   90 (297)
Q Consensus        32 ~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~   90 (297)
                      ||||+|||++|+++++.++++.||+. ...+|++.+|..|++||++||++|.+.|.+..
T Consensus         1 PkrP~~af~lf~~~~~~~~k~~~p~~-~~~~i~~~~~~~W~~l~~~eK~~y~~~a~~~~   58 (69)
T PF00505_consen    1 PKRPPNAFMLFCKEKRAKLKEENPDL-SNKEISKILAQMWKNLSEEEKAPYKEEAEEEK   58 (69)
T ss_dssp             SSSS--HHHHHHHHHHHHHHHHSTTS-THHHHHHHHHHHHHCSHHHHHHHHHHHHHHHH
T ss_pred             CcCCCCHHHHHHHHHHHHHHHHhccc-ccccchhhHHHHHhcCCHHHHHHHHHHHHHHH
Confidence            89999999999999999999999955 59999999999999999999999999998744


No 7  
>smart00398 HMG high mobility group.
Probab=99.37  E-value=9e-13  Score=96.36  Aligned_cols=59  Identities=24%  Similarity=0.377  Sum_probs=55.0

Q ss_pred             CCCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcCC
Q 039418           31 KDRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGG   90 (297)
Q Consensus        31 ~~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~   90 (297)
                      +||||+|||++|++++|+.++++||+ ....++++.+|..|+.||++||++|.++|++..
T Consensus         1 ~pkrp~~~y~~f~~~~r~~~~~~~~~-~~~~~i~~~~~~~W~~l~~~ek~~y~~~a~~~~   59 (70)
T smart00398        1 KPKRPMSAFMLFSQENRAKIKAENPD-LSNAEISKKLGERWKLLSEEEKAPYEEKAKKDK   59 (70)
T ss_pred             CcCCCCcHHHHHHHHHHHHHHHHCcC-CCHHHHHHHHHHHHHcCCHHHHHHHHHHHHHHH
Confidence            58999999999999999999999995 578999999999999999999999999988743


No 8  
>PF09011 HMG_box_2:  HMG-box domain;  InterPro: IPR015101 This domain is predominantly found in Maelstrom homologue proteins. It has no known function. ; GO: 0005634 nucleus; PDB: 2EQZ_A 1V64_A 2CTO_A 1H5P_A 3TQ6_A 3FGH_A 3TMM_A 1J3X_A 2YRQ_A 1AAB_A ....
Probab=99.37  E-value=6.3e-13  Score=99.32  Aligned_cols=59  Identities=25%  Similarity=0.343  Sum_probs=51.0

Q ss_pred             CCCCCCCcchhhhHHHHHHHHHHh-CCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcC
Q 039418           30 SKDRENNHGFISFFAESVRQLKAK-DGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRG   89 (297)
Q Consensus        30 ~~~~r~~~~f~~~~~~~~~~~~~~-~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~   89 (297)
                      +|||||+|||++|+.|++..+++. .+ .....+|.+.+|..|++||++||++|.++|++.
T Consensus         2 ~kpK~~~say~lF~~~~~~~~k~~G~~-~~~~~e~~k~~~~~Wk~Ls~~EK~~Y~~~A~~~   61 (73)
T PF09011_consen    2 KKPKRPPSAYNLFMKEMRKEVKEEGGQ-KQSFREVMKEISERWKSLSEEEKEPYEERAKED   61 (73)
T ss_dssp             SS--SSSSHHHHHHHHHHHHHHHHT-T--SSHHHHHHHHHHHHHHS-HHHHHHHHHHHHHH
T ss_pred             cCCCCCCCHHHHHHHHHHHHHHHhccc-CCCHHHHHHHHHHHHHhcCHHHHHHHHHHHHHH
Confidence            689999999999999999999999 55 778899999999999999999999999999874


No 9  
>cd00084 HMG-box High Mobility Group (HMG)-box is found in a variety of eukaryotic chromosomal proteins and transcription factors. HMGs bind to the minor groove of DNA and have been classified by DNA binding preferences. Two phylogenically distinct groups of Class I proteins bind DNA in a sequence specific fashion and contain a single HMG box. One group (SOX-TCF) includes transcription factors, TCF-1, -3, -4; and also SRY and LEF-1, which bind four-way DNA junctions and duplex DNA targets. The second group (MATA) includes fungal mating type gene products MC, MATA1 and Ste11. Class II and III proteins (HMGB-UBF) bind DNA in a non-sequence specific fashion and contain two or more tandem HMG boxes. Class II members include non-histone chromosomal proteins, HMG1 and HMG2, which bind to bent or distorted DNA such as four-way DNA junctions, synthetic DNA cruciforms, kinked cisplatin-modified DNA, DNA bulges, cross-overs in supercoiled DNA, and can cause looping of linear DNA. Class III member
Probab=99.35  E-value=1.2e-12  Score=94.54  Aligned_cols=59  Identities=20%  Similarity=0.306  Sum_probs=55.1

Q ss_pred             CCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcCCC
Q 039418           32 DRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGGK   91 (297)
Q Consensus        32 ~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~k   91 (297)
                      ||||+|||++|+.|+++.+++.+|+ ....+|.+.+|..|+.||++||++|.++|++...
T Consensus         1 pkrp~~af~~f~~~~~~~~~~~~~~-~~~~~i~~~~~~~W~~l~~~~k~~y~~~a~~~~~   59 (66)
T cd00084           1 PKRPLSAYFLFSQEHRAEVKAENPG-LSVGEISKILGEMWKSLSEEEKKKYEEKAEKDKE   59 (66)
T ss_pred             CCCCCcHHHHHHHHHHHHHHHHCcC-CCHHHHHHHHHHHHHhCCHHHHHHHHHHHHHHHH
Confidence            7999999999999999999999995 6799999999999999999999999999987543


No 10 
>KOG0381 consensus HMG box-containing protein [General function prediction only]
Probab=99.20  E-value=2.4e-11  Score=95.03  Aligned_cols=60  Identities=25%  Similarity=0.314  Sum_probs=55.9

Q ss_pred             CCCCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcCC
Q 039418           30 SKDRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGG   90 (297)
Q Consensus        30 ~~~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~   90 (297)
                      +.||||++|||+|+.+++..++..||+ ..+.+|+|++|..|++|+++||.+|..++.+-.
T Consensus        21 ~~pkrp~sa~~~f~~~~~~~~k~~~p~-~~~~~v~k~~g~~W~~l~~~~k~~y~~ka~~~k   80 (96)
T KOG0381|consen   21 QAPKRPLSAFFLFSSEQRSKIKAENPG-LSVGEVAKALGEMWKNLAEEEKQPYEEKASKLK   80 (96)
T ss_pred             CCCCCCCcHHHHHHHHHHHHHHHhCCC-CCHHHHHHHHHHHHhcCCHHHHHHHHHHHHHHH
Confidence            479999999999999999999999996 899999999999999999999999988887643


No 11 
>PF09331 DUF1985:  Domain of unknown function (DUF1985);  InterPro: IPR015410 This domain is functionally uncharacterised; it is found in a set of Arabidopsis thaliana (Mouse-ear cress) hypothetical proteins. 
Probab=99.03  E-value=4.7e-10  Score=94.48  Aligned_cols=123  Identities=20%  Similarity=0.390  Sum_probs=91.2

Q ss_pred             ceEEECCeEeecCccchhheeccccCCccccccCChhHHH---HHH-hhhccCCCccchHHHHHHHhcC--CCCCchhhh
Q 039418          156 CVLQLNGHELELSPNSFGYIMGVTDGGMPMELQGDSAEVA---AYL-DKFNATSRGINIKTMEDILLTS--KDADNDFKV  229 (297)
Q Consensus       156 ~~~~i~g~~i~iT~~dV~~VLGLP~gG~~v~~~~~~~~~~---~l~-~~~~~~~~~isl~~L~~~ll~~--~~~~d~f~r  229 (297)
                      ..+.++|..|.++..+.+.|+|||++..|-..........   .+- ..++ .+..+++..+.++|...  .+.++.+.-
T Consensus        14 ~W~~~~g~piRfsl~Ef~lvTGL~C~~~p~~~~~~~~~~~~~~~fw~~Lf~-~~~~vtv~dv~~~L~~~~~~~~~~Rlrl   92 (142)
T PF09331_consen   14 IWFVFNGVPIRFSLREFALVTGLNCGPYPKEKKVDKKGKKEKGSFWNKLFG-REEDVTVEDVIAKLKKMKKWDSEDRLRL   92 (142)
T ss_pred             EEEEECCEeeEecHHHHHhhcCCcCCCCCcccchhhccccchhhhhhhhcc-ccccCcHHHHHHHHhhcccCChhhHHHH
Confidence            7889999999999999999999999887766543221111   232 2333 34569999999998654  234444555


Q ss_pred             HHHhhhhcceeCCCCCCccCcchhhhhhccccCcccchhHHHHHHHHHHH
Q 039418          230 AFMLFTLCTLLCPPGGVHISYSFLFTLKDVHSIRNRNWATFCFERLMRGI  279 (297)
Q Consensus       230 ~Fll~~i~~~L~Ptt~~~vs~~yl~~l~D~~~I~~ynW~~~Vld~L~~~l  279 (297)
                      ++++++.|.+++++....|+..++..++|++.+.+|-||.+.++.++++|
T Consensus        93 a~L~~v~gvl~~~~~~~~i~~~~~~~v~Dl~~f~~yPWGr~sF~~~~~sI  142 (142)
T PF09331_consen   93 ALLLFVDGVLIATSKTTKIPKEHLKMVDDLEKFLNYPWGRYSFDMLMKSI  142 (142)
T ss_pred             HHHHhhheeeeccCCCCCCCHHHHHHHhhHHHHhcCCcHHHHHHHHHhcC
Confidence            55555555555555556899999999999999999999999999999874


No 12 
>KOG0527 consensus HMG-box transcription factor [Transcription]
Probab=98.98  E-value=3.4e-10  Score=107.09  Aligned_cols=60  Identities=18%  Similarity=0.248  Sum_probs=56.2

Q ss_pred             CCCCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcCC
Q 039418           30 SKDRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGG   90 (297)
Q Consensus        30 ~~~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~   90 (297)
                      .++|||-|||||+.++.||.+.+.|| .-.+++|+|.+|..||.|+|+||.||++-|++-.
T Consensus        61 ~hIKRPMNAFMVWSq~~RRkma~qnP-~mHNSEISK~LG~~WK~Lse~EKrPFi~EAeRLR  120 (331)
T KOG0527|consen   61 DRIKRPMNAFMVWSQGQRRKLAKQNP-KMHNSEISKRLGAEWKLLSEEEKRPFVDEAERLR  120 (331)
T ss_pred             cccCCCcchhhhhhHHHHHHHHHhCc-chhhHHHHHHHHHHHhhcCHhhhccHHHHHHHHH
Confidence            45799999999999999999999999 5599999999999999999999999999998755


No 13 
>COG5648 NHP6B Chromatin-associated proteins containing the HMG domain [Chromatin structure and dynamics]
Probab=98.97  E-value=6.6e-10  Score=97.67  Aligned_cols=61  Identities=18%  Similarity=0.251  Sum_probs=57.4

Q ss_pred             CCCCCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcCC
Q 039418           29 GSKDRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGG   90 (297)
Q Consensus        29 ~~~~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~   90 (297)
                      +|.||||-||||.|.++-|.++.+.+|.. .+++|+|.+|++||+|++.||+||...+..-+
T Consensus        68 pN~PKRp~sayf~y~~~~R~ei~~~~p~l-~~~e~~k~~~e~WK~Ltd~eke~y~k~~~~~~  128 (211)
T COG5648          68 PNGPKRPLSAYFLYSAENRDEIRKENPKL-TFGEVGKLLSEKWKELTDEEKEPYYKEANSDR  128 (211)
T ss_pred             CCCCCCchhHHHHHHHHHHHHHHHhCCCC-ChHHHHHHHHHHHHhccHhhhhhHHHHHhhHH
Confidence            48899999999999999999999999965 99999999999999999999999999988744


No 14 
>KOG0526 consensus Nucleosome-binding factor SPN, POB3 subunit [Transcription; Replication, recombination and repair; Chromatin structure and dynamics]
Probab=98.21  E-value=8e-07  Score=87.50  Aligned_cols=56  Identities=11%  Similarity=0.243  Sum_probs=51.1

Q ss_pred             CCCCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcCC
Q 039418           30 SKDRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGG   90 (297)
Q Consensus        30 ~~~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~   90 (297)
                      |.||||-||||+|++.-|..+|++   .-.+.+|+|.+|++||.||.  |++|.++|.+-+
T Consensus       534 napkra~sa~m~w~~~~r~~ik~d---gi~~~dv~kk~g~~wk~ms~--k~~we~ka~~dk  589 (615)
T KOG0526|consen  534 NAPKRATSAYMLWLNASRESIKED---GISVGDVAKKAGEKWKQMSA--KEEWEDKAAVDK  589 (615)
T ss_pred             CCCccchhHHHHHHHhhhhhHhhc---CchHHHHHHHHhHHHhhhcc--cchhhHHHHHHH
Confidence            788999999999999999999998   45899999999999999999  788988887643


No 15 
>KOG3248 consensus Transcription factor TCF-4 [Transcription]
Probab=97.49  E-value=7.5e-05  Score=70.13  Aligned_cols=57  Identities=14%  Similarity=0.224  Sum_probs=53.1

Q ss_pred             CCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcCC
Q 039418           33 RENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGG   90 (297)
Q Consensus        33 ~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~   90 (297)
                      |.|.+||++||.|.|+-+-++-- .|..+++-+++|.+|-.||-+|.|.|..-|++..
T Consensus       193 KKPLNAFmlyMKEmRa~vvaEct-lKeSAaiNqiLGrRWH~LSrEEQAKYyElArKer  249 (421)
T KOG3248|consen  193 KKPLNAFMLYMKEMRAKVVAECT-LKESAAINQILGRRWHALSREEQAKYYELARKER  249 (421)
T ss_pred             cccHHHHHHHHHHHHHHHHHHhh-hhhHHHHHHHHhHHHhhhhHHHHHHHHHHHHHHH
Confidence            89999999999999999999988 8999999999999999999999998888887643


No 16 
>KOG0528 consensus HMG-box transcription factor SOX5 [Transcription]
Probab=97.44  E-value=4.6e-05  Score=74.59  Aligned_cols=63  Identities=14%  Similarity=0.181  Sum_probs=53.9

Q ss_pred             CCCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcCCCccc
Q 039418           31 KDRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGGKANV   94 (297)
Q Consensus        31 ~~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~k~~v   94 (297)
                      -+|||-+||||+-.|=|+-.-...| .=.+..++|.+|..||.||..||.||..--.+-+|.+.
T Consensus       325 HIKRPMNAFMVWAkDERRKILqA~P-DMHNSnISKILGSRWKaMSN~eKQPYYEEQaRLSk~Hl  387 (511)
T KOG0528|consen  325 HIKRPMNAFMVWAKDERRKILQAFP-DMHNSNISKILGSRWKAMSNTEKQPYYEEQARLSKLHL  387 (511)
T ss_pred             cccCCcchhhcccchhhhhhhhcCc-cccccchhHHhcccccccccccccchHHHHHHHHHhhh
Confidence            3599999999999999999999999 45889999999999999999999977665555455554


No 17 
>KOG2746 consensus HMG-box transcription factor Capicua and related proteins [Transcription]
Probab=95.48  E-value=0.0041  Score=63.40  Aligned_cols=62  Identities=15%  Similarity=0.096  Sum_probs=56.8

Q ss_pred             CCCCCcchhhhHHHHH--HHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhcCCCccc
Q 039418           32 DRENNHGFISFFAESV--RQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRRGGKANV   94 (297)
Q Consensus        32 ~~r~~~~f~~~~~~~~--~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~~~k~~v   94 (297)
                      +.||-+||+.|.+-+|  ...++.|| |..+..|+|.+|+.|=+|.+.||.-|++-|.+++-++-
T Consensus       182 irrPMnaf~ifskrhr~~g~vhq~~p-n~DNrtIskiLgewWytL~~~Ekq~yhdLa~Qvk~Ahf  245 (683)
T KOG2746|consen  182 IRRPMNAFHIFSKRHRGEGRVHQRHP-NQDNRTISKILGEWWYTLGPNEKQKYHDLAFQVKEAHF  245 (683)
T ss_pred             hhhhhHHHHHHHhhcCCccchhccCc-cccchhHHHHHhhhHhhhCchhhhhHHHHHHHHHHHHh
Confidence            3899999999999999  99999999 88999999999999999999999999988888774443


No 18 
>KOG4715 consensus SWI/SNF-related matrix-associated actin-dependent regulator of chromatin  [Chromatin structure and dynamics]
Probab=95.38  E-value=0.017  Score=54.26  Aligned_cols=61  Identities=21%  Similarity=0.230  Sum_probs=49.5

Q ss_pred             hcCCCC-CCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhc
Q 039418           27 SRGSKD-RENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRR   88 (297)
Q Consensus        27 ~~~~~~-~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~   88 (297)
                      -+|-|| -||.-.||-|+.-.-.+.|++||+. -.=++||.||..|+-|+++||.-|..-=+.
T Consensus        59 pkpPkppekpl~pymrySrkvWd~VkA~nPe~-kLWeiGK~Ig~mW~dLpd~EK~ey~~EYea  120 (410)
T KOG4715|consen   59 PKPPKPPEKPLMPYMRYSRKVWDQVKASNPEL-KLWEIGKIIGGMWLDLPDEEKQEYLNEYEA  120 (410)
T ss_pred             CCCCCCCCcccchhhHHhhhhhhhhhccCcch-HHHHHHHHHHHHHhhCcchHHHHHHHHHHH
Confidence            335444 5778889999999999999999955 678999999999999999999977654433


No 19 
>PF06382 DUF1074:  Protein of unknown function (DUF1074);  InterPro: IPR024460 This family consists of several proteins which appear to be specific to Insecta. The function of this family is unknown.
Probab=94.97  E-value=0.026  Score=48.92  Aligned_cols=48  Identities=19%  Similarity=0.324  Sum_probs=38.2

Q ss_pred             CCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhh
Q 039418           35 NNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSR   87 (297)
Q Consensus        35 ~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~   87 (297)
                      -.+||+-|+.||++.    |. .-+..++-..+...|..||++||.+|...+.
T Consensus        82 TnnaYLNFLReFRrk----h~-~L~p~dlI~~AAraW~rLSe~eK~rYrr~~~  129 (183)
T PF06382_consen   82 TNNAYLNFLREFRRK----HC-GLSPQDLIQRAARAWCRLSEAEKNRYRRMAP  129 (183)
T ss_pred             cchHHHHHHHHHHHH----cc-CCCHHHHHHHHHHHHHhCCHHHHHHHHhhcc
Confidence            357999999888874    44 3455677777889999999999999998654


No 20 
>PF14887 HMG_box_5:  HMG (high mobility group) box 5; PDB: 1L8Y_A 1L8Z_A 2HDZ_A.
Probab=93.06  E-value=0.15  Score=38.43  Aligned_cols=55  Identities=9%  Similarity=0.048  Sum_probs=43.0

Q ss_pred             CCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhc
Q 039418           32 DRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRR   88 (297)
Q Consensus        32 ~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~   88 (297)
                      |--|-+|==.+.+.-+..|-+.+++... ++ .|+....|++|++.||=+|..+|.+
T Consensus         4 PE~PKt~qe~Wqq~vi~dYla~~~~dr~-K~-~kam~~~W~~me~Kekl~WIkKA~E   58 (85)
T PF14887_consen    4 PETPKTAQEIWQQSVIGDYLAKFRNDRK-KA-LKAMEAQWSQMEKKEKLKWIKKAAE   58 (85)
T ss_dssp             S----THHHHHHHHHHHHHHHHTTSTHH-HH-HHHHHHHHHTTGGGHHHHHHHHHHH
T ss_pred             CCCCCCHHHHHHHHHHHHHHHHhhHhHH-HH-HHHHHHHHHHhhhhhhhHHHHHHHH
Confidence            3445566667888899999999996643 33 6699999999999999999999987


No 21 
>PF04690 YABBY:  YABBY protein;  InterPro: IPR006780 YABBY proteins are a group of plant-specific transcription factors involved in the specification of abaxial polarity in lateral organs such as leaves and floral organs [, ].
Probab=91.71  E-value=0.26  Score=42.76  Aligned_cols=44  Identities=14%  Similarity=0.291  Sum_probs=38.4

Q ss_pred             CCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCCh
Q 039418           32 DRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPV   76 (297)
Q Consensus        32 ~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~   76 (297)
                      ..|-||||-.|+.|=++.+|++|| .-.++++=+++.+-|+..+.
T Consensus       122 RqR~psaYn~f~k~ei~rik~~~p-~ishkeaFs~aAknW~h~ph  165 (170)
T PF04690_consen  122 RQRVPSAYNRFMKEEIQRIKAENP-DISHKEAFSAAAKNWAHFPH  165 (170)
T ss_pred             cCCCchhHHHHHHHHHHHHHhcCC-CCCHHHHHHHHHHhhhhCcc
Confidence            368999999999999999999999 55788888888899987653


No 22 
>COG5648 NHP6B Chromatin-associated proteins containing the HMG domain [Chromatin structure and dynamics]
Probab=90.09  E-value=0.16  Score=45.27  Aligned_cols=56  Identities=16%  Similarity=0.119  Sum_probs=47.8

Q ss_pred             CCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhhhhhhhhhc
Q 039418           32 DRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKCQYKFQSRR   88 (297)
Q Consensus        32 ~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~~~~~~~~~   88 (297)
                      +++|+..|+-+-++-|......+| .++....+|++|+.|++|++.-|++|.+.++.
T Consensus       144 ~~~~~~~~~e~~~~~r~~~~~~~~-~~~~~e~~k~~~~~w~el~~skK~~~~~~~Kk  199 (211)
T COG5648         144 NKAPIGPFIENEPKIRPKVEGPSP-DKALVEETKIISKAWSELDESKKKKYIDKYKK  199 (211)
T ss_pred             CCCCCchhhhccHHhccccCCCCc-chhhhHHhhhhhhhhhhhChhhhhHHHHHHHH
Confidence            467777777777778888888888 66888899999999999999999999998875


No 23 
>PF03078 ATHILA:  ATHILA ORF-1 family;  InterPro: IPR004312 ATHILA is a group of Arabidopsis thaliana retrotransposons [] belonging to the Ty3/gypsy family of the long terminal repeat (LTR) class of eukaryotic retrotransposons[, ]. The central region of ATHILA retrotransposons contains two or three open reading frames (ORFs). This family represents the ORF1 product. The function of ORF1 is unknown.
Probab=76.17  E-value=51  Score=33.12  Aligned_cols=168  Identities=17%  Similarity=0.185  Sum_probs=94.5

Q ss_pred             eec-CHHHHHHHHhhcCHHHHHHHHhcCcccccccccccccHHHHHHHHhccc-------c--------CcceEEECCeE
Q 039418          101 TRC-APDRLAALVSHLIEKQRKAVCDIGLGSIIDLKCGRLKRKLCAWLVERID-------T--------ARCVLQLNGHE  164 (297)
Q Consensus       101 trc-S~~~~~~~i~~Ls~~qk~~I~~~GFg~LL~i~~~~l~~~L~~wL~~~~d-------~--------~t~~~~i~g~~  164 (297)
                      ||. ++.-+..+  .|.++-..+++.+|.+.|..++...-+...+..|+..-=       +        ..-+|.|.|..
T Consensus        63 TRyp~~etl~~L--Gl~~dV~~lf~~~gL~~f~~~~~~~Y~eet~qFLaTl~v~~~~~~~~~~~e~~glG~l~F~V~~~~  140 (458)
T PF03078_consen   63 TRYPDPETLQKL--GLLEDVEYLFKKCGLGTFMSYPYPTYPEETRQFLATLKVTFYNPSEPRAKELDGLGYLTFFVYGVE  140 (458)
T ss_pred             cccCCHHHHHHh--ccHHHHHHHHHhcCchhhccCCCCCcHHHHHHhhheeeeeecccccchhhcccCcceEEEEEccee
Confidence            443 33444444  667888889999999999988886655544444443211       1        23567778999


Q ss_pred             eecCccchhheeccccCCccccccCChhHHHHHHhhhccCCCccchHHHHHHHhcCCCCCchhhhHHHhhhhcceeCCCC
Q 039418          165 LELSPNSFGYIMGVTDGGMPMELQGDSAEVAAYLDKFNATSRGINIKTMEDILLTSKDADNDFKVAFMLFTLCTLLCPPG  244 (297)
Q Consensus       165 i~iT~~dV~~VLGLP~gG~~v~~~~~~~~~~~l~~~~~~~~~~isl~~L~~~ll~~~~~~d~f~r~Fll~~i~~~L~Ptt  244 (297)
                      ..+|-.+...++|+|.|+. +...-..++...|-...|... .++...-...      ..-.=+.+|+--+++..|+|..
T Consensus       141 y~lsi~~L~~i~GF~~~~~-i~~~~~~~el~~~W~~ig~~~-p~~~~~~ks~------~Ir~PviRy~hr~iA~tlf~R~  212 (458)
T PF03078_consen  141 YSLSIKHLERIFGFPSGDE-IKPDFDPEELNDFWATIGGGK-PFNSARSKSN------QIRSPVIRYFHRLIANTLFARE  212 (458)
T ss_pred             eeeeHHHHHHHhCCCCccc-cCCCCCchHHHHHHHHhcCCC-cccccccccc------cccChHHHHHHHHHHhhhcccc
Confidence            9999999999999999854 332223344444444444220 0111000010      1112234445555666666665


Q ss_pred             CC-ccCcchhhhh-----------hcc----ccCcccchhHHHHHHHHHH
Q 039418          245 GV-HISYSFLFTL-----------KDV----HSIRNRNWATFCFERLMRG  278 (297)
Q Consensus       245 ~~-~vs~~yl~~l-----------~D~----~~I~~ynW~~~Vld~L~~~  278 (297)
                      .. .|..+-|.++           .|.    .+..+.+-+-..++||...
T Consensus       213 ~~~~v~~~El~~l~~~L~~~Lr~~~~g~~l~~d~~dt~~~~vl~~hL~~y  262 (458)
T PF03078_consen  213 ETGTVRNDELEMLDQALKHLLRRTKDGKLLRGDLNDTNVSMVLLDHLCSY  262 (458)
T ss_pred             ccCceechhHHHHHHHHHHHHHhcCCCccccCcccccchhHHHHHHHHhh
Confidence            44 6776665542           111    1135556666666666654


No 24 
>PF11304 DUF3106:  Protein of unknown function (DUF3106);  InterPro: IPR021455  Some members in this family of proteins are annotated as transmembrane proteins however this cannot be confirmed. Currently no function is known. 
Probab=65.07  E-value=8  Score=30.91  Aligned_cols=21  Identities=14%  Similarity=0.295  Sum_probs=11.1

Q ss_pred             HHhhhcCCChHhhhhhhhhhh
Q 039418           67 IRNAFKNLPVEEKCQYKFQSR   87 (297)
Q Consensus        67 ~g~~wk~ls~~ek~~~~~~~~   87 (297)
                      +.+.|.+||++|++.+..+..
T Consensus        34 ~a~r~~~mspeqq~r~~~rm~   54 (107)
T PF11304_consen   34 IAERWPSMSPEQQQRLRERMR   54 (107)
T ss_pred             HHHHHhcCCHHHHHHHHHHHH
Confidence            555556666665554444433


No 25 
>PF04769 MAT_Alpha1:  Mating-type protein MAT alpha 1;  InterPro: IPR006856 This family includes Saccharomyces cerevisiae (Baker's yeast) mating type protein alpha 1 (P01365 from SWISSPROT). MAT alpha 1 is a transcription activator that activates mating-type alpha-specific genes with the help of the MADS-box containing MCM1 transcription factor, which together bind cooperatively to PQ elements upstream of alpha-specific genes. The MCM1-MATalpha1 complex is required for the proper DNA-bending that is needed for transcriptional activation []. Alpha 1 interacts in vivo with STE12, linking expression of alpha-specific genes to the alpha-pheromone (IPR006742 from INTERPRO) response pathway [].; GO: 0000772 mating pheromone activity, 0003677 DNA binding, 0045895 positive regulation of transcription, mating-type specific, 0005634 nucleus
Probab=61.82  E-value=17  Score=32.38  Aligned_cols=44  Identities=14%  Similarity=0.187  Sum_probs=33.6

Q ss_pred             CCCCCCCCcchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChH
Q 039418           29 GSKDRENNHGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVE   77 (297)
Q Consensus        29 ~~~~~r~~~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~   77 (297)
                      ..+++||.++||.|+.    -|+...|. --.+.++..++..|..=+..
T Consensus        41 ~~~~kr~lN~Fm~FRs----yy~~~~~~-~~Qk~~S~~l~~lW~~dp~k   84 (201)
T PF04769_consen   41 PEKAKRPLNGFMAFRS----YYSPIFPP-LPQKELSGILTKLWEKDPFK   84 (201)
T ss_pred             ccccccchhHHHHHHH----HHHhhcCC-cCHHHHHHHHHHHHhCCccH
Confidence            3567999999998765    45556663 45899999999999984443


No 26 
>PF08073 CHDNT:  CHDNT (NUC034) domain;  InterPro: IPR012958 The CHD N-terminal domain is found in PHD/RING fingers and chromo domain-associated helicases [].; GO: 0003677 DNA binding, 0005524 ATP binding, 0008270 zinc ion binding, 0016818 hydrolase activity, acting on acid anhydrides, in phosphorus-containing anhydrides, 0006355 regulation of transcription, DNA-dependent, 0005634 nucleus
Probab=44.48  E-value=25  Score=24.87  Aligned_cols=39  Identities=8%  Similarity=0.148  Sum_probs=32.0

Q ss_pred             cchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCCh
Q 039418           37 HGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPV   76 (297)
Q Consensus        37 ~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~   76 (297)
                      +.+=.|.+--|++..++||.. .+..+-.-++.|||.-++
T Consensus        14 t~yK~Fsq~vRP~l~~~NPk~-~~sKl~~l~~AKwrEF~~   52 (55)
T PF08073_consen   14 TNYKAFSQHVRPLLAKANPKA-PMSKLMMLLQAKWREFQE   52 (55)
T ss_pred             HHHHHHHHHHHHHHHHHCCCC-cHHHHHHHHHHHHHHHHh
Confidence            446679999999999999954 677778889999997654


No 27 
>PF11943 DUF3460:  Protein of unknown function (DUF3460);  InterPro: IPR021853  This family of proteins are functionally uncharacterised. This protein is found in bacteria. Proteins in this family are about 70 amino acids in length. This protein has a conserved WDK sequence motif. 
Probab=41.25  E-value=57  Score=23.51  Aligned_cols=38  Identities=26%  Similarity=0.357  Sum_probs=26.4

Q ss_pred             HHHHHHHHHhCCCCcchhHHHHHHHhhh--cCCChHhhhhhhh
Q 039418           44 AESVRQLKAKDGRACITNEVRKEIRNAF--KNLPVEEKCQYKF   84 (297)
Q Consensus        44 ~~~~~~~~~~~~~~~~~~~v~k~~g~~w--k~ls~~ek~~~~~   84 (297)
                      ..|..+||++||+.   .+=.+++...|  |-++.+|.+.|.+
T Consensus         8 TqFl~~lk~~~Pel---e~~Q~~GRallWDk~~d~e~~~~~~~   47 (60)
T PF11943_consen    8 TQFLNQLKAKHPEL---EEEQRAGRALLWDKPQDLEEQARFRA   47 (60)
T ss_pred             HHHHHHHHHhCCch---HHHHHHhhHHhcCCCCCHHHHHHHHh
Confidence            46899999999954   44455555444  5788888776543


No 28 
>PF10234 Cluap1:  Clusterin-associated protein-1;  InterPro: IPR019366 This protein of 413 amino acids contains a central coiled-coil domain, possibly the region that binds to clusterin. Cluap1 expression is highest in the nucleus and gradually increases during late S to G2/M phases of the cell cycle and returns to the basal level in the G0/G1 phases. In addition, it is upregulated in colon cancer tissues compared to corresponding non-cancerous mucosa. It thus plays a crucial role in the life of the cell []. 
Probab=39.29  E-value=15  Score=34.22  Aligned_cols=32  Identities=25%  Similarity=0.461  Sum_probs=25.4

Q ss_pred             HhcCccccccccccccc-----HHHHHHHHhccccCc
Q 039418          124 CDIGLGSIIDLKCGRLK-----RKLCAWLVERIDTAR  155 (297)
Q Consensus       124 ~~~GFg~LL~i~~~~l~-----~~L~~wL~~~~d~~t  155 (297)
                      +.+||--++.|.+..-|     -+++.||+.+|||+.
T Consensus         1 R~LGypr~iSmenFrtPNF~LVAeiL~WLv~rydP~~   37 (267)
T PF10234_consen    1 RALGYPRLISMENFRTPNFELVAEILRWLVKRYDPDA   37 (267)
T ss_pred             CCCCCCCCCcHHHcCCCChHHHHHHHHHHHHHcCCCC
Confidence            35799999999775444     478889999999985


No 29 
>COG5202 Predicted membrane protein [Function unknown]
Probab=39.23  E-value=19  Score=35.15  Aligned_cols=27  Identities=26%  Similarity=0.540  Sum_probs=23.8

Q ss_pred             CCchh--hhHHHHHhHhhcCCCCCCCCcc
Q 039418           12 GDFES--SCEMLVDAHRSRGSKDRENNHG   38 (297)
Q Consensus        12 ~~~~~--~~~~~~~~~~~~~~~~~r~~~~   38 (297)
                      .||||  .|||.-+--.+.+-+|+|||++
T Consensus       189 ~dfd~~lt~~mFr~l~e~pyEr~~~P~n~  217 (512)
T COG5202         189 ADFDSWLTCEMFRSLMENPYERPKRPPNG  217 (512)
T ss_pred             chhhhhhhHHHHhhhhcCcccccCCCCCc
Confidence            47787  7999999999999999999975


No 30 
>PF06945 DUF1289:  Protein of unknown function (DUF1289);  InterPro: IPR010710 This family consists of a number of hypothetical bacterial proteins. The aligned region spans around 56 residues and contains 4 highly conserved cysteine residues towards the N terminus. The function of this family is unknown.
Probab=33.56  E-value=21  Score=24.53  Aligned_cols=20  Identities=15%  Similarity=0.218  Sum_probs=15.0

Q ss_pred             hhhcCCChHhhhhhhhhhhc
Q 039418           69 NAFKNLPVEEKCQYKFQSRR   88 (297)
Q Consensus        69 ~~wk~ls~~ek~~~~~~~~~   88 (297)
                      ..|+.||++||..-.++..+
T Consensus        28 ~~W~~~s~~er~~i~~~l~~   47 (51)
T PF06945_consen   28 RDWKSMSDDERRAILARLRA   47 (51)
T ss_pred             HHHhhCCHHHHHHHHHHHHH
Confidence            35999999998866655543


No 31 
>PF01418 HTH_6:  Helix-turn-helix domain, rpiR family;  InterPro: IPR000281 This domain contains a helix-turn-helix motif []. Every member of this family is N-terminal to a SIS domain IPR001347 from INTERPRO. Members of this family are probably regulators of genes involved in phosphosugar metobolism.; GO: 0003700 sequence-specific DNA binding transcription factor activity, 0006355 regulation of transcription, DNA-dependent; PDB: 2O3F_B 3IWF_B.
Probab=31.20  E-value=33  Score=25.32  Aligned_cols=60  Identities=15%  Similarity=0.325  Sum_probs=34.1

Q ss_pred             HHHHhhhcCCChHhhh--hhhhh-hhcCCCccccccceeeecCHHHHHHHHhhcCHHHHHHHHhcCcccccccc
Q 039418           65 KEIRNAFKNLPVEEKC--QYKFQ-SRRGGKANVEKVKFLTRCAPDRLAALVSHLIEKQRKAVCDIGLGSIIDLK  135 (297)
Q Consensus        65 k~~g~~wk~ls~~ek~--~~~~~-~~~~~k~~v~k~~~~trcS~~~~~~~i~~Ls~~qk~~I~~~GFg~LL~i~  135 (297)
                      ..+...+.+||+.|+.  .|.-+ -.+.....+....-.+-.|+..+..+           ++++||.|+-++.
T Consensus         5 ~~i~~~~~~ls~~e~~Ia~yil~~~~~~~~~si~elA~~~~vS~sti~Rf-----------~kkLG~~gf~efk   67 (77)
T PF01418_consen    5 EKIRSQYNSLSPTEKKIADYILENPDEIAFMSISELAEKAGVSPSTIVRF-----------CKKLGFSGFKEFK   67 (77)
T ss_dssp             HHHHHHGGGS-HHHHHHHHHHHH-HHHHCT--HHHHHHHCTS-HHHHHHH-----------HHHCTTTCHHHHH
T ss_pred             HHHHHHHhhCCHHHHHHHHHHHhCHHHHHHccHHHHHHHcCCCHHHHHHH-----------HHHhCCCCHHHHH
Confidence            4556677888988876  33333 22333333344444455566666655           7888999987764


No 32 
>PF13875 DUF4202:  Domain of unknown function (DUF4202)
Probab=28.82  E-value=79  Score=27.90  Aligned_cols=64  Identities=16%  Similarity=0.226  Sum_probs=42.6

Q ss_pred             cCCchhhhHHHHHhHhhcCCCCCCCC-------cchhhhHHHHHHHHHHhCCCCcchhHHHHHHHhhhcCCChHhhh
Q 039418           11 AGDFESSCEMLVDAHRSRGSKDRENN-------HGFISFFAESVRQLKAKDGRACITNEVRKEIRNAFKNLPVEEKC   80 (297)
Q Consensus        11 ~~~~~~~~~~~~~~~~~~~~~~~r~~-------~~f~~~~~~~~~~~~~~~~~~~~~~~v~k~~g~~wk~ls~~ek~   80 (297)
                      +|-=+..|+-....-|.+|  .|+-|       -+=+||++.+-..|.++|...|.+    .++.+-|+-||+.-++
T Consensus        99 ~Gy~~~~i~rV~~lv~K~~--lk~d~e~Q~LEDvacLVFL~~~f~~F~~~~deeK~v----~Il~KTw~KMS~~g~~  169 (185)
T PF13875_consen   99 AGYDEEEIDRVAALVRKEG--LKRDPETQALEDVACLVFLEYYFEDFAAKHDEEKIV----DILRKTWRKMSERGHE  169 (185)
T ss_pred             CCCCHHHHHHHHHHHHhcc--CCCCchHHHHHhhHHHHhHHHHHHHHHhcCCHHHHH----HHHHHHHHHCCHHHHH
Confidence            4444555555555555544  34433       367899999999999999555544    4556679999998554


No 33 
>PF14513 DAG_kinase_N:  Diacylglycerol kinase N-terminus; PDB: 1TUZ_A.
Probab=28.52  E-value=11  Score=31.76  Aligned_cols=72  Identities=15%  Similarity=0.191  Sum_probs=39.1

Q ss_pred             hhhcCCChHhhhhhhhhhhcCCCccccccceeeecCHHHHHHHHhhcCHH-------HHHHHHhcCccccccccc-cccc
Q 039418           69 NAFKNLPVEEKCQYKFQSRRGGKANVEKVKFLTRCAPDRLAALVSHLIEK-------QRKAVCDIGLGSIIDLKC-GRLK  140 (297)
Q Consensus        69 ~~wk~ls~~ek~~~~~~~~~~~k~~v~k~~~~trcS~~~~~~~i~~Ls~~-------qk~~I~~~GFg~LL~i~~-~~l~  140 (297)
                      ++|.+||++|=+.-..-+               .-|.+++.+++..+.++       +.+-|.--||.-+|.+-. ..+|
T Consensus         3 ~~~~~lsp~eF~qLq~y~---------------eys~kklkdvl~eF~~~g~~~~~~~~~~Id~egF~~Fm~~yLe~d~P   67 (138)
T PF14513_consen    3 KEWVSLSPEEFAQLQKYS---------------EYSTKKLKDVLKEFHGDGSLAKYNPEEPIDYEGFKLFMKTYLEVDLP   67 (138)
T ss_dssp             ---S-S-HHHHHHHHHHH---------------HH----HHHHHHHH-HTSGGGGGEETTEE-HHHHHHHHHHHTT-S--
T ss_pred             cceeccCHHHHHHHHHHH---------------HHHHHHHHHHHHHHhcCCcccccCCCCCcCHHHHHHHHHHHHcCCCC
Confidence            579999999855433333               34677888888888644       234677778888888766 5599


Q ss_pred             HHHHHHHHhccccCc
Q 039418          141 RKLCAWLVERIDTAR  155 (297)
Q Consensus       141 ~~L~~wL~~~~d~~t  155 (297)
                      .+||..|---|....
T Consensus        68 ~~lc~hLF~sF~~~~   82 (138)
T PF14513_consen   68 EDLCQHLFLSFQKKP   82 (138)
T ss_dssp             HHHHHHHHHHS----
T ss_pred             HHHHHHHHHHHhCcc
Confidence            999999988877443


No 34 
>cd09071 FAR_C C-terminal domain of fatty acyl CoA reductases. C-terminal domain of fatty acyl CoA reductases, a family of SDR-like proteins. SDRs or short-chain dehydrogenases/reductases are Rossmann-fold NAD(P)H-binding proteins. Many proteins in this FAR_C family may function as fatty acyl-CoA reductases (FARs), acting on medium and long chain fatty acids, and have been reported to be involved in diverse processes such as the biosynthesis of insect pheromones, plant cuticular wax production, and mammalian wax biosynthesis. In Arabidopsis thaliana, proteins with this particular architecture have also been identified as the MALE STERILITY 2 (MS2) gene product, which is implicated in male gametogenesis. Mutations in MS2 inhibit the synthesis of exine (sporopollenin), rendering plants unable to reduce pollen wall fatty acids to corresponding alcohols. The function of this C-terminal domain is unclear.
Probab=28.37  E-value=55  Score=24.52  Aligned_cols=21  Identities=24%  Similarity=0.708  Sum_probs=18.4

Q ss_pred             cCcccchhHHHHHHHHHHHHhh
Q 039418          261 SIRNRNWATFCFERLMRGITRY  282 (297)
Q Consensus       261 ~I~~ynW~~~Vld~L~~~l~k~  282 (297)
                      ++.++||..++.++ +.|+++|
T Consensus        70 D~~~idW~~Y~~~~-~~G~r~y   90 (92)
T cd09071          70 DIRSIDWDDYFENY-IPGLRKY   90 (92)
T ss_pred             CCCCCCHHHHHHHH-HHHHHHH
Confidence            46899999999999 8888876


No 35 
>PF05494 Tol_Tol_Ttg2:  Toluene tolerance, Ttg2 ;  InterPro: IPR008869 Toluene tolerance is mediated by increased cell membrane rigidity resulting from changes in fatty acid and phospholipid compositions, exclusion of toluene from the cell membrane, and removal of intracellular toluene by degradation []. Many proteins are involved in these processes. This family is a transporter which shows similarity to ABC transporters [].; PDB: 2QGU_A.
Probab=26.97  E-value=48  Score=28.08  Aligned_cols=30  Identities=0%  Similarity=0.066  Sum_probs=20.7

Q ss_pred             chhHHHH-HHHhhhcCCChHhhhhhhhhhhc
Q 039418           59 ITNEVRK-EIRNAFKNLPVEEKCQYKFQSRR   88 (297)
Q Consensus        59 ~~~~v~k-~~g~~wk~ls~~ek~~~~~~~~~   88 (297)
                      ....+++ ++|.-|+.+|++|++.|...=++
T Consensus        39 D~~~~ar~~LG~~w~~~s~~q~~~F~~~f~~   69 (170)
T PF05494_consen   39 DFERMARRVLGRYWRKASPAQRQRFVEAFKQ   69 (170)
T ss_dssp             -HHHHHHHHHGGGTTTS-HHHHHHHHHHHHH
T ss_pred             CHHHHHHHHHHHhHhhCCHHHHHHHHHHHHH
Confidence            4444444 78989999999999977665444


No 36 
>PF03457 HA:  Helicase associated domain;  InterPro: IPR005114 This short domain is found in multiple copies in bacterial helicase proteins. The domain is predicted to contain 3 alpha helices. The function of this domain may be to bind nucleic acid.; PDB: 2KTA_A.
Probab=26.40  E-value=42  Score=23.90  Aligned_cols=16  Identities=13%  Similarity=0.304  Sum_probs=11.2

Q ss_pred             hhcCHHHHHHHHhcCc
Q 039418          113 SHLIEKQRKAVCDIGL  128 (297)
Q Consensus       113 ~~Ls~~qk~~I~~~GF  128 (297)
                      ..|+++|.+.++++||
T Consensus        52 g~L~~er~~~L~~lg~   67 (68)
T PF03457_consen   52 GKLTPERIERLDALGF   67 (68)
T ss_dssp             T---HHHHHHHHHHT-
T ss_pred             CCCCHHHHHHHHcCCC
Confidence            4599999999999998


No 37 
>PF07970 COPIIcoated_ERV:  Endoplasmic reticulum vesicle transporter ;  InterPro: IPR012936 This domain occurs in many hypothetical proteins, and also two partially characterised proteins. One of these proteins, PTX1 Q96RQ1 from SWISSPROT, is a homeodomain-containing transcription factor involved in regulating all pituitary hormone genes []. This protein is down regulated in prostate carcinoma []. The other protein, ERGIC-32 Q969X5 from SWISSPROT, is involved in protein transport from the ER to the Golgi [].
Probab=25.82  E-value=71  Score=28.62  Aligned_cols=35  Identities=23%  Similarity=0.310  Sum_probs=23.2

Q ss_pred             hhhHHHHHhHhhcCCCCCCCCcchhhhHHHHHHHHH
Q 039418           16 SSCEMLVDAHRSRGSKDRENNHGFISFFAESVRQLK   51 (297)
Q Consensus        16 ~~~~~~~~~~~~~~~~~~r~~~~f~~~~~~~~~~~~   51 (297)
                      .+||-+.+||+.+|.+++.+- .+=-..+|+.++.+
T Consensus        13 nTC~~V~~ay~~~~w~~~~~~-~~eQC~~~~~~~~~   47 (222)
T PF07970_consen   13 NTCEDVREAYRKKGWAFPDLE-NIEQCRREYVKKIK   47 (222)
T ss_pred             cCHHHHHHHHHHhCCCCCCcc-ccccccchhhhhhh
Confidence            589999999999999776654 33333334333333


No 38 
>TIGR03481 HpnM hopanoid biosynthesis associated membrane protein HpnM. The genomes containing members of this family share the machinery for the biosynthesis of hopanoid lipids. Furthermore, the genes of this family are usually located proximal to other components of this biological process. The proteins are members of the pfam05494 family of putative transporters known as "toluene tolerance protein Ttg2D", although it is unlikely that the members included here have anything to do with toluene per-se.
Probab=25.59  E-value=47  Score=29.33  Aligned_cols=30  Identities=10%  Similarity=0.177  Sum_probs=22.9

Q ss_pred             chhHHHH-HHHhhhcCCChHhhhhhhhhhhc
Q 039418           59 ITNEVRK-EIRNAFKNLPVEEKCQYKFQSRR   88 (297)
Q Consensus        59 ~~~~v~k-~~g~~wk~ls~~ek~~~~~~~~~   88 (297)
                      ....+++ ++|..|+.+|+++|+.|.+.=++
T Consensus        65 Df~~mar~vLG~~W~~~s~~Qr~~F~~~F~~   95 (198)
T TIGR03481        65 DLPAMARLTLGSSWTSLSPEQRRRFIGAFRE   95 (198)
T ss_pred             CHHHHHHHHhhhhhhhCCHHHHHHHHHHHHH
Confidence            4555554 88999999999999977765443


No 39 
>cd07321 Extradiol_Dioxygenase_3A_like Subunit A of Class III extradiol dioxygenases. Extradiol dioxygenases catalyze the incorporation of both atoms of molecular oxygen into substrates using a variety of reaction mechanisms, resulting in the cleavage of aromatic rings.  There are two major groups of dioxygenases according to the cleavage site of the aromatic ring. Intradiol enzymes cleave the aromatic ring between two hydroxyl groups, whereas extradiol enzymes cleave the aromatic ring between a hydroxylated carbon and an adjacent non-hydroxylated carbon. Extradiol dioxygenases can be divided into three classes. Class I and II enzymes are evolutionary related and show sequence similarity, with the two domain class II enzymes evolving from the class I enzyme through gene duplication. Class III enzymes are different in sequence and structure and usually have two subunits, designated A and B, which form a tetramer composed of two copies of each subunit. This model represents subunit A of c
Probab=25.19  E-value=69  Score=23.93  Aligned_cols=32  Identities=16%  Similarity=0.210  Sum_probs=26.1

Q ss_pred             hcCHHHHHHHHhcCcccccccccccccHHHHHHHH
Q 039418          114 HLIEKQRKAVCDIGLGSIIDLKCGRLKRKLCAWLV  148 (297)
Q Consensus       114 ~Ls~~qk~~I~~~GFg~LL~i~~~~l~~~L~~wL~  148 (297)
                      .||++|+++|.+--+.+|+++..   |..++.++.
T Consensus        34 ~Lt~eE~~al~~rD~~~L~~lG~---~~~~l~k~~   65 (77)
T cd07321          34 GLTPEEKAALLARDVGALYVLGV---NPMLLMHFA   65 (77)
T ss_pred             CCCHHHHHHHHcCCHHHHHHcCC---CHHHHHHHH
Confidence            89999999999999999999874   555555554


No 40 
>PRK15117 ABC transporter periplasmic binding protein MlaC; Provisional
Probab=24.93  E-value=47  Score=29.63  Aligned_cols=25  Identities=12%  Similarity=0.114  Sum_probs=20.4

Q ss_pred             HHHHHhhhcCCChHhhhhhhhhhhc
Q 039418           64 RKEIRNAFKNLPVEEKCQYKFQSRR   88 (297)
Q Consensus        64 ~k~~g~~wk~ls~~ek~~~~~~~~~   88 (297)
                      ..++|..|+..|+++|+.|.+.=++
T Consensus        75 ~~vLG~~wr~as~eQr~~F~~~F~~   99 (211)
T PRK15117         75 ALVLGRYYKDATPAQREAYFAAFRE   99 (211)
T ss_pred             HHHhhhhhhhCCHHHHHHHHHHHHH
Confidence            4489999999999999988765544


No 41 
>cd02988 Phd_like_VIAF Phosducin (Phd)-like family, Viral inhibitor of apoptosis (IAP)-associated factor (VIAF) subfamily; VIAF is a Phd-like protein that functions in caspase activation during apoptosis. It was identified as an IAP binding protein through a screen of a human B-cell library using a prototype IAP. VIAF lacks a consensus IAP binding motif and while it does not function as an IAP antagonist, it still plays a regulatory role in the complete activation of caspases. VIAF itself is a substrate for IAP-mediated ubiquitination, suggesting that it may be a target of IAPs in the prevention of cell death. The similarity of VIAF to Phd points to a potential role distinct from apoptosis regulation. Phd functions as a cytosolic regulator of G protein by specifically binding to G protein betagamma (Gbg)-subunits. The C-terminal domain of Phd adopts a thioredoxin fold, but it does not contain a CXXC motif. Phd interacts with G protein beta mostly through the N-terminal helical domain.
Probab=23.80  E-value=48  Score=29.12  Aligned_cols=17  Identities=18%  Similarity=0.120  Sum_probs=14.3

Q ss_pred             HHhHhhcCCCCCCCCcc
Q 039418           22 VDAHRSRGSKDRENNHG   38 (297)
Q Consensus        22 ~~~~~~~~~~~~r~~~~   38 (297)
                      =|++|.+|+-|+|||+.
T Consensus         4 ~di~r~~g~~p~~~~~~   20 (192)
T cd02988           4 NDILRKKGILPPKPPSP   20 (192)
T ss_pred             hHHHHHcCCCCCCCCCC
Confidence            38899999999999743


No 42 
>PF00701 DHDPS:  Dihydrodipicolinate synthetase family;  InterPro: IPR002220 Dihydropicolinate synthase (DHDPS) is the key enzyme in lysine biosynthesis via the diaminopimelate pathway of prokaryotes, some phycomycetes and higher plants. The enzyme catalyses the condensation of L-aspartate-beta- semialdehyde and pyruvate to dihydropicolinic acid via a ping-pong mechanism in which pyruvate binds to the enzyme by forming a Schiff-base with a lysine residue []. Three other proteins are structurally related to DHDPS and probably also act via a similar catalytic mechanism. These are Escherichia coli N-acetylneuraminate lyase (4.1.3.3 from EC) (gene nanA), which catalyzes the condensation of N-acetyl-D-mannosamine and pyruvate to form N-acetylneuraminate; Rhizobium meliloti (Sinorhizobium meliloti) protein mosA [], which is involved in the biosynthesis of the rhizopine 3-o-methyl-scyllo-inosamine; and E. coli hypothetical protein yjhH. The sequences of DHDPS from different sources are well-conserved. The structure takes the form of a homotetramer, in which 2 monomers are related by an approximate 2-fold symmetry []. Each monomer comprises 2 domains: an 8-fold alpha-/beta-barrel, and a C-terminal alpha-helical domain. The fold resembles that of N-acetylneuraminate lyase. The active site lysine is located in the barrel domain, and has access via 2 channels on the C-terminal side of the barrel.; GO: 0016829 lyase activity, 0008152 metabolic process; PDB: 3B4U_B 3S8H_A 3QZE_B 1XXX_F 3L21_F 3IRD_A 3A5F_B 3G0S_B 3DAQ_C 3UQN_A ....
Probab=23.22  E-value=66  Score=29.64  Aligned_cols=102  Identities=13%  Similarity=0.141  Sum_probs=63.9

Q ss_pred             hhcCCChHhhhhhhhhhhcCCCccccccceeeecCHHHHHHHHhhcCHHHHHHHHhcCccccccccccc--c-cHHHHHH
Q 039418           70 AFKNLPVEEKCQYKFQSRRGGKANVEKVKFLTRCAPDRLAALVSHLIEKQRKAVCDIGLGSIIDLKCGR--L-KRKLCAW  146 (297)
Q Consensus        70 ~wk~ls~~ek~~~~~~~~~~~k~~v~k~~~~trcS~~~~~~~i~~Ls~~qk~~I~~~GFg~LL~i~~~~--l-~~~L~~w  146 (297)
                      .+-+||.+||....+-+.+..+.+++--.-....|.....+..+        ..+++|+.+++-++...  . ...+..|
T Consensus        47 E~~~Lt~~Er~~l~~~~~~~~~~~~~vi~gv~~~st~~~i~~a~--------~a~~~Gad~v~v~~P~~~~~s~~~l~~y  118 (289)
T PF00701_consen   47 EFYSLTDEERKELLEIVVEAAAGRVPVIAGVGANSTEEAIELAR--------HAQDAGADAVLVIPPYYFKPSQEELIDY  118 (289)
T ss_dssp             TGGGS-HHHHHHHHHHHHHHHTTSSEEEEEEESSSHHHHHHHHH--------HHHHTT-SEEEEEESTSSSCCHHHHHHH
T ss_pred             ccccCCHHHHHHHHHHHHHHccCceEEEecCcchhHHHHHHHHH--------HHhhcCceEEEEeccccccchhhHHHHH
Confidence            46789999999999888887666655444455557766665543        46789999998886632  2 2456666


Q ss_pred             HHhccccCcceEEECC----eEeecCccchhheeccc
Q 039418          147 LVERIDTARCVLQLNG----HELELSPNSFGYIMGVT  179 (297)
Q Consensus       147 L~~~~d~~t~~~~i~g----~~i~iT~~dV~~VLGLP  179 (297)
                      ..+--+....-+.+.+    ....++++.+..+..+|
T Consensus       119 ~~~ia~~~~~pi~iYn~P~~tg~~ls~~~l~~L~~~~  155 (289)
T PF00701_consen  119 FRAIADATDLPIIIYNNPARTGNDLSPETLARLAKIP  155 (289)
T ss_dssp             HHHHHHHSSSEEEEEEBHHHHSSTSHHHHHHHHHTST
T ss_pred             HHHHHhhcCCCEEEEECCCccccCCCHHHHHHHhcCC
Confidence            5555554445555532    23566666666666555


No 43 
>PF12650 DUF3784:  Domain of unknown function (DUF3784);  InterPro: IPR017259 This group represents an uncharacterised conserved protein.
Probab=23.22  E-value=45  Score=25.65  Aligned_cols=17  Identities=24%  Similarity=0.470  Sum_probs=13.1

Q ss_pred             hhcCCChHhhhhhhhhh
Q 039418           70 AFKNLPVEEKCQYKFQS   86 (297)
Q Consensus        70 ~wk~ls~~ek~~~~~~~   86 (297)
                      -+++||++||+.|-.+.
T Consensus        25 Gyntms~eEk~~~D~~~   41 (97)
T PF12650_consen   25 GYNTMSKEEKEKYDKKK   41 (97)
T ss_pred             hcccCCHHHHHHhhHHH
Confidence            37899999999775543


No 44 
>PF03015 Sterile:  Male sterility protein;  InterPro: IPR004262 This family represents the C-terminal region of the male sterility protein in a number of organisms. The Arabidopsis thaliana male sterility 2 (MS2) protein is involved in male gametogenesis. The MS2 protein shows sequence similarity to a jojoba protein (also a member of this group) that converts wax fatty acids to fatty alcohols. It has been suggested that a possible function of the MS2 protein may be as a fatty acyl reductase in the formation of pollen wall substances [].; GO: 0016620 oxidoreductase activity, acting on the aldehyde or oxo group of donors, NAD or NADP as acceptor, 0055114 oxidation-reduction process
Probab=22.32  E-value=85  Score=23.83  Aligned_cols=54  Identities=17%  Similarity=0.253  Sum_probs=31.9

Q ss_pred             hhhHHHhhhhcceeCCCCCC-----ccCcchhhhhhccccCcccchhHHHHHHHHHHHHhhh
Q 039418          227 FKVAFMLFTLCTLLCPPGGV-----HISYSFLFTLKDVHSIRNRNWATFCFERLMRGITRYK  283 (297)
Q Consensus       227 f~r~Fll~~i~~~L~Ptt~~-----~vs~~yl~~l~D~~~I~~ynW~~~Vld~L~~~l~k~~  283 (297)
                      ...++-.|+.....+.+.+.     ..++..-..+ +. +++++||-.++.++ +.|+++|-
T Consensus        33 ~~~~~~~F~~~eW~F~~~n~~~L~~~l~~~D~~~F-~f-D~~~idW~~Y~~~~-~~G~rkyl   91 (94)
T PF03015_consen   33 ALEVLEYFTTNEWIFDNDNTRRLWERLSPEDREIF-NF-DIRSIDWEEYFRNY-IPGIRKYL   91 (94)
T ss_pred             HHHHHHHHHhCceeecchHHHHHHHhCchhcCcee-cC-CCCCCCHHHHHHHH-HHHHHHHH
Confidence            33444455555555544432     1233333222 12 57899999999999 88998874


No 45 
>smart00271 DnaJ DnaJ molecular chaperone homology domain.
Probab=22.10  E-value=99  Score=20.86  Aligned_cols=36  Identities=17%  Similarity=0.022  Sum_probs=25.0

Q ss_pred             hHHHHHHHHHHhCCCCcc-----hhHHHHHHHhhhcCCChH
Q 039418           42 FFAESVRQLKAKDGRACI-----TNEVRKEIRNAFKNLPVE   77 (297)
Q Consensus        42 ~~~~~~~~~~~~~~~~~~-----~~~v~k~~g~~wk~ls~~   77 (297)
                      -...|++..+.-||+...     ....-..+.+.|..|++.
T Consensus        18 ik~ay~~l~~~~HPD~~~~~~~~~~~~~~~l~~Ay~~L~~~   58 (60)
T smart00271       18 IKKAYRKLALKYHPDKNPGDKEEAEEKFKEINEAYEVLSDP   58 (60)
T ss_pred             HHHHHHHHHHHHCcCCCCCchHHHHHHHHHHHHHHHHHcCC
Confidence            356788888899998876     334555667777777665


No 46 
>PF06628 Catalase-rel:  Catalase-related immune-responsive;  InterPro: IPR010582 Catalases (1.11.1.6 from EC) are antioxidant enzymes that catalyse the conversion of hydrogen peroxide to water and molecular oxygen, serving to protect cells from its toxic effects []. Hydrogen peroxide is produced as a consequence of oxidative cellular metabolism and can be converted to the highly reactive hydroxyl radical via transition metals, this radical being able to damage a wide variety of molecules within a cell, leading to oxidative stress and cell death. Catalases act to neutralise hydrogen peroxide toxicity, and are produced by all aerobic organisms ranging from bacteria to man. Most catalases are mono-functional, haem-containing enzymes, although there are also bifunctional haem-containing peroxidase/catalases (IPR000763 from INTERPRO) that are closely related to plant peroxidases, and non-haem, manganese-containing catalases (IPR007760 from INTERPRO) that are found in bacteria []. This entry represents a small conserved region within catalase enzymes that carries the immune-responsive amphipathic octa-peptide that is recognised by T cells [].; PDB: 2CAH_A 1NM0_A 1H7K_A 1E93_A 1H6N_A 3HB6_A 2CAG_A 1M85_A 1MQF_A 1A4E_C ....
Probab=21.78  E-value=50  Score=23.93  Aligned_cols=23  Identities=17%  Similarity=0.202  Sum_probs=18.1

Q ss_pred             HHHhhhcCCChHhhhhhhhhhhc
Q 039418           66 EIRNAFKNLPVEEKCQYKFQSRR   88 (297)
Q Consensus        66 ~~g~~wk~ls~~ek~~~~~~~~~   88 (297)
                      -.|..|++||++||+.....-..
T Consensus        12 Qa~~ly~~l~~~er~~lv~nia~   34 (68)
T PF06628_consen   12 QARDLYRVLSDEERERLVENIAG   34 (68)
T ss_dssp             HHHHHHHHSSHHHHHHHHHHHHH
T ss_pred             hHHHHHHHCCHHHHHHHHHHHHH
Confidence            46789999999999977765444


No 47 
>PF00226 DnaJ:  DnaJ domain;  InterPro: IPR001623 The prokaryotic heat shock protein DnaJ interacts with the chaperone hsp70-like DnaK protein []. Structurally, the DnaJ protein consists of an N-terminal conserved domain (called 'J' domain) of about 70 amino acids, a glycine-rich region ('G' domain') of about 30 residues, a central domain containing four repeats of a CXXCXGXG motif ('CRR' domain) and a C-terminal region of 120 to 170 residues. Such a structure is shown in the following schematic representation:  +------------+-+-------+-----+-----------+--------------------------------+ | N-terminal | | Gly-R | | CXXCXGXG | C-terminal | +------------+-+-------+-----+-----------+--------------------------------+   It is thought that the 'J' domain of DnaJ mediates the interaction with the dnaK protein and consists of four helices, the second of which has a charged surface that includes at least one pair of basic residues that are essential for interaction with the ATPase domain of Hsp70. The J- and CRR-domains are found in many prokaryotic and eukaryotic proteins [], either together or separately. In yeast, J-domains have been classified into 3 groups; the class III proteins are functionally distinct and do not appear to act as molecular chaperones []. ; GO: 0031072 heat shock protein binding; PDB: 2GUZ_C 2L6L_A 1HDJ_A 2EJ7_A 1FPO_C 2CUG_A 2QSA_A 2OCH_A 3BVO_B 3APQ_A ....
Probab=20.05  E-value=1.1e+02  Score=21.12  Aligned_cols=39  Identities=18%  Similarity=0.052  Sum_probs=30.0

Q ss_pred             hHHHHHHHHHHhCCCCcchh-----HHHHHHHhhhcCCChHhhh
Q 039418           42 FFAESVRQLKAKDGRACITN-----EVRKEIRNAFKNLPVEEKC   80 (297)
Q Consensus        42 ~~~~~~~~~~~~~~~~~~~~-----~v~k~~g~~wk~ls~~ek~   80 (297)
                      -..-|++..+.-||+.....     ..-..+.+.|+-|+++++.
T Consensus        17 ik~~y~~l~~~~HPD~~~~~~~~~~~~~~~i~~Ay~~L~~~~~R   60 (64)
T PF00226_consen   17 IKKAYRRLSKQYHPDKNSGDEAEAEEKFARINEAYEILSDPERR   60 (64)
T ss_dssp             HHHHHHHHHHHTSTTTGTSTHHHHHHHHHHHHHHHHHHHSHHHH
T ss_pred             HHHHHHhhhhccccccchhhhhhhhHHHHHHHHHHHHhCCHHHH
Confidence            35668888899999885544     4777899999999888754


Done!