Query         031973
Match_columns 150
No_of_seqs    112 out of 1110
Neff          7.0 
Searched_HMMs 46136
Date          Fri Mar 29 07:53:51 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/031973.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/031973hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PTZ00199 high mobility group p  99.9 2.2E-23 4.7E-28  144.8  11.3   84   25-109     8-93  (94)
  2 cd01389 MATA_HMG-box MATA_HMG-  99.8 5.2E-21 1.1E-25  127.9   7.9   73   39-112     1-73  (77)
  3 cd01388 SOX-TCF_HMG-box SOX-TC  99.8 1.7E-20 3.7E-25  123.9   8.0   71   39-110     1-71  (72)
  4 PF00505 HMG_box:  HMG (high mo  99.8 8.6E-20 1.9E-24  118.5   8.9   69   40-109     1-69  (69)
  5 cd01390 HMGB-UBF_HMG-box HMGB-  99.8 3.9E-19 8.5E-24  114.4   8.8   65   40-105     1-65  (66)
  6 smart00398 HMG high mobility g  99.8 4.7E-19   1E-23  114.7   9.0   70   39-109     1-70  (70)
  7 PF09011 HMG_box_2:  HMG-box do  99.8 1.4E-18 3.1E-23  115.0   9.0   72   37-109     1-73  (73)
  8 COG5648 NHP6B Chromatin-associ  99.8 2.6E-18 5.6E-23  133.3   7.7   93   25-118    56-148 (211)
  9 cd00084 HMG-box High Mobility   99.7 2.3E-17   5E-22  105.6   8.8   65   40-105     1-65  (66)
 10 KOG0527 HMG-box transcription   99.7 7.7E-18 1.7E-22  139.8   6.8   85   32-117    55-139 (331)
 11 KOG0381 HMG box-containing pro  99.7 1.6E-16 3.5E-21  109.7  10.9   76   36-112    17-95  (96)
 12 KOG0526 Nucleosome-binding fac  99.6 2.4E-15 5.2E-20  129.7   7.7   81   25-110   521-601 (615)
 13 KOG3248 Transcription factor T  99.5 9.2E-14   2E-18  114.5   5.8   79   38-117   190-268 (421)
 14 KOG0528 HMG-box transcription   99.2 7.6E-12 1.6E-16  107.1   3.7   85   29-114   315-399 (511)
 15 KOG4715 SWI/SNF-related matrix  99.2 7.6E-11 1.6E-15   96.7   8.9   80   32-112    57-136 (410)
 16 KOG2746 HMG-box transcription   98.7 1.9E-08 4.1E-13   89.3   4.7   74   30-104   172-247 (683)
 17 PF14887 HMG_box_5:  HMG (high   98.0 3.2E-05 6.9E-10   51.7   7.3   75   39-115     3-77  (85)
 18 PF06382 DUF1074:  Protein of u  97.4 0.00083 1.8E-08   51.5   7.7   50   44-98     83-132 (183)
 19 PF04690 YABBY:  YABBY protein;  97.2 0.00095   2E-08   51.0   5.7   49   34-83    116-164 (170)
 20 COG5648 NHP6B Chromatin-associ  96.9 0.00069 1.5E-08   53.2   2.9   68   38-106   142-209 (211)
 21 PF08073 CHDNT:  CHDNT (NUC034)  95.7   0.014   3E-07   36.6   3.1   39   45-84     14-52  (55)
 22 PF06244 DUF1014:  Protein of u  94.1   0.084 1.8E-06   38.3   3.9   49   36-85     68-117 (122)
 23 PF04769 MAT_Alpha1:  Mating-ty  93.5     0.2 4.4E-06   39.3   5.5   56   34-96     38-93  (201)
 24 KOG3223 Uncharacterized conser  89.7    0.23 4.9E-06   38.8   2.0   55   36-94    160-215 (221)
 25 TIGR03481 HpnM hopanoid biosyn  84.5     2.4 5.1E-05   33.0   5.0   47   66-112    64-112 (198)
 26 PRK15117 ABC transporter perip  79.2     6.1 0.00013   31.0   5.6   49   63-112    66-116 (211)
 27 PF05494 Tol_Tol_Ttg2:  Toluene  75.9       6 0.00013   29.5   4.6   45   66-110    38-84  (170)
 28 PF13875 DUF4202:  Domain of un  60.3      14  0.0003   28.7   3.8   39   46-88    131-169 (185)
 29 PF11304 DUF3106:  Protein of u  52.6      58  0.0013   22.8   5.6   22   73-94     14-35  (107)
 30 COG2854 Ttg2D ABC-type transpo  48.4      22 0.00048   28.0   3.2   43   73-115    78-121 (202)
 31 PF01352 KRAB:  KRAB box;  Inte  37.8      23 0.00049   20.6   1.4   26   69-94      4-30  (41)
 32 PRK09706 transcriptional repre  37.0      84  0.0018   22.3   4.6   43   71-113    88-130 (135)
 33 PF15076 DUF4543:  Domain of un  33.8      20 0.00044   23.3   0.8   22   33-54     25-46  (75)
 34 PF12881 NUT_N:  NUT protein N   32.9 1.3E+02  0.0028   25.5   5.5   54   58-112   243-297 (328)
 35 PRK10363 cpxP periplasmic repr  32.3 1.2E+02  0.0025   23.3   4.8   36   68-103   110-145 (166)
 36 PRK12750 cpxP periplasmic repr  30.9 1.3E+02  0.0028   22.8   4.9   33   73-105   128-160 (170)
 37 PRK12751 cpxP periplasmic stre  30.1 1.2E+02  0.0026   23.0   4.6   31   71-101   119-149 (162)
 38 PF00887 ACBP:  Acyl CoA bindin  24.0 2.2E+02  0.0047   18.7   4.8   54   45-100    28-85  (87)
 39 PF02209 VHP:  Villin headpiece  23.5      45 0.00098   18.9   1.0   10  140-149     3-12  (36)
 40 TIGR00787 dctP tripartite ATP-  22.7 1.7E+02  0.0038   22.9   4.6   28   76-103   213-240 (257)
 41 PRK10236 hypothetical protein;  22.4      80  0.0017   25.5   2.5   25   70-94    117-141 (237)
 42 PHA02819 hypothetical protein;  22.3      47   0.001   21.8   1.0   11  138-148    14-24  (71)
 43 PF06945 DUF1289:  Protein of u  21.9 1.2E+02  0.0027   18.1   2.8   23   67-94     23-45  (51)
 44 PHA02975 hypothetical protein;  20.8      53  0.0011   21.4   1.0   11  138-148    14-24  (69)
 45 KOG1610 Corticosteroid 11-beta  20.7 2.3E+02  0.0051   23.9   5.0   49   50-98    188-248 (322)
 46 PHA02844 putative transmembran  20.5      54  0.0012   21.8   1.0   11  138-148    14-24  (75)

No 1  
>PTZ00199 high mobility group protein; Provisional
Probab=99.90  E-value=2.2e-23  Score=144.82  Aligned_cols=84  Identities=38%  Similarity=0.662  Sum_probs=77.1

Q ss_pred             CccccCCccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCC--HHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHH
Q 031973           25 GKRTAKPKAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKS--VATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYN  102 (150)
Q Consensus        25 ~k~kk~kk~~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~--~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~  102 (150)
                      +.++++++..+||+.|+||+|||||||+++|..|..+||+ ++  +.+|+++||++|+.|++++|.+|+++|..++.+|.
T Consensus         8 ~~~k~~~k~~kdp~~PKrP~sAY~~F~~~~R~~i~~~~P~-~~~~~~evsk~ige~Wk~ls~eeK~~y~~~A~~dk~rY~   86 (94)
T PTZ00199          8 VLVRKNKRKKKDPNAPKRALSAYMFFAKEKRAEIIAENPE-LAKDVAAVGKMVGEAWNKLSEEEKAPYEKKAQEDKVRYE   86 (94)
T ss_pred             ccccccCCCCCCCCCCCCCCcHHHHHHHHHHHHHHHHCcC-CcccHHHHHHHHHHHHHcCCHHHHHHHHHHHHHHHHHHH
Confidence            3444455668999999999999999999999999999999 75  89999999999999999999999999999999999


Q ss_pred             HHHHHHH
Q 031973          103 KNMQDYN  109 (150)
Q Consensus       103 ~~~~~y~  109 (150)
                      .+|..|.
T Consensus        87 ~e~~~Y~   93 (94)
T PTZ00199         87 KEKAEYA   93 (94)
T ss_pred             HHHHHHh
Confidence            9999995


No 2  
>cd01389 MATA_HMG-box MATA_HMG-box, class I member of the HMG-box superfamily of DNA-binding proteins. These proteins contain a single HMG box, and bind the minor groove of DNA in a highly sequence-specific manner. Members include the fungal mating type gene products MC, MATA1 and Ste11.
Probab=99.84  E-value=5.2e-21  Score=127.85  Aligned_cols=73  Identities=27%  Similarity=0.459  Sum_probs=70.5

Q ss_pred             CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHhhc
Q 031973           39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQL  112 (150)
Q Consensus        39 ~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~~~  112 (150)
                      +|+||+||||||+++.|..|+.+||+ +++.+|+++||.+|+.|++++|++|.++|..++++|..++++|+...
T Consensus         1 ~~kRP~naf~lf~~~~r~~~~~~~p~-~~~~eisk~~g~~Wk~ls~eeK~~y~~~A~~~k~~~~~~~p~Yky~p   73 (77)
T cd01389           1 KIPRPRNAFILYRQDKHAQLKTENPG-LTNNEISRIIGRMWRSESPEVKAYYKELAEEEKERHAREYPDYKYTP   73 (77)
T ss_pred             CCCCCCcHHHHHHHHHHHHHHHHCCC-CCHHHHHHHHHHHHhhCCHHHHHHHHHHHHHHHHHHHHHCCCCcccC
Confidence            58999999999999999999999999 99999999999999999999999999999999999999999998754


No 3  
>cd01388 SOX-TCF_HMG-box SOX-TCF_HMG-box, class I member of the HMG-box superfamily of DNA-binding proteins. These proteins contain a single HMG box, and bind the minor groove of DNA in a highly sequence-specific manner. Members include SRY and its homologs in insects and vertebrates, and transcription factor-like proteins, TCF-1, -3, -4, and LEF-1. They appear to bind the minor groove of the A/T C A A A G/C-motif.
Probab=99.83  E-value=1.7e-20  Score=123.91  Aligned_cols=71  Identities=34%  Similarity=0.561  Sum_probs=68.6

Q ss_pred             CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHh
Q 031973           39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNK  110 (150)
Q Consensus        39 ~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~  110 (150)
                      ++|||+|||||||+++|..++.+||+ +++.+|+++||.+|+.|++++|++|.++|..++++|..++++|+.
T Consensus         1 ~iKrP~naf~~F~~~~r~~~~~~~p~-~~~~eisk~l~~~Wk~ls~~eK~~y~~~a~~~k~~y~~~~p~y~y   71 (72)
T cd01388           1 HIKRPMNAFMLFSKRHRRKVLQEYPL-KENRAISKILGDRWKALSNEEKQPYYEEAKKLKELHMKLYPDYKW   71 (72)
T ss_pred             CCCCCCcHHHHHHHHHHHHHHHHCCC-CCHHHHHHHHHHHHHcCCHHHHHHHHHHHHHHHHHHHHHCcCCCC
Confidence            47899999999999999999999999 999999999999999999999999999999999999999999863


No 4  
>PF00505 HMG_box:  HMG (high mobility group) box;  InterPro: IPR000910 High mobility group (HMG or HMGB) proteins are a family of relatively low molecular weight non-histone components in chromatin. HMG1 (also called HMG-T in fish) and HMG2 are two highly related proteins that bind single-stranded DNA preferentially and unwind double-stranded DNA. Although they have no sequence specificity, they have a high affinity for bent or distorted DNA, and bend linear DNA. HMG1 and HMG2 contain two DNA-binding HMG-box domains (A and B) that show structural and functional differences, and have a long acidic C-terminal domain rich in aspartic and glutamic acid residues. The acidic tail modulates the affinity of the tandem HMG boxes in HMG1 and 2 for a variety of DNA targets. HMG1 and 2 appear to play important architectural roles in the assembly of nucleoprotein complexes in a variety of biological processes, for example V(D)J recombination, the initiation of transcription, and DNA repair []. The profile in this entry describing the HMG-domains is much more general than the signature. In addition to the HMG1 and HMG2 proteins, HMG-domains occur in single or multiple copies in the following protein classes; the SOX family of transcription factors; SRY sex determining region Y protein and related proteins []; LEF1 lymphoid enhancer binding factor 1 []; SSRP recombination signal recognition protein; MTF1 mitochondrial transcription factor 1; UBF1/2 nucleolar transcription factors; Abf2 yeast ARS-binding factor []; and Saccharomyces cerevisiae transcription factors Ixr1, Rox1, Nhp6a, Nhp6b and Spp41.; GO: 0003677 DNA binding; PDB: 1I11_A 1J3C_A 1J3D_A 1WZ6_A 1WGF_A 2D7L_A 1GT0_D 3U2B_C 2CRJ_A 2CS1_A ....
Probab=99.82  E-value=8.6e-20  Score=118.54  Aligned_cols=69  Identities=45%  Similarity=0.836  Sum_probs=65.8

Q ss_pred             CCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHH
Q 031973           40 PKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN  109 (150)
Q Consensus        40 PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~  109 (150)
                      |+||+|||+|||.+++..++.+||+ ++..+|+++||.+|+.|++++|.+|.+.|..++..|..+|+.|+
T Consensus         1 PkrP~~af~lf~~~~~~~~k~~~p~-~~~~~i~~~~~~~W~~l~~~eK~~y~~~a~~~~~~y~~~~~~y~   69 (69)
T PF00505_consen    1 PKRPPNAFMLFCKEKRAKLKEENPD-LSNKEISKILAQMWKNLSEEEKAPYKEEAEEEKERYEKEMPEYK   69 (69)
T ss_dssp             SSSS--HHHHHHHHHHHHHHHHSTT-STHHHHHHHHHHHHHCSHHHHHHHHHHHHHHHHHHHHHHHHHHH
T ss_pred             CcCCCCHHHHHHHHHHHHHHHHhcc-cccccchhhHHHHHhcCCHHHHHHHHHHHHHHHHHHHHHHHhcC
Confidence            8999999999999999999999999 99999999999999999999999999999999999999999995


No 5  
>cd01390 HMGB-UBF_HMG-box HMGB-UBF_HMG-box, class II and III members of the HMG-box superfamily of DNA-binding proteins. These proteins bind the minor groove of DNA in a non-sequence specific fashion and contain two or more tandem HMG boxes. Class II members include non-histone chromosomal proteins, HMG1 and HMG2, which bind to bent or distorted DNA such as four-way DNA junctions, synthetic DNA cruciforms, kinked cisplatin-modified DNA, DNA bulges, cross-overs in supercoiled DNA, and can cause looping of linear DNA. Class III members include nucleolar and mitochondrial transcription factors, UBF and mtTF1, which bind four-way DNA junctions.
Probab=99.80  E-value=3.9e-19  Score=114.36  Aligned_cols=65  Identities=51%  Similarity=0.854  Sum_probs=63.6

Q ss_pred             CCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHH
Q 031973           40 PKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNM  105 (150)
Q Consensus        40 PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~  105 (150)
                      |++|+|||+||++++|..++..||+ +++.+|++.||.+|+.|++++|.+|.+.|..++.+|..+|
T Consensus         1 Pkrp~saf~~f~~~~r~~~~~~~p~-~~~~~i~~~~~~~W~~ls~~eK~~y~~~a~~~~~~y~~e~   65 (66)
T cd01390           1 PKRPLSAYFLFSQEQRPKLKKENPD-ASVTEVTKILGEKWKELSEEEKKKYEEKAEKDKERYEKEM   65 (66)
T ss_pred             CCCCCcHHHHHHHHHHHHHHHHCcC-CCHHHHHHHHHHHHHhCCHHHHHHHHHHHHHHHHHHHHhh
Confidence            8999999999999999999999999 9999999999999999999999999999999999999876


No 6  
>smart00398 HMG high mobility group.
Probab=99.80  E-value=4.7e-19  Score=114.71  Aligned_cols=70  Identities=47%  Similarity=0.831  Sum_probs=67.9

Q ss_pred             CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHH
Q 031973           39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN  109 (150)
Q Consensus        39 ~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~  109 (150)
                      +|++|+|||+||++++|..+..+||+ +++.+|++.||.+|+.|++++|.+|.++|..++.+|...++.|.
T Consensus         1 ~pkrp~~~y~~f~~~~r~~~~~~~~~-~~~~~i~~~~~~~W~~l~~~ek~~y~~~a~~~~~~y~~~~~~y~   70 (70)
T smart00398        1 KPKRPMSAFMLFSQENRAKIKAENPD-LSNAEISKKLGERWKLLSEEEKAPYEEKAKKDKERYEEEMPEYK   70 (70)
T ss_pred             CcCCCCcHHHHHHHHHHHHHHHHCcC-CCHHHHHHHHHHHHHcCCHHHHHHHHHHHHHHHHHHHHHHHhcC
Confidence            58999999999999999999999999 99999999999999999999999999999999999999999884


No 7  
>PF09011 HMG_box_2:  HMG-box domain;  InterPro: IPR015101 This domain is predominantly found in Maelstrom homologue proteins. It has no known function. ; GO: 0005634 nucleus; PDB: 2EQZ_A 1V64_A 2CTO_A 1H5P_A 3TQ6_A 3FGH_A 3TMM_A 1J3X_A 2YRQ_A 1AAB_A ....
Probab=99.78  E-value=1.4e-18  Score=114.97  Aligned_cols=72  Identities=49%  Similarity=0.876  Sum_probs=63.8

Q ss_pred             CCCCCCCCChHHHHHHHHHHHHHHh-CCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHH
Q 031973           37 PNKPKRPPSAFFVFMEEFRKQFKEA-HPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN  109 (150)
Q Consensus        37 p~~PKrP~~aY~lF~~~~r~~~k~~-~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~  109 (150)
                      |++||+|+|||+||+.+++..++.. ++. ....++++.|+..|+.||+++|.+|.++|..++.+|..+|..|.
T Consensus         1 p~kpK~~~say~lF~~~~~~~~k~~G~~~-~~~~e~~k~~~~~Wk~Ls~~EK~~Y~~~A~~~k~~y~~e~~~~~   73 (73)
T PF09011_consen    1 PKKPKRPPSAYNLFMKEMRKEVKEEGGQK-QSFREVMKEISERWKSLSEEEKEPYEERAKEDKERYEREMKEWN   73 (73)
T ss_dssp             SSS--SSSSHHHHHHHHHHHHHHHHT-T--SSHHHHHHHHHHHHHHS-HHHHHHHHHHHHHHHHHHHHHHHHH-
T ss_pred             CcCCCCCCCHHHHHHHHHHHHHHHhcccC-CCHHHHHHHHHHHHHhcCHHHHHHHHHHHHHHHHHHHHHHHhcC
Confidence            6899999999999999999999998 665 88999999999999999999999999999999999999999984


No 8  
>COG5648 NHP6B Chromatin-associated proteins containing the HMG domain [Chromatin structure and dynamics]
Probab=99.75  E-value=2.6e-18  Score=133.26  Aligned_cols=93  Identities=33%  Similarity=0.676  Sum_probs=86.2

Q ss_pred             CccccCCccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHH
Q 031973           25 GKRTAKPKAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKN  104 (150)
Q Consensus        25 ~k~kk~kk~~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~  104 (150)
                      ++.+..++..+||+.|+||++|||+|+.++|++|+..+|. +++.+|+++||++|++|++++|++|...|..++++|...
T Consensus        56 ~ksk~~~r~k~dpN~PKRp~sayf~y~~~~R~ei~~~~p~-l~~~e~~k~~~e~WK~Ltd~eke~y~k~~~~~~erYq~e  134 (211)
T COG5648          56 TKSKRLVRKKKDPNGPKRPLSAYFLYSAENRDEIRKENPK-LTFGEVGKLLSEKWKELTDEEKEPYYKEANSDRERYQRE  134 (211)
T ss_pred             hHHHHHHHHhcCCCCCCCchhHHHHHHHHHHHHHHHhCCC-CChHHHHHHHHHHHHhccHhhhhhHHHHHhhHHHHHHHH
Confidence            4445667889999999999999999999999999999999 999999999999999999999999999999999999999


Q ss_pred             HHHHHhhccCCcch
Q 031973          105 MQDYNKQLADGVNA  118 (150)
Q Consensus       105 ~~~y~~~~~~~~~~  118 (150)
                      +..|....+.....
T Consensus       135 k~~y~~k~~~~~~~  148 (211)
T COG5648         135 KEEYNKKLPNKAPI  148 (211)
T ss_pred             HHhhhcccCCCCCC
Confidence            99999988775544


No 9  
>cd00084 HMG-box High Mobility Group (HMG)-box is found in a variety of eukaryotic chromosomal proteins and transcription factors. HMGs bind to the minor groove of DNA and have been classified by DNA binding preferences. Two phylogenically distinct groups of Class I proteins bind DNA in a sequence specific fashion and contain a single HMG box. One group (SOX-TCF) includes transcription factors, TCF-1, -3, -4; and also SRY and LEF-1, which bind four-way DNA junctions and duplex DNA targets. The second group (MATA) includes fungal mating type gene products MC, MATA1 and Ste11. Class II and III proteins (HMGB-UBF) bind DNA in a non-sequence specific fashion and contain two or more tandem HMG boxes. Class II members include non-histone chromosomal proteins, HMG1 and HMG2, which bind to bent or distorted DNA such as four-way DNA junctions, synthetic DNA cruciforms, kinked cisplatin-modified DNA, DNA bulges, cross-overs in supercoiled DNA, and can cause looping of linear DNA. Class III member
Probab=99.73  E-value=2.3e-17  Score=105.57  Aligned_cols=65  Identities=49%  Similarity=0.827  Sum_probs=63.1

Q ss_pred             CCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHH
Q 031973           40 PKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNM  105 (150)
Q Consensus        40 PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~  105 (150)
                      |++|+|||+||+++.+..++..||+ ++..+|++.||.+|+.|++++|.+|.+.|..++..|..++
T Consensus         1 pkrp~~af~~f~~~~~~~~~~~~~~-~~~~~i~~~~~~~W~~l~~~~k~~y~~~a~~~~~~y~~~~   65 (66)
T cd00084           1 PKRPLSAYFLFSQEHRAEVKAENPG-LSVGEISKILGEMWKSLSEEEKKKYEEKAEKDKERYEKEM   65 (66)
T ss_pred             CCCCCcHHHHHHHHHHHHHHHHCcC-CCHHHHHHHHHHHHHhCCHHHHHHHHHHHHHHHHHHHHhh
Confidence            7999999999999999999999999 9999999999999999999999999999999999998875


No 10 
>KOG0527 consensus HMG-box transcription factor [Transcription]
Probab=99.72  E-value=7.7e-18  Score=139.76  Aligned_cols=85  Identities=29%  Similarity=0.539  Sum_probs=79.4

Q ss_pred             ccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHhh
Q 031973           32 KAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQ  111 (150)
Q Consensus        32 k~~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~~  111 (150)
                      .......++||||||||||.+..|.+|..+||. +.+.||+++||.+|+.|++++|.+|.+.|++++..|++++++|+.+
T Consensus        55 ~~k~~~~hIKRPMNAFMVWSq~~RRkma~qnP~-mHNSEISK~LG~~WK~Lse~EKrPFi~EAeRLR~~HmkehPdYKYR  133 (331)
T KOG0527|consen   55 KDKTSTDRIKRPMNAFMVWSQGQRRKLAKQNPK-MHNSEISKRLGAEWKLLSEEEKRPFVDEAERLRAQHMKEYPDYKYR  133 (331)
T ss_pred             cCCCCccccCCCcchhhhhhHHHHHHHHHhCcc-hhhHHHHHHHHHHHhhcCHhhhccHHHHHHHHHHHHHHhCCCcccc
Confidence            345667899999999999999999999999999 9999999999999999999999999999999999999999999998


Q ss_pred             ccCCcc
Q 031973          112 LADGVN  117 (150)
Q Consensus       112 ~~~~~~  117 (150)
                      ......
T Consensus       134 PRRKkk  139 (331)
T KOG0527|consen  134 PRRKKK  139 (331)
T ss_pred             cccccc
Confidence            776554


No 11 
>KOG0381 consensus HMG box-containing protein [General function prediction only]
Probab=99.71  E-value=1.6e-16  Score=109.69  Aligned_cols=76  Identities=47%  Similarity=0.840  Sum_probs=72.0

Q ss_pred             CC--CCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHH-HHHhhc
Q 031973           36 DP--NKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQ-DYNKQL  112 (150)
Q Consensus        36 dp--~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~-~y~~~~  112 (150)
                      +|  +.|++|++||++|+.+.+..++.+||+ ++..+|+++||++|++|++++|.+|...|..++.+|...|. .|+...
T Consensus        17 ~p~~~~pkrp~sa~~~f~~~~~~~~k~~~p~-~~~~~v~k~~g~~W~~l~~~~k~~y~~ka~~~k~~Y~~~~~~~~~~~~   95 (96)
T KOG0381|consen   17 DPNAQAPKRPLSAFFLFSSEQRSKIKAENPG-LSVGEVAKALGEMWKNLAEEEKQPYEEKASKLKEKYEKELAGEYKASL   95 (96)
T ss_pred             CCCCCCCCCCCcHHHHHHHHHHHHHHHhCCC-CCHHHHHHHHHHHHhcCCHHHHHHHHHHHHHHHHHHHHHHHHHHhhcc
Confidence            55  599999999999999999999999999 99999999999999999999999999999999999999999 887653


No 12 
>KOG0526 consensus Nucleosome-binding factor SPN, POB3 subunit [Transcription; Replication, recombination and repair; Chromatin structure and dynamics]
Probab=99.59  E-value=2.4e-15  Score=129.68  Aligned_cols=81  Identities=40%  Similarity=0.678  Sum_probs=74.8

Q ss_pred             CccccCCccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHH
Q 031973           25 GKRTAKPKAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKN  104 (150)
Q Consensus        25 ~k~kk~kk~~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~  104 (150)
                      .+++++.++.+||+.||||++|||||.+..|..|+..  + .+.++|++.+|.+|+.|+.  |.+|.+.|+.++++|+.+
T Consensus       521 ~~~~k~~kk~kdpnapkra~sa~m~w~~~~r~~ik~d--g-i~~~dv~kk~g~~wk~ms~--k~~we~ka~~dk~ry~~e  595 (615)
T KOG0526|consen  521 KEKKKKGKKKKDPNAPKRATSAYMLWLNASRESIKED--G-ISVGDVAKKAGEKWKQMSA--KEEWEDKAAVDKQRYEDE  595 (615)
T ss_pred             hccccCcccCCCCCCCccchhHHHHHHHhhhhhHhhc--C-chHHHHHHHHhHHHhhhcc--cchhhHHHHHHHHHHHHH
Confidence            3444677789999999999999999999999999987  5 8999999999999999998  899999999999999999


Q ss_pred             HHHHHh
Q 031973          105 MQDYNK  110 (150)
Q Consensus       105 ~~~y~~  110 (150)
                      |.+|+.
T Consensus       596 m~~yk~  601 (615)
T KOG0526|consen  596 MKEYKN  601 (615)
T ss_pred             HHhhcC
Confidence            999993


No 13 
>KOG3248 consensus Transcription factor TCF-4 [Transcription]
Probab=99.45  E-value=9.2e-14  Score=114.48  Aligned_cols=79  Identities=23%  Similarity=0.406  Sum_probs=74.6

Q ss_pred             CCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHhhccCCcc
Q 031973           38 NKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQLADGVN  117 (150)
Q Consensus        38 ~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~~~~~~~~  117 (150)
                      .+.|+|+|||||||++.|..|..++.- ....+|.++||++|..|+.++.++|.++|.++++.|.+.++.|.+.......
T Consensus       190 phiKKPLNAFmlyMKEmRa~vvaEctl-KeSAaiNqiLGrRWH~LSrEEQAKYyElArKerqlH~qlYP~WSARdNYgKK  268 (421)
T KOG3248|consen  190 PHIKKPLNAFMLYMKEMRAKVVAECTL-KESAAINQILGRRWHALSREEQAKYYELARKERQLHMQLYPGWSARDNYGKK  268 (421)
T ss_pred             ccccccHHHHHHHHHHHHHHHHHHhhh-hhHHHHHHHHhHHHhhhhHHHHHHHHHHHHHHHHHHHHhcCCcchhhhhhhh
Confidence            488999999999999999999999986 6889999999999999999999999999999999999999999999988754


No 14 
>KOG0528 consensus HMG-box transcription factor SOX5 [Transcription]
Probab=99.21  E-value=7.6e-12  Score=107.10  Aligned_cols=85  Identities=26%  Similarity=0.453  Sum_probs=77.0

Q ss_pred             cCCccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHH
Q 031973           29 AKPKAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDY  108 (150)
Q Consensus        29 k~kk~~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y  108 (150)
                      .++-..+.+.++||||||||+|.++-|..|.+.+|+ +.+..|+++||.+|+.|+..+|++|.+.-..+...|.+.+++|
T Consensus       315 srrg~~ss~PHIKRPMNAFMVWAkDERRKILqA~PD-MHNSnISKILGSRWKaMSN~eKQPYYEEQaRLSk~HlEk~PdY  393 (511)
T KOG0528|consen  315 SRRGRASSEPHIKRPMNAFMVWAKDERRKILQAFPD-MHNSNISKILGSRWKAMSNTEKQPYYEEQARLSKLHLEKYPDY  393 (511)
T ss_pred             cccCcCCCCccccCCcchhhcccchhhhhhhhcCcc-ccccchhHHhcccccccccccccchHHHHHHHHHhhhccCccc
Confidence            335556677899999999999999999999999999 9999999999999999999999999998888888999999999


Q ss_pred             HhhccC
Q 031973          109 NKQLAD  114 (150)
Q Consensus       109 ~~~~~~  114 (150)
                      +.+...
T Consensus       394 rYkPRP  399 (511)
T KOG0528|consen  394 RYKPRP  399 (511)
T ss_pred             ccCCCC
Confidence            987643


No 15 
>KOG4715 consensus SWI/SNF-related matrix-associated actin-dependent regulator of chromatin  [Chromatin structure and dynamics]
Probab=99.20  E-value=7.6e-11  Score=96.73  Aligned_cols=80  Identities=24%  Similarity=0.548  Sum_probs=74.6

Q ss_pred             ccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHhh
Q 031973           32 KAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQ  111 (150)
Q Consensus        32 k~~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~~  111 (150)
                      ...+.|.+|-+|+-+||.|++..|++|+..||. +...+|.++||.+|..|++++|+.|...+...+..|.+.|..|...
T Consensus        57 t~pkpPkppekpl~pymrySrkvWd~VkA~nPe-~kLWeiGK~Ig~mW~dLpd~EK~ey~~EYeaEKieY~~smkayh~s  135 (410)
T KOG4715|consen   57 TRPKPPKPPEKPLMPYMRYSRKVWDQVKASNPE-LKLWEIGKIIGGMWLDLPDEEKQEYLNEYEAEKIEYNESMKAYHNS  135 (410)
T ss_pred             cCCCCCCCCCcccchhhHHhhhhhhhhhccCcc-hHHHHHHHHHHHHHhhCcchHHHHHHHHHHHHHHHHHHHHHHhhCC
Confidence            345567888999999999999999999999999 9999999999999999999999999999999999999999998875


Q ss_pred             c
Q 031973          112 L  112 (150)
Q Consensus       112 ~  112 (150)
                      .
T Consensus       136 p  136 (410)
T KOG4715|consen  136 P  136 (410)
T ss_pred             c
Confidence            4


No 16 
>KOG2746 consensus HMG-box transcription factor Capicua and related proteins [Transcription]
Probab=98.68  E-value=1.9e-08  Score=89.28  Aligned_cols=74  Identities=27%  Similarity=0.478  Sum_probs=69.0

Q ss_pred             CCccCCCCCCCCCCCChHHHHHHHHH--HHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHH
Q 031973           30 KPKAAKDPNKPKRPPSAFFVFMEEFR--KQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKN  104 (150)
Q Consensus        30 ~kk~~~dp~~PKrP~~aY~lF~~~~r--~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~  104 (150)
                      +-...++..+.++|||+|+||++.+|  ..+.+.||+ ..+.-|+++||++|-.|.+.+|+.|+++|.+.++.|.++
T Consensus       172 rspnkr~k~HirrPMnaf~ifskrhr~~g~vhq~~pn-~DNrtIskiLgewWytL~~~Ekq~yhdLa~Qvk~Ahfka  247 (683)
T KOG2746|consen  172 RSPNKRDKDHIRRPMNAFHIFSKRHRGEGRVHQRHPN-QDNRTISKILGEWWYTLGPNEKQKYHDLAFQVKEAHFKA  247 (683)
T ss_pred             CCCCcCcchhhhhhhHHHHHHHhhcCCccchhccCcc-ccchhHHHHHhhhHhhhCchhhhhHHHHHHHHHHHHhhh
Confidence            33556778899999999999999999  899999999 999999999999999999999999999999999999986


No 17 
>PF14887 HMG_box_5:  HMG (high mobility group) box 5; PDB: 1L8Y_A 1L8Z_A 2HDZ_A.
Probab=98.04  E-value=3.2e-05  Score=51.67  Aligned_cols=75  Identities=19%  Similarity=0.387  Sum_probs=61.4

Q ss_pred             CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHhhccCC
Q 031973           39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQLADG  115 (150)
Q Consensus        39 ~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~~~~~~  115 (150)
                      .|..|-+|--||.+.....+...++. ....+ .+.+...|++|++.+|.+|...|.++..+|+..|.+|+...+..
T Consensus         3 lPE~PKt~qe~Wqq~vi~dYla~~~~-dr~K~-~kam~~~W~~me~Kekl~WIkKA~EdqKrYE~el~e~r~~~~~~   77 (85)
T PF14887_consen    3 LPETPKTAQEIWQQSVIGDYLAKFRN-DRKKA-LKAMEAQWSQMEKKEKLKWIKKAAEDQKRYERELREMRSAPADA   77 (85)
T ss_dssp             -S----THHHHHHHHHHHHHHHHTTS-THHHH-HHHHHHHHHTTGGGHHHHHHHHHHHHHHHHHHHHHCCS-CCCTT
T ss_pred             CCCCCCCHHHHHHHHHHHHHHHHhhH-hHHHH-HHHHHHHHHHhhhhhhhHHHHHHHHHHHHHHHHHHHHhcCCCCC
Confidence            57788999999999999999999987 54444 56899999999999999999999999999999999999877643


No 18 
>PF06382 DUF1074:  Protein of unknown function (DUF1074);  InterPro: IPR024460 This family consists of several proteins which appear to be specific to Insecta. The function of this family is unknown.
Probab=97.43  E-value=0.00083  Score=51.49  Aligned_cols=50  Identities=28%  Similarity=0.449  Sum_probs=43.2

Q ss_pred             CChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHH
Q 031973           44 PSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRK   98 (150)
Q Consensus        44 ~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k   98 (150)
                      -++|+-|+.+.+.    .|.+ +...++....+..|..|++.+|..|..++....
T Consensus        83 nnaYLNFLReFRr----kh~~-L~p~dlI~~AAraW~rLSe~eK~rYrr~~~~~~  132 (183)
T PF06382_consen   83 NNAYLNFLREFRR----KHCG-LSPQDLIQRAARAWCRLSEAEKNRYRRMAPSVR  132 (183)
T ss_pred             chHHHHHHHHHHH----HccC-CCHHHHHHHHHHHHHhCCHHHHHHHHhhcchhh
Confidence            3789999988865    6677 999999999999999999999999998766543


No 19 
>PF04690 YABBY:  YABBY protein;  InterPro: IPR006780 YABBY proteins are a group of plant-specific transcription factors involved in the specification of abaxial polarity in lateral organs such as leaves and floral organs [, ].
Probab=97.19  E-value=0.00095  Score=51.04  Aligned_cols=49  Identities=33%  Similarity=0.498  Sum_probs=43.2

Q ss_pred             CCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCC
Q 031973           34 AKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMS   83 (150)
Q Consensus        34 ~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~   83 (150)
                      .+.|.+-.|-++||..|+++.-..|+..+|+ +++.|+....+..|...|
T Consensus       116 ~kPPEKRqR~psaYn~f~k~ei~rik~~~p~-ishkeaFs~aAknW~h~p  164 (170)
T PF04690_consen  116 NKPPEKRQRVPSAYNRFMKEEIQRIKAENPD-ISHKEAFSAAAKNWAHFP  164 (170)
T ss_pred             cCCccccCCCchhHHHHHHHHHHHHHhcCCC-CCHHHHHHHHHHhhhhCc
Confidence            3445555677899999999999999999999 999999999999998876


No 20 
>COG5648 NHP6B Chromatin-associated proteins containing the HMG domain [Chromatin structure and dynamics]
Probab=96.94  E-value=0.00069  Score=53.18  Aligned_cols=68  Identities=19%  Similarity=0.393  Sum_probs=62.2

Q ss_pred             CCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHH
Q 031973           38 NKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQ  106 (150)
Q Consensus        38 ~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~  106 (150)
                      .+|..|..+|+-+...+|+.+...+|+ ....+++++++..|.+|++.-+.+|.+.+..++..|...|+
T Consensus       142 ~~~~~~~~~~~e~~~~~r~~~~~~~~~-~~~~e~~k~~~~~w~el~~skK~~~~~~~Kk~k~~~~~~~~  209 (211)
T COG5648         142 LPNKAPIGPFIENEPKIRPKVEGPSPD-KALVEETKIISKAWSELDESKKKKYIDKYKKLKEEYDSFYP  209 (211)
T ss_pred             cCCCCCCchhhhccHHhccccCCCCcc-hhhhHHhhhhhhhhhhhChhhhhHHHHHHHHHHHHHhhhcc
Confidence            467888889999999999999999998 88999999999999999999999999999999999887664


No 21 
>PF08073 CHDNT:  CHDNT (NUC034) domain;  InterPro: IPR012958 The CHD N-terminal domain is found in PHD/RING fingers and chromo domain-associated helicases [].; GO: 0003677 DNA binding, 0005524 ATP binding, 0008270 zinc ion binding, 0016818 hydrolase activity, acting on acid anhydrides, in phosphorus-containing anhydrides, 0006355 regulation of transcription, DNA-dependent, 0005634 nucleus
Probab=95.69  E-value=0.014  Score=36.62  Aligned_cols=39  Identities=18%  Similarity=0.430  Sum_probs=35.4

Q ss_pred             ChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCCh
Q 031973           45 SAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSE   84 (150)
Q Consensus        45 ~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~   84 (150)
                      +-|-+|.+-.|+.|...||+ +....|...++..|++.+.
T Consensus        14 t~yK~Fsq~vRP~l~~~NPk-~~~sKl~~l~~AKwrEF~~   52 (55)
T PF08073_consen   14 TNYKAFSQHVRPLLAKANPK-APMSKLMMLLQAKWREFQE   52 (55)
T ss_pred             HHHHHHHHHHHHHHHHHCCC-CcHHHHHHHHHHHHHHHHh
Confidence            56889999999999999999 9999999999999987653


No 22 
>PF06244 DUF1014:  Protein of unknown function (DUF1014);  InterPro: IPR010422 This family consists of several hypothetical eukaryotic proteins of unknown function.
Probab=94.05  E-value=0.084  Score=38.34  Aligned_cols=49  Identities=20%  Similarity=0.353  Sum_probs=42.9

Q ss_pred             CCCCCCCCC-ChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChh
Q 031973           36 DPNKPKRPP-SAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSED   85 (150)
Q Consensus        36 dp~~PKrP~-~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~e   85 (150)
                      ...||-|.+ -||.-|...+.+.|+.++|+ +..+++-.+|-.+|...|++
T Consensus        68 ~drHPErR~KAAy~afeE~~Lp~lK~E~Pg-LrlsQ~kq~l~K~w~KSPeN  117 (122)
T PF06244_consen   68 IDRHPERRMKAAYKAFEERRLPELKEENPG-LRLSQYKQMLWKEWQKSPEN  117 (122)
T ss_pred             CCCCcchhHHHHHHHHHHHHhHHHHhhCCC-chHHHHHHHHHHHHhcCCCC
Confidence            345675555 78999999999999999999 99999999999999988865


No 23 
>PF04769 MAT_Alpha1:  Mating-type protein MAT alpha 1;  InterPro: IPR006856 This family includes Saccharomyces cerevisiae (Baker's yeast) mating type protein alpha 1 (P01365 from SWISSPROT). MAT alpha 1 is a transcription activator that activates mating-type alpha-specific genes with the help of the MADS-box containing MCM1 transcription factor, which together bind cooperatively to PQ elements upstream of alpha-specific genes. The MCM1-MATalpha1 complex is required for the proper DNA-bending that is needed for transcriptional activation []. Alpha 1 interacts in vivo with STE12, linking expression of alpha-specific genes to the alpha-pheromone (IPR006742 from INTERPRO) response pathway [].; GO: 0000772 mating pheromone activity, 0003677 DNA binding, 0045895 positive regulation of transcription, mating-type specific, 0005634 nucleus
Probab=93.45  E-value=0.2  Score=39.28  Aligned_cols=56  Identities=20%  Similarity=0.341  Sum_probs=39.7

Q ss_pred             CCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHH
Q 031973           34 AKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEK   96 (150)
Q Consensus        34 ~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~   96 (150)
                      ......++||+|+||+|+.-.-    ...|+ ..-.+++..|+.+|..=+-  |..|.-.|+.
T Consensus        38 ~~~~~~~kr~lN~Fm~FRsyy~----~~~~~-~~Qk~~S~~l~~lW~~dp~--k~~W~l~ak~   93 (201)
T PF04769_consen   38 KRSPEKAKRPLNGFMAFRSYYS----PIFPP-LPQKELSGILTKLWEKDPF--KNKWSLMAKA   93 (201)
T ss_pred             cccccccccchhHHHHHHHHHH----hhcCC-cCHHHHHHHHHHHHhCCcc--HhHHHHHhhh
Confidence            3445678999999999986654    34454 5668999999999997543  4446555543


No 24 
>KOG3223 consensus Uncharacterized conserved protein [Function unknown]
Probab=89.70  E-value=0.23  Score=38.82  Aligned_cols=55  Identities=25%  Similarity=0.477  Sum_probs=45.6

Q ss_pred             CCCCC-CCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHH
Q 031973           36 DPNKP-KRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERA   94 (150)
Q Consensus        36 dp~~P-KrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A   94 (150)
                      |..|| +|=.-||.-|-....+.|+.+||+ +..+++-.+|-.+|...|++   ||.+.+
T Consensus       160 ddrHPEkRmrAA~~afEe~~LPrLK~e~P~-lrlsQ~Kqll~Kew~KsPDN---P~Nq~~  215 (221)
T KOG3223|consen  160 DDRHPEKRMRAAFKAFEEARLPRLKKENPG-LRLSQYKQLLKKEWQKSPDN---PFNQAA  215 (221)
T ss_pred             cccChHHHHHHHHHHHHHhhchhhhhcCCC-ccHHHHHHHHHHHHhhCCCC---hhhHHh
Confidence            44677 444567999999999999999999 99999999999999999976   555543


No 25 
>TIGR03481 HpnM hopanoid biosynthesis associated membrane protein HpnM. The genomes containing members of this family share the machinery for the biosynthesis of hopanoid lipids. Furthermore, the genes of this family are usually located proximal to other components of this biological process. The proteins are members of the pfam05494 family of putative transporters known as "toluene tolerance protein Ttg2D", although it is unlikely that the members included here have anything to do with toluene per-se.
Probab=84.45  E-value=2.4  Score=32.99  Aligned_cols=47  Identities=15%  Similarity=0.442  Sum_probs=39.6

Q ss_pred             CCHHHHHH-HHHHhhccCChhhhhhHHHHHHH-HHHHHHHHHHHHHhhc
Q 031973           66 KSVATVGK-AAGEKWKSMSEDEKAPFVERAEK-RKSDYNKNMQDYNKQL  112 (150)
Q Consensus        66 ~~~~eisk-~l~~~Wk~l~~eeK~~y~~~A~~-~k~~y~~~~~~y~~~~  112 (150)
                      .++..|++ .||..|+.+++++|+.|...... ....|-..+..|....
T Consensus        64 ~Df~~mar~vLG~~W~~~s~~Qr~~F~~~F~~~l~~tY~~~l~~y~~~~  112 (198)
T TIGR03481        64 FDLPAMARLTLGSSWTSLSPEQRRRFIGAFRELSIATYASQFKSYAGER  112 (198)
T ss_pred             CCHHHHHHHHhhhhhhhCCHHHHHHHHHHHHHHHHHHHHHHHHhhcCce
Confidence            56778876 58999999999999999998877 6788888898887653


No 26 
>PRK15117 ABC transporter periplasmic binding protein MlaC; Provisional
Probab=79.16  E-value=6.1  Score=30.98  Aligned_cols=49  Identities=18%  Similarity=0.354  Sum_probs=39.3

Q ss_pred             CCCCCHHHHHH-HHHHhhccCChhhhhhHHHHHHHH-HHHHHHHHHHHHhhc
Q 031973           63 PNNKSVATVGK-AAGEKWKSMSEDEKAPFVERAEKR-KSDYNKNMQDYNKQL  112 (150)
Q Consensus        63 p~~~~~~eisk-~l~~~Wk~l~~eeK~~y~~~A~~~-k~~y~~~~~~y~~~~  112 (150)
                      |. .++..|++ .||..|+.+++++|+.|...-... ..-|...+..|..+.
T Consensus        66 p~-~Df~~~s~~vLG~~wr~as~eQr~~F~~~F~~~Lv~tYa~~l~~y~~q~  116 (211)
T PRK15117         66 PY-VQVKYAGALVLGRYYKDATPAQREAYFAAFREYLKQAYGQALAMYHGQT  116 (211)
T ss_pred             cc-CCHHHHHHHHhhhhhhhCCHHHHHHHHHHHHHHHHHHHHHHHHHhCCce
Confidence            44 67777766 589999999999999999866655 568889999997653


No 27 
>PF05494 Tol_Tol_Ttg2:  Toluene tolerance, Ttg2 ;  InterPro: IPR008869 Toluene tolerance is mediated by increased cell membrane rigidity resulting from changes in fatty acid and phospholipid compositions, exclusion of toluene from the cell membrane, and removal of intracellular toluene by degradation []. Many proteins are involved in these processes. This family is a transporter which shows similarity to ABC transporters [].; PDB: 2QGU_A.
Probab=75.86  E-value=6  Score=29.49  Aligned_cols=45  Identities=20%  Similarity=0.451  Sum_probs=34.0

Q ss_pred             CCHHHHHHH-HHHhhccCChhhhhhHHHHHHHH-HHHHHHHHHHHHh
Q 031973           66 KSVATVGKA-AGEKWKSMSEDEKAPFVERAEKR-KSDYNKNMQDYNK  110 (150)
Q Consensus        66 ~~~~eisk~-l~~~Wk~l~~eeK~~y~~~A~~~-k~~y~~~~~~y~~  110 (150)
                      .++..|++. ||..|+.+++++++.|....... ...|...+..|..
T Consensus        38 ~D~~~~ar~~LG~~w~~~s~~q~~~F~~~f~~~l~~~Y~~~l~~y~~   84 (170)
T PF05494_consen   38 FDFERMARRVLGRYWRKASPAQRQRFVEAFKQLLVRTYAKRLDEYSG   84 (170)
T ss_dssp             B-HHHHHHHHHGGGTTTS-HHHHHHHHHHHHHHHHHHHHHHHHT-SS
T ss_pred             CCHHHHHHHHHHHhHhhCCHHHHHHHHHHHHHHHHHHHHHHHHhhCC
Confidence            677777765 77899999999999999876665 5678888888875


No 28 
>PF13875 DUF4202:  Domain of unknown function (DUF4202)
Probab=60.30  E-value=14  Score=28.70  Aligned_cols=39  Identities=26%  Similarity=0.485  Sum_probs=32.8

Q ss_pred             hHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhh
Q 031973           46 AFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKA   88 (150)
Q Consensus        46 aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~   88 (150)
                      +-++|...+...+...|.    ...+..+|...|+.||+..++
T Consensus       131 acLVFL~~~f~~F~~~~d----eeK~v~Il~KTw~KMS~~g~~  169 (185)
T PF13875_consen  131 ACLVFLEYYFEDFAAKHD----EEKIVDILRKTWRKMSERGHE  169 (185)
T ss_pred             HHHHhHHHHHHHHHhcCC----HHHHHHHHHHHHHHCCHHHHH
Confidence            467899999999988774    368889999999999998875


No 29 
>PF11304 DUF3106:  Protein of unknown function (DUF3106);  InterPro: IPR021455  Some members in this family of proteins are annotated as transmembrane proteins however this cannot be confirmed. Currently no function is known. 
Probab=52.60  E-value=58  Score=22.76  Aligned_cols=22  Identities=18%  Similarity=0.582  Sum_probs=9.7

Q ss_pred             HHHHHhhccCChhhhhhHHHHH
Q 031973           73 KAAGEKWKSMSEDEKAPFVERA   94 (150)
Q Consensus        73 k~l~~~Wk~l~~eeK~~y~~~A   94 (150)
                      .-+...|+.|++..+..+...|
T Consensus        14 ~pl~~~W~~l~~~qr~k~l~~a   35 (107)
T PF11304_consen   14 APLAERWNSLPPEQRRKWLQIA   35 (107)
T ss_pred             HHHHHHHhcCCHHHHHHHHHHH
Confidence            3344444444444444444333


No 30 
>COG2854 Ttg2D ABC-type transport system involved in resistance to organic solvents, auxiliary component [Secondary metabolites biosynthesis, transport, and catabolism]
Probab=48.43  E-value=22  Score=27.99  Aligned_cols=43  Identities=19%  Similarity=0.352  Sum_probs=35.6

Q ss_pred             HHHHHhhccCChhhhhhHHHHHHHH-HHHHHHHHHHHHhhccCC
Q 031973           73 KAAGEKWKSMSEDEKAPFVERAEKR-KSDYNKNMQDYNKQLADG  115 (150)
Q Consensus        73 k~l~~~Wk~l~~eeK~~y~~~A~~~-k~~y~~~~~~y~~~~~~~  115 (150)
                      ..||.-|+.+++++++.|....... ...|-..+..|+.+...-
T Consensus        78 ~vLGk~~k~aspeQ~~~F~~aF~~yl~q~Y~~aL~~Y~~q~~~v  121 (202)
T COG2854          78 LVLGKYYKTASPEQRQAFFKAFRTYLEQTYGQALLDYKGQTLKV  121 (202)
T ss_pred             HHhccccccCCHHHHHHHHHHHHHHHHHHHHHHHHHccCCCcee
Confidence            3488999999999999999876665 677999999999877543


No 31 
>PF01352 KRAB:  KRAB box;  InterPro: IPR001909 The Krueppel-associated box (KRAB) is a domain of around 75 amino acids that is found in the N-terminal part of about one third of eukaryotic Krueppel-type C2H2 zinc finger proteins (ZFPs) []. It is enriched in charged amino acids and can be divided into subregions A and B, which are predicted to fold into two amphipathic alpha-helices. The KRAB A and B boxes can be separated by variable spacer segments and many KRAB proteins contain only the A box []. The functions currently known for members of the KRAB-containing protein family include transcriptional repression of RNA polymerase I, II, and III promoters, binding and splicing of RNA, and control of nucleolus function. The KRAB domain functions as a transcriptional repressor when tethered to the template DNA by a DNA-binding domain. A sequence of 45 amino acids in the KRAB A subdomain has been shown to be necessary and sufficient for transcriptional repression. The B box does not repress by itself but does potentiate the repression exerted by the KRAB A subdomain [, ]. Gene silencing requires the binding of the KRAB domain to the RING-B box-coiled coil (RBCC) domain of the KAP-1/TIF1-beta corepressor. As KAP-1 binds to the heterochromatin proteins HP1, it has been proposed that the KRAB-ZFP-bound target gene could be silenced following recruitment to heterochromatin [, ]. KRAB-ZFPs probably constitute the single largest class of transcription factors within the human genome []. Although the function of KRAB-ZFPs is largely unknown, they appear to play important roles during cell differentiation and development. The KRAB domain is generally encoded by two exons. The regions coded by the two exons are known as KRAB-A and KRAB-B.; GO: 0003676 nucleic acid binding, 0006355 regulation of transcription, DNA-dependent, 0005622 intracellular; PDB: 1V65_A.
Probab=37.81  E-value=23  Score=20.56  Aligned_cols=26  Identities=15%  Similarity=0.311  Sum_probs=14.9

Q ss_pred             HHHHHHHH-HhhccCChhhhhhHHHHH
Q 031973           69 ATVGKAAG-EKWKSMSEDEKAPFVERA   94 (150)
Q Consensus        69 ~eisk~l~-~~Wk~l~~eeK~~y~~~A   94 (150)
                      .+|+--++ +.|..|.+.+|..|.+.-
T Consensus         4 ~Dvav~fs~eEW~~L~~~Qk~ly~dvm   30 (41)
T PF01352_consen    4 EDVAVYFSQEEWELLDPAQKNLYRDVM   30 (41)
T ss_dssp             ---TT---HHHHHTS-HHHHHHHHHHH
T ss_pred             EEEEEEcChhhcccccceecccchhHH
Confidence            34444444 669999999998887644


No 32 
>PRK09706 transcriptional repressor DicA; Reviewed
Probab=36.99  E-value=84  Score=22.34  Aligned_cols=43  Identities=19%  Similarity=0.233  Sum_probs=37.2

Q ss_pred             HHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHhhcc
Q 031973           71 VGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQLA  113 (150)
Q Consensus        71 isk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~~~~  113 (150)
                      -...|-..|+.|+++++.............|...+.+|-....
T Consensus        88 ~~~~ll~~~~~L~~~~~~~~l~~l~~~~~~~~~~~~~~~~~~~  130 (135)
T PRK09706         88 DQKELLELFDALPESEQDAQLSEMRARVENFNKLFEELLKARK  130 (135)
T ss_pred             HHHHHHHHHHHCCHHHHHHHHHHHHHHHHHHHHHHHHHHHHHh
Confidence            3467889999999999999999999999999999988876543


No 33 
>PF15076 DUF4543:  Domain of unknown function (DUF4543)
Probab=33.83  E-value=20  Score=23.34  Aligned_cols=22  Identities=14%  Similarity=0.546  Sum_probs=17.8

Q ss_pred             cCCCCCCCCCCCChHHHHHHHH
Q 031973           33 AAKDPNKPKRPPSAFFVFMEEF   54 (150)
Q Consensus        33 ~~~dp~~PKrP~~aY~lF~~~~   54 (150)
                      +...|++|.-||.-||+|++.-
T Consensus        25 r~~K~GfpdepmrE~ml~l~~L   46 (75)
T PF15076_consen   25 RPRKPGFPDEPMREYMLHLQAL   46 (75)
T ss_pred             CCCCCCCCcchHHHHHHHHHHH
Confidence            3556899999999999998643


No 34 
>PF12881 NUT_N:  NUT protein N terminus;  InterPro: IPR024309 This domain is found in the N-terminal region of Nuclear Testis (NUT) proteins. It is also found in FAM22, which are a family of uncharacterised mammalian proteins.
Probab=32.87  E-value=1.3e+02  Score=25.45  Aligned_cols=54  Identities=24%  Similarity=0.205  Sum_probs=38.2

Q ss_pred             HHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHH-HHHHHHHHHhhc
Q 031973           58 FKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSD-YNKNMQDYNKQL  112 (150)
Q Consensus        58 ~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~-y~~~~~~y~~~~  112 (150)
                      |....|. ++..|-..+.-+.|...+.-+|..|.++|.+-.+= -+++|+.-+.++
T Consensus       243 Lar~kPt-MtlEeGl~ra~qEW~~~SnfdRmifyemaekFmEFEaeEEmq~q~lq~  297 (328)
T PF12881_consen  243 LARLKPT-MTLEEGLWRAVQEWQHTSNFDRMIFYEMAEKFMEFEAEEEMQIQKLQL  297 (328)
T ss_pred             HHhcCCC-ccHHHHHHHHHHHhhccccccHHHHHHHHHHHccCCcHHHHHHHHHHH
Confidence            4445565 77778777788999999999999999999887531 124555544443


No 35 
>PRK10363 cpxP periplasmic repressor CpxP; Reviewed
Probab=32.26  E-value=1.2e+02  Score=23.27  Aligned_cols=36  Identities=11%  Similarity=0.304  Sum_probs=27.5

Q ss_pred             HHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHH
Q 031973           68 VATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNK  103 (150)
Q Consensus        68 ~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~  103 (150)
                      ..++.+.-.++++-|+|++|..|.....+-...+..
T Consensus       110 ~Vem~k~~nqmy~lLTPEQKaq~~~~~~~rm~~~~~  145 (166)
T PRK10363        110 QVEMAKVRNQMYRLLTPEQQAVLNEKHQQRMEQLRD  145 (166)
T ss_pred             HHHHHHHHHHHHHhCCHHHHHHHHHHHHHHHHHHHH
Confidence            345666677899999999999998877766655543


No 36 
>PRK12750 cpxP periplasmic repressor CpxP; Reviewed
Probab=30.93  E-value=1.3e+02  Score=22.82  Aligned_cols=33  Identities=18%  Similarity=0.320  Sum_probs=25.6

Q ss_pred             HHHHHhhccCChhhhhhHHHHHHHHHHHHHHHH
Q 031973           73 KAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNM  105 (150)
Q Consensus        73 k~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~  105 (150)
                      +...+++..|++++|..|.+....-...+...+
T Consensus       128 ~~~~~~~~vLTpEQRak~~e~~~~r~~~~~~~~  160 (170)
T PRK12750        128 EKRHQMLSILTPEQKAKFQELQQERMQECQDKM  160 (170)
T ss_pred             HHHHHHHHhCCHHHHHHHHHHHHHHHHHHHHHH
Confidence            335568999999999999988777766666655


No 37 
>PRK12751 cpxP periplasmic stress adaptor protein CpxP; Reviewed
Probab=30.12  E-value=1.2e+02  Score=22.97  Aligned_cols=31  Identities=10%  Similarity=0.314  Sum_probs=22.3

Q ss_pred             HHHHHHHhhccCChhhhhhHHHHHHHHHHHH
Q 031973           71 VGKAAGEKWKSMSEDEKAPFVERAEKRKSDY  101 (150)
Q Consensus        71 isk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y  101 (150)
                      +.+...++++.|++++|..|.+...+-....
T Consensus       119 ~~~~~~qmy~lLTPEQra~l~~~~e~r~~~~  149 (162)
T PRK12751        119 MAKVRNQMYNLLTPEQKEALNKKHQERIEKL  149 (162)
T ss_pred             HHHHHHHHHHcCCHHHHHHHHHHHHHHHHHH
Confidence            3444567889999999999988665554433


No 38 
>PF00887 ACBP:  Acyl CoA binding protein;  InterPro: IPR000582 Acyl-CoA-binding protein (ACBP) is a small (10 Kd) protein that binds medium- and long-chain acyl-CoA esters with very high affinity and may function as an intracellular carrier of acyl-CoA esters []. ACBP is also known as diazepam binding inhibitor (DBI) or endozepine (EP) because of its ability to displace diazepam from the benzodiazepine (BZD) recognition site located on the GABA type A receptor. It is therefore possible that this protein also acts as a neuropeptide to modulate the action of the GABA receptor []. ACBP is a highly conserved protein of about 90 residues that is found in all four eukaryotic kingdoms, Animalia, Plantae, Fungi and Protista, and in some eubacterial species []. Although ACBP occurs as a completely independent protein, intact ACB domains have been identified in a number of large, multifunctional proteins in a variety of eukaryotic species. These include large membrane-associated proteins with N-terminal ACB domains, multifunctional enzymes with both ACB and peroxisomal enoyl-CoA Delta(3), Delta(2)-enoyl-CoA isomerase domains, and proteins with both an ACB domain and ankyrin repeats (IPR002110 from INTERPRO) []. The ACB domain consists of four alpha-helices arranged in a bowl shape with a highly exposed acyl-CoA-binding site. The ligand is bound through specific interactions with residues on the protein, most notably several conserved positive charges that interact with the phosphate group on the adenosine-3'phosphate moiety, and the acyl chain is sandwiched between the hydrophobic surfaces of CoA and the protein []. Other proteins containing an ACB domain include:   Endozepine-like peptide (ELP) (gene DBIL5) from mouse []. ELP is a testis-specific ACBP homologue that may be involved in the energy metabolism of the mature sperm. MA-DBI, a transmembrane protein of unknown function which has been found in mammals. MA-DBI contains a N-terminal ACB domain. DRS-1 [], a human protein of unknown function that contains a N-terminal ACB domain and a C-terminal enoyl-CoA isomerase/hydratase domain.  ; GO: 0000062 fatty-acyl-CoA binding; PDB: 2CB8_A 2FJ9_A 2LBB_A 1ST7_A 3EPY_B 2FDQ_C 1NTI_A 1HB8_A 1ACA_A 1NVL_A ....
Probab=24.00  E-value=2.2e+02  Score=18.66  Aligned_cols=54  Identities=11%  Similarity=0.261  Sum_probs=28.9

Q ss_pred             ChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCCh----hhhhhHHHHHHHHHHH
Q 031973           45 SAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSE----DEKAPFVERAEKRKSD  100 (150)
Q Consensus        45 ~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~----eeK~~y~~~A~~~k~~  100 (150)
                      .-|-||.+.....+....|+..+....  .--.-|+.+..    +-+..|.+........
T Consensus        28 ~LYalyKQAt~Gd~~~~~P~~~d~~~~--~K~~AW~~l~gms~~eA~~~Yi~~v~~~~~~   85 (87)
T PF00887_consen   28 ELYALYKQATHGDCDTPRPGFFDIEGR--AKWDAWKALKGMSKEEAMREYIELVEELIPK   85 (87)
T ss_dssp             HHHHHHHHHHTSS--S-CTTTTCHHHH--HHHHHHHTTTTTHHHHHHHHHHHHHHHHHHH
T ss_pred             HHHHHHHHHHhCCCcCCCCcchhHHHH--HHHHHHHHccCCCHHHHHHHHHHHHHHHHHh
Confidence            348888888877666666763333333  33466887763    3344555555544433


No 39 
>PF02209 VHP:  Villin headpiece domain;  InterPro: IPR003128 Villin is an F-actin bundling protein involved in the maintenance of the microvilli of the absorptive epithelia. The villin-type "headpiece" domain is a modular motif found at the extreme C terminus of larger "core" domains in over 25 cytoskeletal proteins in plants and animals, often in assocation with the Gelsolin repeat. Although the headpiece is classified as an F-actin-binding domain, it has been shown that not all headpiece domains are intrinsically F-actin-binding motifs, surface charge distribution may be an important element for F-actin recognition []. An autonomously folding, 35 residue, thermostable subdomain (HP36) of the full-length 76 amino acid residue villin headpiece, is the smallest known example of a cooperatively folded domain of a naturally occurring protein. The structure of HP36, as determined by NMR spectroscopy, consists of three short helices surrounding a tightly packed hydrophobic core []. ; GO: 0003779 actin binding, 0007010 cytoskeleton organization; PDB: 1ZV6_A 1QZP_A 1UND_A 2PPZ_A 3TJW_B 1YU8_X 2JM0_A 1WY4_A 3MYC_A 1YU5_X ....
Probab=23.47  E-value=45  Score=18.89  Aligned_cols=10  Identities=30%  Similarity=0.381  Sum_probs=7.3

Q ss_pred             hhHHHHhhhc
Q 031973          140 SGEVFLQLLT  149 (150)
Q Consensus       140 ~deef~~~~~  149 (150)
                      +||||..+|.
T Consensus         3 sd~dF~~vFg   12 (36)
T PF02209_consen    3 SDEDFEKVFG   12 (36)
T ss_dssp             -HHHHHHHHS
T ss_pred             CHHHHHHHHC
Confidence            6889998873


No 40 
>TIGR00787 dctP tripartite ATP-independent periplasmic transporter solute receptor, DctP family. TRAP-T (Tripartite ATP-independent Periplasmic Transporter) family proteins generally consist of three components, and these systems have so far been found in Gram-negative bacteria, Gram-postive bacteria and archaea. The best characterized example is the DctPQM system of Rhodobacter capsulatus, a C4 dicarboxylate (malate, fumarate, succinate) transporter. This model represents the DctP family, one of at least three major families of extracytoplasmic solute receptor for TRAP family transporters. Other are the SnoM family (see pfam03480) and TAXI (TRAP-associated extracytoplasmic immunogenic) family.
Probab=22.72  E-value=1.7e+02  Score=22.85  Aligned_cols=28  Identities=29%  Similarity=0.306  Sum_probs=21.4

Q ss_pred             HHhhccCChhhhhhHHHHHHHHHHHHHH
Q 031973           76 GEKWKSMSEDEKAPFVERAEKRKSDYNK  103 (150)
Q Consensus        76 ~~~Wk~l~~eeK~~y~~~A~~~k~~y~~  103 (150)
                      ...|..||++.|....+.+...-.....
T Consensus       213 ~~~~~~L~~e~q~~i~~a~~~~~~~~~~  240 (257)
T TIGR00787       213 KAFWKSLPPDLQAVVKEAAKEAGEYQRK  240 (257)
T ss_pred             HHHHhcCCHHHHHHHHHHHHHHHHHHHH
Confidence            4779999999999998877766544443


No 41 
>PRK10236 hypothetical protein; Provisional
Probab=22.39  E-value=80  Score=25.51  Aligned_cols=25  Identities=24%  Similarity=0.506  Sum_probs=20.2

Q ss_pred             HHHHHHHHhhccCChhhhhhHHHHH
Q 031973           70 TVGKAAGEKWKSMSEDEKAPFVERA   94 (150)
Q Consensus        70 eisk~l~~~Wk~l~~eeK~~y~~~A   94 (150)
                      -+.+.+...|..||+++++.+...-
T Consensus       117 il~kll~~a~~kms~eE~~~L~~~l  141 (237)
T PRK10236        117 LLEQFLRNTWKKMDEEHKQEFLHAV  141 (237)
T ss_pred             HHHHHHHHHHHHCCHHHHHHHHHHH
Confidence            3677899999999999998776543


No 42 
>PHA02819 hypothetical protein; Provisional
Probab=22.26  E-value=47  Score=21.81  Aligned_cols=11  Identities=18%  Similarity=0.371  Sum_probs=7.6

Q ss_pred             cchhHHHHhhh
Q 031973          138 EGSGEVFLQLL  148 (150)
Q Consensus       138 ~~~deef~~~~  148 (150)
                      +.+||+|+++|
T Consensus        14 sS~DdDFnnFI   24 (71)
T PHA02819         14 SSSDDDFNNFI   24 (71)
T ss_pred             CCchhHHHHHH
Confidence            35677787776


No 43 
>PF06945 DUF1289:  Protein of unknown function (DUF1289);  InterPro: IPR010710 This family consists of a number of hypothetical bacterial proteins. The aligned region spans around 56 residues and contains 4 highly conserved cysteine residues towards the N terminus. The function of this family is unknown.
Probab=21.91  E-value=1.2e+02  Score=18.14  Aligned_cols=23  Identities=35%  Similarity=0.744  Sum_probs=16.7

Q ss_pred             CHHHHHHHHHHhhccCChhhhhhHHHHH
Q 031973           67 SVATVGKAAGEKWKSMSEDEKAPFVERA   94 (150)
Q Consensus        67 ~~~eisk~l~~~Wk~l~~eeK~~y~~~A   94 (150)
                      +..||..     |..|++++|.......
T Consensus        23 T~dEI~~-----W~~~s~~er~~i~~~l   45 (51)
T PF06945_consen   23 TLDEIRD-----WKSMSDDERRAILARL   45 (51)
T ss_pred             cHHHHHH-----HhhCCHHHHHHHHHHH
Confidence            5667754     9999999987655433


No 44 
>PHA02975 hypothetical protein; Provisional
Probab=20.77  E-value=53  Score=21.45  Aligned_cols=11  Identities=18%  Similarity=0.371  Sum_probs=7.3

Q ss_pred             cchhHHHHhhh
Q 031973          138 EGSGEVFLQLL  148 (150)
Q Consensus       138 ~~~deef~~~~  148 (150)
                      +..||+|+++|
T Consensus        14 sS~DdDF~nFI   24 (69)
T PHA02975         14 ESNDSDFEDFI   24 (69)
T ss_pred             CCChHHHHHHH
Confidence            34677777776


No 45 
>KOG1610 consensus Corticosteroid 11-beta-dehydrogenase and related short chain-type dehydrogenases [Secondary metabolites biosynthesis, transport and catabolism; General function prediction only]
Probab=20.67  E-value=2.3e+02  Score=23.93  Aligned_cols=49  Identities=16%  Similarity=0.342  Sum_probs=35.1

Q ss_pred             HHHHHHHHHHHhC-------CCC-----CCHHHHHHHHHHhhccCChhhhhhHHHHHHHHH
Q 031973           50 FMEEFRKQFKEAH-------PNN-----KSVATVGKAAGEKWKSMSEDEKAPFVERAEKRK   98 (150)
Q Consensus        50 F~~~~r~~~k~~~-------p~~-----~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k   98 (150)
                      |+...|.++..-.       |+.     .....+.+.+.++|..||++.|+.|=+.+-.+.
T Consensus       188 f~D~lR~EL~~fGV~VsiiePG~f~T~l~~~~~~~~~~~~~w~~l~~e~k~~YGedy~~~~  248 (322)
T KOG1610|consen  188 FSDSLRRELRPFGVKVSIIEPGFFKTNLANPEKLEKRMKEIWERLPQETKDEYGEDYFEDY  248 (322)
T ss_pred             HHHHHHHHHHhcCcEEEEeccCccccccCChHHHHHHHHHHHhcCCHHHHHHHHHHHHHHH
Confidence            6666676665322       321     245788899999999999999999987665543


No 46 
>PHA02844 putative transmembrane protein; Provisional
Probab=20.48  E-value=54  Score=21.76  Aligned_cols=11  Identities=18%  Similarity=0.347  Sum_probs=7.4

Q ss_pred             cchhHHHHhhh
Q 031973          138 EGSGEVFLQLL  148 (150)
Q Consensus       138 ~~~deef~~~~  148 (150)
                      +.+||+|+++|
T Consensus        14 sS~DdDFnnFI   24 (75)
T PHA02844         14 SSENEDFNNFI   24 (75)
T ss_pred             CCchHHHHHHH
Confidence            34677777776


Done!