Query 031973
Match_columns 150
No_of_seqs 112 out of 1110
Neff 7.0
Searched_HMMs 46136
Date Fri Mar 29 07:53:51 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/031973.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/031973hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PTZ00199 high mobility group p 99.9 2.2E-23 4.7E-28 144.8 11.3 84 25-109 8-93 (94)
2 cd01389 MATA_HMG-box MATA_HMG- 99.8 5.2E-21 1.1E-25 127.9 7.9 73 39-112 1-73 (77)
3 cd01388 SOX-TCF_HMG-box SOX-TC 99.8 1.7E-20 3.7E-25 123.9 8.0 71 39-110 1-71 (72)
4 PF00505 HMG_box: HMG (high mo 99.8 8.6E-20 1.9E-24 118.5 8.9 69 40-109 1-69 (69)
5 cd01390 HMGB-UBF_HMG-box HMGB- 99.8 3.9E-19 8.5E-24 114.4 8.8 65 40-105 1-65 (66)
6 smart00398 HMG high mobility g 99.8 4.7E-19 1E-23 114.7 9.0 70 39-109 1-70 (70)
7 PF09011 HMG_box_2: HMG-box do 99.8 1.4E-18 3.1E-23 115.0 9.0 72 37-109 1-73 (73)
8 COG5648 NHP6B Chromatin-associ 99.8 2.6E-18 5.6E-23 133.3 7.7 93 25-118 56-148 (211)
9 cd00084 HMG-box High Mobility 99.7 2.3E-17 5E-22 105.6 8.8 65 40-105 1-65 (66)
10 KOG0527 HMG-box transcription 99.7 7.7E-18 1.7E-22 139.8 6.8 85 32-117 55-139 (331)
11 KOG0381 HMG box-containing pro 99.7 1.6E-16 3.5E-21 109.7 10.9 76 36-112 17-95 (96)
12 KOG0526 Nucleosome-binding fac 99.6 2.4E-15 5.2E-20 129.7 7.7 81 25-110 521-601 (615)
13 KOG3248 Transcription factor T 99.5 9.2E-14 2E-18 114.5 5.8 79 38-117 190-268 (421)
14 KOG0528 HMG-box transcription 99.2 7.6E-12 1.6E-16 107.1 3.7 85 29-114 315-399 (511)
15 KOG4715 SWI/SNF-related matrix 99.2 7.6E-11 1.6E-15 96.7 8.9 80 32-112 57-136 (410)
16 KOG2746 HMG-box transcription 98.7 1.9E-08 4.1E-13 89.3 4.7 74 30-104 172-247 (683)
17 PF14887 HMG_box_5: HMG (high 98.0 3.2E-05 6.9E-10 51.7 7.3 75 39-115 3-77 (85)
18 PF06382 DUF1074: Protein of u 97.4 0.00083 1.8E-08 51.5 7.7 50 44-98 83-132 (183)
19 PF04690 YABBY: YABBY protein; 97.2 0.00095 2E-08 51.0 5.7 49 34-83 116-164 (170)
20 COG5648 NHP6B Chromatin-associ 96.9 0.00069 1.5E-08 53.2 2.9 68 38-106 142-209 (211)
21 PF08073 CHDNT: CHDNT (NUC034) 95.7 0.014 3E-07 36.6 3.1 39 45-84 14-52 (55)
22 PF06244 DUF1014: Protein of u 94.1 0.084 1.8E-06 38.3 3.9 49 36-85 68-117 (122)
23 PF04769 MAT_Alpha1: Mating-ty 93.5 0.2 4.4E-06 39.3 5.5 56 34-96 38-93 (201)
24 KOG3223 Uncharacterized conser 89.7 0.23 4.9E-06 38.8 2.0 55 36-94 160-215 (221)
25 TIGR03481 HpnM hopanoid biosyn 84.5 2.4 5.1E-05 33.0 5.0 47 66-112 64-112 (198)
26 PRK15117 ABC transporter perip 79.2 6.1 0.00013 31.0 5.6 49 63-112 66-116 (211)
27 PF05494 Tol_Tol_Ttg2: Toluene 75.9 6 0.00013 29.5 4.6 45 66-110 38-84 (170)
28 PF13875 DUF4202: Domain of un 60.3 14 0.0003 28.7 3.8 39 46-88 131-169 (185)
29 PF11304 DUF3106: Protein of u 52.6 58 0.0013 22.8 5.6 22 73-94 14-35 (107)
30 COG2854 Ttg2D ABC-type transpo 48.4 22 0.00048 28.0 3.2 43 73-115 78-121 (202)
31 PF01352 KRAB: KRAB box; Inte 37.8 23 0.00049 20.6 1.4 26 69-94 4-30 (41)
32 PRK09706 transcriptional repre 37.0 84 0.0018 22.3 4.6 43 71-113 88-130 (135)
33 PF15076 DUF4543: Domain of un 33.8 20 0.00044 23.3 0.8 22 33-54 25-46 (75)
34 PF12881 NUT_N: NUT protein N 32.9 1.3E+02 0.0028 25.5 5.5 54 58-112 243-297 (328)
35 PRK10363 cpxP periplasmic repr 32.3 1.2E+02 0.0025 23.3 4.8 36 68-103 110-145 (166)
36 PRK12750 cpxP periplasmic repr 30.9 1.3E+02 0.0028 22.8 4.9 33 73-105 128-160 (170)
37 PRK12751 cpxP periplasmic stre 30.1 1.2E+02 0.0026 23.0 4.6 31 71-101 119-149 (162)
38 PF00887 ACBP: Acyl CoA bindin 24.0 2.2E+02 0.0047 18.7 4.8 54 45-100 28-85 (87)
39 PF02209 VHP: Villin headpiece 23.5 45 0.00098 18.9 1.0 10 140-149 3-12 (36)
40 TIGR00787 dctP tripartite ATP- 22.7 1.7E+02 0.0038 22.9 4.6 28 76-103 213-240 (257)
41 PRK10236 hypothetical protein; 22.4 80 0.0017 25.5 2.5 25 70-94 117-141 (237)
42 PHA02819 hypothetical protein; 22.3 47 0.001 21.8 1.0 11 138-148 14-24 (71)
43 PF06945 DUF1289: Protein of u 21.9 1.2E+02 0.0027 18.1 2.8 23 67-94 23-45 (51)
44 PHA02975 hypothetical protein; 20.8 53 0.0011 21.4 1.0 11 138-148 14-24 (69)
45 KOG1610 Corticosteroid 11-beta 20.7 2.3E+02 0.0051 23.9 5.0 49 50-98 188-248 (322)
46 PHA02844 putative transmembran 20.5 54 0.0012 21.8 1.0 11 138-148 14-24 (75)
No 1
>PTZ00199 high mobility group protein; Provisional
Probab=99.90 E-value=2.2e-23 Score=144.82 Aligned_cols=84 Identities=38% Similarity=0.662 Sum_probs=77.1
Q ss_pred CccccCCccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCC--HHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHH
Q 031973 25 GKRTAKPKAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKS--VATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYN 102 (150)
Q Consensus 25 ~k~kk~kk~~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~--~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~ 102 (150)
+.++++++..+||+.|+||+|||||||+++|..|..+||+ ++ +.+|+++||++|+.|++++|.+|+++|..++.+|.
T Consensus 8 ~~~k~~~k~~kdp~~PKrP~sAY~~F~~~~R~~i~~~~P~-~~~~~~evsk~ige~Wk~ls~eeK~~y~~~A~~dk~rY~ 86 (94)
T PTZ00199 8 VLVRKNKRKKKDPNAPKRALSAYMFFAKEKRAEIIAENPE-LAKDVAAVGKMVGEAWNKLSEEEKAPYEKKAQEDKVRYE 86 (94)
T ss_pred ccccccCCCCCCCCCCCCCCcHHHHHHHHHHHHHHHHCcC-CcccHHHHHHHHHHHHHcCCHHHHHHHHHHHHHHHHHHH
Confidence 3444455668999999999999999999999999999999 75 89999999999999999999999999999999999
Q ss_pred HHHHHHH
Q 031973 103 KNMQDYN 109 (150)
Q Consensus 103 ~~~~~y~ 109 (150)
.+|..|.
T Consensus 87 ~e~~~Y~ 93 (94)
T PTZ00199 87 KEKAEYA 93 (94)
T ss_pred HHHHHHh
Confidence 9999995
No 2
>cd01389 MATA_HMG-box MATA_HMG-box, class I member of the HMG-box superfamily of DNA-binding proteins. These proteins contain a single HMG box, and bind the minor groove of DNA in a highly sequence-specific manner. Members include the fungal mating type gene products MC, MATA1 and Ste11.
Probab=99.84 E-value=5.2e-21 Score=127.85 Aligned_cols=73 Identities=27% Similarity=0.459 Sum_probs=70.5
Q ss_pred CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHhhc
Q 031973 39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQL 112 (150)
Q Consensus 39 ~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~~~ 112 (150)
+|+||+||||||+++.|..|+.+||+ +++.+|+++||.+|+.|++++|++|.++|..++++|..++++|+...
T Consensus 1 ~~kRP~naf~lf~~~~r~~~~~~~p~-~~~~eisk~~g~~Wk~ls~eeK~~y~~~A~~~k~~~~~~~p~Yky~p 73 (77)
T cd01389 1 KIPRPRNAFILYRQDKHAQLKTENPG-LTNNEISRIIGRMWRSESPEVKAYYKELAEEEKERHAREYPDYKYTP 73 (77)
T ss_pred CCCCCCcHHHHHHHHHHHHHHHHCCC-CCHHHHHHHHHHHHhhCCHHHHHHHHHHHHHHHHHHHHHCCCCcccC
Confidence 58999999999999999999999999 99999999999999999999999999999999999999999998754
No 3
>cd01388 SOX-TCF_HMG-box SOX-TCF_HMG-box, class I member of the HMG-box superfamily of DNA-binding proteins. These proteins contain a single HMG box, and bind the minor groove of DNA in a highly sequence-specific manner. Members include SRY and its homologs in insects and vertebrates, and transcription factor-like proteins, TCF-1, -3, -4, and LEF-1. They appear to bind the minor groove of the A/T C A A A G/C-motif.
Probab=99.83 E-value=1.7e-20 Score=123.91 Aligned_cols=71 Identities=34% Similarity=0.561 Sum_probs=68.6
Q ss_pred CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHh
Q 031973 39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNK 110 (150)
Q Consensus 39 ~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~ 110 (150)
++|||+|||||||+++|..++.+||+ +++.+|+++||.+|+.|++++|++|.++|..++++|..++++|+.
T Consensus 1 ~iKrP~naf~~F~~~~r~~~~~~~p~-~~~~eisk~l~~~Wk~ls~~eK~~y~~~a~~~k~~y~~~~p~y~y 71 (72)
T cd01388 1 HIKRPMNAFMLFSKRHRRKVLQEYPL-KENRAISKILGDRWKALSNEEKQPYYEEAKKLKELHMKLYPDYKW 71 (72)
T ss_pred CCCCCCcHHHHHHHHHHHHHHHHCCC-CCHHHHHHHHHHHHHcCCHHHHHHHHHHHHHHHHHHHHHCcCCCC
Confidence 47899999999999999999999999 999999999999999999999999999999999999999999863
No 4
>PF00505 HMG_box: HMG (high mobility group) box; InterPro: IPR000910 High mobility group (HMG or HMGB) proteins are a family of relatively low molecular weight non-histone components in chromatin. HMG1 (also called HMG-T in fish) and HMG2 are two highly related proteins that bind single-stranded DNA preferentially and unwind double-stranded DNA. Although they have no sequence specificity, they have a high affinity for bent or distorted DNA, and bend linear DNA. HMG1 and HMG2 contain two DNA-binding HMG-box domains (A and B) that show structural and functional differences, and have a long acidic C-terminal domain rich in aspartic and glutamic acid residues. The acidic tail modulates the affinity of the tandem HMG boxes in HMG1 and 2 for a variety of DNA targets. HMG1 and 2 appear to play important architectural roles in the assembly of nucleoprotein complexes in a variety of biological processes, for example V(D)J recombination, the initiation of transcription, and DNA repair []. The profile in this entry describing the HMG-domains is much more general than the signature. In addition to the HMG1 and HMG2 proteins, HMG-domains occur in single or multiple copies in the following protein classes; the SOX family of transcription factors; SRY sex determining region Y protein and related proteins []; LEF1 lymphoid enhancer binding factor 1 []; SSRP recombination signal recognition protein; MTF1 mitochondrial transcription factor 1; UBF1/2 nucleolar transcription factors; Abf2 yeast ARS-binding factor []; and Saccharomyces cerevisiae transcription factors Ixr1, Rox1, Nhp6a, Nhp6b and Spp41.; GO: 0003677 DNA binding; PDB: 1I11_A 1J3C_A 1J3D_A 1WZ6_A 1WGF_A 2D7L_A 1GT0_D 3U2B_C 2CRJ_A 2CS1_A ....
Probab=99.82 E-value=8.6e-20 Score=118.54 Aligned_cols=69 Identities=45% Similarity=0.836 Sum_probs=65.8
Q ss_pred CCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHH
Q 031973 40 PKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN 109 (150)
Q Consensus 40 PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~ 109 (150)
|+||+|||+|||.+++..++.+||+ ++..+|+++||.+|+.|++++|.+|.+.|..++..|..+|+.|+
T Consensus 1 PkrP~~af~lf~~~~~~~~k~~~p~-~~~~~i~~~~~~~W~~l~~~eK~~y~~~a~~~~~~y~~~~~~y~ 69 (69)
T PF00505_consen 1 PKRPPNAFMLFCKEKRAKLKEENPD-LSNKEISKILAQMWKNLSEEEKAPYKEEAEEEKERYEKEMPEYK 69 (69)
T ss_dssp SSSS--HHHHHHHHHHHHHHHHSTT-STHHHHHHHHHHHHHCSHHHHHHHHHHHHHHHHHHHHHHHHHHH
T ss_pred CcCCCCHHHHHHHHHHHHHHHHhcc-cccccchhhHHHHHhcCCHHHHHHHHHHHHHHHHHHHHHHHhcC
Confidence 8999999999999999999999999 99999999999999999999999999999999999999999995
No 5
>cd01390 HMGB-UBF_HMG-box HMGB-UBF_HMG-box, class II and III members of the HMG-box superfamily of DNA-binding proteins. These proteins bind the minor groove of DNA in a non-sequence specific fashion and contain two or more tandem HMG boxes. Class II members include non-histone chromosomal proteins, HMG1 and HMG2, which bind to bent or distorted DNA such as four-way DNA junctions, synthetic DNA cruciforms, kinked cisplatin-modified DNA, DNA bulges, cross-overs in supercoiled DNA, and can cause looping of linear DNA. Class III members include nucleolar and mitochondrial transcription factors, UBF and mtTF1, which bind four-way DNA junctions.
Probab=99.80 E-value=3.9e-19 Score=114.36 Aligned_cols=65 Identities=51% Similarity=0.854 Sum_probs=63.6
Q ss_pred CCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHH
Q 031973 40 PKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNM 105 (150)
Q Consensus 40 PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~ 105 (150)
|++|+|||+||++++|..++..||+ +++.+|++.||.+|+.|++++|.+|.+.|..++.+|..+|
T Consensus 1 Pkrp~saf~~f~~~~r~~~~~~~p~-~~~~~i~~~~~~~W~~ls~~eK~~y~~~a~~~~~~y~~e~ 65 (66)
T cd01390 1 PKRPLSAYFLFSQEQRPKLKKENPD-ASVTEVTKILGEKWKELSEEEKKKYEEKAEKDKERYEKEM 65 (66)
T ss_pred CCCCCcHHHHHHHHHHHHHHHHCcC-CCHHHHHHHHHHHHHhCCHHHHHHHHHHHHHHHHHHHHhh
Confidence 8999999999999999999999999 9999999999999999999999999999999999999876
No 6
>smart00398 HMG high mobility group.
Probab=99.80 E-value=4.7e-19 Score=114.71 Aligned_cols=70 Identities=47% Similarity=0.831 Sum_probs=67.9
Q ss_pred CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHH
Q 031973 39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN 109 (150)
Q Consensus 39 ~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~ 109 (150)
+|++|+|||+||++++|..+..+||+ +++.+|++.||.+|+.|++++|.+|.++|..++.+|...++.|.
T Consensus 1 ~pkrp~~~y~~f~~~~r~~~~~~~~~-~~~~~i~~~~~~~W~~l~~~ek~~y~~~a~~~~~~y~~~~~~y~ 70 (70)
T smart00398 1 KPKRPMSAFMLFSQENRAKIKAENPD-LSNAEISKKLGERWKLLSEEEKAPYEEKAKKDKERYEEEMPEYK 70 (70)
T ss_pred CcCCCCcHHHHHHHHHHHHHHHHCcC-CCHHHHHHHHHHHHHcCCHHHHHHHHHHHHHHHHHHHHHHHhcC
Confidence 58999999999999999999999999 99999999999999999999999999999999999999999884
No 7
>PF09011 HMG_box_2: HMG-box domain; InterPro: IPR015101 This domain is predominantly found in Maelstrom homologue proteins. It has no known function. ; GO: 0005634 nucleus; PDB: 2EQZ_A 1V64_A 2CTO_A 1H5P_A 3TQ6_A 3FGH_A 3TMM_A 1J3X_A 2YRQ_A 1AAB_A ....
Probab=99.78 E-value=1.4e-18 Score=114.97 Aligned_cols=72 Identities=49% Similarity=0.876 Sum_probs=63.8
Q ss_pred CCCCCCCCChHHHHHHHHHHHHHHh-CCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHH
Q 031973 37 PNKPKRPPSAFFVFMEEFRKQFKEA-HPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN 109 (150)
Q Consensus 37 p~~PKrP~~aY~lF~~~~r~~~k~~-~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~ 109 (150)
|++||+|+|||+||+.+++..++.. ++. ....++++.|+..|+.||+++|.+|.++|..++.+|..+|..|.
T Consensus 1 p~kpK~~~say~lF~~~~~~~~k~~G~~~-~~~~e~~k~~~~~Wk~Ls~~EK~~Y~~~A~~~k~~y~~e~~~~~ 73 (73)
T PF09011_consen 1 PKKPKRPPSAYNLFMKEMRKEVKEEGGQK-QSFREVMKEISERWKSLSEEEKEPYEERAKEDKERYEREMKEWN 73 (73)
T ss_dssp SSS--SSSSHHHHHHHHHHHHHHHHT-T--SSHHHHHHHHHHHHHHS-HHHHHHHHHHHHHHHHHHHHHHHHH-
T ss_pred CcCCCCCCCHHHHHHHHHHHHHHHhcccC-CCHHHHHHHHHHHHHhcCHHHHHHHHHHHHHHHHHHHHHHHhcC
Confidence 6899999999999999999999998 665 88999999999999999999999999999999999999999984
No 8
>COG5648 NHP6B Chromatin-associated proteins containing the HMG domain [Chromatin structure and dynamics]
Probab=99.75 E-value=2.6e-18 Score=133.26 Aligned_cols=93 Identities=33% Similarity=0.676 Sum_probs=86.2
Q ss_pred CccccCCccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHH
Q 031973 25 GKRTAKPKAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKN 104 (150)
Q Consensus 25 ~k~kk~kk~~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~ 104 (150)
++.+..++..+||+.|+||++|||+|+.++|++|+..+|. +++.+|+++||++|++|++++|++|...|..++++|...
T Consensus 56 ~ksk~~~r~k~dpN~PKRp~sayf~y~~~~R~ei~~~~p~-l~~~e~~k~~~e~WK~Ltd~eke~y~k~~~~~~erYq~e 134 (211)
T COG5648 56 TKSKRLVRKKKDPNGPKRPLSAYFLYSAENRDEIRKENPK-LTFGEVGKLLSEKWKELTDEEKEPYYKEANSDRERYQRE 134 (211)
T ss_pred hHHHHHHHHhcCCCCCCCchhHHHHHHHHHHHHHHHhCCC-CChHHHHHHHHHHHHhccHhhhhhHHHHHhhHHHHHHHH
Confidence 4445667889999999999999999999999999999999 999999999999999999999999999999999999999
Q ss_pred HHHHHhhccCCcch
Q 031973 105 MQDYNKQLADGVNA 118 (150)
Q Consensus 105 ~~~y~~~~~~~~~~ 118 (150)
+..|....+.....
T Consensus 135 k~~y~~k~~~~~~~ 148 (211)
T COG5648 135 KEEYNKKLPNKAPI 148 (211)
T ss_pred HHhhhcccCCCCCC
Confidence 99999988775544
No 9
>cd00084 HMG-box High Mobility Group (HMG)-box is found in a variety of eukaryotic chromosomal proteins and transcription factors. HMGs bind to the minor groove of DNA and have been classified by DNA binding preferences. Two phylogenically distinct groups of Class I proteins bind DNA in a sequence specific fashion and contain a single HMG box. One group (SOX-TCF) includes transcription factors, TCF-1, -3, -4; and also SRY and LEF-1, which bind four-way DNA junctions and duplex DNA targets. The second group (MATA) includes fungal mating type gene products MC, MATA1 and Ste11. Class II and III proteins (HMGB-UBF) bind DNA in a non-sequence specific fashion and contain two or more tandem HMG boxes. Class II members include non-histone chromosomal proteins, HMG1 and HMG2, which bind to bent or distorted DNA such as four-way DNA junctions, synthetic DNA cruciforms, kinked cisplatin-modified DNA, DNA bulges, cross-overs in supercoiled DNA, and can cause looping of linear DNA. Class III member
Probab=99.73 E-value=2.3e-17 Score=105.57 Aligned_cols=65 Identities=49% Similarity=0.827 Sum_probs=63.1
Q ss_pred CCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHH
Q 031973 40 PKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNM 105 (150)
Q Consensus 40 PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~ 105 (150)
|++|+|||+||+++.+..++..||+ ++..+|++.||.+|+.|++++|.+|.+.|..++..|..++
T Consensus 1 pkrp~~af~~f~~~~~~~~~~~~~~-~~~~~i~~~~~~~W~~l~~~~k~~y~~~a~~~~~~y~~~~ 65 (66)
T cd00084 1 PKRPLSAYFLFSQEHRAEVKAENPG-LSVGEISKILGEMWKSLSEEEKKKYEEKAEKDKERYEKEM 65 (66)
T ss_pred CCCCCcHHHHHHHHHHHHHHHHCcC-CCHHHHHHHHHHHHHhCCHHHHHHHHHHHHHHHHHHHHhh
Confidence 7999999999999999999999999 9999999999999999999999999999999999998875
No 10
>KOG0527 consensus HMG-box transcription factor [Transcription]
Probab=99.72 E-value=7.7e-18 Score=139.76 Aligned_cols=85 Identities=29% Similarity=0.539 Sum_probs=79.4
Q ss_pred ccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHhh
Q 031973 32 KAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQ 111 (150)
Q Consensus 32 k~~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~~ 111 (150)
.......++||||||||||.+..|.+|..+||. +.+.||+++||.+|+.|++++|.+|.+.|++++..|++++++|+.+
T Consensus 55 ~~k~~~~hIKRPMNAFMVWSq~~RRkma~qnP~-mHNSEISK~LG~~WK~Lse~EKrPFi~EAeRLR~~HmkehPdYKYR 133 (331)
T KOG0527|consen 55 KDKTSTDRIKRPMNAFMVWSQGQRRKLAKQNPK-MHNSEISKRLGAEWKLLSEEEKRPFVDEAERLRAQHMKEYPDYKYR 133 (331)
T ss_pred cCCCCccccCCCcchhhhhhHHHHHHHHHhCcc-hhhHHHHHHHHHHHhhcCHhhhccHHHHHHHHHHHHHHhCCCcccc
Confidence 345667899999999999999999999999999 9999999999999999999999999999999999999999999998
Q ss_pred ccCCcc
Q 031973 112 LADGVN 117 (150)
Q Consensus 112 ~~~~~~ 117 (150)
......
T Consensus 134 PRRKkk 139 (331)
T KOG0527|consen 134 PRRKKK 139 (331)
T ss_pred cccccc
Confidence 776554
No 11
>KOG0381 consensus HMG box-containing protein [General function prediction only]
Probab=99.71 E-value=1.6e-16 Score=109.69 Aligned_cols=76 Identities=47% Similarity=0.840 Sum_probs=72.0
Q ss_pred CC--CCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHH-HHHhhc
Q 031973 36 DP--NKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQ-DYNKQL 112 (150)
Q Consensus 36 dp--~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~-~y~~~~ 112 (150)
+| +.|++|++||++|+.+.+..++.+||+ ++..+|+++||++|++|++++|.+|...|..++.+|...|. .|+...
T Consensus 17 ~p~~~~pkrp~sa~~~f~~~~~~~~k~~~p~-~~~~~v~k~~g~~W~~l~~~~k~~y~~ka~~~k~~Y~~~~~~~~~~~~ 95 (96)
T KOG0381|consen 17 DPNAQAPKRPLSAFFLFSSEQRSKIKAENPG-LSVGEVAKALGEMWKNLAEEEKQPYEEKASKLKEKYEKELAGEYKASL 95 (96)
T ss_pred CCCCCCCCCCCcHHHHHHHHHHHHHHHhCCC-CCHHHHHHHHHHHHhcCCHHHHHHHHHHHHHHHHHHHHHHHHHHhhcc
Confidence 55 599999999999999999999999999 99999999999999999999999999999999999999999 887653
No 12
>KOG0526 consensus Nucleosome-binding factor SPN, POB3 subunit [Transcription; Replication, recombination and repair; Chromatin structure and dynamics]
Probab=99.59 E-value=2.4e-15 Score=129.68 Aligned_cols=81 Identities=40% Similarity=0.678 Sum_probs=74.8
Q ss_pred CccccCCccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHH
Q 031973 25 GKRTAKPKAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKN 104 (150)
Q Consensus 25 ~k~kk~kk~~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~ 104 (150)
.+++++.++.+||+.||||++|||||.+..|..|+.. + .+.++|++.+|.+|+.|+. |.+|.+.|+.++++|+.+
T Consensus 521 ~~~~k~~kk~kdpnapkra~sa~m~w~~~~r~~ik~d--g-i~~~dv~kk~g~~wk~ms~--k~~we~ka~~dk~ry~~e 595 (615)
T KOG0526|consen 521 KEKKKKGKKKKDPNAPKRATSAYMLWLNASRESIKED--G-ISVGDVAKKAGEKWKQMSA--KEEWEDKAAVDKQRYEDE 595 (615)
T ss_pred hccccCcccCCCCCCCccchhHHHHHHHhhhhhHhhc--C-chHHHHHHHHhHHHhhhcc--cchhhHHHHHHHHHHHHH
Confidence 3444677789999999999999999999999999987 5 8999999999999999998 899999999999999999
Q ss_pred HHHHHh
Q 031973 105 MQDYNK 110 (150)
Q Consensus 105 ~~~y~~ 110 (150)
|.+|+.
T Consensus 596 m~~yk~ 601 (615)
T KOG0526|consen 596 MKEYKN 601 (615)
T ss_pred HHhhcC
Confidence 999993
No 13
>KOG3248 consensus Transcription factor TCF-4 [Transcription]
Probab=99.45 E-value=9.2e-14 Score=114.48 Aligned_cols=79 Identities=23% Similarity=0.406 Sum_probs=74.6
Q ss_pred CCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHhhccCCcc
Q 031973 38 NKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQLADGVN 117 (150)
Q Consensus 38 ~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~~~~~~~~ 117 (150)
.+.|+|+|||||||++.|..|..++.- ....+|.++||++|..|+.++.++|.++|.++++.|.+.++.|.+.......
T Consensus 190 phiKKPLNAFmlyMKEmRa~vvaEctl-KeSAaiNqiLGrRWH~LSrEEQAKYyElArKerqlH~qlYP~WSARdNYgKK 268 (421)
T KOG3248|consen 190 PHIKKPLNAFMLYMKEMRAKVVAECTL-KESAAINQILGRRWHALSREEQAKYYELARKERQLHMQLYPGWSARDNYGKK 268 (421)
T ss_pred ccccccHHHHHHHHHHHHHHHHHHhhh-hhHHHHHHHHhHHHhhhhHHHHHHHHHHHHHHHHHHHHhcCCcchhhhhhhh
Confidence 488999999999999999999999986 6889999999999999999999999999999999999999999999988754
No 14
>KOG0528 consensus HMG-box transcription factor SOX5 [Transcription]
Probab=99.21 E-value=7.6e-12 Score=107.10 Aligned_cols=85 Identities=26% Similarity=0.453 Sum_probs=77.0
Q ss_pred cCCccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHH
Q 031973 29 AKPKAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDY 108 (150)
Q Consensus 29 k~kk~~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y 108 (150)
.++-..+.+.++||||||||+|.++-|..|.+.+|+ +.+..|+++||.+|+.|+..+|++|.+.-..+...|.+.+++|
T Consensus 315 srrg~~ss~PHIKRPMNAFMVWAkDERRKILqA~PD-MHNSnISKILGSRWKaMSN~eKQPYYEEQaRLSk~HlEk~PdY 393 (511)
T KOG0528|consen 315 SRRGRASSEPHIKRPMNAFMVWAKDERRKILQAFPD-MHNSNISKILGSRWKAMSNTEKQPYYEEQARLSKLHLEKYPDY 393 (511)
T ss_pred cccCcCCCCccccCCcchhhcccchhhhhhhhcCcc-ccccchhHHhcccccccccccccchHHHHHHHHHhhhccCccc
Confidence 335556677899999999999999999999999999 9999999999999999999999999998888888999999999
Q ss_pred HhhccC
Q 031973 109 NKQLAD 114 (150)
Q Consensus 109 ~~~~~~ 114 (150)
+.+...
T Consensus 394 rYkPRP 399 (511)
T KOG0528|consen 394 RYKPRP 399 (511)
T ss_pred ccCCCC
Confidence 987643
No 15
>KOG4715 consensus SWI/SNF-related matrix-associated actin-dependent regulator of chromatin [Chromatin structure and dynamics]
Probab=99.20 E-value=7.6e-11 Score=96.73 Aligned_cols=80 Identities=24% Similarity=0.548 Sum_probs=74.6
Q ss_pred ccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHhh
Q 031973 32 KAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQ 111 (150)
Q Consensus 32 k~~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~~ 111 (150)
...+.|.+|-+|+-+||.|++..|++|+..||. +...+|.++||.+|..|++++|+.|...+...+..|.+.|..|...
T Consensus 57 t~pkpPkppekpl~pymrySrkvWd~VkA~nPe-~kLWeiGK~Ig~mW~dLpd~EK~ey~~EYeaEKieY~~smkayh~s 135 (410)
T KOG4715|consen 57 TRPKPPKPPEKPLMPYMRYSRKVWDQVKASNPE-LKLWEIGKIIGGMWLDLPDEEKQEYLNEYEAEKIEYNESMKAYHNS 135 (410)
T ss_pred cCCCCCCCCCcccchhhHHhhhhhhhhhccCcc-hHHHHHHHHHHHHHhhCcchHHHHHHHHHHHHHHHHHHHHHHhhCC
Confidence 345567888999999999999999999999999 9999999999999999999999999999999999999999998875
Q ss_pred c
Q 031973 112 L 112 (150)
Q Consensus 112 ~ 112 (150)
.
T Consensus 136 p 136 (410)
T KOG4715|consen 136 P 136 (410)
T ss_pred c
Confidence 4
No 16
>KOG2746 consensus HMG-box transcription factor Capicua and related proteins [Transcription]
Probab=98.68 E-value=1.9e-08 Score=89.28 Aligned_cols=74 Identities=27% Similarity=0.478 Sum_probs=69.0
Q ss_pred CCccCCCCCCCCCCCChHHHHHHHHH--HHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHH
Q 031973 30 KPKAAKDPNKPKRPPSAFFVFMEEFR--KQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKN 104 (150)
Q Consensus 30 ~kk~~~dp~~PKrP~~aY~lF~~~~r--~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~ 104 (150)
+-...++..+.++|||+|+||++.+| ..+.+.||+ ..+.-|+++||++|-.|.+.+|+.|+++|.+.++.|.++
T Consensus 172 rspnkr~k~HirrPMnaf~ifskrhr~~g~vhq~~pn-~DNrtIskiLgewWytL~~~Ekq~yhdLa~Qvk~Ahfka 247 (683)
T KOG2746|consen 172 RSPNKRDKDHIRRPMNAFHIFSKRHRGEGRVHQRHPN-QDNRTISKILGEWWYTLGPNEKQKYHDLAFQVKEAHFKA 247 (683)
T ss_pred CCCCcCcchhhhhhhHHHHHHHhhcCCccchhccCcc-ccchhHHHHHhhhHhhhCchhhhhHHHHHHHHHHHHhhh
Confidence 33556778899999999999999999 899999999 999999999999999999999999999999999999986
No 17
>PF14887 HMG_box_5: HMG (high mobility group) box 5; PDB: 1L8Y_A 1L8Z_A 2HDZ_A.
Probab=98.04 E-value=3.2e-05 Score=51.67 Aligned_cols=75 Identities=19% Similarity=0.387 Sum_probs=61.4
Q ss_pred CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHhhccCC
Q 031973 39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQLADG 115 (150)
Q Consensus 39 ~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~~~~~~ 115 (150)
.|..|-+|--||.+.....+...++. ....+ .+.+...|++|++.+|.+|...|.++..+|+..|.+|+...+..
T Consensus 3 lPE~PKt~qe~Wqq~vi~dYla~~~~-dr~K~-~kam~~~W~~me~Kekl~WIkKA~EdqKrYE~el~e~r~~~~~~ 77 (85)
T PF14887_consen 3 LPETPKTAQEIWQQSVIGDYLAKFRN-DRKKA-LKAMEAQWSQMEKKEKLKWIKKAAEDQKRYERELREMRSAPADA 77 (85)
T ss_dssp -S----THHHHHHHHHHHHHHHHTTS-THHHH-HHHHHHHHHTTGGGHHHHHHHHHHHHHHHHHHHHHCCS-CCCTT
T ss_pred CCCCCCCHHHHHHHHHHHHHHHHhhH-hHHHH-HHHHHHHHHHhhhhhhhHHHHHHHHHHHHHHHHHHHHhcCCCCC
Confidence 57788999999999999999999987 54444 56899999999999999999999999999999999999877643
No 18
>PF06382 DUF1074: Protein of unknown function (DUF1074); InterPro: IPR024460 This family consists of several proteins which appear to be specific to Insecta. The function of this family is unknown.
Probab=97.43 E-value=0.00083 Score=51.49 Aligned_cols=50 Identities=28% Similarity=0.449 Sum_probs=43.2
Q ss_pred CChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHH
Q 031973 44 PSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRK 98 (150)
Q Consensus 44 ~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k 98 (150)
-++|+-|+.+.+. .|.+ +...++....+..|..|++.+|..|..++....
T Consensus 83 nnaYLNFLReFRr----kh~~-L~p~dlI~~AAraW~rLSe~eK~rYrr~~~~~~ 132 (183)
T PF06382_consen 83 NNAYLNFLREFRR----KHCG-LSPQDLIQRAARAWCRLSEAEKNRYRRMAPSVR 132 (183)
T ss_pred chHHHHHHHHHHH----HccC-CCHHHHHHHHHHHHHhCCHHHHHHHHhhcchhh
Confidence 3789999988865 6677 999999999999999999999999998766543
No 19
>PF04690 YABBY: YABBY protein; InterPro: IPR006780 YABBY proteins are a group of plant-specific transcription factors involved in the specification of abaxial polarity in lateral organs such as leaves and floral organs [, ].
Probab=97.19 E-value=0.00095 Score=51.04 Aligned_cols=49 Identities=33% Similarity=0.498 Sum_probs=43.2
Q ss_pred CCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCC
Q 031973 34 AKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMS 83 (150)
Q Consensus 34 ~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~ 83 (150)
.+.|.+-.|-++||..|+++.-..|+..+|+ +++.|+....+..|...|
T Consensus 116 ~kPPEKRqR~psaYn~f~k~ei~rik~~~p~-ishkeaFs~aAknW~h~p 164 (170)
T PF04690_consen 116 NKPPEKRQRVPSAYNRFMKEEIQRIKAENPD-ISHKEAFSAAAKNWAHFP 164 (170)
T ss_pred cCCccccCCCchhHHHHHHHHHHHHHhcCCC-CCHHHHHHHHHHhhhhCc
Confidence 3445555677899999999999999999999 999999999999998876
No 20
>COG5648 NHP6B Chromatin-associated proteins containing the HMG domain [Chromatin structure and dynamics]
Probab=96.94 E-value=0.00069 Score=53.18 Aligned_cols=68 Identities=19% Similarity=0.393 Sum_probs=62.2
Q ss_pred CCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHH
Q 031973 38 NKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQ 106 (150)
Q Consensus 38 ~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~ 106 (150)
.+|..|..+|+-+...+|+.+...+|+ ....+++++++..|.+|++.-+.+|.+.+..++..|...|+
T Consensus 142 ~~~~~~~~~~~e~~~~~r~~~~~~~~~-~~~~e~~k~~~~~w~el~~skK~~~~~~~Kk~k~~~~~~~~ 209 (211)
T COG5648 142 LPNKAPIGPFIENEPKIRPKVEGPSPD-KALVEETKIISKAWSELDESKKKKYIDKYKKLKEEYDSFYP 209 (211)
T ss_pred cCCCCCCchhhhccHHhccccCCCCcc-hhhhHHhhhhhhhhhhhChhhhhHHHHHHHHHHHHHhhhcc
Confidence 467888889999999999999999998 88999999999999999999999999999999999887664
No 21
>PF08073 CHDNT: CHDNT (NUC034) domain; InterPro: IPR012958 The CHD N-terminal domain is found in PHD/RING fingers and chromo domain-associated helicases [].; GO: 0003677 DNA binding, 0005524 ATP binding, 0008270 zinc ion binding, 0016818 hydrolase activity, acting on acid anhydrides, in phosphorus-containing anhydrides, 0006355 regulation of transcription, DNA-dependent, 0005634 nucleus
Probab=95.69 E-value=0.014 Score=36.62 Aligned_cols=39 Identities=18% Similarity=0.430 Sum_probs=35.4
Q ss_pred ChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCCh
Q 031973 45 SAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSE 84 (150)
Q Consensus 45 ~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~ 84 (150)
+-|-+|.+-.|+.|...||+ +....|...++..|++.+.
T Consensus 14 t~yK~Fsq~vRP~l~~~NPk-~~~sKl~~l~~AKwrEF~~ 52 (55)
T PF08073_consen 14 TNYKAFSQHVRPLLAKANPK-APMSKLMMLLQAKWREFQE 52 (55)
T ss_pred HHHHHHHHHHHHHHHHHCCC-CcHHHHHHHHHHHHHHHHh
Confidence 56889999999999999999 9999999999999987653
No 22
>PF06244 DUF1014: Protein of unknown function (DUF1014); InterPro: IPR010422 This family consists of several hypothetical eukaryotic proteins of unknown function.
Probab=94.05 E-value=0.084 Score=38.34 Aligned_cols=49 Identities=20% Similarity=0.353 Sum_probs=42.9
Q ss_pred CCCCCCCCC-ChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChh
Q 031973 36 DPNKPKRPP-SAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSED 85 (150)
Q Consensus 36 dp~~PKrP~-~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~e 85 (150)
...||-|.+ -||.-|...+.+.|+.++|+ +..+++-.+|-.+|...|++
T Consensus 68 ~drHPErR~KAAy~afeE~~Lp~lK~E~Pg-LrlsQ~kq~l~K~w~KSPeN 117 (122)
T PF06244_consen 68 IDRHPERRMKAAYKAFEERRLPELKEENPG-LRLSQYKQMLWKEWQKSPEN 117 (122)
T ss_pred CCCCcchhHHHHHHHHHHHHhHHHHhhCCC-chHHHHHHHHHHHHhcCCCC
Confidence 345675555 78999999999999999999 99999999999999988865
No 23
>PF04769 MAT_Alpha1: Mating-type protein MAT alpha 1; InterPro: IPR006856 This family includes Saccharomyces cerevisiae (Baker's yeast) mating type protein alpha 1 (P01365 from SWISSPROT). MAT alpha 1 is a transcription activator that activates mating-type alpha-specific genes with the help of the MADS-box containing MCM1 transcription factor, which together bind cooperatively to PQ elements upstream of alpha-specific genes. The MCM1-MATalpha1 complex is required for the proper DNA-bending that is needed for transcriptional activation []. Alpha 1 interacts in vivo with STE12, linking expression of alpha-specific genes to the alpha-pheromone (IPR006742 from INTERPRO) response pathway [].; GO: 0000772 mating pheromone activity, 0003677 DNA binding, 0045895 positive regulation of transcription, mating-type specific, 0005634 nucleus
Probab=93.45 E-value=0.2 Score=39.28 Aligned_cols=56 Identities=20% Similarity=0.341 Sum_probs=39.7
Q ss_pred CCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHH
Q 031973 34 AKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEK 96 (150)
Q Consensus 34 ~~dp~~PKrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~ 96 (150)
......++||+|+||+|+.-.- ...|+ ..-.+++..|+.+|..=+- |..|.-.|+.
T Consensus 38 ~~~~~~~kr~lN~Fm~FRsyy~----~~~~~-~~Qk~~S~~l~~lW~~dp~--k~~W~l~ak~ 93 (201)
T PF04769_consen 38 KRSPEKAKRPLNGFMAFRSYYS----PIFPP-LPQKELSGILTKLWEKDPF--KNKWSLMAKA 93 (201)
T ss_pred cccccccccchhHHHHHHHHHH----hhcCC-cCHHHHHHHHHHHHhCCcc--HhHHHHHhhh
Confidence 3445678999999999986654 34454 5668999999999997543 4446555543
No 24
>KOG3223 consensus Uncharacterized conserved protein [Function unknown]
Probab=89.70 E-value=0.23 Score=38.82 Aligned_cols=55 Identities=25% Similarity=0.477 Sum_probs=45.6
Q ss_pred CCCCC-CCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHH
Q 031973 36 DPNKP-KRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERA 94 (150)
Q Consensus 36 dp~~P-KrP~~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A 94 (150)
|..|| +|=.-||.-|-....+.|+.+||+ +..+++-.+|-.+|...|++ ||.+.+
T Consensus 160 ddrHPEkRmrAA~~afEe~~LPrLK~e~P~-lrlsQ~Kqll~Kew~KsPDN---P~Nq~~ 215 (221)
T KOG3223|consen 160 DDRHPEKRMRAAFKAFEEARLPRLKKENPG-LRLSQYKQLLKKEWQKSPDN---PFNQAA 215 (221)
T ss_pred cccChHHHHHHHHHHHHHhhchhhhhcCCC-ccHHHHHHHHHHHHhhCCCC---hhhHHh
Confidence 44677 444567999999999999999999 99999999999999999976 555543
No 25
>TIGR03481 HpnM hopanoid biosynthesis associated membrane protein HpnM. The genomes containing members of this family share the machinery for the biosynthesis of hopanoid lipids. Furthermore, the genes of this family are usually located proximal to other components of this biological process. The proteins are members of the pfam05494 family of putative transporters known as "toluene tolerance protein Ttg2D", although it is unlikely that the members included here have anything to do with toluene per-se.
Probab=84.45 E-value=2.4 Score=32.99 Aligned_cols=47 Identities=15% Similarity=0.442 Sum_probs=39.6
Q ss_pred CCHHHHHH-HHHHhhccCChhhhhhHHHHHHH-HHHHHHHHHHHHHhhc
Q 031973 66 KSVATVGK-AAGEKWKSMSEDEKAPFVERAEK-RKSDYNKNMQDYNKQL 112 (150)
Q Consensus 66 ~~~~eisk-~l~~~Wk~l~~eeK~~y~~~A~~-~k~~y~~~~~~y~~~~ 112 (150)
.++..|++ .||..|+.+++++|+.|...... ....|-..+..|....
T Consensus 64 ~Df~~mar~vLG~~W~~~s~~Qr~~F~~~F~~~l~~tY~~~l~~y~~~~ 112 (198)
T TIGR03481 64 FDLPAMARLTLGSSWTSLSPEQRRRFIGAFRELSIATYASQFKSYAGER 112 (198)
T ss_pred CCHHHHHHHHhhhhhhhCCHHHHHHHHHHHHHHHHHHHHHHHHhhcCce
Confidence 56778876 58999999999999999998877 6788888898887653
No 26
>PRK15117 ABC transporter periplasmic binding protein MlaC; Provisional
Probab=79.16 E-value=6.1 Score=30.98 Aligned_cols=49 Identities=18% Similarity=0.354 Sum_probs=39.3
Q ss_pred CCCCCHHHHHH-HHHHhhccCChhhhhhHHHHHHHH-HHHHHHHHHHHHhhc
Q 031973 63 PNNKSVATVGK-AAGEKWKSMSEDEKAPFVERAEKR-KSDYNKNMQDYNKQL 112 (150)
Q Consensus 63 p~~~~~~eisk-~l~~~Wk~l~~eeK~~y~~~A~~~-k~~y~~~~~~y~~~~ 112 (150)
|. .++..|++ .||..|+.+++++|+.|...-... ..-|...+..|..+.
T Consensus 66 p~-~Df~~~s~~vLG~~wr~as~eQr~~F~~~F~~~Lv~tYa~~l~~y~~q~ 116 (211)
T PRK15117 66 PY-VQVKYAGALVLGRYYKDATPAQREAYFAAFREYLKQAYGQALAMYHGQT 116 (211)
T ss_pred cc-CCHHHHHHHHhhhhhhhCCHHHHHHHHHHHHHHHHHHHHHHHHHhCCce
Confidence 44 67777766 589999999999999999866655 568889999997653
No 27
>PF05494 Tol_Tol_Ttg2: Toluene tolerance, Ttg2 ; InterPro: IPR008869 Toluene tolerance is mediated by increased cell membrane rigidity resulting from changes in fatty acid and phospholipid compositions, exclusion of toluene from the cell membrane, and removal of intracellular toluene by degradation []. Many proteins are involved in these processes. This family is a transporter which shows similarity to ABC transporters [].; PDB: 2QGU_A.
Probab=75.86 E-value=6 Score=29.49 Aligned_cols=45 Identities=20% Similarity=0.451 Sum_probs=34.0
Q ss_pred CCHHHHHHH-HHHhhccCChhhhhhHHHHHHHH-HHHHHHHHHHHHh
Q 031973 66 KSVATVGKA-AGEKWKSMSEDEKAPFVERAEKR-KSDYNKNMQDYNK 110 (150)
Q Consensus 66 ~~~~eisk~-l~~~Wk~l~~eeK~~y~~~A~~~-k~~y~~~~~~y~~ 110 (150)
.++..|++. ||..|+.+++++++.|....... ...|...+..|..
T Consensus 38 ~D~~~~ar~~LG~~w~~~s~~q~~~F~~~f~~~l~~~Y~~~l~~y~~ 84 (170)
T PF05494_consen 38 FDFERMARRVLGRYWRKASPAQRQRFVEAFKQLLVRTYAKRLDEYSG 84 (170)
T ss_dssp B-HHHHHHHHHGGGTTTS-HHHHHHHHHHHHHHHHHHHHHHHHT-SS
T ss_pred CCHHHHHHHHHHHhHhhCCHHHHHHHHHHHHHHHHHHHHHHHHhhCC
Confidence 677777765 77899999999999999876665 5678888888875
No 28
>PF13875 DUF4202: Domain of unknown function (DUF4202)
Probab=60.30 E-value=14 Score=28.70 Aligned_cols=39 Identities=26% Similarity=0.485 Sum_probs=32.8
Q ss_pred hHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCChhhhh
Q 031973 46 AFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKA 88 (150)
Q Consensus 46 aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~ 88 (150)
+-++|...+...+...|. ...+..+|...|+.||+..++
T Consensus 131 acLVFL~~~f~~F~~~~d----eeK~v~Il~KTw~KMS~~g~~ 169 (185)
T PF13875_consen 131 ACLVFLEYYFEDFAAKHD----EEKIVDILRKTWRKMSERGHE 169 (185)
T ss_pred HHHHhHHHHHHHHHhcCC----HHHHHHHHHHHHHHCCHHHHH
Confidence 467899999999988774 368889999999999998875
No 29
>PF11304 DUF3106: Protein of unknown function (DUF3106); InterPro: IPR021455 Some members in this family of proteins are annotated as transmembrane proteins however this cannot be confirmed. Currently no function is known.
Probab=52.60 E-value=58 Score=22.76 Aligned_cols=22 Identities=18% Similarity=0.582 Sum_probs=9.7
Q ss_pred HHHHHhhccCChhhhhhHHHHH
Q 031973 73 KAAGEKWKSMSEDEKAPFVERA 94 (150)
Q Consensus 73 k~l~~~Wk~l~~eeK~~y~~~A 94 (150)
.-+...|+.|++..+..+...|
T Consensus 14 ~pl~~~W~~l~~~qr~k~l~~a 35 (107)
T PF11304_consen 14 APLAERWNSLPPEQRRKWLQIA 35 (107)
T ss_pred HHHHHHHhcCCHHHHHHHHHHH
Confidence 3344444444444444444333
No 30
>COG2854 Ttg2D ABC-type transport system involved in resistance to organic solvents, auxiliary component [Secondary metabolites biosynthesis, transport, and catabolism]
Probab=48.43 E-value=22 Score=27.99 Aligned_cols=43 Identities=19% Similarity=0.352 Sum_probs=35.6
Q ss_pred HHHHHhhccCChhhhhhHHHHHHHH-HHHHHHHHHHHHhhccCC
Q 031973 73 KAAGEKWKSMSEDEKAPFVERAEKR-KSDYNKNMQDYNKQLADG 115 (150)
Q Consensus 73 k~l~~~Wk~l~~eeK~~y~~~A~~~-k~~y~~~~~~y~~~~~~~ 115 (150)
..||.-|+.+++++++.|....... ...|-..+..|+.+...-
T Consensus 78 ~vLGk~~k~aspeQ~~~F~~aF~~yl~q~Y~~aL~~Y~~q~~~v 121 (202)
T COG2854 78 LVLGKYYKTASPEQRQAFFKAFRTYLEQTYGQALLDYKGQTLKV 121 (202)
T ss_pred HHhccccccCCHHHHHHHHHHHHHHHHHHHHHHHHHccCCCcee
Confidence 3488999999999999999876665 677999999999877543
No 31
>PF01352 KRAB: KRAB box; InterPro: IPR001909 The Krueppel-associated box (KRAB) is a domain of around 75 amino acids that is found in the N-terminal part of about one third of eukaryotic Krueppel-type C2H2 zinc finger proteins (ZFPs) []. It is enriched in charged amino acids and can be divided into subregions A and B, which are predicted to fold into two amphipathic alpha-helices. The KRAB A and B boxes can be separated by variable spacer segments and many KRAB proteins contain only the A box []. The functions currently known for members of the KRAB-containing protein family include transcriptional repression of RNA polymerase I, II, and III promoters, binding and splicing of RNA, and control of nucleolus function. The KRAB domain functions as a transcriptional repressor when tethered to the template DNA by a DNA-binding domain. A sequence of 45 amino acids in the KRAB A subdomain has been shown to be necessary and sufficient for transcriptional repression. The B box does not repress by itself but does potentiate the repression exerted by the KRAB A subdomain [, ]. Gene silencing requires the binding of the KRAB domain to the RING-B box-coiled coil (RBCC) domain of the KAP-1/TIF1-beta corepressor. As KAP-1 binds to the heterochromatin proteins HP1, it has been proposed that the KRAB-ZFP-bound target gene could be silenced following recruitment to heterochromatin [, ]. KRAB-ZFPs probably constitute the single largest class of transcription factors within the human genome []. Although the function of KRAB-ZFPs is largely unknown, they appear to play important roles during cell differentiation and development. The KRAB domain is generally encoded by two exons. The regions coded by the two exons are known as KRAB-A and KRAB-B.; GO: 0003676 nucleic acid binding, 0006355 regulation of transcription, DNA-dependent, 0005622 intracellular; PDB: 1V65_A.
Probab=37.81 E-value=23 Score=20.56 Aligned_cols=26 Identities=15% Similarity=0.311 Sum_probs=14.9
Q ss_pred HHHHHHHH-HhhccCChhhhhhHHHHH
Q 031973 69 ATVGKAAG-EKWKSMSEDEKAPFVERA 94 (150)
Q Consensus 69 ~eisk~l~-~~Wk~l~~eeK~~y~~~A 94 (150)
.+|+--++ +.|..|.+.+|..|.+.-
T Consensus 4 ~Dvav~fs~eEW~~L~~~Qk~ly~dvm 30 (41)
T PF01352_consen 4 EDVAVYFSQEEWELLDPAQKNLYRDVM 30 (41)
T ss_dssp ---TT---HHHHHTS-HHHHHHHHHHH
T ss_pred EEEEEEcChhhcccccceecccchhHH
Confidence 34444444 669999999998887644
No 32
>PRK09706 transcriptional repressor DicA; Reviewed
Probab=36.99 E-value=84 Score=22.34 Aligned_cols=43 Identities=19% Similarity=0.233 Sum_probs=37.2
Q ss_pred HHHHHHHhhccCChhhhhhHHHHHHHHHHHHHHHHHHHHhhcc
Q 031973 71 VGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQLA 113 (150)
Q Consensus 71 isk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~~~y~~~~~ 113 (150)
-...|-..|+.|+++++.............|...+.+|-....
T Consensus 88 ~~~~ll~~~~~L~~~~~~~~l~~l~~~~~~~~~~~~~~~~~~~ 130 (135)
T PRK09706 88 DQKELLELFDALPESEQDAQLSEMRARVENFNKLFEELLKARK 130 (135)
T ss_pred HHHHHHHHHHHCCHHHHHHHHHHHHHHHHHHHHHHHHHHHHHh
Confidence 3467889999999999999999999999999999988876543
No 33
>PF15076 DUF4543: Domain of unknown function (DUF4543)
Probab=33.83 E-value=20 Score=23.34 Aligned_cols=22 Identities=14% Similarity=0.546 Sum_probs=17.8
Q ss_pred cCCCCCCCCCCCChHHHHHHHH
Q 031973 33 AAKDPNKPKRPPSAFFVFMEEF 54 (150)
Q Consensus 33 ~~~dp~~PKrP~~aY~lF~~~~ 54 (150)
+...|++|.-||.-||+|++.-
T Consensus 25 r~~K~GfpdepmrE~ml~l~~L 46 (75)
T PF15076_consen 25 RPRKPGFPDEPMREYMLHLQAL 46 (75)
T ss_pred CCCCCCCCcchHHHHHHHHHHH
Confidence 3556899999999999998643
No 34
>PF12881 NUT_N: NUT protein N terminus; InterPro: IPR024309 This domain is found in the N-terminal region of Nuclear Testis (NUT) proteins. It is also found in FAM22, which are a family of uncharacterised mammalian proteins.
Probab=32.87 E-value=1.3e+02 Score=25.45 Aligned_cols=54 Identities=24% Similarity=0.205 Sum_probs=38.2
Q ss_pred HHHhCCCCCCHHHHHHHHHHhhccCChhhhhhHHHHHHHHHHH-HHHHHHHHHhhc
Q 031973 58 FKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSD-YNKNMQDYNKQL 112 (150)
Q Consensus 58 ~k~~~p~~~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~-y~~~~~~y~~~~ 112 (150)
|....|. ++..|-..+.-+.|...+.-+|..|.++|.+-.+= -+++|+.-+.++
T Consensus 243 Lar~kPt-MtlEeGl~ra~qEW~~~SnfdRmifyemaekFmEFEaeEEmq~q~lq~ 297 (328)
T PF12881_consen 243 LARLKPT-MTLEEGLWRAVQEWQHTSNFDRMIFYEMAEKFMEFEAEEEMQIQKLQL 297 (328)
T ss_pred HHhcCCC-ccHHHHHHHHHHHhhccccccHHHHHHHHHHHccCCcHHHHHHHHHHH
Confidence 4445565 77778777788999999999999999999887531 124555544443
No 35
>PRK10363 cpxP periplasmic repressor CpxP; Reviewed
Probab=32.26 E-value=1.2e+02 Score=23.27 Aligned_cols=36 Identities=11% Similarity=0.304 Sum_probs=27.5
Q ss_pred HHHHHHHHHHhhccCChhhhhhHHHHHHHHHHHHHH
Q 031973 68 VATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNK 103 (150)
Q Consensus 68 ~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~ 103 (150)
..++.+.-.++++-|+|++|..|.....+-...+..
T Consensus 110 ~Vem~k~~nqmy~lLTPEQKaq~~~~~~~rm~~~~~ 145 (166)
T PRK10363 110 QVEMAKVRNQMYRLLTPEQQAVLNEKHQQRMEQLRD 145 (166)
T ss_pred HHHHHHHHHHHHHhCCHHHHHHHHHHHHHHHHHHHH
Confidence 345666677899999999999998877766655543
No 36
>PRK12750 cpxP periplasmic repressor CpxP; Reviewed
Probab=30.93 E-value=1.3e+02 Score=22.82 Aligned_cols=33 Identities=18% Similarity=0.320 Sum_probs=25.6
Q ss_pred HHHHHhhccCChhhhhhHHHHHHHHHHHHHHHH
Q 031973 73 KAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNM 105 (150)
Q Consensus 73 k~l~~~Wk~l~~eeK~~y~~~A~~~k~~y~~~~ 105 (150)
+...+++..|++++|..|.+....-...+...+
T Consensus 128 ~~~~~~~~vLTpEQRak~~e~~~~r~~~~~~~~ 160 (170)
T PRK12750 128 EKRHQMLSILTPEQKAKFQELQQERMQECQDKM 160 (170)
T ss_pred HHHHHHHHhCCHHHHHHHHHHHHHHHHHHHHHH
Confidence 335568999999999999988777766666655
No 37
>PRK12751 cpxP periplasmic stress adaptor protein CpxP; Reviewed
Probab=30.12 E-value=1.2e+02 Score=22.97 Aligned_cols=31 Identities=10% Similarity=0.314 Sum_probs=22.3
Q ss_pred HHHHHHHhhccCChhhhhhHHHHHHHHHHHH
Q 031973 71 VGKAAGEKWKSMSEDEKAPFVERAEKRKSDY 101 (150)
Q Consensus 71 isk~l~~~Wk~l~~eeK~~y~~~A~~~k~~y 101 (150)
+.+...++++.|++++|..|.+...+-....
T Consensus 119 ~~~~~~qmy~lLTPEQra~l~~~~e~r~~~~ 149 (162)
T PRK12751 119 MAKVRNQMYNLLTPEQKEALNKKHQERIEKL 149 (162)
T ss_pred HHHHHHHHHHcCCHHHHHHHHHHHHHHHHHH
Confidence 3444567889999999999988665554433
No 38
>PF00887 ACBP: Acyl CoA binding protein; InterPro: IPR000582 Acyl-CoA-binding protein (ACBP) is a small (10 Kd) protein that binds medium- and long-chain acyl-CoA esters with very high affinity and may function as an intracellular carrier of acyl-CoA esters []. ACBP is also known as diazepam binding inhibitor (DBI) or endozepine (EP) because of its ability to displace diazepam from the benzodiazepine (BZD) recognition site located on the GABA type A receptor. It is therefore possible that this protein also acts as a neuropeptide to modulate the action of the GABA receptor []. ACBP is a highly conserved protein of about 90 residues that is found in all four eukaryotic kingdoms, Animalia, Plantae, Fungi and Protista, and in some eubacterial species []. Although ACBP occurs as a completely independent protein, intact ACB domains have been identified in a number of large, multifunctional proteins in a variety of eukaryotic species. These include large membrane-associated proteins with N-terminal ACB domains, multifunctional enzymes with both ACB and peroxisomal enoyl-CoA Delta(3), Delta(2)-enoyl-CoA isomerase domains, and proteins with both an ACB domain and ankyrin repeats (IPR002110 from INTERPRO) []. The ACB domain consists of four alpha-helices arranged in a bowl shape with a highly exposed acyl-CoA-binding site. The ligand is bound through specific interactions with residues on the protein, most notably several conserved positive charges that interact with the phosphate group on the adenosine-3'phosphate moiety, and the acyl chain is sandwiched between the hydrophobic surfaces of CoA and the protein []. Other proteins containing an ACB domain include: Endozepine-like peptide (ELP) (gene DBIL5) from mouse []. ELP is a testis-specific ACBP homologue that may be involved in the energy metabolism of the mature sperm. MA-DBI, a transmembrane protein of unknown function which has been found in mammals. MA-DBI contains a N-terminal ACB domain. DRS-1 [], a human protein of unknown function that contains a N-terminal ACB domain and a C-terminal enoyl-CoA isomerase/hydratase domain. ; GO: 0000062 fatty-acyl-CoA binding; PDB: 2CB8_A 2FJ9_A 2LBB_A 1ST7_A 3EPY_B 2FDQ_C 1NTI_A 1HB8_A 1ACA_A 1NVL_A ....
Probab=24.00 E-value=2.2e+02 Score=18.66 Aligned_cols=54 Identities=11% Similarity=0.261 Sum_probs=28.9
Q ss_pred ChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHhhccCCh----hhhhhHHHHHHHHHHH
Q 031973 45 SAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSE----DEKAPFVERAEKRKSD 100 (150)
Q Consensus 45 ~aY~lF~~~~r~~~k~~~p~~~~~~eisk~l~~~Wk~l~~----eeK~~y~~~A~~~k~~ 100 (150)
.-|-||.+.....+....|+..+.... .--.-|+.+.. +-+..|.+........
T Consensus 28 ~LYalyKQAt~Gd~~~~~P~~~d~~~~--~K~~AW~~l~gms~~eA~~~Yi~~v~~~~~~ 85 (87)
T PF00887_consen 28 ELYALYKQATHGDCDTPRPGFFDIEGR--AKWDAWKALKGMSKEEAMREYIELVEELIPK 85 (87)
T ss_dssp HHHHHHHHHHTSS--S-CTTTTCHHHH--HHHHHHHTTTTTHHHHHHHHHHHHHHHHHHH
T ss_pred HHHHHHHHHHhCCCcCCCCcchhHHHH--HHHHHHHHccCCCHHHHHHHHHHHHHHHHHh
Confidence 348888888877666666763333333 33466887763 3344555555544433
No 39
>PF02209 VHP: Villin headpiece domain; InterPro: IPR003128 Villin is an F-actin bundling protein involved in the maintenance of the microvilli of the absorptive epithelia. The villin-type "headpiece" domain is a modular motif found at the extreme C terminus of larger "core" domains in over 25 cytoskeletal proteins in plants and animals, often in assocation with the Gelsolin repeat. Although the headpiece is classified as an F-actin-binding domain, it has been shown that not all headpiece domains are intrinsically F-actin-binding motifs, surface charge distribution may be an important element for F-actin recognition []. An autonomously folding, 35 residue, thermostable subdomain (HP36) of the full-length 76 amino acid residue villin headpiece, is the smallest known example of a cooperatively folded domain of a naturally occurring protein. The structure of HP36, as determined by NMR spectroscopy, consists of three short helices surrounding a tightly packed hydrophobic core []. ; GO: 0003779 actin binding, 0007010 cytoskeleton organization; PDB: 1ZV6_A 1QZP_A 1UND_A 2PPZ_A 3TJW_B 1YU8_X 2JM0_A 1WY4_A 3MYC_A 1YU5_X ....
Probab=23.47 E-value=45 Score=18.89 Aligned_cols=10 Identities=30% Similarity=0.381 Sum_probs=7.3
Q ss_pred hhHHHHhhhc
Q 031973 140 SGEVFLQLLT 149 (150)
Q Consensus 140 ~deef~~~~~ 149 (150)
+||||..+|.
T Consensus 3 sd~dF~~vFg 12 (36)
T PF02209_consen 3 SDEDFEKVFG 12 (36)
T ss_dssp -HHHHHHHHS
T ss_pred CHHHHHHHHC
Confidence 6889998873
No 40
>TIGR00787 dctP tripartite ATP-independent periplasmic transporter solute receptor, DctP family. TRAP-T (Tripartite ATP-independent Periplasmic Transporter) family proteins generally consist of three components, and these systems have so far been found in Gram-negative bacteria, Gram-postive bacteria and archaea. The best characterized example is the DctPQM system of Rhodobacter capsulatus, a C4 dicarboxylate (malate, fumarate, succinate) transporter. This model represents the DctP family, one of at least three major families of extracytoplasmic solute receptor for TRAP family transporters. Other are the SnoM family (see pfam03480) and TAXI (TRAP-associated extracytoplasmic immunogenic) family.
Probab=22.72 E-value=1.7e+02 Score=22.85 Aligned_cols=28 Identities=29% Similarity=0.306 Sum_probs=21.4
Q ss_pred HHhhccCChhhhhhHHHHHHHHHHHHHH
Q 031973 76 GEKWKSMSEDEKAPFVERAEKRKSDYNK 103 (150)
Q Consensus 76 ~~~Wk~l~~eeK~~y~~~A~~~k~~y~~ 103 (150)
...|..||++.|....+.+...-.....
T Consensus 213 ~~~~~~L~~e~q~~i~~a~~~~~~~~~~ 240 (257)
T TIGR00787 213 KAFWKSLPPDLQAVVKEAAKEAGEYQRK 240 (257)
T ss_pred HHHHhcCCHHHHHHHHHHHHHHHHHHHH
Confidence 4779999999999998877766544443
No 41
>PRK10236 hypothetical protein; Provisional
Probab=22.39 E-value=80 Score=25.51 Aligned_cols=25 Identities=24% Similarity=0.506 Sum_probs=20.2
Q ss_pred HHHHHHHHhhccCChhhhhhHHHHH
Q 031973 70 TVGKAAGEKWKSMSEDEKAPFVERA 94 (150)
Q Consensus 70 eisk~l~~~Wk~l~~eeK~~y~~~A 94 (150)
-+.+.+...|..||+++++.+...-
T Consensus 117 il~kll~~a~~kms~eE~~~L~~~l 141 (237)
T PRK10236 117 LLEQFLRNTWKKMDEEHKQEFLHAV 141 (237)
T ss_pred HHHHHHHHHHHHCCHHHHHHHHHHH
Confidence 3677899999999999998776543
No 42
>PHA02819 hypothetical protein; Provisional
Probab=22.26 E-value=47 Score=21.81 Aligned_cols=11 Identities=18% Similarity=0.371 Sum_probs=7.6
Q ss_pred cchhHHHHhhh
Q 031973 138 EGSGEVFLQLL 148 (150)
Q Consensus 138 ~~~deef~~~~ 148 (150)
+.+||+|+++|
T Consensus 14 sS~DdDFnnFI 24 (71)
T PHA02819 14 SSSDDDFNNFI 24 (71)
T ss_pred CCchhHHHHHH
Confidence 35677787776
No 43
>PF06945 DUF1289: Protein of unknown function (DUF1289); InterPro: IPR010710 This family consists of a number of hypothetical bacterial proteins. The aligned region spans around 56 residues and contains 4 highly conserved cysteine residues towards the N terminus. The function of this family is unknown.
Probab=21.91 E-value=1.2e+02 Score=18.14 Aligned_cols=23 Identities=35% Similarity=0.744 Sum_probs=16.7
Q ss_pred CHHHHHHHHHHhhccCChhhhhhHHHHH
Q 031973 67 SVATVGKAAGEKWKSMSEDEKAPFVERA 94 (150)
Q Consensus 67 ~~~eisk~l~~~Wk~l~~eeK~~y~~~A 94 (150)
+..||.. |..|++++|.......
T Consensus 23 T~dEI~~-----W~~~s~~er~~i~~~l 45 (51)
T PF06945_consen 23 TLDEIRD-----WKSMSDDERRAILARL 45 (51)
T ss_pred cHHHHHH-----HhhCCHHHHHHHHHHH
Confidence 5667754 9999999987655433
No 44
>PHA02975 hypothetical protein; Provisional
Probab=20.77 E-value=53 Score=21.45 Aligned_cols=11 Identities=18% Similarity=0.371 Sum_probs=7.3
Q ss_pred cchhHHHHhhh
Q 031973 138 EGSGEVFLQLL 148 (150)
Q Consensus 138 ~~~deef~~~~ 148 (150)
+..||+|+++|
T Consensus 14 sS~DdDF~nFI 24 (69)
T PHA02975 14 ESNDSDFEDFI 24 (69)
T ss_pred CCChHHHHHHH
Confidence 34677777776
No 45
>KOG1610 consensus Corticosteroid 11-beta-dehydrogenase and related short chain-type dehydrogenases [Secondary metabolites biosynthesis, transport and catabolism; General function prediction only]
Probab=20.67 E-value=2.3e+02 Score=23.93 Aligned_cols=49 Identities=16% Similarity=0.342 Sum_probs=35.1
Q ss_pred HHHHHHHHHHHhC-------CCC-----CCHHHHHHHHHHhhccCChhhhhhHHHHHHHHH
Q 031973 50 FMEEFRKQFKEAH-------PNN-----KSVATVGKAAGEKWKSMSEDEKAPFVERAEKRK 98 (150)
Q Consensus 50 F~~~~r~~~k~~~-------p~~-----~~~~eisk~l~~~Wk~l~~eeK~~y~~~A~~~k 98 (150)
|+...|.++..-. |+. .....+.+.+.++|..||++.|+.|=+.+-.+.
T Consensus 188 f~D~lR~EL~~fGV~VsiiePG~f~T~l~~~~~~~~~~~~~w~~l~~e~k~~YGedy~~~~ 248 (322)
T KOG1610|consen 188 FSDSLRRELRPFGVKVSIIEPGFFKTNLANPEKLEKRMKEIWERLPQETKDEYGEDYFEDY 248 (322)
T ss_pred HHHHHHHHHHhcCcEEEEeccCccccccCChHHHHHHHHHHHhcCCHHHHHHHHHHHHHHH
Confidence 6666676665322 321 245788899999999999999999987665543
No 46
>PHA02844 putative transmembrane protein; Provisional
Probab=20.48 E-value=54 Score=21.76 Aligned_cols=11 Identities=18% Similarity=0.347 Sum_probs=7.4
Q ss_pred cchhHHHHhhh
Q 031973 138 EGSGEVFLQLL 148 (150)
Q Consensus 138 ~~~deef~~~~ 148 (150)
+.+||+|+++|
T Consensus 14 sS~DdDFnnFI 24 (75)
T PHA02844 14 SSENEDFNNFI 24 (75)
T ss_pred CCchHHHHHHH
Confidence 34677777776
Done!