Query 024359
Match_columns 268
No_of_seqs 142 out of 263
Neff 3.8
Searched_HMMs 46136
Date Fri Mar 29 03:58:45 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/024359.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/024359hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 KOG0489 Transcription factor z 99.8 5.5E-20 1.2E-24 167.4 4.2 63 6-80 158-220 (261)
2 KOG0842 Transcription factor t 99.8 1.7E-19 3.6E-24 169.0 6.6 68 3-82 149-216 (307)
3 KOG0488 Transcription factor B 99.8 3.3E-19 7E-24 166.8 5.6 67 2-80 167-233 (309)
4 KOG0487 Transcription factor A 99.7 9.7E-18 2.1E-22 157.3 3.8 63 4-78 232-294 (308)
5 KOG0843 Transcription factor E 99.7 8.3E-17 1.8E-21 142.0 5.2 63 6-80 101-163 (197)
6 KOG0492 Transcription factor M 99.6 8.8E-17 1.9E-21 144.8 4.6 61 7-79 144-204 (246)
7 KOG0485 Transcription factor N 99.6 5.2E-16 1.1E-20 140.8 6.0 62 7-80 104-165 (268)
8 KOG0850 Transcription factor D 99.6 1.1E-15 2.3E-20 139.0 7.0 66 2-79 117-182 (245)
9 PF00046 Homeobox: Homeobox do 99.6 1.6E-15 3.4E-20 106.9 5.6 57 8-76 1-57 (57)
10 KOG0484 Transcription factor P 99.6 2.2E-15 4.8E-20 123.9 3.8 59 8-78 18-76 (125)
11 KOG0493 Transcription factor E 99.5 1.7E-14 3.6E-19 134.0 5.1 61 8-80 247-307 (342)
12 smart00389 HOX Homeodomain. DN 99.5 5.5E-14 1.2E-18 97.8 5.6 56 8-75 1-56 (56)
13 KOG0491 Transcription factor B 99.4 1.9E-14 4.1E-19 126.1 1.4 62 7-80 100-161 (194)
14 cd00086 homeodomain Homeodomai 99.4 1.9E-13 4.1E-18 95.4 6.0 57 9-77 2-58 (59)
15 KOG2251 Homeobox transcription 99.4 4.3E-13 9.2E-18 121.4 9.5 64 4-79 34-97 (228)
16 KOG0494 Transcription factor C 99.4 2.6E-13 5.6E-18 126.0 5.1 60 8-79 142-201 (332)
17 KOG0848 Transcription factor C 99.4 1.1E-13 2.4E-18 128.5 2.1 60 9-80 201-260 (317)
18 KOG0844 Transcription factor E 99.4 1.3E-13 2.8E-18 130.3 2.2 61 7-79 181-241 (408)
19 TIGR01565 homeo_ZF_HD homeobox 99.4 5.3E-13 1.1E-17 98.4 5.0 52 8-71 2-57 (58)
20 KOG0483 Transcription factor H 99.3 1.2E-12 2.6E-17 116.6 4.1 60 8-79 51-110 (198)
21 KOG0847 Transcription factor, 99.2 6E-12 1.3E-16 114.9 2.7 62 7-80 167-228 (288)
22 COG5576 Homeodomain-containing 99.1 7.5E-11 1.6E-15 101.6 5.6 61 8-80 52-112 (156)
23 KOG0486 Transcription factor P 99.1 8.5E-11 1.8E-15 111.3 3.4 60 7-78 112-171 (351)
24 KOG0490 Transcription factor, 99.0 1.9E-10 4.1E-15 99.0 4.2 62 6-79 59-120 (235)
25 KOG4577 Transcription factor L 98.9 1.4E-09 3E-14 102.7 4.4 60 6-77 166-225 (383)
26 KOG0849 Transcription factor P 98.6 3.4E-08 7.4E-13 94.3 4.1 61 6-78 175-235 (354)
27 KOG3802 Transcription factor O 98.4 2.7E-07 5.9E-12 89.7 4.2 65 2-79 289-354 (398)
28 KOG0775 Transcription factor S 98.2 6.3E-06 1.4E-10 77.5 9.0 53 14-78 183-235 (304)
29 PF05920 Homeobox_KN: Homeobox 98.1 2.7E-06 5.8E-11 58.4 3.0 38 26-73 3-40 (40)
30 KOG2252 CCAAT displacement pro 97.9 3.8E-05 8.3E-10 77.5 8.7 57 6-74 419-475 (558)
31 KOG0490 Transcription factor, 97.8 2.3E-05 4.9E-10 67.6 3.8 62 6-79 152-213 (235)
32 KOG0774 Transcription factor P 97.2 0.00057 1.2E-08 64.6 5.3 64 7-80 188-252 (334)
33 KOG1168 Transcription factor A 96.6 0.00076 1.7E-08 64.6 1.4 63 6-80 308-370 (385)
34 PF15057 DUF4537: Domain of un 95.6 0.077 1.7E-06 44.1 8.3 100 147-264 6-105 (124)
35 PF11717 Tudor-knot: RNA bindi 95.1 0.014 3E-07 41.9 2.1 40 151-192 12-51 (55)
36 KOG1146 Homeobox protein [Gene 94.6 0.028 6.1E-07 62.2 3.5 59 7-77 903-961 (1406)
37 KOG0773 Transcription factor M 94.4 0.036 7.8E-07 52.0 3.3 58 7-74 239-297 (342)
38 PF11569 Homez: Homeodomain le 93.3 0.074 1.6E-06 39.6 2.6 42 19-72 10-51 (56)
39 PLN00104 MYST -like histone ac 92.9 0.33 7.1E-06 48.8 7.3 51 147-199 62-115 (450)
40 cd00024 CHROMO Chromatin organ 92.6 0.073 1.6E-06 36.5 1.7 37 156-192 3-40 (55)
41 cd04508 TUDOR Tudor domains ar 90.8 0.23 5.1E-06 33.4 2.6 36 146-186 5-40 (48)
42 PF04218 CENP-B_N: CENP-B N-te 90.1 0.64 1.4E-05 33.2 4.5 46 8-70 1-46 (53)
43 smart00298 CHROMO Chromatin or 89.0 0.33 7E-06 33.0 2.2 37 156-192 2-38 (55)
44 smart00333 TUDOR Tudor domain. 88.6 0.52 1.1E-05 32.7 3.0 43 145-196 9-51 (57)
45 PF00385 Chromo: Chromo (CHRro 88.6 0.16 3.5E-06 35.4 0.4 37 156-192 1-39 (55)
46 PF02820 MBT: mbt repeat; Int 88.6 0.33 7.2E-06 36.4 2.1 45 143-192 1-45 (73)
47 smart00561 MBT Present in Dros 87.5 0.62 1.4E-05 37.2 3.2 46 142-192 31-76 (96)
48 PF12824 MRP-L20: Mitochondria 86.6 1.4 3E-05 38.7 5.2 56 10-69 82-137 (164)
49 PF05641 Agenet: Agenet domain 85.9 2.4 5.1E-05 31.4 5.4 41 210-258 1-42 (68)
50 smart00333 TUDOR Tudor domain. 85.8 2.5 5.5E-05 29.2 5.3 45 209-264 2-46 (57)
51 PF12148 DUF3590: Protein of u 84.1 1.3 2.9E-05 35.4 3.5 70 146-218 3-74 (85)
52 PF11717 Tudor-knot: RNA bindi 83.2 5.8 0.00013 28.3 6.2 40 210-257 1-40 (55)
53 smart00743 Agenet Tudor-like d 82.3 4.8 0.0001 28.6 5.5 46 209-264 2-49 (61)
54 KOG3623 Homeobox transcription 80.5 7.2 0.00016 42.2 8.2 52 14-78 564-615 (1007)
55 PF05641 Agenet: Agenet domain 74.2 3.2 6.8E-05 30.8 2.6 40 152-197 17-62 (68)
56 PF02796 HTH_7: Helix-turn-hel 71.9 4 8.8E-05 27.9 2.5 35 2-50 1-35 (45)
57 PF06003 SMN: Survival motor n 71.6 2.8 6E-05 39.1 2.2 42 145-190 75-116 (264)
58 PF13565 HTH_32: Homeodomain-l 71.0 15 0.00033 26.6 5.7 53 5-66 24-76 (77)
59 smart00743 Agenet Tudor-like d 68.2 5.3 0.00012 28.3 2.6 36 145-185 9-46 (61)
60 cd04508 TUDOR Tudor domains ar 62.7 24 0.00053 23.4 4.9 39 213-261 1-40 (48)
61 PF00567 TUDOR: Tudor domain; 57.0 10 0.00022 28.5 2.5 51 149-207 62-117 (121)
62 PF04967 HTH_10: HTH DNA bindi 55.0 16 0.00036 26.5 3.2 34 14-50 1-37 (53)
63 cd00569 HTH_Hin_like Helix-tur 54.5 34 0.00075 19.2 4.1 31 13-50 5-35 (42)
64 PF06003 SMN: Survival motor n 54.2 23 0.00051 33.0 4.9 48 208-264 67-114 (264)
65 PF13551 HTH_29: Winged helix- 53.3 45 0.00097 25.2 5.6 53 8-68 52-109 (112)
66 PF12148 DUF3590: Protein of u 52.5 31 0.00068 27.7 4.7 36 224-260 8-43 (85)
67 PF13873 Myb_DNA-bind_5: Myb/S 50.4 27 0.00058 25.8 3.8 62 12-77 3-75 (78)
68 PTZ00064 histone acetyltransfe 47.6 16 0.00034 38.0 2.8 38 170-209 147-184 (552)
69 PF11523 DUF3223: Protein of u 47.0 22 0.00047 27.4 3.0 28 233-262 41-69 (76)
70 PF00249 Myb_DNA-binding: Myb- 42.4 72 0.0016 21.6 4.7 45 12-69 2-46 (48)
71 smart00027 EH Eps15 homology d 41.0 65 0.0014 24.6 4.8 45 14-68 4-51 (96)
72 PF11516 DUF3220: Protein of u 39.7 13 0.00029 30.1 0.8 19 51-70 22-40 (106)
73 smart00717 SANT SANT SWI3, AD 39.6 56 0.0012 20.6 3.7 44 12-69 2-45 (49)
74 PF01527 HTH_Tnp_1: Transposas 38.2 42 0.00092 24.1 3.2 46 9-70 2-47 (76)
75 PF10668 Phage_terminase: Phag 36.9 23 0.0005 26.6 1.6 30 24-67 14-43 (60)
76 PF09465 LBR_tudor: Lamin-B re 36.6 1.2E+02 0.0025 22.8 5.2 40 209-257 5-44 (55)
77 PF13518 HTH_28: Helix-turn-he 34.4 32 0.00069 22.9 1.9 23 38-70 14-36 (52)
78 COG3413 Predicted DNA binding 34.3 43 0.00093 29.5 3.2 35 13-50 155-192 (215)
79 PTZ00183 centrin; Provisional 33.9 1.4E+02 0.0031 23.5 5.9 41 9-49 6-49 (158)
80 PRK03975 tfx putative transcri 33.5 58 0.0012 28.1 3.7 49 12-78 5-53 (141)
81 TIGR01321 TrpR trp operon repr 33.3 27 0.00058 28.4 1.6 56 13-69 32-92 (94)
82 COG3458 Acetyl esterase (deace 33.0 59 0.0013 31.8 4.1 93 140-239 56-163 (321)
83 PLN00104 MYST -like histone ac 32.6 1E+02 0.0023 31.5 5.9 48 208-258 52-99 (450)
84 PF03672 UPF0154: Uncharacteri 31.8 80 0.0017 24.2 3.8 36 20-65 20-55 (64)
85 PF04717 Phage_base_V: Phage-r 31.7 91 0.002 23.2 4.2 50 170-224 9-58 (79)
86 cd06171 Sigma70_r4 Sigma70, re 30.4 38 0.00082 21.5 1.7 44 13-73 10-53 (55)
87 PF15057 DUF4537: Domain of un 30.3 95 0.0021 25.8 4.4 40 213-262 1-40 (124)
88 PRK10072 putative transcriptio 29.6 37 0.00081 27.3 1.8 24 39-72 49-72 (96)
89 PF13936 HTH_38: Helix-turn-he 29.6 46 0.001 22.7 2.1 32 12-50 3-34 (44)
90 PF07930 DAP_B: D-aminopeptida 29.5 31 0.00067 28.0 1.3 73 149-232 11-83 (88)
91 PRK07539 NADH dehydrogenase su 28.3 1.2E+02 0.0026 25.8 4.9 20 31-50 35-54 (154)
92 PF05506 DUF756: Domain of unk 28.0 54 0.0012 25.0 2.4 23 150-180 66-88 (89)
93 PTZ00184 calmodulin; Provision 28.0 1.7E+02 0.0037 22.6 5.4 37 13-49 4-43 (149)
94 COG2944 Predicted transcriptio 27.9 78 0.0017 26.3 3.5 39 14-71 44-82 (104)
95 cd00167 SANT 'SWI3, ADA2, N-Co 27.5 1.2E+02 0.0025 18.8 3.6 43 13-69 1-43 (45)
96 PRK07571 bidirectional hydroge 26.0 1.4E+02 0.003 26.4 4.9 38 13-50 15-68 (169)
97 PRK04980 hypothetical protein; 25.6 1.2E+02 0.0027 24.9 4.2 32 209-241 31-62 (102)
98 PF01343 Peptidase_S49: Peptid 22.2 1.6E+02 0.0034 24.7 4.4 54 2-70 68-121 (154)
99 TIGR02607 antidote_HigA addict 21.6 2.1E+02 0.0046 20.5 4.5 19 31-49 42-60 (78)
100 PF08880 QLQ: QLQ; InterPro: 21.5 70 0.0015 21.8 1.7 16 13-28 2-17 (37)
101 PF00196 GerE: Bacterial regul 21.4 1.4E+02 0.0031 20.8 3.4 45 12-74 2-46 (58)
102 PRK00523 hypothetical protein; 20.9 1.6E+02 0.0035 23.1 3.8 36 20-65 28-63 (72)
103 PRK15451 tRNA cmo(5)U34 methyl 20.9 2E+02 0.0042 25.7 5.0 47 13-72 189-235 (247)
104 PF14773 VIGSSK: Helicase-asso 20.6 46 0.001 25.3 0.8 15 251-265 36-50 (61)
105 PF04545 Sigma70_r4: Sigma-70, 20.3 1.9E+02 0.004 19.5 3.7 39 13-68 4-42 (50)
No 1
>KOG0489 consensus Transcription factor zerknullt and related HOX domain proteins [General function prediction only]
Probab=99.79 E-value=5.5e-20 Score=167.44 Aligned_cols=63 Identities=21% Similarity=0.298 Sum_probs=58.8
Q ss_pred CCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcccC
Q 024359 6 SNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSIK 80 (268)
Q Consensus 6 s~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~~ 80 (268)
..++.||.||..|+.||||+|.. |+||++..|.+||..|+|+ |.|||||||||||||||....
T Consensus 158 ~~kR~RtayT~~QllELEkEFhf--N~YLtR~RRiEiA~~L~Lt----------ErQIKIWFQNRRMK~Kk~~k~ 220 (261)
T KOG0489|consen 158 KSKRRRTAFTRYQLLELEKEFHF--NKYLTRSRRIEIAHALNLT----------ERQIKIWFQNRRMKWKKENKA 220 (261)
T ss_pred CCCCCCcccchhhhhhhhhhhcc--ccccchHHHHHHHhhcchh----------HHHHHHHHHHHHHHHHHhhcc
Confidence 46899999999999999999999 6999999999999999965 699999999999999988854
No 2
>KOG0842 consensus Transcription factor tinman/NKX2-3, contains HOX domain [Transcription]
Probab=99.78 E-value=1.7e-19 Score=169.03 Aligned_cols=68 Identities=25% Similarity=0.296 Sum_probs=61.1
Q ss_pred CCCCCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcccCCC
Q 024359 3 RPPSNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSIKSP 82 (268)
Q Consensus 3 rPps~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~~~p 82 (268)
-+..+|++|..||+.||.|||+.|++ ++||+..+|+.||..|+|+ +|||||||||||||.||+-.--.
T Consensus 149 ~~~~kRKrRVLFSqAQV~ELERRFrq--QRYLSAPERE~LA~~LrLT----------~TQVKIWFQNrRYK~KR~~~dk~ 216 (307)
T KOG0842|consen 149 GKRKKRKRRVLFSQAQVYELERRFRQ--QRYLSAPEREHLASSLRLT----------PTQVKIWFQNRRYKTKRQQKDKA 216 (307)
T ss_pred ccccccccccccchhHHHHHHHHHHh--hhccccHhHHHHHHhcCCC----------chheeeeeecchhhhhhhhhhhh
Confidence 35577999999999999999999999 4999999999999999965 59999999999999998876433
No 3
>KOG0488 consensus Transcription factor BarH and related HOX domain proteins [General function prediction only]
Probab=99.76 E-value=3.3e-19 Score=166.77 Aligned_cols=67 Identities=22% Similarity=0.315 Sum_probs=61.9
Q ss_pred CCCCCCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcccC
Q 024359 2 GRPPSNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSIK 80 (268)
Q Consensus 2 GrPps~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~~ 80 (268)
+.|+..++.||.||..||.+|||.|+.+ +||+..+|.+||.+|||+ +.|||+||||||||||+.+..
T Consensus 167 ~~pkK~RksRTaFT~~Ql~~LEkrF~~Q--KYLS~~DR~~LA~~LgLT----------daQVKtWfQNRRtKWKrq~a~ 233 (309)
T KOG0488|consen 167 STPKKRRKSRTAFSDHQLFELEKRFEKQ--KYLSVADRIELAASLGLT----------DAQVKTWFQNRRTKWKRQTAE 233 (309)
T ss_pred CCCcccccchhhhhHHHHHHHHHHHHHh--hcccHHHHHHHHHHcCCc----------hhhHHHHHhhhhHHHHHHHHh
Confidence 4677889999999999999999999995 999999999999999965 699999999999999998764
No 4
>KOG0487 consensus Transcription factor Abd-B, contains HOX domain [Transcription]
Probab=99.69 E-value=9.7e-18 Score=157.28 Aligned_cols=63 Identities=21% Similarity=0.269 Sum_probs=58.7
Q ss_pred CCCCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcc
Q 024359 4 PPSNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKS 78 (268)
Q Consensus 4 Pps~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~ 78 (268)
+.+.|+.|.-+|+.||.||||+|.- |.||+++.|-+|++.|||| +.||||||||||||.||..
T Consensus 232 ~~~~RKKRcPYTK~QtlELEkEFlf--N~YitkeKR~ElSr~lNLT----------eRQVKIWFQNRRMK~KK~~ 294 (308)
T KOG0487|consen 232 ARRGRKKRCPYTKHQTLELEKEFLF--NMYITKEKRLELSRTLNLT----------ERQVKIWFQNRRMKEKKVN 294 (308)
T ss_pred ccccccccCCchHHHHHHHHHHHHH--HHHHhHHHHHHHHHhcccc----------hhheeeeehhhhhHHhhhh
Confidence 4567999999999999999999999 6999999999999999965 6999999999999999877
No 5
>KOG0843 consensus Transcription factor EMX1 and related HOX domain proteins [Transcription]
Probab=99.65 E-value=8.3e-17 Score=142.02 Aligned_cols=63 Identities=25% Similarity=0.278 Sum_probs=58.3
Q ss_pred CCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcccC
Q 024359 6 SNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSIK 80 (268)
Q Consensus 6 s~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~~ 80 (268)
..++.||.||.+|+..||..|+. ++|+...+|.+||..|||| ++|||+||||||+|.||...+
T Consensus 101 ~~kr~RT~ft~~Ql~~LE~~F~~--~~Yvvg~eR~~LA~~L~Ls----------etQVkvWFQNRRtk~kr~~~e 163 (197)
T KOG0843|consen 101 RPKRIRTAFTPEQLLKLEHAFEG--NQYVVGAERKQLAQSLSLS----------ETQVKVWFQNRRTKHKRMQQE 163 (197)
T ss_pred CCCccccccCHHHHHHHHHHHhc--CCeeechHHHHHHHHcCCC----------hhHhhhhhhhhhHHHHHHHHH
Confidence 45789999999999999999999 6999999999999999976 599999999999999988764
No 6
>KOG0492 consensus Transcription factor MSH, contains HOX domain [General function prediction only]
Probab=99.65 E-value=8.8e-17 Score=144.79 Aligned_cols=61 Identities=23% Similarity=0.277 Sum_probs=56.4
Q ss_pred CCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhccc
Q 024359 7 NGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSI 79 (268)
Q Consensus 7 ~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~ 79 (268)
+|+|||.||..||..||+.|++. +||+..+|.+++.+|+| |++||||||||||+|-|+-.-
T Consensus 144 nRkPRtPFTtqQLlaLErkfrek--qYLSiaEraefSsSL~L----------TeTqVKIWFQNRRAKaKRlQe 204 (246)
T KOG0492|consen 144 NRKPRTPFTTQQLLALERKFREK--QYLSIAERAEFSSSLEL----------TETQVKIWFQNRRAKAKRLQE 204 (246)
T ss_pred CCCCCCCCCHHHHHHHHHHHhHh--hhhhHHHHHhhhhhhhh----------hhhheehhhhhhhHHHHHHHH
Confidence 68999999999999999999995 99999999999999995 679999999999999887653
No 7
>KOG0485 consensus Transcription factor NKX-5.1/HMX1, contains HOX domain [Transcription]
Probab=99.61 E-value=5.2e-16 Score=140.80 Aligned_cols=62 Identities=24% Similarity=0.261 Sum_probs=57.9
Q ss_pred CCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcccC
Q 024359 7 NGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSIK 80 (268)
Q Consensus 7 ~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~~ 80 (268)
++++||.|+.+||.+||-.|+. .+||+..+|..||.+|.| ||+||||||||||.|||++...
T Consensus 104 KKktRTvFSraQV~qLEs~Fe~--krYLSsaeRa~LA~sLqL----------TETQVKIWFQNRRnKwKRq~aa 165 (268)
T KOG0485|consen 104 KKKTRTVFSRAQVFQLESTFEL--KRYLSSAERAGLAASLQL----------TETQVKIWFQNRRNKWKRQYAA 165 (268)
T ss_pred cccchhhhhHHHHHHHHHHHHH--HhhhhHHHHhHHHHhhhh----------hhhhhhhhhhhhhHHHHHHHhh
Confidence 5789999999999999999999 499999999999999995 6799999999999999998765
No 8
>KOG0850 consensus Transcription factor DLX and related proteins with LIM Zn-binding and HOX domains [Transcription]
Probab=99.60 E-value=1.1e-15 Score=138.96 Aligned_cols=66 Identities=20% Similarity=0.267 Sum_probs=60.8
Q ss_pred CCCCCCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhccc
Q 024359 2 GRPPSNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSI 79 (268)
Q Consensus 2 GrPps~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~ 79 (268)
|++..-|+|||.|+..||+.|-+.|++ .+||...+|.+||..|||+ .+||||||||||.|.||...
T Consensus 117 gk~KK~RKPRTIYSS~QLqaL~rRFQk--TQYLALPERAeLAAsLGLT----------QTQVKIWFQNrRSK~KKl~k 182 (245)
T KOG0850|consen 117 GKGKKVRKPRTIYSSLQLQALNRRFQQ--TQYLALPERAELAASLGLT----------QTQVKIWFQNRRSKFKKLKK 182 (245)
T ss_pred CCcccccCCcccccHHHHHHHHHHHhh--cchhcCcHHHHHHHHhCCc----------hhHhhhhhhhhHHHHHHHHh
Confidence 667777999999999999999999999 5999999999999999976 59999999999999988764
No 9
>PF00046 Homeobox: Homeobox domain not present here.; InterPro: IPR001356 The homeobox domain was first identified in a number of drosophila homeotic and segmentation proteins, but is now known to be well-conserved in many other animals, including vertebrates [, , ]. Hox genes encode homeodomain-containing transcriptional regulators that operate differential genetic programs along the anterior-posterior axis of animal bodies []. The domain binds DNA through a helix-turn-helix (HTH) structure. The HTH motif is characterised by two alpha-helices, which make intimate contacts with the DNA and are joined by a short turn. The second helix binds to DNA via a number of hydrogen bonds and hydrophobic interactions, which occur between specific side chains and the exposed bases and thymine methyl groups within the major groove of the DNA []. The first helix helps to stabilise the structure. The motif is very similar in sequence and structure in a wide range of DNA-binding proteins (e.g., cro and repressor proteins, homeotic proteins, etc.). One of the principal differences between HTH motifs in these different proteins arises from the stereo-chemical requirement for glycine in the turn which is needed to avoid steric interference of the beta-carbon with the main chain: for cro and repressor proteins the glycine appears to be mandatory, while for many of the homeotic and other DNA-binding proteins the requirement is relaxed.; GO: 0003700 sequence-specific DNA binding transcription factor activity, 0043565 sequence-specific DNA binding, 0006355 regulation of transcription, DNA-dependent; PDB: 2DA3_A 1LFB_A 2LFB_A 2ECB_A 2DA5_A 3D1N_O 3A03_A 2XSD_C 3CMY_A 1AHD_P ....
Probab=99.59 E-value=1.6e-15 Score=106.90 Aligned_cols=57 Identities=35% Similarity=0.502 Sum_probs=53.3
Q ss_pred CCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhh
Q 024359 8 GGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRA 76 (268)
Q Consensus 8 ~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kk 76 (268)
+++|+.||..|+..||..|.. ++||+.+.++.||..+|++ ..||++||||||.+.|+
T Consensus 1 kr~r~~~t~~q~~~L~~~f~~--~~~p~~~~~~~la~~l~l~----------~~~V~~WF~nrR~k~kk 57 (57)
T PF00046_consen 1 KRKRTRFTKEQLKVLEEYFQE--NPYPSKEEREELAKELGLT----------ERQVKNWFQNRRRKEKK 57 (57)
T ss_dssp SSSSSSSSHHHHHHHHHHHHH--SSSCHHHHHHHHHHHHTSS----------HHHHHHHHHHHHHHHHH
T ss_pred CcCCCCCCHHHHHHHHHHHHH--hcccccccccccccccccc----------ccccccCHHHhHHHhCc
Confidence 478999999999999999999 6999999999999999976 59999999999999875
No 10
>KOG0484 consensus Transcription factor PHOX2/ARIX, contains HOX domain [Transcription]
Probab=99.55 E-value=2.2e-15 Score=123.92 Aligned_cols=59 Identities=27% Similarity=0.369 Sum_probs=54.4
Q ss_pred CCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcc
Q 024359 8 GGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKS 78 (268)
Q Consensus 8 ~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~ 78 (268)
++-||.||..||.|||++|.+ .+||+.-.|++||-++.| |+..||+||||||+|.||.-
T Consensus 18 RRIRTTFTS~QLkELErvF~E--THYPDIYTREEiA~kidL----------TEARVQVWFQNRRAKfRKQE 76 (125)
T KOG0484|consen 18 RRIRTTFTSAQLKELERVFAE--THYPDIYTREEIALKIDL----------TEARVQVWFQNRRAKFRKQE 76 (125)
T ss_pred hhhhhhhhHHHHHHHHHHHHh--hcCCcchhHHHHHHhhhh----------hHHHHHHHHHhhHHHHHHHH
Confidence 788999999999999999999 499999999999999995 57999999999999987654
No 11
>KOG0493 consensus Transcription factor Engrailed, contains HOX domain [General function prediction only]
Probab=99.50 E-value=1.7e-14 Score=133.98 Aligned_cols=61 Identities=23% Similarity=0.359 Sum_probs=57.3
Q ss_pred CCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcccC
Q 024359 8 GGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSIK 80 (268)
Q Consensus 8 ~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~~ 80 (268)
++|||.||.+||++|...|++ |+||+...||+||.+|+|. |.||||||||+|+|.||.+..
T Consensus 247 KRPRTAFtaeQL~RLK~EF~e--nRYlTEqRRQ~La~ELgLN----------EsQIKIWFQNKRAKiKKsTgs 307 (342)
T KOG0493|consen 247 KRPRTAFTAEQLQRLKAEFQE--NRYLTEQRRQELAQELGLN----------ESQIKIWFQNKRAKIKKSTGS 307 (342)
T ss_pred cCccccccHHHHHHHHHHHhh--hhhHHHHHHHHHHHHhCcC----------HHHhhHHhhhhhhhhhhccCC
Confidence 789999999999999999999 6999999999999999976 599999999999999987754
No 12
>smart00389 HOX Homeodomain. DNA-binding factors that are involved in the transcriptional regulation of key developmental processes
Probab=99.48 E-value=5.5e-14 Score=97.85 Aligned_cols=56 Identities=39% Similarity=0.525 Sum_probs=50.9
Q ss_pred CCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhh
Q 024359 8 GGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIR 75 (268)
Q Consensus 8 ~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~K 75 (268)
+++|+.||++|+..||..|.. +.||+...++.||..+|++ .+||++||+|||++.+
T Consensus 1 ~k~r~~~~~~~~~~L~~~f~~--~~~P~~~~~~~la~~~~l~----------~~qV~~WF~nrR~~~~ 56 (56)
T smart00389 1 RRKRTSFTPEQLEELEKEFQK--NPYPSREEREELAAKLGLS----------ERQVKVWFQNRRAKWK 56 (56)
T ss_pred CCCCCcCCHHHHHHHHHHHHh--CCCCCHHHHHHHHHHHCcC----------HHHHHHhHHHHhhccC
Confidence 357888999999999999999 5899999999999999976 5999999999998753
No 13
>KOG0491 consensus Transcription factor BSH, contains HOX domain [General function prediction only]
Probab=99.45 E-value=1.9e-14 Score=126.12 Aligned_cols=62 Identities=24% Similarity=0.292 Sum_probs=57.3
Q ss_pred CCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcccC
Q 024359 7 NGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSIK 80 (268)
Q Consensus 7 ~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~~ 80 (268)
.++.|+.|+..|+.-||+.|+.+ +||+..+|++||..|||| ++|||+||||||||.||...+
T Consensus 100 r~K~Rtvfs~~ql~~l~~rFe~Q--rYLS~~e~~ELan~L~LS----------~~QVKTWFQNrRMK~Kk~~r~ 161 (194)
T KOG0491|consen 100 RRKARTVFSDPQLSGLEKRFERQ--RYLSTPERQELANALSLS----------ETQVKTWFQNRRMKHKKQQRN 161 (194)
T ss_pred hhhhcccccCccccccHHHHhhh--hhcccHHHHHHHHHhhhh----------HHHHHHHHHHHHHHHHHHHhc
Confidence 47889999999999999999995 999999999999999987 599999999999999987754
No 14
>cd00086 homeodomain Homeodomain; DNA binding domains involved in the transcriptional regulation of key eukaryotic developmental processes; may bind to DNA as monomers or as homo- and/or heterodimers, in a sequence-specific manner.
Probab=99.44 E-value=1.9e-13 Score=95.37 Aligned_cols=57 Identities=35% Similarity=0.556 Sum_probs=52.3
Q ss_pred CCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhc
Q 024359 9 GPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAK 77 (268)
Q Consensus 9 ~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk 77 (268)
+.+..|+..|+..||+.|.. +.||+...++.||..+|++ .+||++||+|||.+.++.
T Consensus 2 ~~r~~~~~~~~~~Le~~f~~--~~~P~~~~~~~la~~~~l~----------~~qV~~WF~nrR~~~~~~ 58 (59)
T cd00086 2 RKRTRFTPEQLEELEKEFEK--NPYPSREEREELAKELGLT----------ERQVKIWFQNRRAKLKRS 58 (59)
T ss_pred CCCCcCCHHHHHHHHHHHHh--CCCCCHHHHHHHHHHHCcC----------HHHHHHHHHHHHHHHhcc
Confidence 57889999999999999999 6999999999999999976 599999999999997653
No 15
>KOG2251 consensus Homeobox transcription factor [Transcription]
Probab=99.44 E-value=4.3e-13 Score=121.38 Aligned_cols=64 Identities=22% Similarity=0.338 Sum_probs=57.6
Q ss_pred CCCCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhccc
Q 024359 4 PPSNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSI 79 (268)
Q Consensus 4 Pps~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~ 79 (268)
|...++.||+||..|+.+||++|.+ .+|||...|++||.+|||. +.+||+||.|||+|+|+...
T Consensus 34 pRkqRRERTtFtr~QlevLe~LF~k--TqYPDv~~rEelAlklnLp----------eSrVqVWFKNRRAK~r~qq~ 97 (228)
T KOG2251|consen 34 PRKQRRERTTFTRKQLEVLEALFAK--TQYPDVFMREELALKLNLP----------ESRVQVWFKNRRAKCRRQQQ 97 (228)
T ss_pred chhcccccceecHHHHHHHHHHHHh--hcCccHHHHHHHHHHhCCc----------hhhhhhhhccccchhhHhhh
Confidence 4456899999999999999999999 5999999999999999976 58899999999999876554
No 16
>KOG0494 consensus Transcription factor CHX10 and related HOX domain proteins [General function prediction only]
Probab=99.39 E-value=2.6e-13 Score=126.00 Aligned_cols=60 Identities=25% Similarity=0.280 Sum_probs=54.4
Q ss_pred CCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhccc
Q 024359 8 GGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSI 79 (268)
Q Consensus 8 ~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~ 79 (268)
|+-||.||..|+.+||+.|++. +|||...|+-||.++.|. |..|++||||||+||||+-.
T Consensus 142 Rh~RTiFT~~Qle~LEkaFkea--HYPDv~Are~la~ktelp----------EDRIqVWfQNRRAKWRk~Ek 201 (332)
T KOG0494|consen 142 RHFRTIFTSYQLEELEKAFKEA--HYPDVYAREMLADKTELP----------EDRIQVWFQNRRAKWRKTEK 201 (332)
T ss_pred ccccchhhHHHHHHHHHHHhhc--cCccHHHHHHHhhhccCc----------hhhhhHHhhhhhHHhhhhhh
Confidence 3348999999999999999995 999999999999999975 59999999999999998754
No 17
>KOG0848 consensus Transcription factor Caudal, contains HOX domain [Transcription]
Probab=99.38 E-value=1.1e-13 Score=128.55 Aligned_cols=60 Identities=25% Similarity=0.246 Sum_probs=54.3
Q ss_pred CCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcccC
Q 024359 9 GPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSIK 80 (268)
Q Consensus 9 ~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~~ 80 (268)
+.|..+|..|-+||||+|.. .+|++.....+||..|+|| |.||||||||||+|.||...|
T Consensus 201 KYRvVYTDhQRLELEKEfh~--SryITirRKSELA~~LgLs----------ERQVKIWFQNRRAKERK~nKK 260 (317)
T KOG0848|consen 201 KYRVVYTDHQRLELEKEFHT--SRYITIRRKSELAATLGLS----------ERQVKIWFQNRRAKERKDNKK 260 (317)
T ss_pred ceeEEecchhhhhhhhhhcc--ccceeeehhHHHHHhhCcc----------HhhhhHhhhhhhHHHHHHHHH
Confidence 45678999999999999999 5999999999999999976 599999999999998877654
No 18
>KOG0844 consensus Transcription factor EVX1, contains HOX domain [Transcription]
Probab=99.38 E-value=1.3e-13 Score=130.34 Aligned_cols=61 Identities=20% Similarity=0.255 Sum_probs=55.9
Q ss_pred CCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhccc
Q 024359 7 NGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSI 79 (268)
Q Consensus 7 ~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~ 79 (268)
-++.||.||.+||++|||.|-. +.|.++..|-+||..|||. |+-||+||||||||.|+.-.
T Consensus 181 mRRYRTAFTReQIaRLEKEFyr--ENYVSRprRcELAAaLNLP----------EtTIKVWFQNRRMKDKRQRl 241 (408)
T KOG0844|consen 181 MRRYRTAFTREQIARLEKEFYR--ENYVSRPRRCELAAALNLP----------ETTIKVWFQNRRMKDKRQRL 241 (408)
T ss_pred HHHHHhhhhHHHHHHHHHHHHH--hccccCchhhhHHHhhCCC----------cceeehhhhhchhhhhhhhh
Confidence 3789999999999999999988 5799999999999999986 59999999999999887753
No 19
>TIGR01565 homeo_ZF_HD homeobox domain, ZF-HD class. This model represents a class of homoebox domain that differs substantially from the typical homoebox domain described in pfam model pfam00046. It is found in both C4 and C3 plants.
Probab=99.38 E-value=5.3e-13 Score=98.39 Aligned_cols=52 Identities=15% Similarity=0.264 Sum_probs=48.6
Q ss_pred CCCccccCHHHHHHHHHHHHhccCCC----CCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcch
Q 024359 8 GGPAFRFNPAEVTEMEGILQEHHNAM----PSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRR 71 (268)
Q Consensus 8 ~~pRt~FT~~Qv~eLEk~F~~~~~~y----p~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR 71 (268)
+++||.||.+|+.+||+.|+. ++| |+...+++||..+|++ +.+||+||||-+
T Consensus 2 kR~RT~Ft~~Q~~~Le~~fe~--~~y~~~~~~~~~r~~la~~lgl~----------~~vvKVWfqN~k 57 (58)
T TIGR01565 2 KRRRTKFTAEQKEKMRDFAEK--LGWKLKDKRREEVREFCEEIGVT----------RKVFKVWMHNNK 57 (58)
T ss_pred CCCCCCCCHHHHHHHHHHHHH--cCCCCCCCCHHHHHHHHHHhCCC----------HHHeeeecccCC
Confidence 689999999999999999999 599 9999999999999976 599999999964
No 20
>KOG0483 consensus Transcription factor HEX, contains HOX and HALZ domains [Transcription]
Probab=99.31 E-value=1.2e-12 Score=116.63 Aligned_cols=60 Identities=27% Similarity=0.345 Sum_probs=54.5
Q ss_pred CCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhccc
Q 024359 8 GGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSI 79 (268)
Q Consensus 8 ~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~ 79 (268)
.....+||.+|+..||+.|+. +.||....+..||..|||. +.||.+||||||++||.|..
T Consensus 51 ~~kk~Rlt~eQ~~~LE~~F~~--~~~L~p~~K~~LAk~LgL~----------pRQVavWFQNRRARwK~kql 110 (198)
T KOG0483|consen 51 KGKKRRLTSEQVKFLEKSFES--EKKLEPERKKKLAKELGLQ----------PRQVAVWFQNRRARWKTKQL 110 (198)
T ss_pred ccccccccHHHHHHhHHhhcc--ccccChHHHHHHHHhhCCC----------hhHHHHHHhhccccccchhh
Confidence 456679999999999999999 5999999999999999976 59999999999999998764
No 21
>KOG0847 consensus Transcription factor, contains HOX domain [Transcription]
Probab=99.20 E-value=6e-12 Score=114.89 Aligned_cols=62 Identities=23% Similarity=0.280 Sum_probs=56.1
Q ss_pred CCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcccC
Q 024359 7 NGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSIK 80 (268)
Q Consensus 7 ~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~~ 80 (268)
+...|.+|+-.||..||+-|++ .+||-..+|-+||..+|+ ++.||++||||||+|||||..-
T Consensus 167 rk~srPTf~g~qi~~le~~feq--tkylaG~~ra~lA~~lgm----------teSqvkVWFQNRRTKWRKkhAa 228 (288)
T KOG0847|consen 167 RKQSRPTFTGHQIYQLERKFEQ--TKYLAGADRAQLAQELNM----------TESQVKVWFQNRRTKWRKKHAA 228 (288)
T ss_pred ccccCCCccchhhhhhhhhhhh--hhcccchhHHHhhccccc----------cHHHHHHHHhcchhhhhhhhcc
Confidence 3456778999999999999999 499999999999999995 5799999999999999999863
No 22
>COG5576 Homeodomain-containing transcription factor [Transcription]
Probab=99.12 E-value=7.5e-11 Score=101.61 Aligned_cols=61 Identities=25% Similarity=0.303 Sum_probs=56.2
Q ss_pred CCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcccC
Q 024359 8 GGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSIK 80 (268)
Q Consensus 8 ~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~~ 80 (268)
++.|++-|..|+.-||+.|+. ++||+...|++|+..+|+++ +-||+||||||++.|++...
T Consensus 52 ~~~r~R~t~~Q~~vL~~~F~i--~p~Ps~~~r~~L~~~lnm~~----------ksVqIWFQNkR~~~k~~~~~ 112 (156)
T COG5576 52 KSKRRRTTDEQLMVLEREFEI--NPYPSSITRIKLSLLLNMPP----------KSVQIWFQNKRAKEKKKRSG 112 (156)
T ss_pred cccceechHHHHHHHHHHhcc--CCCCCHHHHHHHHHhcCCCh----------hhhhhhhchHHHHHHHhccc
Confidence 567889999999999999999 69999999999999999764 99999999999999988863
No 23
>KOG0486 consensus Transcription factor PTX1, contains HOX domain [Transcription]
Probab=99.05 E-value=8.5e-11 Score=111.32 Aligned_cols=60 Identities=23% Similarity=0.346 Sum_probs=56.0
Q ss_pred CCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcc
Q 024359 7 NGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKS 78 (268)
Q Consensus 7 ~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~ 78 (268)
.++.|+-||..|++|||..|+. |+||+-+.|++||--.||+ |+.|.+||.|||+||||+-
T Consensus 112 qrrQrthFtSqqlqele~tF~r--NrypdMstrEEIavwtNlT----------E~rvrvwfknrrakwrkrE 171 (351)
T KOG0486|consen 112 QRRQRTHFTSQQLQELEATFQR--NRYPDMSTREEIAVWTNLT----------EARVRVWFKNRRAKWRKRE 171 (351)
T ss_pred hhhhhhhhHHHHHHHHHHHHhh--ccCCccchhhHHHhhcccc----------chhhhhhcccchhhhhhhh
Confidence 4788999999999999999999 7999999999999999965 6999999999999999874
No 24
>KOG0490 consensus Transcription factor, contains HOX domain [General function prediction only]
Probab=99.03 E-value=1.9e-10 Score=99.00 Aligned_cols=62 Identities=24% Similarity=0.310 Sum_probs=57.0
Q ss_pred CCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhccc
Q 024359 6 SNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSI 79 (268)
Q Consensus 6 s~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~ 79 (268)
..++.|+.||..|+.+||++|+.. +||+...|+.||..++++ +..|++||||||+||++...
T Consensus 59 ~~rr~rt~~~~~ql~~ler~f~~~--h~Pd~~~r~~la~~~~~~----------e~rVqvwFqnrrak~r~~~~ 120 (235)
T KOG0490|consen 59 SKRCARCKFTISQLDELERAFEKV--HLPCFACRECLALLLTGD----------EFRVQVWFQNRRAKDRKEER 120 (235)
T ss_pred cccccCCCCCcCHHHHHHHhhcCC--CcCccchHHHHhhcCCCC----------eeeeehhhhhhcHhhhhhhc
Confidence 468899999999999999999994 999999999999999965 59999999999999998764
No 25
>KOG4577 consensus Transcription factor LIM3, contains LIM and HOX domains [Transcription]
Probab=98.88 E-value=1.4e-09 Score=102.70 Aligned_cols=60 Identities=23% Similarity=0.365 Sum_probs=54.7
Q ss_pred CCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhc
Q 024359 6 SNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAK 77 (268)
Q Consensus 6 s~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk 77 (268)
++++|||+.|..||+-|..+|+.+ .-|.+-+|++|+...||. ...||+||||||+|.|+-
T Consensus 166 ~nKRPRTTItAKqLETLK~AYn~S--pKPARHVREQLsseTGLD----------MRVVQVWFQNRRAKEKRL 225 (383)
T KOG4577|consen 166 SNKRPRTTITAKQLETLKQAYNTS--PKPARHVREQLSSETGLD----------MRVVQVWFQNRRAKEKRL 225 (383)
T ss_pred ccCCCcceeeHHHHHHHHHHhcCC--CchhHHHHHHhhhccCcc----------eeehhhhhhhhhHHHHhh
Confidence 468999999999999999999985 899999999999999976 489999999999997653
No 26
>KOG0849 consensus Transcription factor PRD and related proteins, contain PAX and HOX domains [Transcription]
Probab=98.59 E-value=3.4e-08 Score=94.29 Aligned_cols=61 Identities=25% Similarity=0.330 Sum_probs=55.8
Q ss_pred CCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcc
Q 024359 6 SNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKS 78 (268)
Q Consensus 6 s~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~ 78 (268)
+.++.|+.||+.|+..||+.|+.. +||+...|++||.+.+++ +..|++||||||.+++|..
T Consensus 175 ~~rr~rtsft~~Q~~~le~~f~rt--~yP~i~~Re~La~~i~l~----------e~riqvwf~nrra~~rr~~ 235 (354)
T KOG0849|consen 175 GGRRNRTSFSPSQLEALEECFQRT--PYPDIVGRETLAKETGLP----------EPRVQVWFQNRRAKWRRQH 235 (354)
T ss_pred cccccccccccchHHHHHHHhcCC--CCCchhhHHHHhhhccCC----------chHHHHHHhhhhhhhhhcc
Confidence 356778999999999999999994 799999999999999976 5999999999999998877
No 27
>KOG3802 consensus Transcription factor OCT-1, contains POU and HOX domains [Transcription]
Probab=98.38 E-value=2.7e-07 Score=89.74 Aligned_cols=65 Identities=20% Similarity=0.270 Sum_probs=57.8
Q ss_pred CCCCCCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccch-hhhhhhcchhhhhhccc
Q 024359 2 GRPPSNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQ-VWNWFQNRRYAIRAKSI 79 (268)
Q Consensus 2 GrPps~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQ-Vk~WFQNRR~k~Kkk~~ 79 (268)
|-.+-+|+.||.|+......||+.|.. |.-|+.+++-.||++|+|- |+ |++||=|||.|.||.+.
T Consensus 289 ~a~~RkRKKRTSie~~vr~aLE~~F~~--npKPt~qEIt~iA~~L~le-----------KEVVRVWFCNRRQkeKR~~~ 354 (398)
T KOG3802|consen 289 GAQSRKRKKRTSIEVNVRGALEKHFLK--NPKPTSQEITHIAESLQLE-----------KEVVRVWFCNRRQKEKRITP 354 (398)
T ss_pred hccccccccccceeHHHHHHHHHHHHh--CCCCCHHHHHHHHHHhccc-----------cceEEEEeeccccccccCCC
Confidence 344457999999999999999999999 6999999999999999984 55 88999999999987764
No 28
>KOG0775 consensus Transcription factor SIX and related HOX domain proteins [Transcription]
Probab=98.20 E-value=6.3e-06 Score=77.52 Aligned_cols=53 Identities=32% Similarity=0.405 Sum_probs=43.8
Q ss_pred cCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcc
Q 024359 14 FNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKS 78 (268)
Q Consensus 14 FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~ 78 (268)
|-..--.-|-..|.+ +.||+..+..+||++.||+ .+||-|||.|||++-|.-.
T Consensus 183 FKekSR~~LrewY~~--~~YPsp~eKReLA~aTgLt----------~tQVsNWFKNRRQRDRa~~ 235 (304)
T KOG0775|consen 183 FKEKSRSLLREWYLQ--NPYPSPREKRELAEATGLT----------ITQVSNWFKNRRQRDRAAA 235 (304)
T ss_pred hhHhhHHHHHHHHhc--CCCCChHHHHHHHHHhCCc----------hhhhhhhhhhhhhhhhhcc
Confidence 444444678888886 7999999999999999976 4999999999999987433
No 29
>PF05920 Homeobox_KN: Homeobox KN domain; InterPro: IPR008422 This entry represents a homeobox transcription factor KN domain conserved from fungi to human and plants [].; GO: 0003677 DNA binding, 0006355 regulation of transcription, DNA-dependent, 0005634 nucleus; PDB: 3K2A_B 2LK2_A 1X2N_A 2DMN_A.
Probab=98.09 E-value=2.7e-06 Score=58.40 Aligned_cols=38 Identities=42% Similarity=0.551 Sum_probs=30.3
Q ss_pred HHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhh
Q 024359 26 LQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYA 73 (268)
Q Consensus 26 F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k 73 (268)
+++..+.||+.++++.||...|+| .+||.+||-|.|.+
T Consensus 3 ~~h~~nPYPs~~ek~~L~~~tgls----------~~Qi~~WF~NaRrR 40 (40)
T PF05920_consen 3 LEHLHNPYPSKEEKEELAKQTGLS----------RKQISNWFINARRR 40 (40)
T ss_dssp HHTTTSGS--HHHHHHHHHHHTS-----------HHHHHHHHHHHHHH
T ss_pred HHHCCCCCCCHHHHHHHHHHcCCC----------HHHHHHHHHHhHcc
Confidence 455668999999999999999976 49999999998853
No 30
>KOG2252 consensus CCAAT displacement protein and related homeoproteins [Transcription]
Probab=97.91 E-value=3.8e-05 Score=77.55 Aligned_cols=57 Identities=25% Similarity=0.414 Sum_probs=51.8
Q ss_pred CCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhh
Q 024359 6 SNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAI 74 (268)
Q Consensus 6 s~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~ 74 (268)
..++||+.||..|..-|-.+|++ +++|+++.++.|+..|||.. .-|.|||-|=|.+.
T Consensus 419 ~~KKPRlVfTd~QkrTL~aiFke--~~RPS~Emq~tIS~qL~L~~----------sTV~NfFmNaRRRs 475 (558)
T KOG2252|consen 419 QTKKPRLVFTDIQKRTLQAIFKE--NKRPSREMQETISQQLNLEL----------STVINFFMNARRRS 475 (558)
T ss_pred cCCCceeeecHHHHHHHHHHHhc--CCCCCHHHHHHHHHHhCCcH----------HHHHHHHHhhhhhc
Confidence 35889999999999999999999 69999999999999999875 88999999976664
No 31
>KOG0490 consensus Transcription factor, contains HOX domain [General function prediction only]
Probab=97.76 E-value=2.3e-05 Score=67.61 Aligned_cols=62 Identities=24% Similarity=0.383 Sum_probs=55.9
Q ss_pred CCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhccc
Q 024359 6 SNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSI 79 (268)
Q Consensus 6 s~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~ 79 (268)
..+++++.|+..|+..|+..|.. +.+|+...++.|+..+++++ ..|++||||+|.+.++...
T Consensus 152 ~~~~~~~~~~~~~~~~~~~~~~~--~~~P~~~~~~~l~~~~~~~~----------~~~q~~~~~~~~~~~~~~~ 213 (235)
T KOG0490|consen 152 KPRRPRTTFTENQLEVLETVFRA--TPKPDADDREQLAEETGLSE----------RVIQVWFQNRRAKLRKHKR 213 (235)
T ss_pred ccCCCccccccchhHhhhhcccC--CCCCchhhHHHHHHhcCCCh----------hhhhhhcccHHHHHHhhcc
Confidence 35788999999999999999999 58999999999999999764 8899999999999987764
No 32
>KOG0774 consensus Transcription factor PBX and related HOX domain proteins [Transcription]
Probab=97.18 E-value=0.00057 Score=64.58 Aligned_cols=64 Identities=23% Similarity=0.302 Sum_probs=54.6
Q ss_pred CCCCccccCHHHHHHHHHHHH-hccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcccC
Q 024359 7 NGGPAFRFNPAEVTEMEGILQ-EHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSIK 80 (268)
Q Consensus 7 ~~~pRt~FT~~Qv~eLEk~F~-~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~~ 80 (268)
.|+.|-.|++.-...|-+.|- +.+|.||+.+..++||.+-|++ -.||-+||-|+|.+.||-..+
T Consensus 188 arRKRRNFsK~aTeiLneyF~~h~~nPYPSee~K~eLAkqCnIt----------vsQvsnwfgnkrIrykK~~~k 252 (334)
T KOG0774|consen 188 ARRKRRNFSKQATEILNEYFYSHLSNPYPSEEAKEELAKQCNIT----------VSQVSNWFGNKRIRYKKNMGK 252 (334)
T ss_pred HHHhhcccchhHHHHHHHHHHHhcCCCCCcHHHHHHHHHHcCce----------ehhhccccccceeehhhhhhh
Confidence 477888999999999988765 5578999999999999999965 599999999999988776544
No 33
>KOG1168 consensus Transcription factor ACJ6/BRN-3, contains POU and HOX domains [Transcription]
Probab=96.64 E-value=0.00076 Score=64.55 Aligned_cols=63 Identities=22% Similarity=0.280 Sum_probs=55.9
Q ss_pred CCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcccC
Q 024359 6 SNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKSIK 80 (268)
Q Consensus 6 s~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~~~ 80 (268)
.++|.||....-|-..||..|..+ .-|+.+.+..||++|.|- -..|++||=|-|.|+|+...+
T Consensus 308 ekKRKRTSIAAPEKRsLEayFavQ--PRPS~EkIAaIAekLDLK----------KNVVRVWFCNQRQKQKRm~~S 370 (385)
T KOG1168|consen 308 EKKRKRTSIAAPEKRSLEAYFAVQ--PRPSGEKIAAIAEKLDLK----------KNVVRVWFCNQRQKQKRMKRS 370 (385)
T ss_pred ccccccccccCcccccHHHHhccC--CCCchhHHHHHHHhhhhh----------hceEEEEeeccHHHHHHhhhh
Confidence 358899999999999999999994 899999999999999964 477999999999999986643
No 34
>PF15057 DUF4537: Domain of unknown function (DUF4537)
Probab=95.64 E-value=0.077 Score=44.11 Aligned_cols=100 Identities=19% Similarity=0.241 Sum_probs=67.5
Q ss_pred eeccCCCceeehhhhhhcccccCCCCeEEEEecCCCCCccceeccccccccccccCcccccccccCCceEEEEeecCccc
Q 024359 147 AKSARDGAWYDVSAFLAQRNFDTADPEVQVRFAGFGAEEDEWVNIKRHVRQRSLPCEASECVAVLPGDLILCFQEGKDQA 226 (268)
Q Consensus 147 a~S~~D~AWYdv~~fl~~R~l~~ge~ev~Vrf~gFg~eedewvnv~~~vR~rS~ple~~eC~~v~~Gd~vlcf~e~~~~a 226 (268)
||+..||-+|-.. .++. + ....+.|.| ...+-+.+... -=|++.+..|+.|++||-||+-.+..+ .
T Consensus 6 AR~~~DG~YY~Gt-V~~~-~---~~~~~lV~f---~~~~~~~v~~~-----~iI~~~~~~~~~L~~GD~VLA~~~~~~-~ 71 (124)
T PF15057_consen 6 ARREEDGFYYPGT-VKKC-V---SSGQFLVEF---DDGDTQEVPIS-----DIIALSDAMRHSLQVGDKVLAPWEPDD-C 71 (124)
T ss_pred EeeCCCCcEEeEE-EEEc-c---CCCEEEEEE---CCCCEEEeChH-----HeEEccCcccCcCCCCCEEEEecCcCC-C
Confidence 7899999888753 2222 2 447899997 33344444443 235788889999999999999977664 5
Q ss_pred eeeeeEEEeeeeccCCCCcceeEEEEEEccCCcccccc
Q 024359 227 LYFDAHVLDAQRRRHDVRGCRCRFLVRYDHDQSEVATT 264 (268)
Q Consensus 227 ly~DA~V~~i~r~~Hd~~~C~C~F~Vr~~hd~sEe~v~ 264 (268)
.|+-|.|+..-.++ ....=.++|+|-.+. .+.||
T Consensus 72 ~Y~Pg~V~~~~~~~---~~~~~~~~V~f~ng~-~~~vp 105 (124)
T PF15057_consen 72 RYGPGTVIAGPERR---ASEDKEYTVRFYNGK-TAKVP 105 (124)
T ss_pred EEeCEEEEECcccc---ccCCceEEEEEECCC-CCccc
Confidence 59999999876555 333335666665444 33343
No 35
>PF11717 Tudor-knot: RNA binding activity-knot of a chromodomain ; PDB: 2EKO_A 2RO0_A 2RNZ_A 1WGS_A 3E9G_A 3E9F_A 2K3X_A 2K3Y_A 2EFI_A 2F5K_F ....
Probab=95.15 E-value=0.014 Score=41.91 Aligned_cols=40 Identities=33% Similarity=0.761 Sum_probs=32.1
Q ss_pred CCCceeehhhhhhcccccCCCCeEEEEecCCCCCccceeccc
Q 024359 151 RDGAWYDVSAFLAQRNFDTADPEVQVRFAGFGAEEDEWVNIK 192 (268)
Q Consensus 151 ~D~AWYdv~~fl~~R~l~~ge~ev~Vrf~gFg~eedewvnv~ 192 (268)
.+|.||...+.- -|. ..|+.+.+|||.|+..--||||...
T Consensus 12 ~~~~~y~A~I~~-~r~-~~~~~~YyVHY~g~nkR~DeWV~~~ 51 (55)
T PF11717_consen 12 KDGQWYEAKILD-IRE-KNGEPEYYVHYQGWNKRLDEWVPES 51 (55)
T ss_dssp TTTEEEEEEEEE-EEE-CTTCEEEEEEETTSTGCC-EEEETT
T ss_pred CCCcEEEEEEEE-EEe-cCCCEEEEEEcCCCCCCceeeecHH
Confidence 589999987543 333 6777999999999999999999875
No 36
>KOG1146 consensus Homeobox protein [General function prediction only]
Probab=94.58 E-value=0.028 Score=62.21 Aligned_cols=59 Identities=14% Similarity=0.191 Sum_probs=52.9
Q ss_pred CCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhc
Q 024359 7 NGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAK 77 (268)
Q Consensus 7 ~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk 77 (268)
.+..|+.|+..||..|-.+|... .|+.-++++.|-..++++. ..|+.||||-|.|.|+.
T Consensus 903 r~a~~~~~~d~qlk~i~~~~~~q--~~~~~~~~E~l~~~~~~~~----------~~i~vw~qna~~~s~k~ 961 (1406)
T KOG1146|consen 903 RRAYRTQESDLQLKIIKACYEAQ--RTPTMQECEVLEEPIGLPK----------RVIQVWFQNARAKSKKA 961 (1406)
T ss_pred hhhhccchhHHHHHHHHHHHhhc--cCChHHHHHhhcccccCCc----------chhHHhhhhhhhhhhhh
Confidence 37789999999999999999994 9999999999999999874 77899999999997754
No 37
>KOG0773 consensus Transcription factor MEIS1 and related HOX domain proteins [Transcription]
Probab=94.37 E-value=0.036 Score=51.96 Aligned_cols=58 Identities=28% Similarity=0.316 Sum_probs=47.7
Q ss_pred CCCCccccCHHHHHHHHHHHHh-ccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhh
Q 024359 7 NGGPAFRFNPAEVTEMEGILQE-HHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAI 74 (268)
Q Consensus 7 ~~~pRt~FT~~Qv~eLEk~F~~-~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~ 74 (268)
..++.-.|....+..|+.-+.+ ....||+......||.+.|++. .||.+||-|.|.+.
T Consensus 239 ~~r~~~~lP~~a~~ilr~Wl~~h~~~PYPse~~K~~La~~TGLs~----------~Qv~NWFINaR~R~ 297 (342)
T KOG0773|consen 239 KWRPQRGLPKEAVSILRAWLFEHLLHPYPSDDEKLMLAKQTGLSR----------PQVSNWFINARVRL 297 (342)
T ss_pred CCCCCCCCCHHHHHHHHHHHHHhccCCCCcchhccccchhcCCCc----------ccCCchhhhccccc
Confidence 4566678988888888876555 4447999999999999999765 99999999998774
No 38
>PF11569 Homez: Homeodomain leucine-zipper encoding, Homez; PDB: 2YS9_A.
Probab=93.34 E-value=0.074 Score=39.57 Aligned_cols=42 Identities=29% Similarity=0.403 Sum_probs=30.3
Q ss_pred HHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchh
Q 024359 19 VTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRY 72 (268)
Q Consensus 19 v~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~ 72 (268)
+.=|++.|..+ +.|.....+.|.++-++| ..||+.||--|+.
T Consensus 10 ~~pL~~Yy~~h--~~L~E~DL~~L~~kS~ms----------~qqVr~WFa~~~~ 51 (56)
T PF11569_consen 10 IQPLEDYYLKH--KQLQEEDLDELCDKSRMS----------YQQVRDWFAERMQ 51 (56)
T ss_dssp -HHHHHHHHHT------TTHHHHHHHHTT------------HHHHHHHHHHHS-
T ss_pred hHHHHHHHHHc--CCccHhhHHHHHHHHCCC----------HHHHHHHHHHhcc
Confidence 46699999986 899999999999999976 4999999987643
No 39
>PLN00104 MYST -like histone acetyltransferase; Provisional
Probab=92.92 E-value=0.33 Score=48.83 Aligned_cols=51 Identities=22% Similarity=0.391 Sum_probs=37.9
Q ss_pred eeccCCCceeehhhhhhcccc---cCCCCeEEEEecCCCCCccceecccccccccc
Q 024359 147 AKSARDGAWYDVSAFLAQRNF---DTADPEVQVRFAGFGAEEDEWVNIKRHVRQRS 199 (268)
Q Consensus 147 a~S~~D~AWYdv~~fl~~R~l---~~ge~ev~Vrf~gFg~eedewvnv~~~vR~rS 199 (268)
|+-..||.||. +..+.-|.. ..|+.+.+|||.||..--||||+.- +|...+
T Consensus 62 a~~~~Dg~~~~-A~VI~~R~~~~~~~~~~~YYVHY~g~nrRlDEWV~~~-rLdls~ 115 (450)
T PLN00104 62 CRWRFDGKYHP-VKVIERRRGGSGGPNDYEYYVHYTEFNRRLDEWVKLE-QLDLDT 115 (450)
T ss_pred EEECCCCCEEE-EEEEEEeccCCCCCCCceEEEEEecCCccHhhccCHh-hccccc
Confidence 44555999998 556666653 2355789999999999999999976 554444
No 40
>cd00024 CHROMO Chromatin organization modifier (chromo) domain is a conserved region of around 50 amino acids found in a variety of chromosomal proteins, which appear to play a role in the functional organization of the eukaryotic nucleus. Experimental evidence implicates the chromo domain in the binding activity of these proteins to methylated histone tails and maybe RNA. May occur as single instance, in a tandem arrangement or followd by a related "chromo shadow" domain.
Probab=92.64 E-value=0.073 Score=36.50 Aligned_cols=37 Identities=27% Similarity=0.548 Sum_probs=30.5
Q ss_pred eehhhhhhcccccC-CCCeEEEEecCCCCCccceeccc
Q 024359 156 YDVSAFLAQRNFDT-ADPEVQVRFAGFGAEEDEWVNIK 192 (268)
Q Consensus 156 Ydv~~fl~~R~l~~-ge~ev~Vrf~gFg~eedewvnv~ 192 (268)
|.|...|.+|.... |..++.|++.|++..+++|+...
T Consensus 3 ~~ve~Il~~r~~~~~~~~~y~VkW~g~~~~~~tWe~~~ 40 (55)
T cd00024 3 YEVEKILDHRKKKDGGEYEYLVKWKGYSYSEDTWEPEE 40 (55)
T ss_pred ceEeeeeeeeecCCCCcEEEEEEECCCCCccCccccHH
Confidence 44566778887765 77999999999999999998764
No 41
>cd04508 TUDOR Tudor domains are found in many eukaryotic organisms and have been implicated in protein-protein interactions in which methylated protein substrates bind to these domains. For example, the Tudor domain of Survival of Motor Neuron (SMN) binds to symmetrically dimethylated arginines of arginine-glycine (RG) rich sequences found in the C-terminal tails of Sm proteins. The SMN protein is linked to spinal muscular atrophy. Another example is the tandem tudor domains of 53BP1, which bind to histone H4 specifically dimethylated at Lys20 (H4-K20me2). 53BP1 is a key transducer of the DNA damage checkpoint signal.
Probab=90.79 E-value=0.23 Score=33.38 Aligned_cols=36 Identities=33% Similarity=0.530 Sum_probs=26.7
Q ss_pred EeeccCCCceeehhhhhhcccccCCCCeEEEEecCCCCCcc
Q 024359 146 EAKSARDGAWYDVSAFLAQRNFDTADPEVQVRFAGFGAEED 186 (268)
Q Consensus 146 Ea~S~~D~AWYdv~~fl~~R~l~~ge~ev~Vrf~gFg~eed 186 (268)
-|+...||.||.+.+.-- .++..+.|.|..||+.+.
T Consensus 5 ~a~~~~d~~wyra~V~~~-----~~~~~~~V~f~DyG~~~~ 40 (48)
T cd04508 5 LAKYSDDGKWYRAKITSI-----LSDGKVEVFFVDYGNTEV 40 (48)
T ss_pred EEEECCCCeEEEEEEEEE-----CCCCcEEEEEEcCCCcEE
Confidence 356677999999774321 126889999999999864
No 42
>PF04218 CENP-B_N: CENP-B N-terminal DNA-binding domain; InterPro: IPR006695 Centromere Protein B (CENP-B) is a DNA-binding protein localized to the centromere. Within the N-terminal 125 residues, there is a DNA-binding region, which binds to a corresponding 17bp CENP-B box sequence. CENP-B dimers either bind two separate DNA molecules or alternatively, they may bind two CENP-B boxes on one DNA molecule, with the intervening stretch of DNA forming a loop structure. The CENP-B DNA-binding domain consists of two repeating domains, RP1 and RP2. This family corresponds to RP1 has been shown to consist of four helices in a helix-turn-helix structure [].; GO: 0003677 DNA binding, 0000775 chromosome, centromeric region; PDB: 1BW6_A 1HLV_A 2ELH_A.
Probab=90.14 E-value=0.64 Score=33.25 Aligned_cols=46 Identities=20% Similarity=0.127 Sum_probs=35.1
Q ss_pred CCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcc
Q 024359 8 GGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNR 70 (268)
Q Consensus 8 ~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNR 70 (268)
+++|..+|..|-.++=+.++.. . -..+||..||++ ..+|..|..||
T Consensus 1 krkR~~LTl~eK~~iI~~~e~g--~-----s~~~ia~~fgv~----------~sTv~~I~K~k 46 (53)
T PF04218_consen 1 KRKRKSLTLEEKLEIIKRLEEG--E-----SKRDIAREFGVS----------RSTVSTILKNK 46 (53)
T ss_dssp SSSSSS--HHHHHHHHHHHHCT--T------HHHHHHHHT------------CCHHHHHHHCH
T ss_pred CCCCccCCHHHHHHHHHHHHcC--C-----CHHHHHHHhCCC----------HHHHHHHHHhH
Confidence 4678899999998888888772 3 588999999976 49999999996
No 43
>smart00298 CHROMO Chromatin organization modifier domain.
Probab=89.01 E-value=0.33 Score=33.03 Aligned_cols=37 Identities=27% Similarity=0.517 Sum_probs=30.5
Q ss_pred eehhhhhhcccccCCCCeEEEEecCCCCCccceeccc
Q 024359 156 YDVSAFLAQRNFDTADPEVQVRFAGFGAEEDEWVNIK 192 (268)
Q Consensus 156 Ydv~~fl~~R~l~~ge~ev~Vrf~gFg~eedewvnv~ 192 (268)
|.|.-.|.+|+...|..++.|+|.|+...++.|+...
T Consensus 2 ~~v~~Il~~r~~~~~~~~ylVkW~g~~~~~~tW~~~~ 38 (55)
T smart00298 2 YEVEKILDHRWKKKGELEYLVKWKGYSYSEDTWEPEE 38 (55)
T ss_pred cchheeeeeeecCCCcEEEEEEECCCCCccCceeeHH
Confidence 3466667787667788999999999999999999764
No 44
>smart00333 TUDOR Tudor domain. Domain of unknown function present in several RNA-binding proteins. 10 copies in the Drosophila Tudor protein. Initial proposal that the survival motor neuron gene product contain a Tudor domain are corroborated by more recent database search techniques such as PSI-BLAST (unpublished).
Probab=88.61 E-value=0.52 Score=32.72 Aligned_cols=43 Identities=28% Similarity=0.509 Sum_probs=30.5
Q ss_pred EEeeccCCCceeehhhhhhcccccCCCCeEEEEecCCCCCccceeccccccc
Q 024359 145 FEAKSARDGAWYDVSAFLAQRNFDTADPEVQVRFAGFGAEEDEWVNIKRHVR 196 (268)
Q Consensus 145 fEa~S~~D~AWYdv~~fl~~R~l~~ge~ev~Vrf~gFg~eedewvnv~~~vR 196 (268)
-.|+- .||.||.+.+ ++. .++..+.|.|..||+. +|++.. .+|
T Consensus 9 ~~a~~-~d~~wyra~I-~~~----~~~~~~~V~f~D~G~~--~~v~~~-~l~ 51 (57)
T smart00333 9 VAARW-EDGEWYRARI-IKV----DGEQLYEVFFIDYGNE--EVVPPS-DLR 51 (57)
T ss_pred EEEEe-CCCCEEEEEE-EEE----CCCCEEEEEEECCCcc--EEEeHH-Hee
Confidence 34556 7999999853 333 2338899999999998 488755 444
No 45
>PF00385 Chromo: Chromo (CHRromatin Organisation MOdifier) domain; InterPro: IPR023780 The CHROMO (CHRromatin Organization MOdifier) domain [, , , ] is a conserved region of around 60 amino acids, originally identified in Drosophila modifiers of variegation. These are proteins that alter the structure of chromatin to the condensed morphology of heterochromatin, a cytologically visible condition where gene expression is repressed. In one of these proteins, Polycomb, the chromo domain has been shown to be important for chromatin targeting. Proteins that contain a chromo domain appear to fall into 3 classes. The first class includes proteins having an N-terminal chromo domain followed by a region termed the chromo shadow domain, with weak but significant sequence similarity to the N-terminal chromo domain,[], eg. Drosophila and human heterochromatin protein Su(var)205 (HP1). The second class includes proteins with a single chromo domain, eg. Drosophila protein Polycomb (Pc); mammalian modifier 3; human Mi-2 autoantigen and several yeast and Caenorhabditis elegans hypothetical proteins. In the third class paired tandem chromo domains are found, eg. in mammalian DNA-binding/helicase proteins CHD-1 to CHD-4 and yeast protein CHD1. Functional dissections of chromo domain proteins suggests a mechanistic role for chromo domains in targeting chromo domain proteins to specific regions of the nucleus. The mechanism of targeting may involve protein-protein and/or protein/nucleic acid interactions. Hence, several line of evidence show that the HP1 chromo domain is a methyl-specific histone binding module, whereas the chromo domain of two protein components of the drosophila dosage compensation complex, MSL3 and MOF, contain chromo domains that bind to RNA in vitro []. The high resolution structures of HP1-family protein chromo and chromo shadow domain reveal a conserved chromo domain fold motif consisting of three beta strands packed against an alpha helix. The chromo domain fold belongs to the OB (oligonucleotide/oligosaccharide binding)-fold class found in a variety of prokaryotic and eukaryotic nucleic acid binding protein [].; PDB: 2H1E_B 3MWY_W 2DY8_A 1KNE_A 1KNA_A 1Q3L_A 2EE1_A 1AP0_A 1GUW_A 1X3P_A ....
Probab=88.59 E-value=0.16 Score=35.38 Aligned_cols=37 Identities=24% Similarity=0.495 Sum_probs=31.5
Q ss_pred eehhhhhhcccccCCC--CeEEEEecCCCCCccceeccc
Q 024359 156 YDVSAFLAQRNFDTAD--PEVQVRFAGFGAEEDEWVNIK 192 (268)
Q Consensus 156 Ydv~~fl~~R~l~~ge--~ev~Vrf~gFg~eedewvnv~ 192 (268)
|-|...|.||+...|. .+++|++.|++.+++.|.+..
T Consensus 1 ~~Ve~Il~~r~~~~~~~~~~ylVkW~g~~~~~~tWe~~~ 39 (55)
T PF00385_consen 1 YEVERILDHRVVKGGNKVYEYLVKWKGYPYSENTWEPEE 39 (55)
T ss_dssp EEEEEEEEEEEETTEESEEEEEEEETTSSGGGEEEEEGG
T ss_pred CEEEEEEEEEEeCCCcccEEEEEEECCCCCCCCeEeeHH
Confidence 5567788999777775 499999999999999999865
No 46
>PF02820 MBT: mbt repeat; InterPro: IPR004092 The function of the malignant brain tumor (MBT) repeat is unknown, but is found in a number of nuclear proteins involved in transcriptional repression. The repeat contains a completely conserved glutamate at its amino terminus that may be important for function. The crystal structure of the two MBT repeats of human SCM-like 2 protein has been reported. Each repeat consists of an extended "arm" and a globular core. The arm of the first repeat packs against the core of the second repeat and vice versa. The structure of the core-interacting part of each arm consists of an N-terminal alpha-helix and a turn of 310 helix connected by a short beta-strand. The core consists of an Src homology 3-like five-stranded beta-barrel followed by a C-terminal alpha-helix and another short beta-strand. Each arm interacts with its partner core in a similar way, with the orientation of the N-terminal helix relative to the barrel varying slightly. There are also extensive interactions between the two barrels [].; GO: 0006355 regulation of transcription, DNA-dependent, 0005634 nucleus; PDB: 2P0K_A 3F70_B 3CEY_A 2VYT_A 2BIV_A 1OI1_A 3OQ5_A 2RJE_B 2RJD_A 2RI3_A ....
Probab=88.56 E-value=0.33 Score=36.39 Aligned_cols=45 Identities=24% Similarity=0.461 Sum_probs=36.7
Q ss_pred ceEEeeccCCCceeehhhhhhcccccCCCCeEEEEecCCCCCccceeccc
Q 024359 143 MEFEAKSARDGAWYDVSAFLAQRNFDTADPEVQVRFAGFGAEEDEWVNIK 192 (268)
Q Consensus 143 ~efEa~S~~D~AWYdv~~fl~~R~l~~ge~ev~Vrf~gFg~eedewvnv~ 192 (268)
|-+||....+...+=|++...- -| ..|+|+|.|+.+++|.|+++.
T Consensus 1 MkLEa~d~~~~~~~~vAtV~~v----~g-~~l~v~~dg~~~~~d~w~~~~ 45 (73)
T PF02820_consen 1 MKLEAVDPRNPSLICVATVVKV----CG-GRLLVRYDGWDDDYDFWCHID 45 (73)
T ss_dssp EEEEEEETTECCEEEEEEEEEE----ET-TEEEEEETTSTGGGEEEEETT
T ss_pred CeEEEECCCCCCeEEEEEEEEE----eC-CEEEEEEcCCCCCccEEEECC
Confidence 5689999999888878776644 24 449999999999999999975
No 47
>smart00561 MBT Present in Drosophila Scm, l(3)mbt, and vertebrate SCML2. Present in Drosophila Scm, l(3)mbt, and vertebrate SCML2. These proteins are involved in transcriptional regulation.
Probab=87.54 E-value=0.62 Score=37.24 Aligned_cols=46 Identities=20% Similarity=0.363 Sum_probs=39.1
Q ss_pred cceEEeeccCCCceeehhhhhhcccccCCCCeEEEEecCCCCCccceeccc
Q 024359 142 FMEFEAKSARDGAWYDVSAFLAQRNFDTADPEVQVRFAGFGAEEDEWVNIK 192 (268)
Q Consensus 142 ~~efEa~S~~D~AWYdv~~fl~~R~l~~ge~ev~Vrf~gFg~eedewvnv~ 192 (268)
.|-+||...++-..+=|++...-. | ..|+|+|.|+.+..|.|+++.
T Consensus 31 GmkLEavD~~~~~~i~vAtV~~v~----g-~~l~v~~dg~~~~~D~W~~~~ 76 (96)
T smart00561 31 GMKLEAVDPRNPSLICVATVVEVK----G-YRLLLHFDGWDDKYDFWCDAD 76 (96)
T ss_pred CCEEEEECCCCCceEEEEEEEEEE----C-CEEEEEEccCCCcCCEEEECC
Confidence 578999999998888888766542 5 689999999999999999986
No 48
>PF12824 MRP-L20: Mitochondrial ribosomal protein subunit L20; InterPro: IPR024388 Ribosomes are the particles that catalyse mRNA-directed protein synthesis in all organisms. The codons of the mRNA are exposed on the ribosome to allow tRNA binding. This leads to the incorporation of amino acids into the growing polypeptide chain in accordance with the genetic information. Incoming amino acid monomers enter the ribosomal A site in the form of aminoacyl-tRNAs complexed with elongation factor Tu (EF-Tu) and GTP. The growing polypeptide chain, situated in the P site as peptidyl-tRNA, is then transferred to aminoacyl-tRNA and the new peptidyl-tRNA, extended by one residue, is translocated to the P site with the aid the elongation factor G (EF-G) and GTP as the deacylated tRNA is released from the ribosome through one or more exit sites [, ]. About 2/3 of the mass of the ribosome consists of RNA and 1/3 of protein. The proteins are named in accordance with the subunit of the ribosome which they belong to - the small (S1 to S31) and the large (L1 to L44). Usually they decorate the rRNA cores of the subunits. Many ribosomal proteins, particularly those of the large subunit, are composed of a globular, surfaced-exposed domain with long finger-like projections that extend into the rRNA core to stabilise its structure. Most of the proteins interact with multiple RNA elements, often from different domains. In the large subunit, about 1/3 of the 23S rRNA nucleotides are at least in van der Waal's contact with protein, and L22 interacts with all six domains of the 23S rRNA. Proteins S4 and S7, which initiate assembly of the 16S rRNA, are located at junctions of five and four RNA helices, respectively. In this way proteins serve to organise and stabilise the rRNA tertiary structure. While the crucial activities of decoding and peptide transfer are RNA based, proteins play an active role in functions that may have evolved to streamline the process of protein synthesis. In addition to their function in the ribosome, many ribosomal proteins have some function 'outside' the ribosome [, ]. This entry represents the essential mitochondrial ribosomal protein L20 family from fungi [].
Probab=86.62 E-value=1.4 Score=38.72 Aligned_cols=56 Identities=20% Similarity=0.254 Sum_probs=42.2
Q ss_pred CccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhc
Q 024359 10 PAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQN 69 (268)
Q Consensus 10 pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQN 69 (268)
....+|+++|+||-++=.+ -|..-.+.+||++||+|+-..+-+.=...|-+.+-+.
T Consensus 82 k~y~Lt~e~i~Eir~LR~~----DP~~wTr~~LAkkF~~S~~fV~~v~~~~~e~~~~~~~ 137 (164)
T PF12824_consen 82 KKYHLTPEDIQEIRRLRAE----DPEKWTRKKLAKKFNCSPLFVSMVAPAPKEKKKEMEA 137 (164)
T ss_pred ccccCCHHHHHHHHHHHHc----CchHhhHHHHHHHhCCCHHHHHHhcCCCHHHHHHHHH
Confidence 4578999999999988776 4777899999999999986666555445454444333
No 49
>PF05641 Agenet: Agenet domain; InterPro: IPR008395 This domain is related to the TUDOR domain IPR008191 from INTERPRO []. The function of the agenet domain is unknown. This signature matches one of the two Agenet domains in the FMR proteins [].; GO: 0003723 RNA binding; PDB: 2BKD_N 3O8V_A 3KUF_A 3H8Z_A.
Probab=85.90 E-value=2.4 Score=31.43 Aligned_cols=41 Identities=27% Similarity=0.384 Sum_probs=27.0
Q ss_pred ccCCceEE-EEeecCccceeeeeEEEeeeeccCCCCcceeEEEEEEccCC
Q 024359 210 VLPGDLIL-CFQEGKDQALYFDAHVLDAQRRRHDVRGCRCRFLVRYDHDQ 258 (268)
Q Consensus 210 v~~Gd~vl-cf~e~~~~aly~DA~V~~i~r~~Hd~~~C~C~F~Vr~~hd~ 258 (268)
+++|+.|- +.++.+-...||-|.|+++.... +++|+|++=.
T Consensus 1 F~~G~~VEV~s~e~g~~gaWf~a~V~~~~~~~--------~~~V~Y~~~~ 42 (68)
T PF05641_consen 1 FKKGDEVEVSSDEDGFRGAWFPATVLKENGDD--------KYLVEYDDLP 42 (68)
T ss_dssp --TT-EEEEEE-SBTT--EEEEEEEEEEETT---------EEEEEETT-S
T ss_pred CCCCCEEEEEEcCCCCCcEEEEEEEEEeCCCc--------EEEEEECCcc
Confidence 36899995 55566668999999999987655 8999996533
No 50
>smart00333 TUDOR Tudor domain. Domain of unknown function present in several RNA-binding proteins. 10 copies in the Drosophila Tudor protein. Initial proposal that the survival motor neuron gene product contain a Tudor domain are corroborated by more recent database search techniques such as PSI-BLAST (unpublished).
Probab=85.85 E-value=2.5 Score=29.19 Aligned_cols=45 Identities=11% Similarity=0.158 Sum_probs=32.0
Q ss_pred cccCCceEEEEeecCccceeeeeEEEeeeeccCCCCcceeEEEEEEccCCcccccc
Q 024359 209 AVLPGDLILCFQEGKDQALYFDAHVLDAQRRRHDVRGCRCRFLVRYDHDQSEVATT 264 (268)
Q Consensus 209 ~v~~Gd~vlcf~e~~~~aly~DA~V~~i~r~~Hd~~~C~C~F~Vr~~hd~sEe~v~ 264 (268)
..++|+.+++.. ++..||.|+|+++... -.+.|.|....+++-|+
T Consensus 2 ~~~~G~~~~a~~---~d~~wyra~I~~~~~~--------~~~~V~f~D~G~~~~v~ 46 (57)
T smart00333 2 TFKVGDKVAARW---EDGEWYRARIIKVDGE--------QLYEVFFIDYGNEEVVP 46 (57)
T ss_pred CCCCCCEEEEEe---CCCCEEEEEEEEECCC--------CEEEEEEECCCccEEEe
Confidence 357898888776 2588999999999742 23567787755555544
No 51
>PF12148 DUF3590: Protein of unknown function (DUF3590); InterPro: IPR021991 This domain is found in eukaryotes, and is typically between 83 and 97 amino acids in length. It is found in association with PF00097 from PFAM, PF02182 from PFAM, PF00628 from PFAM, PF00240 from PFAM. There are two conserved sequence motifs: RAR and NYN. The domain is part of the protein NIRF which has zinc finger and ubiquitinating domains. The function of this domain is likely to be mainly structural, however this has not been confirmed. ; PDB: 3DB4_A 3ASK_A 3DB3_A 2L3R_A.
Probab=84.09 E-value=1.3 Score=35.43 Aligned_cols=70 Identities=13% Similarity=0.288 Sum_probs=41.1
Q ss_pred EeeccCCCceeehhhhhhccccc--CCCCeEEEEecCCCCCccceeccccccccccccCcccccccccCCceEEE
Q 024359 146 EAKSARDGAWYDVSAFLAQRNFD--TADPEVQVRFAGFGAEEDEWVNIKRHVRQRSLPCEASECVAVLPGDLILC 218 (268)
Q Consensus 146 Ea~S~~D~AWYdv~~fl~~R~l~--~ge~ev~Vrf~gFg~eedewvnv~~~vR~rS~ple~~eC~~v~~Gd~vlc 218 (268)
-||+...|||++....-.++--. ..+.-..|.|.+|.+..-.=+.++ .||+|..-+= .=..|.+|+.|..
T Consensus 3 D~~d~~~gAWfEa~i~~i~~~~~~~~e~viYhIkyddype~gvv~~~~~-~iRpRARt~l--~w~~L~VG~~VMv 74 (85)
T PF12148_consen 3 DARDRNMGAWFEAQIVTITKKCMSDDEDVIYHIKYDDYPENGVVEMRSK-DIRPRARTIL--KWDELKVGQVVMV 74 (85)
T ss_dssp EEE-TTT-EEEEEEEEEEEES-SSSSTTEEEEEEETT-GGG-EEEEEGG-GEEE---SBE---GGG--TT-EEEE
T ss_pred ccccCCCcceEEEEEEEeeccCCCCCCCEEEEEEeccCCCcCceecccc-cccceeeEec--cHHhCCcccEEEE
Confidence 37888899999977655554332 235667899999987776666666 8888876543 3457889999874
No 52
>PF11717 Tudor-knot: RNA binding activity-knot of a chromodomain ; PDB: 2EKO_A 2RO0_A 2RNZ_A 1WGS_A 3E9G_A 3E9F_A 2K3X_A 2K3Y_A 2EFI_A 2F5K_F ....
Probab=83.23 E-value=5.8 Score=28.34 Aligned_cols=40 Identities=20% Similarity=0.478 Sum_probs=28.8
Q ss_pred ccCCceEEEEeecCccceeeeeEEEeeeeccCCCCcceeEEEEEEccC
Q 024359 210 VLPGDLILCFQEGKDQALYFDAHVLDAQRRRHDVRGCRCRFLVRYDHD 257 (268)
Q Consensus 210 v~~Gd~vlcf~e~~~~aly~DA~V~~i~r~~Hd~~~C~C~F~Vr~~hd 257 (268)
+..|+.|.|.. .+..+|.|+|++|..+... =.|.|-|.--
T Consensus 1 ~~vG~~v~~~~---~~~~~y~A~I~~~r~~~~~-----~~YyVHY~g~ 40 (55)
T PF11717_consen 1 FEVGEKVLCKY---KDGQWYEAKILDIREKNGE-----PEYYVHYQGW 40 (55)
T ss_dssp --TTEEEEEEE---TTTEEEEEEEEEEEECTTC-----EEEEEEETTS
T ss_pred CCcCCEEEEEE---CCCcEEEEEEEEEEecCCC-----EEEEEEcCCC
Confidence 46899999999 3578999999999885433 3466666543
No 53
>smart00743 Agenet Tudor-like domain present in plant sequences. Domain in plant sequences with possible chromatin-associated functions.
Probab=82.33 E-value=4.8 Score=28.55 Aligned_cols=46 Identities=22% Similarity=0.395 Sum_probs=34.8
Q ss_pred cccCCceEEEEeecCccceeeeeEEEeeeeccCCCCcceeEEEEEEcc--CCcccccc
Q 024359 209 AVLPGDLILCFQEGKDQALYFDAHVLDAQRRRHDVRGCRCRFLVRYDH--DQSEVATT 264 (268)
Q Consensus 209 ~v~~Gd~vlcf~e~~~~aly~DA~V~~i~r~~Hd~~~C~C~F~Vr~~h--d~sEe~v~ 264 (268)
..+.||.|-++... +.-||-|.|+++.. .. +|.|+|+. ...++.|+
T Consensus 2 ~~~~G~~Ve~~~~~--~~~W~~a~V~~~~~-----~~---~~~V~~~~~~~~~~e~v~ 49 (61)
T smart00743 2 DFKKGDRVEVFSKE--EDSWWEAVVTKVLG-----DG---KYLVRYLTESEPLKETVD 49 (61)
T ss_pred CcCCCCEEEEEECC--CCEEEEEEEEEECC-----CC---EEEEEECCCCcccEEEEe
Confidence 46799999988864 57899999998875 22 37999988 55555554
No 54
>KOG3623 consensus Homeobox transcription factor SIP1 [Transcription]
Probab=80.46 E-value=7.2 Score=42.18 Aligned_cols=52 Identities=21% Similarity=0.319 Sum_probs=39.5
Q ss_pred cCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcc
Q 024359 14 FNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKS 78 (268)
Q Consensus 14 FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~ 78 (268)
|++. +.-|...|.- |..|+.++..++|...|++- .-|+.||+|++.+.....
T Consensus 564 ~~~p-~sllkayyal--n~~ps~eelskia~qvglp~----------~vvk~wfE~~~a~e~sv~ 615 (1007)
T KOG3623|consen 564 FNHP-TSLLKAYYAL--NGLPSEEELSKIAQQVGLPF----------AVVKAWFEDEEAEEMSVE 615 (1007)
T ss_pred cCCc-HHHHHHHHHh--cCCCCHHHHHHHHHHhcccH----------HHHHHHHHhhhhhhhhhc
Confidence 4444 3444455555 69999999999999999863 669999999999865443
No 55
>PF05641 Agenet: Agenet domain; InterPro: IPR008395 This domain is related to the TUDOR domain IPR008191 from INTERPRO []. The function of the agenet domain is unknown. This signature matches one of the two Agenet domains in the FMR proteins [].; GO: 0003723 RNA binding; PDB: 2BKD_N 3O8V_A 3KUF_A 3H8Z_A.
Probab=74.24 E-value=3.2 Score=30.75 Aligned_cols=40 Identities=25% Similarity=0.678 Sum_probs=22.7
Q ss_pred CCceeehhhhhhcccccCCCCeEEEEecCCCCCcc------ceecccccccc
Q 024359 152 DGAWYDVSAFLAQRNFDTADPEVQVRFAGFGAEED------EWVNIKRHVRQ 197 (268)
Q Consensus 152 D~AWYdv~~fl~~R~l~~ge~ev~Vrf~gFg~eed------ewvnv~~~vR~ 197 (268)
.||||.+.+.-.. ++..+.|+|..+..+++ |||+.+ ++|+
T Consensus 17 ~gaWf~a~V~~~~-----~~~~~~V~Y~~~~~~~~~~~~l~e~V~~~-~iRP 62 (68)
T PF05641_consen 17 RGAWFPATVLKEN-----GDDKYLVEYDDLPDEDGESPPLKEWVDAR-RIRP 62 (68)
T ss_dssp --EEEEEEEEEEE-----TT-EEEEEETT-SS--------EEEEEGG-GEEE
T ss_pred CcEEEEEEEEEeC-----CCcEEEEEECCcccccccccccEEEechh-eEEC
Confidence 6799998744322 22399999998888854 566665 4554
No 56
>PF02796 HTH_7: Helix-turn-helix domain of resolvase; InterPro: IPR006120 Site-specific recombination plays an important role in DNA rearrangement in prokaryotic organisms. Two types of site-specific recombination are known to occur: Recombination between inverted repeats resulting in the reversal of a DNA segment. Recombination between repeat sequences on two DNA molecules resulting in their cointegration, or between repeats on one DNA molecule resulting in the excision of a DNA fragment. Site-specific recombination is characterised by a strand exchange mechanism that requires no DNA synthesis or high energy cofactor; the phosphodiester bond energy is conserved in a phospho-protein linkage during strand cleavage and re-ligation. Two unrelated families of recombinases are currently known []. The first, called the 'phage integrase' family, groups a number of bacterial phage and yeast plasmid enzymes. The second [], called the 'resolvase' family, groups enzymes which share the following structural characteristics: an N-terminal catalytic and dimerization domain that contains a conserved serine residue involved in the transient covalent attachment to DNA IPR006119 from INTERPRO, and a C-terminal helix-turn-helix DNA-binding domain. ; GO: 0000150 recombinase activity, 0003677 DNA binding, 0006310 DNA recombination; PDB: 1ZR2_A 2GM4_B 1RES_A 1ZR4_A 1RET_A 1GDT_B 2R0Q_C 1JKP_C 1IJW_C 1JJ6_C ....
Probab=71.85 E-value=4 Score=27.88 Aligned_cols=35 Identities=26% Similarity=0.577 Sum_probs=25.2
Q ss_pred CCCCCCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCc
Q 024359 2 GRPPSNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESP 50 (268)
Q Consensus 2 GrPps~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~ 50 (268)
||||. +++++++++-+++.+ + ....+||+.||+|.
T Consensus 1 GRp~~-------~~~~~~~~i~~l~~~--G-----~si~~IA~~~gvsr 35 (45)
T PF02796_consen 1 GRPPK-------LSKEQIEEIKELYAE--G-----MSIAEIAKQFGVSR 35 (45)
T ss_dssp SSSSS-------SSHCCHHHHHHHHHT--T-------HHHHHHHTTS-H
T ss_pred CcCCC-------CCHHHHHHHHHHHHC--C-----CCHHHHHHHHCcCH
Confidence 67765 566678888888887 2 34789999999874
No 57
>PF06003 SMN: Survival motor neuron protein (SMN); InterPro: IPR010304 This family consists of several eukaryotic survival motor neuron (SMN) proteins. The Survival of Motor Neurons (SMN) protein, the product of the spinal muscular atrophy-determining gene, is part of a large macromolecular complex (SMN complex) that functions in the assembly of spliceosomal small nuclear ribonucleoproteins (snRNPs). The SMN complex functions as a specificity factor essential for the efficient assembly of Sm proteins on U snRNAs and likely protects cells from illicit, and potentially deleterious, non-specific binding of Sm proteins to RNAs.; GO: 0003723 RNA binding, 0006397 mRNA processing, 0005634 nucleus, 0005737 cytoplasm; PDB: 1MHN_A 4A4G_A 3S6N_M 4A4E_A 1G5V_A 4A4H_A 4A4F_A 2D9T_A.
Probab=71.64 E-value=2.8 Score=39.08 Aligned_cols=42 Identities=26% Similarity=0.400 Sum_probs=29.4
Q ss_pred EEeeccCCCceeehhhhhhcccccCCCCeEEEEecCCCCCccceec
Q 024359 145 FEAKSARDGAWYDVSAFLAQRNFDTADPEVQVRFAGFGAEEDEWVN 190 (268)
Q Consensus 145 fEa~S~~D~AWYdv~~fl~~R~l~~ge~ev~Vrf~gFg~eedewvn 190 (268)
=.|.-+.||-||...+---+ ...+.+.|+|.|||+.|+.++.
T Consensus 75 C~A~~s~Dg~~Y~A~I~~i~----~~~~~~~V~f~gYgn~e~v~l~ 116 (264)
T PF06003_consen 75 CMAVYSEDGQYYPATIESID----EEDGTCVVVFTGYGNEEEVNLS 116 (264)
T ss_dssp EEEE-TTTSSEEEEEEEEEE----TTTTEEEEEETTTTEEEEEEGG
T ss_pred EEEEECCCCCEEEEEEEEEc----CCCCEEEEEEcccCCeEeeehh
Confidence 35666889999998754422 2235788999999999865543
No 58
>PF13565 HTH_32: Homeodomain-like domain
Probab=71.03 E-value=15 Score=26.63 Aligned_cols=53 Identities=28% Similarity=0.429 Sum_probs=36.4
Q ss_pred CCCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhh
Q 024359 5 PSNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNW 66 (268)
Q Consensus 5 ps~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~W 66 (268)
|..++|+. +.++.+.|.+++.++ ...-.....+.|++.||.+. .++...|..|
T Consensus 24 ~~~Grp~~--~~e~~~~i~~~~~~~-p~wt~~~i~~~L~~~~g~~~------~~S~~tv~R~ 76 (77)
T PF13565_consen 24 PRPGRPRK--DPEQRERIIALIEEH-PRWTPREIAEYLEEEFGISV------RVSRSTVYRI 76 (77)
T ss_pred CCCCCCCC--cHHHHHHHHHHHHhC-CCCCHHHHHHHHHHHhCCCC------CccHhHHHHh
Confidence 55677766 777779999999984 23445577888999988541 2355666554
No 59
>smart00743 Agenet Tudor-like domain present in plant sequences. Domain in plant sequences with possible chromatin-associated functions.
Probab=68.16 E-value=5.3 Score=28.30 Aligned_cols=36 Identities=19% Similarity=0.311 Sum_probs=24.8
Q ss_pred EEeeccCCCceeehhhhhhcccccCCCCeEEEEecC--CCCCc
Q 024359 145 FEAKSARDGAWYDVSAFLAQRNFDTADPEVQVRFAG--FGAEE 185 (268)
Q Consensus 145 fEa~S~~D~AWYdv~~fl~~R~l~~ge~ev~Vrf~g--Fg~ee 185 (268)
-||++..||+||...+.- ++ ++..+.|+|.+ +|+.+
T Consensus 9 Ve~~~~~~~~W~~a~V~~---~~--~~~~~~V~~~~~~~~~~e 46 (61)
T smart00743 9 VEVFSKEEDSWWEAVVTK---VL--GDGKYLVRYLTESEPLKE 46 (61)
T ss_pred EEEEECCCCEEEEEEEEE---EC--CCCEEEEEECCCCcccEE
Confidence 356666699999876542 22 24679999999 65444
No 60
>cd04508 TUDOR Tudor domains are found in many eukaryotic organisms and have been implicated in protein-protein interactions in which methylated protein substrates bind to these domains. For example, the Tudor domain of Survival of Motor Neuron (SMN) binds to symmetrically dimethylated arginines of arginine-glycine (RG) rich sequences found in the C-terminal tails of Sm proteins. The SMN protein is linked to spinal muscular atrophy. Another example is the tandem tudor domains of 53BP1, which bind to histone H4 specifically dimethylated at Lys20 (H4-K20me2). 53BP1 is a key transducer of the DNA damage checkpoint signal.
Probab=62.68 E-value=24 Score=23.38 Aligned_cols=39 Identities=23% Similarity=0.337 Sum_probs=25.3
Q ss_pred CceEEEEeecCccceeeeeEEEeeeeccCCCCcceeEEEEEEcc-CCccc
Q 024359 213 GDLILCFQEGKDQALYFDAHVLDAQRRRHDVRGCRCRFLVRYDH-DQSEV 261 (268)
Q Consensus 213 Gd~vlcf~e~~~~aly~DA~V~~i~r~~Hd~~~C~C~F~Vr~~h-d~sEe 261 (268)
|+++++.-. ++..||-|.|+++.. .-.+.|.|.. +++|.
T Consensus 1 G~~c~a~~~--~d~~wyra~V~~~~~--------~~~~~V~f~DyG~~~~ 40 (48)
T cd04508 1 GDLCLAKYS--DDGKWYRAKITSILS--------DGKVEVFFVDYGNTEV 40 (48)
T ss_pred CCEEEEEEC--CCCeEEEEEEEEECC--------CCcEEEEEEcCCCcEE
Confidence 455554433 358999999999974 2235677766 55543
No 61
>PF00567 TUDOR: Tudor domain; InterPro: IPR008191 There are multiple copies of this domain in the Drosophila melanogaster tudor protein and it has been identified in several RNA-binding proteins []. Although the function of this domain is unknown, in Drosophila melanogaster the tudor protein is required during oogenesis for the formation of primordial germ cells and for normal abdominal segmentation [].; PDB: 3NTI_A 3NTK_B 3NTH_A 2DIQ_A 3FDR_A 3PNW_O 3S6W_A 3PMT_A 2WAC_A 2O4X_A ....
Probab=56.96 E-value=10 Score=28.54 Aligned_cols=51 Identities=29% Similarity=0.585 Sum_probs=35.1
Q ss_pred ccCCCceeehhhhhhcccccCCCCeEEEEecCCCCCccceeccccccc-----cccccCccccc
Q 024359 149 SARDGAWYDVSAFLAQRNFDTADPEVQVRFAGFGAEEDEWVNIKRHVR-----QRSLPCEASEC 207 (268)
Q Consensus 149 S~~D~AWYdv~~fl~~R~l~~ge~ev~Vrf~gFg~eedewvnv~~~vR-----~rS~ple~~eC 207 (268)
...||.||-+.+ ....++..+.|.|..||..+- ++.. .+| ...+|.++..|
T Consensus 62 ~~~~~~w~Ra~I-----~~~~~~~~~~V~~iD~G~~~~--v~~~-~l~~l~~~~~~~P~~a~~~ 117 (121)
T PF00567_consen 62 VSEDGRWYRAVI-----TVDIDENQYKVFLIDYGNTEK--VSAS-DLRPLPPEFASLPPQAIKC 117 (121)
T ss_dssp ETTTSEEEEEEE-----EEEECTTEEEEEETTTTEEEE--EEGG-GEEE--HHHCSSSSSCEEE
T ss_pred EecCCceeeEEE-----EEecccceeEEEEEecCceEE--EcHH-HhhhhCHHHhhCChhhEEE
Confidence 467999999886 245677999999999998764 5544 222 23355555555
No 62
>PF04967 HTH_10: HTH DNA binding domain; InterPro: IPR007050 Numerous bacterial transcription regulatory proteins bind DNA via a helix-turn-helix (HTH) motif. This entry represents the HTH DNA binding domain found in Halobacterium salinarium (Halobacterium halobium) and described as a putative bacterio-opsin activator.
Probab=54.99 E-value=16 Score=26.55 Aligned_cols=34 Identities=18% Similarity=0.136 Sum_probs=29.0
Q ss_pred cCHHHHHHHHHHHHhccCCCCC---HHHHHHHHHHhCCCc
Q 024359 14 FNPAEVTEMEGILQEHHNAMPS---REILVALAEKFSESP 50 (268)
Q Consensus 14 FT~~Qv~eLEk~F~~~~~~yp~---~~~rq~LA~~fnlS~ 50 (268)
+|+.|...|...++. .|.+ .....+||+.||+|.
T Consensus 1 LT~~Q~e~L~~A~~~---GYfd~PR~~tl~elA~~lgis~ 37 (53)
T PF04967_consen 1 LTDRQREILKAAYEL---GYFDVPRRITLEELAEELGISK 37 (53)
T ss_pred CCHHHHHHHHHHHHc---CCCCCCCcCCHHHHHHHhCCCH
Confidence 589999999999998 5544 577899999999984
No 63
>cd00569 HTH_Hin_like Helix-turn-helix domain of Hin and related proteins, a family of DNA-binding domains unique to bacteria and represented by the Hin protein of Salmonella. The basic HTH domain is a simple fold comprised of three core helices that form a right-handed helical bundle. The principal DNA-protein interface is formed by the third helix, the recognition helix, inserting itself into the major groove of the DNA. A diverse array of HTH domains participate in a variety of functions that depend on their DNA-binding properties. HTH_Hin represents one of the simplest versions of the HTH domains; the characterization of homologous relationships between various sequence-diverse HTH domain families remains difficult. The Hin recombinase induces the site-specific inversion of a chromosomal DNA segment containing a promoter, which controls the alternate expression of two genes by reversibly switching orientation. The Hin recombinase consists of a single polypeptide chain containing a D
Probab=54.50 E-value=34 Score=19.17 Aligned_cols=31 Identities=16% Similarity=0.379 Sum_probs=21.7
Q ss_pred ccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCc
Q 024359 13 RFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESP 50 (268)
Q Consensus 13 ~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~ 50 (268)
.|+..+...+...+.. .+ ...++|+.|+++.
T Consensus 5 ~~~~~~~~~i~~~~~~---~~----s~~~ia~~~~is~ 35 (42)
T cd00569 5 KLTPEQIEEARRLLAA---GE----SVAEIARRLGVSR 35 (42)
T ss_pred cCCHHHHHHHHHHHHc---CC----CHHHHHHHHCCCH
Confidence 3677777777777654 22 4678999999764
No 64
>PF06003 SMN: Survival motor neuron protein (SMN); InterPro: IPR010304 This family consists of several eukaryotic survival motor neuron (SMN) proteins. The Survival of Motor Neurons (SMN) protein, the product of the spinal muscular atrophy-determining gene, is part of a large macromolecular complex (SMN complex) that functions in the assembly of spliceosomal small nuclear ribonucleoproteins (snRNPs). The SMN complex functions as a specificity factor essential for the efficient assembly of Sm proteins on U snRNAs and likely protects cells from illicit, and potentially deleterious, non-specific binding of Sm proteins to RNAs.; GO: 0003723 RNA binding, 0006397 mRNA processing, 0005634 nucleus, 0005737 cytoplasm; PDB: 1MHN_A 4A4G_A 3S6N_M 4A4E_A 1G5V_A 4A4H_A 4A4F_A 2D9T_A.
Probab=54.21 E-value=23 Score=33.01 Aligned_cols=48 Identities=15% Similarity=0.227 Sum_probs=30.6
Q ss_pred ccccCCceEEEEeecCccceeeeeEEEeeeeccCCCCcceeEEEEEEccCCcccccc
Q 024359 208 VAVLPGDLILCFQEGKDQALYFDAHVLDAQRRRHDVRGCRCRFLVRYDHDQSEVATT 264 (268)
Q Consensus 208 ~~v~~Gd~vlcf~e~~~~aly~DA~V~~i~r~~Hd~~~C~C~F~Vr~~hd~sEe~v~ 264 (268)
..-++||..++.- .++.+||.|.|.+|... ...| +|+|+.-+.+|.|.
T Consensus 67 ~~WkvGd~C~A~~--s~Dg~~Y~A~I~~i~~~---~~~~----~V~f~gYgn~e~v~ 114 (264)
T PF06003_consen 67 KKWKVGDKCMAVY--SEDGQYYPATIESIDEE---DGTC----VVVFTGYGNEEEVN 114 (264)
T ss_dssp T---TT-EEEEE---TTTSSEEEEEEEEEETT---TTEE----EEEETTTTEEEEEE
T ss_pred cCCCCCCEEEEEE--CCCCCEEEEEEEEEcCC---CCEE----EEEEcccCCeEeee
Confidence 4688999988874 33468999999999521 2334 49998877666654
No 65
>PF13551 HTH_29: Winged helix-turn helix
Probab=53.31 E-value=45 Score=25.25 Aligned_cols=53 Identities=19% Similarity=0.203 Sum_probs=33.4
Q ss_pred CCCccccCHHHHHHHHHHHHhccC---CCCCHHH-HHHH-HHHhCCCccccCCcccccchhhhhhh
Q 024359 8 GGPAFRFNPAEVTEMEGILQEHHN---AMPSREI-LVAL-AEKFSESPERKGKIMVQMKQVWNWFQ 68 (268)
Q Consensus 8 ~~pRt~FT~~Qv~eLEk~F~~~~~---~yp~~~~-rq~L-A~~fnlS~~RaGK~~lt~kQVk~WFQ 68 (268)
++++..+|+++.+.|.+.+.+... ...+... .+.| .+.+++ .++...|+.|++
T Consensus 52 g~~~~~l~~~~~~~l~~~~~~~p~~g~~~~t~~~l~~~l~~~~~~~--------~~s~~ti~r~L~ 109 (112)
T PF13551_consen 52 GRPRKRLSEEQRAQLIELLRENPPEGRSRWTLEELAEWLIEEEFGI--------DVSPSTIRRILK 109 (112)
T ss_pred CCCCCCCCHHHHHHHHHHHHHCCCCCCCcccHHHHHHHHHHhccCc--------cCCHHHHHHHHH
Confidence 455555999999999999998421 0233333 3335 444454 367788888875
No 66
>PF12148 DUF3590: Protein of unknown function (DUF3590); InterPro: IPR021991 This domain is found in eukaryotes, and is typically between 83 and 97 amino acids in length. It is found in association with PF00097 from PFAM, PF02182 from PFAM, PF00628 from PFAM, PF00240 from PFAM. There are two conserved sequence motifs: RAR and NYN. The domain is part of the protein NIRF which has zinc finger and ubiquitinating domains. The function of this domain is likely to be mainly structural, however this has not been confirmed. ; PDB: 3DB4_A 3ASK_A 3DB3_A 2L3R_A.
Probab=52.53 E-value=31 Score=27.68 Aligned_cols=36 Identities=11% Similarity=0.328 Sum_probs=26.7
Q ss_pred ccceeeeeEEEeeeeccCCCCcceeEEEEEEccCCcc
Q 024359 224 DQALYFDAHVLDAQRRRHDVRGCRCRFLVRYDHDQSE 260 (268)
Q Consensus 224 ~~aly~DA~V~~i~r~~Hd~~~C~C~F~Vr~~hd~sE 260 (268)
....||+|.|+.|.++- ....+.+.+-|.|+.....
T Consensus 8 ~~gAWfEa~i~~i~~~~-~~~~e~viYhIkyddype~ 43 (85)
T PF12148_consen 8 NMGAWFEAQIVTITKKC-MSDDEDVIYHIKYDDYPEN 43 (85)
T ss_dssp TT-EEEEEEEEEEEES--SSSSTTEEEEEEETT-GGG
T ss_pred CCcceEEEEEEEeeccC-CCCCCCEEEEEEeccCCCc
Confidence 34789999999999664 4445999999999977644
No 67
>PF13873 Myb_DNA-bind_5: Myb/SANT-like DNA-binding domain
Probab=50.38 E-value=27 Score=25.77 Aligned_cols=62 Identities=16% Similarity=0.198 Sum_probs=42.6
Q ss_pred cccCHHHHHHHHHHHHhccCCCCC-----------HHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhc
Q 024359 12 FRFNPAEVTEMEGILQEHHNAMPS-----------REILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAK 77 (268)
Q Consensus 12 t~FT~~Qv~eLEk~F~~~~~~yp~-----------~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk 77 (268)
..||.+|...|-.++..+....-+ ...=++||..||.-. | ..=++.|++..++|=+...|++
T Consensus 3 ~~fs~~E~~~Lv~~v~~~~~il~~k~~~~~~~~~k~~~W~~I~~~lN~~~---~-~~Rs~~~lkkkW~nlk~~~Kk~ 75 (78)
T PF13873_consen 3 PNFSEEEKEILVELVEKHKDILENKFSDSVSNKEKRKAWEEIAEELNALG---P-GKRSWKQLKKKWKNLKSKAKKK 75 (78)
T ss_pred CCCCHHHHHHHHHHHHHhHHHHhcccccHHHHHHHHHHHHHHHHHHHhcC---C-CCCCHHHHHHHHHHHHHHHHHH
Confidence 479999998888887764211111 233468999999632 2 3678899998899877776654
No 68
>PTZ00064 histone acetyltransferase; Provisional
Probab=47.64 E-value=16 Score=37.96 Aligned_cols=38 Identities=26% Similarity=0.354 Sum_probs=26.6
Q ss_pred CCCeEEEEecCCCCCccceeccccccccccccCccccccc
Q 024359 170 ADPEVQVRFAGFGAEEDEWVNIKRHVRQRSLPCEASECVA 209 (268)
Q Consensus 170 ge~ev~Vrf~gFg~eedewvnv~~~vR~rS~ple~~eC~~ 209 (268)
|+-|.+|||.||.-.-||||.-. ++.... +.+..++..
T Consensus 147 ~~~eyYVHy~g~nrRlD~WV~~~-ri~~~~-~~~~~~~~~ 184 (552)
T PTZ00064 147 EDYEFYVHFRGLNRRLDRWVKGK-DIKLSF-DVEELNDPN 184 (552)
T ss_pred CCeEEEEEecCcCchHhhhcChh-hccccc-ccccccccc
Confidence 55799999999999999999965 444322 334444443
No 69
>PF11523 DUF3223: Protein of unknown function (DUF3223); InterPro: IPR021602 This family of proteins has no known function. ; PDB: 2K0M_A.
Probab=47.05 E-value=22 Score=27.40 Aligned_cols=28 Identities=36% Similarity=0.449 Sum_probs=20.5
Q ss_pred EEeeeeccCCCCc-ceeEEEEEEccCCcccc
Q 024359 233 VLDAQRRRHDVRG-CRCRFLVRYDHDQSEVA 262 (268)
Q Consensus 233 V~~i~r~~Hd~~~-C~C~F~Vr~~hd~sEe~ 262 (268)
|..|+-+.|...+ .||.|+||= |+|+|-
T Consensus 41 i~~i~V~~hp~~~~srCF~vvR~--DGs~~D 69 (76)
T PF11523_consen 41 IDHIMVRKHPEFKDSRCFFVVRT--DGSEED 69 (76)
T ss_dssp EEEEEEEESSSS---EEEEEEET--TS-EEE
T ss_pred eeeEEEeecCCCCcceEEEEEEe--CCCeee
Confidence 6788889998875 999999995 677654
No 70
>PF00249 Myb_DNA-binding: Myb-like DNA-binding domain; InterPro: IPR014778 The retroviral oncogene v-myb, and its cellular counterpart c-myb, encode nuclear DNA-binding proteins. These belong to the SANT domain family that specifically recognise the sequence YAAC(G/T)G [, ]. In myb, one of the most conserved regions consisting of three tandem repeats has been shown to be involved in DNA-binding [].; PDB: 1X41_A 2XAF_B 2XAG_B 2XAH_B 2UXN_B 2Y48_B 2XAQ_B 2X0L_B 2IW5_B 2XAJ_B ....
Probab=42.43 E-value=72 Score=21.59 Aligned_cols=45 Identities=13% Similarity=0.126 Sum_probs=32.6
Q ss_pred cccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhc
Q 024359 12 FRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQN 69 (268)
Q Consensus 12 t~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQN 69 (268)
-.||++|...|.+++...+.. .=..||..++.+ =|..|++.=|+|
T Consensus 2 ~~Wt~eE~~~l~~~v~~~g~~-----~W~~Ia~~~~~~--------Rt~~qc~~~~~~ 46 (48)
T PF00249_consen 2 GPWTEEEDEKLLEAVKKYGKD-----NWKKIAKRMPGG--------RTAKQCRSRYQN 46 (48)
T ss_dssp -SS-HHHHHHHHHHHHHSTTT-----HHHHHHHHHSSS--------STHHHHHHHHHH
T ss_pred CCCCHHHHHHHHHHHHHhCCc-----HHHHHHHHcCCC--------CCHHHHHHHHHh
Confidence 469999999999999997533 678899998821 245888865554
No 71
>smart00027 EH Eps15 homology domain. Pair of EF hand motifs that recognise proteins containing Asn-Pro-Phe (NPF) sequences.
Probab=41.05 E-value=65 Score=24.62 Aligned_cols=45 Identities=9% Similarity=0.130 Sum_probs=32.1
Q ss_pred cCHHHHHHHHHHHHh---ccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhh
Q 024359 14 FNPAEVTEMEGILQE---HHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQ 68 (268)
Q Consensus 14 FT~~Qv~eLEk~F~~---~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQ 68 (268)
+|++|+.++.++|.. .++.+++..+..++-..++++ ..+|+.+|.
T Consensus 4 ls~~~~~~l~~~F~~~D~d~~G~Is~~el~~~l~~~~~~----------~~ev~~i~~ 51 (96)
T smart00027 4 ISPEDKAKYEQIFRSLDKNQDGTVTGAQAKPILLKSGLP----------QTLLAKIWN 51 (96)
T ss_pred CCHHHHHHHHHHHHHhCCCCCCeEeHHHHHHHHHHcCCC----------HHHHHHHHH
Confidence 688999999999887 345678887777766666654 355665553
No 72
>PF11516 DUF3220: Protein of unknown function (DUF3120); InterPro: IPR021597 This family of proteins with unknown function appears to be restricted to Bordetella. ; PDB: 2JPF_A.
Probab=39.69 E-value=13 Score=30.11 Aligned_cols=19 Identities=37% Similarity=0.522 Sum_probs=14.6
Q ss_pred cccCCcccccchhhhhhhcc
Q 024359 51 ERKGKIMVQMKQVWNWFQNR 70 (268)
Q Consensus 51 ~RaGK~~lt~kQVk~WFQNR 70 (268)
-|+|+++|+- .||.|.||=
T Consensus 22 lragsmalqg-dvkvwmqnl 40 (106)
T PF11516_consen 22 LRAGSMALQG-DVKVWMQNL 40 (106)
T ss_dssp -SSSSSSS-H-HHHHHHHHH
T ss_pred hhhhhhHhcc-cHHHHHHHH
Confidence 4899999974 599999993
No 73
>smart00717 SANT SANT SWI3, ADA2, N-CoR and TFIIIB'' DNA-binding domains.
Probab=39.64 E-value=56 Score=20.60 Aligned_cols=44 Identities=9% Similarity=0.091 Sum_probs=31.8
Q ss_pred cccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhc
Q 024359 12 FRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQN 69 (268)
Q Consensus 12 t~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQN 69 (268)
..||++|...|.+.+...+. ..-..||..|+- =|..||+..|.+
T Consensus 2 ~~Wt~~E~~~l~~~~~~~g~-----~~w~~Ia~~~~~---------rt~~~~~~~~~~ 45 (49)
T smart00717 2 GEWTEEEDELLIELVKKYGK-----NNWEKIAKELPG---------RTAEQCRERWNN 45 (49)
T ss_pred CCCCHHHHHHHHHHHHHHCc-----CCHHHHHHHcCC---------CCHHHHHHHHHH
Confidence 46999999999999999642 234678888871 145788766554
No 74
>PF01527 HTH_Tnp_1: Transposase; InterPro: IPR002514 Transposase proteins are necessary for efficient DNA transposition. This family consists of various Escherichia coli insertion elements and other bacterial transposases some of which are members of the IS3 family. This region includes a helix-turn-helix motif (HTH) at the N terminus followed by a leucine zipper (LZ) motif. The LZ motif has been shown to mediate oligomerisation of the transposase components in IS911 []. More information about these proteins can be found at Protein of the Month: Transposase [].; GO: 0003677 DNA binding, 0004803 transposase activity, 0006313 transposition, DNA-mediated; PDB: 2JN6_A 2RN7_A.
Probab=38.16 E-value=42 Score=24.12 Aligned_cols=46 Identities=20% Similarity=0.323 Sum_probs=29.5
Q ss_pred CCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcc
Q 024359 9 GPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNR 70 (268)
Q Consensus 9 ~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNR 70 (268)
+.+..||+++-..+=+.+... .....+||..+|+++ .++.+|-+-=
T Consensus 2 ~~r~~ys~e~K~~~v~~~~~~------g~sv~~va~~~gi~~----------~~l~~W~~~~ 47 (76)
T PF01527_consen 2 RKRRRYSPEFKLQAVREYLES------GESVSEVAREYGISP----------STLYNWRKQY 47 (76)
T ss_dssp -SS----HHHHHHHHHHHHHH------HCHHHHHHHHHTS-H----------HHHHHHHHHH
T ss_pred CCCCCCCHHHHHHHHHHHHHC------CCceEeeeccccccc----------ccccHHHHHH
Confidence 356789999887776655332 357889999999765 8999996543
No 75
>PF10668 Phage_terminase: Phage terminase small subunit; InterPro: IPR018925 This entry describes the terminase small subunit from Enterococcus phage phiFL1A, related proteins in other bacteriophage, and prophage regions of bacterial genomes. Packaging of double-stranded viral DNA concatemers requires interaction of the prohead with virus DNA. This process is mediated by a phage-encoded DNA recognition and terminase protein. The terminase enzymes described so far, which are hetero-oligomers composed of a small and a large subunit, do not have a significant level of sequence homology. The small terminase subunit is thought to form a nucleoprotein structure that helps to position the terminase large subunit at the packaging initiation site [].
Probab=36.89 E-value=23 Score=26.63 Aligned_cols=30 Identities=27% Similarity=0.420 Sum_probs=21.1
Q ss_pred HHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhh
Q 024359 24 GILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWF 67 (268)
Q Consensus 24 k~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WF 67 (268)
++|.++++.. ...+||++||.| +.||..|=
T Consensus 14 e~y~~~~g~i----~lkdIA~~Lgvs----------~~tIr~WK 43 (60)
T PF10668_consen 14 EIYKESNGKI----KLKDIAEKLGVS----------ESTIRKWK 43 (60)
T ss_pred HHHHHhCCCc----cHHHHHHHHCCC----------HHHHHHHh
Confidence 4566654333 456799999976 49999983
No 76
>PF09465 LBR_tudor: Lamin-B receptor of TUDOR domain; InterPro: IPR019023 The Lamin-B receptor is a chromatin and lamin binding protein in the inner nuclear membrane. It is one of the integral inner nuclear envelope membrane proteins responsible for targeting nuclear membranes to chromatin, being a downstream effector of Ran, a small Ras-like nuclear GTPase which regulates NE assembly. Lamin-B receptor interacts with importin beta, a Ran-binding protein, thereby directly contributing to the fusion of membrane vesicles and the formation of the nuclear envelope []. ; PDB: 2L8D_A 2DIG_A.
Probab=36.57 E-value=1.2e+02 Score=22.76 Aligned_cols=40 Identities=23% Similarity=0.543 Sum_probs=25.9
Q ss_pred cccCCceEEEEeecCccceeeeeEEEeeeeccCCCCcceeEEEEEEccC
Q 024359 209 AVLPGDLILCFQEGKDQALYFDAHVLDAQRRRHDVRGCRCRFLVRYDHD 257 (268)
Q Consensus 209 ~v~~Gd~vlcf~e~~~~aly~DA~V~~i~r~~Hd~~~C~C~F~Vr~~hd 257 (268)
+.-.|+.|...=-+ .++||.|.|++.-.+.| ...|.|..+
T Consensus 5 k~~~Ge~V~~rWP~--s~lYYe~kV~~~d~~~~-------~y~V~Y~DG 44 (55)
T PF09465_consen 5 KFAIGEVVMVRWPG--SSLYYEGKVLSYDSKSD-------RYTVLYEDG 44 (55)
T ss_dssp SS-SS-EEEEE-TT--TS-EEEEEEEEEETTTT-------EEEEEETTS
T ss_pred cccCCCEEEEECCC--CCcEEEEEEEEecccCc-------eEEEEEcCC
Confidence 45578888776544 58999999999766555 457788753
No 77
>PF13518 HTH_28: Helix-turn-helix domain
Probab=34.41 E-value=32 Score=22.94 Aligned_cols=23 Identities=22% Similarity=0.492 Sum_probs=18.7
Q ss_pred HHHHHHHHhCCCccccCCcccccchhhhhhhcc
Q 024359 38 ILVALAEKFSESPERKGKIMVQMKQVWNWFQNR 70 (268)
Q Consensus 38 ~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNR 70 (268)
...++|..||+|. .+|..|.+.=
T Consensus 14 s~~~~a~~~gis~----------~tv~~w~~~y 36 (52)
T PF13518_consen 14 SVREIAREFGISR----------STVYRWIKRY 36 (52)
T ss_pred CHHHHHHHHCCCH----------hHHHHHHHHH
Confidence 4667999999764 9999998763
No 78
>COG3413 Predicted DNA binding protein [General function prediction only]
Probab=34.28 E-value=43 Score=29.54 Aligned_cols=35 Identities=14% Similarity=0.097 Sum_probs=30.0
Q ss_pred ccCHHHHHHHHHHHHhccCCCCC---HHHHHHHHHHhCCCc
Q 024359 13 RFNPAEVTEMEGILQEHHNAMPS---REILVALAEKFSESP 50 (268)
Q Consensus 13 ~FT~~Qv~eLEk~F~~~~~~yp~---~~~rq~LA~~fnlS~ 50 (268)
.+|..|++.|-.+|+. .|.+ +....+||+.||.|+
T Consensus 155 ~LTdrQ~~vL~~A~~~---GYFd~PR~~~l~dLA~~lGISk 192 (215)
T COG3413 155 DLTDRQLEVLRLAYKM---GYFDYPRRVSLKDLAKELGISK 192 (215)
T ss_pred cCCHHHHHHHHHHHHc---CCCCCCccCCHHHHHHHhCCCH
Confidence 6999999999999998 5654 466789999999984
No 79
>PTZ00183 centrin; Provisional
Probab=33.90 E-value=1.4e+02 Score=23.54 Aligned_cols=41 Identities=5% Similarity=-0.027 Sum_probs=31.6
Q ss_pred CCccccCHHHHHHHHHHHHh---ccCCCCCHHHHHHHHHHhCCC
Q 024359 9 GPAFRFNPAEVTEMEGILQE---HHNAMPSREILVALAEKFSES 49 (268)
Q Consensus 9 ~pRt~FT~~Qv~eLEk~F~~---~~~~yp~~~~rq~LA~~fnlS 49 (268)
--+..|++.|+.+++++|.. .++.+++..+...+-..+++.
T Consensus 6 ~~~~~~~~~~~~~~~~~F~~~D~~~~G~i~~~e~~~~l~~~g~~ 49 (158)
T PTZ00183 6 SERPGLTEDQKKEIREAFDLFDTDGSGTIDPKELKVAMRSLGFE 49 (158)
T ss_pred cccCCCCHHHHHHHHHHHHHhCCCCCCcccHHHHHHHHHHhCCC
Confidence 34567999999999999986 345788887777777776643
No 80
>PRK03975 tfx putative transcriptional regulator; Provisional
Probab=33.54 E-value=58 Score=28.10 Aligned_cols=49 Identities=10% Similarity=-0.016 Sum_probs=37.1
Q ss_pred cccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhhhhcc
Q 024359 12 FRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAIRAKS 78 (268)
Q Consensus 12 t~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~Kkk~ 78 (268)
..+|+.|.+.|+-.++. -..++||+.||+|. ..|+.|-++-+.+.++.-
T Consensus 5 ~~Lt~rqreVL~lr~~G--------lTq~EIAe~LGiS~----------~tVs~ie~ra~kkLr~~~ 53 (141)
T PRK03975 5 SFLTERQIEVLRLRERG--------LTQQEIADILGTSR----------ANVSSIEKRARENIEKAR 53 (141)
T ss_pred cCCCHHHHHHHHHHHcC--------CCHHHHHHHHCCCH----------HHHHHHHHHHHHHHHHHH
Confidence 46788888888774322 24689999999874 889999998888766544
No 81
>TIGR01321 TrpR trp operon repressor, proteobacterial. This model represents TrpR, the repressor of the trp operon. It is found so far only in the gamma subdivision of the proteobacteria and in Chlamydia trachomatis. All members belong to species capable of tryptophan biosynthesis.
Probab=33.30 E-value=27 Score=28.43 Aligned_cols=56 Identities=9% Similarity=0.077 Sum_probs=38.2
Q ss_pred ccCHHHHHHHHHHHHhccCCCC-CHHHHHHHHHHhCCCcccc--CCcccc--cchhhhhhhc
Q 024359 13 RFNPAEVTEMEGILQEHHNAMP-SREILVALAEKFSESPERK--GKIMVQ--MKQVWNWFQN 69 (268)
Q Consensus 13 ~FT~~Qv~eLEk~F~~~~~~yp-~~~~rq~LA~~fnlS~~Ra--GK~~lt--~kQVk~WFQN 69 (268)
.+|++|+..|...|.-.+ .-+ ..-...+||+++|.|..-. |.-.|+ +.+++.|.+.
T Consensus 32 lLTp~E~~~l~~R~~i~~-~Ll~~~~tQrEIa~~lGiS~atIsR~sn~lk~~~~~~~~~l~~ 92 (94)
T TIGR01321 32 ILTRSEREDLGDRIRIVN-ELLNGNMSQREIASKLGVSIATITRGSNNLKTMDPNFKQFLRK 92 (94)
T ss_pred hCCHHHHHHHHHHHHHHH-HHHhCCCCHHHHHHHhCCChhhhhHHHhhcccCCHHHHHHHHh
Confidence 489999999999888752 111 2345778999999875322 555666 6667777653
No 82
>COG3458 Acetyl esterase (deacetylase) [Secondary metabolites biosynthesis, transport, and catabolism]
Probab=33.03 E-value=59 Score=31.78 Aligned_cols=93 Identities=23% Similarity=0.366 Sum_probs=66.5
Q ss_pred CCcceEEe-eccCCCceeehhhhhhcccccCCCCeEEEEecCCCCCccceec-----------cccccccccccCccccc
Q 024359 140 STFMEFEA-KSARDGAWYDVSAFLAQRNFDTADPEVQVRFAGFGAEEDEWVN-----------IKRHVRQRSLPCEASEC 207 (268)
Q Consensus 140 ~~~~efEa-~S~~D~AWYdv~~fl~~R~l~~ge~ev~Vrf~gFg~eedewvn-----------v~~~vR~rS~ple~~eC 207 (268)
-.+|.|+. +-.+=.+||=+.. .++|..-..|+|.||+.--.+|-+ +.-.+|=.|.--+++-|
T Consensus 56 ~ydvTf~g~~g~rI~gwlvlP~------~~~~~~P~vV~fhGY~g~~g~~~~~l~wa~~Gyavf~MdvRGQg~~~~dt~~ 129 (321)
T COG3458 56 VYDVTFTGYGGARIKGWLVLPR------HEKGKLPAVVQFHGYGGRGGEWHDMLHWAVAGYAVFVMDVRGQGSSSQDTAD 129 (321)
T ss_pred EEEEEEeccCCceEEEEEEeec------ccCCccceEEEEeeccCCCCCccccccccccceeEEEEecccCCCccccCCC
Confidence 36777762 1123357998772 234778899999999998888733 45577888877777777
Q ss_pred cccc---CCceEEEEeecCccceeeeeEEEeeeec
Q 024359 208 VAVL---PGDLILCFQEGKDQALYFDAHVLDAQRR 239 (268)
Q Consensus 208 ~~v~---~Gd~vlcf~e~~~~aly~DA~V~~i~r~ 239 (268)
...- ||-.+.+..+++| .+||=-.++|+.|.
T Consensus 130 ~p~~~s~pG~mtrGilD~kd-~yyyr~v~~D~~~a 163 (321)
T COG3458 130 PPGGPSDPGFMTRGILDRKD-TYYYRGVFLDAVRA 163 (321)
T ss_pred CCCCCcCCceeEeecccCCC-ceEEeeehHHHHHH
Confidence 7665 7888889999886 77887777777654
No 83
>PLN00104 MYST -like histone acetyltransferase; Provisional
Probab=32.62 E-value=1e+02 Score=31.46 Aligned_cols=48 Identities=15% Similarity=0.270 Sum_probs=32.1
Q ss_pred ccccCCceEEEEeecCccceeeeeEEEeeeeccCCCCcceeEEEEEEccCC
Q 024359 208 VAVLPGDLILCFQEGKDQALYFDAHVLDAQRRRHDVRGCRCRFLVRYDHDQ 258 (268)
Q Consensus 208 ~~v~~Gd~vlcf~e~~~~aly~DA~V~~i~r~~Hd~~~C~C~F~Vr~~hd~ 258 (268)
..+..|+.|+|+.... -.||.|.|+++.+..-...+ .-.+-|.|...|
T Consensus 52 ~~~~VGekVla~~~~D--g~~~~A~VI~~R~~~~~~~~-~~~YYVHY~g~n 99 (450)
T PLN00104 52 LPLEVGTRVMCRWRFD--GKYHPVKVIERRRGGSGGPN-DYEYYVHYTEFN 99 (450)
T ss_pred ceeccCCEEEEEECCC--CCEEEEEEEEEeccCCCCCC-CceEEEEEecCC
Confidence 3466999999998643 57889999999763300111 115888888654
No 84
>PF03672 UPF0154: Uncharacterised protein family (UPF0154); InterPro: IPR005359 The proteins in this entry are functionally uncharacterised.
Probab=31.84 E-value=80 Score=24.20 Aligned_cols=36 Identities=25% Similarity=0.324 Sum_probs=31.6
Q ss_pred HHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhh
Q 024359 20 TEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWN 65 (268)
Q Consensus 20 ~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~ 65 (268)
..||+.|++ |..++.+....+....|-.| +++||+.
T Consensus 20 ~~~~k~l~~--NPpine~mir~M~~QMG~kp--------Sekqi~Q 55 (64)
T PF03672_consen 20 KYMEKQLKE--NPPINEKMIRAMMMQMGRKP--------SEKQIKQ 55 (64)
T ss_pred HHHHHHHHH--CCCCCHHHHHHHHHHhCCCc--------cHHHHHH
Confidence 468999988 69999999999999999876 8888874
No 85
>PF04717 Phage_base_V: Phage-related baseplate assembly protein; InterPro: IPR006531 This domain occurs in a family of phage (and bacteriocin) proteins related to the phage P2 V gene product, which forms the small spike at the tip of the tail []. Homologs in general are annotated as baseplate assembly protein V. At least one member is encoded within a region of Pectobacterium carotovorum (Erwinia carotovora) described as a bacteriocin, a phage tail-derived module able to kill bacteria closely related to the host strain. It is also found in Vgr-related proteins. Genes encoding type VI secretion systems (T6SS) are widely distributed in pathogenic Gram-negative bacterial species. In Vibrio cholerae, T6SS have been found to secrete three related proteins extracellularly, VgrG-1, VgrG-2, and VgrG-3. VgrG-1 can covalently cross-link actin in vitro, and this activity was used to demonstrate that V. cholerae can translocate VgrG-1 into macrophages by a T6SS-dependent mechanism. VgrG-related proteins likely assemble into a trimeric complex that is analogous to that formed by the two trimeric proteins gp27 and gp5 that make up the baseplate "tail spike" of Escherichia coli bacteriophage T4. The VgrG components of the T6SS apparatus might assemble a "cell-puncturing device" analogous to phage tail spikes to deliver effector protein domains through membranes of target host cells []. Gp5 is an integral component of the virion baseplate of bacteriophage T4. T4 Gp5 consists of 3 domains connected via long linkers: the N-terminal oligosaccharide/oligonucleotide-binding (OB)-fold domain, the middle lysozyme domain, and the C-terminal triplestranded-helix. The equivalent of the Gp5 OB-fold domain in the structure of VgrG is the domain of unknown function comprising residues 380-470 and conserved in all known VgrGs. This entry represents the OB-fold domain which consists of a 5-stranded antiparallel-barrel with a Greek-key topology [].; PDB: 3AQJ_C 3QR8_A 2P5Z_X.
Probab=31.74 E-value=91 Score=23.20 Aligned_cols=50 Identities=22% Similarity=0.314 Sum_probs=25.9
Q ss_pred CCCeEEEEecCCCCCccceeccccccccccccCcccccccccCCceEEEEeecCc
Q 024359 170 ADPEVQVRFAGFGAEEDEWVNIKRHVRQRSLPCEASECVAVLPGDLILCFQEGKD 224 (268)
Q Consensus 170 ge~ev~Vrf~gFg~eedewvnv~~~vR~rS~ple~~eC~~v~~Gd~vlcf~e~~~ 224 (268)
+++.+||+|..-++..--|+.+-. .++- ......-..+||.|+|...++|
T Consensus 9 ~~grvrV~~~~~~~~~s~Wl~~~~---~~ag--~~g~~~~P~iGeqV~v~~~~Gd 58 (79)
T PF04717_consen 9 DKGRVRVRFPDDGDIVSDWLPVLQ---PRAG--GWGFWFPPEIGEQVLVLFPGGD 58 (79)
T ss_dssp TTTEEEEE-B-CTTEEEEEEEE-----S-BS--SSB------TT-EEEEEEGGCT
T ss_pred CCCEEEEEEecCCCccceEEEeee---hhcc--CCeeEccCCCCcEEEEEccCCc
Confidence 357899999545555556887652 1111 4445566689999998888774
No 86
>cd06171 Sigma70_r4 Sigma70, region (SR) 4 refers to the most C-terminal of four conserved domains found in Escherichia coli (Ec) sigma70, the main housekeeping sigma, and related sigma-factors (SFs). A SF is a dissociable subunit of RNA polymerase, it directs bacterial or plastid core RNA polymerase to specific promoter elements located upstream of transcription initiation points. The SR4 of Ec sigma70 and other essential primary SFs contact promoter sequences located 35 base-pairs upstream of the initiation point, recognizing a 6-base-pair -35 consensus TTGACA. Sigma70 related SFs also include SFs which are dispensable for bacterial cell growth for example Ec sigmaS, SFs which activate regulons in response to a specific signal for example heat-shock Ec sigmaH, and a group of SFs which includes the extracytoplasmic function (ECF) SFs and is typified by Ec sigmaE which contains SR2 and -4 only. ECF SFs direct the transcription of genes that regulate various responses including periplas
Probab=30.45 E-value=38 Score=21.48 Aligned_cols=44 Identities=14% Similarity=-0.003 Sum_probs=29.2
Q ss_pred ccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhh
Q 024359 13 RFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYA 73 (268)
Q Consensus 13 ~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k 73 (268)
.+++.+...++..|.+ .-...++|+.+|+|. ..|..|.+.-+.+
T Consensus 10 ~l~~~~~~~~~~~~~~-------~~~~~~ia~~~~~s~----------~~i~~~~~~~~~~ 53 (55)
T cd06171 10 KLPEREREVILLRFGE-------GLSYEEIAEILGISR----------STVRQRLHRALKK 53 (55)
T ss_pred hCCHHHHHHHHHHHhc-------CCCHHHHHHHHCcCH----------HHHHHHHHHHHHH
Confidence 3566666666655533 235778899999764 8899888765443
No 87
>PF15057 DUF4537: Domain of unknown function (DUF4537)
Probab=30.28 E-value=95 Score=25.78 Aligned_cols=40 Identities=23% Similarity=0.478 Sum_probs=30.0
Q ss_pred CceEEEEeecCccceeeeeEEEeeeeccCCCCcceeEEEEEEccCCcccc
Q 024359 213 GDLILCFQEGKDQALYFDAHVLDAQRRRHDVRGCRCRFLVRYDHDQSEVA 262 (268)
Q Consensus 213 Gd~vlcf~e~~~~aly~DA~V~~i~r~~Hd~~~C~C~F~Vr~~hd~sEe~ 262 (268)
|..|++-.+. +..||=+.|.+... ...|+|.|++++.++.
T Consensus 1 g~~VlAR~~~--DG~YY~GtV~~~~~--------~~~~lV~f~~~~~~~v 40 (124)
T PF15057_consen 1 GQKVLARREE--DGFYYPGTVKKCVS--------SGQFLVEFDDGDTQEV 40 (124)
T ss_pred CCeEEEeeCC--CCcEEeEEEEEccC--------CCEEEEEECCCCEEEe
Confidence 6788887763 47899999998873 2469999977766643
No 88
>PRK10072 putative transcriptional regulator; Provisional
Probab=29.65 E-value=37 Score=27.33 Aligned_cols=24 Identities=21% Similarity=0.249 Sum_probs=18.6
Q ss_pred HHHHHHHhCCCccccCCcccccchhhhhhhcchh
Q 024359 39 LVALAEKFSESPERKGKIMVQMKQVWNWFQNRRY 72 (268)
Q Consensus 39 rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~ 72 (268)
..+||+.+|+| ..-|..|.+.+|.
T Consensus 49 Q~elA~~lGvS----------~~TVs~WE~G~r~ 72 (96)
T PRK10072 49 IDDFARVLGVS----------VAMVKEWESRRVK 72 (96)
T ss_pred HHHHHHHhCCC----------HHHHHHHHcCCCC
Confidence 56778888865 4779999999854
No 89
>PF13936 HTH_38: Helix-turn-helix domain; PDB: 2W48_A.
Probab=29.64 E-value=46 Score=22.67 Aligned_cols=32 Identities=19% Similarity=0.389 Sum_probs=15.2
Q ss_pred cccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCc
Q 024359 12 FRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESP 50 (268)
Q Consensus 12 t~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~ 50 (268)
..||.+|..+++.++++ ..-..+||+.||.|+
T Consensus 3 ~~Lt~~eR~~I~~l~~~-------G~s~~~IA~~lg~s~ 34 (44)
T PF13936_consen 3 KHLTPEERNQIEALLEQ-------GMSIREIAKRLGRSR 34 (44)
T ss_dssp ---------HHHHHHCS----------HHHHHHHTT--H
T ss_pred cchhhhHHHHHHHHHHc-------CCCHHHHHHHHCcCc
Confidence 46899999999988765 245677999999875
No 90
>PF07930 DAP_B: D-aminopeptidase, domain B; InterPro: IPR012856 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. D-aminopeptidase (Q9ZBA9 from SWISSPROT) is a dimeric enzyme with each monomer being composed of three domains. Domain B is organised to form a beta barrel made up of eight antiparallel beta strands. It is connected to domain A, the catalytic domain, by an eight-residue sequence, and also interacts with both domains A and C via non-covalent bonds. Domain B probably functions in maintaining domain C in a good position to interact with the catalytic domain []. This domain is found in peptidases that belong to MEROPS peptidase family S12 (D-Ala-D-Ala carboxypeptidase B family, clan ME).; GO: 0004177 aminopeptidase activity; PDB: 1EI5_A.
Probab=29.54 E-value=31 Score=28.04 Aligned_cols=73 Identities=25% Similarity=0.290 Sum_probs=44.2
Q ss_pred ccCCCceeehhhhhhcccccCCCCeEEEEecCCCCCccceeccccccccccccCcccccccccCCceEEEEeecCcccee
Q 024359 149 SARDGAWYDVSAFLAQRNFDTADPEVQVRFAGFGAEEDEWVNIKRHVRQRSLPCEASECVAVLPGDLILCFQEGKDQALY 228 (268)
Q Consensus 149 S~~D~AWYdv~~fl~~R~l~~ge~ev~Vrf~gFg~eedewvnv~~~vR~rS~ple~~eC~~v~~Gd~vlcf~e~~~~aly 228 (268)
..=.|.|||=.+=|.-|+=.-|++.|+|||.+. +|-+++-..=|.+|. --..++.||.|--- +.++.+-
T Consensus 11 ~~W~G~wLD~etgL~l~i~~~~~G~~~~rya~~----pE~l~~~~~~~a~s~-----~~~~~rDGd~l~m~--R~~ENlt 79 (88)
T PF07930_consen 11 PAWFGSWLDPETGLVLRIEDAGQGRVKLRYATS----PEMLDLVSENEARSS-----GTVLRRDGDMLRME--RLDENLT 79 (88)
T ss_dssp GGG-EEEE-TTT--EEEEEE-STTEEEEE-SSS-----EEEEEEETTEEE-S-----S-EEEEETTEEEEE--EGGGTEE
T ss_pred CCcceeeEcCCCceEEEeecCCCceEEEEecCC----CceeeccCCCcccCc-----ceEEEEcCCeEEEe--ecccceE
Confidence 355799999999999999999999999998764 566676656677665 34567788877532 2333444
Q ss_pred eeeE
Q 024359 229 FDAH 232 (268)
Q Consensus 229 ~DA~ 232 (268)
-+++
T Consensus 80 l~~~ 83 (88)
T PF07930_consen 80 LNMK 83 (88)
T ss_dssp EEEE
T ss_pred EEee
Confidence 4443
No 91
>PRK07539 NADH dehydrogenase subunit E; Validated
Probab=28.34 E-value=1.2e+02 Score=25.77 Aligned_cols=20 Identities=15% Similarity=0.170 Sum_probs=18.2
Q ss_pred CCCCCHHHHHHHHHHhCCCc
Q 024359 31 NAMPSREILVALAEKFSESP 50 (268)
Q Consensus 31 ~~yp~~~~rq~LA~~fnlS~ 50 (268)
..|++.+..+.+|+.+|+++
T Consensus 35 ~g~ip~~~~~~iA~~l~v~~ 54 (154)
T PRK07539 35 RGWVPDEAIEAVADYLGMPA 54 (154)
T ss_pred hCCCCHHHHHHHHHHhCcCH
Confidence 57999999999999999875
No 92
>PF05506 DUF756: Domain of unknown function (DUF756); InterPro: IPR008475 This domain is found, normally as a tandem repeat, at the C terminus of bacterial phospholipase C proteins.; GO: 0004629 phospholipase C activity, 0016042 lipid catabolic process
Probab=27.99 E-value=54 Score=25.05 Aligned_cols=23 Identities=39% Similarity=0.718 Sum_probs=16.2
Q ss_pred cCCCceeehhhhhhcccccCCCCeEEEEecC
Q 024359 150 ARDGAWYDVSAFLAQRNFDTADPEVQVRFAG 180 (268)
Q Consensus 150 ~~D~AWYdv~~fl~~R~l~~ge~ev~Vrf~g 180 (268)
...+-|||+.+.... + ..=||+|
T Consensus 66 ~~s~gwYDl~v~~~~-------~-F~rr~aG 88 (89)
T PF05506_consen 66 AASGGWYDLTVTGPN-------G-FLRRFAG 88 (89)
T ss_pred cCCCCcEEEEEEcCC-------C-EEEEecC
Confidence 668899999865533 2 6667766
No 93
>PTZ00184 calmodulin; Provisional
Probab=27.96 E-value=1.7e+02 Score=22.59 Aligned_cols=37 Identities=5% Similarity=0.171 Sum_probs=28.0
Q ss_pred ccCHHHHHHHHHHHHhc---cCCCCCHHHHHHHHHHhCCC
Q 024359 13 RFNPAEVTEMEGILQEH---HNAMPSREILVALAEKFSES 49 (268)
Q Consensus 13 ~FT~~Qv~eLEk~F~~~---~~~yp~~~~rq~LA~~fnlS 49 (268)
-+|..++.++.+.|... +..+++..+...+...++.+
T Consensus 4 ~~~~~~~~~~~~~F~~~D~~~~G~i~~~e~~~~l~~~~~~ 43 (149)
T PTZ00184 4 QLTEEQIAEFKEAFSLFDKDGDGTITTKELGTVMRSLGQN 43 (149)
T ss_pred ccCHHHHHHHHHHHHHHcCCCCCcCCHHHHHHHHHHhCCC
Confidence 47889999999998763 45678887777777777654
No 94
>COG2944 Predicted transcriptional regulator [Transcription]
Probab=27.92 E-value=78 Score=26.30 Aligned_cols=39 Identities=23% Similarity=0.358 Sum_probs=31.5
Q ss_pred cCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcch
Q 024359 14 FNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRR 71 (268)
Q Consensus 14 FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR 71 (268)
+++.||.+|-+.+.-+ +..+|..||.|. .-|+.|=|+|+
T Consensus 44 ls~~eIk~iRe~~~lS---------Q~vFA~~L~vs~----------~Tv~~WEqGr~ 82 (104)
T COG2944 44 LSPTEIKAIREKLGLS---------QPVFARYLGVSV----------STVRKWEQGRK 82 (104)
T ss_pred CCHHHHHHHHHHhCCC---------HHHHHHHHCCCH----------HHHHHHHcCCc
Confidence 7888888888877764 578899999763 55999999983
No 95
>cd00167 SANT 'SWI3, ADA2, N-CoR and TFIIIB' DNA-binding domains. Tandem copies of the domain bind telomeric DNA tandem repeatsas part of the capping complex. Binding is sequence dependent for repeats which contain the G/C rich motif [C2-3 A (CA)1-6]. The domain is also found in regulatory transcriptional repressor complexes where it also binds DNA.
Probab=27.53 E-value=1.2e+02 Score=18.85 Aligned_cols=43 Identities=12% Similarity=0.117 Sum_probs=30.1
Q ss_pred ccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhc
Q 024359 13 RFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQN 69 (268)
Q Consensus 13 ~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQN 69 (268)
.||.+|...|.+.+...+. ..=..||+.++.- +..||+.-|+|
T Consensus 1 ~Wt~eE~~~l~~~~~~~g~-----~~w~~Ia~~~~~r---------s~~~~~~~~~~ 43 (45)
T cd00167 1 PWTEEEDELLLEAVKKYGK-----NNWEKIAKELPGR---------TPKQCRERWRN 43 (45)
T ss_pred CCCHHHHHHHHHHHHHHCc-----CCHHHHHhHcCCC---------CHHHHHHHHHH
Confidence 3799999999999998642 2346788888631 45777754443
No 96
>PRK07571 bidirectional hydrogenase complex protein HoxE; Reviewed
Probab=26.00 E-value=1.4e+02 Score=26.41 Aligned_cols=38 Identities=11% Similarity=0.178 Sum_probs=29.1
Q ss_pred ccCHHHHHHHHHHHHhcc----------------CCCCCHHHHHHHHHHhCCCc
Q 024359 13 RFNPAEVTEMEGILQEHH----------------NAMPSREILVALAEKFSESP 50 (268)
Q Consensus 13 ~FT~~Qv~eLEk~F~~~~----------------~~yp~~~~rq~LA~~fnlS~ 50 (268)
.|+.+++++++++..... ..|++.+..+.+|+.||+++
T Consensus 15 ~~~~~~~~~i~~ii~~~~~~~~~li~~L~~iQ~~~GyIp~e~~~~iA~~l~v~~ 68 (169)
T PRK07571 15 PSGDKRFKVLEATMKRNQYRQDALIEVLHKAQELFGYLERDLLLYVARQLKLPL 68 (169)
T ss_pred cCcHHHHHHHHHHHHHcCCCHHHHHHHHHHHHHHcCCCCHHHHHHHHHHhCcCH
Confidence 466777777776555433 47999999999999999875
No 97
>PRK04980 hypothetical protein; Provisional
Probab=25.57 E-value=1.2e+02 Score=24.93 Aligned_cols=32 Identities=16% Similarity=0.192 Sum_probs=25.5
Q ss_pred cccCCceEEEEeecCccceeeeeEEEeeeeccC
Q 024359 209 AVLPGDLILCFQEGKDQALYFDAHVLDAQRRRH 241 (268)
Q Consensus 209 ~v~~Gd~vlcf~e~~~~aly~DA~V~~i~r~~H 241 (268)
..+|||.|..+.-+. ...|++++|++|...+-
T Consensus 31 ~~~~G~~~~V~~~e~-g~~~c~ieI~sV~~i~f 62 (102)
T PRK04980 31 HFKPGDVLRVGTFED-DRYFCTIEVLSVSPVTF 62 (102)
T ss_pred CCCCCCEEEEEECCC-CcEEEEEEEEEEEEEeh
Confidence 467999999865555 48999999999987653
No 98
>PF01343 Peptidase_S49: Peptidase family S49 peptidase classification.; InterPro: IPR002142 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This group of serine peptidases belong to MEROPS peptidase family S49 (protease IV family, clan S-). The predicted active site serine for members of this family occurs in a transmembrane domain. The domain defines sequences in viruses, archaea, bacteria and plants. These sequences are variously annotated in the different taxonomic groups, examples are: Viruses: capsid protein Archaea: proteinase IV homolog Bacteria: proteinase IV, sohB, SppA, pfaP, putative protease Plants: SppA, protease IV This group also contains proteins classified as non-peptidase homologues that either have been found experimentally to be without peptidase activity, or lack amino acid residues that are believed to be essential for the catalytic activity of peptidases. Related proteins, non-peptidase homologs and unclassified S49 members are also to be found in IPR002810 from INTERPRO.; GO: 0008233 peptidase activity, 0006508 proteolysis; PDB: 3RST_B 3BEZ_D 3BF0_A.
Probab=22.19 E-value=1.6e+02 Score=24.66 Aligned_cols=54 Identities=22% Similarity=0.212 Sum_probs=40.3
Q ss_pred CCCCCCCCCccccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcc
Q 024359 2 GRPPSNGGPAFRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNR 70 (268)
Q Consensus 2 GrPps~~~pRt~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNR 70 (268)
|.-.+..-++..+|+++.+.|++.+... -..+...+|+.=+++. .+|..|++++
T Consensus 68 g~~K~~~~~~~~~s~~~r~~~~~~l~~~-----~~~f~~~Va~~R~~~~----------~~v~~~~~~~ 121 (154)
T PF01343_consen 68 GEYKSAGFPRDPMSEEERENLQELLDEL-----YDQFVNDVAEGRGLSP----------DDVEEIADGG 121 (154)
T ss_dssp STTCCCCCTTSS--HHHHHHHHHHHHHH-----HHHHHHHHHHHHTS-H----------HHHHCHHCCH
T ss_pred CccccccCcCCCCCHHHHHHHHHHHHHH-----HHHHHHHHHHccCCCH----------HHHHHHHhhc
Confidence 3344555688899999999999999884 2678899999888664 7899999885
No 99
>TIGR02607 antidote_HigA addiction module antidote protein, HigA family. Members of this family form a distinct clade within the larger family HTH_3 of helix-turn-helix proteins, described by Pfam model pfam01381. Members of this clade are strictly bacterial and nearly always shorter than 110 amino acids. This family includes the characterized member HigA, without which the killer protein HigB cannot be cloned. The hig (host inhibition of growth) system is noted to be unusual in that killer protein is uncoded by the upstream member of the gene pair.
Probab=21.60 E-value=2.1e+02 Score=20.50 Aligned_cols=19 Identities=16% Similarity=0.202 Sum_probs=14.9
Q ss_pred CCCCCHHHHHHHHHHhCCC
Q 024359 31 NAMPSREILVALAEKFSES 49 (268)
Q Consensus 31 ~~yp~~~~rq~LA~~fnlS 49 (268)
+..++.+...+||+.||++
T Consensus 42 ~~~~~~~~~~~l~~~l~v~ 60 (78)
T TIGR02607 42 RRGITADMALRLAKALGTS 60 (78)
T ss_pred CCCCCHHHHHHHHHHcCCC
Confidence 4567888888888888876
No 100
>PF08880 QLQ: QLQ; InterPro: IPR014978 QLQ is named after the conserved Gln, Leu, Gln motif. QLQ is found at the N terminus of SWI2/SNF2 protein, which has been shown to be involved in protein-protein interactions. QLQ has been postulated to be involved in mediating protein interactions []. ; GO: 0005524 ATP binding, 0016818 hydrolase activity, acting on acid anhydrides, in phosphorus-containing anhydrides, 0006355 regulation of transcription, DNA-dependent, 0005634 nucleus
Probab=21.49 E-value=70 Score=21.80 Aligned_cols=16 Identities=25% Similarity=0.557 Sum_probs=12.6
Q ss_pred ccCHHHHHHHHHHHHh
Q 024359 13 RFNPAEVTEMEGILQE 28 (268)
Q Consensus 13 ~FT~~Qv~eLEk~F~~ 28 (268)
.||++|+.+||.-..-
T Consensus 2 ~FT~~Ql~~L~~Qi~a 17 (37)
T PF08880_consen 2 PFTPAQLQELRAQILA 17 (37)
T ss_pred CCCHHHHHHHHHHHHH
Confidence 5999999999974433
No 101
>PF00196 GerE: Bacterial regulatory proteins, luxR family; InterPro: IPR000792 This domain is a DNA-binding, helix-turn-helix (HTH) domain of about 65 amino acids, present in transcription regulators of the LuxR/FixJ family of response regulators. The domain is named after Vibrio fischeri luxR, a transcriptional activator for quorum-sensing control of luminescence. LuxR-type HTH domain proteins occur in a variety of organisms. The DNA-binding HTH domain is usually located in the C-terminal region; the N-terminal region often containing an autoinducer-binding domain or a response regulatory domain. Most luxR-type regulators act as transcription activators, but some can be repressors or have a dual role for different sites. LuxR-type HTH regulators control a wide variety of activities in various biological processes. The luxR-type, DNA-binding HTH domain forms a four-helical bundle structure. The HTH motif comprises the second and third helices, known as the scaffold and recognition helix, respectively. The HTH binds DNA in the major groove, where the N-terminal part of the recognition helix makes most of the DNA contacts. The fourth helix is involved in dimerisation of gerE and traR. Signalling events by one of the four activation mechanisms described below lead to multimerisation of the regulator. The regulators bind DNA as multimers [, , ]. LuxR-type HTH proteins can be activated by one of four different mechanisms: 1) Regulators which belong to a two-component sensory transduction system where the protein is activated by its phosphorylation, generally on an aspartate residue, by a transmembrane kinase [, ]. Some proteins that belong to this category are: Rhizobiaceae fixJ (global regulator inducing expression of nitrogen-fixation genes in microaerobiosis) Escherichia coli and Salmonella typhimurium uhpA (activates hexose phosphate transport gene uhpT) E. coli narL and narP (activate nitrate reductase operon) Enterobacteria rcsB (regulation of exopolysaccharide biosynthesis in enteric and plant pathogenesis) Bordetella pertussis bvgA (virulence factor) Bacillus subtilis coma (involved in expression of late-expressing competence genes) 2) Regulators which are activated, or in very rare cases repressed, when bound to N-acyl homoserine lactones, which are used as quorum sensing molecules in a variety of Gram-negative bacteria []: V. fischeri luxR (activates bioluminescence operon) Agrobacterium tumefaciens traR (regulation of Ti plasmid transfer) Erwinia carotovora carR (control of carbapenem antibiotics biosynthesis) E. carotovora expR (virulence factor for soft rot disease; activates plant tissue macerating enzyme genes) Pseudomonas aeruginosa lasR (activates elastase gene lasB) Erwinia chrysanthemi echR and Erwinia stewartii esaR Pseudomonas chlororaphis phzR (positive regulator of phenazine antibiotic production) Pseudomonas aeruginosa rhlR (activates rhlAB operon and lasB gene) 3) Autonomous effector domain regulators, without a regulatory domain, represented by gerE []. B. subtilis gerE (transcription activator and repressor for the regulation of spore formation) 4) Multiple ligand-binding regulators, exemplified by malT []. E. coli malT (activates maltose operon; MalT binds ATP and maltotriose); GO: 0003700 sequence-specific DNA binding transcription factor activity, 0043565 sequence-specific DNA binding, 0006355 regulation of transcription, DNA-dependent, 0005622 intracellular; PDB: 3SZT_A 3CLO_A 1H0M_A 1L3L_A 3C57_B 1ZLK_B 1ZLJ_H 3C3W_B 1RNL_A 1ZG1_A ....
Probab=21.43 E-value=1.4e+02 Score=20.76 Aligned_cols=45 Identities=16% Similarity=0.037 Sum_probs=33.7
Q ss_pred cccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchhhh
Q 024359 12 FRFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRYAI 74 (268)
Q Consensus 12 t~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~k~ 74 (268)
..||+.|+.-|.-+..-. ...++|+.+|+|+ +-|..+..|=+.|.
T Consensus 2 ~~LT~~E~~vl~~l~~G~--------~~~eIA~~l~is~----------~tV~~~~~~i~~Kl 46 (58)
T PF00196_consen 2 PSLTERELEVLRLLAQGM--------SNKEIAEELGISE----------KTVKSHRRRIMKKL 46 (58)
T ss_dssp GSS-HHHHHHHHHHHTTS---------HHHHHHHHTSHH----------HHHHHHHHHHHHHH
T ss_pred CccCHHHHHHHHHHHhcC--------CcchhHHhcCcch----------hhHHHHHHHHHHHh
Confidence 368999999998887764 4689999999764 88998877755554
No 102
>PRK00523 hypothetical protein; Provisional
Probab=20.95 E-value=1.6e+02 Score=23.14 Aligned_cols=36 Identities=14% Similarity=0.232 Sum_probs=31.3
Q ss_pred HHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhh
Q 024359 20 TEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWN 65 (268)
Q Consensus 20 ~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~ 65 (268)
..||+.|++ |+.++.+....+....|-.| +++||+.
T Consensus 28 k~~~k~l~~--NPpine~mir~M~~QMGqKP--------Sekki~Q 63 (72)
T PRK00523 28 KMFKKQIRE--NPPITENMIRAMYMQMGRKP--------SESQIKQ 63 (72)
T ss_pred HHHHHHHHH--CcCCCHHHHHHHHHHhCCCc--------cHHHHHH
Confidence 468999999 69999999999999999876 7888874
No 103
>PRK15451 tRNA cmo(5)U34 methyltransferase; Provisional
Probab=20.87 E-value=2e+02 Score=25.73 Aligned_cols=47 Identities=19% Similarity=0.213 Sum_probs=36.2
Q ss_pred ccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhhcchh
Q 024359 13 RFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQNRRY 72 (268)
Q Consensus 13 ~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQNRR~ 72 (268)
.+|..|+.++.+.++.. -...+.+.-.+|.++-|. ++|..|||+--.
T Consensus 189 g~s~~ei~~~~~~~~~~-~~~~~~~~~~~~L~~aGF------------~~v~~~~~~~~f 235 (247)
T PRK15451 189 GYSELEISQKRSMLENV-MLTDSVETHKARLHKAGF------------EHSELWFQCFNF 235 (247)
T ss_pred CCCHHHHHHHHHHHHhh-cccCCHHHHHHHHHHcCc------------hhHHHHHHHHhH
Confidence 67888888887777663 244588888889999886 679999998543
No 104
>PF14773 VIGSSK: Helicase-associated putative binding domain, C-terminal
Probab=20.61 E-value=46 Score=25.35 Aligned_cols=15 Identities=40% Similarity=0.521 Sum_probs=12.2
Q ss_pred EEEEccCCccccccc
Q 024359 251 LVRYDHDQSEVATTS 265 (268)
Q Consensus 251 ~Vr~~hd~sEe~v~~ 265 (268)
-|.|.|+|+|.+-+|
T Consensus 36 gV~YtH~N~eVIGsS 50 (61)
T PF14773_consen 36 GVEYTHSNQEVIGSS 50 (61)
T ss_pred ceeeeecCcceeccH
Confidence 488999999887766
No 105
>PF04545 Sigma70_r4: Sigma-70, region 4; InterPro: IPR007630 The bacterial core RNA polymerase complex, which consists of five subunits, is sufficient for transcription elongation and termination but is unable to initiate transcription. Transcription initiation from promoter elements requires a sixth, dissociable subunit called a sigma factor, which reversibly associates with the core RNA polymerase complex to form a holoenzyme []. RNA polymerase recruits alternative sigma factors as a means of switching on specific regulons. Most bacteria express a multiplicity of sigma factors. Two of these factors, sigma-70 (gene rpoD), generally known as the major or primary sigma factor, and sigma-54 (gene rpoN or ntrA) direct the transcription of a wide variety of genes. The other sigma factors, known as alternative sigma factors, are required for the transcription of specific subsets of genes. With regard to sequence similarity, sigma factors can be grouped into two classes, the sigma-54 and sigma-70 families. Sequence alignments of the sigma70 family members reveal four conserved regions that can be further divided into subregions eg. sub-region 2.2, which may be involved in the binding of the sigma factor to the core RNA polymerase; and sub-region 4.2, which seems to harbor a DNA-binding 'helix-turn-helix' motif involved in binding the conserved -35 region of promoters recognised by the major sigma factors [, ]. Region 4 of sigma-70 like sigma-factors is involved in binding to the -35 promoter element via a helix-turn-helix motif []. Due to the way Pfam works, the threshold has been set artificially high to prevent overlaps with other helix-turn-helix families. Therefore there are many false negatives.; GO: 0003677 DNA binding, 0003700 sequence-specific DNA binding transcription factor activity, 0016987 sigma factor activity, 0006352 transcription initiation, DNA-dependent, 0006355 regulation of transcription, DNA-dependent; PDB: 2P7V_B 3IYD_F 1TLH_B 1KU7_A 1RIO_H 3N97_A 1KU3_A 1RP3_C 1SC5_A 1NR3_A ....
Probab=20.30 E-value=1.9e+02 Score=19.51 Aligned_cols=39 Identities=21% Similarity=0.176 Sum_probs=26.6
Q ss_pred ccCHHHHHHHHHHHHhccCCCCCHHHHHHHHHHhCCCccccCCcccccchhhhhhh
Q 024359 13 RFNPAEVTEMEGILQEHHNAMPSREILVALAEKFSESPERKGKIMVQMKQVWNWFQ 68 (268)
Q Consensus 13 ~FT~~Qv~eLEk~F~~~~~~yp~~~~rq~LA~~fnlS~~RaGK~~lt~kQVk~WFQ 68 (268)
.+++.|..-|.-.|-+ .-..+++|+.+|+|. ..|+.+..
T Consensus 4 ~L~~~er~vi~~~y~~-------~~t~~eIa~~lg~s~----------~~V~~~~~ 42 (50)
T PF04545_consen 4 QLPPREREVIRLRYFE-------GLTLEEIAERLGISR----------STVRRILK 42 (50)
T ss_dssp TS-HHHHHHHHHHHTS-------T-SHHHHHHHHTSCH----------HHHHHHHH
T ss_pred hCCHHHHHHHHHHhcC-------CCCHHHHHHHHCCcH----------HHHHHHHH
Confidence 4677777777777744 345789999999874 66776544
Done!