Query 032798
Match_columns 133
No_of_seqs 214 out of 1032
Neff 7.5
Searched_HMMs 46136
Date Fri Mar 29 06:05:55 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/032798.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/032798hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PTZ00199 high mobility group p 99.9 1.6E-25 3.6E-30 149.3 11.3 80 29-109 12-93 (94)
2 cd01389 MATA_HMG-box MATA_HMG- 99.9 2.7E-22 5.9E-27 128.9 7.8 72 39-111 1-72 (77)
3 cd01388 SOX-TCF_HMG-box SOX-TC 99.9 9.1E-22 2E-26 125.0 7.7 70 39-109 1-70 (72)
4 PF00505 HMG_box: HMG (high mo 99.9 5.6E-21 1.2E-25 119.6 9.1 69 40-109 1-69 (69)
5 cd01390 HMGB-UBF_HMG-box HMGB- 99.8 1.8E-20 3.8E-25 116.2 9.0 65 40-105 1-65 (66)
6 PF09011 HMG_box_2: HMG-box do 99.8 2.1E-20 4.6E-25 118.9 9.1 72 37-109 1-73 (73)
7 smart00398 HMG high mobility g 99.8 2.4E-20 5.1E-25 116.4 9.0 70 39-109 1-70 (70)
8 COG5648 NHP6B Chromatin-associ 99.8 2.2E-20 4.8E-25 138.6 7.2 89 29-118 60-148 (211)
9 KOG0381 HMG box-containing pro 99.8 2.8E-18 6.1E-23 114.0 11.0 76 36-112 17-95 (96)
10 cd00084 HMG-box High Mobility 99.8 1.8E-18 3.8E-23 106.7 9.0 65 40-105 1-65 (66)
11 KOG0527 HMG-box transcription 99.8 7.8E-19 1.7E-23 139.6 5.8 79 32-111 55-133 (331)
12 KOG0526 Nucleosome-binding fac 99.7 1E-17 2.2E-22 138.1 8.6 80 25-109 521-600 (615)
13 KOG3248 Transcription factor T 99.3 2.1E-12 4.6E-17 102.0 6.7 72 39-111 191-262 (421)
14 KOG4715 SWI/SNF-related matrix 99.3 1.9E-11 4E-16 96.0 7.6 78 32-110 57-134 (410)
15 KOG0528 HMG-box transcription 99.2 1.6E-11 3.4E-16 100.8 2.9 77 34-111 320-396 (511)
16 KOG2746 HMG-box transcription 98.7 1.2E-08 2.6E-13 86.7 4.7 85 19-104 161-247 (683)
17 PF14887 HMG_box_5: HMG (high 98.5 8.7E-07 1.9E-11 56.6 7.2 75 39-115 3-77 (85)
18 PF04690 YABBY: YABBY protein; 97.6 0.00021 4.5E-09 52.4 5.6 49 34-83 116-164 (170)
19 PF06382 DUF1074: Protein of u 97.5 0.00041 8.9E-09 50.9 7.1 49 44-97 83-131 (183)
20 COG5648 NHP6B Chromatin-associ 97.3 0.00023 5E-09 53.5 2.9 68 38-106 142-209 (211)
21 PF08073 CHDNT: CHDNT (NUC034) 96.4 0.0039 8.5E-08 37.4 3.0 40 44-84 13-52 (55)
22 PF04769 MAT_Alpha1: Mating-ty 94.7 0.14 3E-06 38.6 6.5 56 33-95 37-92 (201)
23 PF06244 DUF1014: Protein of u 94.0 0.093 2E-06 36.5 3.9 49 36-85 69-117 (122)
24 TIGR03481 HpnM hopanoid biosyn 90.5 0.89 1.9E-05 34.0 5.7 46 66-111 64-111 (198)
25 PRK15117 ABC transporter perip 87.6 2 4.3E-05 32.4 5.8 46 66-111 68-115 (211)
26 PF12881 NUT_N: NUT protein N 87.3 3.5 7.5E-05 33.1 7.1 72 44-117 229-302 (328)
27 KOG3223 Uncharacterized conser 84.1 1.6 3.5E-05 32.7 3.7 53 38-94 162-215 (221)
28 PF05494 Tol_Tol_Ttg2: Toluene 82.6 5.1 0.00011 28.7 5.8 45 66-110 38-84 (170)
29 COG2854 Ttg2D ABC-type transpo 76.7 3.5 7.6E-05 31.1 3.4 48 73-120 78-126 (202)
30 PF13875 DUF4202: Domain of un 72.1 9.5 0.00021 28.4 4.7 39 46-88 131-169 (185)
31 PF11304 DUF3106: Protein of u 69.7 23 0.00051 23.7 5.9 41 70-110 11-58 (107)
32 PF01352 KRAB: KRAB box; Inte 60.1 7.1 0.00015 21.8 1.6 28 68-95 3-31 (41)
33 PF06945 DUF1289: Protein of u 53.3 21 0.00045 20.7 2.9 25 67-96 23-47 (51)
34 PRK09706 transcriptional repre 50.7 54 0.0012 22.4 5.2 43 70-112 87-129 (135)
35 PRK12750 cpxP periplasmic repr 46.8 64 0.0014 23.5 5.2 32 74-105 129-160 (170)
36 PRK12751 cpxP periplasmic stre 45.8 55 0.0012 23.8 4.7 29 73-101 121-149 (162)
37 PRK10236 hypothetical protein; 45.2 21 0.00047 27.6 2.6 25 71-95 118-142 (237)
38 PF12650 DUF3784: Domain of un 42.6 19 0.0004 23.4 1.7 16 78-93 25-40 (97)
39 PRK10363 cpxP periplasmic repr 41.3 85 0.0018 23.0 5.1 33 70-102 112-144 (166)
40 COG1638 DctP TRAP-type C4-dica 40.4 76 0.0016 25.6 5.2 35 76-110 244-278 (332)
41 TIGR00787 dctP tripartite ATP- 40.1 81 0.0018 23.8 5.2 28 76-103 213-240 (257)
42 KOG1610 Corticosteroid 11-beta 38.6 85 0.0018 25.4 5.1 51 49-99 187-249 (322)
43 PF09164 VitD-bind_III: Vitami 35.1 1.1E+02 0.0025 19.0 4.9 32 45-77 9-40 (68)
44 PF00887 ACBP: Acyl CoA bindin 34.6 1.2E+02 0.0026 19.1 5.0 53 47-101 30-86 (87)
45 cd07081 ALDH_F20_ACDH_EutE-lik 32.7 1.2E+02 0.0026 25.3 5.4 41 70-110 6-46 (439)
46 PF03480 SBP_bac_7: Bacterial 31.2 1.1E+02 0.0023 23.5 4.6 39 76-114 213-255 (286)
47 KOG2880 SMAD6 interacting prot 30.7 2.3E+02 0.005 23.6 6.4 64 44-110 52-117 (424)
48 PF15581 Imm35: Immunity prote 30.6 1.1E+02 0.0023 20.2 3.7 24 67-90 31-54 (93)
49 PF09655 Nitr_red_assoc: Conse 30.5 1.1E+02 0.0024 21.9 4.2 45 77-121 33-80 (144)
50 smart00271 DnaJ DnaJ molecular 29.3 94 0.002 17.6 3.2 33 53-85 21-58 (60)
51 cd00225 API3 Ascaris pepsin in 28.7 91 0.002 22.5 3.5 29 77-106 26-54 (159)
52 PF01297 TroA: Periplasmic sol 28.7 1.5E+02 0.0034 22.2 5.1 49 68-116 101-149 (256)
53 KOG0493 Transcription factor E 27.7 2E+02 0.0044 22.9 5.5 24 38-64 244-267 (342)
54 KOG3838 Mannose lectin ERGIC-5 27.6 66 0.0014 27.1 2.9 37 82-118 269-305 (497)
55 cd07133 ALDH_CALDH_CalB Conife 27.1 1.9E+02 0.0042 23.9 5.7 40 70-109 5-44 (434)
56 cd07132 ALDH_F3AB Aldehyde deh 26.6 1.8E+02 0.004 24.1 5.5 40 70-109 5-44 (443)
57 PF05388 Carbpep_Y_N: Carboxyp 26.5 1.2E+02 0.0026 20.7 3.7 29 68-96 45-73 (113)
58 cd01145 TroA_c Periplasmic bin 25.6 2.4E+02 0.0051 20.6 5.5 48 68-115 117-164 (203)
59 PF06394 Pepsin-I3: Pepsin inh 25.5 96 0.0021 19.7 2.8 27 80-114 38-64 (76)
60 KOG1827 Chromatin remodeling c 25.3 4.4 9.6E-05 35.5 -4.4 44 43-87 552-595 (629)
61 cd01137 PsaA Metal binding pro 24.9 2.3E+02 0.005 22.0 5.5 47 68-114 126-172 (287)
62 cd07122 ALDH_F20_ACDH Coenzyme 24.5 2E+02 0.0043 24.0 5.3 40 71-110 7-46 (436)
63 cd08317 Death_ank Death domain 24.4 38 0.00081 21.4 0.8 19 66-84 5-23 (84)
64 PF02026 RyR: RyR domain; Int 24.2 71 0.0015 20.9 2.1 19 79-97 61-79 (94)
65 PRK10455 periplasmic protein; 24.0 1.6E+02 0.0036 21.2 4.2 25 72-96 120-144 (161)
66 PF13945 NST1: Salt tolerance 23.1 2E+02 0.0044 21.5 4.6 26 68-93 100-125 (190)
67 cd07085 ALDH_F6_MMSDH Methylma 23.0 2.3E+02 0.0049 23.7 5.4 37 72-108 47-83 (478)
68 cd07087 ALDH_F3-13-14_CALDH-li 23.0 2.4E+02 0.0051 23.2 5.5 40 70-109 5-44 (426)
69 PF07813 LTXXQ: LTXXQ motif fa 22.5 1.5E+02 0.0033 18.4 3.5 25 69-93 75-99 (100)
70 PF09791 Oxidored-like: Oxidor 22.2 1.4E+02 0.0031 17.2 2.9 15 94-108 31-45 (48)
71 PF08367 M16C_assoc: Peptidase 21.8 1.8E+02 0.004 22.0 4.3 32 69-100 13-44 (248)
72 PRK13252 betaine aldehyde dehy 21.1 2.5E+02 0.0054 23.5 5.3 38 72-109 53-90 (488)
73 cd07150 ALDH_VaniDH_like Pseud 21.0 2.5E+02 0.0055 23.1 5.3 37 72-108 30-66 (451)
74 PHA03102 Small T antigen; Revi 20.7 1.5E+02 0.0032 21.3 3.4 36 52-87 26-61 (153)
75 PTZ00037 DnaJ_C chaperone prot 20.6 2.2E+02 0.0048 23.8 4.8 42 51-92 46-88 (421)
76 TIGR02664 nitr_red_assoc conse 20.2 2.5E+02 0.0054 20.1 4.4 44 77-120 33-79 (145)
77 cd07152 ALDH_BenzADH NAD-depen 20.1 2.8E+02 0.0061 22.8 5.4 37 72-108 22-58 (443)
No 1
>PTZ00199 high mobility group protein; Provisional
Probab=99.93 E-value=1.6e-25 Score=149.30 Aligned_cols=80 Identities=40% Similarity=0.726 Sum_probs=75.3
Q ss_pred CCcccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCC--HHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHH
Q 032798 29 AKPKAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKS--VATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQ 106 (133)
Q Consensus 29 ~k~k~~~dp~~PKrP~say~lF~~~~r~~~k~~~p~~~~--~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~ 106 (133)
++++..+||+.|+||+|||++|++++|..|..+||+ ++ +.+|+++||++|+.||+++|.+|+++|+.++.+|..+|.
T Consensus 12 ~~~k~~kdp~~PKrP~sAY~~F~~~~R~~i~~~~P~-~~~~~~evsk~ige~Wk~ls~eeK~~y~~~A~~dk~rY~~e~~ 90 (94)
T PTZ00199 12 KNKRKKKDPNAPKRALSAYMFFAKEKRAEIIAENPE-LAKDVAAVGKMVGEAWNKLSEEEKAPYEKKAQEDKVRYEKEKA 90 (94)
T ss_pred ccCCCCCCCCCCCCCCcHHHHHHHHHHHHHHHHCcC-CcccHHHHHHHHHHHHHcCCHHHHHHHHHHHHHHHHHHHHHHH
Confidence 345668999999999999999999999999999999 64 899999999999999999999999999999999999999
Q ss_pred HHH
Q 032798 107 DYN 109 (133)
Q Consensus 107 ~y~ 109 (133)
+|+
T Consensus 91 ~Y~ 93 (94)
T PTZ00199 91 EYA 93 (94)
T ss_pred HHh
Confidence 995
No 2
>cd01389 MATA_HMG-box MATA_HMG-box, class I member of the HMG-box superfamily of DNA-binding proteins. These proteins contain a single HMG box, and bind the minor groove of DNA in a highly sequence-specific manner. Members include the fungal mating type gene products MC, MATA1 and Ste11.
Probab=99.87 E-value=2.7e-22 Score=128.88 Aligned_cols=72 Identities=28% Similarity=0.476 Sum_probs=69.4
Q ss_pred CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHHh
Q 032798 39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQ 111 (133)
Q Consensus 39 ~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~k 111 (133)
+|+||+||||||+++.|..++.+||+ +++.+|+++||++|+.|++++|++|.++|+.++++|..++++|+..
T Consensus 1 ~~kRP~naf~lf~~~~r~~~~~~~p~-~~~~eisk~~g~~Wk~ls~eeK~~y~~~A~~~k~~~~~~~p~Yky~ 72 (77)
T cd01389 1 KIPRPRNAFILYRQDKHAQLKTENPG-LTNNEISRIIGRMWRSESPEVKAYYKELAEEEKERHAREYPDYKYT 72 (77)
T ss_pred CCCCCCcHHHHHHHHHHHHHHHHCCC-CCHHHHHHHHHHHHhhCCHHHHHHHHHHHHHHHHHHHHHCCCCccc
Confidence 58999999999999999999999999 9999999999999999999999999999999999999999999753
No 3
>cd01388 SOX-TCF_HMG-box SOX-TCF_HMG-box, class I member of the HMG-box superfamily of DNA-binding proteins. These proteins contain a single HMG box, and bind the minor groove of DNA in a highly sequence-specific manner. Members include SRY and its homologs in insects and vertebrates, and transcription factor-like proteins, TCF-1, -3, -4, and LEF-1. They appear to bind the minor groove of the A/T C A A A G/C-motif.
Probab=99.86 E-value=9.1e-22 Score=124.98 Aligned_cols=70 Identities=34% Similarity=0.585 Sum_probs=67.7
Q ss_pred CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHH
Q 032798 39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN 109 (133)
Q Consensus 39 ~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~ 109 (133)
+.|||+|||++|+++.|..++.+||+ +++.+|+++||++|+.||+++|++|.++|+.++++|.+++++|+
T Consensus 1 ~iKrP~naf~~F~~~~r~~~~~~~p~-~~~~eisk~l~~~Wk~ls~~eK~~y~~~a~~~k~~y~~~~p~y~ 70 (72)
T cd01388 1 HIKRPMNAFMLFSKRHRRKVLQEYPL-KENRAISKILGDRWKALSNEEKQPYYEEAKKLKELHMKLYPDYK 70 (72)
T ss_pred CCCCCCcHHHHHHHHHHHHHHHHCCC-CCHHHHHHHHHHHHHcCCHHHHHHHHHHHHHHHHHHHHHCcCCC
Confidence 36899999999999999999999999 99999999999999999999999999999999999999999985
No 4
>PF00505 HMG_box: HMG (high mobility group) box; InterPro: IPR000910 High mobility group (HMG or HMGB) proteins are a family of relatively low molecular weight non-histone components in chromatin. HMG1 (also called HMG-T in fish) and HMG2 are two highly related proteins that bind single-stranded DNA preferentially and unwind double-stranded DNA. Although they have no sequence specificity, they have a high affinity for bent or distorted DNA, and bend linear DNA. HMG1 and HMG2 contain two DNA-binding HMG-box domains (A and B) that show structural and functional differences, and have a long acidic C-terminal domain rich in aspartic and glutamic acid residues. The acidic tail modulates the affinity of the tandem HMG boxes in HMG1 and 2 for a variety of DNA targets. HMG1 and 2 appear to play important architectural roles in the assembly of nucleoprotein complexes in a variety of biological processes, for example V(D)J recombination, the initiation of transcription, and DNA repair []. The profile in this entry describing the HMG-domains is much more general than the signature. In addition to the HMG1 and HMG2 proteins, HMG-domains occur in single or multiple copies in the following protein classes; the SOX family of transcription factors; SRY sex determining region Y protein and related proteins []; LEF1 lymphoid enhancer binding factor 1 []; SSRP recombination signal recognition protein; MTF1 mitochondrial transcription factor 1; UBF1/2 nucleolar transcription factors; Abf2 yeast ARS-binding factor []; and Saccharomyces cerevisiae transcription factors Ixr1, Rox1, Nhp6a, Nhp6b and Spp41.; GO: 0003677 DNA binding; PDB: 1I11_A 1J3C_A 1J3D_A 1WZ6_A 1WGF_A 2D7L_A 1GT0_D 3U2B_C 2CRJ_A 2CS1_A ....
Probab=99.85 E-value=5.6e-21 Score=119.56 Aligned_cols=69 Identities=45% Similarity=0.836 Sum_probs=65.7
Q ss_pred CCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHH
Q 032798 40 PKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN 109 (133)
Q Consensus 40 PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~ 109 (133)
|+||+|||++|+.+.+..++.+||+ ++..+|+++||++|++|++++|.+|.+.|..++.+|..++++|+
T Consensus 1 PkrP~~af~lf~~~~~~~~k~~~p~-~~~~~i~~~~~~~W~~l~~~eK~~y~~~a~~~~~~y~~~~~~y~ 69 (69)
T PF00505_consen 1 PKRPPNAFMLFCKEKRAKLKEENPD-LSNKEISKILAQMWKNLSEEEKAPYKEEAEEEKERYEKEMPEYK 69 (69)
T ss_dssp SSSS--HHHHHHHHHHHHHHHHSTT-STHHHHHHHHHHHHHCSHHHHHHHHHHHHHHHHHHHHHHHHHHH
T ss_pred CcCCCCHHHHHHHHHHHHHHHHhcc-cccccchhhHHHHHhcCCHHHHHHHHHHHHHHHHHHHHHHHhcC
Confidence 8999999999999999999999999 99999999999999999999999999999999999999999995
No 5
>cd01390 HMGB-UBF_HMG-box HMGB-UBF_HMG-box, class II and III members of the HMG-box superfamily of DNA-binding proteins. These proteins bind the minor groove of DNA in a non-sequence specific fashion and contain two or more tandem HMG boxes. Class II members include non-histone chromosomal proteins, HMG1 and HMG2, which bind to bent or distorted DNA such as four-way DNA junctions, synthetic DNA cruciforms, kinked cisplatin-modified DNA, DNA bulges, cross-overs in supercoiled DNA, and can cause looping of linear DNA. Class III members include nucleolar and mitochondrial transcription factors, UBF and mtTF1, which bind four-way DNA junctions.
Probab=99.84 E-value=1.8e-20 Score=116.20 Aligned_cols=65 Identities=51% Similarity=0.854 Sum_probs=63.7
Q ss_pred CCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHH
Q 032798 40 PKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNM 105 (133)
Q Consensus 40 PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~ 105 (133)
||+|+|||++|+++.|..++.+||+ +++.+|++.||++|+.||+++|.+|.+.|++++.+|..+|
T Consensus 1 Pkrp~saf~~f~~~~r~~~~~~~p~-~~~~~i~~~~~~~W~~ls~~eK~~y~~~a~~~~~~y~~e~ 65 (66)
T cd01390 1 PKRPLSAYFLFSQEQRPKLKKENPD-ASVTEVTKILGEKWKELSEEEKKKYEEKAEKDKERYEKEM 65 (66)
T ss_pred CCCCCcHHHHHHHHHHHHHHHHCcC-CCHHHHHHHHHHHHHhCCHHHHHHHHHHHHHHHHHHHHhh
Confidence 8999999999999999999999999 9999999999999999999999999999999999999886
No 6
>PF09011 HMG_box_2: HMG-box domain; InterPro: IPR015101 This domain is predominantly found in Maelstrom homologue proteins. It has no known function. ; GO: 0005634 nucleus; PDB: 2EQZ_A 1V64_A 2CTO_A 1H5P_A 3TQ6_A 3FGH_A 3TMM_A 1J3X_A 2YRQ_A 1AAB_A ....
Probab=99.84 E-value=2.1e-20 Score=118.93 Aligned_cols=72 Identities=49% Similarity=0.876 Sum_probs=63.8
Q ss_pred CCCCCCCCChHHHHHHHHHHHHHHh-CCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHH
Q 032798 37 PNKPKRPPSAFFVFMEEFRKQFKEA-HPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN 109 (133)
Q Consensus 37 p~~PKrP~say~lF~~~~r~~~k~~-~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~ 109 (133)
|++||+|+|||+||+.+++..++.+ ++. ..+.++++.|+..|++||++||.+|.++|++++++|..+|..|+
T Consensus 1 p~kpK~~~say~lF~~~~~~~~k~~G~~~-~~~~e~~k~~~~~Wk~Ls~~EK~~Y~~~A~~~k~~y~~e~~~~~ 73 (73)
T PF09011_consen 1 PKKPKRPPSAYNLFMKEMRKEVKEEGGQK-QSFREVMKEISERWKSLSEEEKEPYEERAKEDKERYEREMKEWN 73 (73)
T ss_dssp SSS--SSSSHHHHHHHHHHHHHHHHT-T--SSHHHHHHHHHHHHHHS-HHHHHHHHHHHHHHHHHHHHHHHHH-
T ss_pred CcCCCCCCCHHHHHHHHHHHHHHHhcccC-CCHHHHHHHHHHHHHhcCHHHHHHHHHHHHHHHHHHHHHHHhcC
Confidence 6899999999999999999999988 665 78999999999999999999999999999999999999999985
No 7
>smart00398 HMG high mobility group.
Probab=99.84 E-value=2.4e-20 Score=116.42 Aligned_cols=70 Identities=47% Similarity=0.831 Sum_probs=67.8
Q ss_pred CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHH
Q 032798 39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN 109 (133)
Q Consensus 39 ~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~ 109 (133)
+|++|+|||++|+++.|..+..+||+ +++.+|++.||.+|+.|++++|.+|.++|+.++++|..+++.|+
T Consensus 1 ~pkrp~~~y~~f~~~~r~~~~~~~~~-~~~~~i~~~~~~~W~~l~~~ek~~y~~~a~~~~~~y~~~~~~y~ 70 (70)
T smart00398 1 KPKRPMSAFMLFSQENRAKIKAENPD-LSNAEISKKLGERWKLLSEEEKAPYEEKAKKDKERYEEEMPEYK 70 (70)
T ss_pred CcCCCCcHHHHHHHHHHHHHHHHCcC-CCHHHHHHHHHHHHHcCCHHHHHHHHHHHHHHHHHHHHHHHhcC
Confidence 58999999999999999999999999 99999999999999999999999999999999999999999883
No 8
>COG5648 NHP6B Chromatin-associated proteins containing the HMG domain [Chromatin structure and dynamics]
Probab=99.82 E-value=2.2e-20 Score=138.61 Aligned_cols=89 Identities=35% Similarity=0.671 Sum_probs=82.8
Q ss_pred CCcccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHH
Q 032798 29 AKPKAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDY 108 (133)
Q Consensus 29 ~k~k~~~dp~~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y 108 (133)
...++.+|||.||||+||||+|+.++|.+++.++|+ +++.++++++|++|++|++++|++|...|..++++|+.++..|
T Consensus 60 ~~~r~k~dpN~PKRp~sayf~y~~~~R~ei~~~~p~-l~~~e~~k~~~e~WK~Ltd~eke~y~k~~~~~~erYq~ek~~y 138 (211)
T COG5648 60 RLVRKKKDPNGPKRPLSAYFLYSAENRDEIRKENPK-LTFGEVGKLLSEKWKELTDEEKEPYYKEANSDRERYQREKEEY 138 (211)
T ss_pred HHHHHhcCCCCCCCchhHHHHHHHHHHHHHHHhCCC-CChHHHHHHHHHHHHhccHhhhhhHHHHHhhHHHHHHHHHHhh
Confidence 446788999999999999999999999999999999 8999999999999999999999999999999999999999999
Q ss_pred HHhhhhhhcc
Q 032798 109 NKQLVIFFGI 118 (133)
Q Consensus 109 ~~k~~~~~~~ 118 (133)
..++.....+
T Consensus 139 ~~k~~~~~~~ 148 (211)
T COG5648 139 NKKLPNKAPI 148 (211)
T ss_pred hcccCCCCCC
Confidence 9987654443
No 9
>KOG0381 consensus HMG box-containing protein [General function prediction only]
Probab=99.78 E-value=2.8e-18 Score=113.99 Aligned_cols=76 Identities=47% Similarity=0.840 Sum_probs=72.4
Q ss_pred CC--CCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHH-HHHHhh
Q 032798 36 DP--NKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQ-DYNKQL 112 (133)
Q Consensus 36 dp--~~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~-~y~~k~ 112 (133)
|| +.|++|++||++|+.+.+..++.+||+ +++.+|++++|++|++|++++|.+|...|..++++|..+|. .|+..+
T Consensus 17 ~p~~~~pkrp~sa~~~f~~~~~~~~k~~~p~-~~~~~v~k~~g~~W~~l~~~~k~~y~~ka~~~k~~Y~~~~~~~~~~~~ 95 (96)
T KOG0381|consen 17 DPNAQAPKRPLSAFFLFSSEQRSKIKAENPG-LSVGEVAKALGEMWKNLAEEEKQPYEEKASKLKEKYEKELAGEYKASL 95 (96)
T ss_pred CCCCCCCCCCCcHHHHHHHHHHHHHHHhCCC-CCHHHHHHHHHHHHhcCCHHHHHHHHHHHHHHHHHHHHHHHHHHhhcc
Confidence 66 599999999999999999999999999 99999999999999999999999999999999999999999 887654
No 10
>cd00084 HMG-box High Mobility Group (HMG)-box is found in a variety of eukaryotic chromosomal proteins and transcription factors. HMGs bind to the minor groove of DNA and have been classified by DNA binding preferences. Two phylogenically distinct groups of Class I proteins bind DNA in a sequence specific fashion and contain a single HMG box. One group (SOX-TCF) includes transcription factors, TCF-1, -3, -4; and also SRY and LEF-1, which bind four-way DNA junctions and duplex DNA targets. The second group (MATA) includes fungal mating type gene products MC, MATA1 and Ste11. Class II and III proteins (HMGB-UBF) bind DNA in a non-sequence specific fashion and contain two or more tandem HMG boxes. Class II members include non-histone chromosomal proteins, HMG1 and HMG2, which bind to bent or distorted DNA such as four-way DNA junctions, synthetic DNA cruciforms, kinked cisplatin-modified DNA, DNA bulges, cross-overs in supercoiled DNA, and can cause looping of linear DNA. Class III member
Probab=99.78 E-value=1.8e-18 Score=106.74 Aligned_cols=65 Identities=49% Similarity=0.827 Sum_probs=63.0
Q ss_pred CCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHH
Q 032798 40 PKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNM 105 (133)
Q Consensus 40 PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~ 105 (133)
|++|+|||++|+++.|..+..+||+ +++.+|++.||.+|+.|++++|.+|.+.|+.++.+|..++
T Consensus 1 pkrp~~af~~f~~~~~~~~~~~~~~-~~~~~i~~~~~~~W~~l~~~~k~~y~~~a~~~~~~y~~~~ 65 (66)
T cd00084 1 PKRPLSAYFLFSQEHRAEVKAENPG-LSVGEISKILGEMWKSLSEEEKKKYEEKAEKDKERYEKEM 65 (66)
T ss_pred CCCCCcHHHHHHHHHHHHHHHHCcC-CCHHHHHHHHHHHHHhCCHHHHHHHHHHHHHHHHHHHHhh
Confidence 7999999999999999999999999 9999999999999999999999999999999999998875
No 11
>KOG0527 consensus HMG-box transcription factor [Transcription]
Probab=99.76 E-value=7.8e-19 Score=139.56 Aligned_cols=79 Identities=32% Similarity=0.602 Sum_probs=74.8
Q ss_pred ccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHHh
Q 032798 32 KAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQ 111 (133)
Q Consensus 32 k~~~dp~~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~k 111 (133)
..+...++.||||||||+|.+..|.+|..+||+ +.+.||+++||.+|+.|+++||.+|+++|++++..|+++.++|+.+
T Consensus 55 ~~k~~~~hIKRPMNAFMVWSq~~RRkma~qnP~-mHNSEISK~LG~~WK~Lse~EKrPFi~EAeRLR~~HmkehPdYKYR 133 (331)
T KOG0527|consen 55 KDKTSTDRIKRPMNAFMVWSQGQRRKLAKQNPK-MHNSEISKRLGAEWKLLSEEEKRPFVDEAERLRAQHMKEYPDYKYR 133 (331)
T ss_pred cCCCCccccCCCcchhhhhhHHHHHHHHHhCcc-hhhHHHHHHHHHHHhhcCHhhhccHHHHHHHHHHHHHHhCCCcccc
Confidence 345667899999999999999999999999999 9999999999999999999999999999999999999999999774
No 12
>KOG0526 consensus Nucleosome-binding factor SPN, POB3 subunit [Transcription; Replication, recombination and repair; Chromatin structure and dynamics]
Probab=99.73 E-value=1e-17 Score=138.13 Aligned_cols=80 Identities=40% Similarity=0.684 Sum_probs=74.6
Q ss_pred CCCCCCcccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHH
Q 032798 25 GKRTAKPKAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKN 104 (133)
Q Consensus 25 ~kk~~k~k~~~dp~~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e 104 (133)
+++.++.++.+|||+|||++||||+|++..|..++.+ + .++++|++.+|++|+.|+. |.+|++.|+.++++|+.+
T Consensus 521 ~~~~k~~kk~kdpnapkra~sa~m~w~~~~r~~ik~d--g-i~~~dv~kk~g~~wk~ms~--k~~we~ka~~dk~ry~~e 595 (615)
T KOG0526|consen 521 KEKKKKGKKKKDPNAPKRATSAYMLWLNASRESIKED--G-ISVGDVAKKAGEKWKQMSA--KEEWEDKAAVDKQRYEDE 595 (615)
T ss_pred hccccCcccCCCCCCCccchhHHHHHHHhhhhhHhhc--C-chHHHHHHHHhHHHhhhcc--cchhhHHHHHHHHHHHHH
Confidence 3344677889999999999999999999999999987 5 8999999999999999999 999999999999999999
Q ss_pred HHHHH
Q 032798 105 MQDYN 109 (133)
Q Consensus 105 ~~~y~ 109 (133)
|.+|+
T Consensus 596 m~~yk 600 (615)
T KOG0526|consen 596 MKEYK 600 (615)
T ss_pred HHhhc
Confidence 99998
No 13
>KOG3248 consensus Transcription factor TCF-4 [Transcription]
Probab=99.34 E-value=2.1e-12 Score=101.99 Aligned_cols=72 Identities=24% Similarity=0.456 Sum_probs=67.3
Q ss_pred CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHHh
Q 032798 39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQ 111 (133)
Q Consensus 39 ~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~k 111 (133)
..|+|+||||+|++|.|..|..++-- ....+|.++||++|..||.||..+|.++|+++++-+.+.+..|-+.
T Consensus 191 hiKKPLNAFmlyMKEmRa~vvaEctl-KeSAaiNqiLGrRWH~LSrEEQAKYyElArKerqlH~qlYP~WSAR 262 (421)
T KOG3248|consen 191 HIKKPLNAFMLYMKEMRAKVVAECTL-KESAAINQILGRRWHALSREEQAKYYELARKERQLHMQLYPGWSAR 262 (421)
T ss_pred cccccHHHHHHHHHHHHHHHHHHhhh-hhHHHHHHHHhHHHhhhhHHHHHHHHHHHHHHHHHHHHhcCCcchh
Confidence 67999999999999999999999875 6788999999999999999999999999999999999999988664
No 14
>KOG4715 consensus SWI/SNF-related matrix-associated actin-dependent regulator of chromatin [Chromatin structure and dynamics]
Probab=99.26 E-value=1.9e-11 Score=96.05 Aligned_cols=78 Identities=24% Similarity=0.571 Sum_probs=73.3
Q ss_pred ccCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHH
Q 032798 32 KAAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNK 110 (133)
Q Consensus 32 k~~~dp~~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~ 110 (133)
...+.|.+|-+|+-+||.|++..+++++..||+ +..-+|.++||.+|..|+++||+.|...++.++..|.+.|..|..
T Consensus 57 t~pkpPkppekpl~pymrySrkvWd~VkA~nPe-~kLWeiGK~Ig~mW~dLpd~EK~ey~~EYeaEKieY~~smkayh~ 134 (410)
T KOG4715|consen 57 TRPKPPKPPEKPLMPYMRYSRKVWDQVKASNPE-LKLWEIGKIIGGMWLDLPDEEKQEYLNEYEAEKIEYNESMKAYHN 134 (410)
T ss_pred cCCCCCCCCCcccchhhHHhhhhhhhhhccCcc-hHHHHHHHHHHHHHhhCcchHHHHHHHHHHHHHHHHHHHHHHhhC
Confidence 344567888999999999999999999999999 999999999999999999999999999999999999999999965
No 15
>KOG0528 consensus HMG-box transcription factor SOX5 [Transcription]
Probab=99.16 E-value=1.6e-11 Score=100.78 Aligned_cols=77 Identities=29% Similarity=0.510 Sum_probs=70.8
Q ss_pred CCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHHh
Q 032798 34 AKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQ 111 (133)
Q Consensus 34 ~~dp~~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~k 111 (133)
...+...||||||||+|.++.|..|...+|| +.+..|+++||.+|+.|+..||++|.++-.++-..|.+..++|+.+
T Consensus 320 ~ss~PHIKRPMNAFMVWAkDERRKILqA~PD-MHNSnISKILGSRWKaMSN~eKQPYYEEQaRLSk~HlEk~PdYrYk 396 (511)
T KOG0528|consen 320 ASSEPHIKRPMNAFMVWAKDERRKILQAFPD-MHNSNISKILGSRWKAMSNTEKQPYYEEQARLSKLHLEKYPDYRYK 396 (511)
T ss_pred CCCCccccCCcchhhcccchhhhhhhhcCcc-ccccchhHHhcccccccccccccchHHHHHHHHHhhhccCcccccC
Confidence 3445677999999999999999999999999 8899999999999999999999999999888888999999999775
No 16
>KOG2746 consensus HMG-box transcription factor Capicua and related proteins [Transcription]
Probab=98.73 E-value=1.2e-08 Score=86.73 Aligned_cols=85 Identities=24% Similarity=0.439 Sum_probs=73.6
Q ss_pred ccCccCCCCCCCcccCCCCCCCCCCCChHHHHHHHHH--HHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHH
Q 032798 19 SKGARAGKRTAKPKAAKDPNKPKRPPSAFFVFMEEFR--KQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEK 96 (133)
Q Consensus 19 ~~~~~~~kk~~k~k~~~dp~~PKrP~say~lF~~~~r--~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~ 96 (133)
+.......+..+...+++..+.++|||||++|++.+| ..+...||+ .++..|+++||+.|-.|.+.||+.|.++|.+
T Consensus 161 kEqdsSs~kdgrspnkr~k~HirrPMnaf~ifskrhr~~g~vhq~~pn-~DNrtIskiLgewWytL~~~Ekq~yhdLa~Q 239 (683)
T KOG2746|consen 161 KEQDSSSEKDGRSPNKRDKDHIRRPMNAFHIFSKRHRGEGRVHQRHPN-QDNRTISKILGEWWYTLGPNEKQKYHDLAFQ 239 (683)
T ss_pred hhhccccccccCCCCcCcchhhhhhhHHHHHHHhhcCCccchhccCcc-ccchhHHHHHhhhHhhhCchhhhhHHHHHHH
Confidence 3333444445555667788889999999999999999 889999999 9999999999999999999999999999999
Q ss_pred HHHHHHHH
Q 032798 97 RKSDYNKN 104 (133)
Q Consensus 97 ~k~~y~~e 104 (133)
.++.|.++
T Consensus 240 vk~Ahfka 247 (683)
T KOG2746|consen 240 VKEAHFKA 247 (683)
T ss_pred HHHHHhhh
Confidence 99998876
No 17
>PF14887 HMG_box_5: HMG (high mobility group) box 5; PDB: 1L8Y_A 1L8Z_A 2HDZ_A.
Probab=98.50 E-value=8.7e-07 Score=56.56 Aligned_cols=75 Identities=16% Similarity=0.327 Sum_probs=60.6
Q ss_pred CCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHHhhhhh
Q 032798 39 KPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQLVIF 115 (133)
Q Consensus 39 ~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~k~~~~ 115 (133)
-|-.|.+|--+|.+.....+...++. ....+ .+.+...|++|++.+|.+|...|.++..+|+.+|.+|+.-....
T Consensus 3 lPE~PKt~qe~Wqq~vi~dYla~~~~-dr~K~-~kam~~~W~~me~Kekl~WIkKA~EdqKrYE~el~e~r~~~~~~ 77 (85)
T PF14887_consen 3 LPETPKTAQEIWQQSVIGDYLAKFRN-DRKKA-LKAMEAQWSQMEKKEKLKWIKKAAEDQKRYERELREMRSAPADA 77 (85)
T ss_dssp -S----THHHHHHHHHHHHHHHHTTS-THHHH-HHHHHHHHHTTGGGHHHHHHHHHHHHHHHHHHHHHCCS-CCCTT
T ss_pred CCCCCCCHHHHHHHHHHHHHHHHhhH-hHHHH-HHHHHHHHHHhhhhhhhHHHHHHHHHHHHHHHHHHHHhcCCCCC
Confidence 47788999999999999999999987 44444 56889999999999999999999999999999999998765543
No 18
>PF04690 YABBY: YABBY protein; InterPro: IPR006780 YABBY proteins are a group of plant-specific transcription factors involved in the specification of abaxial polarity in lateral organs such as leaves and floral organs [, ].
Probab=97.56 E-value=0.00021 Score=52.38 Aligned_cols=49 Identities=33% Similarity=0.498 Sum_probs=43.3
Q ss_pred CCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCC
Q 032798 34 AKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMS 83 (133)
Q Consensus 34 ~~dp~~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls 83 (133)
.+.|.+-.|-+|||..|+++....|+..+|+ ++..|.....+..|...+
T Consensus 116 ~kPPEKRqR~psaYn~f~k~ei~rik~~~p~-ishkeaFs~aAknW~h~p 164 (170)
T PF04690_consen 116 NKPPEKRQRVPSAYNRFMKEEIQRIKAENPD-ISHKEAFSAAAKNWAHFP 164 (170)
T ss_pred cCCccccCCCchhHHHHHHHHHHHHHhcCCC-CCHHHHHHHHHHhhhhCc
Confidence 3445556678999999999999999999999 999999999999998765
No 19
>PF06382 DUF1074: Protein of unknown function (DUF1074); InterPro: IPR024460 This family consists of several proteins which appear to be specific to Insecta. The function of this family is unknown.
Probab=97.55 E-value=0.00041 Score=50.92 Aligned_cols=49 Identities=29% Similarity=0.440 Sum_probs=42.4
Q ss_pred CChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHH
Q 032798 44 PSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKR 97 (133)
Q Consensus 44 ~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~ 97 (133)
-+||+-|+.+++. .|.+ ++..|+....+..|..|++.+|..|..++...
T Consensus 83 nnaYLNFLReFRr----kh~~-L~p~dlI~~AAraW~rLSe~eK~rYrr~~~~~ 131 (183)
T PF06382_consen 83 NNAYLNFLREFRR----KHCG-LSPQDLIQRAARAWCRLSEAEKNRYRRMAPSV 131 (183)
T ss_pred chHHHHHHHHHHH----HccC-CCHHHHHHHHHHHHHhCCHHHHHHHHhhcchh
Confidence 5789999998876 4566 89999999999999999999999999876543
No 20
>COG5648 NHP6B Chromatin-associated proteins containing the HMG domain [Chromatin structure and dynamics]
Probab=97.25 E-value=0.00023 Score=53.46 Aligned_cols=68 Identities=19% Similarity=0.393 Sum_probs=61.1
Q ss_pred CCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHH
Q 032798 38 NKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQ 106 (133)
Q Consensus 38 ~~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~ 106 (133)
.+++.|..+|+-+-...|+.+...+|+ ....+++++++..|.+|++.-|.+|.+.+..++..|...++
T Consensus 142 ~~~~~~~~~~~e~~~~~r~~~~~~~~~-~~~~e~~k~~~~~w~el~~skK~~~~~~~Kk~k~~~~~~~~ 209 (211)
T COG5648 142 LPNKAPIGPFIENEPKIRPKVEGPSPD-KALVEETKIISKAWSELDESKKKKYIDKYKKLKEEYDSFYP 209 (211)
T ss_pred cCCCCCCchhhhccHHhccccCCCCcc-hhhhHHhhhhhhhhhhhChhhhhHHHHHHHHHHHHHhhhcc
Confidence 466788888999989999999999998 88999999999999999999999999999999998876654
No 21
>PF08073 CHDNT: CHDNT (NUC034) domain; InterPro: IPR012958 The CHD N-terminal domain is found in PHD/RING fingers and chromo domain-associated helicases [].; GO: 0003677 DNA binding, 0005524 ATP binding, 0008270 zinc ion binding, 0016818 hydrolase activity, acting on acid anhydrides, in phosphorus-containing anhydrides, 0006355 regulation of transcription, DNA-dependent, 0005634 nucleus
Probab=96.41 E-value=0.0039 Score=37.40 Aligned_cols=40 Identities=18% Similarity=0.400 Sum_probs=35.9
Q ss_pred CChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCCh
Q 032798 44 PSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSE 84 (133)
Q Consensus 44 ~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~ 84 (133)
++.|-+|.+..|+.+...||+ .....+...++.+|++.++
T Consensus 13 lt~yK~Fsq~vRP~l~~~NPk-~~~sKl~~l~~AKwrEF~~ 52 (55)
T PF08073_consen 13 LTNYKAFSQHVRPLLAKANPK-APMSKLMMLLQAKWREFQE 52 (55)
T ss_pred HHHHHHHHHHHHHHHHHHCCC-CcHHHHHHHHHHHHHHHHh
Confidence 356889999999999999999 9999999999999987553
No 22
>PF04769 MAT_Alpha1: Mating-type protein MAT alpha 1; InterPro: IPR006856 This family includes Saccharomyces cerevisiae (Baker's yeast) mating type protein alpha 1 (P01365 from SWISSPROT). MAT alpha 1 is a transcription activator that activates mating-type alpha-specific genes with the help of the MADS-box containing MCM1 transcription factor, which together bind cooperatively to PQ elements upstream of alpha-specific genes. The MCM1-MATalpha1 complex is required for the proper DNA-bending that is needed for transcriptional activation []. Alpha 1 interacts in vivo with STE12, linking expression of alpha-specific genes to the alpha-pheromone (IPR006742 from INTERPRO) response pathway [].; GO: 0000772 mating pheromone activity, 0003677 DNA binding, 0045895 positive regulation of transcription, mating-type specific, 0005634 nucleus
Probab=94.69 E-value=0.14 Score=38.56 Aligned_cols=56 Identities=20% Similarity=0.340 Sum_probs=41.3
Q ss_pred cCCCCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHH
Q 032798 33 AAKDPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAE 95 (133)
Q Consensus 33 ~~~dp~~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~ 95 (133)
.......++||+|+||.|..-.- ...|+ ....+++..|+..|..=+- |..|.-+|.
T Consensus 37 ~~~~~~~~kr~lN~Fm~FRsyy~----~~~~~-~~Qk~~S~~l~~lW~~dp~--k~~W~l~ak 92 (201)
T PF04769_consen 37 RKRSPEKAKRPLNGFMAFRSYYS----PIFPP-LPQKELSGILTKLWEKDPF--KNKWSLMAK 92 (201)
T ss_pred ccccccccccchhHHHHHHHHHH----hhcCC-cCHHHHHHHHHHHHhCCcc--HhHHHHHhh
Confidence 34456678999999999988765 33455 6778999999999987433 555665554
No 23
>PF06244 DUF1014: Protein of unknown function (DUF1014); InterPro: IPR010422 This family consists of several hypothetical eukaryotic proteins of unknown function.
Probab=93.99 E-value=0.093 Score=36.52 Aligned_cols=49 Identities=20% Similarity=0.365 Sum_probs=41.1
Q ss_pred CCCCCCCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChh
Q 032798 36 DPNKPKRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSED 85 (133)
Q Consensus 36 dp~~PKrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~ 85 (133)
|.-|-+|-.-||.-|.....+.++.++|+ +-.+++-.+|-..|..-+++
T Consensus 69 drHPErR~KAAy~afeE~~Lp~lK~E~Pg-LrlsQ~kq~l~K~w~KSPeN 117 (122)
T PF06244_consen 69 DRHPERRMKAAYKAFEERRLPELKEENPG-LRLSQYKQMLWKEWQKSPEN 117 (122)
T ss_pred CCCcchhHHHHHHHHHHHHhHHHHhhCCC-chHHHHHHHHHHHHhcCCCC
Confidence 33333455578999999999999999999 99999999999999887753
No 24
>TIGR03481 HpnM hopanoid biosynthesis associated membrane protein HpnM. The genomes containing members of this family share the machinery for the biosynthesis of hopanoid lipids. Furthermore, the genes of this family are usually located proximal to other components of this biological process. The proteins are members of the pfam05494 family of putative transporters known as "toluene tolerance protein Ttg2D", although it is unlikely that the members included here have anything to do with toluene per-se.
Probab=90.54 E-value=0.89 Score=33.95 Aligned_cols=46 Identities=15% Similarity=0.468 Sum_probs=39.4
Q ss_pred CCHHHHHH-HHHHHhcCCChhhhHHHHHHHHH-HHHHHHHHHHHHHHh
Q 032798 66 KSVATVGK-AAGEKWKSMSEDEKAPFVERAEK-RKSDYNKNMQDYNKQ 111 (133)
Q Consensus 66 ~~~~ei~k-~l~~~Wk~ls~~eK~~y~~~A~~-~k~~y~~e~~~y~~k 111 (133)
.++..+++ .||..|+.+|+++|+.|.+.... ....|-..+..|...
T Consensus 64 ~Df~~mar~vLG~~W~~~s~~Qr~~F~~~F~~~l~~tY~~~l~~y~~~ 111 (198)
T TIGR03481 64 FDLPAMARLTLGSSWTSLSPEQRRRFIGAFRELSIATYASQFKSYAGE 111 (198)
T ss_pred CCHHHHHHHHhhhhhhhCCHHHHHHHHHHHHHHHHHHHHHHHHhhcCc
Confidence 56788877 57999999999999999999888 677888888888764
No 25
>PRK15117 ABC transporter periplasmic binding protein MlaC; Provisional
Probab=87.56 E-value=2 Score=32.40 Aligned_cols=46 Identities=17% Similarity=0.342 Sum_probs=38.3
Q ss_pred CCHHHHHH-HHHHHhcCCChhhhHHHHHHHHHH-HHHHHHHHHHHHHh
Q 032798 66 KSVATVGK-AAGEKWKSMSEDEKAPFVERAEKR-KSDYNKNMQDYNKQ 111 (133)
Q Consensus 66 ~~~~ei~k-~l~~~Wk~ls~~eK~~y~~~A~~~-k~~y~~e~~~y~~k 111 (133)
.++..+++ .||..|+.+|+++|..|.+..... ..-|-..+..|..+
T Consensus 68 ~Df~~~s~~vLG~~wr~as~eQr~~F~~~F~~~Lv~tYa~~l~~y~~q 115 (211)
T PRK15117 68 VQVKYAGALVLGRYYKDATPAQREAYFAAFREYLKQAYGQALAMYHGQ 115 (211)
T ss_pred CCHHHHHHHHhhhhhhhCCHHHHHHHHHHHHHHHHHHHHHHHHHhCCc
Confidence 56777766 579999999999999999988774 56788899999764
No 26
>PF12881 NUT_N: NUT protein N terminus; InterPro: IPR024309 This domain is found in the N-terminal region of Nuclear Testis (NUT) proteins. It is also found in FAM22, which are a family of uncharacterised mammalian proteins.
Probab=87.25 E-value=3.5 Score=33.13 Aligned_cols=72 Identities=21% Similarity=0.206 Sum_probs=51.0
Q ss_pred CChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHH--HHHHHHHHhhhhhhc
Q 032798 44 PSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYN--KNMQDYNKQLVIFFG 117 (133)
Q Consensus 44 ~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~--~e~~~y~~k~~~~~~ 117 (133)
..|+.+|..-....+....|. ++..|-....-+.|...|.-+|-.|.++|++-.+ |+ +||+.-+-++.....
T Consensus 229 ~EAlSCFLIpvLrsLar~kPt-MtlEeGl~ra~qEW~~~SnfdRmifyemaekFmE-FEaeEEmq~q~lq~~~g~~ 302 (328)
T PF12881_consen 229 AEALSCFLIPVLRSLARLKPT-MTLEEGLWRAVQEWQHTSNFDRMIFYEMAEKFME-FEAEEEMQIQKLQLMNGSQ 302 (328)
T ss_pred hhhhhhhHHHHHHHHHhcCCC-ccHHHHHHHHHHHhhccccccHHHHHHHHHHHcc-CCcHHHHHHHHHHHhcCCC
Confidence 345555555555555566777 7888888888899999999999999999998753 33 466655555444333
No 27
>KOG3223 consensus Uncharacterized conserved protein [Function unknown]
Probab=84.09 E-value=1.6 Score=32.71 Aligned_cols=53 Identities=25% Similarity=0.469 Sum_probs=43.8
Q ss_pred CCC-CCCCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHH
Q 032798 38 NKP-KRPPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERA 94 (133)
Q Consensus 38 ~~P-KrP~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A 94 (133)
.+| +|=.-||.-|-....+.|+.+||+ +..+++-.+|-.+|..-++. ||.+.+
T Consensus 162 rHPEkRmrAA~~afEe~~LPrLK~e~P~-lrlsQ~Kqll~Kew~KsPDN---P~Nq~~ 215 (221)
T KOG3223|consen 162 RHPEKRMRAAFKAFEEARLPRLKKENPG-LRLSQYKQLLKKEWQKSPDN---PFNQAA 215 (221)
T ss_pred cChHHHHHHHHHHHHHhhchhhhhcCCC-ccHHHHHHHHHHHHhhCCCC---hhhHHh
Confidence 455 555677999999999999999999 99999999999999988875 555443
No 28
>PF05494 Tol_Tol_Ttg2: Toluene tolerance, Ttg2 ; InterPro: IPR008869 Toluene tolerance is mediated by increased cell membrane rigidity resulting from changes in fatty acid and phospholipid compositions, exclusion of toluene from the cell membrane, and removal of intracellular toluene by degradation []. Many proteins are involved in these processes. This family is a transporter which shows similarity to ABC transporters [].; PDB: 2QGU_A.
Probab=82.60 E-value=5.1 Score=28.71 Aligned_cols=45 Identities=20% Similarity=0.451 Sum_probs=33.6
Q ss_pred CCHHHHHHH-HHHHhcCCChhhhHHHHHHHHHH-HHHHHHHHHHHHH
Q 032798 66 KSVATVGKA-AGEKWKSMSEDEKAPFVERAEKR-KSDYNKNMQDYNK 110 (133)
Q Consensus 66 ~~~~ei~k~-l~~~Wk~ls~~eK~~y~~~A~~~-k~~y~~e~~~y~~ 110 (133)
.++..+++. ||.-|+.+|+++++.|.+...+. ...|-..+..|..
T Consensus 38 ~D~~~~ar~~LG~~w~~~s~~q~~~F~~~f~~~l~~~Y~~~l~~y~~ 84 (170)
T PF05494_consen 38 FDFERMARRVLGRYWRKASPAQRQRFVEAFKQLLVRTYAKRLDEYSG 84 (170)
T ss_dssp B-HHHHHHHHHGGGTTTS-HHHHHHHHHHHHHHHHHHHHHHHHT-SS
T ss_pred CCHHHHHHHHHHHhHhhCCHHHHHHHHHHHHHHHHHHHHHHHHhhCC
Confidence 567777765 67899999999999999887764 5667778887765
No 29
>COG2854 Ttg2D ABC-type transport system involved in resistance to organic solvents, auxiliary component [Secondary metabolites biosynthesis, transport, and catabolism]
Probab=76.67 E-value=3.5 Score=31.11 Aligned_cols=48 Identities=17% Similarity=0.266 Sum_probs=38.5
Q ss_pred HHHHHHhcCCChhhhHHHHHHHHHH-HHHHHHHHHHHHHhhhhhhcccc
Q 032798 73 KAAGEKWKSMSEDEKAPFVERAEKR-KSDYNKNMQDYNKQLVIFFGIIV 120 (133)
Q Consensus 73 k~l~~~Wk~ls~~eK~~y~~~A~~~-k~~y~~e~~~y~~k~~~~~~~~~ 120 (133)
..||.-|+.+|+++++.|....... ...|-..+..|+.+........+
T Consensus 78 ~vLGk~~k~aspeQ~~~F~~aF~~yl~q~Y~~aL~~Y~~q~~~v~~~~~ 126 (202)
T COG2854 78 LVLGKYYKTASPEQRQAFFKAFRTYLEQTYGQALLDYKGQTLKVKPSRP 126 (202)
T ss_pred HHhccccccCCHHHHHHHHHHHHHHHHHHHHHHHHHccCCCceeCCCcc
Confidence 3478999999999999999887764 66789999999988665544443
No 30
>PF13875 DUF4202: Domain of unknown function (DUF4202)
Probab=72.10 E-value=9.5 Score=28.41 Aligned_cols=39 Identities=26% Similarity=0.485 Sum_probs=33.4
Q ss_pred hHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhH
Q 032798 46 AFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKA 88 (133)
Q Consensus 46 ay~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~ 88 (133)
+.++|...+.+.+...|. -..+..+|..-|+.||+..++
T Consensus 131 acLVFL~~~f~~F~~~~d----eeK~v~Il~KTw~KMS~~g~~ 169 (185)
T PF13875_consen 131 ACLVFLEYYFEDFAAKHD----EEKIVDILRKTWRKMSERGHE 169 (185)
T ss_pred HHHHhHHHHHHHHHhcCC----HHHHHHHHHHHHHHCCHHHHH
Confidence 578999999999998883 457888899999999998775
No 31
>PF11304 DUF3106: Protein of unknown function (DUF3106); InterPro: IPR021455 Some members in this family of proteins are annotated as transmembrane proteins however this cannot be confirmed. Currently no function is known.
Probab=69.73 E-value=23 Score=23.75 Aligned_cols=41 Identities=15% Similarity=0.428 Sum_probs=22.2
Q ss_pred HHHHHHHHHhcCCChhhhHHHHHHHH-------HHHHHHHHHHHHHHH
Q 032798 70 TVGKAAGEKWKSMSEDEKAPFVERAE-------KRKSDYNKNMQDYNK 110 (133)
Q Consensus 70 ei~k~l~~~Wk~ls~~eK~~y~~~A~-------~~k~~y~~e~~~y~~ 110 (133)
++..-+...|+.|++..+..+...|. ++..+....|..|..
T Consensus 11 ~~L~pl~~~W~~l~~~qr~k~l~~a~r~~~mspeqq~r~~~rm~~W~~ 58 (107)
T PF11304_consen 11 QALAPLAERWNSLPPEQRRKWLQIAERWPSMSPEQQQRLRERMRRWAA 58 (107)
T ss_pred HHHHHHHHHHhcCCHHHHHHHHHHHHHHhcCCHHHHHHHHHHHHHHHh
Confidence 44455556666666666665555543 244455555555543
No 32
>PF01352 KRAB: KRAB box; InterPro: IPR001909 The Krueppel-associated box (KRAB) is a domain of around 75 amino acids that is found in the N-terminal part of about one third of eukaryotic Krueppel-type C2H2 zinc finger proteins (ZFPs) []. It is enriched in charged amino acids and can be divided into subregions A and B, which are predicted to fold into two amphipathic alpha-helices. The KRAB A and B boxes can be separated by variable spacer segments and many KRAB proteins contain only the A box []. The functions currently known for members of the KRAB-containing protein family include transcriptional repression of RNA polymerase I, II, and III promoters, binding and splicing of RNA, and control of nucleolus function. The KRAB domain functions as a transcriptional repressor when tethered to the template DNA by a DNA-binding domain. A sequence of 45 amino acids in the KRAB A subdomain has been shown to be necessary and sufficient for transcriptional repression. The B box does not repress by itself but does potentiate the repression exerted by the KRAB A subdomain [, ]. Gene silencing requires the binding of the KRAB domain to the RING-B box-coiled coil (RBCC) domain of the KAP-1/TIF1-beta corepressor. As KAP-1 binds to the heterochromatin proteins HP1, it has been proposed that the KRAB-ZFP-bound target gene could be silenced following recruitment to heterochromatin [, ]. KRAB-ZFPs probably constitute the single largest class of transcription factors within the human genome []. Although the function of KRAB-ZFPs is largely unknown, they appear to play important roles during cell differentiation and development. The KRAB domain is generally encoded by two exons. The regions coded by the two exons are known as KRAB-A and KRAB-B.; GO: 0003676 nucleic acid binding, 0006355 regulation of transcription, DNA-dependent, 0005622 intracellular; PDB: 1V65_A.
Probab=60.10 E-value=7.1 Score=21.76 Aligned_cols=28 Identities=14% Similarity=0.257 Sum_probs=15.7
Q ss_pred HHHHHHHHH-HHhcCCChhhhHHHHHHHH
Q 032798 68 VATVGKAAG-EKWKSMSEDEKAPFVERAE 95 (133)
Q Consensus 68 ~~ei~k~l~-~~Wk~ls~~eK~~y~~~A~ 95 (133)
|.||+--++ +.|..|.+.+|.-|.+.-.
T Consensus 3 f~Dvav~fs~eEW~~L~~~Qk~ly~dvm~ 31 (41)
T PF01352_consen 3 FEDVAVYFSQEEWELLDPAQKNLYRDVML 31 (41)
T ss_dssp ----TT---HHHHHTS-HHHHHHHHHHHH
T ss_pred EEEEEEEcChhhcccccceecccchhHHH
Confidence 445554454 5699999999999986543
No 33
>PF06945 DUF1289: Protein of unknown function (DUF1289); InterPro: IPR010710 This family consists of a number of hypothetical bacterial proteins. The aligned region spans around 56 residues and contains 4 highly conserved cysteine residues towards the N terminus. The function of this family is unknown.
Probab=53.29 E-value=21 Score=20.71 Aligned_cols=25 Identities=32% Similarity=0.684 Sum_probs=18.0
Q ss_pred CHHHHHHHHHHHhcCCChhhhHHHHHHHHH
Q 032798 67 SVATVGKAAGEKWKSMSEDEKAPFVERAEK 96 (133)
Q Consensus 67 ~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~ 96 (133)
+..||. .|..|++++|.........
T Consensus 23 T~dEI~-----~W~~~s~~er~~i~~~l~~ 47 (51)
T PF06945_consen 23 TLDEIR-----DWKSMSDDERRAILARLRA 47 (51)
T ss_pred cHHHHH-----HHhhCCHHHHHHHHHHHHH
Confidence 456664 4999999998877654443
No 34
>PRK09706 transcriptional repressor DicA; Reviewed
Probab=50.68 E-value=54 Score=22.42 Aligned_cols=43 Identities=19% Similarity=0.235 Sum_probs=36.5
Q ss_pred HHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHHhh
Q 032798 70 TVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQL 112 (133)
Q Consensus 70 ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~k~ 112 (133)
+-...|-..|+.|+++++.......+...+.|..-+++|-...
T Consensus 87 ~~~~~ll~~~~~L~~~~~~~~l~~l~~~~~~~~~~~~~~~~~~ 129 (135)
T PRK09706 87 EDQKELLELFDALPESEQDAQLSEMRARVENFNKLFEELLKAR 129 (135)
T ss_pred HHHHHHHHHHHHCCHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 3456788999999999999999999999999988888886653
No 35
>PRK12750 cpxP periplasmic repressor CpxP; Reviewed
Probab=46.82 E-value=64 Score=23.52 Aligned_cols=32 Identities=19% Similarity=0.317 Sum_probs=25.1
Q ss_pred HHHHHhcCCChhhhHHHHHHHHHHHHHHHHHH
Q 032798 74 AAGEKWKSMSEDEKAPFVERAEKRKSDYNKNM 105 (133)
Q Consensus 74 ~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~ 105 (133)
..-+.+..|++++|..|.++-.+-.+.+.+.+
T Consensus 129 ~~~~~~~vLTpEQRak~~e~~~~r~~~~~~~~ 160 (170)
T PRK12750 129 KRHQMLSILTPEQKAKFQELQQERMQECQDKM 160 (170)
T ss_pred HHHHHHHhCCHHHHHHHHHHHHHHHHHHHHHH
Confidence 34467999999999999988877766666655
No 36
>PRK12751 cpxP periplasmic stress adaptor protein CpxP; Reviewed
Probab=45.78 E-value=55 Score=23.78 Aligned_cols=29 Identities=10% Similarity=0.312 Sum_probs=21.3
Q ss_pred HHHHHHhcCCChhhhHHHHHHHHHHHHHH
Q 032798 73 KAAGEKWKSMSEDEKAPFVERAEKRKSDY 101 (133)
Q Consensus 73 k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y 101 (133)
+..-+++..|++++|..|.+..++-..+.
T Consensus 121 ~~~~qmy~lLTPEQra~l~~~~e~r~~~~ 149 (162)
T PRK12751 121 KVRNQMYNLLTPEQKEALNKKHQERIEKL 149 (162)
T ss_pred HHHHHHHHcCCHHHHHHHHHHHHHHHHHH
Confidence 44467789999999999987666554443
No 37
>PRK10236 hypothetical protein; Provisional
Probab=45.25 E-value=21 Score=27.57 Aligned_cols=25 Identities=24% Similarity=0.559 Sum_probs=20.5
Q ss_pred HHHHHHHHhcCCChhhhHHHHHHHH
Q 032798 71 VGKAAGEKWKSMSEDEKAPFVERAE 95 (133)
Q Consensus 71 i~k~l~~~Wk~ls~~eK~~y~~~A~ 95 (133)
+.+.+...|..||++|++.+.+.-.
T Consensus 118 l~kll~~a~~kms~eE~~~L~~~l~ 142 (237)
T PRK10236 118 LEQFLRNTWKKMDEEHKQEFLHAVD 142 (237)
T ss_pred HHHHHHHHHHHCCHHHHHHHHHHHh
Confidence 4778899999999999988875433
No 38
>PF12650 DUF3784: Domain of unknown function (DUF3784); InterPro: IPR017259 This group represents an uncharacterised conserved protein.
Probab=42.60 E-value=19 Score=23.42 Aligned_cols=16 Identities=25% Similarity=0.584 Sum_probs=13.4
Q ss_pred HhcCCChhhhHHHHHH
Q 032798 78 KWKSMSEDEKAPFVER 93 (133)
Q Consensus 78 ~Wk~ls~~eK~~y~~~ 93 (133)
=|+.||++||+.|.+.
T Consensus 25 Gyntms~eEk~~~D~~ 40 (97)
T PF12650_consen 25 GYNTMSKEEKEKYDKK 40 (97)
T ss_pred hcccCCHHHHHHhhHH
Confidence 3899999999999653
No 39
>PRK10363 cpxP periplasmic repressor CpxP; Reviewed
Probab=41.33 E-value=85 Score=23.00 Aligned_cols=33 Identities=12% Similarity=0.341 Sum_probs=24.8
Q ss_pred HHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHH
Q 032798 70 TVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYN 102 (133)
Q Consensus 70 ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~ 102 (133)
+..+.-.++++-|++|+|..|++..++-...+.
T Consensus 112 em~k~~nqmy~lLTPEQKaq~~~~~~~rm~~~~ 144 (166)
T PRK10363 112 EMAKVRNQMYRLLTPEQQAVLNEKHQQRMEQLR 144 (166)
T ss_pred HHHHHHHHHHHhCCHHHHHHHHHHHHHHHHHHH
Confidence 345555689999999999999877766655553
No 40
>COG1638 DctP TRAP-type C4-dicarboxylate transport system, periplasmic component [Carbohydrate transport and metabolism]
Probab=40.35 E-value=76 Score=25.56 Aligned_cols=35 Identities=17% Similarity=0.419 Sum_probs=26.4
Q ss_pred HHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHH
Q 032798 76 GEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNK 110 (133)
Q Consensus 76 ~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~ 110 (133)
...|..||+++++...+.+.+......+...+.+.
T Consensus 244 ~~~w~~L~~e~q~il~~aa~e~~~~~~~~~~~~e~ 278 (332)
T COG1638 244 KAFWDSLPEEDQTILLEAAKEAAEEQRKLVEELED 278 (332)
T ss_pred HHHHhcCCHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 57899999999999998888776555554444443
No 41
>TIGR00787 dctP tripartite ATP-independent periplasmic transporter solute receptor, DctP family. TRAP-T (Tripartite ATP-independent Periplasmic Transporter) family proteins generally consist of three components, and these systems have so far been found in Gram-negative bacteria, Gram-postive bacteria and archaea. The best characterized example is the DctPQM system of Rhodobacter capsulatus, a C4 dicarboxylate (malate, fumarate, succinate) transporter. This model represents the DctP family, one of at least three major families of extracytoplasmic solute receptor for TRAP family transporters. Other are the SnoM family (see pfam03480) and TAXI (TRAP-associated extracytoplasmic immunogenic) family.
Probab=40.14 E-value=81 Score=23.82 Aligned_cols=28 Identities=29% Similarity=0.306 Sum_probs=21.4
Q ss_pred HHHhcCCChhhhHHHHHHHHHHHHHHHH
Q 032798 76 GEKWKSMSEDEKAPFVERAEKRKSDYNK 103 (133)
Q Consensus 76 ~~~Wk~ls~~eK~~y~~~A~~~k~~y~~ 103 (133)
...|..||++.|+-..+.+.+.......
T Consensus 213 ~~~~~~L~~e~q~~i~~a~~~~~~~~~~ 240 (257)
T TIGR00787 213 KAFWKSLPPDLQAVVKEAAKEAGEYQRK 240 (257)
T ss_pred HHHHhcCCHHHHHHHHHHHHHHHHHHHH
Confidence 4679999999999998877766444333
No 42
>KOG1610 consensus Corticosteroid 11-beta-dehydrogenase and related short chain-type dehydrogenases [Secondary metabolites biosynthesis, transport and catabolism; General function prediction only]
Probab=38.59 E-value=85 Score=25.44 Aligned_cols=51 Identities=16% Similarity=0.342 Sum_probs=35.8
Q ss_pred HHHHHHHHHHHHh-------CCC-----CCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHH
Q 032798 49 VFMEEFRKQFKEA-------HPN-----NKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKS 99 (133)
Q Consensus 49 lF~~~~r~~~k~~-------~p~-----~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~ 99 (133)
.|+...|.++..= -|+ ......+.+.+.+.|..|+++.|+.|-+.+-.+..
T Consensus 187 af~D~lR~EL~~fGV~VsiiePG~f~T~l~~~~~~~~~~~~~w~~l~~e~k~~YGedy~~~~~ 249 (322)
T KOG1610|consen 187 AFSDSLRRELRPFGVKVSIIEPGFFKTNLANPEKLEKRMKEIWERLPQETKDEYGEDYFEDYK 249 (322)
T ss_pred HHHHHHHHHHHhcCcEEEEeccCccccccCChHHHHHHHHHHHhcCCHHHHHHHHHHHHHHHH
Confidence 4666677666421 122 12357888999999999999999999877765533
No 43
>PF09164 VitD-bind_III: Vitamin D binding protein, domain III; InterPro: IPR015247 This domain is predominantly found in Vitamin D binding proteins, and adopts a multihelical structure. It is required for formation of an actin 'clamp', allowing the protein to bind to actin []. ; PDB: 1MA9_A 1KW2_A 1KXP_D 1J7E_A 1J78_A 1LOT_A.
Probab=35.11 E-value=1.1e+02 Score=18.95 Aligned_cols=32 Identities=6% Similarity=0.283 Sum_probs=23.4
Q ss_pred ChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHH
Q 032798 45 SAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGE 77 (133)
Q Consensus 45 say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~ 77 (133)
+.|.-|-+...+.++...|+ .+..++..++.+
T Consensus 9 ~tFtEyKKrL~e~l~~k~P~-at~~~l~~lve~ 40 (68)
T PF09164_consen 9 NTFTEYKKRLAERLRAKLPD-ATPTELKELVEK 40 (68)
T ss_dssp S-HHHHHHHHHHHHHHH-TT-S-HHHHHHHHHH
T ss_pred ccHHHHHHHHHHHHHHHCCC-CCHHHHHHHHHH
Confidence 45777888888899999999 888888777654
No 44
>PF00887 ACBP: Acyl CoA binding protein; InterPro: IPR000582 Acyl-CoA-binding protein (ACBP) is a small (10 Kd) protein that binds medium- and long-chain acyl-CoA esters with very high affinity and may function as an intracellular carrier of acyl-CoA esters []. ACBP is also known as diazepam binding inhibitor (DBI) or endozepine (EP) because of its ability to displace diazepam from the benzodiazepine (BZD) recognition site located on the GABA type A receptor. It is therefore possible that this protein also acts as a neuropeptide to modulate the action of the GABA receptor []. ACBP is a highly conserved protein of about 90 residues that is found in all four eukaryotic kingdoms, Animalia, Plantae, Fungi and Protista, and in some eubacterial species []. Although ACBP occurs as a completely independent protein, intact ACB domains have been identified in a number of large, multifunctional proteins in a variety of eukaryotic species. These include large membrane-associated proteins with N-terminal ACB domains, multifunctional enzymes with both ACB and peroxisomal enoyl-CoA Delta(3), Delta(2)-enoyl-CoA isomerase domains, and proteins with both an ACB domain and ankyrin repeats (IPR002110 from INTERPRO) []. The ACB domain consists of four alpha-helices arranged in a bowl shape with a highly exposed acyl-CoA-binding site. The ligand is bound through specific interactions with residues on the protein, most notably several conserved positive charges that interact with the phosphate group on the adenosine-3'phosphate moiety, and the acyl chain is sandwiched between the hydrophobic surfaces of CoA and the protein []. Other proteins containing an ACB domain include: Endozepine-like peptide (ELP) (gene DBIL5) from mouse []. ELP is a testis-specific ACBP homologue that may be involved in the energy metabolism of the mature sperm. MA-DBI, a transmembrane protein of unknown function which has been found in mammals. MA-DBI contains a N-terminal ACB domain. DRS-1 [], a human protein of unknown function that contains a N-terminal ACB domain and a C-terminal enoyl-CoA isomerase/hydratase domain. ; GO: 0000062 fatty-acyl-CoA binding; PDB: 2CB8_A 2FJ9_A 2LBB_A 1ST7_A 3EPY_B 2FDQ_C 1NTI_A 1HB8_A 1ACA_A 1NVL_A ....
Probab=34.63 E-value=1.2e+02 Score=19.12 Aligned_cols=53 Identities=15% Similarity=0.359 Sum_probs=28.4
Q ss_pred HHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCC----hhhhHHHHHHHHHHHHHH
Q 032798 47 FFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMS----EDEKAPFVERAEKRKSDY 101 (133)
Q Consensus 47 y~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls----~~eK~~y~~~A~~~k~~y 101 (133)
|-+|.+.....+....|...++....+ -..|+.|. ++-+..|.+...+....|
T Consensus 30 YalyKQAt~Gd~~~~~P~~~d~~~~~K--~~AW~~l~gms~~eA~~~Yi~~v~~~~~~~ 86 (87)
T PF00887_consen 30 YALYKQATHGDCDTPRPGFFDIEGRAK--WDAWKALKGMSKEEAMREYIELVEELIPKY 86 (87)
T ss_dssp HHHHHHHHTSS--S-CTTTTCHHHHHH--HHHHHTTTTTHHHHHHHHHHHHHHHHHHHH
T ss_pred HHHHHHHHhCCCcCCCCcchhHHHHHH--HHHHHHccCCCHHHHHHHHHHHHHHHHHhc
Confidence 666777665555555565334444333 35687776 344556666666555444
No 45
>cd07081 ALDH_F20_ACDH_EutE-like Coenzyme A acylating aldehyde dehydrogenase (ACDH), Ethanolamine utilization protein EutE, and related proteins. Coenzyme A acylating aldehyde dehydrogenase (ACDH), an NAD+ and CoA-dependent acetaldehyde dehydrogenase, acetylating (EC=1.2.1.10), functions as a single enzyme (such as the Ethanolamine utilization protein, EutE, in Salmonella typhimurium) or as part of a multifunctional enzyme to convert acetaldehyde into acetyl-CoA. The E. coli aldehyde-alcohol dehydrogenase includes the functional domains, alcohol dehydrogenase (ADH), ACDH, and pyruvate-formate-lyase deactivase; and the Entamoeba histolytica aldehyde-alcohol dehydrogenase 2 (ALDH20A1) includes the functional domains ADH and ACDH, and may be critical enzymes in the fermentative pathway.
Probab=32.66 E-value=1.2e+02 Score=25.29 Aligned_cols=41 Identities=12% Similarity=0.105 Sum_probs=32.3
Q ss_pred HHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHH
Q 032798 70 TVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNK 110 (133)
Q Consensus 70 ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~ 110 (133)
+.++..-..|..++.++|..+...+.+..+++..++.....
T Consensus 6 ~~A~~A~~~W~~~~~~~R~~iL~~~a~~l~~~~~ela~~~~ 46 (439)
T cd07081 6 AAAKVAQQGLSCKSQEMVDLIFRAAAEAAEDARIDLAKLAV 46 (439)
T ss_pred HHHHHHHHHHhhCCHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 34455557899999999999998888888888888876633
No 46
>PF03480 SBP_bac_7: Bacterial extracellular solute-binding protein, family 7; InterPro: IPR018389 This family of proteins are involved in binding extracellular solutes for transport across the bacterial cytoplasmic membrane. This family includes a C4-dicarboxylate-binding protein DctP [, ] and the sialic acid-binding protein SiaP. The structure of the SiaP receptor has revealed an overall topology similar to ATP binding cassette ESR (extracytoplasmic solute receptors) proteins []. Upon binding of sialic acid, SiaP undergoes domain closure about a hinge region and kinking of an alpha-helix hinge component [].; GO: 0006810 transport, 0030288 outer membrane-bounded periplasmic space; PDB: 2HZK_C 2HZL_B 2HPG_C 2XWI_A 2XWK_A 2WX9_A 2CEY_A 2WYP_A 3B50_A 2CEX_B ....
Probab=31.18 E-value=1.1e+02 Score=23.53 Aligned_cols=39 Identities=13% Similarity=0.315 Sum_probs=24.8
Q ss_pred HHHhcCCChhhhHHHHHHHHHHHHHH----HHHHHHHHHhhhh
Q 032798 76 GEKWKSMSEDEKAPFVERAEKRKSDY----NKNMQDYNKQLVI 114 (133)
Q Consensus 76 ~~~Wk~ls~~eK~~y~~~A~~~k~~y----~~e~~~y~~k~~~ 114 (133)
.+.|..||++.|+-..+.+.+....+ .....+..+.+.+
T Consensus 213 ~~~w~~L~~e~q~~l~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 255 (286)
T PF03480_consen 213 KDWWDSLPDEDQEALDDAADEAEARAREYYEAEDEEALKELEE 255 (286)
T ss_dssp HHHHHHS-HHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
T ss_pred HHHHhcCCHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 36799999999999998776654433 3344444444444
No 47
>KOG2880 consensus SMAD6 interacting protein AMSH, contains JAB/MPN/Mov34 domain [Signal transduction mechanisms]
Probab=30.66 E-value=2.3e+02 Score=23.60 Aligned_cols=64 Identities=17% Similarity=0.252 Sum_probs=34.6
Q ss_pred CChHHHHHHHHH--HHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHH
Q 032798 44 PSAFFVFMEEFR--KQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNK 110 (133)
Q Consensus 44 ~say~lF~~~~r--~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~ 110 (133)
-+||+||.+-.- -+-...||| . +.+-...-...+.|-++.-..-.++..+...+|..+..+|..
T Consensus 52 enafvLy~ry~tLfiEkipkHrD-y--~s~k~ek~d~~~klk~~~~p~~deL~~~ll~rY~~eyn~y~~ 117 (424)
T KOG2880|consen 52 ENAFVLYLRYITLFIEKIPKHRD-Y--RSVKPEKEDIRKKLKEEAFPRIDELKAKLLKRYNVEYNEYDH 117 (424)
T ss_pred chhhhHHHHHHHHHHHhcccCcc-h--hhhchhHHHHHHHHHHHhhhhHHHHHHHHHHHHhhHHHHHHH
Confidence 467777765331 111234555 2 233333333334444555555566777777888888887765
No 48
>PF15581 Imm35: Immunity protein 35
Probab=30.64 E-value=1.1e+02 Score=20.19 Aligned_cols=24 Identities=8% Similarity=0.329 Sum_probs=18.4
Q ss_pred CHHHHHHHHHHHhcCCChhhhHHH
Q 032798 67 SVATVGKAAGEKWKSMSEDEKAPF 90 (133)
Q Consensus 67 ~~~ei~k~l~~~Wk~ls~~eK~~y 90 (133)
+..-+...|.+.|+.|++++=..-
T Consensus 31 ~i~~l~~lIe~eWRGl~~~qV~~k 54 (93)
T PF15581_consen 31 TIRNLESLIEHEWRGLPEEQVLYK 54 (93)
T ss_pred HHHHHHHHHHHHHcCCCHHHHHHH
Confidence 456778889999999998765433
No 49
>PF09655 Nitr_red_assoc: Conserved nitrate reductase-associated protein (Nitr_red_assoc); InterPro: IPR013481 Proteins in this entry are found in the Cyanobacteria, and are mostly encoded near nitrate reductase and molybdopterin biosynthesis genes. Molybdopterin guanine dinucleotide is a cofactor for nitrate reductase. These proteins are sometimes annotated as nitrate reductase-associated proteins, though their function is unknown.
Probab=30.52 E-value=1.1e+02 Score=21.86 Aligned_cols=45 Identities=13% Similarity=0.248 Sum_probs=32.7
Q ss_pred HHhcCCChhhhHHHHHHH---HHHHHHHHHHHHHHHHhhhhhhccccc
Q 032798 77 EKWKSMSEDEKAPFVERA---EKRKSDYNKNMQDYNKQLVIFFGIIVV 121 (133)
Q Consensus 77 ~~Wk~ls~~eK~~y~~~A---~~~k~~y~~e~~~y~~k~~~~~~~~~~ 121 (133)
..|..|+.+||+...+.. ..+.+.|...+.+.-..+.......+.
T Consensus 33 ~~W~~l~~~eRq~Lv~~pc~t~~ei~~yr~~L~~li~~~~~~~~~~l~ 80 (144)
T PF09655_consen 33 SHWQQLSQEERQQLVDLPCDTPEEIQNYREFLQELIRTHAGGPAKDLP 80 (144)
T ss_pred HHHhcCCHHHHHHHHcCCCCCHHHHHHHHHHHHHHHHHHhCCCcccCC
Confidence 569999999999988765 445667888777777666655544443
No 50
>smart00271 DnaJ DnaJ molecular chaperone homology domain.
Probab=29.29 E-value=94 Score=17.56 Aligned_cols=33 Identities=21% Similarity=0.347 Sum_probs=18.6
Q ss_pred HHHHHHHHhCCCCCC-----HHHHHHHHHHHhcCCChh
Q 032798 53 EFRKQFKEAHPNNKS-----VATVGKAAGEKWKSMSED 85 (133)
Q Consensus 53 ~~r~~~k~~~p~~~~-----~~ei~k~l~~~Wk~ls~~ 85 (133)
..+...+.-||+... ..+....|.+.|..|.+.
T Consensus 21 ay~~l~~~~HPD~~~~~~~~~~~~~~~l~~Ay~~L~~~ 58 (60)
T smart00271 21 AYRKLALKYHPDKNPGDKEEAEEKFKEINEAYEVLSDP 58 (60)
T ss_pred HHHHHHHHHCcCCCCCchHHHHHHHHHHHHHHHHHcCC
Confidence 344555666888333 335556666666666543
No 51
>cd00225 API3 Ascaris pepsin inhibitor-3 (API3); protein inhibitor that reversibly inhibits aspartic proteinase cathepsin E, and gastric enzymes pepsin and gastricsin.
Probab=28.74 E-value=91 Score=22.55 Aligned_cols=29 Identities=14% Similarity=0.334 Sum_probs=16.9
Q ss_pred HHhcCCChhhhHHHHHHHHHHHHHHHHHHH
Q 032798 77 EKWKSMSEDEKAPFVERAEKRKSDYNKNMQ 106 (133)
Q Consensus 77 ~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~ 106 (133)
-.|+.|+.+|.+.+. .+..+-..|+.++.
T Consensus 26 ~~lReLt~~Eq~el~-~y~~d~~~yK~~~k 54 (159)
T cd00225 26 FPLRELTPDEQQELA-QYVEDVADYKEEVK 54 (159)
T ss_pred ceeeeCCHHHHHHHH-HHHHHHHHHHHHHH
Confidence 469999999865543 33333444544444
No 52
>PF01297 TroA: Periplasmic solute binding protein family; InterPro: IPR006127 This is a family of ABC transporter metal-binding lipoproteins. An example is the periplasmic zinc-binding protein TroA P96116 from SWISSPROT that interacts with an ATP-binding cassette transport system in Treponema pallidum and plays a role in the transport of zinc across the cytoplasmic membrane. Related proteins are found in both Gram-positive and Gram-negative bacteria. ; GO: 0046872 metal ion binding, 0030001 metal ion transport; PDB: 2PS9_A 2PS0_A 2OSV_A 2OGW_A 2PS3_A 2PRS_B 3MFQ_C 3GI1_B 2OV3_A 1PQ4_A ....
Probab=28.73 E-value=1.5e+02 Score=22.21 Aligned_cols=49 Identities=12% Similarity=0.179 Sum_probs=38.1
Q ss_pred HHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHHhhhhhh
Q 032798 68 VATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQLVIFF 116 (133)
Q Consensus 68 ~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~k~~~~~ 116 (133)
...++..|++....+.++.+..|.+-++....+-..--..+...+....
T Consensus 101 ~~~~~~~Ia~~L~~~~P~~~~~y~~N~~~~~~~L~~l~~~~~~~~~~~~ 149 (256)
T PF01297_consen 101 AKKMAEAIADALSELDPANKDYYEKNAEKYLKELDELDAEIKEKLAKLP 149 (256)
T ss_dssp HHHHHHHHHHHHHHHTGGGHHHHHHHHHHHHHHHHHHHHHHHHHHTTSS
T ss_pred HHHHHHHHHHHHHHhCccchHHHHHHHHHHHHHHHHHHHHHHHHhhccc
Confidence 4566777888888889999999999888877777777777777666544
No 53
>KOG0493 consensus Transcription factor Engrailed, contains HOX domain [General function prediction only]
Probab=27.75 E-value=2e+02 Score=22.89 Aligned_cols=24 Identities=29% Similarity=0.451 Sum_probs=16.0
Q ss_pred CCCCCCCChHHHHHHHHHHHHHHhCCC
Q 032798 38 NKPKRPPSAFFVFMEEFRKQFKEAHPN 64 (133)
Q Consensus 38 ~~PKrP~say~lF~~~~r~~~k~~~p~ 64 (133)
+--|||.+||. .++...++.+.-.
T Consensus 244 ~eeKRPRTAFt---aeQL~RLK~EF~e 267 (342)
T KOG0493|consen 244 KEEKRPRTAFT---AEQLQRLKAEFQE 267 (342)
T ss_pred chhcCcccccc---HHHHHHHHHHHhh
Confidence 34589999954 6666667665543
No 54
>KOG3838 consensus Mannose lectin ERGIC-53, involved in glycoprotein traffic [Intracellular trafficking, secretion, and vesicular transport]
Probab=27.57 E-value=66 Score=27.07 Aligned_cols=37 Identities=27% Similarity=0.245 Sum_probs=26.3
Q ss_pred CChhhhHHHHHHHHHHHHHHHHHHHHHHHhhhhhhcc
Q 032798 82 MSEDEKAPFVERAEKRKSDYNKNMQDYNKQLVIFFGI 118 (133)
Q Consensus 82 ls~~eK~~y~~~A~~~k~~y~~e~~~y~~k~~~~~~~ 118 (133)
+.+.+|++|.++.+.....|+.+.++|.+.+.+..+.
T Consensus 269 ~qe~ek~kyqeEfe~~q~elek~k~efkk~hpd~~~e 305 (497)
T KOG3838|consen 269 MQELEKAKYQEEFEWAQLELEKRKDEFKKSHPDAQGE 305 (497)
T ss_pred hhHHHHHHHHHHHHHHHHHHhhhHhhhccCCchhhcc
Confidence 3455777888888888888888888887766654443
No 55
>cd07133 ALDH_CALDH_CalB Coniferyl aldehyde dehydrogenase-like. Coniferyl aldehyde dehydrogenase (CALDH, EC=1.2.1.68) of Pseudomonas sp. strain HR199 (CalB) which catalyzes the NAD+-dependent oxidation of coniferyl aldehyde to ferulic acid, and similar sequences, are present in this CD.
Probab=27.05 E-value=1.9e+02 Score=23.87 Aligned_cols=40 Identities=18% Similarity=0.012 Sum_probs=29.8
Q ss_pred HHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHH
Q 032798 70 TVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN 109 (133)
Q Consensus 70 ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~ 109 (133)
+.++..-..|..++..+|..+.....+..+.+..++....
T Consensus 5 ~~a~~a~~~w~~~~~~~R~~~L~~~a~~l~~~~~el~~~~ 44 (434)
T cd07133 5 ERQKAAFLANPPPSLEERRDRLDRLKALLLDNQDALAEAI 44 (434)
T ss_pred HHHHHHHHhcCCCCHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 4445556679999999999888887777777777776543
No 56
>cd07132 ALDH_F3AB Aldehyde dehydrogenase family 3 members A1, A2, and B1 and related proteins. NAD(P)+-dependent, aldehyde dehydrogenase, family 3 members A1 and B1 (ALDH3A1, ALDH3B1, EC=1.2.1.5) and fatty aldehyde dehydrogenase, family 3 member A2 (ALDH3A2, EC=1.2.1.3), and similar sequences are included in this CD. Human ALDH3A1 is a homodimer with a critical role in cellular defense against oxidative stress; it catalyzes the oxidation of various cellular membrane lipid-derived aldehydes. Corneal crystalline ALDH3A1 protects the cornea and underlying lens against UV-induced oxidative stress. Human ALDH3A2, a microsomal homodimer, catalyzes the oxidation of long-chain aliphatic aldehydes to fatty acids. Human ALDH3B1 is highly expressed in the kidney and liver and catalyzes the oxidation of various medium- and long-chain saturated and unsaturated aliphatic aldehydes.
Probab=26.63 E-value=1.8e+02 Score=24.10 Aligned_cols=40 Identities=8% Similarity=-0.039 Sum_probs=30.8
Q ss_pred HHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHH
Q 032798 70 TVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN 109 (133)
Q Consensus 70 ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~ 109 (133)
+.++..-..|..++..+|..+........+.+..++.+-.
T Consensus 5 ~~A~~A~~~w~~~~~~~R~~~L~~~a~~l~~~~~~l~~~~ 44 (443)
T cd07132 5 RRAREAFSSGKTRPLEFRIQQLEALLRMLEENEDEIVEAL 44 (443)
T ss_pred HHHHHHHHhcCCCCHHHHHHHHHHHHHHHHHhHHHHHHHH
Confidence 4455556779999999999999888887777777776543
No 57
>PF05388 Carbpep_Y_N: Carboxypeptidase Y pro-peptide; InterPro: IPR008442 This signature is found at the N terminus of carboxypeptidase Y, which belong to MEROPS peptidase family S10. This region contains the signal peptide and pro-peptide regions [,].; GO: 0004185 serine-type carboxypeptidase activity, 0005773 vacuole
Probab=26.47 E-value=1.2e+02 Score=20.66 Aligned_cols=29 Identities=24% Similarity=0.241 Sum_probs=24.4
Q ss_pred HHHHHHHHHHHhcCCChhhhHHHHHHHHH
Q 032798 68 VATVGKAAGEKWKSMSEDEKAPFVERAEK 96 (133)
Q Consensus 68 ~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~ 96 (133)
+..+++.+++.++.|+.+-|+.|.++...
T Consensus 45 ~~~~~~~l~e~l~~Lt~e~k~~W~E~~~~ 73 (113)
T PF05388_consen 45 LEKISKYLNEPLKSLTSEAKALWDEMMLL 73 (113)
T ss_pred HHHHHHHHHHHHhhccHHHHHHHHHHHHH
Confidence 45666778899999999999999988765
No 58
>cd01145 TroA_c Periplasmic binding protein TroA_c. These proteins are predicted to function as initial receptors in the ABC metal ion uptake in eubacteria and archaea. They belong to the TroA superfamily of helical backbone metal receptor proteins that share a distinct fold and ligand binding mechanism. A typical TroA protein is comprised of two globular subdomains connected by a single helix and can bind their ligands in the cleft between these domains.
Probab=25.62 E-value=2.4e+02 Score=20.61 Aligned_cols=48 Identities=15% Similarity=0.251 Sum_probs=36.5
Q ss_pred HHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHHhhhhh
Q 032798 68 VATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQLVIF 115 (133)
Q Consensus 68 ~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~k~~~~ 115 (133)
...++..|++....+.++.+..|.+-++....+-..-...+...+...
T Consensus 117 ~~~~a~~I~~~L~~~dP~~~~~y~~N~~~~~~~l~~l~~~~~~~l~~~ 164 (203)
T cd01145 117 APALAKALADALIELDPSEQEEYKENLRVFLAKLNKLLREWERQFEGL 164 (203)
T ss_pred HHHHHHHHHHHHHHhCcccHHHHHHHHHHHHHHHHHHHHHHHHHhhcc
Confidence 466777788888889999999999888877777666666666665543
No 59
>PF06394 Pepsin-I3: Pepsin inhibitor-3-like repeated domain; InterPro: IPR010480 Peptide proteinase inhibitors can be found as single domain proteins or as single or multiple domains within proteins; these are referred to as either simple or compound inhibitors, respectively. In many cases they are synthesised as part of a larger precursor protein, either as a prepropeptide or as an N-terminal domain associated with an inactive peptidase or zymogen. This domain prevents access of the substrate to the active site. Removal of the N-terminal inhibitor domain either by interaction with a second peptidase or by autocatalytic cleavage activates the zymogen. Other inhibitors interact direct with proteinases using a simple noncovalent lock and key mechanism; while yet others use a conformational change-based trapping mechanism that depends on their structural and thermodynamic properties. The members of this group of proteins belong to MEROPS inhibitor family I33, clan IR; the nematode aspartyl protease inhibitors or Aspins. They are restricted to parasitic nematode species. Structural features common to the nematode Aspins include the presence of a signal peptide sequence and the conservation of all four cysteine residues in the mature protein. The Y[V.A]RDLT sequence motif has been suggested as being of crucial functional importance in several filarial nematode inhibitors [], this sequence is not conserved in Tco-API-1 from Trichostrongylus colubriformis (Black scour worm) and it has been demonstrated that Tco-API-1, is not an Aspin as it does not inhibit porcine pepsin []. Related inhibitors from Onchocerca volvulus, Ov33 [] and Ascaris suum (Pig roundworm), PI-3 [] inhibit the in vitro activity of aspartyl proteases such as pepsin and cathepsin E (MEROPS peptidase family A1). Aspin may facilitate the safe passage of the eggs of Ascaris through the host stomach without digestion by pepsin [, ]. The other parasitic nematodes known to express homologous proteins do not pass through the stomach of their hosts []. Several proteins in the family are potent allergens in mammals. The three-dimensional structures of pepsin inhibitor-3 (PI-3) from A. suum and of the complex between PI-3 and porcine pepsin at 1. 75 A and 2.45 A resolution, respectively, have revealed the mechanism of aspartic protease inhibition. PI-3 has a new fold consisting of two identical domains, each comprising an antiparallel beta-sheet flanked by an alpha-helix. In the enzyme-inhibitor complex, the N-terminal beta-strand of PI-3 pairs with one strand of the 'active site flap' (residues 70-82) of pepsin, thus forming an eight-stranded beta-sheet that spans the two proteins. PI-3 has a novel mode of inhibition, using its N-terminal residues to occupy and therefore block the first three binding pockets in pepsin for substrate residues C-terminal to the scissile bond (S1'-S3') [].; PDB: 1F32_A 1F34_B.
Probab=25.47 E-value=96 Score=19.73 Aligned_cols=27 Identities=26% Similarity=0.555 Sum_probs=16.8
Q ss_pred cCCChhhhHHHHHHHHHHHHHHHHHHHHHHHhhhh
Q 032798 80 KSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQLVI 114 (133)
Q Consensus 80 k~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~k~~~ 114 (133)
+.|+++|+... ..|.+++..|...+..
T Consensus 38 R~Lt~~E~~eL--------~~y~~~v~~y~~~l~~ 64 (76)
T PF06394_consen 38 RDLTPDEQQEL--------KTYQKKVAAYKEQLQQ 64 (76)
T ss_dssp EE--HHHHHHH--------HHHHHHHHHHHHHHTT
T ss_pred ccCCHHHHHHH--------HHHHHHHHHHHHHHHH
Confidence 45666665443 5788888888877654
No 60
>KOG1827 consensus Chromatin remodeling complex RSC, subunit RSC1/Polybromo and related proteins [Chromatin structure and dynamics; Transcription]
Probab=25.26 E-value=4.4 Score=35.48 Aligned_cols=44 Identities=18% Similarity=0.316 Sum_probs=39.9
Q ss_pred CCChHHHHHHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhh
Q 032798 43 PPSAFFVFMEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEK 87 (133)
Q Consensus 43 P~say~lF~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK 87 (133)
-.++|+.|+.+.+..+...+|+ ..+++++.++|..|..|+...|
T Consensus 552 ~~~~~~~~s~~~~~~~~~~np~-v~~~~~~~~vg~~~~~lp~~~k 595 (629)
T KOG1827|consen 552 SPEPYILDSIENRTIIWFENPT-VGFGEVSIIVGNDWDKLPNINK 595 (629)
T ss_pred CCccccccccccCceeeeeCCC-cccceeEEeecCCcccCccccc
Confidence 5788999999999999999999 8999999999999999994443
No 61
>cd01137 PsaA Metal binding protein PsaA. These proteins have been shown to function as initial receptors in ABC transport of Mn2+ and as surface adhesins in some eubacterial species. They belong to the TroA superfamily of periplasmic metal binding proteins that share a distinct fold and ligand binding mechanism. A typical TroA protein is comprised of two globular subdomains connected by a single helix and can bind the metal ion in the cleft between these domains. In addition, these proteins sometimes have a low complexity region containing a metal-binding histidine-rich motif (repetitive HDH sequence).
Probab=24.94 E-value=2.3e+02 Score=22.01 Aligned_cols=47 Identities=6% Similarity=-0.004 Sum_probs=36.9
Q ss_pred HHHHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHHhhhh
Q 032798 68 VATVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNKQLVI 114 (133)
Q Consensus 68 ~~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~k~~~ 114 (133)
...+...|++....+.++.+..|.+-++....+-..-...|+..+..
T Consensus 126 ~~~~a~~Ia~~L~~~dP~~~~~y~~N~~~~~~~L~~l~~~~~~~l~~ 172 (287)
T cd01137 126 AIIYVKNIAKALSEADPANAETYQKNAAAYKAKLKALDEWAKAKFAT 172 (287)
T ss_pred HHHHHHHHHHHHHHHCcccHHHHHHHHHHHHHHHHHHHHHHHHHHhc
Confidence 56677778888888899999999988888777776666677777665
No 62
>cd07122 ALDH_F20_ACDH Coenzyme A acylating aldehyde dehydrogenase (ACDH), ALDH family 20-like. Coenzyme A acylating aldehyde dehydrogenase (ACDH, EC=1.2.1.10), an NAD+ and CoA-dependent acetaldehyde dehydrogenase, functions as a single enzyme (such as the Ethanolamine utilization protein, EutE, in Salmonella typhimurium) or as part of a multifunctional enzyme to convert acetaldehyde into acetyl-CoA . The E. coli aldehyde-alcohol dehydrogenase includes the functional domains, alcohol dehydrogenase (ADH), ACDH, and pyruvate-formate-lyase deactivase; and the Entamoeba histolytica aldehyde-alcohol dehydrogenase 2 (ALDH20A1) includes the functional domains ADH and ACDH and may be critical enzymes in the fermentative pathway.
Probab=24.49 E-value=2e+02 Score=24.03 Aligned_cols=40 Identities=13% Similarity=0.211 Sum_probs=31.1
Q ss_pred HHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHHH
Q 032798 71 VGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYNK 110 (133)
Q Consensus 71 i~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~~ 110 (133)
.++..-..|..++.++|..+...+.+..+++..++.....
T Consensus 7 ~A~~A~~~W~~~~~~eR~~~L~~~a~~l~~~~eela~~~~ 46 (436)
T cd07122 7 RARKAQREFATFSQEQVDKIVEAVAWAAADAAEELAKMAV 46 (436)
T ss_pred HHHHHHHHHHhCCHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 3444556799999999999998888888888888776643
No 63
>cd08317 Death_ank Death domain associated with Ankyrins. Death Domain (DD) associated with Ankyrins. Ankyrins are modular proteins comprising three conserved domains, an N-terminal membrane-binding domain containing ANK repeats, a spectrin-binding domain and a C-terminal DD. Ankyrins function as adaptor proteins and they interact, through ANK repeats, with structurally diverse membrane proteins, including ion channels/pumps, calcium release channels, and cell adhesion molecules. They play critical roles in the proper expression and membrane localization of these proteins. In mammals, this family includes ankyrin-R for restricted (or ANK1), ankyrin-B for broadly expressed (or ANK2) and ankyrin-G for general or giant (or ANK3). They are expressed in different combinations in many tissues and play non-overlapping functions. In general, DDs are protein-protein interaction domains found in a variety of domain architectures. Their common feature is that they form homodimers by self-associati
Probab=24.42 E-value=38 Score=21.41 Aligned_cols=19 Identities=16% Similarity=0.513 Sum_probs=16.2
Q ss_pred CCHHHHHHHHHHHhcCCCh
Q 032798 66 KSVATVGKAAGEKWKSMSE 84 (133)
Q Consensus 66 ~~~~ei~k~l~~~Wk~ls~ 84 (133)
..+..|+..||..|+.|..
T Consensus 5 ~~l~~ia~~lG~dW~~LAr 23 (84)
T cd08317 5 IRLADISNLLGSDWPQLAR 23 (84)
T ss_pred chHHHHHHHHhhHHHHHHH
Confidence 6788999999999987654
No 64
>PF02026 RyR: RyR domain; InterPro: IPR003032 This domain is called RyR for Ryanodine receptor []. The domain is found in four copies in the ryanodine receptor. The function of this domain is unknown.; PDB: 4ETV_A 3RQR_A 4ETT_A 4ERT_A 4ESU_A 4ETU_A 4ERV_A 3NRT_E.
Probab=24.18 E-value=71 Score=20.90 Aligned_cols=19 Identities=21% Similarity=0.323 Sum_probs=14.7
Q ss_pred hcCCChhhhHHHHHHHHHH
Q 032798 79 WKSMSEDEKAPFVERAEKR 97 (133)
Q Consensus 79 Wk~ls~~eK~~y~~~A~~~ 97 (133)
|..|++++|..+.+.+.+.
T Consensus 61 y~~L~e~eK~~dr~~~~e~ 79 (94)
T PF02026_consen 61 YDELSEEEKEKDRDMVRET 79 (94)
T ss_dssp GGGS-HHHHHHHHHHHHHH
T ss_pred hhhCCHHHHHHhHHHHHHH
Confidence 8889999998888777664
No 65
>PRK10455 periplasmic protein; Reviewed
Probab=24.02 E-value=1.6e+02 Score=21.16 Aligned_cols=25 Identities=16% Similarity=0.294 Sum_probs=18.0
Q ss_pred HHHHHHHhcCCChhhhHHHHHHHHH
Q 032798 72 GKAAGEKWKSMSEDEKAPFVERAEK 96 (133)
Q Consensus 72 ~k~l~~~Wk~ls~~eK~~y~~~A~~ 96 (133)
.+.-..++..|++++|..|.+..++
T Consensus 120 ~~~~~qiy~vLTPEQr~q~~~~~ek 144 (161)
T PRK10455 120 METQNKIYNVLTPEQKKQFNANFEK 144 (161)
T ss_pred HHHHHHHHHhCCHHHHHHHHHHHHH
Confidence 3344567899999999988865543
No 66
>PF13945 NST1: Salt tolerance down-regulator
Probab=23.06 E-value=2e+02 Score=21.50 Aligned_cols=26 Identities=27% Similarity=0.434 Sum_probs=21.1
Q ss_pred HHHHHHHHHHHhcCCChhhhHHHHHH
Q 032798 68 VATVGKAAGEKWKSMSEDEKAPFVER 93 (133)
Q Consensus 68 ~~ei~k~l~~~Wk~ls~~eK~~y~~~ 93 (133)
..+....|-+-|-.|+++||......
T Consensus 100 s~eEre~LkeFW~SL~eeERr~LVkI 125 (190)
T PF13945_consen 100 SQEEREKLKEFWESLSEEERRSLVKI 125 (190)
T ss_pred hHHHHHHHHHHHHccCHHHHHHHHHh
Confidence 44666789999999999999877654
No 67
>cd07085 ALDH_F6_MMSDH Methylmalonate semialdehyde dehydrogenase and ALDH family members 6A1 and 6B2. Methylmalonate semialdehyde dehydrogenase (MMSDH, EC=1.2.1.27) [acylating] from Bacillus subtilis is involved in valine metabolism and catalyses the NAD+- and CoA-dependent oxidation of methylmalonate semialdehyde into propionyl-CoA. Mitochondrial human MMSDH ALDH6A1 and Arabidopsis MMSDH ALDH6B2 are also present in this CD.
Probab=23.01 E-value=2.3e+02 Score=23.68 Aligned_cols=37 Identities=11% Similarity=0.121 Sum_probs=28.3
Q ss_pred HHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHH
Q 032798 72 GKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDY 108 (133)
Q Consensus 72 ~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y 108 (133)
++.....|..++.++|..+...+....+++..++..-
T Consensus 47 A~~A~~~w~~~~~~~R~~~L~~~a~~l~~~~~el~~~ 83 (478)
T cd07085 47 AKAAFPAWSATPVLKRQQVMFKFRQLLEENLDELARL 83 (478)
T ss_pred HHHHHHHHhcCCHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 3444567999999999999988887777777666653
No 68
>cd07087 ALDH_F3-13-14_CALDH-like ALDH subfamily: Coniferyl aldehyde dehydrogenase, ALDH families 3, 13, and 14, and other related proteins. ALDH subfamily which includes NAD(P)+-dependent, aldehyde dehydrogenase, family 3 member A1 and B1 (ALDH3A1, ALDH3B1, EC=1.2.1.5) and fatty aldehyde dehydrogenase, family 3 member A2 (ALDH3A2, EC=1.2.1.3), and also plant ALDH family members ALDH3F1, ALDH3H1, and ALDH3I1, fungal ALDH14 (YMR110C) and the protozoan family 13 member (ALDH13), as well as coniferyl aldehyde dehydrogenases (CALDH, EC=1.2.1.68), and other similar sequences, such as the Pseudomonas putida benzaldehyde dehydrogenase I that is involved in the metabolism of mandelate.
Probab=22.98 E-value=2.4e+02 Score=23.20 Aligned_cols=40 Identities=5% Similarity=-0.078 Sum_probs=29.8
Q ss_pred HHHHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHH
Q 032798 70 TVGKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN 109 (133)
Q Consensus 70 ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~ 109 (133)
+.++..-..|..++..+|..+...+.+..+++..++.+..
T Consensus 5 ~~a~~a~~~w~~~~~~~R~~~L~~~a~~l~~~~~el~~~~ 44 (426)
T cd07087 5 ARLRETFLTGKTRSLEWRKAQLKALKRMLTENEEEIAAAL 44 (426)
T ss_pred HHHHHHHHhcCCCCHHHHHHHHHHHHHHHHHhHHHHHHHH
Confidence 3345555679999999999988888877777777766553
No 69
>PF07813 LTXXQ: LTXXQ motif family protein; InterPro: IPR012899 This five residue motif is found in a number of bacterial proteins bearing similarity to the protein CpxP (P32158 from SWISSPROT). This is a periplasmic protein that aids in combating extracytoplasmic protein-mediated toxicity, and may also be involved in the response to alkaline pH []. Another member of this family, Spy (P77754 from SWISSPROT) is also a periplasmic protein that may be involved in the response to stress []. The homology between CpxP and Spy may indicate that these two proteins are functionally related []. The motif is found repeated twice in many members of this entry. ; GO: 0042597 periplasmic space; PDB: 3ITF_B 3QZC_B 3OEO_D 3O39_A.
Probab=22.52 E-value=1.5e+02 Score=18.40 Aligned_cols=25 Identities=16% Similarity=0.234 Sum_probs=18.3
Q ss_pred HHHHHHHHHHhcCCChhhhHHHHHH
Q 032798 69 ATVGKAAGEKWKSMSEDEKAPFVER 93 (133)
Q Consensus 69 ~ei~k~l~~~Wk~ls~~eK~~y~~~ 93 (133)
..+.......+..|++++|..|..+
T Consensus 75 ~~~~~~~~~~~~vLt~eQk~~~~~l 99 (100)
T PF07813_consen 75 EERAKAQHALYAVLTPEQKEKFDQL 99 (100)
T ss_dssp HHHHHHHHHHHTTS-HHHHHHHHHH
T ss_pred HHHHHHHHHHHhcCCHHHHHHHHHh
Confidence 3455666788999999999988754
No 70
>PF09791 Oxidored-like: Oxidoreductase-like protein, N-terminal; InterPro: IPR019180 This entry represents the N-terminal domain of various oxidoreductase-like proteins whose exact function is, as yet, unknown.
Probab=22.24 E-value=1.4e+02 Score=17.15 Aligned_cols=15 Identities=7% Similarity=0.348 Sum_probs=6.4
Q ss_pred HHHHHHHHHHHHHHH
Q 032798 94 AEKRKSDYNKNMQDY 108 (133)
Q Consensus 94 A~~~k~~y~~e~~~y 108 (133)
+.++.++|...++.+
T Consensus 31 Y~eel~~y~~~~~~~ 45 (48)
T PF09791_consen 31 YAEELEEYREALAAW 45 (48)
T ss_pred HHHHHHHHHHHHHHH
Confidence 333344444444444
No 71
>PF08367 M16C_assoc: Peptidase M16C associated; InterPro: IPR013578 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Metalloproteases are the most diverse of the four main types of protease, with more than 50 families identified to date. In these enzymes, a divalent cation, usually zinc, activates the water molecule. The metal ion is held in place by amino acid ligands, usually three in number. The known metal ligands are His, Glu, Asp or Lys and at least one other residue is required for catalysis, which may play an electrophillic role. Of the known metalloproteases, around half contain an HEXXH motif, which has been shown in crystallographic studies to form part of the metal-binding site []. The HEXXH motif is relatively common, but can be more stringently defined for metalloproteases as 'abXHEbbHbc', where 'a' is most often valine or threonine and forms part of the S1' subsite in thermolysin and neprilysin, 'b' is an uncharged residue, and 'c' a hydrophobic residue. Proline is never found in this site, possibly because it would break the helical structure adopted by this motif in metalloproteases []. This domain appears in eukaryotes as well as bacteria and tends to be found near the C terminus of metalloproteases and related sequences belonging to MEROPS peptidase family M16 (subfamily M16C, clan ME). These include: eupitrilysin, falcilysin, PreP peptidase, CYM1 peptidase and subfamily M16C non-peptidase homologues.; GO: 0008237 metallopeptidase activity, 0008270 zinc ion binding, 0006508 proteolysis; PDB: 2FGE_B 3S5I_A 3S5H_A 3S5M_A 3S5K_A.
Probab=21.84 E-value=1.8e+02 Score=21.97 Aligned_cols=32 Identities=22% Similarity=0.250 Sum_probs=25.8
Q ss_pred HHHHHHHHHHhcCCChhhhHHHHHHHHHHHHH
Q 032798 69 ATVGKAAGEKWKSMSEDEKAPFVERAEKRKSD 100 (133)
Q Consensus 69 ~ei~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~ 100 (133)
.+..+.|.+.+..|++++++...+.+++.++.
T Consensus 13 ~~e~~~L~~~k~~Ls~~e~~~i~~~~~~L~~~ 44 (248)
T PF08367_consen 13 EEEKEKLAAYKASLSEEEKEKIIEQTKELKER 44 (248)
T ss_dssp HHHHHHHHHHHHCS-HHHHHHHHHHHHHHHHH
T ss_pred HHHHHHHHHHHhhCCHHHHHHHHHHHHHHHHH
Confidence 46678899999999999999999888887543
No 72
>PRK13252 betaine aldehyde dehydrogenase; Provisional
Probab=21.14 E-value=2.5e+02 Score=23.54 Aligned_cols=38 Identities=18% Similarity=0.351 Sum_probs=29.2
Q ss_pred HHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHHH
Q 032798 72 GKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDYN 109 (133)
Q Consensus 72 ~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y~ 109 (133)
++.....|..++.++|..+...+....+.+..++..-.
T Consensus 53 A~~a~~~w~~~~~~~R~~~L~~~a~~l~~~~~ela~~~ 90 (488)
T PRK13252 53 AKQGQKIWAAMTAMERSRILRRAVDILRERNDELAALE 90 (488)
T ss_pred HHHHHHHHhcCCHHHHHHHHHHHHHHHHHhHHHHHHHH
Confidence 44456789999999999998877777777777766543
No 73
>cd07150 ALDH_VaniDH_like Pseudomonas putida vanillin dehydrogenase-like. Vanillin dehydrogenase (Vdh, VaniDH) involved in the metabolism of ferulic acid and other related sequences are included in this CD. The E. coli vanillin dehydrogenase (LigV) preferred NAD+ to NADP+ and exhibited a broad substrate preference, including vanillin, benzaldehyde, protocatechualdehyde, m-anisaldehyde, and p-hydroxybenzaldehyde.
Probab=20.97 E-value=2.5e+02 Score=23.07 Aligned_cols=37 Identities=14% Similarity=0.220 Sum_probs=27.6
Q ss_pred HHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHH
Q 032798 72 GKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDY 108 (133)
Q Consensus 72 ~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y 108 (133)
++..-..|..++.++|..+...+.+..+.+..++.+-
T Consensus 30 A~~A~~~w~~~~~~~R~~~L~~~a~~l~~~~~ela~~ 66 (451)
T cd07150 30 AYDAFPAWAATTPSERERILLKAAEIMERRADDLIDL 66 (451)
T ss_pred HHHHHHHHhcCCHHHHHHHHHHHHHHHHHhHHHHHHH
Confidence 3444567999999999999887777777777665543
No 74
>PHA03102 Small T antigen; Reviewed
Probab=20.67 E-value=1.5e+02 Score=21.30 Aligned_cols=36 Identities=19% Similarity=0.268 Sum_probs=23.9
Q ss_pred HHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhh
Q 032798 52 EEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEK 87 (133)
Q Consensus 52 ~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK 87 (133)
+..|...+.-|||.....+..+.|.+.|..|++..+
T Consensus 26 kAYr~la~~~HPDkgg~~e~~k~in~Ay~~L~d~~~ 61 (153)
T PHA03102 26 KAYLRKCLEFHPDKGGDEEKMKELNTLYKKFRESVK 61 (153)
T ss_pred HHHHHHHHHHCcCCCchhHHHHHHHHHHHHHhhHHH
Confidence 455666677899843445667777777777776543
No 75
>PTZ00037 DnaJ_C chaperone protein; Provisional
Probab=20.64 E-value=2.2e+02 Score=23.79 Aligned_cols=42 Identities=21% Similarity=0.295 Sum_probs=30.5
Q ss_pred HHHHHHHHHHhCCCCCCHHHHHHHHHHHhcCCChhhhH-HHHH
Q 032798 51 MEEFRKQFKEAHPNNKSVATVGKAAGEKWKSMSEDEKA-PFVE 92 (133)
Q Consensus 51 ~~~~r~~~k~~~p~~~~~~ei~k~l~~~Wk~ls~~eK~-~y~~ 92 (133)
-+.+|...++-|||.....+..+.|.+.|.-|++.+|. .|..
T Consensus 46 KkAYrkla~k~HPDk~~~~e~F~~i~~AYevLsD~~kR~~YD~ 88 (421)
T PTZ00037 46 KKAYRKLAIKHHPDKGGDPEKFKEISRAYEVLSDPEKRKIYDE 88 (421)
T ss_pred HHHHHHHHHHHCCCCCchHHHHHHHHHHHHHhccHHHHHHHhh
Confidence 45666667788999434457888999999999987755 4543
No 76
>TIGR02664 nitr_red_assoc conserved hypothetical protein. Most members of this protein family are found in the Cyanobacteria, and these mostly near nitrate reductase genes and molybdopterin biosynthesis genes. We note that molybdopterin guanine dinucleotide is a cofactor for nitrate reductase. This protein is sometimes annotated as nitrate reductase-associated protein. Its function is unknown.
Probab=20.24 E-value=2.5e+02 Score=20.11 Aligned_cols=44 Identities=14% Similarity=0.244 Sum_probs=29.8
Q ss_pred HHhcCCChhhhHHHHHHH---HHHHHHHHHHHHHHHHhhhhhhcccc
Q 032798 77 EKWKSMSEDEKAPFVERA---EKRKSDYNKNMQDYNKQLVIFFGIIV 120 (133)
Q Consensus 77 ~~Wk~ls~~eK~~y~~~A---~~~k~~y~~e~~~y~~k~~~~~~~~~ 120 (133)
..|..|+.+||+...+.. ..+...|..-+.+.-..+.......+
T Consensus 33 ~hW~~ls~~eRq~Lv~~pc~t~~e~~~yr~~L~~l~~~~a~~~~~~l 79 (145)
T TIGR02664 33 EHWQQLTQAEREELVRLPCDTAEVIDPYREYLRDLLRTHADTPPSDL 79 (145)
T ss_pred HHHhhCCHHHHHHHHhCccCCHHHHHHHHHHHHHHHHHHcCCCCcCC
Confidence 569999999999998765 23345677766666655554444433
No 77
>cd07152 ALDH_BenzADH NAD-dependent benzaldehyde dehydrogenase II-like. NAD-dependent, benzaldehyde dehydrogenase II (XylC, BenzADH, EC=1.2.1.28) is involved in the oxidation of benzyl alcohol to benzoate. In Acinetobacter calcoaceticus, this process is carried out by the chromosomally encoded, benzyl alcohol dehydrogenase (xylB) and benzaldehyde dehydrogenase II (xylC) enzymes; whereas in Pseudomonas putida they are encoded by TOL plasmids.
Probab=20.07 E-value=2.8e+02 Score=22.78 Aligned_cols=37 Identities=22% Similarity=0.409 Sum_probs=28.1
Q ss_pred HHHHHHHhcCCChhhhHHHHHHHHHHHHHHHHHHHHH
Q 032798 72 GKAAGEKWKSMSEDEKAPFVERAEKRKSDYNKNMQDY 108 (133)
Q Consensus 72 ~k~l~~~Wk~ls~~eK~~y~~~A~~~k~~y~~e~~~y 108 (133)
++..-..|..++.++|..+...+.+....+..++...
T Consensus 22 A~~A~~~w~~~~~~~R~~~L~~~a~~l~~~~~ela~~ 58 (443)
T cd07152 22 AAAAQRAWAATPPRERAAVLRRAADLLEEHADEIADW 58 (443)
T ss_pred HHHHHHHHhcCCHHHHHHHHHHHHHHHHHhHHHHHHH
Confidence 4444568999999999999988777777777666643
Done!