Query 033983
Match_columns 106
No_of_seqs 104 out of 253
Neff 4.5
Searched_HMMs 46136
Date Fri Mar 29 08:29:12 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/033983.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/033983hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 KOG2712 Transcriptional coacti 100.0 4.4E-35 9.5E-40 206.9 7.2 103 3-106 2-106 (108)
2 PF02229 PC4: Transcriptional 99.9 2.7E-25 5.9E-30 140.2 5.3 54 42-95 2-56 (56)
3 COG4443 Uncharacterized protei 93.0 0.059 1.3E-06 36.0 1.6 40 54-96 30-70 (72)
4 TIGR02530 flg_new flagellar op 72.2 4.5 9.7E-05 28.4 2.8 22 78-99 27-48 (96)
5 COG1448 TyrB Aspartate/tyrosin 64.5 5.8 0.00012 34.0 2.5 19 76-96 184-202 (396)
6 KOG3064 RNA-binding nuclear pr 61.4 17 0.00037 30.1 4.6 52 40-102 40-97 (303)
7 PF08880 QLQ: QLQ; InterPro: 60.5 8.2 0.00018 22.4 2.0 16 82-97 2-17 (37)
8 PF10815 ComZ: ComZ; InterPro 58.5 11 0.00023 24.2 2.4 22 77-98 22-43 (56)
9 PF12651 RHH_3: Ribbon-helix-h 53.4 20 0.00043 21.1 2.9 19 79-97 5-23 (44)
10 COG5129 MAK16 Nuclear protein 51.2 36 0.00077 27.9 4.8 52 40-102 39-96 (303)
11 PF01707 Peptidase_C9: Peptida 50.9 3.3 7E-05 32.6 -1.1 14 79-92 62-75 (202)
12 cd04754 Commd6 COMM_Domain con 50.0 63 0.0014 22.3 5.3 43 58-102 42-85 (86)
13 PF13101 DUF3945: Protein of u 44.5 14 0.00031 23.0 1.3 15 79-93 31-45 (59)
14 PF08988 DUF1895: Protein of u 40.4 44 0.00094 21.8 3.2 21 83-103 38-58 (68)
15 KOG1412 Aspartate aminotransfe 39.8 22 0.00049 30.4 2.2 21 74-96 188-208 (410)
16 TIGR02501 type_III_yscE type I 39.0 45 0.00098 21.6 3.1 21 83-103 37-57 (67)
17 smart00712 PUR DNA/RNA-binding 38.9 1E+02 0.0022 19.6 6.0 49 39-101 12-61 (63)
18 PF13487 HD_5: HD domain; PDB: 38.1 50 0.0011 20.2 3.1 24 82-105 2-25 (64)
19 PF04358 DsrC: DsrC like prote 35.9 23 0.00049 25.0 1.4 18 77-94 33-50 (109)
20 PF06526 DUF1107: Protein of u 35.5 24 0.00051 23.1 1.3 45 55-106 19-64 (64)
21 COG2921 Uncharacterized conser 34.8 36 0.00078 23.7 2.2 32 68-99 44-84 (90)
22 cd07999 GH7_CBH_EG Glycosyl hy 31.7 56 0.0012 28.1 3.3 41 47-87 256-307 (386)
23 PF06831 H2TH: Formamidopyrimi 31.7 47 0.001 22.2 2.4 31 72-102 51-82 (92)
24 PF02866 Ldh_1_C: lactate/mala 31.1 65 0.0014 23.2 3.2 44 60-104 126-169 (174)
25 TIGR03342 dsrC_tusE_dsvC sulfu 30.2 62 0.0013 22.9 2.8 17 78-94 33-49 (108)
26 TIGR02675 tape_meas_nterm tape 30.0 66 0.0014 20.8 2.8 22 83-104 31-52 (75)
27 PF09655 Nitr_red_assoc: Conse 30.0 21 0.00046 26.7 0.5 17 77-93 102-118 (144)
28 PF11580 DUF3239: Protein of u 28.3 34 0.00074 25.0 1.3 19 82-100 110-128 (128)
29 COG3530 Uncharacterized protei 27.4 30 0.00066 23.0 0.8 29 53-82 16-48 (71)
30 PF08743 Nse4_C: Nse4 C-termin 27.2 48 0.001 22.2 1.8 16 81-96 69-84 (93)
31 PRK11508 sulfur transfer prote 26.9 76 0.0017 22.5 2.8 18 78-95 34-51 (109)
32 cd01277 HINT_subgroup HINT (hi 26.5 90 0.0019 20.0 3.0 24 82-105 50-73 (103)
33 PF03102 NeuB: NeuB family; I 26.3 66 0.0014 25.4 2.7 25 79-103 213-237 (241)
34 PF11006 DUF2845: Protein of u 26.3 1.2E+02 0.0025 20.0 3.5 28 39-66 58-85 (87)
35 TIGR03586 PseI pseudaminic aci 26.1 74 0.0016 26.2 3.0 24 81-104 236-259 (327)
36 PF08848 DUF1818: Domain of un 26.0 1E+02 0.0023 22.3 3.4 25 81-105 29-53 (117)
37 PF00840 Glyco_hydro_7: Glycos 25.3 90 0.002 27.2 3.5 47 47-93 280-339 (433)
38 COG2920 DsrC Dissimilatory sul 24.4 50 0.0011 23.8 1.5 20 73-92 31-50 (111)
39 cd02679 MIT_spastin MIT: domai 24.0 53 0.0012 21.9 1.5 31 72-104 41-71 (79)
40 PRK13398 3-deoxy-7-phosphohept 22.9 81 0.0018 25.1 2.6 24 79-102 243-266 (266)
41 PF11325 DUF3127: Domain of un 22.5 88 0.0019 21.2 2.4 17 51-67 65-81 (84)
42 PRK06223 malate dehydrogenase; 22.4 1.1E+02 0.0025 23.8 3.3 45 60-106 263-307 (307)
43 PF14164 YqzH: YqzH-like prote 22.3 96 0.0021 20.3 2.4 18 81-98 24-41 (64)
44 TIGR03569 NeuB_NnaB N-acetylne 21.8 97 0.0021 25.6 3.0 25 80-104 236-260 (329)
45 PRK08673 3-deoxy-7-phosphohept 21.7 1E+02 0.0022 25.6 3.1 25 80-104 310-334 (335)
46 PRK13396 3-deoxy-7-phosphohept 21.3 1.1E+02 0.0023 25.8 3.1 25 80-104 319-343 (352)
47 cd01275 FHIT FHIT (fragile his 21.1 1.2E+02 0.0027 20.5 3.0 40 56-104 33-72 (126)
48 COG2089 SpsE Sialic acid synth 20.8 1.1E+02 0.0024 26.0 3.1 25 79-103 247-271 (347)
49 PRK10287 thiosulfate:cyanide s 20.4 64 0.0014 21.9 1.4 34 55-90 17-50 (104)
No 1
>KOG2712 consensus Transcriptional coactivator [Transcription]
Probab=100.00 E-value=4.4e-35 Score=206.88 Aligned_cols=103 Identities=50% Similarity=0.827 Sum_probs=85.0
Q ss_pred CCCcccccccccCC--CCCCCCCCCCCCCCCCCCCCCCCcEEEEcCCceEEEEeeeCCceEEEeEEEEecCCeecCcccc
Q 033983 3 GKGKRKEEEEYDSD--GSVDGHAPPKKASKTDSSDDSDDIVVCEISKNRRVSVRNWQGKVWVDIREFYVKEGKKFPGKKG 80 (106)
Q Consensus 3 ~~~k~k~~~~~~sd--~~~~~~~~~Kk~~~~~~~~~~~~~~~~~Ls~~rrVtV~~FkG~~~VdIREyY~kdGe~~PgKKG 80 (106)
.++.++....+.++ ++...++|+++..+.. .+++++.++|+|+++|||||++|+|+.||||||||.++|+++||+||
T Consensus 2 s~~~~~~~~~r~~~~~~~~~~~a~~~~v~k~~-d~~s~~~~i~~l~~~RrVtV~eFkGk~~VdIREyY~kdG~mlPgkKG 80 (108)
T KOG2712|consen 2 SSSSRKDVDSRVDKKLKEKKSHAPNKKVEKPK-DDDSEDDNIFNLGKNRRVTVREFKGKILVDIREYYVKDGKMLPGKKG 80 (108)
T ss_pred ccccccCccccccccccchhhhCCCccccCcc-cCCcCccceeecCCceEEehhhcCCceEEehhHhhhccCccccCccc
Confidence 34555555444444 4566677776655532 22566678999999999999999999999999999999999999999
Q ss_pred eecCHHHHHHHHHhHHHHHHHhhcCC
Q 033983 81 ISLSVDQWNTLRDHVEEINKALGDNS 106 (106)
Q Consensus 81 ISL~~eqw~~L~~~~~~Id~ai~~~~ 106 (106)
||||++||..|++++++||+||.+|+
T Consensus 81 ISLs~~qW~~Lk~~~~eId~Al~~l~ 106 (108)
T KOG2712|consen 81 ISLSLEQWSKLKEHIEEIDKALRKLS 106 (108)
T ss_pred cccCHHHHHHHHHHHHHHHHHHHHhc
Confidence 99999999999999999999999885
No 2
>PF02229 PC4: Transcriptional Coactivator p15 (PC4); InterPro: IPR003173 p15 has a bipartite structure composed of an amino-terminal regulatory domain and a carboxy-terminal cryptic DNA-binding domain []. The DNA-binding activity of the carboxy-terminal is disguised by the amino-terminal p15 domain. Activity is controlled by protein kinases that target the regulatory domain.; GO: 0003677 DNA binding, 0003713 transcription coactivator activity, 0006355 regulation of transcription, DNA-dependent; PDB: 3PM7_B 2LTD_A 2LTT_B 3OBH_B 2L3A_B 2PHE_B 1PCF_B 2C62_B.
Probab=99.92 E-value=2.7e-25 Score=140.17 Aligned_cols=54 Identities=46% Similarity=0.861 Sum_probs=49.2
Q ss_pred EEEcCCceEEEEeeeCCceEEEeEEEEec-CCeecCcccceecCHHHHHHHHHhH
Q 033983 42 VCEISKNRRVSVRNWQGKVWVDIREFYVK-EGKKFPGKKGISLSVDQWNTLRDHV 95 (106)
Q Consensus 42 ~~~Ls~~rrVtV~~FkG~~~VdIREyY~k-dGe~~PgKKGISL~~eqw~~L~~~~ 95 (106)
+|+++.+++|+|++|+|++|||||+||.+ +|+|+||+|||||+++||.+|++++
T Consensus 2 ~~~~~~~~rv~v~~fkG~~~vdIRe~y~~~~g~~~P~kKGIsL~~~q~~~l~~~l 56 (56)
T PF02229_consen 2 IKNLGEKRRVSVSEFKGKPYVDIREWYEKKDGEWKPTKKGISLTPEQWKELKEAL 56 (56)
T ss_dssp EETTEEEEEEEEEEETTSEEEEEEEEETTSSS-EEEEEEEEEE-HHHHHHHHHH-
T ss_pred cccCCCeEEEEEEEeCCeEEEEEEeeEEcCCCcCcCcCCEEEcCHHHHHHHHhhC
Confidence 57899999999999999999999999997 8999999999999999999999874
No 3
>COG4443 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=93.04 E-value=0.059 Score=35.95 Aligned_cols=40 Identities=30% Similarity=0.604 Sum_probs=31.0
Q ss_pred eeeCCce-EEEeEEEEecCCeecCcccceecCHHHHHHHHHhHH
Q 033983 54 RNWQGKV-WVDIREFYVKEGKKFPGKKGISLSVDQWNTLRDHVE 96 (106)
Q Consensus 54 ~~FkG~~-~VdIREyY~kdGe~~PgKKGISL~~eqw~~L~~~~~ 96 (106)
-.|+|++ -.|||.|-.+.- +-| |||+|+.+++..|++.+.
T Consensus 30 vSwNg~~~KyDiR~Wspdh~--KMG-KGiTLt~eE~~~l~d~l~ 70 (72)
T COG4443 30 VSWNGRPPKYDIRAWSPDHS--KMG-KGITLTNEEFKALKDLLN 70 (72)
T ss_pred cccCCCCCcCcccccCcchh--hhc-CceeecHHHHHHHHHHHh
Confidence 3578876 789999976532 335 899999999999988764
No 4
>TIGR02530 flg_new flagellar operon protein. Members of this family are found in a subset of bacterial flagellar operons, generally between genes designated flgD and flgE, in species as diverse as Bacillus halodurans and various other Firmicutes, Geobacter sulfurreducens, and Bdellovibrio bacteriovorus. The specific molecular function is unknown.
Probab=72.22 E-value=4.5 Score=28.38 Aligned_cols=22 Identities=36% Similarity=0.720 Sum_probs=18.5
Q ss_pred ccceecCHHHHHHHHHhHHHHH
Q 033983 78 KKGISLSVDQWNTLRDHVEEIN 99 (106)
Q Consensus 78 KKGISL~~eqw~~L~~~~~~Id 99 (106)
..||+|+.++|..|.+++....
T Consensus 27 ~R~I~l~~~~~~~i~~av~~A~ 48 (96)
T TIGR02530 27 ERNISINPDDWKKLLEAVEEAE 48 (96)
T ss_pred HcCCCCCHHHHHHHHHHHHHHH
Confidence 4799999999999988877654
No 5
>COG1448 TyrB Aspartate/tyrosine/aromatic aminotransferase [Amino acid transport and metabolism]
Probab=64.50 E-value=5.8 Score=34.04 Aligned_cols=19 Identities=37% Similarity=0.800 Sum_probs=16.8
Q ss_pred CcccceecCHHHHHHHHHhHH
Q 033983 76 PGKKGISLSVDQWNTLRDHVE 96 (106)
Q Consensus 76 PgKKGISL~~eqw~~L~~~~~ 96 (106)
|| ||-||.+||..|.+.+.
T Consensus 184 PT--G~D~t~~qW~~l~~~~~ 202 (396)
T COG1448 184 PT--GIDPTEEQWQELADLIK 202 (396)
T ss_pred CC--CCCCCHHHHHHHHHHHH
Confidence 76 99999999999987765
No 6
>KOG3064 consensus RNA-binding nuclear protein (MAK16) containing a distinct C4 Zn-finger [RNA processing and modification]
Probab=61.37 E-value=17 Score=30.05 Aligned_cols=52 Identities=17% Similarity=0.463 Sum_probs=36.9
Q ss_pred cEEEEcCCceEEEEeeeCCceEEEeEEEEecCCeecCcccceecCHHHHHHHH------HhHHHHHHHh
Q 033983 40 IVVCEISKNRRVSVRNWQGKVWVDIREFYVKEGKKFPGKKGISLSVDQWNTLR------DHVEEINKAL 102 (106)
Q Consensus 40 ~~~~~Ls~~rrVtV~~FkG~~~VdIREyY~kdGe~~PgKKGISL~~eqw~~L~------~~~~~Id~ai 102 (106)
.+.|.|.+-|++||++=+|..|+-+-. | --..++...|+.++ .++..|++-|
T Consensus 40 R~SCPLANSrYATVre~~g~~yLymKt---------~--ERaH~P~klwErikLSkNyekALeQIde~L 97 (303)
T KOG3064|consen 40 RSSCPLANSRYATVREENGVLYLYMKT---------I--ERAHMPRKLWERIKLSKNYEKALEQIDEQL 97 (303)
T ss_pred cccCcCccccceeEeecCCEEEEEEec---------h--hhhcCcHHHHHHHhcchhHHHHHHHHHHHH
Confidence 467999999999999999999975433 1 23446666777654 5566666544
No 7
>PF08880 QLQ: QLQ; InterPro: IPR014978 QLQ is named after the conserved Gln, Leu, Gln motif. QLQ is found at the N terminus of SWI2/SNF2 protein, which has been shown to be involved in protein-protein interactions. QLQ has been postulated to be involved in mediating protein interactions []. ; GO: 0005524 ATP binding, 0016818 hydrolase activity, acting on acid anhydrides, in phosphorus-containing anhydrides, 0006355 regulation of transcription, DNA-dependent, 0005634 nucleus
Probab=60.45 E-value=8.2 Score=22.44 Aligned_cols=16 Identities=19% Similarity=0.300 Sum_probs=13.7
Q ss_pred ecCHHHHHHHHHhHHH
Q 033983 82 SLSVDQWNTLRDHVEE 97 (106)
Q Consensus 82 SL~~eqw~~L~~~~~~ 97 (106)
++|++||..|+..+-.
T Consensus 2 ~FT~~Ql~~L~~Qi~a 17 (37)
T PF08880_consen 2 PFTPAQLQELRAQILA 17 (37)
T ss_pred CCCHHHHHHHHHHHHH
Confidence 5899999999998764
No 8
>PF10815 ComZ: ComZ; InterPro: IPR024558 ComZ, which contains a leucine zipper motif, negatively regulates transcription of the ComG operon [].
Probab=58.51 E-value=11 Score=24.19 Aligned_cols=22 Identities=32% Similarity=0.519 Sum_probs=19.0
Q ss_pred cccceecCHHHHHHHHHhHHHH
Q 033983 77 GKKGISLSVDQWNTLRDHVEEI 98 (106)
Q Consensus 77 gKKGISL~~eqw~~L~~~~~~I 98 (106)
-++||-|+.++...|.+.+-.+
T Consensus 22 ~k~GIeLsme~~qP~m~L~~~V 43 (56)
T PF10815_consen 22 DKKGIELSMEMLQPLMQLLTKV 43 (56)
T ss_pred HHcCccCCHHHHHHHHHHHHHH
Confidence 3689999999999998887765
No 9
>PF12651 RHH_3: Ribbon-helix-helix domain
Probab=53.37 E-value=20 Score=21.13 Aligned_cols=19 Identities=26% Similarity=0.378 Sum_probs=16.0
Q ss_pred cceecCHHHHHHHHHhHHH
Q 033983 79 KGISLSVDQWNTLRDHVEE 97 (106)
Q Consensus 79 KGISL~~eqw~~L~~~~~~ 97 (106)
=+++|+.+++..|.+...+
T Consensus 5 ~t~~l~~el~~~L~~ls~~ 23 (44)
T PF12651_consen 5 FTFSLDKELYEKLKELSEE 23 (44)
T ss_pred EEEecCHHHHHHHHHHHHH
Confidence 3789999999999987655
No 10
>COG5129 MAK16 Nuclear protein with HMG-like acidic region [General function prediction only]
Probab=51.19 E-value=36 Score=27.92 Aligned_cols=52 Identities=19% Similarity=0.530 Sum_probs=36.6
Q ss_pred cEEEEcCCceEEEEeeeCCceEEEeEEEEecCCeecCcccceecCHHHHHHHH------HhHHHHHHHh
Q 033983 40 IVVCEISKNRRVSVRNWQGKVWVDIREFYVKEGKKFPGKKGISLSVDQWNTLR------DHVEEINKAL 102 (106)
Q Consensus 40 ~~~~~Ls~~rrVtV~~FkG~~~VdIREyY~kdGe~~PgKKGISL~~eqw~~L~------~~~~~Id~ai 102 (106)
.+.|.|.+.|++||+.-.|+.|+-+.+ | --...+...|+.++ .++.+||+.|
T Consensus 39 RqSCPLANSrYATVr~dngkLyLymKt---------p--ERaH~P~klwerIkLSkNY~kAL~QIde~L 96 (303)
T COG5129 39 RQSCPLANSRYATVRADNGKLYLYMKT---------P--ERAHVPRKLWERIKLSKNYEKALKQIDESL 96 (303)
T ss_pred cccCcCccCcceEEEecCCEEEEEecC---------h--hhccCcHHHHHHHHhhhhHHHHHHHHHHHH
Confidence 568999999999999999999874332 2 23456667777664 4555666544
No 11
>PF01707 Peptidase_C9: Peptidase family C9; InterPro: IPR002620 The family of alphaviruses includes 26 known members. They infect a variety of hosts including mosquitoes, birds, rodents and other mammals with worldwide distribution. Alphaviruses also pose a potential threat to human health in many area. For example, Venezuelan Equine Encephalitis Virus (VEEV) causes encephalitis in humans as well as livestock in Central and South America, and some variants of Sinbis Virus (SIN) and Semliki Forest Virus (SFV) have been found to cause fever and arthritis in humans []. Alphaviruses possess a single-stranded RNA genome of approximately 12 kb. The genomic RNA of alphaviruses is translated into two polyproteins that, respectively, encode structural proteins and nonstructural proteins. The nonstructural proteins may be translated as one or two polyproteins, nsp123 or nsp1234, depending on the virus. These polyproteins are cleaved to generate nsp1, nsp2, nsp3 and nsp4 by a protease activity that resides within nsp2 []. The nsp2 protein of alphaviruses has multiple enzymatic acivities. Its N-terminal domain has been shown to possess ATPase and GTPase activity, RNA helicase activity and RNA 5'-triphosphatase activity []. The C-terminal nsp2pro domain of nsp2 is responsible for the regulation of 26S subgenome RNA synthesis, switching between negative- and positive-strand RNA synthesis, targeting nsp2 for nuclear transport and proteolytic processing of the nonstructural polyprotein [, ]. The nsp2pro domain is a member of peptidase family C9 of clan CA. The nsp2pro domain consists of two distinct subdomains. The nsp2pro N-terminal subdomain is largely alpha-helical and contains the catalytic dyad cysteine and histidine residues organised in a protein fold that differs significantly from any known cysteine protease or protein folds. The nsp2pro C-terminal subdomain displays structural similarity to S-adenosyl- L-methionine-dependent RNA methyltransferases and provides essential elements that contribute to substrate recognition and may also regulate the structure of the substrate binding cleft []. This entry represents the nsp2pro domain.; PDB: 3TRK_A 2HWK_A.
Probab=50.90 E-value=3.3 Score=32.60 Aligned_cols=14 Identities=50% Similarity=0.947 Sum_probs=8.8
Q ss_pred cceecCHHHHHHHH
Q 033983 79 KGISLSVDQWNTLR 92 (106)
Q Consensus 79 KGISL~~eqw~~L~ 92 (106)
-||.||.+||+.|-
T Consensus 62 AGI~LT~~qW~~l~ 75 (202)
T PF01707_consen 62 AGIQLTAEQWSTLF 75 (202)
T ss_dssp TT----HHHHCCCH
T ss_pred cCcccCHHHHHHHh
Confidence 69999999999886
No 12
>cd04754 Commd6 COMM_Domain containing protein 6. The COMM Domain is found at the C-terminus of a variety of proteins; presumably all COMM_Domain containing proteins are located in the nucleus and the COMM domain plays a role in protein-protein interactions. Several family members have been shown to bind and inhibit NF-kappaB.
Probab=50.02 E-value=63 Score=22.29 Aligned_cols=43 Identities=12% Similarity=0.315 Sum_probs=35.4
Q ss_pred CceEEEeEEEEec-CCeecCcccceecCHHHHHHHHHhHHHHHHHh
Q 033983 58 GKVWVDIREFYVK-EGKKFPGKKGISLSVDQWNTLRDHVEEINKAL 102 (106)
Q Consensus 58 G~~~VdIREyY~k-dGe~~PgKKGISL~~eqw~~L~~~~~~Id~ai 102 (106)
|.+||.+--=..+ +|...| +-+-||.+|+..|...+.++.+.|
T Consensus 42 ~~Pfl~l~L~V~~~~G~~~~--~~~EmTlpEFq~f~~~~~~~~a~l 85 (86)
T cd04754 42 NSPYVAVTLKVADPSGQVVT--KSFEMTIPEFQNFSRQFKEMAAVL 85 (86)
T ss_pred CCceEEEEEEEEccCCCccc--eEEEEcHHHHHHHHHHHHHHHHhc
Confidence 7889887665555 788877 599999999999999988887654
No 13
>PF13101 DUF3945: Protein of unknown function (DUF3945)
Probab=44.46 E-value=14 Score=23.02 Aligned_cols=15 Identities=47% Similarity=0.705 Sum_probs=13.5
Q ss_pred cceecCHHHHHHHHH
Q 033983 79 KGISLSVDQWNTLRD 93 (106)
Q Consensus 79 KGISL~~eqw~~L~~ 93 (106)
+|+.||++|.+.|++
T Consensus 31 ~g~~Ls~~q~~~L~~ 45 (59)
T PF13101_consen 31 KGVELSPEQKEDLRE 45 (59)
T ss_pred cCccCCHHHHHHHHC
Confidence 799999999999875
No 14
>PF08988 DUF1895: Protein of unknown function (DUF1895); InterPro: IPR015081 The YscE protein, produced by the pathogen Yersinia, assumes a secondary structure composed of two anti-parallel alpha-helices separated by a flexible loop. The function of this protein is, as yet, unknown. ; PDB: 1ZW0_B 2P58_A 2UWJ_E 2Q1K_D 3PH0_B.
Probab=40.38 E-value=44 Score=21.82 Aligned_cols=21 Identities=14% Similarity=0.352 Sum_probs=18.7
Q ss_pred cCHHHHHHHHHhHHHHHHHhh
Q 033983 83 LSVDQWNTLRDHVEEINKALG 103 (106)
Q Consensus 83 L~~eqw~~L~~~~~~Id~ai~ 103 (106)
++|+||..+....+.|..|++
T Consensus 38 ~~P~eyQq~q~~~~AieAA~~ 58 (68)
T PF08988_consen 38 GTPQEYQQLQQQYDAIEAAIA 58 (68)
T ss_dssp SSHHHHHHHHHHHHHHHHHHH
T ss_pred CCHHHHHHHHHHHHHHHHHHH
Confidence 689999999999999998875
No 15
>KOG1412 consensus Aspartate aminotransferase/Glutamic oxaloacetic transaminase AAT2/GOT1 [Amino acid transport and metabolism]
Probab=39.82 E-value=22 Score=30.44 Aligned_cols=21 Identities=24% Similarity=0.642 Sum_probs=17.0
Q ss_pred ecCcccceecCHHHHHHHHHhHH
Q 033983 74 KFPGKKGISLSVDQWNTLRDHVE 96 (106)
Q Consensus 74 ~~PgKKGISL~~eqw~~L~~~~~ 96 (106)
.-|+ ||-.|.|||.++.+.|.
T Consensus 188 hNPT--GmDPT~EQW~qia~vik 208 (410)
T KOG1412|consen 188 HNPT--GMDPTREQWKQIADVIK 208 (410)
T ss_pred cCCC--CCCCCHHHHHHHHHHHH
Confidence 3465 99999999999977664
No 16
>TIGR02501 type_III_yscE type III secretion system protein, YseE family. Members of this family are found exclusively in type III secretion appparatus gene clusters in bacteria. Those bacteria with a protein from this family tend to target animal cells, as does Yersinia pestis. This protein is small (about 70 amino acids) and not well characterized.
Probab=38.99 E-value=45 Score=21.62 Aligned_cols=21 Identities=14% Similarity=0.143 Sum_probs=17.7
Q ss_pred cCHHHHHHHHHhHHHHHHHhh
Q 033983 83 LSVDQWNTLRDHVEEINKALG 103 (106)
Q Consensus 83 L~~eqw~~L~~~~~~Id~ai~ 103 (106)
.+|+||..|...+..++.|++
T Consensus 37 ~tp~qYq~l~~~~~A~~aA~~ 57 (67)
T TIGR02501 37 GDPQQYQEWQLLADAIEAAIK 57 (67)
T ss_pred CCHHHHHHHHHHHHHHHHHHH
Confidence 489999999888888888875
No 17
>smart00712 PUR DNA/RNA-binding repeats in PUR-alpha/beta/gamma and in hypothetical proteins from spirochetes and the Bacteroides-Cytophaga-Flexibacter bacteria.
Probab=38.92 E-value=1e+02 Score=19.61 Aligned_cols=49 Identities=24% Similarity=0.486 Sum_probs=32.7
Q ss_pred CcEEEEcCCceEEEEeeeCCceEEEeEEEEec-CCeecCcccceecCHHHHHHHHHhHHHHHHH
Q 033983 39 DIVVCEISKNRRVSVRNWQGKVWVDIREFYVK-EGKKFPGKKGISLSVDQWNTLRDHVEEINKA 101 (106)
Q Consensus 39 ~~~~~~Ls~~rrVtV~~FkG~~~VdIREyY~k-dGe~~PgKKGISL~~eqw~~L~~~~~~Id~a 101 (106)
-.++|+|..++| | .|+-|-|- + .+ ++-=|.|+.+.|..+++++.++-+-
T Consensus 12 k~fyfDvk~N~r-------G-~fLrIsE~--~~~~----~r~~I~lp~~~~~~F~~~l~~~~~~ 61 (63)
T smart00712 12 KRFYFDVKENRR-------G-RFLRISEV--KNNG----GRSSITVPEQGAAEFRDALNKLIEK 61 (63)
T ss_pred cEEEEEecccCC-------c-cEEEEEEe--cCCC----CceEEEEEHHHHHHHHHHHHHHHHh
Confidence 456677765543 4 55555552 2 11 2678999999999999998876543
No 18
>PF13487 HD_5: HD domain; PDB: 3TMD_A 3TM8_B 3TMC_A 3TMB_B.
Probab=38.11 E-value=50 Score=20.25 Aligned_cols=24 Identities=17% Similarity=0.209 Sum_probs=17.2
Q ss_pred ecCHHHHHHHHHhHHHHHHHhhcC
Q 033983 82 SLSVDQWNTLRDHVEEINKALGDN 105 (106)
Q Consensus 82 SL~~eqw~~L~~~~~~Id~ai~~~ 105 (106)
.||++||..++.+...--+.|.++
T Consensus 2 ~Lt~~e~~~~~~Hp~~~~~~l~~~ 25 (64)
T PF13487_consen 2 KLTPEEREIIQQHPEYGAELLSQI 25 (64)
T ss_dssp GS-HHHHHHHHHHHHHHHHHHTT-
T ss_pred CCCHHHHHHHHHHHHHHHHHHHcc
Confidence 489999999998877666666543
No 19
>PF04358 DsrC: DsrC like protein; InterPro: IPR007453 DsrC (P45573 from SWISSPROT) has been observed to co-purify with Desulphovibrio vulgaris dissimilatory sulphite reductase []. However, DsrC appears to be only loosely associated to the sulphite reductase, which suggests that it may not be an integral part of the dissimilatory sulphite reductase. Many proteins in this entry are found in organisms such as Escherichia coli and Haemophilus influenzae which do not contain dissimilatory sulphite reductases but can synthesise assimilatory sirohaem sulphite and nitrite reductases. It is speculated that DsrC may be involved in the assembly, folding or stabilisation of sirohaem proteins []. The strictly conserved cysteine in the C terminus suggests that DsrC may have a catalytic function in the metabolism of sulphur compounds []. Also included in this entry is TusE, a partner to TusBCD in a sulphur relay system for 2-thiouridine biosynthesis, a tRNA base modification process. Many proteins in this entry are annotated as the third (gamma) subunit of dissimilatory sulphite reductase ; PDB: 2V4J_F 2A5W_C 1SAU_A 1JI8_A 1YX3_A.
Probab=35.86 E-value=23 Score=24.95 Aligned_cols=18 Identities=28% Similarity=0.645 Sum_probs=11.3
Q ss_pred cccceecCHHHHHHHHHh
Q 033983 77 GKKGISLSVDQWNTLRDH 94 (106)
Q Consensus 77 gKKGISL~~eqw~~L~~~ 94 (106)
-.-||.||.++|+.+.-.
T Consensus 33 ~~egI~Ltd~HW~vI~fl 50 (109)
T PF04358_consen 33 KEEGIELTDEHWEVIRFL 50 (109)
T ss_dssp HCTT-S--HHHHHHHHHH
T ss_pred HHcCCCCCHHHHHHHHHH
Confidence 356999999999887543
No 20
>PF06526 DUF1107: Protein of unknown function (DUF1107); InterPro: IPR009491 This family consists of several short, hypothetical bacterial proteins of unknown function.; PDB: 2JRO_A.
Probab=35.46 E-value=24 Score=23.07 Aligned_cols=45 Identities=24% Similarity=0.383 Sum_probs=25.7
Q ss_pred eeCCceEEE-eEEEEecCCeecCcccceecCHHHHHHHHHhHHHHHHHhhcCC
Q 033983 55 NWQGKVWVD-IREFYVKEGKKFPGKKGISLSVDQWNTLRDHVEEINKALGDNS 106 (106)
Q Consensus 55 ~FkG~~~Vd-IREyY~kdGe~~PgKKGISL~~eqw~~L~~~~~~Id~ai~~~~ 106 (106)
-|+|..||+ |-.|--++|..++-++ ... .-...+.+|+.+|..|+
T Consensus 19 lF~Gr~~I~g~G~feFd~Gkillp~~----~~~---~~~~~~~EiN~~I~~L~ 64 (64)
T PF06526_consen 19 LFRGRIYIKGIGAFEFDNGKILLPKK----ADK---RHLSVMSEINQEIRRLS 64 (64)
T ss_dssp H-SEEEEETTTEEEEEETTEE---SS------H---HHHHHHHHHHHHHHHH-
T ss_pred HccceEEEEecccEEEcCCEEeCCcc----ccH---HHHHHHHHHHHHHHhcC
Confidence 388999885 5555448888775332 223 34455788888887764
No 21
>COG2921 Uncharacterized conserved protein [Function unknown]
Probab=34.79 E-value=36 Score=23.72 Aligned_cols=32 Identities=19% Similarity=0.268 Sum_probs=21.3
Q ss_pred EecCCeecCcccc---------eecCHHHHHHHHHhHHHHH
Q 033983 68 YVKEGKKFPGKKG---------ISLSVDQWNTLRDHVEEIN 99 (106)
Q Consensus 68 Y~kdGe~~PgKKG---------ISL~~eqw~~L~~~~~~Id 99 (106)
|..-=.|+|+.|| +..+.||.+.|-..+.+++
T Consensus 44 ~~~~~~~k~SSkGnY~svsI~i~A~~~EQ~e~ly~eL~~~~ 84 (90)
T COG2921 44 YTPRVSWKPSSKGNYLSVSITIRATNIEQVEALYRELRKHE 84 (90)
T ss_pred cCceeeeccCCCCceEEEEEEEEECCHHHHHHHHHHHhhCC
Confidence 3333457888888 4567788888877666543
No 22
>cd07999 GH7_CBH_EG Glycosyl hydrolase family 7. Glycosyl hydrolase family 7 contains eukaryotic endoglucanases (EGs) and cellobiohydrolases (CBHs) that hydrolyze glycosidic bonds using a double-displacement mechanism. This leads to a net retention of the conformation at the anomeric carbon. Both enzymes work synergistically in the degradation of cellulose,which is the main component of plant cell wall, and is composed of beta-1,4 linked glycosyl units. EG cleaves the beta-1,4 linkages of cellulose and CBH cleaves off cellobiose disaccharide units from the reducing end of the chain. In general, the O-glycosyl hydrolases are a widespread group of enzymes that hydrolyze the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A glycosyl hydrolase classification system based on sequence similarity has led to the definition of more than 95 different families inlcuding glycoside hydrolase family 7.
Probab=31.75 E-value=56 Score=28.10 Aligned_cols=41 Identities=22% Similarity=0.379 Sum_probs=29.4
Q ss_pred CceEEEEeeeC---CceEEEeEEEEecCCeecCccc----c----eecCHHH
Q 033983 47 KNRRVSVRNWQ---GKVWVDIREFYVKEGKKFPGKK----G----ISLSVDQ 87 (106)
Q Consensus 47 ~~rrVtV~~Fk---G~~~VdIREyY~kdGe~~PgKK----G----ISL~~eq 87 (106)
.+++-.|..|- |-.|..||-+|..+|+..|.-+ | =||+.+-
T Consensus 256 ~k~fTVVTQFit~~~G~LteIrR~YVQ~GkvI~n~~~~~~g~~~~~sitd~f 307 (386)
T cd07999 256 SKPFTVVTQFVTNDGGKLTEIKRLYIQNGKVIESAVVNIEGIPPGNSITDDF 307 (386)
T ss_pred CCCeEEEEEeEeCCCCCcceeeEEEEECCEEEeCCCccccCCCCCCccCHHH
Confidence 34555678896 4589999999999998876442 3 3777764
No 23
>PF06831 H2TH: Formamidopyrimidine-DNA glycosylase H2TH domain; InterPro: IPR015886 This entry represents a helix-2turn-helix DNA-binding domain found in DNA glycosylase/AP lyase enzymes, which are involved in base excision repair of DNA damaged by oxidation or by mutagenic agents. Most damage to bases in DNA is repaired by the base excision repair pathway []. These enzymes are primarily from bacteria, and have both DNA glycosylase activity (3.2.2 from EC) and AP lyase activity (4.2.99.18 from EC). Examples include formamidopyrimidine-DNA glycosylases (Fpg; MutM) and endonuclease VIII (Nei). Formamidopyrimidine-DNA glycosylases (Fpg, MutM) is a trifunctional DNA base excision repair enzyme that removes a wide range of oxidation-damaged bases (N-glycosylase activity; 3.2.2.23 from EC) and cleaves both the 3'- and 5'-phosphodiester bonds of the resulting apurinic/apyrimidinic site (AP lyase activity; 4.2.99.18 from EC). Fpg has a preference for oxidised purines, excising oxidized purine bases such as 7,8-dihydro-8-oxoguanine (8-oxoG). ITs AP (apurinic/apyrimidinic) lyase activity introduces nicks in the DNA strand, cleaving the DNA backbone by beta-delta elimination to generate a single-strand break at the site of the removed base with both 3'- and 5'-phosphates. Fpg is a monomer composed of 2 domains connected by a flexible hinge []. The two DNA-binding motifs (a zinc finger and the helix-two-turns-helix motifs) suggest that the oxidized base is flipped out from double-stranded DNA in the binding mode and excised by a catalytic mechanism similar to that of bifunctional base excision repair enzymes []. Fpg binds one ion of zinc at the C terminus, which contains four conserved and essential cysteines [, ]. Endonuclease VIII (Nei) has the same enzyme activities as Fpg above (3.2.2 from EC, 4.2.99.18 from EC), but with a preference for oxidized pyrimidines, such as thymine glycol, 5,6-dihydrouracil and 5,6-dihydrothymine []. These protein contains three structural domains: an N-terminal catalytic core domain, a central helix-two turn-helix (H2TH) module and a C-terminal zinc finger []. The N-terminal catalytic domain and the C-terminal zinc finger straddle the DNA with the long axis of the protein oriented roughly orthogonal to the helical axis of the DNA. Residues that contact DNA are located in the catalytic domain and in a beta-hairpin loop formed by the zinc finger []. This entry represents the central domain containing the DNA-binding helix-two turn-helix domain [].; GO: 0003684 damaged DNA binding, 0003906 DNA-(apurinic or apyrimidinic site) lyase activity, 0008270 zinc ion binding, 0016799 hydrolase activity, hydrolyzing N-glycosyl compounds, 0006289 nucleotide-excision repair; PDB: 3GQ3_A 3JR5_A 3SAT_A 3GPX_A 2F5Q_A 3SBJ_A 3U6S_A 3SAU_A 3SAR_A 2F5P_A ....
Probab=31.73 E-value=47 Score=22.19 Aligned_cols=31 Identities=19% Similarity=0.389 Sum_probs=24.7
Q ss_pred CeecCcccceecCHHHHHHHHHhHHHH-HHHh
Q 033983 72 GKKFPGKKGISLSVDQWNTLRDHVEEI-NKAL 102 (106)
Q Consensus 72 Ge~~PgKKGISL~~eqw~~L~~~~~~I-d~ai 102 (106)
-..+|..+.-+|+.+||..|.+++..| ..||
T Consensus 51 a~i~P~~~~~~L~~~~~~~l~~~~~~vl~~ai 82 (92)
T PF06831_consen 51 AGIHPERPASSLSEEELRRLHEAIKRVLREAI 82 (92)
T ss_dssp TTB-TTSBGGGSHHHHHHHHHHHHHHHHHHHH
T ss_pred cCCCccCccccCCHHHHHHHHHHHHHHHHHHH
Confidence 568899999999999999998887765 4444
No 24
>PF02866 Ldh_1_C: lactate/malate dehydrogenase, alpha/beta C-terminal domain Prosite entry for lactate dehydrogenase Prosite entry for malate dehydrogenase; InterPro: IPR022383 L-lactate dehydrogenases are metabolic enzymes which catalyse the conversion of L-lactate to pyruvate, the last step in anaerobic glycolysis []. L-lactate dehydrogenase is also found as a lens crystallin in bird and crocodile eyes. L-2-hydroxyisocaproate dehydrogenases are also members of the family. Malate dehydrogenases catalyse the interconversion of malate to oxaloacetate []. The enzyme participates in the citric acid cycle. This entry represents the C-terminal, and is thought to be an is an unusual alpha+beta fold.; GO: 0016616 oxidoreductase activity, acting on the CH-OH group of donors, NAD or NADP as acceptor, 0055114 oxidation-reduction process; PDB: 4MDH_B 5MDH_A 1GV0_A 1GUZ_D 2EWD_B 2FRM_D 2FNZ_B 2FN7_B 2FM3_A 1LTH_T ....
Probab=31.12 E-value=65 Score=23.24 Aligned_cols=44 Identities=18% Similarity=0.308 Sum_probs=34.4
Q ss_pred eEEEeEEEEecCCeecCcccceecCHHHHHHHHHhHHHHHHHhhc
Q 033983 60 VWVDIREFYVKEGKKFPGKKGISLSVDQWNTLRDHVEEINKALGD 104 (106)
Q Consensus 60 ~~VdIREyY~kdGe~~PgKKGISL~~eqw~~L~~~~~~Id~ai~~ 104 (106)
+|+.+---..++|-+.-= .++.|++++.+.|.+++..|.+.+++
T Consensus 126 v~~s~P~~ig~~Gv~~i~-~~~~L~~~E~~~l~~sa~~l~~~i~~ 169 (174)
T PF02866_consen 126 VYFSVPVVIGKNGVEKIV-EDLPLSEEEQEKLKESAKELKKEIEK 169 (174)
T ss_dssp EEEEEEEEEETTEEEEEE-CSBSSTHHHHHHHHHHHHHHHHHHHH
T ss_pred ceecceEEEcCCeeEEEe-CCCCCCHHHHHHHHHHHHHHHHHHHH
Confidence 667777666678866651 24789999999999999999888764
No 25
>TIGR03342 dsrC_tusE_dsvC sulfur relay protein, TusE/DsrC/DsvC family. Members of this protein family may be described as TusE, a partner to TusBCD in a sulfur relay system for 2-thiouridine biosynthesis, a tRNA base modification process. Other members are DsrC, a functionally similar protein in species where the sulfur relay system exists primarily for sulfur metabolism rather than tRNA base modification. Some members of this family are known explicitly as the gamma subunit of sulfite reductases.
Probab=30.23 E-value=62 Score=22.86 Aligned_cols=17 Identities=24% Similarity=0.579 Sum_probs=13.6
Q ss_pred ccceecCHHHHHHHHHh
Q 033983 78 KKGISLSVDQWNTLRDH 94 (106)
Q Consensus 78 KKGISL~~eqw~~L~~~ 94 (106)
.-||.||.++|+.+.-.
T Consensus 33 ~egieLT~~Hw~vI~~l 49 (108)
T TIGR03342 33 EEGIELTEAHWEVINFL 49 (108)
T ss_pred HcCCCCCHHHHHHHHHH
Confidence 56999999999876543
No 26
>TIGR02675 tape_meas_nterm tape measure domain. Proteins containing this domain are strictly bacterial, including bacteriophage and prophage regions of bacterial genomes. Most members are 800 to 1800 amino acids long, making them among the longest predicted proteins of their respective phage genomes, where they are encoded in tail protein regions. This roughly 80-residue domain described here usually begins between residue 100 and 250. Many members are known or predicted to act as phage tail tape measure proteins, a minor tail component that regulates tail length.
Probab=30.03 E-value=66 Score=20.76 Aligned_cols=22 Identities=23% Similarity=0.270 Sum_probs=18.6
Q ss_pred cCHHHHHHHHHhHHHHHHHhhc
Q 033983 83 LSVDQWNTLRDHVEEINKALGD 104 (106)
Q Consensus 83 L~~eqw~~L~~~~~~Id~ai~~ 104 (106)
|+.++|+.|.+.++.+-.+|.+
T Consensus 31 v~~ee~n~~~e~~p~~~~~lAk 52 (75)
T TIGR02675 31 LRGEEINSLLEALPGALQALAK 52 (75)
T ss_pred ccHHHHHHHHHHhHHHHHHHHH
Confidence 7889999999999988777753
No 27
>PF09655 Nitr_red_assoc: Conserved nitrate reductase-associated protein (Nitr_red_assoc); InterPro: IPR013481 Proteins in this entry are found in the Cyanobacteria, and are mostly encoded near nitrate reductase and molybdopterin biosynthesis genes. Molybdopterin guanine dinucleotide is a cofactor for nitrate reductase. These proteins are sometimes annotated as nitrate reductase-associated proteins, though their function is unknown.
Probab=30.00 E-value=21 Score=26.66 Aligned_cols=17 Identities=29% Similarity=0.731 Sum_probs=14.1
Q ss_pred cccceecCHHHHHHHHH
Q 033983 77 GKKGISLSVDQWNTLRD 93 (106)
Q Consensus 77 gKKGISL~~eqw~~L~~ 93 (106)
...||.++++||..|-.
T Consensus 102 ~~~gv~~t~~qW~~L~p 118 (144)
T PF09655_consen 102 QEFGVPLTLEQWAALTP 118 (144)
T ss_pred HHcCCCCCHHHHhcCCH
Confidence 45799999999998854
No 28
>PF11580 DUF3239: Protein of unknown function (DUF3239); InterPro: IPR021632 This entry contains possible membrane proteins, however this cannot be confirmed. Currently they have no known function. ; PDB: 3C8I_B.
Probab=28.28 E-value=34 Score=24.96 Aligned_cols=19 Identities=11% Similarity=0.639 Sum_probs=15.0
Q ss_pred ecCHHHHHHHHHhHHHHHH
Q 033983 82 SLSVDQWNTLRDHVEEINK 100 (106)
Q Consensus 82 SL~~eqw~~L~~~~~~Id~ 100 (106)
+++.+||+.|..+++.|++
T Consensus 110 aIp~~eW~~L~~~~~r~~~ 128 (128)
T PF11580_consen 110 AIPQEEWRQLEKNLKRLEQ 128 (128)
T ss_dssp HS-HHHHHHHHHHGGGHHH
T ss_pred hCCHHHHHHHHHHHhhhcC
Confidence 5788999999999887763
No 29
>COG3530 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=27.37 E-value=30 Score=22.98 Aligned_cols=29 Identities=38% Similarity=0.673 Sum_probs=20.8
Q ss_pred EeeeCCceEEEeEEEEe----cCCeecCccccee
Q 033983 53 VRNWQGKVWVDIREFYV----KEGKKFPGKKGIS 82 (106)
Q Consensus 53 V~~FkG~~~VdIREyY~----kdGe~~PgKKGIS 82 (106)
...|+|+++||+-|-|. ..| .-||+-|.-
T Consensus 16 FGKYqGR~liDLPe~YLlWFarkg-FP~G~lG~L 48 (71)
T COG3530 16 FGKYQGRVLIDLPEEYLLWFARKG-FPPGKLGRL 48 (71)
T ss_pred cccccceeeecCCHHHHHHHHHhC-CCchHHHHH
Confidence 35799999999999664 345 556666543
No 30
>PF08743 Nse4_C: Nse4 C-terminal; InterPro: IPR014854 Nse4 is a component of the Smc5/6 DNA repair complex. It forms interactions with Smc5 and Nse1 [].
Probab=27.18 E-value=48 Score=22.16 Aligned_cols=16 Identities=25% Similarity=0.657 Sum_probs=13.8
Q ss_pred eecCHHHHHHHHHhHH
Q 033983 81 ISLSVDQWNTLRDHVE 96 (106)
Q Consensus 81 ISL~~eqw~~L~~~~~ 96 (106)
++|+.++|+.|.+...
T Consensus 69 ~~ld~~~W~~li~~~~ 84 (93)
T PF08743_consen 69 LSLDYEDWQELIEKYN 84 (93)
T ss_pred EEcCHHHHHHHHHHhC
Confidence 7999999999988653
No 31
>PRK11508 sulfur transfer protein TusE; Provisional
Probab=26.91 E-value=76 Score=22.46 Aligned_cols=18 Identities=39% Similarity=0.671 Sum_probs=13.9
Q ss_pred ccceecCHHHHHHHHHhH
Q 033983 78 KKGISLSVDQWNTLRDHV 95 (106)
Q Consensus 78 KKGISL~~eqw~~L~~~~ 95 (106)
.-||.||.++|+.+.-.-
T Consensus 34 ~egieLT~~HW~VI~~lR 51 (109)
T PRK11508 34 NEGISLSPEHWEVVRFVR 51 (109)
T ss_pred HhCCCCCHHHHHHHHHHH
Confidence 469999999998765433
No 32
>cd01277 HINT_subgroup HINT (histidine triad nucleotide-binding protein) subgroup: Members of this CD belong to the superfamily of histidine triad hydrolases that act on alpha-phosphate of ribonucleotides. This subgroup includes members from all three forms of cellular life. Although the biochemical function has not been characterised for many of the members of this subgroup, the proteins from Yeast have been shown to be involved in secretion, peroxisome formation and gene expression.
Probab=26.55 E-value=90 Score=19.97 Aligned_cols=24 Identities=17% Similarity=0.235 Sum_probs=20.6
Q ss_pred ecCHHHHHHHHHhHHHHHHHhhcC
Q 033983 82 SLSVDQWNTLRDHVEEINKALGDN 105 (106)
Q Consensus 82 SL~~eqw~~L~~~~~~Id~ai~~~ 105 (106)
.|+.++|..|...+..+..++.++
T Consensus 50 ~l~~~e~~~l~~~~~~v~~~l~~~ 73 (103)
T cd01277 50 DLDPEELAELILAAKKVARALKKA 73 (103)
T ss_pred hCCHHHHHHHHHHHHHHHHHHHHh
Confidence 489999999999999998888753
No 33
>PF03102 NeuB: NeuB family; InterPro: IPR013132 NeuB is the prokaryotic N-acetylneuraminic acid synthase (Neu5Ac). It catalyses the direct formation of Neu5Ac (the most common sialic acid) by condensation of phosphoenolpyruvate (PEP) and N-acetylmannosamine (ManNAc). This reaction has only been observed in prokaryotes; eukaryotes synthesise the 9-phosphate form, Neu5Ac-9-P, and utilise ManNAc-6-P instead of ManNAc. Such eukaryotic enzymes are not present in this family []. This family also contains SpsE spore coat polysaccharide biosynthesis proteins.; GO: 0016051 carbohydrate biosynthetic process; PDB: 3G8R_B 1XUU_A 1XUZ_A 3CM4_A 2ZDR_A 1VLI_A 2WQP_A.
Probab=26.31 E-value=66 Score=25.35 Aligned_cols=25 Identities=36% Similarity=0.547 Sum_probs=21.9
Q ss_pred cceecCHHHHHHHHHhHHHHHHHhh
Q 033983 79 KGISLSVDQWNTLRDHVEEINKALG 103 (106)
Q Consensus 79 KGISL~~eqw~~L~~~~~~Id~ai~ 103 (106)
-..||.|+|+..|.+.+..+..|+.
T Consensus 213 h~~Sl~p~el~~lv~~ir~~~~alG 237 (241)
T PF03102_consen 213 HKFSLEPDELKQLVRDIREVEKALG 237 (241)
T ss_dssp GCCCB-HHHHHHHHHHHHHHHHHCS
T ss_pred hhhcCCHHHHHHHHHHHHHHHHHcC
Confidence 4689999999999999999999885
No 34
>PF11006 DUF2845: Protein of unknown function (DUF2845); InterPro: IPR021268 This bacterial family of proteins has no known function.
Probab=26.28 E-value=1.2e+02 Score=19.96 Aligned_cols=28 Identities=11% Similarity=0.137 Sum_probs=24.1
Q ss_pred CcEEEEcCCceEEEEeeeCCceEEEeEE
Q 033983 39 DIVVCEISKNRRVSVRNWQGKVWVDIRE 66 (106)
Q Consensus 39 ~~~~~~Ls~~rrVtV~~FkG~~~VdIRE 66 (106)
+.++.+.+.++...+-.|.|-.++.|+.
T Consensus 58 E~W~Yn~Gp~~~~~~l~f~~Gkl~~I~~ 85 (87)
T PF11006_consen 58 EEWTYNFGPNGFMQILTFENGKLVRIES 85 (87)
T ss_pred eEEEEeCCCCCcEEEEEEECCEEEEEEe
Confidence 3567778999999999999999999973
No 35
>TIGR03586 PseI pseudaminic acid synthase.
Probab=26.05 E-value=74 Score=26.23 Aligned_cols=24 Identities=33% Similarity=0.471 Sum_probs=22.0
Q ss_pred eecCHHHHHHHHHhHHHHHHHhhc
Q 033983 81 ISLSVDQWNTLRDHVEEINKALGD 104 (106)
Q Consensus 81 ISL~~eqw~~L~~~~~~Id~ai~~ 104 (106)
.||+|+|+..|+..+..|..++..
T Consensus 236 ~Sl~p~e~~~lv~~ir~~~~~lg~ 259 (327)
T TIGR03586 236 FSLEPDEFKALVKEVRNAWLALGE 259 (327)
T ss_pred ccCCHHHHHHHHHHHHHHHHHhCC
Confidence 789999999999999999998864
No 36
>PF08848 DUF1818: Domain of unknown function (DUF1818); InterPro: IPR014947 This entry represents a small family of uncharacterised cyanobacterial proteins. ; PDB: 2IT9_A 2NVN_A.
Probab=26.05 E-value=1e+02 Score=22.28 Aligned_cols=25 Identities=12% Similarity=0.290 Sum_probs=21.3
Q ss_pred eecCHHHHHHHHHhHHHHHHHhhcC
Q 033983 81 ISLSVDQWNTLRDHVEEINKALGDN 105 (106)
Q Consensus 81 ISL~~eqw~~L~~~~~~Id~ai~~~ 105 (106)
|-||..+|+.|...+..+.+.+..+
T Consensus 29 iELT~~E~~~f~~Ll~~L~~q~~~i 53 (117)
T PF08848_consen 29 IELTEAEFNDFCRLLQQLAEQMQAI 53 (117)
T ss_dssp EEE-HHHHHHHHHHHHHHHHHHHCC
T ss_pred eeecHHHHHHHHHHHHHHHHHHHHH
Confidence 8899999999999999998887654
No 37
>PF00840 Glyco_hydro_7: Glycosyl hydrolase family 7; InterPro: IPR001722 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. Glycoside hydrolase family 7 GH7 from CAZY comprises enzymes with several known activities; endoglucanase (3.2.1.4 from EC); cellobiohydrolase (3.2.1.91 from EC). These enzymes were formerly known as cellulase family C. Exoglucanases and cellobiohydrolases [] play a role in the conversion of cellulose to glucose by cutting the dissaccharide cellobiose from the nonreducing end of the cellulose polymer chain. Structurally, cellulases and xylanases generally consist of a catalytic domain joined to a cellulose-binding domain (CBD) via a linker region that is rich in proline and/or hydroxy-amino acids. In type I exoglucanases, the CBD domain is found at the C-terminal extremity of these enzyme (this short domain forms a hairpin loop structure stabilised by 2 disulphide bridges).; GO: 0004553 hydrolase activity, hydrolyzing O-glycosyl compounds, 0005975 carbohydrate metabolic process; PDB: 2Y9N_A 2Y9L_A 2RFW_D 2RFZ_A 2RFY_D 2RG0_C 1EG1_C 1OVW_D 2OVW_A 4OVW_A ....
Probab=25.28 E-value=90 Score=27.23 Aligned_cols=47 Identities=21% Similarity=0.347 Sum_probs=30.2
Q ss_pred CceEEEEeeeCCc-----eEEEeEEEEecCCeecCccc----c----eecCHHHHHHHHH
Q 033983 47 KNRRVSVRNWQGK-----VWVDIREFYVKEGKKFPGKK----G----ISLSVDQWNTLRD 93 (106)
Q Consensus 47 ~~rrVtV~~FkG~-----~~VdIREyY~kdGe~~PgKK----G----ISL~~eqw~~L~~ 93 (106)
.+++-.|..|-.. .|+-||-||..+|+..+..+ | =||+.+-=.+-+.
T Consensus 280 tkkfTVVTQFit~~~t~G~L~EIrR~YVQnGkvI~n~~~~~~g~~~~nsItd~fC~~~~~ 339 (433)
T PF00840_consen 280 TKKFTVVTQFITDDGTTGDLSEIRRLYVQNGKVIQNPKVNIPGLPGFNSITDEFCSAQKS 339 (433)
T ss_dssp TSEEEEEEEEEETTSSTS-EEEEEEEEEETTEEEESSSEESTTSESSSSBSHHHHHHHHH
T ss_pred CCccEEEEEeecCCCCccccceeeEEEEECCEEEeCCCcccCCCCCCCccCHHHHhhhcc
Confidence 4455557778654 49999999999998775432 2 2577664444444
No 38
>COG2920 DsrC Dissimilatory sulfite reductase (desulfoviridin), gamma subunit [Inorganic ion transport and metabolism]
Probab=24.44 E-value=50 Score=23.80 Aligned_cols=20 Identities=25% Similarity=0.768 Sum_probs=15.7
Q ss_pred eecCcccceecCHHHHHHHH
Q 033983 73 KKFPGKKGISLSVDQWNTLR 92 (106)
Q Consensus 73 e~~PgKKGISL~~eqw~~L~ 92 (106)
+++--.-||.||.++|+.++
T Consensus 31 e~lA~~e~i~LT~eHWevv~ 50 (111)
T COG2920 31 EALAEREGIELTEEHWEVVR 50 (111)
T ss_pred HHHHHHhccCccHHHHHHHH
Confidence 44555679999999998764
No 39
>cd02679 MIT_spastin MIT: domain contained within Microtubule Interacting and Trafficking molecules. This MIT domain sub-family is found in the AAA protein spastin, a probable ATPase involved in the assembly or function of nuclear protein complexes; spastins might also be involved in microtubule dynamics. The molecular function of the MIT domain is unclear.
Probab=23.98 E-value=53 Score=21.88 Aligned_cols=31 Identities=19% Similarity=0.258 Sum_probs=22.2
Q ss_pred CeecCcccceecCHHHHHHHHHhHHHHHHHhhc
Q 033983 72 GKKFPGKKGISLSVDQWNTLRDHVEEINKALGD 104 (106)
Q Consensus 72 Ge~~PgKKGISL~~eqw~~L~~~~~~Id~ai~~ 104 (106)
|--.|.. +.-+-+||+..+.....+..++.+
T Consensus 41 g~ai~~~--~~~~~~~w~~ar~~~~Km~~~~~~ 71 (79)
T cd02679 41 GIAVPVP--SAGVGSQWERARRLQQKMKTNLNM 71 (79)
T ss_pred HcCCCCC--cccccHHHHHHHHHHHHHHHHHHH
Confidence 4444542 455669999999999888887754
No 40
>PRK13398 3-deoxy-7-phosphoheptulonate synthase; Provisional
Probab=22.95 E-value=81 Score=25.10 Aligned_cols=24 Identities=25% Similarity=0.400 Sum_probs=20.9
Q ss_pred cceecCHHHHHHHHHhHHHHHHHh
Q 033983 79 KGISLSVDQWNTLRDHVEEINKAL 102 (106)
Q Consensus 79 KGISL~~eqw~~L~~~~~~Id~ai 102 (106)
--.||+++++..|.+.+..|.+++
T Consensus 243 ~~~sl~p~~l~~l~~~i~~~~~~~ 266 (266)
T PRK13398 243 ARQTLNFEEMKELVDELKPMAKAL 266 (266)
T ss_pred hhhcCCHHHHHHHHHHHHHHHhhC
Confidence 458899999999999999988764
No 41
>PF11325 DUF3127: Domain of unknown function (DUF3127); InterPro: IPR021474 This bacterial family of proteins has no known function.
Probab=22.49 E-value=88 Score=21.22 Aligned_cols=17 Identities=29% Similarity=0.812 Sum_probs=14.3
Q ss_pred EEEeeeCCceEEEeEEE
Q 033983 51 VSVRNWQGKVWVDIREF 67 (106)
Q Consensus 51 VtV~~FkG~~~VdIREy 67 (106)
+.-|+|.|+-|.|||-|
T Consensus 65 i~~RE~~gr~fn~i~aW 81 (84)
T PF11325_consen 65 IEGREWNGRWFNSIRAW 81 (84)
T ss_pred eeccEecceEeeEeEEE
Confidence 34589999999999986
No 42
>PRK06223 malate dehydrogenase; Reviewed
Probab=22.42 E-value=1.1e+02 Score=23.76 Aligned_cols=45 Identities=18% Similarity=0.192 Sum_probs=33.0
Q ss_pred eEEEeEEEEecCCeecCcccceecCHHHHHHHHHhHHHHHHHhhcCC
Q 033983 60 VWVDIREFYVKEGKKFPGKKGISLSVDQWNTLRDHVEEINKALGDNS 106 (106)
Q Consensus 60 ~~VdIREyY~kdGe~~PgKKGISL~~eqw~~L~~~~~~Id~ai~~~~ 106 (106)
.++-+--...++|-..- -.+.|+.++.+.|.+....|.+.+++++
T Consensus 263 ~~~s~P~~i~~~Gv~~i--~~~~l~~~e~~~l~~s~~~l~~~~~~~~ 307 (307)
T PRK06223 263 VYVGVPVKLGKNGVEKI--IELELDDEEKAAFDKSVEAVKKLIEALK 307 (307)
T ss_pred eEEEeEEEEeCCeEEEE--eCCCCCHHHHHHHHHHHHHHHHHHHhcC
Confidence 45555555555554333 2478999999999999999999988763
No 43
>PF14164 YqzH: YqzH-like protein
Probab=22.27 E-value=96 Score=20.27 Aligned_cols=18 Identities=33% Similarity=0.739 Sum_probs=14.9
Q ss_pred eecCHHHHHHHHHhHHHH
Q 033983 81 ISLSVDQWNTLRDHVEEI 98 (106)
Q Consensus 81 ISL~~eqw~~L~~~~~~I 98 (106)
+.|+.++|+.|.+.+..+
T Consensus 24 ~pls~~E~~~L~~~i~~~ 41 (64)
T PF14164_consen 24 MPLSDEEWEELCKHIQER 41 (64)
T ss_pred CCCCHHHHHHHHHHHHHH
Confidence 449999999999887664
No 44
>TIGR03569 NeuB_NnaB N-acetylneuraminate synthase. This family is a subset of the Pfam model pfam03102 and is believed to include only authentic NeuB N-acetylneuraminate (sialic acid) synthase enzymes. The majority of the genes identified by this model are observed adjacent to both the NeuA and NeuC genes which together effect the biosynthesis of CMP-N-acetylneuraminate from UDP-N-acetylglucosamine.
Probab=21.85 E-value=97 Score=25.57 Aligned_cols=25 Identities=32% Similarity=0.517 Sum_probs=22.5
Q ss_pred ceecCHHHHHHHHHhHHHHHHHhhc
Q 033983 80 GISLSVDQWNTLRDHVEEINKALGD 104 (106)
Q Consensus 80 GISL~~eqw~~L~~~~~~Id~ai~~ 104 (106)
-.||+++|+..|.+.+..+..++..
T Consensus 236 ~~Sl~p~el~~lv~~ir~~~~~lG~ 260 (329)
T TIGR03569 236 KASLEPDELKEMVQGIRNVEKALGD 260 (329)
T ss_pred hhcCCHHHHHHHHHHHHHHHHHcCC
Confidence 5899999999999999999998863
No 45
>PRK08673 3-deoxy-7-phosphoheptulonate synthase; Reviewed
Probab=21.66 E-value=1e+02 Score=25.57 Aligned_cols=25 Identities=28% Similarity=0.460 Sum_probs=22.1
Q ss_pred ceecCHHHHHHHHHhHHHHHHHhhc
Q 033983 80 GISLSVDQWNTLRDHVEEINKALGD 104 (106)
Q Consensus 80 GISL~~eqw~~L~~~~~~Id~ai~~ 104 (106)
-.||+++++..|++.+..|.+++.+
T Consensus 310 ~~sl~p~e~~~lv~~i~~i~~~~g~ 334 (335)
T PRK08673 310 PQSLTPEEFEELMKKLRAIAEALGR 334 (335)
T ss_pred hhcCCHHHHHHHHHHHHHHHHHhCC
Confidence 4789999999999999999998864
No 46
>PRK13396 3-deoxy-7-phosphoheptulonate synthase; Provisional
Probab=21.33 E-value=1.1e+02 Score=25.82 Aligned_cols=25 Identities=28% Similarity=0.454 Sum_probs=21.9
Q ss_pred ceecCHHHHHHHHHhHHHHHHHhhc
Q 033983 80 GISLSVDQWNTLRDHVEEINKALGD 104 (106)
Q Consensus 80 GISL~~eqw~~L~~~~~~Id~ai~~ 104 (106)
--||+++++..|.+.+..|..++.+
T Consensus 319 ~qsl~p~~~~~l~~~i~~i~~~~g~ 343 (352)
T PRK13396 319 PQSLTPDRFDRLMQELAVIGKTVGR 343 (352)
T ss_pred hhcCCHHHHHHHHHHHHHHHHHhCC
Confidence 3679999999999999999998864
No 47
>cd01275 FHIT FHIT (fragile histidine family): FHIT proteins, related to the HIT family carry a motif HxHxH/Qxx (x, is a hydrophobic amino acid), On the basis of sequence, substrate specificity, structure, evolution and mechanism, HIT proteins are classified into three branches: the Hint branch, which consists of adenosine 5' -monophosphoramide hydrolases, the Fhit branch, that consists of diadenosine polyphosphate hydrolases, and the GalT branch consisting of specific nucloside monophosphate transferases. Fhit plays a very important role in the development of tumours. Infact, Fhit deletions are among the earliest and most frequent genetic alterations in the development of tumours.
Probab=21.14 E-value=1.2e+02 Score=20.49 Aligned_cols=40 Identities=23% Similarity=0.052 Sum_probs=29.4
Q ss_pred eCCceEEEeEEEEecCCeecCcccceecCHHHHHHHHHhHHHHHHHhhc
Q 033983 56 WQGKVWVDIREFYVKEGKKFPGKKGISLSVDQWNTLRDHVEEINKALGD 104 (106)
Q Consensus 56 FkG~~~VdIREyY~kdGe~~PgKKGISL~~eqw~~L~~~~~~Id~ai~~ 104 (106)
+.|.++|=-|+.+.. =..|++++|..|...+..+..+|++
T Consensus 33 ~~gh~lIiPk~H~~~---------~~~L~~~e~~~l~~~~~~v~~~l~~ 72 (126)
T cd01275 33 NPGHVLVVPYRHVPR---------LEDLTPEEIADLFKLVQLAMKALKV 72 (126)
T ss_pred CCCcEEEEeccccCC---------hhhCCHHHHHHHHHHHHHHHHHHHH
Confidence 456666666655432 2348999999999999988888875
No 48
>COG2089 SpsE Sialic acid synthase [Cell envelope biogenesis, outer membrane]
Probab=20.82 E-value=1.1e+02 Score=25.97 Aligned_cols=25 Identities=32% Similarity=0.638 Sum_probs=22.7
Q ss_pred cceecCHHHHHHHHHhHHHHHHHhh
Q 033983 79 KGISLSVDQWNTLRDHVEEINKALG 103 (106)
Q Consensus 79 KGISL~~eqw~~L~~~~~~Id~ai~ 103 (106)
--+||.|++|..|++++.++..||.
T Consensus 247 ~~fSldP~efk~mv~~ir~~~~alG 271 (347)
T COG2089 247 HAFSLDPDEFKEMVDAIRQVEKALG 271 (347)
T ss_pred cceecCHHHHHHHHHHHHHHHHHhC
Confidence 4589999999999999999999885
No 49
>PRK10287 thiosulfate:cyanide sulfurtransferase; Provisional
Probab=20.42 E-value=64 Score=21.89 Aligned_cols=34 Identities=15% Similarity=0.378 Sum_probs=23.1
Q ss_pred eeCCceEEEeEEEEecCCeecCcccceecCHHHHHH
Q 033983 55 NWQGKVWVDIREFYVKEGKKFPGKKGISLSVDQWNT 90 (106)
Q Consensus 55 ~FkG~~~VdIREyY~kdGe~~PgKKGISL~~eqw~~ 90 (106)
.|.-..+||||+-=+=.+.-.|| -|+++..++..
T Consensus 17 ~~~~~~lIDvR~~~ef~~ghIpG--AiniP~~~l~~ 50 (104)
T PRK10287 17 VFAAEHWIDVRVPEQYQQEHVQG--AINIPLKEVKE 50 (104)
T ss_pred ccCCCEEEECCCHHHHhcCCCCc--cEECCHHHHHH
Confidence 38889999999932213456787 47888666543
Done!