Query 012348
Match_columns 465
No_of_seqs 133 out of 372
Neff 3.6
Searched_HMMs 46136
Date Fri Mar 29 01:41:45 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/012348.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/012348hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF02362 B3: B3 DNA binding do 99.7 8.2E-17 1.8E-21 132.0 11.1 98 359-460 1-99 (100)
2 PF03754 DUF313: Domain of unk 98.2 3.3E-06 7E-11 74.9 7.0 80 353-432 18-114 (114)
3 PF09217 EcoRII-N: Restriction 97.7 0.00011 2.3E-09 68.6 7.8 89 356-444 7-110 (156)
4 smart00249 PHD PHD zinc finger 69.7 2.2 4.8E-05 29.7 0.9 25 77-101 10-34 (47)
5 PF10844 DUF2577: Protein of u 66.2 17 0.00037 31.4 5.8 82 354-458 16-99 (100)
6 PRK03760 hypothetical protein; 39.3 59 0.0013 29.1 4.8 29 417-445 88-116 (117)
7 PF13248 zf-ribbon_3: zinc-rib 38.0 15 0.00033 24.6 0.8 15 77-91 12-26 (26)
8 PF08922 DUF1905: Domain of un 36.6 1E+02 0.0022 25.6 5.6 79 359-444 1-79 (80)
9 PF02643 DUF192: Uncharacteriz 36.2 82 0.0018 27.4 5.2 52 393-444 49-107 (108)
10 KOG4718 Non-SMC (structural ma 33.4 16 0.00034 36.7 0.3 18 82-99 195-212 (235)
11 PF04014 Antitoxin-MazE: Antid 30.5 57 0.0012 24.2 2.8 23 427-449 13-35 (47)
12 TIGR01643 YD_repeat_2x YD repe 30.2 76 0.0016 22.4 3.3 22 392-413 4-25 (42)
13 COG1998 RPS31 Ribosomal protei 29.5 27 0.00058 27.8 1.0 13 82-94 20-41 (51)
14 PF03120 DNA_ligase_OB: NAD-de 29.4 41 0.00088 28.8 2.1 22 427-448 42-63 (82)
15 PF09297 zf-NADH-PPase: NADH p 27.6 31 0.00067 24.0 0.9 25 58-91 5-31 (32)
16 cd05829 Sortase_E Sortase E (S 27.2 1.2E+02 0.0025 27.6 4.8 39 418-456 49-94 (144)
17 PF12760 Zn_Tnp_IS1595: Transp 24.5 41 0.00089 25.1 1.2 26 58-90 20-46 (46)
18 cd04459 Rho_CSD Rho_CSD: Rho p 23.4 71 0.0015 26.4 2.4 26 428-453 34-60 (68)
19 TIGR03784 marine_sortase sorta 23.3 84 0.0018 29.9 3.2 81 371-459 51-133 (174)
20 TIGR00375 conserved hypothetic 22.8 36 0.00079 36.2 0.8 44 54-107 238-282 (374)
21 PRK09570 rpoH DNA-directed RNA 22.3 1E+02 0.0022 26.4 3.3 32 427-459 44-75 (79)
22 PF01191 RNA_pol_Rpb5_C: RNA p 22.3 1.2E+02 0.0026 25.6 3.6 32 427-459 41-72 (74)
23 PF05593 RHS_repeat: RHS Repea 22.2 1.4E+02 0.003 21.3 3.4 23 391-413 3-25 (38)
No 1
>PF02362 B3: B3 DNA binding domain; InterPro: IPR003340 Two DNA binding proteins, RAV1 and RAV2 from Arabidopsis thaliana contain two distinct amino acid sequence domains found only in higher plant species. The N-terminal regions of RAV1 and RAV2 are homologous to the AP2 DNA-binding domain (see IPR001471 from INTERPRO) present in a family of transcription factors, while the C-terminal region exhibits homology to the highly conserved C-terminal domain, designated B3, of VP1/ABI3 transcription factors []. The AP2 and B3-like domains of RAV1 bind autonomously to the CAACA and CACCTG motifs, respectively, and together achieve a high affinity and specificity of binding. It has been suggested that the AP2 and B3-like domains of RAV1 are connected by a highly flexible structure enabling the two domains to bind to the CAACA and CACCTG motifs in various spacings and orientations [].; GO: 0003677 DNA binding, 0006355 regulation of transcription, DNA-dependent; PDB: 1WID_A 1YEL_A.
Probab=99.71 E-value=8.2e-17 Score=131.96 Aligned_cols=98 Identities=22% Similarity=0.325 Sum_probs=71.9
Q ss_pred EEEecccccCCCCCcEEeehhhhhhcCCCCCCCCCceEEEEeCCCCeEEeEEEEcCCCCCCceeec-CchhHhhhcCCCC
Q 012348 359 FEKMLSASDAGRIGRLVLPKKCAEAYFPPISQPEGLPLKVQDSKGKEWIFQFRFWPNNNSRMYVLE-GVTPCIQNMQLQA 437 (465)
Q Consensus 359 F~KvLT~SDVg~lgRLVIPK~~AEa~FP~L~~~~Gv~L~veD~~Gk~W~Frfrfw~NnsSR~YVLt-GW~~FVrsK~Lqa 437 (465)
|.|+|+++|+.+..+|+||++.+++|. .....++.|.++|..|++|.+++.++ +.++.|+|+ ||.+||++++|++
T Consensus 1 F~K~l~~s~~~~~~~l~iP~~f~~~~~--~~~~~~~~v~l~~~~g~~W~v~~~~~--~~~~~~~l~~GW~~Fv~~n~L~~ 76 (100)
T PF02362_consen 1 FFKVLKPSDVSSSCRLIIPKEFAKKHG--GNKRKSREVTLKDPDGRSWPVKLKYR--KNSGRYYLTGGWKKFVRDNGLKE 76 (100)
T ss_dssp EEEE--TTCCCCTT-EEE-HHHHTTTS----SS--CEEEEEETTTEEEEEEEEEE--CCTTEEEEETTHHHHHHHCT--T
T ss_pred CEEEEEccCcCCCCEEEeCHHHHHHhC--CCcCCCeEEEEEeCCCCEEEEEEEEE--ccCCeEEECCCHHHHHHHcCCCC
Confidence 899999999998889999999999882 22235789999999999999999987 445568886 9999999999999
Q ss_pred CCEEEEeecCCCCceEEEEEeee
Q 012348 438 GDIGNQPKSQESYEMFSFMWKVH 460 (465)
Q Consensus 438 GDtVtF~R~~ngg~rf~I~~rr~ 460 (465)
||.|+|+...++...+.+.+.++
T Consensus 77 GD~~~F~~~~~~~~~~~v~i~~~ 99 (100)
T PF02362_consen 77 GDVCVFELIGNSNFTLKVHIFRK 99 (100)
T ss_dssp T-EEEEEE-SSSCE-EEEEEE--
T ss_pred CCEEEEEEecCCCceEEEEEEEC
Confidence 99999999876555567777654
No 2
>PF03754 DUF313: Domain of unknown function (DUF313) ; InterPro: IPR005508 This is a family of proteins from Arabidopsis thaliana (Mouse-ear cress) with uncharacterised function.
Probab=98.22 E-value=3.3e-06 Score=74.89 Aligned_cols=80 Identities=21% Similarity=0.433 Sum_probs=63.9
Q ss_pred CcccceEEEecccccCC-CCCcEEeehhhhhh--cCCC-----C-------CCCCCceEEEEeCCCCeEEeEEEEcCC-C
Q 012348 353 SVITPLFEKMLSASDAG-RIGRLVLPKKCAEA--YFPP-----I-------SQPEGLPLKVQDSKGKEWIFQFRFWPN-N 416 (465)
Q Consensus 353 s~~~~LF~KvLT~SDVg-~lgRLVIPK~~AEa--~FP~-----L-------~~~~Gv~L~veD~~Gk~W~Frfrfw~N-n 416 (465)
.....+++|+|++|||. ..+||.||...... +|=+ + ....|+.+.+.|..+++|..+++.|.- +
T Consensus 18 ~d~kli~~K~L~~tDv~~~qsRLsmP~~qi~~~dFLt~eE~~~i~~~~~~~~~~~Gv~V~lvdp~~~~~~m~lkkW~mg~ 97 (114)
T PF03754_consen 18 EDPKLIIEKTLFKTDVDPHQSRLSMPFNQIIDNDFLTEEEKRIIKEEKKNNDKKKGVEVILVDPSLRKWTMRLKKWNMGN 97 (114)
T ss_pred CCCeEEEeeeecccCCCCCCceeeccHHHhcccccCCHHHHHHHHHhhccCcccCCceEEEECCcCcEEEEEEEEecccC
Confidence 45578999999999999 48899999987633 2211 2 235689999999999999999999965 4
Q ss_pred CCCceeec-CchhHhhh
Q 012348 417 NSRMYVLE-GVTPCIQN 432 (465)
Q Consensus 417 sSR~YVLt-GW~~FVrs 432 (465)
..-.|+|. ||.+.|.+
T Consensus 98 ~~~~YvL~~gWn~VV~~ 114 (114)
T PF03754_consen 98 GTSNYVLNSGWNKVVED 114 (114)
T ss_pred CceEEEEEcChHhhccC
Confidence 45689996 99999863
No 3
>PF09217 EcoRII-N: Restriction endonuclease EcoRII, N-terminal; InterPro: IPR023372 There are four classes of restriction endonucleases: types I, II,III and IV. All types of enzymes recognise specific short DNA sequences and carry out the endonucleolytic cleavage of DNA to give specific double-stranded fragments with terminal 5'-phosphates. They differ in their recognition sequence, subunit composition, cleavage position, and cofactor requirements [, ], as summarised below: Type I enzymes (3.1.21.3 from EC) cleave at sites remote from recognition site; require both ATP and S-adenosyl-L-methionine to function; multifunctional protein with both restriction and methylase (2.1.1.72 from EC) activities. Type II enzymes (3.1.21.4 from EC) cleave within or at short specific distances from recognition site; most require magnesium; single function (restriction) enzymes independent of methylase. Type III enzymes (3.1.21.5 from EC) cleave at sites a short distance from recognition site; require ATP (but doesn't hydrolyse it); S-adenosyl-L-methionine stimulates reaction but is not required; exists as part of a complex with a modification methylase methylase (2.1.1.72 from EC). Type IV enzymes target methylated DNA. Type II restriction endonucleases (3.1.21.4 from EC) are components of prokaryotic DNA restriction-modification mechanisms that protect the organism against invading foreign DNA. These site-specific deoxyribonucleases catalyse the endonucleolytic cleavage of DNA to give specific double-stranded fragments with terminal 5'-phosphates. Of the 3000 restriction endonucleases that have been characterised, most are homodimeric or tetrameric enzymes that cleave target DNA at sequence-specific sites close to the recognition site. For homodimeric enzymes, the recognition site is usually a palindromic sequence 4-8 bp in length. Most enzymes require magnesium ions as a cofactor for catalysis. Although they can vary in their mode of recognition, many restriction endonucleases share a similar structural core comprising four beta-strands and one alpha-helix, as well as a similar mechanism of cleavage, suggesting a common ancestral origin []. However, there is still considerable diversity amongst restriction endonucleases [, ]. The target site recognition process triggers large conformational changes of the enzyme and the target DNA, leading to the activation of the catalytic centres. Like other DNA binding proteins, restriction enzymes are capable of non-specific DNA binding as well, which is the prerequisite for efficient target site location by facilitated diffusion. Non-specific binding usually does not involve interactions with the bases but only with the DNA backbone []. This entry represents the N-terminal effector-binding domain of the type II restriction endonuclease EcoRII, which has a DNA recognition fold, allowing for binding to 5'-CCWGG sequences. It assumes a structure composed of an eight-stranded beta-sheet with the strands in the order of b2, b5, b4, b3, b7, b6, b1 and b8. They are mostly antiparallel to each other except that b3 is parallel to b7. Alternatively, it may also be viewed as consisting of two mini beta-sheets of four antiparallel beta-strands, sheet I from beta-strands b2, b5, b4, b3 and sheet II from strands b7, b6, b1, b8, folded into an open mixed beta-barrel with a novel topology. Sheet I has a simple Greek key motif while sheet II does not []. The domain represented by this entry is only found in bacterial proteins.; PDB: 3HQF_A 1NA6_A.
Probab=97.73 E-value=0.00011 Score=68.63 Aligned_cols=89 Identities=21% Similarity=0.301 Sum_probs=56.5
Q ss_pred cceEEEecccccCCC----CCcEEeehhhhhhcCCCCCCCCC----ceEEEEeCCC--CeEEeEEEEcCC----CCCCce
Q 012348 356 TPLFEKMLSASDAGR----IGRLVLPKKCAEAYFPPISQPEG----LPLKVQDSKG--KEWIFQFRFWPN----NNSRMY 421 (465)
Q Consensus 356 ~~LF~KvLT~SDVg~----lgRLVIPK~~AEa~FP~L~~~~G----v~L~veD~~G--k~W~Frfrfw~N----nsSR~Y 421 (465)
...|.|.||+.|++. ..+++|||..++.+||.+...++ +.|.+.+..+ ..|.+||+|..| +.+..|
T Consensus 7 ~~~~~K~LSaNDtGaTGgHQaGiyIpk~~~~~lFp~~~~~~~~Np~~~~~~~~~s~~~~~~~~r~iYYnn~~~~gTRNE~ 86 (156)
T PF09217_consen 7 WAIYCKRLSANDTGATGGHQAGIYIPKSAAELLFPSINHTKEENPDIWLKARWQSHFVTDSQVRFIYYNNRLFGGTRNEY 86 (156)
T ss_dssp EEEEEEE--CCCCTTTSSS--EEEE-HHHHHHH-GGG-SSSSSS-EEEEEEEETTTT---EEEEEEEE-CCCTTSS--EE
T ss_pred eEEEEEEccCCCCCCcCcccceeEecccHHHHhCCCCCcccccCCceeEEEEECCCCccceeEEEEEEcccccCCCcCce
Confidence 468999999999994 34799999999999988765433 7788888866 678899999955 346779
Q ss_pred eecCchhHhhhcC-CCCCCEEEEe
Q 012348 422 VLEGVTPCIQNMQ-LQAGDIGNQP 444 (465)
Q Consensus 422 VLtGW~~FVrsK~-LqaGDtVtF~ 444 (465)
-||+|+....-.+ =.+||.++|-
T Consensus 87 RIT~~G~~~~~~~~~~tGaL~vla 110 (156)
T PF09217_consen 87 RITRFGRGFPLQNPENTGALLVLA 110 (156)
T ss_dssp EEE---TTSGGG-GGGTT-EEEEE
T ss_pred EEeeecCCCccCCccccccEEEEE
Confidence 9999986665333 3689987765
No 4
>smart00249 PHD PHD zinc finger. The plant homeodomain (PHD) finger is a C4HC3 zinc-finger-like motif found in nuclear proteins thought to be involved in epigenetics and chromatin-mediated transcriptional regulation. The PHD finger binds two zinc ions using the so-called 'cross-brace' motif and is thus structurally related to the PF10844 DUF2577: Protein of unknown function (DUF2577); InterPro: IPR022555 This family of proteins has no known function
Probab=66.22 E-value=17 Score=31.36 Aligned_cols=82 Identities=12% Similarity=0.140 Sum_probs=50.1
Q ss_pred cccceEEEecccccCC-C-CCcEEeehhhhhhcCCCCCCCCCceEEEEeCCCCeEEeEEEEcCCCCCCceeecCchhHhh
Q 012348 354 VITPLFEKMLSASDAG-R-IGRLVLPKKCAEAYFPPISQPEGLPLKVQDSKGKEWIFQFRFWPNNNSRMYVLEGVTPCIQ 431 (465)
Q Consensus 354 ~~~~LF~KvLT~SDVg-~-lgRLVIPK~~AEa~FP~L~~~~Gv~L~veD~~Gk~W~Frfrfw~NnsSR~YVLtGW~~FVr 431 (465)
.....|-++++.+-+. + .++++||++.. +++..-......+.+....+.. .. .+.-
T Consensus 16 p~~i~~G~V~s~~PL~I~i~~~liL~~~~L--~i~~~l~~~~~~~~~~~~~~~~------------~~--------~i~~ 73 (100)
T PF10844_consen 16 PVDIVIGTVVSVPPLKIKIDQKLILDKDFL--IIPELLKDYTRDITIEHNSETD------------NI--------TITF 73 (100)
T ss_pred CceeEEEEEEecccEEEEECCeEEEchHHE--EeehhccceEEEEEEecccccc------------ce--------eEEE
Confidence 3444899999999854 2 34599998754 5565323333444444332110 00 0445
Q ss_pred hcCCCCCCEEEEeecCCCCceEEEEEe
Q 012348 432 NMQLQAGDIGNQPKSQESYEMFSFMWK 458 (465)
Q Consensus 432 sK~LqaGDtVtF~R~~ngg~rf~I~~r 458 (465)
...|++||.|...|.+.+ -+|.|=-|
T Consensus 74 ~~~Lk~GD~V~ll~~~~g-Q~yiVlDk 99 (100)
T PF10844_consen 74 TDGLKVGDKVLLLRVQGG-QKYIVLDK 99 (100)
T ss_pred ecCCcCCCEEEEEEecCC-CEEEEEEe
Confidence 578999999999997766 57776443
No 6
>PRK03760 hypothetical protein; Provisional
Probab=39.26 E-value=59 Score=29.13 Aligned_cols=29 Identities=21% Similarity=0.323 Sum_probs=21.3
Q ss_pred CCCceeecCchhHhhhcCCCCCCEEEEee
Q 012348 417 NSRMYVLEGVTPCIQNMQLQAGDIGNQPK 445 (465)
Q Consensus 417 sSR~YVLtGW~~FVrsK~LqaGDtVtF~R 445 (465)
..-.|||+==.-++.+.++++||.|.|-+
T Consensus 88 ~~a~~VLEl~aG~~~~~gi~~Gd~v~~~~ 116 (117)
T PRK03760 88 KPARYIIEGPVGKIRVLKVEVGDEIEWID 116 (117)
T ss_pred ccceEEEEeCCChHHHcCCCCCCEEEEee
Confidence 35669997222335789999999998866
No 7
>PF13248 zf-ribbon_3: zinc-ribbon domain
Probab=38.00 E-value=15 Score=24.60 Aligned_cols=15 Identities=20% Similarity=0.662 Sum_probs=12.5
Q ss_pred CCCCccccccCCCce
Q 012348 77 NASGWRCCESCGKRV 91 (465)
Q Consensus 77 ~~sGWR~C~~C~Krl 91 (465)
.+.++|-|..||.+|
T Consensus 12 ~~~~~~fC~~CG~~L 26 (26)
T PF13248_consen 12 IDPDAKFCPNCGAKL 26 (26)
T ss_pred CCcccccChhhCCCC
Confidence 477899999999875
No 8
>PF08922 DUF1905: Domain of unknown function (DUF1905); InterPro: IPR015018 This family consist of hypothetical bacterial proteins. ; PDB: 2D9R_A.
Probab=36.61 E-value=1e+02 Score=25.63 Aligned_cols=79 Identities=19% Similarity=0.257 Sum_probs=39.6
Q ss_pred EEEecccccCCCCCcEEeehhhhhhcCCCCCCCCCceEEEEeCCCCeEEeEEEEcCCCCCCceeecCchhHhhhcCCCCC
Q 012348 359 FEKMLSASDAGRIGRLVLPKKCAEAYFPPISQPEGLPLKVQDSKGKEWIFQFRFWPNNNSRMYVLEGVTPCIQNMQLQAG 438 (465)
Q Consensus 359 F~KvLT~SDVg~lgRLVIPK~~AEa~FP~L~~~~Gv~L~veD~~Gk~W~Frfrfw~NnsSR~YVLtGW~~FVrsK~LqaG 438 (465)
|+..|-+.+-+ -..+.||.+-++++-.. +-..+.+.++ ..|.+|+= +..+ ...+.|+|-==.+.-++-++.+|
T Consensus 1 F~a~l~~~~~~-~~fv~vP~~v~~~l~~~--~~g~v~V~~t-I~g~~~~~--sl~p-~g~G~~~Lpv~~~vRk~~g~~~G 73 (80)
T PF08922_consen 1 FTATLWKGEGG-WTFVEVPFDVAEELGEG--GWGRVPVRGT-IDGHPWRT--SLFP-MGNGGYILPVKAAVRKAIGKEAG 73 (80)
T ss_dssp EEEE-EE-TTS--EEEE--S-HHHHH--S----S-EEEEEE-ETTEEEEE--EEEE-SSTT-EEEEE-HHHHHHHT--TT
T ss_pred CeEEEEecCCc-eEEEEeCHHHHHHhccc--cCCceEEEEE-ECCEEEEE--EEEE-CCCCCEEEEEcHHHHHHcCCCCC
Confidence 45555555443 23577899888876332 1123555554 36666655 5555 23466777423466788899999
Q ss_pred CEEEEe
Q 012348 439 DIGNQP 444 (465)
Q Consensus 439 DtVtF~ 444 (465)
|+|.+.
T Consensus 74 d~V~v~ 79 (80)
T PF08922_consen 74 DTVEVT 79 (80)
T ss_dssp SEEEEE
T ss_pred CEEEEE
Confidence 999863
No 9
>PF02643 DUF192: Uncharacterized ACR, COG1430; InterPro: IPR003795 This entry describes proteins of unknown function.; PDB: 3M7A_B 3PJY_B.
Probab=36.18 E-value=82 Score=27.42 Aligned_cols=52 Identities=21% Similarity=0.200 Sum_probs=27.6
Q ss_pred CceEEEEeCCCCeEEeEEEEcCC-------CCCCceeecCchhHhhhcCCCCCCEEEEe
Q 012348 393 GLPLKVQDSKGKEWIFQFRFWPN-------NNSRMYVLEGVTPCIQNMQLQAGDIGNQP 444 (465)
Q Consensus 393 Gv~L~veD~~Gk~W~Frfrfw~N-------nsSR~YVLtGW~~FVrsK~LqaGDtVtF~ 444 (465)
.+.|.+.|..|+.=....-..|. ...-.|||+-=..++..+++++||.|.|-
T Consensus 49 pLDi~fld~~g~Vv~i~~~~~P~~~~~~~~~~~a~~vLE~~aG~~~~~~i~~Gd~v~~~ 107 (108)
T PF02643_consen 49 PLDIAFLDSDGRVVKIERMVPPWRTYPCPSYKPARYVLELPAGWFEKLGIKVGDRVRIE 107 (108)
T ss_dssp -EEEEEE-TTSBEEEEEEEE-TT--S-EEECCEECEEEEEETTHHHHHT--TT-EEE--
T ss_pred eEEEEEECCCCeEEEEEccCCCCccCCCCCCCccCEEEEcCCCchhhcCCCCCCEEEec
Confidence 35666677766654443322111 12357999844556789999999999873
No 10
>KOG4718 consensus Non-SMC (structural maintenance of chromosomes) element 1 protein (NSE1) [Chromatin structure and dynamics]
Probab=33.42 E-value=16 Score=36.72 Aligned_cols=18 Identities=39% Similarity=0.829 Sum_probs=15.5
Q ss_pred cccccCCCceeechhhhh
Q 012348 82 RCCESCGKRVHCGCITSV 99 (465)
Q Consensus 82 R~C~~C~KrlHCGCI~S~ 99 (465)
+.|.+||-|.|||||.--
T Consensus 195 ~rCg~c~i~~h~~c~qty 212 (235)
T KOG4718|consen 195 IRCGSCNIQYHRGCIQTY 212 (235)
T ss_pred eccCcccchhhhHHHHHH
Confidence 568899999999999753
No 11
>PF04014 Antitoxin-MazE: Antidote-toxin recognition MazE; InterPro: IPR007159 This domain is found in AbrB from Bacillus subtilis. The product of the abrB gene is an ambiactive repressor and activator of the transcription of genes expressed during the transition state between vegetative growth and the onset of stationary phase and sporulation []. AbrB is thought to interact directly with the transcription initiation regions of genes under its control []. AbrB contains a helix-turn-helix structure, but this domain ends before the helix-turn-helix begins []. The product of the B. subtilis gene spoVT is another member of this family and is also a transcriptional regulator []. DNA-binding activity in this AbrB homologue requires hexamerisation []. Another family member has been isolated from the Sulfolobus solfataricus and has been identified as a homologue of bacterial repressor-like proteins. The Escherichia coli family member SohA or Prl1F appears to be bifunctional and is able to regulate its own expression as well as relieve the export block imposed by high-level synthesis of beta-galactosidase hybrid proteins [].; PDB: 2L66_A 2GLW_A 3TND_D 2W1T_B 2RO5_B 2FY9_A 2RO3_B 1UB4_C 3ZVK_G 1YFB_B ....
Probab=30.51 E-value=57 Score=24.22 Aligned_cols=23 Identities=13% Similarity=0.156 Sum_probs=19.7
Q ss_pred hhHhhhcCCCCCCEEEEeecCCC
Q 012348 427 TPCIQNMQLQAGDIGNQPKSQES 449 (465)
Q Consensus 427 ~~FVrsK~LqaGDtVtF~R~~ng 449 (465)
.++.+..+|++||.|.|.-..++
T Consensus 13 k~~~~~l~l~~Gd~v~i~~~~~g 35 (47)
T PF04014_consen 13 KEIREKLGLKPGDEVEIEVEGDG 35 (47)
T ss_dssp HHHHHHTTSSTTTEEEEEEETTS
T ss_pred HHHHHHcCCCCCCEEEEEEeCCC
Confidence 36778889999999999988775
No 12
>TIGR01643 YD_repeat_2x YD repeat (two copies). This model describes two tandem copies of a 21-residue extracellular repeat found in Gram-negative, Gram-positive, and animal proteins. The repeat is named for a YD dipeptide, the most strongly conserved motif of the repeat. These repeats appear in general to be involved in binding carbohydrate; the chicken teneurin-1 YD-repeat region has been shown to bind heparin.
Probab=30.24 E-value=76 Score=22.36 Aligned_cols=22 Identities=14% Similarity=0.149 Sum_probs=18.1
Q ss_pred CCceEEEEeCCCCeEEeEEEEc
Q 012348 392 EGLPLKVQDSKGKEWIFQFRFW 413 (465)
Q Consensus 392 ~Gv~L~veD~~Gk~W~Frfrfw 413 (465)
.|..+.+.|..|..|+|.|--.
T Consensus 4 ~g~l~~~~~p~G~~~~~~YD~~ 25 (42)
T TIGR01643 4 AGRLTGSTDADGTTTRYTYDAA 25 (42)
T ss_pred CCCEEEEECCCCCEEEEEECCC
Confidence 4778899999999999987643
No 13
>COG1998 RPS31 Ribosomal protein S27AE [Translation, ribosomal structure and biogenesis]
Probab=29.53 E-value=27 Score=27.84 Aligned_cols=13 Identities=54% Similarity=1.349 Sum_probs=8.7
Q ss_pred cccccCC---------Cceeec
Q 012348 82 RCCESCG---------KRVHCG 94 (465)
Q Consensus 82 R~C~~C~---------KrlHCG 94 (465)
|.|..|| +|+|||
T Consensus 20 ~~CPrCG~gvfmA~H~dR~~CG 41 (51)
T COG1998 20 RFCPRCGPGVFMADHKDRWACG 41 (51)
T ss_pred ccCCCCCCcchhhhcCceeEec
Confidence 4566666 488887
No 14
>PF03120 DNA_ligase_OB: NAD-dependent DNA ligase OB-fold domain; InterPro: IPR004150 DNA ligases catalyse the crucial step of joining the breaks in duplex DNA during DNA replication, repair and recombination, utilizing either ATP or NAD(+) as a cofactor []. This family is a small domain found after the adenylation domain DNA_ligase_N in NAD+-dependent ligases (IPR001679 from INTERPRO). OB-fold domains generally are involved in nucleic acid binding.; GO: 0003911 DNA ligase (NAD+) activity, 0006260 DNA replication, 0006281 DNA repair; PDB: 2OWO_A 1TAE_A 3UQ8_A 1DGS_A 1V9P_B 3SGI_A.
Probab=29.37 E-value=41 Score=28.81 Aligned_cols=22 Identities=14% Similarity=0.290 Sum_probs=17.7
Q ss_pred hhHhhhcCCCCCCEEEEeecCC
Q 012348 427 TPCIQNMQLQAGDIGNQPKSQE 448 (465)
Q Consensus 427 ~~FVrsK~LqaGDtVtF~R~~n 448 (465)
.+|+++++|..||.|.++|.-+
T Consensus 42 ~~~i~~~~i~~Gd~V~V~raGd 63 (82)
T PF03120_consen 42 YDYIKELDIRIGDTVLVTRAGD 63 (82)
T ss_dssp HHHHHHTT-BBT-EEEEEEETT
T ss_pred HHHHHHcCCCCCCEEEEEECCC
Confidence 6899999999999999999654
No 15
>PF09297 zf-NADH-PPase: NADH pyrophosphatase zinc ribbon domain; InterPro: IPR015376 This domain has a zinc ribbon structure and is often found between two NUDIX domains.; GO: 0016787 hydrolase activity, 0046872 metal ion binding; PDB: 1VK6_A 2GB5_A.
Probab=27.55 E-value=31 Score=23.97 Aligned_cols=25 Identities=32% Similarity=0.800 Sum_probs=13.3
Q ss_pred hh-ccccccccccccccccCCCCCc-cccccCCCce
Q 012348 58 LC-VYRSIYEEGRFCDTFHVNASGW-RCCESCGKRV 91 (465)
Q Consensus 58 LC-~CgsayE~~~FCd~FH~~~sGW-R~C~~C~Krl 91 (465)
-| +||+.-+ ..+.|| |-|.+||...
T Consensus 5 fC~~CG~~t~---------~~~~g~~r~C~~Cg~~~ 31 (32)
T PF09297_consen 5 FCGRCGAPTK---------PAPGGWARRCPSCGHEH 31 (32)
T ss_dssp B-TTT--BEE---------E-SSSS-EEESSSS-EE
T ss_pred ccCcCCcccc---------CCCCcCEeECCCCcCEe
Confidence 36 7777543 345677 6799998753
No 16
>cd05829 Sortase_E Sortase E (SrtE) is a membrane transpeptidase found in gram-positive bacteria that cleaves surface proteins at a cell sorting motif and catalyzes a transpeptidation reaction in which the surface protein substrate is covalently linked to peptidoglycan for display on the bacterial surface. Sortases are grouped into different classes and subfamilies based on sequence, membrane topology, genomic positioning, and cleavage site preference. The function of Sortase E is unknown. In two different sortase families, the N-terminus either functions as both a signal peptide for secretion and a stop-transfer signal for membrane anchoring, or it contains a signal peptide only and the C-terminus serves as a membrane anchor. Most gram-positive bacteria contain more than one sortase and it is thought that the different sortases anchor different surface protein classes. The sortase domain is a modified beta-barrel flanked by two (SrtA) or three (SrtB) short alpha-helices.
Probab=27.16 E-value=1.2e+02 Score=27.64 Aligned_cols=39 Identities=15% Similarity=0.112 Sum_probs=26.1
Q ss_pred CCceeec---Cch----hHhhhcCCCCCCEEEEeecCCCCceEEEE
Q 012348 418 SRMYVLE---GVT----PCIQNMQLQAGDIGNQPKSQESYEMFSFM 456 (465)
Q Consensus 418 SR~YVLt---GW~----~FVrsK~LqaGDtVtF~R~~ngg~rf~I~ 456 (465)
.+.++|. +|. .|-+=.+|++||.|.+.........|.+.
T Consensus 49 ~Gn~viaGH~~~~g~~~~F~~L~~l~~GD~I~v~~~~g~~~~Y~V~ 94 (144)
T cd05829 49 KGTAVLAGHVDSRGGPAVFFRLGDLRKGDKVEVTRADGQTATFRVD 94 (144)
T ss_pred CCCEEEEEecCCCCCChhhcchhcCCCCCEEEEEECCCCEEEEEEe
Confidence 4566773 332 38888999999999998854433444443
No 17
>PF12760 Zn_Tnp_IS1595: Transposase zinc-ribbon domain; InterPro: IPR024442 This zinc binding domain is found in a range of transposase proteins such as ISSPO8, ISSOD11, ISRSSP2 etc. It may be a zinc-binding beta ribbon domain that could bind DNA.
Probab=24.46 E-value=41 Score=25.10 Aligned_cols=26 Identities=23% Similarity=0.645 Sum_probs=17.9
Q ss_pred hh-ccccccccccccccccCCCCCccccccCCCc
Q 012348 58 LC-VYRSIYEEGRFCDTFHVNASGWRCCESCGKR 90 (465)
Q Consensus 58 LC-~CgsayE~~~FCd~FH~~~sGWR~C~~C~Kr 90 (465)
-| +||+. +.+.....+-..|..|+|+
T Consensus 20 ~CP~Cg~~-------~~~~~~~~~~~~C~~C~~q 46 (46)
T PF12760_consen 20 VCPHCGST-------KHYRLKTRGRYRCKACRKQ 46 (46)
T ss_pred CCCCCCCe-------eeEEeCCCCeEECCCCCCc
Confidence 48 99985 2233344788899999875
No 18
>cd04459 Rho_CSD Rho_CSD: Rho protein cold-shock domain (CSD). Rho protein is a transcription termination factor in most bacteria. In bacteria, there are two distinct mechanisms for mRNA transcription termination. In intrinsic termination, RNA polymerase and nascent mRNA are released from DNA template by an mRNA stem loop structure, which resembles the transcription termination mechanism used by eukaryotic pol III. The second mechanism is mediated by Rho factor. Rho factor terminates transcription by using energy from ATP hydrolysis to forcibly dissociate the transcripts from RNA polymerase. Rho protein contains an N-terminal S1-like domain, which binds single-stranded RNA. Rho has a C-terminal ATPase domain which hydrolyzes ATP to provide energy to strip RNA polymerase and mRNA from the DNA template. Rho functions as a homohexamer.
Probab=23.37 E-value=71 Score=26.39 Aligned_cols=26 Identities=19% Similarity=0.272 Sum_probs=19.0
Q ss_pred hHhhhcCCCCCCEEEE-eecCCCCceE
Q 012348 428 PCIQNMQLQAGDIGNQ-PKSQESYEMF 453 (465)
Q Consensus 428 ~FVrsK~LqaGDtVtF-~R~~ngg~rf 453 (465)
.-||..+|+.||.|.= -|...++++|
T Consensus 34 ~~Irr~~LR~GD~V~G~vr~p~~~ek~ 60 (68)
T cd04459 34 SQIRRFNLRTGDTVVGQIRPPKEGERY 60 (68)
T ss_pred HHHHHhCCCCCCEEEEEEeCCCCCCCc
Confidence 5789999999999874 4544444554
No 19
>TIGR03784 marine_sortase sortase, marine proteobacterial type. Members of this protein family are sortase enzymes, cysteine transpeptidases involved in protein sorting activities. Members of this family tend to be found in proteobacteria, rather than in Gram-positive bacteria where sortases attach proteins to the Gram-positive cell wall or participate in pilin cross-linking. Many species with this sortase appear to contain a signal target sequence, a protein with a Vault protein inter-alpha-trypsin domain (pfam08487) and a von Willebrand factor type A domain (pfam00092), encoded by an adjacent gene. These sortases are designated subfamily 6 according to Comfort and Clubb (2004).
Probab=23.30 E-value=84 Score=29.90 Aligned_cols=81 Identities=15% Similarity=0.149 Sum_probs=47.0
Q ss_pred CCcEEeehhhhhhcCCCCCCCCCceEEEEeCCCCeEEeEEEEcCCCCCCceeec--CchhHhhhcCCCCCCEEEEeecCC
Q 012348 371 IGRLVLPKKCAEAYFPPISQPEGLPLKVQDSKGKEWIFQFRFWPNNNSRMYVLE--GVTPCIQNMQLQAGDIGNQPKSQE 448 (465)
Q Consensus 371 lgRLVIPK~~AEa~FP~L~~~~Gv~L~veD~~Gk~W~Frfrfw~NnsSR~YVLt--GW~~FVrsK~LqaGDtVtF~R~~n 448 (465)
.+||.||+--.+ +|-+.+..+..|.+ | .-.+..+-.| +..+.++|. .-+.|-+=.+|+.||.|.+.....
T Consensus 51 va~L~IP~lg~~--~~V~~G~s~~~L~~----G-~G~~~~t~~P-G~~Gn~VIAGHrdt~F~~L~~L~~GD~I~v~~~~g 122 (174)
T TIGR03784 51 VAKLSAPRLGAS--LYVLAGASGRNLAF----G-PGHMLATAQP-GAQGNSVIAGHRDTHFAFLQELRPGDVIRLQTPDG 122 (174)
T ss_pred eEEEEEccCCCc--eeEEecCCHHHHhh----e-eEEecCCCCC-CCCCcEEEEeeCCccCCChhhCCCCCEEEEEECCC
Confidence 579999985432 45554444333321 1 1112222233 234677884 334699999999999999987655
Q ss_pred CCceEEEEEee
Q 012348 449 SYEMFSFMWKV 459 (465)
Q Consensus 449 gg~rf~I~~rr 459 (465)
...+|.+.-.+
T Consensus 123 ~~~~Y~V~~~~ 133 (174)
T TIGR03784 123 QWQSYQVTATR 133 (174)
T ss_pred eEEEEEEeEEE
Confidence 43456665443
No 20
>TIGR00375 conserved hypothetical protein TIGR00375. The member of this family from Methanococcus jannaschii, MJ0043, is considerably longer and appears to contain an intein N-terminal to the region of homology.
Probab=22.83 E-value=36 Score=36.24 Aligned_cols=44 Identities=23% Similarity=0.330 Sum_probs=29.5
Q ss_pred HHHHhh-ccccccccccccccccCCCCCccccccCCCceeechhhhhhhhhhhhc
Q 012348 54 LTLILC-VYRSIYEEGRFCDTFHVNASGWRCCESCGKRVHCGCITSVHAFTLLDA 107 (465)
Q Consensus 54 ~~~~LC-~CgsayE~~~FCd~FH~~~sGWR~C~~C~KrlHCGCI~S~~~~~lLD~ 107 (465)
|+-+-| +|+.-|+.. | ..+-||+ |. |||+|-=| |.--.-||=|.
T Consensus 238 Yh~~~c~~C~~~~~~~---~---~~~~~~~-Cp-CG~~i~~G--V~~Rv~eLad~ 282 (374)
T TIGR00375 238 YHQTACEACGEPAVSE---D---AETACAN-CP-CGGRIKKG--VSDRLRELSDQ 282 (374)
T ss_pred cchhhhcccCCcCCch---h---hhhcCCC-CC-CCCcceec--hHHHHHHHhcC
Confidence 678889 998877643 1 2233788 88 99998766 45555566563
No 21
>PRK09570 rpoH DNA-directed RNA polymerase subunit H; Reviewed
Probab=22.29 E-value=1e+02 Score=26.37 Aligned_cols=32 Identities=9% Similarity=0.262 Sum_probs=23.8
Q ss_pred hhHhhhcCCCCCCEEEEeecCCCCceEEEEEee
Q 012348 427 TPCIQNMQLQAGDIGNQPKSQESYEMFSFMWKV 459 (465)
Q Consensus 427 ~~FVrsK~LqaGDtVtF~R~~ngg~rf~I~~rr 459 (465)
-+.++.-+++.||+|-+.|..... .-.+.||.
T Consensus 44 DPv~r~~g~k~GdVvkI~R~S~ta-G~~v~YR~ 75 (79)
T PRK09570 44 DPVVKAIGAKPGDVIKIVRKSPTA-GEAVYYRL 75 (79)
T ss_pred ChhhhhcCCCCCCEEEEEECCCCC-CccEEEEE
Confidence 367888899999999999986542 22566664
No 22
>PF01191 RNA_pol_Rpb5_C: RNA polymerase Rpb5, C-terminal domain; InterPro: IPR000783 Prokaryotes contain a single DNA-dependent RNA polymerase (RNAP; 2.7.7.6 from EC) that is responsible for the transcription of all genes, while eukaryotes have three classes of RNAPs (I-III) that transcribe different sets of genes. Each class of RNA polymerase is an assemblage of ten to twelve different polypeptides. Certain subunits of RNAPs, including RPB5 (POLR2E in mammals), are common to all three eukaryotic polymerases. RPB5 plays a role in the transcription activation process. Eukaryotic RPB5 has a bipartite structure consisting of a unique N-terminal region (IPR005571 from INTERPRO), plus a C-terminal region that is structurally homologous to the prokaryotic RPB5 homologue, subunit H (gene rpoH) [, , , ]. This entry represents prokaryotic subunit H and the C-terminal domain of eukaryotic RPB5, which share a two-layer alpha/beta fold, with a core structure of beta/alpha/beta/alpha/beta(2). ; GO: 0003677 DNA binding, 0003899 DNA-directed RNA polymerase activity, 0006351 transcription, DNA-dependent; PDB: 1EIK_A 2Y0S_Z 1DZF_A 3GTG_E 2VUM_E 3GTP_E 3GTO_E 3S17_E 3S1R_E 1I3Q_E ....
Probab=22.29 E-value=1.2e+02 Score=25.57 Aligned_cols=32 Identities=13% Similarity=0.225 Sum_probs=23.0
Q ss_pred hhHhhhcCCCCCCEEEEeecCCCCceEEEEEee
Q 012348 427 TPCIQNMQLQAGDIGNQPKSQESYEMFSFMWKV 459 (465)
Q Consensus 427 ~~FVrsK~LqaGDtVtF~R~~ngg~rf~I~~rr 459 (465)
.+.++..+++.||+|-+.|..... .-.+.||.
T Consensus 41 DPv~r~~g~k~GdVvkI~R~S~ta-G~~v~YR~ 72 (74)
T PF01191_consen 41 DPVARYLGAKPGDVVKIIRKSETA-GEYVTYRL 72 (74)
T ss_dssp SHHHHHTT--TTSEEEEEEEETTT-SEEEEEEE
T ss_pred ChhhhhcCCCCCCEEEEEecCCCC-CCcEEEEE
Confidence 478888999999999999977652 34666663
No 23
>PF05593 RHS_repeat: RHS Repeat; InterPro: IPR006530 These sequences contain two tandem copies of a 21-residue extracellular repeat that is found in Gram-negative, Gram-positive, and animal proteins. The repeat is named for a YD dipeptide, the most strongly conserved motif of the repeat. These repeats appear in general to be involved in binding carbohydrate; the chicken teneurin-1 YD-repeat region has been shown to bind heparin [, , ].
Probab=22.18 E-value=1.4e+02 Score=21.29 Aligned_cols=23 Identities=17% Similarity=0.268 Sum_probs=18.1
Q ss_pred CCCceEEEEeCCCCeEEeEEEEc
Q 012348 391 PEGLPLKVQDSKGKEWIFQFRFW 413 (465)
Q Consensus 391 ~~Gv~L~veD~~Gk~W~Frfrfw 413 (465)
..|..+.+.|..|.+|+|.|--.
T Consensus 3 ~~G~l~~~~d~~G~~~~y~YD~~ 25 (38)
T PF05593_consen 3 ANGRLTSVTDPDGRTTRYTYDAA 25 (38)
T ss_pred CCCCEEEEEcCCCCEEEEEECCC
Confidence 35778889999999998776544
Done!