Query 025018
Match_columns 259
No_of_seqs 193 out of 1004
Neff 5.1
Searched_HMMs 46136
Date Fri Mar 29 09:18:56 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/025018.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/025018hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF04970 LRAT: Lecithin retino 100.0 7.1E-33 1.5E-37 225.1 9.0 122 9-163 3-124 (125)
2 PRK11479 hypothetical protein; 98.5 5.5E-07 1.2E-11 83.5 8.2 40 5-44 57-107 (274)
3 TIGR02219 phage_NlpC_fam putat 98.4 3.8E-07 8.1E-12 75.6 5.7 41 6-46 70-111 (134)
4 PRK10838 spr outer membrane li 98.3 7.1E-07 1.5E-11 78.7 5.7 42 4-46 120-161 (190)
5 PF08405 Calici_PP_N: Viral po 98.3 1.4E-06 3.1E-11 81.7 7.7 37 8-46 4-40 (358)
6 PF00877 NLPC_P60: NlpC/P60 fa 98.0 6.5E-06 1.4E-10 64.4 3.5 37 7-44 46-82 (105)
7 PF05708 DUF830: Orthopoxvirus 97.9 6.1E-05 1.3E-09 62.5 8.4 96 12-159 1-118 (158)
8 COG0791 Spr Cell wall-associat 97.9 1.5E-05 3.3E-10 68.9 4.9 42 5-46 131-173 (197)
9 PRK13914 invasion associated s 97.8 2E-05 4.4E-10 78.1 5.3 41 4-45 418-458 (481)
10 PRK10030 hypothetical protein; 97.6 0.00031 6.8E-09 62.1 8.0 85 9-147 17-118 (197)
11 PRK11470 hypothetical protein; 96.7 0.0063 1.4E-07 54.4 7.9 87 11-147 7-108 (200)
12 TIGR02594 conserved hypothetic 96.5 0.0034 7.5E-08 52.1 4.4 39 5-47 68-109 (129)
13 PF05608 DUF778: Protein of un 96.1 0.033 7.2E-07 47.1 8.3 38 123-162 76-113 (136)
14 PF05903 Peptidase_C97: PPPDE 95.9 0.01 2.2E-07 50.3 4.2 34 127-160 84-119 (151)
15 PF05382 Amidase_5: Bacterioph 92.8 0.17 3.7E-06 43.1 4.5 39 7-45 67-111 (145)
16 KOG0324 Uncharacterized conser 92.4 0.076 1.7E-06 47.9 2.0 34 125-158 85-120 (214)
17 COG3863 Uncharacterized distan 92.4 0.26 5.7E-06 44.3 5.2 41 6-46 72-127 (231)
18 PF06672 DUF1175: Protein of u 89.6 0.78 1.7E-05 41.6 5.6 37 9-45 132-173 (216)
19 PF05257 CHAP: CHAP domain; I 78.0 2.5 5.3E-05 33.8 3.2 29 9-37 59-88 (124)
20 KOG3150 Uncharacterized conser 75.2 10 0.00022 33.3 6.3 37 122-160 91-127 (182)
21 PF06940 DUF1287: Domain of un 73.9 7 0.00015 34.1 5.1 41 6-47 100-146 (164)
22 PF10030 DUF2272: Uncharacteri 62.6 15 0.00032 32.6 4.8 45 3-47 84-145 (183)
23 PF03658 Ub-RnfH: RnfH family 60.9 3.7 8E-05 32.1 0.7 25 1-25 49-74 (84)
24 COG3738 Uncharacterized protei 55.3 24 0.00053 31.3 4.9 38 9-47 137-180 (200)
25 COG3234 Uncharacterized protei 48.0 23 0.0005 31.7 3.6 33 9-43 138-170 (215)
26 PF07313 DUF1460: Protein of u 46.0 34 0.00073 31.0 4.5 37 11-47 152-193 (216)
27 PF01052 SpoA: Surface present 43.4 34 0.00075 25.0 3.5 31 11-43 27-57 (77)
28 KOG4577 Transcription factor L 43.4 9.1 0.0002 36.6 0.4 55 68-137 60-115 (383)
29 PF05820 DUF845: Baculovirus p 42.5 20 0.00043 29.8 2.2 24 130-153 91-114 (119)
30 cd04482 RPA2_OBF_like RPA2_OBF 37.0 78 0.0017 24.4 4.7 8 68-75 84-91 (91)
31 PF08007 Cupin_4: Cupin superf 36.9 25 0.00055 33.0 2.3 33 8-45 175-207 (319)
32 TIGR02480 fliN flagellar motor 35.4 56 0.0012 24.4 3.6 31 10-42 26-56 (77)
33 PF13387 DUF4105: Domain of un 35.1 38 0.00082 29.0 2.9 15 142-156 128-142 (176)
34 COG2850 Uncharacterized conser 34.7 15 0.00034 35.9 0.5 34 6-44 176-209 (383)
35 PF11730 DUF3297: Protein of u 32.0 18 0.00039 27.4 0.4 44 213-257 16-60 (71)
36 KOG3416 Predicted nucleic acid 29.7 43 0.00094 28.3 2.3 13 11-23 60-72 (134)
37 KOG3706 Uncharacterized conser 29.7 27 0.00059 35.7 1.3 39 4-45 376-414 (629)
38 COG2914 Uncharacterized protei 27.8 43 0.00093 27.0 1.9 25 1-25 52-77 (99)
39 PRK11032 hypothetical protein; 26.7 51 0.0011 28.6 2.3 9 69-77 143-151 (160)
40 PF06887 DUF1265: Protein of u 25.1 61 0.0013 22.9 2.0 36 147-189 3-38 (48)
41 PF00122 E1-E2_ATPase: E1-E2 A 24.9 37 0.00081 29.3 1.2 19 7-25 46-64 (230)
42 PRK06033 hypothetical protein; 24.4 1.2E+02 0.0026 23.3 3.8 31 11-43 26-56 (83)
43 PF10077 DUF2314: Uncharacteri 24.2 75 0.0016 26.4 2.8 35 2-38 68-103 (133)
44 PRK03187 tgl transglutaminase; 22.7 1.1E+02 0.0023 29.0 3.7 29 11-39 164-198 (272)
45 PF11948 DUF3465: Protein of u 22.2 76 0.0017 26.8 2.5 30 11-47 84-113 (131)
46 PRK01777 hypothetical protein; 22.1 60 0.0013 25.6 1.8 25 1-25 52-77 (95)
47 cd05834 HDGF_related The PWWP 21.9 65 0.0014 24.6 1.9 18 12-31 2-19 (83)
48 COG0272 Lig NAD-dependent DNA 21.0 93 0.002 32.9 3.3 19 7-25 362-380 (667)
49 PF12671 Amidase_6: Putative a 20.9 1.2E+02 0.0026 25.4 3.4 23 14-36 99-122 (157)
No 1
>PF04970 LRAT: Lecithin retinol acyltransferase; InterPro: IPR007053 This entry represents a conserved sequence region found in proteins from viruses, bacteria and eukaryotes. It contains a well-conserved NCEHF motif, though its function in these proteins is unknown.; PDB: 2KYT_A 4DOT_A 4FA0_A.
Probab=99.98 E-value=7.1e-33 Score=225.08 Aligned_cols=122 Identities=35% Similarity=0.628 Sum_probs=80.8
Q ss_pred CCCCCCCCCEEEEeecCcccceEEEEEcCCEEEEeCCCCCccccccccccccccCCCCcccCCCCcCccCCCCceEEccc
Q 025018 9 ERNEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVHFRPERNLIVGAETSSETQNSILPSSCLIFPDCGFRQPNSGVILSCL 88 (259)
Q Consensus 9 ~~~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH~~~~~~~~~g~~t~l~~~~s~~p~~~~~~~~cg~~~~~~gVv~s~L 88 (259)
+.++|+|||||+++|.. |+|||||+|+++|||+.++.+...++. ...++.......|+.++|
T Consensus 3 ~~~~~~~GD~I~~~r~~--y~H~gIYvG~~~ViH~~~~~~~~~~~~----------------~~~~~~~~~~~~V~~~~l 64 (125)
T PF04970_consen 3 DKKRLKPGDHIEVPRGL--YEHWGIYVGDGEVIHFSGPGEISVSNR----------------SSICGFSKKKAEVKKDSL 64 (125)
T ss_dssp ---S--TT-EEEEEETT--EEEEEEEEETTEEEEEE-S-SSS-SSS----------------SGGGGT--S-EEEEEEEH
T ss_pred cccCCCCCCEEEEecCC--ccEEEEEecCCeEEEeccccccccccc----------------ccccceecCCCEEEEEEh
Confidence 35789999999999996 999999999999999997654211111 112333444567999999
Q ss_pred hhhcCCCceEEEeeccCcceeeehccCCcccccCCCCHHHHHHHHHHHhhcCCcccccccCchhHHHHHhhhCcc
Q 025018 89 DCFLGNGSLYCFEYGVAPSVFLAKVRGGTCTTATSDPPETVIHRAMYLLQNGFGNYNVFQNNCEDFALYCRTGLL 163 (259)
Q Consensus 89 ~~Fl~G~~l~~f~Y~vs~~~flak~rggtC~~~~~~p~eeVV~RA~~~L~~G~g~YnL~~NNCEHFA~~CktGl~ 163 (259)
++|+.|..+++..|- + ...+++++++|++||+++|++++ +|||++|||||||+|||||..
T Consensus 65 ~~~~~~~~~~v~~~~----------~----~~~~~~~~~~iv~rA~~~lg~~~-~Y~l~~nNCEhFa~~c~tG~~ 124 (125)
T PF04970_consen 65 EEFAQGRKVRVNNYL----------D----HRYKPFPPEEIVERAESRLGKEF-EYNLLFNNCEHFATWCRTGKS 124 (125)
T ss_dssp HHHHTTSEEEE--GG----------G----GTS--S-HHHHHHHHHHTTT-EE-SS---HHHHHHHHHHHHHS--
T ss_pred HHhcCCCEEEEEecC----------C----ccCCCCCHHHHHHHHHHHHcCCC-ccCCCcCCHHHHHHHHHcCCC
Confidence 999999987764331 1 34779999999999999996545 999999999999999999964
No 2
>PRK11479 hypothetical protein; Provisional
Probab=98.46 E-value=5.5e-07 Score=83.46 Aligned_cols=40 Identities=23% Similarity=0.443 Sum_probs=33.7
Q ss_pred CcccCCCCCCCCCEEEEeec-----------CcccceEEEEEcCCEEEEeC
Q 025018 5 TNRVERNEIKAGDHIYTYRA-----------VFAYSHHGIYVGGSKVVHFR 44 (259)
Q Consensus 5 ~~~v~~~~lk~GD~I~~~r~-----------~~~y~H~GIYvG~g~VIH~~ 44 (259)
+++|+.+++||||+|++... .-.++|.|||+|+++|||++
T Consensus 57 g~~Vs~~~LqpGDLVFfst~t~~S~~Ik~~T~s~~SHVgIylGdg~vIEA~ 107 (274)
T PRK11479 57 IKEITAPDLKPGDLLFSSSLGVTSFGIRVFSTSSVSHVAIYLGENNVAEAT 107 (274)
T ss_pred CcccChhhCCCCCEEEEecCCccccceecccCCCCcEEEEEecCCeEEEcC
Confidence 56899999999999998532 12479999999999999984
No 3
>TIGR02219 phage_NlpC_fam putative phage cell wall peptidase, NlpC/P60 family. Members of this family show sequence similarity to members of the NlpC/P60 family described by Pfam model pfam00877 and by Anantharaman and Aravind (PubMed:12620121). The NlpC/P60 family includes a number of characterized bacterial cell wall hydrolases. Members of this related family are all found in prophage regions of bacterial genomes.
Probab=98.43 E-value=3.8e-07 Score=75.57 Aligned_cols=41 Identities=17% Similarity=0.120 Sum_probs=33.5
Q ss_pred cccCCCCCCCCCEEEEeec-CcccceEEEEEcCCEEEEeCCC
Q 025018 6 NRVERNEIKAGDHIYTYRA-VFAYSHHGIYVGGSKVVHFRPE 46 (259)
Q Consensus 6 ~~v~~~~lk~GD~I~~~r~-~~~y~H~GIYvG~g~VIH~~~~ 46 (259)
.+|+++++||||+|+|.-. +....|.|||+|++++||.+..
T Consensus 70 ~~v~~~~~qpGDlvff~~~~~~~~~HvGIy~G~g~~iHa~~~ 111 (134)
T TIGR02219 70 VPVPCDAAQPGDVLVFRWRPGAAAKHAAIAASPTRFIHAYDG 111 (134)
T ss_pred cccchhcCCCCCEEEEeeCCCCCCcEEEEEeCCCcEEEECCC
Confidence 4678899999999999632 2125899999999999999864
No 4
>PRK10838 spr outer membrane lipoprotein; Provisional
Probab=98.34 E-value=7.1e-07 Score=78.72 Aligned_cols=42 Identities=26% Similarity=0.515 Sum_probs=35.3
Q ss_pred CCcccCCCCCCCCCEEEEeecCcccceEEEEEcCCEEEEeCCC
Q 025018 4 LTNRVERNEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVHFRPE 46 (259)
Q Consensus 4 ~~~~v~~~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH~~~~ 46 (259)
.+.+|++++++|||+|+|.... ...|+|||+||+++||.+..
T Consensus 120 ~g~~V~~~~lqpGDLVfF~~~~-~~~HVGIyiGng~~IHAs~~ 161 (190)
T PRK10838 120 MGKSVSRSKLRTGDLVLFRAGS-TGRHVGIYIGNNQFVHASTS 161 (190)
T ss_pred cCcCcccCCCCCCcEEEECCCC-CCCEEEEEecCCEEEEeCCC
Confidence 4678999999999999986433 24799999999999999764
No 5
>PF08405 Calici_PP_N: Viral polyprotein N-terminal; InterPro: IPR013614 This domain is found at the N terminus of non-structural viral polyproteins of the Caliciviridae subfamily. ; GO: 0003968 RNA-directed RNA polymerase activity, 0004197 cysteine-type endopeptidase activity, 0017111 nucleoside-triphosphatase activity, 0044419 interspecies interaction between organisms
Probab=98.33 E-value=1.4e-06 Score=81.75 Aligned_cols=37 Identities=19% Similarity=0.212 Sum_probs=31.3
Q ss_pred cCCCCCCCCCEEEEeecCcccceEEEEEcCCEEEEeCCC
Q 025018 8 VERNEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVHFRPE 46 (259)
Q Consensus 8 v~~~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH~~~~ 46 (259)
.+..++++|++|+++-+. +.|+|||+|+|+++-..++
T Consensus 4 ~~a~EP~~GsilE~~eG~--~yHYaIYi~~G~~lgvh~p 40 (358)
T PF08405_consen 4 MPAREPLIGSILEMDEGD--IYHYAIYIGKGLVLGVHSP 40 (358)
T ss_pred CCCCCCCCCceEEEecCe--eEEEEEEecCCeEEeecCc
Confidence 467899999999999987 7899999999999744443
No 6
>PF00877 NLPC_P60: NlpC/P60 family; InterPro: IPR000064 The Escherichia coli NLPC/Listeria P60 domain occurs at the C terminus of a number of different bacterial and viral proteins. The viral proteins are either described as tail assembly proteins or Gp19. In bacteria, the proteins are variously described as being putative tail component of prophage, invasin, invasion associated protein, putative lipoprotein, cell wall hydrolase, or putative endopeptidase. The E. coli NLPC/Listeria P60 domain is contained within the boundaries of the cysteine peptidase domain that defines the MEROPS peptidase family C40 (clan C-). A type example being dipeptidyl-peptidase VI from Bacillus sphaericus and gamma-glutamyl-diamino acid-endopeptidase precursor from Lactococcus lactis 3.4.19.11 from EC. This group also contains proteins classified as non-peptidase homologues in that they either have been found experimentally to be without peptidase activity, or lack amino acid residues that are believed to be essential for the catalytic activity of peptidases in the C40 family. ; PDB: 3PVQ_B 3GT2_A 3NPF_B 2K1G_A 3I86_A 3S0Q_A 2XIV_A 3PBC_A 3NE0_A 3M1U_B ....
Probab=97.96 E-value=6.5e-06 Score=64.36 Aligned_cols=37 Identities=38% Similarity=0.561 Sum_probs=32.3
Q ss_pred ccCCCCCCCCCEEEEeecCcccceEEEEEcCCEEEEeC
Q 025018 7 RVERNEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVHFR 44 (259)
Q Consensus 7 ~v~~~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH~~ 44 (259)
.++.++++|||+|++.. .....|.|||+|++++||..
T Consensus 46 ~~~~~~~~pGDlif~~~-~~~~~Hvgiy~g~~~~iha~ 82 (105)
T PF00877_consen 46 RVPISELQPGDLIFFKG-GGGISHVGIYLGDGKFIHAS 82 (105)
T ss_dssp HEEGGG-TTTEEEEEEG-TGGEEEEEEEEETTEEEEEE
T ss_pred ccchhcCCcccEEEEeC-CccCCEeEEEEeCCeEEEeC
Confidence 48899999999999987 33479999999999999999
No 7
>PF05708 DUF830: Orthopoxvirus protein of unknown function (DUF830); PDB: 2IF6_B 3KW0_C.
Probab=97.90 E-value=6.1e-05 Score=62.53 Aligned_cols=96 Identities=23% Similarity=0.346 Sum_probs=57.5
Q ss_pred CCCCCCEEEEeecC-----------cccceEEEEEcCC----EEEEeCCCCCccccccccccccccCCCCcccCCCCcCc
Q 025018 12 EIKAGDHIYTYRAV-----------FAYSHHGIYVGGS----KVVHFRPERNLIVGAETSSETQNSILPSSCLIFPDCGF 76 (259)
Q Consensus 12 ~lk~GD~I~~~r~~-----------~~y~H~GIYvG~g----~VIH~~~~~~~~~g~~t~l~~~~s~~p~~~~~~~~cg~ 76 (259)
+||+||+|++.... ..|.|.|||++++ .|+|+....
T Consensus 1 ~l~~GDIil~~~~~~~s~~i~~~t~~~~~HvgI~~~~~~~~~~viea~~~~----------------------------- 51 (158)
T PF05708_consen 1 KLQTGDIILTRGKSSLSKAIRPVTSSPYSHVGIVIGDEGQEPYVIEATPGD----------------------------- 51 (158)
T ss_dssp ---TT-EEEEEE-SCCHHHHHHHHTSS--EEEEEEEETTE-EEEEEEETTT-----------------------------
T ss_pred CCCCeeEEEEECCchHHHHHHHHhCCCCCEEEEEEecCCCceEEEEeccCC-----------------------------
Confidence 58999999996532 2489999999987 689985422
Q ss_pred cCCCCceEEccchhhcCC-CceEEEeeccCcceeeehccCCcccccCCCCHHHHHHHHHHHhhcCCcccccc------cC
Q 025018 77 RQPNSGVILSCLDCFLGN-GSLYCFEYGVAPSVFLAKVRGGTCTTATSDPPETVIHRAMYLLQNGFGNYNVF------QN 149 (259)
Q Consensus 77 ~~~~~gVv~s~L~~Fl~G-~~l~~f~Y~vs~~~flak~rggtC~~~~~~p~eeVV~RA~~~L~~G~g~YnL~------~N 149 (259)
||....|+.|+.. +.+.++.+... ....-.+.+++.|.+++ |. .|++. .-
T Consensus 52 -----Gv~~~~l~~~~~~~~~~~V~r~~~~---------------~~~~~~~~~~~~a~~~~--g~-~Y~~~~~~~~~~~ 108 (158)
T PF05708_consen 52 -----GVRLEPLSDFLKRNEKIAVYRLKDP---------------LSEEQRQKAAEFAKSYI--GK-PYDFNFSLDDDRF 108 (158)
T ss_dssp -----CEEEEECHHHHHCCCEEEEEEECCG---------------TTCHHHHHHHHHHHCCT--TS--B-CC-HCCSSSB
T ss_pred -----CeEEeeHHHHhcCCceEEEEEECCC---------------CCHHHHHHHHHHHHHHc--CC-CccccccCCCCCE
Confidence 5788899999884 44444322211 01223556777777777 43 77776 34
Q ss_pred chhHHHHHhh
Q 025018 150 NCEDFALYCR 159 (259)
Q Consensus 150 NCEHFA~~Ck 159 (259)
.|=.|+..|-
T Consensus 109 yCSelV~~~y 118 (158)
T PF05708_consen 109 YCSELVAEAY 118 (158)
T ss_dssp -HHHHHHHHH
T ss_pred EcHHHHHHHH
Confidence 6888888885
No 8
>COG0791 Spr Cell wall-associated hydrolases (invasion-associated proteins) [Cell envelope biogenesis, outer membrane]
Probab=97.90 E-value=1.5e-05 Score=68.90 Aligned_cols=42 Identities=21% Similarity=0.468 Sum_probs=35.8
Q ss_pred CcccCCCCCCCCCEEEEeec-CcccceEEEEEcCCEEEEeCCC
Q 025018 5 TNRVERNEIKAGDHIYTYRA-VFAYSHHGIYVGGSKVVHFRPE 46 (259)
Q Consensus 5 ~~~v~~~~lk~GD~I~~~r~-~~~y~H~GIYvG~g~VIH~~~~ 46 (259)
+.+|+..+++|||+|+|... .....|.|||+|+|++||.+..
T Consensus 131 g~~v~~~~~~~GDlvff~~~~~~~~~Hvgiy~g~g~~iha~~~ 173 (197)
T COG0791 131 GTAVDDSDLQPGDLVFFNTGGGSSANHVGIYLGNGQFIHAAGS 173 (197)
T ss_pred cCccChhhCCCCCEEEEecCCCCCCCeEEEEecCCeEEecCCC
Confidence 67888999999999999862 3347899999999999999764
No 9
>PRK13914 invasion associated secreted endopeptidase; Provisional
Probab=97.84 E-value=2e-05 Score=78.10 Aligned_cols=41 Identities=29% Similarity=0.537 Sum_probs=34.1
Q ss_pred CCcccCCCCCCCCCEEEEeecCcccceEEEEEcCCEEEEeCC
Q 025018 4 LTNRVERNEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVHFRP 45 (259)
Q Consensus 4 ~~~~v~~~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH~~~ 45 (259)
.+.+|+.++++|||+|||.... ...|+|||+|+|++||...
T Consensus 418 ~G~~Vs~selqpGDLVFF~~~~-~~~HVGIYiGnG~~IHA~~ 458 (481)
T PRK13914 418 STTRISESQAKPGDLVFFDYGS-GISHVGIYVGNGQMINAQD 458 (481)
T ss_pred cCcccccccCCCCCEEEeCCCC-CCCEEEEEeCCCEEEEcCC
Confidence 3678999999999999996433 2579999999999999753
No 10
>PRK10030 hypothetical protein; Provisional
Probab=97.55 E-value=0.00031 Score=62.09 Aligned_cols=85 Identities=19% Similarity=0.245 Sum_probs=53.9
Q ss_pred CCCCCCCCCEEEEeecC-----------cccceEEEEEcC---CEEEEeCCCCCccccccccccccccCCCCcccCCCCc
Q 025018 9 ERNEIKAGDHIYTYRAV-----------FAYSHHGIYVGG---SKVVHFRPERNLIVGAETSSETQNSILPSSCLIFPDC 74 (259)
Q Consensus 9 ~~~~lk~GD~I~~~r~~-----------~~y~H~GIYvG~---g~VIH~~~~~~~~~g~~t~l~~~~s~~p~~~~~~~~c 74 (259)
...++++||+|++.-.. -.|+|.||+++. -.|+|+.+
T Consensus 17 ~~~~l~~GDlif~~g~~~~s~aI~~~T~s~~SHVGIi~~~~~~~~ViEAv~----------------------------- 67 (197)
T PRK10030 17 FAWQPQTGDIIFQISRSSQSKAIQLATHSDYSHTGMIVKRNKKPYVFEAVG----------------------------- 67 (197)
T ss_pred hhcCCCCCCEEEEeCCCcHhHHHhHhhCCCCceEEEEEEECCcEEEEEecC-----------------------------
Confidence 44589999999985421 249999998873 25888742
Q ss_pred CccCCCCceEEccchhhcCCC---ceEEEeeccCcceeeehccCCcccccCCCCHHHHHHHHHHHhhcCCcccccc
Q 025018 75 GFRQPNSGVILSCLDCFLGNG---SLYCFEYGVAPSVFLAKVRGGTCTTATSDPPETVIHRAMYLLQNGFGNYNVF 147 (259)
Q Consensus 75 g~~~~~~gVv~s~L~~Fl~G~---~l~~f~Y~vs~~~flak~rggtC~~~~~~p~eeVV~RA~~~L~~G~g~YnL~ 147 (259)
+|+.++|+.|++-. .+.++++.. . ..+...+.+++.|.+.+ |. .||+.
T Consensus 68 -------~V~~~pL~~Fl~~~~~~~~~V~Rl~~--~-------------lt~~~~~~li~~A~~~l--Gk-pYD~~ 118 (197)
T PRK10030 68 -------PVKYTPLKQWIAHGEKGKYVVRRLEN--G-------------LSVEQQQKLAQTAKRYL--GK-PYDFY 118 (197)
T ss_pred -------ceEEEEHHHHhhcCccCcEEEEEeCC--C-------------CCHHHHHHHHHHHHHHc--CC-CCCcc
Confidence 37888999999643 222211110 0 01123456777888888 54 78865
No 11
>PRK11470 hypothetical protein; Provisional
Probab=96.74 E-value=0.0063 Score=54.35 Aligned_cols=87 Identities=17% Similarity=0.223 Sum_probs=52.4
Q ss_pred CCCCCCCEEEEeec-----------CcccceEEEEEcC---C-EEEEeCCCCCccccccccccccccCCCCcccCCCCcC
Q 025018 11 NEIKAGDHIYTYRA-----------VFAYSHHGIYVGG---S-KVVHFRPERNLIVGAETSSETQNSILPSSCLIFPDCG 75 (259)
Q Consensus 11 ~~lk~GD~I~~~r~-----------~~~y~H~GIYvG~---g-~VIH~~~~~~~~~g~~t~l~~~~s~~p~~~~~~~~cg 75 (259)
.+++.||+|+..-. +.-++|.||.++. + .|+|...+
T Consensus 7 ~~l~~GDLvF~~~~~~~~~aI~~aT~s~~sHvGII~~~~~~~~~VlEA~~~----------------------------- 57 (200)
T PRK11470 7 AEYEIGDIVFTCIGAALFGQISAASNCWSNHVGIIIGHNGEDFLVAESRVP----------------------------- 57 (200)
T ss_pred CCCCCCCEEEEeCCcchhHHHHhccCCccceEEEEEEEcCCceEEEEecCC-----------------------------
Confidence 58999999998631 1346899999843 2 66776431
Q ss_pred ccCCCCceEEccchhhcCCCceEEEeeccCcceeeehccCCcccccCCCCHHHHHHHHHHHhhcCCcccccc
Q 025018 76 FRQPNSGVILSCLDCFLGNGSLYCFEYGVAPSVFLAKVRGGTCTTATSDPPETVIHRAMYLLQNGFGNYNVF 147 (259)
Q Consensus 76 ~~~~~~gVv~s~L~~Fl~G~~l~~f~Y~vs~~~flak~rggtC~~~~~~p~eeVV~RA~~~L~~G~g~YnL~ 147 (259)
+|+.+.|+.|++-+.- ..+.+++.. ....++-...+++.|+++|++ .||.-
T Consensus 58 ------~vr~TpLs~fi~r~~~--------g~i~v~Rl~----~~l~~~~~~~~~~~A~~~lGk---pYD~~ 108 (200)
T PRK11470 58 ------LSTVTTLSRFIKRSAN--------QRYAIKRLD----AGLTEQQKQRIVEQVPSRLRK---LYHTG 108 (200)
T ss_pred ------ceEEeEHHHHHhcCcC--------ceEEEEEec----CCCCHHHHHHHHHHHHHHcCC---CCCCc
Confidence 3577899999975331 111222221 011222345588899999943 55553
No 12
>TIGR02594 conserved hypothetical protein TIGR02594. Members of this protein family known so far are restricted to the bacteria, and for the most to the proteobacteria. The function is unknown.
Probab=96.53 E-value=0.0034 Score=52.09 Aligned_cols=39 Identities=15% Similarity=0.123 Sum_probs=29.2
Q ss_pred CcccCCCCCCCCCEEEEeecCcccceEEEEEcCC-E--EEEeCCCC
Q 025018 5 TNRVERNEIKAGDHIYTYRAVFAYSHHGIYVGGS-K--VVHFRPER 47 (259)
Q Consensus 5 ~~~v~~~~lk~GD~I~~~r~~~~y~H~GIYvG~g-~--VIH~~~~~ 47 (259)
+.+++ +++|||+|+|++.. ..|+|||+|++ . .||.-++.
T Consensus 68 G~~v~--~p~~GDiv~f~~~~--~~HVGi~~g~~~~~g~i~~lgGN 109 (129)
T TIGR02594 68 GTKLS--KPAYGCIAVKRRGG--GGHVGFVVGKDKQTGTIIVLGGN 109 (129)
T ss_pred CCcCC--CCCccEEEEEECCC--CCEEEEEEeEcCCCCEEEEeeCC
Confidence 44444 78999999998766 67999999964 2 57766643
No 13
>PF05608 DUF778: Protein of unknown function (DUF778); InterPro: IPR008496 This family consists of several eukaryotic proteins of unknown function.
Probab=96.14 E-value=0.033 Score=47.05 Aligned_cols=38 Identities=18% Similarity=0.387 Sum_probs=31.1
Q ss_pred CCCHHHHHHHHHHHhhcCCcccccccCchhHHHHHhhhCc
Q 025018 123 SDPPETVIHRAMYLLQNGFGNYNVFQNNCEDFALYCRTGL 162 (259)
Q Consensus 123 ~~p~eeVV~RA~~~L~~G~g~YnL~~NNCEHFA~~CktGl 162 (259)
...=|+.|++|...- +.+.||||.+||.+|+..|..-.
T Consensus 76 ~~~wD~Av~~a~~~y--~~r~yNlf~~NCHSfVA~aLN~m 113 (136)
T PF05608_consen 76 AESWDDAVQKASEEY--KHRMYNLFTDNCHSFVANALNRM 113 (136)
T ss_pred HHHHHHHHHHHHHHH--hhCceeeeccCcHHHHHHHHHhc
Confidence 345678899998877 34699999999999999998844
No 14
>PF05903 Peptidase_C97: PPPDE putative peptidase domain; InterPro: IPR008580 This domain consists of the N-terminal portion of several eukaryotic sequences. The function of this domain is unknown.; PDB: 2WP7_A 3EBQ_A.
Probab=95.91 E-value=0.01 Score=50.30 Aligned_cols=34 Identities=21% Similarity=0.411 Sum_probs=22.6
Q ss_pred HHHHHHHHHHhhcCC--cccccccCchhHHHHHhhh
Q 025018 127 ETVIHRAMYLLQNGF--GNYNVFQNNCEDFALYCRT 160 (259)
Q Consensus 127 eeVV~RA~~~L~~G~--g~YnL~~NNCEHFA~~Ckt 160 (259)
++-+++....|+..+ ..|||+.+||-||+...-.
T Consensus 84 ~~~~~~~l~~l~~~~~~~~Y~Ll~~NCNhFs~~l~~ 119 (151)
T PF05903_consen 84 EEEFEEILRSLSREFTGDSYHLLNRNCNHFSDALCQ 119 (151)
T ss_dssp HHHHHHHHHHHHTT-SGGG-BTTTBSHHHHHHHHHH
T ss_pred HHHHHHHHHHHHhhccCCcchhhhhhhhHHHHHHHH
Confidence 344555556665433 5999999999999976644
No 15
>PF05382 Amidase_5: Bacteriophage peptidoglycan hydrolase ; InterPro: IPR008044 This entry is represented by Bacteriophage SFi21, lysin (Cell wall hydrolase; 3.5.1.28 from EC). At least one of proteins in this entry, the Pal protein from the pneumococcal bacteriophage Dp-1 (O03979 from SWISSPROT) has been shown to be an N-acetylmuramoyl-L-alanine amidase []. According to the known modular structure of this and other peptidoglycan hydrolases from the pneumococcal system, the active site should reside within this domain while a C-terminal domain binds to the choline residues of the cell wall teichoic acids [, ].
Probab=92.80 E-value=0.17 Score=43.11 Aligned_cols=39 Identities=23% Similarity=0.376 Sum_probs=31.0
Q ss_pred ccCCC---CCCCCCEEEEeecC---cccceEEEEEcCCEEEEeCC
Q 025018 7 RVERN---EIKAGDHIYTYRAV---FAYSHHGIYVGGSKVVHFRP 45 (259)
Q Consensus 7 ~v~~~---~lk~GD~I~~~r~~---~~y~H~GIYvG~g~VIH~~~ 45 (259)
+|+.. ++|+||++...+.+ ..+-|.||+++..++||..-
T Consensus 67 ~I~~~~~~~~q~GDI~I~g~~g~S~G~~GHtgif~~~~~iIhc~y 111 (145)
T PF05382_consen 67 KISENVDWNLQRGDIFIWGRRGNSAGAGGHTGIFMDNDTIIHCNY 111 (145)
T ss_pred EeccCCcccccCCCEEEEcCCCCCCCCCCeEEEEeCCCcEEEecC
Confidence 45544 89999999875532 24789999999999999985
No 16
>KOG0324 consensus Uncharacterized conserved protein [Function unknown]
Probab=92.44 E-value=0.076 Score=47.93 Aligned_cols=34 Identities=21% Similarity=0.444 Sum_probs=26.1
Q ss_pred CHHHHHHHHHHHhhcCC--cccccccCchhHHHHHh
Q 025018 125 PPETVIHRAMYLLQNGF--GNYNVFQNNCEDFALYC 158 (259)
Q Consensus 125 p~eeVV~RA~~~L~~G~--g~YnL~~NNCEHFA~~C 158 (259)
-+++.+++-+..|.+.+ ..|||+.+||-||+.-.
T Consensus 85 ~~~~~v~~~le~L~~ey~G~~YhL~~kNCNHFsn~l 120 (214)
T KOG0324|consen 85 LTEDDVRRILEELSEEYRGNSYHLLTKNCNHFSNEL 120 (214)
T ss_pred CCHHHHHHHHHHHHhhcCCceehhhhhccchhHHHH
Confidence 45667888887776533 39999999999998654
No 17
>COG3863 Uncharacterized distant relative of cell wall-associated hydrolases [Function unknown]
Probab=92.39 E-value=0.26 Score=44.29 Aligned_cols=41 Identities=27% Similarity=0.556 Sum_probs=30.9
Q ss_pred cccCCCCCCCCCEEEEe---ecC------------cccceEEEEEcCCEEEEeCCC
Q 025018 6 NRVERNEIKAGDHIYTY---RAV------------FAYSHHGIYVGGSKVVHFRPE 46 (259)
Q Consensus 6 ~~v~~~~lk~GD~I~~~---r~~------------~~y~H~GIYvG~g~VIH~~~~ 46 (259)
++.++.-++|||.++.. |.+ ..|-|.|+|.|.++++...+.
T Consensus 72 ~~~dr~v~~~gd~~~gdyPTr~g~i~~t~~~~~~~~H~gHagmy~~a~~~VEs~ps 127 (231)
T COG3863 72 NNLDRSVLQPGDILLGDYPTRGGAIWLTDTFGNIVGHWGHAGMYIGAGQMVESWPS 127 (231)
T ss_pred hhhhhhhcCCcchhhccCCCCcceEEEEcccccccccccceEEEEcCCcEEeeccC
Confidence 45678889999998872 111 136788999999999988775
No 18
>PF06672 DUF1175: Protein of unknown function (DUF1175); InterPro: IPR009558 This family consists of several hypothetical bacterial proteins of around 210 residues in length. The function of this family is unknown.
Probab=89.55 E-value=0.78 Score=41.58 Aligned_cols=37 Identities=24% Similarity=0.287 Sum_probs=26.6
Q ss_pred CCCCCCCCCEEEEeecCcc-cceEEEEEcC----CEEEEeCC
Q 025018 9 ERNEIKAGDHIYTYRAVFA-YSHHGIYVGG----SKVVHFRP 45 (259)
Q Consensus 9 ~~~~lk~GD~I~~~r~~~~-y~H~GIYvG~----g~VIH~~~ 45 (259)
+.+..+|||+|++...... ..|.-||+|+ .-|-|-.+
T Consensus 132 dl~~A~pGDL~Ff~~~d~~~pfHlMI~~g~~~~~~ivYHTG~ 173 (216)
T PF06672_consen 132 DLEQARPGDLLFFHQGDDQMPFHLMIWVGRDAPPWIVYHTGP 173 (216)
T ss_pred hhhhcCCCcEEEecCCCCCcceEEEEEEcCCcceEEEEecCC
Confidence 3678999999988765421 3499999998 55555443
No 19
>PF05257 CHAP: CHAP domain; InterPro: IPR007921 The CHAP (cysteine, histidine-dependent amidohydrolases/peptidases) domain is a region between 110 and 140 amino acids that is found in proteins from bacteria, bacteriophages, archaea and eukaryotes of the Trypanosomidae family. Many of these proteins are uncharacterised, but it has been proposed that they may function mainly in peptidoglycan hydrolysis. The CHAP domain is found in a wide range of protein architectures; it is commonly associated with bacterial type SH3 domains and with several families of amidase domains. It has been suggested that CHAP domain containing proteins utilise a catalytic cysteine residue in a nucleophilic-attack mechanism [, ]. The CHAP domain contains two invariant residues, a cysteine and a histidine. These residues form part of the putative active site of CHAP domain containing proteins. Secondary structure predictions show that the CHAP domain belongs to the alpha + beta structural class, with the N-terminal half largely containing predicted alpha helices and the C-terminal half principally composed of predicted beta strands [, ]. Some proteins known to contain a CHAP domain are listed below: Bacterial and trypanosomal glutathionylspermidine amidases. A variety of bacterial autolysins. A Nocardia aerocolonigenes putative esterase. Streptococcus pneumoniae choline-binding protein D. Methanosarcina mazei protein MM2478, a putative chloride channel. Several phage-encoded peptidoglycan hydrolases. Cysteine peptidases belonging to MEROPS peptidase family C51 (D-alanyl-glycyl endopeptidase, clan CA). ; PDB: 2LRJ_A 2VPM_B 2VOB_B 2VPS_A 2K3A_A 2IO9_A 2IO8_A 2IOB_A 2IOA_B 2IO7_B ....
Probab=77.96 E-value=2.5 Score=33.76 Aligned_cols=29 Identities=17% Similarity=0.078 Sum_probs=18.6
Q ss_pred CCCCCCCCCEEEEe-ecCcccceEEEEEcC
Q 025018 9 ERNEIKAGDHIYTY-RAVFAYSHHGIYVGG 37 (259)
Q Consensus 9 ~~~~lk~GD~I~~~-r~~~~y~H~GIYvG~ 37 (259)
....++|||++.+. .....|-|.||..+-
T Consensus 59 ~~~~P~~Gdivv~~~~~~~~~GHVaIV~~v 88 (124)
T PF05257_consen 59 TGSTPQPGDIVVWDSGSGGGYGHVAIVESV 88 (124)
T ss_dssp ECS---TTEEEEEEECTTTTT-EEEEEEEE
T ss_pred cCcccccceEEEeccCCCCCCCeEEEEEEE
Confidence 45789999999993 333458999999763
No 20
>KOG3150 consensus Uncharacterized conserved protein [Function unknown]
Probab=75.23 E-value=10 Score=33.27 Aligned_cols=37 Identities=14% Similarity=0.203 Sum_probs=30.5
Q ss_pred CCCCHHHHHHHHHHHhhcCCcccccccCchhHHHHHhhh
Q 025018 122 TSDPPETVIHRAMYLLQNGFGNYNVFQNNCEDFALYCRT 160 (259)
Q Consensus 122 ~~~p~eeVV~RA~~~L~~G~g~YnL~~NNCEHFA~~Ckt 160 (259)
.+..-|+.|+.|...- +.+.|||+..||+-|+.-|..
T Consensus 91 g~~~wD~Av~~as~~y--~hr~hNi~cdNCHShVA~aLn 127 (182)
T KOG3150|consen 91 GARTWDNAVSKASREY--KHRTHNIFCDNCHSHVANALN 127 (182)
T ss_pred CCchHHHHHHHHHHHh--hhcccceeeccHHHHHHHHHH
Confidence 4556788899988877 457999999999999988765
No 21
>PF06940 DUF1287: Domain of unknown function (DUF1287); InterPro: IPR009706 This family consists of several hypothetical bacterial proteins of around 200 residues in length. The function of this family is unknown.
Probab=73.93 E-value=7 Score=34.15 Aligned_cols=41 Identities=17% Similarity=0.183 Sum_probs=30.2
Q ss_pred cccCCCCCCCCCEEEEeecCcccceEEEEEc----CC--EEEEeCCCC
Q 025018 6 NRVERNEIKAGDHIYTYRAVFAYSHHGIYVG----GS--KVVHFRPER 47 (259)
Q Consensus 6 ~~v~~~~lk~GD~I~~~r~~~~y~H~GIYvG----~g--~VIH~~~~~ 47 (259)
..+..++.+|||+|.+...+ .-.|.||... +| .|||..+..
T Consensus 100 ~~~~~~~~q~GDIVtw~l~~-~~~HIgIVSd~r~~~G~p~viHNiG~g 146 (164)
T PF06940_consen 100 TDINPEDWQPGDIVTWRLPG-GLPHIGIVSDRRSKDGVPLVIHNIGPG 146 (164)
T ss_pred CCCChhhcCCCCEEEEeCCC-CCCeEEEEeCCcCCCCCEEEEEecCCC
Confidence 34455899999999764333 3689999985 34 899998865
No 22
>PF10030 DUF2272: Uncharacterized protein conserved in bacteria (DUF2272); InterPro: IPR019262 This is a domain of unknown function found in proteins of unknown function.
Probab=62.59 E-value=15 Score=32.55 Aligned_cols=45 Identities=20% Similarity=0.101 Sum_probs=32.9
Q ss_pred CCCcccCCCCCCCCCEEEEeecC-------------cccceEEEEEc----CCEEEEeCCCC
Q 025018 3 LLTNRVERNEIKAGDHIYTYRAV-------------FAYSHHGIYVG----GSKVVHFRPER 47 (259)
Q Consensus 3 ~~~~~v~~~~lk~GD~I~~~r~~-------------~~y~H~GIYvG----~g~VIH~~~~~ 47 (259)
++..+.....+++||+|...|.. ..-+|.+|.|. ++..+..-++.
T Consensus 84 ~~~~~~~~y~P~~GDlIc~~R~~~~~~~~~~~~~~~~~~~HcdIVVa~~~~d~~~v~~IGGN 145 (183)
T PF10030_consen 84 FRARDPAEYKPRPGDLICYDRGRSKTYDFASLPTSGGFPSHCDIVVAVNVVDGRTVTTIGGN 145 (183)
T ss_pred ccccCcCCCCCCCCCEEEecCCCCcccchhhhccCCCCCCceeEEEeeccCCCCEEEEEcCc
Confidence 34566778899999999998854 13589999987 44666666544
No 23
>PF03658 Ub-RnfH: RnfH family Ubiquitin; InterPro: IPR005346 This is a small family of proteins of unknown function.; PDB: 2HJ1_B.
Probab=60.91 E-value=3.7 Score=32.06 Aligned_cols=25 Identities=24% Similarity=0.599 Sum_probs=13.9
Q ss_pred CCCCCcccC-CCCCCCCCEEEEeecC
Q 025018 1 MGLLTNRVE-RNEIKAGDHIYTYRAV 25 (259)
Q Consensus 1 mg~~~~~v~-~~~lk~GD~I~~~r~~ 25 (259)
+|+||+.+. ...|+.||-|+++|..
T Consensus 49 vGIfGk~~~~d~~L~~GDRVEIYRPL 74 (84)
T PF03658_consen 49 VGIFGKLVKLDTVLRDGDRVEIYRPL 74 (84)
T ss_dssp EEEEE-S--TT-B--TT-EEEEE-S-
T ss_pred eeeeeeEcCCCCcCCCCCEEEEeccC
Confidence 588999884 4679999999999975
No 24
>COG3738 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=55.31 E-value=24 Score=31.28 Aligned_cols=38 Identities=24% Similarity=0.407 Sum_probs=28.4
Q ss_pred CCCCCCCCCEEEEeecCcccceEEEEEcC----C--EEEEeCCCC
Q 025018 9 ERNEIKAGDHIYTYRAVFAYSHHGIYVGG----S--KVVHFRPER 47 (259)
Q Consensus 9 ~~~~lk~GD~I~~~r~~~~y~H~GIYvG~----g--~VIH~~~~~ 47 (259)
+.+..+|||+| .||....-.|.||...+ | .|||.-+..
T Consensus 137 ~~s~y~aGDIv-sWRLdngl~HiGv~sd~~~~~g~plViHNIGaG 180 (200)
T COG3738 137 DPSDYQAGDIV-SWRLDNGLAHIGVVSDGFTRDGTPLVIHNIGAG 180 (200)
T ss_pred CccccCCCceE-EEEcCCCCceeEEEecCCCCCCCeEEEeecCCC
Confidence 45788999998 46754447899998752 3 889988754
No 25
>COG3234 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=48.02 E-value=23 Score=31.68 Aligned_cols=33 Identities=24% Similarity=0.287 Sum_probs=25.3
Q ss_pred CCCCCCCCCEEEEeecCcccceEEEEEcCCEEEEe
Q 025018 9 ERNEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVHF 43 (259)
Q Consensus 9 ~~~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH~ 43 (259)
+.++..|||+++|..+- -+|--|++|.=-+.|-
T Consensus 138 dvnqAlPGDl~ffdqgd--dqHLMIwmgr~i~YHT 170 (215)
T COG3234 138 DVNQALPGDLIFFDQGD--DQHLMIWMGRYIAYHT 170 (215)
T ss_pred hhhhhCCCcEEEEecCC--ceEEEEEecceEEEec
Confidence 34678899999998876 5899999994444443
No 26
>PF07313 DUF1460: Protein of unknown function (DUF1460); InterPro: IPR010846 This family consists of several hypothetical bacterial proteins of around 260 residues in length. The function of this family is unknown.; PDB: 2P1G_B 2IM9_A.
Probab=46.02 E-value=34 Score=31.01 Aligned_cols=37 Identities=30% Similarity=0.291 Sum_probs=24.7
Q ss_pred CCCCCCCEEEEeec--CcccceEEEEEc--CC-EEEEeCCCC
Q 025018 11 NEIKAGDHIYTYRA--VFAYSHHGIYVG--GS-KVVHFRPER 47 (259)
Q Consensus 11 ~~lk~GD~I~~~r~--~~~y~H~GIYvG--~g-~VIH~~~~~ 47 (259)
++++.||+|-+... +-..+|.||.+= ++ .+.|+++..
T Consensus 152 ~~i~~GDiI~i~t~~~GLDvsH~Giav~~~~~l~l~hASs~~ 193 (216)
T PF07313_consen 152 SQIKNGDIIAIVTNIKGLDVSHVGIAVWKNDGLHLRHASSLH 193 (216)
T ss_dssp TTS-TT-EEEEEEECTTECEEEEEEEEEETTEEEEEEEETTT
T ss_pred hcCCCCCEEEEEeCCCCCceeeEEEEEEECCeEEEEeCCCCC
Confidence 78999999998763 345899998884 33 446666544
No 27
>PF01052 SpoA: Surface presentation of antigens (SPOA); InterPro: IPR001543 Proteins in this group are involved in a secretory pathway responsible for the surface presentation of invasion plasmid antigen needed for the entry of Salmonella and other species into mammalian cells [, ].They could play a role in preserving the translocation competence of the IPA antigens and are required for secretion of the three IPA proteins []. The C-terminal region of flagellar motor switch proteins FliN and FliM is also included in this entry. ; PDB: 3UEP_A 1O9Y_B 1YAB_A.
Probab=43.44 E-value=34 Score=24.99 Aligned_cols=31 Identities=19% Similarity=0.177 Sum_probs=21.6
Q ss_pred CCCCCCCEEEEeecCcccceEEEEEcCCEEEEe
Q 025018 11 NEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVHF 43 (259)
Q Consensus 11 ~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH~ 43 (259)
.++++||+|.+.... ..+.-+|+++-.+.+.
T Consensus 27 ~~L~~Gdvi~l~~~~--~~~v~l~v~g~~~~~g 57 (77)
T PF01052_consen 27 LNLKVGDVIPLDKPA--DEPVELRVNGQPIFRG 57 (77)
T ss_dssp HC--TT-EEEECCES--STEEEEEETTEEEEEE
T ss_pred hcCCCCCEEEeCCCC--CCCEEEEECCEEEEEE
Confidence 579999999998875 6899999976555443
No 28
>KOG4577 consensus Transcription factor LIM3, contains LIM and HOX domains [Transcription]
Probab=43.40 E-value=9.1 Score=36.58 Aligned_cols=55 Identities=33% Similarity=0.727 Sum_probs=36.7
Q ss_pred ccCCCCcCccCCCCceEEccchhhcCCCceEEEeeccCcceeeehccCCcccc-cCCCCHHHHHHHHHHHh
Q 025018 68 CLIFPDCGFRQPNSGVILSCLDCFLGNGSLYCFEYGVAPSVFLAKVRGGTCTT-ATSDPPETVIHRAMYLL 137 (259)
Q Consensus 68 ~~~~~~cg~~~~~~gVv~s~L~~Fl~G~~l~~f~Y~vs~~~flak~rggtC~~-~~~~p~eeVV~RA~~~L 137 (259)
|..|.+|-.+... .||+.++.+|..+ -|+-+ -|..|+. ..-.||.+||+||...+
T Consensus 60 CLkCs~C~~qL~d--------rCFsR~~s~yCke------dFfKr-fGTKCsaC~~GIpPtqVVRkAqd~V 115 (383)
T KOG4577|consen 60 CLKCSDCHDQLAD--------RCFSREGSVYCKE------DFFKR-FGTKCSACQEGIPPTQVVRKAQDFV 115 (383)
T ss_pred hcchhhhhhHHHH--------HHhhcCCceeehH------HHHHH-hCCcchhhcCCCChHHHHHHhhcce
Confidence 5666777555443 5899999999732 23322 3556644 44589999999998554
No 29
>PF05820 DUF845: Baculovirus protein of unknown function (DUF845); InterPro: IPR008563 This entry is represented by Autographa californica nuclear polyhedrosis virus (AcMNPV), Orf81; it is a family of uncharacterised viral proteins.
Probab=42.45 E-value=20 Score=29.80 Aligned_cols=24 Identities=25% Similarity=0.477 Sum_probs=16.6
Q ss_pred HHHHHHHhhcCCcccccccCchhH
Q 025018 130 IHRAMYLLQNGFGNYNVFQNNCED 153 (259)
Q Consensus 130 V~RA~~~L~~G~g~YnL~~NNCEH 153 (259)
.++-...--+|+..+|+.++|||-
T Consensus 91 ck~eL~~~vegEn~FNiaf~NCEs 114 (119)
T PF05820_consen 91 CKEELRKFVEGENNFNIAFQNCES 114 (119)
T ss_pred HHHHHHHHHhccccceeeeccchh
Confidence 333333333588899999999995
No 30
>cd04482 RPA2_OBF_like RPA2_OBF_like: A subgroup of uncharacterized archaeal OB folds with similarity to the OB fold of the central ssDNA-binding domain (DBD)-D of human RPA2 (also called RPA32). RPA2 is a subunit of Replication protein A (RPA). RPA is a nuclear ssDNA-binding protein (SSB) which appears to be involved in all aspects of DNA metabolism including replication, recombination, and repair. RPA also mediates specific interactions of various nuclear proteins. In animals, plants, and fungi, RPA is a heterotrimer with subunits of 70KDa (RPA1), 32kDa (RPA2), and 14 KDa (RPA3). The major DNA binding activity of RPA is associated with RPA1 DBD-A and DBD-B; RPA2 DBD-D is a weak ssDNA-binding domain. RPA2 DBD-D is also involved in trimerization. The ssDNA binding mechanism is believed to be multistep and to involve conformational change. N-terminal to human RPA2 DBD-D is a domain containing all the known phosphorylation sites of RPA. Human RPA2 is phosphorylated in a cell cycle depende
Probab=37.00 E-value=78 Score=24.37 Aligned_cols=8 Identities=38% Similarity=1.016 Sum_probs=6.2
Q ss_pred ccCCCCcC
Q 025018 68 CLIFPDCG 75 (259)
Q Consensus 68 ~~~~~~cg 75 (259)
.+.||+|+
T Consensus 84 np~C~~C~ 91 (91)
T cd04482 84 NPVCPKCG 91 (91)
T ss_pred CCcCCCCC
Confidence 56799985
No 31
>PF08007 Cupin_4: Cupin superfamily protein; InterPro: IPR022777 This signature represents primarily the cupin fold found in JmjC transcription factors. The fold is also found in lysine-specific demethylase NO66.; PDB: 2XDV_A 1VRB_B 4DIQ_B.
Probab=36.88 E-value=25 Score=32.98 Aligned_cols=33 Identities=24% Similarity=0.381 Sum_probs=21.0
Q ss_pred cCCCCCCCCCEEEEeecCcccceEEEEEcCCEEEEeCC
Q 025018 8 VERNEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVHFRP 45 (259)
Q Consensus 8 v~~~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH~~~ 45 (259)
+..-.|+|||++|+.|+- .|.+.=.+ .=+|++-
T Consensus 175 ~~~~~L~pGD~LYlPrG~---~H~~~~~~--~S~hltv 207 (319)
T PF08007_consen 175 VEEVVLEPGDVLYLPRGW---WHQAVTTD--PSLHLTV 207 (319)
T ss_dssp SEEEEE-TT-EEEE-TT----EEEEEESS---EEEEEE
T ss_pred eEEEEECCCCEEEECCCc---cCCCCCCC--CceEEEE
Confidence 334568999999998875 79998877 4456653
No 32
>TIGR02480 fliN flagellar motor switch protein FliN. Proteins that consist largely of the domain described by this model can be designated flagellar motor switch protein FliN. Longer proteins in which this region is a C-terminal domain typically are designated FliY. More distantly related sequences, outside the scope of this family, are associated with type III secretion and include the surface presentation of antigens protein SpaO required or invasion of host cells by Salmonella enterica.
Probab=35.44 E-value=56 Score=24.36 Aligned_cols=31 Identities=16% Similarity=0.118 Sum_probs=24.1
Q ss_pred CCCCCCCCEEEEeecCcccceEEEEEcCCEEEE
Q 025018 10 RNEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVH 42 (259)
Q Consensus 10 ~~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH 42 (259)
..++++||+|.+.+.. ..+.-||+++-.+..
T Consensus 26 ll~L~~Gdvi~L~~~~--~~~v~l~v~g~~~~~ 56 (77)
T TIGR02480 26 LLKLGEGSVIELDKLA--GEPLDILVNGRLIAR 56 (77)
T ss_pred HhcCCCCCEEEcCCCC--CCcEEEEECCEEEEE
Confidence 3579999999998755 578999998765543
No 33
>PF13387 DUF4105: Domain of unknown function (DUF4105)
Probab=35.11 E-value=38 Score=28.97 Aligned_cols=15 Identities=33% Similarity=0.585 Sum_probs=11.7
Q ss_pred cccccccCchhHHHH
Q 025018 142 GNYNVFQNNCEDFAL 156 (259)
Q Consensus 142 g~YnL~~NNCEHFA~ 156 (259)
-.|+.+.+||=--..
T Consensus 128 ~~Y~f~~~NCat~i~ 142 (176)
T PF13387_consen 128 YRYNFFTDNCATRIR 142 (176)
T ss_pred eeehhhhcchHHHHH
Confidence 489999999965443
No 34
>COG2850 Uncharacterized conserved protein [Function unknown]
Probab=34.68 E-value=15 Score=35.95 Aligned_cols=34 Identities=15% Similarity=0.339 Sum_probs=25.1
Q ss_pred cccCCCCCCCCCEEEEeecCcccceEEEEEcCCEEEEeC
Q 025018 6 NRVERNEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVHFR 44 (259)
Q Consensus 6 ~~v~~~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH~~ 44 (259)
..+....+.|||++|+... +.|+||-.++. .||+
T Consensus 176 ~~~~d~vlepGDiLYiPp~---~~H~gvae~dc--~tyS 209 (383)
T COG2850 176 EPDIDEVLEPGDILYIPPG---FPHYGVAEDDC--MTYS 209 (383)
T ss_pred CchhhhhcCCCceeecCCC---CCcCCcccccc--ccee
Confidence 3456678999999999765 48999987543 4444
No 35
>PF11730 DUF3297: Protein of unknown function (DUF3297); InterPro: IPR021724 This family is expressed in Proteobacteria and Actinobacteria. The function is not known.
Probab=31.96 E-value=18 Score=27.41 Aligned_cols=44 Identities=20% Similarity=0.422 Sum_probs=30.7
Q ss_pred hhhhhhcccc-ccceeeehhhhhhhcCcccccccchhhccccCccc
Q 025018 213 SRYATDIGVR-SDVIKVAVEDLAVNLGWLSRHEETSEENKSSNQLI 257 (259)
Q Consensus 213 ~r~~~dig~r-~d~~kv~~e~l~~~~~~~~~~~~~~~~~~~~~~~~ 257 (259)
.-+..|||+| +++-|-.||+--..-||-. +...++...+-+|++
T Consensus 16 ~~l~~~iGIrfng~Er~nVeEYciSEGWvr-v~~gka~DR~G~Pl~ 60 (71)
T PF11730_consen 16 EVLERGIGIRFNGKERTNVEEYCISEGWVR-VAAGKALDRRGNPLT 60 (71)
T ss_pred HHHhcCcceEECCeEcccceeEeccCCEEE-eecCcccccCCCeeE
Confidence 4467899999 7888999999888889966 333344344444443
No 36
>KOG3416 consensus Predicted nucleic acid binding protein [General function prediction only]
Probab=29.73 E-value=43 Score=28.33 Aligned_cols=13 Identities=31% Similarity=0.215 Sum_probs=9.9
Q ss_pred CCCCCCCEEEEee
Q 025018 11 NEIKAGDHIYTYR 23 (259)
Q Consensus 11 ~~lk~GD~I~~~r 23 (259)
.-++|||+|.+.+
T Consensus 60 ~~~~PGDIirLt~ 72 (134)
T KOG3416|consen 60 CLIQPGDIIRLTG 72 (134)
T ss_pred cccCCccEEEecc
Confidence 4578999998754
No 37
>KOG3706 consensus Uncharacterized conserved protein [Function unknown]
Probab=29.70 E-value=27 Score=35.70 Aligned_cols=39 Identities=21% Similarity=0.214 Sum_probs=26.8
Q ss_pred CCcccCCCCCCCCCEEEEeecCcccceEEEEEcCCEEEEeCC
Q 025018 4 LTNRVERNEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVHFRP 45 (259)
Q Consensus 4 ~~~~v~~~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH~~~ 45 (259)
+|.+|-..-|+|||+|||.|+. -|-++--..-.=.|.+-
T Consensus 376 lgePV~e~vle~GDllYfPRG~---IHQA~t~~~vHSlHvTl 414 (629)
T KOG3706|consen 376 LGEPVHEFVLEPGDLLYFPRGT---IHQADTPALVHSLHVTL 414 (629)
T ss_pred hCCchHHhhcCCCcEEEecCcc---eeeccccchhceeEEEe
Confidence 4577888889999999999975 46665533333355543
No 38
>COG2914 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=27.76 E-value=43 Score=27.00 Aligned_cols=25 Identities=28% Similarity=0.667 Sum_probs=20.2
Q ss_pred CCCCCcccCC-CCCCCCCEEEEeecC
Q 025018 1 MGLLTNRVER-NEIKAGDHIYTYRAV 25 (259)
Q Consensus 1 mg~~~~~v~~-~~lk~GD~I~~~r~~ 25 (259)
.|++|+++.. .+++-||-|+++|..
T Consensus 52 ~GI~~k~~kl~~~l~dgDRVEIyRPL 77 (99)
T COG2914 52 VGIYSKPVKLDDELHDGDRVEIYRPL 77 (99)
T ss_pred eeEEccccCccccccCCCEEEEeccc
Confidence 3788888744 668999999999976
No 39
>PRK11032 hypothetical protein; Provisional
Probab=26.69 E-value=51 Score=28.64 Aligned_cols=9 Identities=33% Similarity=0.944 Sum_probs=6.4
Q ss_pred cCCCCcCcc
Q 025018 69 LIFPDCGFR 77 (259)
Q Consensus 69 ~~~~~cg~~ 77 (259)
++||.|+..
T Consensus 143 ~pCp~C~~~ 151 (160)
T PRK11032 143 PLCPKCGHD 151 (160)
T ss_pred CCCCCCCCC
Confidence 478888653
No 40
>PF06887 DUF1265: Protein of unknown function (DUF1265); InterPro: IPR009676 This family represents a conserved region approximately 50 residues long within a number of proteins of unknown function that seem to be restricted to Caenorhabditis elegans.
Probab=25.06 E-value=61 Score=22.87 Aligned_cols=36 Identities=28% Similarity=0.460 Sum_probs=25.4
Q ss_pred ccCchhHHHHHhhhCccccccCCcccchhhhHHhhhhHHHHhh
Q 025018 147 FQNNCEDFALYCRTGLLIVDRQGVGSSGQASSVIGAPLAAILS 189 (259)
Q Consensus 147 ~~NNCEHFA~~CktGl~~~~~~g~grSgQa~s~l~~~l~a~~s 189 (259)
+..|||+|..-|.- +. +..-.|..++..+..|.+++
T Consensus 3 L~kN~EDl~YV~nm-Li------vA~d~~f~~v~~~C~Atii~ 38 (48)
T PF06887_consen 3 LVKNHEDLMYVCNM-LI------VAHDARFGNVQNCCIATIIS 38 (48)
T ss_pred HHHhhhhHHHHHhH-he------eeccccchHHHHHHHHHHHH
Confidence 45799999988865 32 34556777777777777664
No 41
>PF00122 E1-E2_ATPase: E1-E2 ATPase p-type cation-transporting ATPase superfamily signature H+-transporting ATPase (proton pump) signature sodium/potassium-transporting ATPase signature; InterPro: IPR008250 ATPases (or ATP synthases) are membrane-bound enzyme complexes/ion transporters that combine ATP synthesis and/or hydrolysis with the transport of protons across a membrane. ATPases can harness the energy from a proton gradient, using the flux of ions across the membrane via the ATPase proton channel to drive the synthesis of ATP. Some ATPases work in reverse, using the energy from the hydrolysis of ATP to create a proton gradient. There are different types of ATPases, which can differ in function (ATP synthesis and/or hydrolysis), structure (e.g., F-, V- and A-ATPases, which contain rotary motors) and in the type of ions they transport [, ]. The different types include: F-ATPases (F1F0-ATPases), which are found in mitochondria, chloroplasts and bacterial plasma membranes where they are the prime producers of ATP, using the proton gradient generated by oxidative phosphorylation (mitochondria) or photosynthesis (chloroplasts). V-ATPases (V1V0-ATPases), which are primarily found in eukaryotic vacuoles and catalyse ATP hydrolysis to transport solutes and lower pH in organelles. A-ATPases (A1A0-ATPases), which are found in Archaea and function like F-ATPases (though with respect to their structure and some inhibitor responses, A-ATPases are more closely related to the V-ATPases). P-ATPases (E1E2-ATPases), which are found in bacteria and in eukaryotic plasma membranes and organelles, and function to transport a variety of different ions across membranes. E-ATPases, which are cell-surface enzymes that hydrolyse a range of NTPs, including extracellular ATP. P-ATPases (sometime known as E1-E2 ATPases) (3.6.3.- from EC) are found in bacteria and in a number of eukaryotic plasma membranes and organelles []. P-ATPases function to transport a variety of different compounds, including ions and phospholipids, across a membrane using ATP hydrolysis for energy. There are many different classes of P-ATPases, each of which transports a specific type of ion: H+, Na+, K+, Mg2+, Ca2+, Ag+ and Ag2+, Zn2+, Co2+, Pb2+, Ni2+, Cd2+, Cu+ and Cu2+. P-ATPases can be composed of one or two polypeptides, and can usually assume two main conformations called E1 and E2. This entry represents the actuator (A) domain, and some transmembrane helices found in P-type ATPases []. It contains the TGES-loop which is essential for the metal ion binding which results in tight association between the A and P (phosphorylation) domains []. It does not contain the phosphorylation site. It is thought that the large movement of the actuator domain, which is transmitted to the transmembrane helices, is essential to the long distance coupling between formation/decomposition of the acyl phosphate in the cytoplasmic P-domain and the changes in the ion-binding sites buried deep in the membranous region []. This domain has a modulatory effect on the phosphoenzyme processing steps through its nucleotide binding [],[]. P-type (or E1-E2-type) ATPases that form an aspartyl phosphate intermediate in the course of ATP hydrolysis, can be divided into 4 major groups []: (1) Ca2+-transporting ATPases; (2) Na+/K+- and gastric H+/K+-transporting ATPases; (3) plasma membrane H+-transporting ATPases (proton pumps) of plants, fungi and lower eukaryotes; and (4) all bacterial P-type ATPases, except the g2+-ATPase of Salmonella typhimurium, which is more similar to the eukaryotic sequences. However, great variety of sequence analysis methods results in diversity of classification. More information about this protein can be found at Protein of the Month: ATP Synthases [].; GO: 0000166 nucleotide binding, 0046872 metal ion binding; PDB: 2XZB_A 1MHS_B 3TLM_A 3A3Y_A 2ZXE_A 3NAL_A 3NAM_A 3NAN_A 2YJ6_B 2IYE_A ....
Probab=24.91 E-value=37 Score=29.32 Aligned_cols=19 Identities=21% Similarity=0.300 Sum_probs=15.2
Q ss_pred ccCCCCCCCCCEEEEeecC
Q 025018 7 RVERNEIKAGDHIYTYRAV 25 (259)
Q Consensus 7 ~v~~~~lk~GD~I~~~r~~ 25 (259)
+++.++++|||+|.+..+.
T Consensus 46 ~i~~~~L~~GDiI~l~~g~ 64 (230)
T PF00122_consen 46 KIPSSELVPGDIIILKAGD 64 (230)
T ss_dssp EEEGGGT-TTSEEEEETTE
T ss_pred cchHhhccceeeeeccccc
Confidence 6788999999999997654
No 42
>PRK06033 hypothetical protein; Validated
Probab=24.38 E-value=1.2e+02 Score=23.26 Aligned_cols=31 Identities=10% Similarity=-0.075 Sum_probs=23.5
Q ss_pred CCCCCCCEEEEeecCcccceEEEEEcCCEEEEe
Q 025018 11 NEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVHF 43 (259)
Q Consensus 11 ~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH~ 43 (259)
-++++||+|...+.. -...-+|+++-.+...
T Consensus 26 L~L~~GDVI~L~~~~--~~~v~v~V~~~~~f~g 56 (83)
T PRK06033 26 LRMGRGAVIPLDATE--ADEVWILANNHPIARG 56 (83)
T ss_pred hCCCCCCEEEeCCCC--CCcEEEEECCEEEEEE
Confidence 579999999997754 4678899987655443
No 43
>PF10077 DUF2314: Uncharacterized protein conserved in bacteria (DUF2314); InterPro: IPR018756 This domain of unkown function is found in various bacterial hypothetical proteins, as well as putative ankyrin repeat proteins.
Probab=24.17 E-value=75 Score=26.43 Aligned_cols=35 Identities=29% Similarity=0.319 Sum_probs=27.2
Q ss_pred CCCCc-ccCCCCCCCCCEEEEeecCcccceEEEEEcCC
Q 025018 2 GLLTN-RVERNEIKAGDHIYTYRAVFAYSHHGIYVGGS 38 (259)
Q Consensus 2 g~~~~-~v~~~~lk~GD~I~~~r~~~~y~H~GIYvG~g 38 (259)
|.|.| |.....++.||.|.+.... .+=|-+|.++.
T Consensus 68 G~L~N~P~~i~~v~~Gd~v~~~~~~--IsDWm~~~~g~ 103 (133)
T PF10077_consen 68 GVLDNEPYYITNVKEGDRVSFPIED--ISDWMIYEDGR 103 (133)
T ss_pred EEEecCCcccCCCCCCCEEEEChHH--eeEeEEEECCc
Confidence 44555 7788999999999998876 68888887544
No 44
>PRK03187 tgl transglutaminase; Provisional
Probab=22.69 E-value=1.1e+02 Score=28.99 Aligned_cols=29 Identities=24% Similarity=0.449 Sum_probs=21.1
Q ss_pred CCCCCCCEEEEeecCcc--cc----eEEEEEcCCE
Q 025018 11 NEIKAGDHIYTYRAVFA--YS----HHGIYVGGSK 39 (259)
Q Consensus 11 ~~lk~GD~I~~~r~~~~--y~----H~GIYvG~g~ 39 (259)
..+-|||.+||...-+. -. -..||+|+|.
T Consensus 164 ~~~~PGD~vYFkNPd~~p~tp~WqGeNaiyLgn~~ 198 (272)
T PRK03187 164 GDFLPGDCVYFKNPDFNPATPEWQGENVIYLGNGL 198 (272)
T ss_pred CCCCCCcEEEecCCCCCCCCCcccceeEEEecCCc
Confidence 67889999999765432 12 3579999984
No 45
>PF11948 DUF3465: Protein of unknown function (DUF3465); InterPro: IPR021856 This family of proteins are functionally uncharacterised. This protein is found in bacteria. Proteins in this family are typically between 131 to 151 amino acids in length. This protein has a conserved HWTH sequence motif.
Probab=22.24 E-value=76 Score=26.83 Aligned_cols=30 Identities=20% Similarity=0.366 Sum_probs=21.4
Q ss_pred CCCCCCCEEEEeecCcccceEEEEEcCCEEEEeCCCC
Q 025018 11 NEIKAGDHIYTYRAVFAYSHHGIYVGGSKVVHFRPER 47 (259)
Q Consensus 11 ~~lk~GD~I~~~r~~~~y~H~GIYvG~g~VIH~~~~~ 47 (259)
..|++||.|+|.- . | .|--.|.|||++-.+
T Consensus 84 p~l~~GD~V~f~G-e--Y----e~n~kggvIHWTH~d 113 (131)
T PF11948_consen 84 PWLQKGDQVEFYG-E--Y----EWNPKGGVIHWTHHD 113 (131)
T ss_pred cCcCCCCEEEEEE-E--E----EECCCCCEEEeeccC
Confidence 4589999999842 2 2 444578999999643
No 46
>PRK01777 hypothetical protein; Validated
Probab=22.14 E-value=60 Score=25.63 Aligned_cols=25 Identities=20% Similarity=0.598 Sum_probs=19.3
Q ss_pred CCCCCcccC-CCCCCCCCEEEEeecC
Q 025018 1 MGLLTNRVE-RNEIKAGDHIYTYRAV 25 (259)
Q Consensus 1 mg~~~~~v~-~~~lk~GD~I~~~r~~ 25 (259)
+|++|+.++ ...|+.||-|+++|..
T Consensus 52 vgI~Gk~v~~d~~L~dGDRVeIyrPL 77 (95)
T PRK01777 52 VGIYSRPAKLTDVLRDGDRVEIYRPL 77 (95)
T ss_pred EEEeCeECCCCCcCCCCCEEEEecCC
Confidence 366777664 4679999999998875
No 47
>cd05834 HDGF_related The PWWP domain is an essential part of the Hepatoma Derived Growth Factor (HDGF) family of proteins, and is necessary for DNA binding by HDGF. This family of endogenous nuclear-targeted mitogens includes HRP (HDGF-related proteins 1, 2, 3, 4, or HPR1, HPR2, HPR3, HPR4, respectively) and lens epithelium-derived growth factor, LEDGF. Members of the HDGF family have been linked to human diseases, and HDGF is a prognostic factor in several types of cancer. The PWWP domain, named for a conserved Pro-Trp-Trp-Pro motif, is a small domain consisting of 100-150 amino acids. The PWWP domain is found in numerous proteins that are involved in cell division, growth and differentiation. Most PWWP-domain proteins seem to be nuclear, often DNA-binding, proteins that function as transcription factors regulating a variety of developmental processes.
Probab=21.86 E-value=65 Score=24.59 Aligned_cols=18 Identities=28% Similarity=0.470 Sum_probs=14.8
Q ss_pred CCCCCCEEEEeecCcccceE
Q 025018 12 EIKAGDHIYTYRAVFAYSHH 31 (259)
Q Consensus 12 ~lk~GD~I~~~r~~~~y~H~ 31 (259)
++++||+|+-.-.+ |.+|
T Consensus 2 ~f~~GdlVwaK~kG--yp~W 19 (83)
T cd05834 2 QFKAGDLVFAKVKG--YPAW 19 (83)
T ss_pred CCCCCCEEEEecCC--CCCC
Confidence 68999999988777 6666
No 48
>COG0272 Lig NAD-dependent DNA ligase (contains BRCT domain type II) [DNA replication, recombination, and repair]
Probab=21.04 E-value=93 Score=32.86 Aligned_cols=19 Identities=26% Similarity=0.559 Sum_probs=16.8
Q ss_pred ccCCCCCCCCCEEEEeecC
Q 025018 7 RVERNEIKAGDHIYTYRAV 25 (259)
Q Consensus 7 ~v~~~~lk~GD~I~~~r~~ 25 (259)
.|.+..+++||.|++.|.+
T Consensus 362 ~I~rkdIrIGDtV~V~kAG 380 (667)
T COG0272 362 EIKRKDIRIGDTVVVRKAG 380 (667)
T ss_pred HHHhcCCCCCCEEEEEecC
Confidence 5678999999999999976
No 49
>PF12671 Amidase_6: Putative amidase domain
Probab=20.90 E-value=1.2e+02 Score=25.44 Aligned_cols=23 Identities=26% Similarity=0.202 Sum_probs=0.0
Q ss_pred CCCCEEEEeecCcc-cceEEEEEc
Q 025018 14 KAGDHIYTYRAVFA-YSHHGIYVG 36 (259)
Q Consensus 14 k~GD~I~~~r~~~~-y~H~GIYvG 36 (259)
.+||+|.+...... +.|.+|.++
T Consensus 99 ~~GDvi~~~~~~~g~~~Hs~iVt~ 122 (157)
T PF12671_consen 99 NPGDVIQYDWSGDGRYDHSMIVTD 122 (157)
T ss_pred cCCCEEEEEeCCCCcEeeEEEEEE
Done!