Query 026044
Match_columns 244
No_of_seqs 102 out of 119
Neff 3.0
Searched_HMMs 46136
Date Fri Mar 29 03:04:25 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/026044.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/026044hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF05212 DUF707: Protein of un 100.0 6.5E-86 1.4E-90 598.7 13.8 173 68-244 3-175 (294)
2 PF14538 Raptor_N: Raptor N-te 68.1 2.1 4.5E-05 36.4 0.5 47 137-195 90-144 (154)
3 PF12996 DUF3880: DUF based on 65.8 5.8 0.00013 29.5 2.5 25 180-214 13-37 (79)
4 PF12621 DUF3779: Phosphate me 52.6 22 0.00048 27.8 3.8 54 173-231 32-87 (95)
5 PRK05325 hypothetical protein; 46.0 14 0.0003 36.3 2.1 28 123-152 296-325 (401)
6 PHA03165 hypothetical protein; 42.6 22 0.00047 26.3 2.2 32 30-70 23-54 (57)
7 PF04285 DUF444: Protein of un 42.4 17 0.00038 35.8 2.2 28 124-153 321-350 (421)
8 TIGR02877 spore_yhbH sporulati 38.5 23 0.0005 34.6 2.3 27 124-152 277-305 (371)
9 PRK13863 type IV secretion sys 36.5 56 0.0012 32.9 4.6 83 108-217 82-178 (446)
10 PF13778 DUF4174: Domain of un 36.0 25 0.00054 28.2 1.8 38 122-159 63-102 (118)
11 PF07862 Nif11: Nitrogen fixat 33.0 35 0.00077 23.1 1.9 21 201-221 27-47 (49)
12 PF11057 Cortexin: Cortexin of 31.8 55 0.0012 26.1 3.0 23 27-49 30-52 (81)
13 PF15018 InaF-motif: TRP-inter 30.2 20 0.00042 24.9 0.3 8 188-195 28-35 (38)
14 KOG2431 1, 2-alpha-mannosidase 29.1 56 0.0012 33.4 3.3 92 28-132 13-106 (546)
15 PF09665 RE_Alw26IDE: Type II 28.2 32 0.00069 35.1 1.5 28 175-203 362-393 (511)
16 PF06679 DUF1180: Protein of u 27.8 1.1E+02 0.0024 26.8 4.5 24 24-47 95-118 (163)
17 PF07172 GRP: Glycine rich pro 27.6 77 0.0017 25.2 3.3 13 26-38 4-16 (95)
18 CHL00123 rps6 ribosomal protei 22.8 57 0.0012 25.5 1.7 37 183-219 5-43 (97)
19 TIGR03798 ocin_TIGR03798 bacte 22.2 69 0.0015 23.0 1.9 24 201-224 25-48 (64)
20 cd04185 GT_2_like_b Subfamily 22.0 81 0.0018 25.1 2.5 38 184-221 78-115 (202)
21 PF07745 Glyco_hydro_53: Glyco 20.7 56 0.0012 31.2 1.5 93 88-196 39-145 (332)
22 PF12849 PBP_like_2: PBP super 20.7 58 0.0013 28.0 1.5 25 144-168 121-146 (281)
23 PRK05637 anthranilate synthase 20.5 84 0.0018 27.4 2.5 49 135-184 136-190 (208)
24 cd04186 GT_2_like_c Subfamily 20.4 1E+02 0.0023 22.9 2.7 25 185-209 74-98 (166)
25 PF07976 Phe_hydrox_dim: Pheno 20.1 2.1E+02 0.0045 24.2 4.7 74 75-158 33-125 (169)
No 1
>PF05212 DUF707: Protein of unknown function (DUF707); InterPro: IPR007877 This family consists of uncharacterised proteins from Arabidopsis thaliana.
Probab=100.00 E-value=6.5e-86 Score=598.72 Aligned_cols=173 Identities=60% Similarity=1.082 Sum_probs=169.0
Q ss_pred cccCCCCcCCCCCCceecCCCcceecCCCCCCCCcccCCCCCccEEEEeecccccccHHHHHhhhCCCCeEEEEEEecCc
Q 026044 68 SRFSSGRLKSLPRGIVQARSDLELRPLWSTSSSRKKFGVYSNRNLLAIPAGIKQKDNVDAIVRKFLPENFTVILFHYDGD 147 (244)
Q Consensus 68 ~~~~p~g~e~LP~gIV~~~Sdl~lr~Lwg~p~~~~~~~~~~~k~Llam~VGikQK~~Vd~~V~KF~~~nF~vmLFHYDG~ 147 (244)
.+++|+|+|+||+|||+++|||+||||||+|+++. +.++|||||||||||||++||++|+|| ++||+||||||||+
T Consensus 3 ~~~~p~g~e~Lp~giv~~~sd~~~r~lw~~p~~~~---~~~~k~Lla~~VG~kqk~~vd~~v~Kf-~~nF~i~LfhYDg~ 78 (294)
T PF05212_consen 3 VPCNPRGAERLPPGIVVRESDLELRPLWGNPSEDL---PKKPKYLLAMTVGIKQKDNVDAIVKKF-SDNFDIMLFHYDGR 78 (294)
T ss_pred cCCCCCccccCCCCccccCCCceeeecCCCccccc---cCCCceEEEEEecHHHHhhhhHHHhhh-ccCceEEEEEecCC
Confidence 46899999999999999999999999999999887 358899999999999999999999999 99999999999999
Q ss_pred cccccccccCCceEEEEEecccchhhhcccCCCccccccceEEEeccccccCCCChhHHHHHHHHhCCccccCccCCCCC
Q 026044 148 VNAWRGLDWSNKAIHIAAQNQTKWWFAKRFLHPDVVSNYDYIFLWDEDLGVENFDPRRYLEIVKSEGFEISQPALDPNST 227 (244)
Q Consensus 148 vd~W~dleWs~~aIHVsa~~QtKWWfAKRFLHPdiVa~YeYiFlWDEDLgve~F~~~rYl~Ivk~~gLEISQPaLd~~~~ 227 (244)
||+|+|||||++||||+++|||||||||||||||||++||||||||||||||||||+|||+|||+|||||||||||+++|
T Consensus 79 vd~w~~~~ws~~aiHv~~~kqtKww~akrfLHPdiv~~YdYiflwDeDL~vd~f~~~ry~~Ivk~~gLeISQPALd~~~~ 158 (294)
T PF05212_consen 79 VDEWDDFEWSDRAIHVSARKQTKWWFAKRFLHPDIVAPYDYIFLWDEDLGVDHFDINRYFEIVKKEGLEISQPALDPDSS 158 (294)
T ss_pred cCchhhcccccceEEEEeccceEEeehhhhcChhhhccceeEEecCCccCcCcCCHHHHHHHHHHhCCcccCcccCCCCc
Confidence 99999999999999999999999999999999999999999999999999999999999999999999999999999999
Q ss_pred ceeeeeeeecCCcccCC
Q 026044 228 EIHHKFTIRARTKKFHR 244 (244)
Q Consensus 228 ~ihh~iT~R~~~~~vHr 244 (244)
++||+||+|++.++|||
T Consensus 159 ~~~~~iT~R~~~~~vhr 175 (294)
T PF05212_consen 159 EIHHPITKRRPDSEVHR 175 (294)
T ss_pred eeeeeEEeecCCceeEe
Confidence 99999999999999997
No 2
>PF14538 Raptor_N: Raptor N-terminal CASPase like domain
Probab=68.12 E-value=2.1 Score=36.36 Aligned_cols=47 Identities=32% Similarity=0.613 Sum_probs=28.4
Q ss_pred eEEEEEEecCccccccccccCCceEEEEEecccchhhhcccCCCccccccc--------eEEEeccc
Q 026044 137 FTVILFHYDGDVNAWRGLDWSNKAIHIAAQNQTKWWFAKRFLHPDVVSNYD--------YIFLWDED 195 (244)
Q Consensus 137 F~vmLFHYDG~vd~W~dleWs~~aIHVsa~~QtKWWfAKRFLHPdiVa~Ye--------YiFlWDED 195 (244)
-.=+||||-|. .++. -..+..=|-|-|.+-.-.-+.-|| -||+||++
T Consensus 90 ~~RvLFHYnGh-----GvP~-------Pt~~GeIw~f~~~~tqyip~si~dL~~~lg~Psi~V~DC~ 144 (154)
T PF14538_consen 90 DERVLFHYNGH-----GVPR-------PTENGEIWVFNKNYTQYIPLSIYDLQSWLGSPSIYVFDCS 144 (154)
T ss_pred CceEEEEECCC-----CCCC-------CCCCCeEEEEcCCCCcceEEEHHHHHHhcCCCEEEEEECC
Confidence 37899999993 2222 122233455666665444455454 48999987
No 3
>PF12996 DUF3880: DUF based on E. rectale Gene description (DUF3880); InterPro: IPR024542 This entry represents proteins of unknown function. The Eubacterium rectale gene appears to be upregulated in the presence of Bacteroides thetaiotaomicron compared to growth in pure culture [].
Probab=65.78 E-value=5.8 Score=29.51 Aligned_cols=25 Identities=32% Similarity=0.743 Sum_probs=18.5
Q ss_pred CccccccceEEEeccccccCCCChhHHHHHHHHhC
Q 026044 180 PDVVSNYDYIFLWDEDLGVENFDPRRYLEIVKSEG 214 (244)
Q Consensus 180 PdiVa~YeYiFlWDEDLgve~F~~~rYl~Ivk~~g 214 (244)
..+...|+|||+||++ .++-.|+.|
T Consensus 13 ~~i~~~~~~iFt~D~~----------~~~~~~~~G 37 (79)
T PF12996_consen 13 YSIANSYDYIFTFDRS----------FVEEYRNLG 37 (79)
T ss_pred hhhCCCCCEEEEECHH----------HHHHHHHcC
Confidence 3677899999999974 455556666
No 4
>PF12621 DUF3779: Phosphate metabolism protein ; InterPro: IPR022257 This domain family is found in eukaryotes, and is approximately 100 amino acids in length. The family is found in association with PF02714 from PFAM. There are two completely conserved residues (W and D) that may be functionally important. This family is likely to be involved in phosphate metabolism however there is little accompanying literature to confirm this.
Probab=52.56 E-value=22 Score=27.75 Aligned_cols=54 Identities=26% Similarity=0.444 Sum_probs=45.1
Q ss_pred hhcccCCCccccccceEEEeccccccCCCChhHHHHHHHHhCCccccCc--cCCCCCceee
Q 026044 173 FAKRFLHPDVVSNYDYIFLWDEDLGVENFDPRRYLEIVKSEGFEISQPA--LDPNSTEIHH 231 (244)
Q Consensus 173 fAKRFLHPdiVa~YeYiFlWDEDLgve~F~~~rYl~Ivk~~gLEISQPa--Ld~~~~~ihh 231 (244)
-..-|+||.+-++--.|-|--.++|| .+.=++-.++.|++||.-+ ||. +|.+.+
T Consensus 32 ~~~ay~~Pa~~~~~P~lWIP~D~~Gv----S~~ei~~~~~~~v~~Sd~gA~lde-kgkv~~ 87 (95)
T PF12621_consen 32 HKHAYLHPAVSAPQPILWIPRDPLGV----SRQEIEETRKVGVPISDEGATLDE-KGKVVW 87 (95)
T ss_pred HHhccCCHhHcCCCCeEEeecCCCCC----CHHHHHHhhcCCeEEECCCeEEcc-CCCEEE
Confidence 35679999999999999999999999 4567788899999999887 676 566766
No 5
>PRK05325 hypothetical protein; Provisional
Probab=46.02 E-value=14 Score=36.31 Aligned_cols=28 Identities=21% Similarity=0.621 Sum_probs=24.1
Q ss_pred ccHHHHHhh-hCCCCeEEEEEEe-cCcccccc
Q 026044 123 DNVDAIVRK-FLPENFTVILFHY-DGDVNAWR 152 (244)
Q Consensus 123 ~~Vd~~V~K-F~~~nF~vmLFHY-DG~vd~W~ 152 (244)
+.++.||++ |+++.+.|..||. || |.|.
T Consensus 296 ~l~~eIi~~rYpp~~wNIY~f~aSDG--DNw~ 325 (401)
T PRK05325 296 KLALEIIEERYPPAEWNIYAFQASDG--DNWS 325 (401)
T ss_pred HHHHHHHHhhCCHhHCeeEEEEcccC--CCcC
Confidence 346788885 9999999999997 88 8887
No 6
>PHA03165 hypothetical protein; Provisional
Probab=42.60 E-value=22 Score=26.26 Aligned_cols=32 Identities=25% Similarity=0.491 Sum_probs=24.7
Q ss_pred HHHHHHHHHHHHHhhhhhhhhhhhhhhhccCCCcccccccc
Q 026044 30 FMAIMCTVMLFVVYRTTYYQYKQTEMEAKFSPFDISKGSRF 70 (244)
Q Consensus 30 ~~~~~c~v~~f~~~~~~~~q~~~~~~~~~~~~~~~~~~~~~ 70 (244)
...+++.+++|++|..+- ...+||++.-.++|
T Consensus 23 yilvvafvlaflvysdfl---------snlspfgeilsspc 54 (57)
T PHA03165 23 YILVVAFVLAFLVYSDFL---------SNLSPFGEILSSPC 54 (57)
T ss_pred ehhHHHHHHHHHHHHHHH---------hccCchhhhhcCcc
Confidence 467788899999999887 66788887666554
No 7
>PF04285 DUF444: Protein of unknown function (DUF444); InterPro: IPR006698 This entry is represented by Thermus phage phiYS40, Orf56. The characteristics of the protein distribution suggest prophage matches in addition to the phage matches [].
Probab=42.36 E-value=17 Score=35.82 Aligned_cols=28 Identities=25% Similarity=0.730 Sum_probs=23.8
Q ss_pred cHHHHHhh-hCCCCeEEEEEEe-cCccccccc
Q 026044 124 NVDAIVRK-FLPENFTVILFHY-DGDVNAWRG 153 (244)
Q Consensus 124 ~Vd~~V~K-F~~~nF~vmLFHY-DG~vd~W~d 153 (244)
.++.||++ |++++++|..||. || |.|.+
T Consensus 321 l~~~ii~erypp~~wNiY~~~~SDG--DN~~~ 350 (421)
T PF04285_consen 321 LALEIIEERYPPSDWNIYVFHASDG--DNWSS 350 (421)
T ss_pred HHHHHHHhhCChhhceeeeEEcccC--ccccC
Confidence 46778886 9999999999998 88 88873
No 8
>TIGR02877 spore_yhbH sporulation protein YhbH. This protein family, typified by YhbH in Bacillus subtilis, is found in nearly every endospore-forming bacterium and in no other genome (but note that the trusted cutoff score is set high to exclude a single high-scoring sequence from Nitrosococcus oceani ATCC 19707, which is classified in the Gammaproteobacteria). The gene in Bacillus subtilis was shown to be in the regulon of the sporulation sigma factor, sigma-E, and its mutation was shown to create a sporulation defect.
Probab=38.52 E-value=23 Score=34.62 Aligned_cols=27 Identities=22% Similarity=0.662 Sum_probs=23.0
Q ss_pred cHHHHHh-hhCCCCeEEEEEEe-cCcccccc
Q 026044 124 NVDAIVR-KFLPENFTVILFHY-DGDVNAWR 152 (244)
Q Consensus 124 ~Vd~~V~-KF~~~nF~vmLFHY-DG~vd~W~ 152 (244)
..+.||+ +|+++.+.|..||. || |.|.
T Consensus 277 l~~eII~~rYpp~~wNIY~f~aSDG--DNw~ 305 (371)
T TIGR02877 277 KALEIIDERYNPARYNIYAFHFSDG--DNLT 305 (371)
T ss_pred HHHHHHHhhCChhhCeeEEEEcccC--CCcc
Confidence 3566776 79999999999998 88 8887
No 9
>PRK13863 type IV secretion system T-DNA border endonuclease VirD2; Provisional
Probab=36.52 E-value=56 Score=32.90 Aligned_cols=83 Identities=19% Similarity=0.343 Sum_probs=50.8
Q ss_pred CCccEEEEeecccccccHHH----HHhhhCCC----Ce-EEEEEEecCccccccccccCCceEEEEEe---cccchhhhc
Q 026044 108 SNRNLLAIPAGIKQKDNVDA----IVRKFLPE----NF-TVILFHYDGDVNAWRGLDWSNKAIHIAAQ---NQTKWWFAK 175 (244)
Q Consensus 108 ~~k~Llam~VGikQK~~Vd~----~V~KF~~~----nF-~vmLFHYDG~vd~W~dleWs~~aIHVsa~---~QtKWWfAK 175 (244)
...-+|.|+.|-.+.+..++ +-++|++. +| -|+-||-|-. .--+||++. +--|=|
T Consensus 82 T~NIVLSMPaGTd~eAVrdAARefA~E~FgsG~~G~~~dYV~AlH~D~d----------HPHVHLvVnrRd~~G~~~--- 148 (446)
T PRK13863 82 TTHIIVSFPAGTSQVAAYAASREWAAEMFGSGAGGGRYNYLTAFHIDRD----------HPHLHVVVNRRELLGHGW--- 148 (446)
T ss_pred eEEEEEeCCCCCCHHHHHHHHHHHHHHHhCCCCCCCceeEEEEEecCCC----------CCeEEEEEEeecCCCCce---
Confidence 33468999999777665552 33556542 44 3678997761 456899988 444423
Q ss_pred ccCCCccccccceEEEe--ccccccCCCChhHHHHHHHHhCCcc
Q 026044 176 RFLHPDVVSNYDYIFLW--DEDLGVENFDPRRYLEIVKSEGFEI 217 (244)
Q Consensus 176 RFLHPdiVa~YeYiFlW--DEDLgve~F~~~rYl~Ivk~~gLEI 217 (244)
++|+ ..|+.++.+ -+.|-++.+++|++.
T Consensus 149 -------------lri~~rk~dlNld~~-Re~FAE~LRe~GIea 178 (446)
T PRK13863 149 -------------LKISRRHPQLNYDAL-RIKMAEISLRHGIVL 178 (446)
T ss_pred -------------eeecCCCccccHHHH-HHHHHHHHHhcCcee
Confidence 2222 123332222 257999999999985
No 10
>PF13778 DUF4174: Domain of unknown function (DUF4174)
Probab=36.01 E-value=25 Score=28.16 Aligned_cols=38 Identities=24% Similarity=0.409 Sum_probs=30.8
Q ss_pred cccHHHHHhhhC--CCCeEEEEEEecCccccccccccCCc
Q 026044 122 KDNVDAIVRKFL--PENFTVILFHYDGDVNAWRGLDWSNK 159 (244)
Q Consensus 122 K~~Vd~~V~KF~--~~nF~vmLFHYDG~vd~W~dleWs~~ 159 (244)
...+..+-++|. .++|+++|.-.||.|-.+..-+|+-+
T Consensus 63 ~~~~~~lr~~l~~~~~~f~~vLiGKDG~vK~r~~~p~~~~ 102 (118)
T PF13778_consen 63 PEDIQALRKRLRIPPGGFTVVLIGKDGGVKLRWPEPIDPE 102 (118)
T ss_pred HHHHHHHHHHhCCCCCceEEEEEeCCCcEEEecCCCCCHH
Confidence 345678888887 78999999999999988877766544
No 11
>PF07862 Nif11: Nitrogen fixation protein of unknown function; InterPro: IPR012903 This domain is found in the cyanobacteria, and the nitrogen-fixing proteobacterium Azotobacter vinelandii and may be involved in nitrogen fixation, but no role has been assigned [].
Probab=33.05 E-value=35 Score=23.10 Aligned_cols=21 Identities=10% Similarity=0.518 Sum_probs=18.2
Q ss_pred CChhHHHHHHHHhCCccccCc
Q 026044 201 FDPRRYLEIVKSEGFEISQPA 221 (244)
Q Consensus 201 F~~~rYl~Ivk~~gLEISQPa 221 (244)
-+++..++|++++|.+||.--
T Consensus 27 ~~~~e~~~lA~~~Gy~ft~~e 47 (49)
T PF07862_consen 27 QNPEEVVALAREAGYDFTEEE 47 (49)
T ss_pred CCHHHHHHHHHHcCCCCCHHH
Confidence 389999999999999998643
No 12
>PF11057 Cortexin: Cortexin of kidney; InterPro: IPR020066 Cortexin is a neuron-specific, 82-residue membrane protein which is found especially in vertebrate brain cortex tissue. It may mediate extracellular or intracellular signalling of cortical neurons during forebrain development. Cortexin is present at significant levels in the foetal brain, suggesting that it may be important to neurons of both the developing and adult cerebral cortex. Cortexin has a conserved single membrane-spanning region in the middle of each sequence []. In humans, there is selective expression of Cortexin 3 (CTXN3) in the kidney as well as the brain []. This entry contains Cortexins 1, 2 and 3.; GO: 0031224 intrinsic to membrane
Probab=31.79 E-value=55 Score=26.09 Aligned_cols=23 Identities=13% Similarity=0.417 Sum_probs=19.3
Q ss_pred hhhHHHHHHHHHHHHHhhhhhhh
Q 026044 27 QLQFMAIMCTVMLFVVYRTTYYQ 49 (244)
Q Consensus 27 ~~~~~~~~c~v~~f~~~~~~~~q 49 (244)
.+-|+.++|+.+++++.|++.+-
T Consensus 30 ~faFV~~L~~fL~~liVRCfrIl 52 (81)
T PF11057_consen 30 AFAFVGLLCLFLGLLIVRCFRIL 52 (81)
T ss_pred eehHHHHHHHHHHHHHHHHHHHH
Confidence 35678999999999999999854
No 13
>PF15018 InaF-motif: TRP-interacting helix
Probab=30.22 E-value=20 Score=24.86 Aligned_cols=8 Identities=75% Similarity=1.825 Sum_probs=4.6
Q ss_pred eEEEeccc
Q 026044 188 YIFLWDED 195 (244)
Q Consensus 188 YiFlWDED 195 (244)
|+|+||.+
T Consensus 28 Y~f~W~p~ 35 (38)
T PF15018_consen 28 YIFFWDPD 35 (38)
T ss_pred HheeeCCC
Confidence 56666554
No 14
>KOG2431 consensus 1, 2-alpha-mannosidase [Carbohydrate transport and metabolism]
Probab=29.10 E-value=56 Score=33.35 Aligned_cols=92 Identities=18% Similarity=0.199 Sum_probs=46.4
Q ss_pred hhHHHHHHHHHHHHHhhhhhhhhhhhhhhhccCCCccccccccCCCCcCCCCCCceecCCCcceecCCCCCCCCcccCCC
Q 026044 28 LQFMAIMCTVMLFVVYRTTYYQYKQTEMEAKFSPFDISKGSRFSSGRLKSLPRGIVQARSDLELRPLWSTSSSRKKFGVY 107 (244)
Q Consensus 28 ~~~~~~~c~v~~f~~~~~~~~q~~~~~~~~~~~~~~~~~~~~~~p~g~e~LP~gIV~~~Sdl~lr~Lwg~p~~~~~~~~~ 107 (244)
+.|.+.+|+.+++.+|...+ ..+ +-..|-...+.....-++++.|||++-+..+.-+...- +..+.+....
T Consensus 13 ilf~~~~~~~v~l~~~~~~~----~p~--~~~~~~~~~~t~~~~~~sa~~l~p~~~~~~~~~~~~~p---~~~~~~~~~i 83 (546)
T KOG2431|consen 13 ILFILAFLLFVLLLLYINPA----NPA--ELPNPQSGQKTKRGGQRSAENLPPDLPQQSATDEQEAP---KEGDPNRTVI 83 (546)
T ss_pred HHHHHHHHHHHHHHHhcCCC----Chh--hcCCccccchhhhhcccCcccCCCCcchhhchhhccCC---ccCCCCCcce
Confidence 56777777777665555421 111 11111111122234567888899988877776665432 1122211000
Q ss_pred CCccEEEEe--ecccccccHHHHHhhh
Q 026044 108 SNRNLLAIP--AGIKQKDNVDAIVRKF 132 (244)
Q Consensus 108 ~~k~Llam~--VGikQK~~Vd~~V~KF 132 (244)
...-+ .+-.||+.|++...-|
T Consensus 84 ----~~~~Ptg~nerq~avv~aF~haW 106 (546)
T KOG2431|consen 84 ----SFRGPTGLNERQKAVVDAFLHAW 106 (546)
T ss_pred ----eecCCCchhHHHHHHHHHHHHHH
Confidence 00002 3667888888877766
No 15
>PF09665 RE_Alw26IDE: Type II restriction endonuclease (RE_Alw26IDE); InterPro: IPR014328 There are four classes of restriction endonucleases: types I, II,III and IV. All types of enzymes recognise specific short DNA sequences and carry out the endonucleolytic cleavage of DNA to give specific double-stranded fragments with terminal 5'-phosphates. They differ in their recognition sequence, subunit composition, cleavage position, and cofactor requirements [, ], as summarised below: Type I enzymes (3.1.21.3 from EC) cleave at sites remote from recognition site; require both ATP and S-adenosyl-L-methionine to function; multifunctional protein with both restriction and methylase (2.1.1.72 from EC) activities. Type II enzymes (3.1.21.4 from EC) cleave within or at short specific distances from recognition site; most require magnesium; single function (restriction) enzymes independent of methylase. Type III enzymes (3.1.21.5 from EC) cleave at sites a short distance from recognition site; require ATP (but doesn't hydrolyse it); S-adenosyl-L-methionine stimulates reaction but is not required; exists as part of a complex with a modification methylase methylase (2.1.1.72 from EC). Type IV enzymes target methylated DNA. Type II restriction endonucleases (3.1.21.4 from EC) are components of prokaryotic DNA restriction-modification mechanisms that protect the organism against invading foreign DNA. These site-specific deoxyribonucleases catalyse the endonucleolytic cleavage of DNA to give specific double-stranded fragments with terminal 5'-phosphates. Of the 3000 restriction endonucleases that have been characterised, most are homodimeric or tetrameric enzymes that cleave target DNA at sequence-specific sites close to the recognition site. For homodimeric enzymes, the recognition site is usually a palindromic sequence 4-8 bp in length. Most enzymes require magnesium ions as a cofactor for catalysis. Although they can vary in their mode of recognition, many restriction endonucleases share a similar structural core comprising four beta-strands and one alpha-helix, as well as a similar mechanism of cleavage, suggesting a common ancestral origin []. However, there is still considerable diversity amongst restriction endonucleases [, ]. The target site recognition process triggers large conformational changes of the enzyme and the target DNA, leading to the activation of the catalytic centres. Like other DNA binding proteins, restriction enzymes are capable of non-specific DNA binding as well, which is the prerequisite for efficient target site location by facilitated diffusion. Non-specific binding usually does not involve interactions with the bases but only with the DNA backbone []. This entry represents type II restriction endonucleases of the Alw26I/Eco31I/Esp3I family [], whose recognition sequences are 5'-GTCTC-3' (Alw26I), 5'-GGTCTC-3' (Eco31I) and 5'-CGTCTC-3' (Esp3I).
Probab=28.18 E-value=32 Score=35.08 Aligned_cols=28 Identities=32% Similarity=0.561 Sum_probs=21.4
Q ss_pred cccCCCccccccceEE--Eeccc--cccCCCCh
Q 026044 175 KRFLHPDVVSNYDYIF--LWDED--LGVENFDP 203 (244)
Q Consensus 175 KRFLHPdiVa~YeYiF--lWDED--Lgve~F~~ 203 (244)
--||||.. +.|+|.| +|-++ +...++.+
T Consensus 362 ~t~L~~~Y-a~y~y~Fe~~~~~~~~~~~~~i~~ 393 (511)
T PF09665_consen 362 ATFLKPEY-ANYDYTFEGLNISNHLTQYKSIYK 393 (511)
T ss_pred HHHhchhh-hhccceeccccccccccccccccc
Confidence 57899999 9999999 56566 55556666
No 16
>PF06679 DUF1180: Protein of unknown function (DUF1180); InterPro: IPR009565 This entry consists of several hypothetical eukaryotic proteins thought to be membrane proteins. Their function is unknown.
Probab=27.77 E-value=1.1e+02 Score=26.78 Aligned_cols=24 Identities=21% Similarity=0.285 Sum_probs=16.8
Q ss_pred eeehhhHHHHHHHHHHHHHhhhhh
Q 026044 24 KMKQLQFMAIMCTVMLFVVYRTTY 47 (244)
Q Consensus 24 ~~~~~~~~~~~c~v~~f~~~~~~~ 47 (244)
+.-++-++++.++++++||.+++-
T Consensus 95 ~R~~~Vl~g~s~l~i~yfvir~~R 118 (163)
T PF06679_consen 95 KRALYVLVGLSALAILYFVIRTFR 118 (163)
T ss_pred hhhHHHHHHHHHHHHHHHHHHHHh
Confidence 444455677778888888888765
No 17
>PF07172 GRP: Glycine rich protein family; InterPro: IPR010800 This family consists of glycine rich proteins. Some of them may be involved in resistance to environmental stress [].
Probab=27.58 E-value=77 Score=25.21 Aligned_cols=13 Identities=8% Similarity=0.419 Sum_probs=5.2
Q ss_pred ehhhHHHHHHHHH
Q 026044 26 KQLQFMAIMCTVM 38 (244)
Q Consensus 26 ~~~~~~~~~c~v~ 38 (244)
|.|.+++|+-+++
T Consensus 4 K~~llL~l~LA~l 16 (95)
T PF07172_consen 4 KAFLLLGLLLAAL 16 (95)
T ss_pred hHHHHHHHHHHHH
Confidence 4344444433333
No 18
>CHL00123 rps6 ribosomal protein S6; Validated
Probab=22.81 E-value=57 Score=25.49 Aligned_cols=37 Identities=19% Similarity=0.439 Sum_probs=31.4
Q ss_pred ccccceEEEeccccccCCCCh--hHHHHHHHHhCCcccc
Q 026044 183 VSNYDYIFLWDEDLGVENFDP--RRYLEIVKSEGFEISQ 219 (244)
Q Consensus 183 Va~YeYiFlWDEDLgve~F~~--~rYl~Ivk~~gLEISQ 219 (244)
+..||-+||.+.|+.=|.... ++|-+++.++|-+|-.
T Consensus 5 mr~YE~~~Il~p~l~e~~~~~~~~~~~~~i~~~gg~i~~ 43 (97)
T CHL00123 5 LNKYETMYLLKPDLNEEELLKWIENYKKLLRKRGAKNIS 43 (97)
T ss_pred ccceeEEEEECCCCCHHHHHHHHHHHHHHHHHCCCEEEE
Confidence 356999999999998887774 8899999999988743
No 19
>TIGR03798 ocin_TIGR03798 bacteriocin propeptide, TIGR03798 family. This model describes a conserved, fairly long (about 65 residue) propeptide region for a family of putative microcins, that is, bacteriocins of small size. Members of the seed alignment tend to have the Gly-Gly motif as the last two residues of the matched region. This is a cleavage site for a combination processing/export ABC transporter with a peptidase domain.
Probab=22.24 E-value=69 Score=23.03 Aligned_cols=24 Identities=33% Similarity=0.527 Sum_probs=20.8
Q ss_pred CChhHHHHHHHHhCCccccCccCC
Q 026044 201 FDPRRYLEIVKSEGFEISQPALDP 224 (244)
Q Consensus 201 F~~~rYl~Ivk~~gLEISQPaLd~ 224 (244)
=+|+..++|++++|.+||.--|+.
T Consensus 25 ~~~e~~~~lA~~~Gf~ft~~el~~ 48 (64)
T TIGR03798 25 EDPEDRVAIAKEAGFEFTGEDLKE 48 (64)
T ss_pred CCHHHHHHHHHHcCCCCCHHHHHH
Confidence 468999999999999999887754
No 20
>cd04185 GT_2_like_b Subfamily of Glycosyltransferase Family GT2 of unknown function. GT-2 includes diverse families of glycosyltransferases with a common GT-A type structural fold, which has two tightly associated beta/alpha/beta domains that tend to form a continuous central sheet of at least eight beta-strands. These are enzymes that catalyze the transfer of sugar moieties from activated donor molecules to specific acceptor molecules, forming glycosidic bonds. Glycosyltransferases have been classified into more than 90 distinct sequence based families.
Probab=22.04 E-value=81 Score=25.10 Aligned_cols=38 Identities=21% Similarity=0.335 Sum_probs=26.4
Q ss_pred cccceEEEeccccccCCCChhHHHHHHHHhCCccccCc
Q 026044 184 SNYDYIFLWDEDLGVENFDPRRYLEIVKSEGFEISQPA 221 (244)
Q Consensus 184 a~YeYiFlWDEDLgve~F~~~rYl~Ivk~~gLEISQPa 221 (244)
+.+||+++-|.|..++.=.-++.++.+++.+..+..|.
T Consensus 78 ~~~d~v~~ld~D~~~~~~~l~~l~~~~~~~~~~~~~~~ 115 (202)
T cd04185 78 LGYDWIWLMDDDAIPDPDALEKLLAYADKDNPQFLAPL 115 (202)
T ss_pred cCCCEEEEeCCCCCcChHHHHHHHHHHhcCCceEecce
Confidence 57999999999998865444556666655555555554
No 21
>PF07745 Glyco_hydro_53: Glycosyl hydrolase family 53; InterPro: IPR011683 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. This domain is found in family 53 of the glycosyl hydrolase classification []. These enzymes are endo-1,4- beta-galactanases (3.2.1.89 from EC). The structure of this domain is known [] and has a TIM barrel fold.; GO: 0015926 glucosidase activity; PDB: 1HJQ_A 1HJS_A 1HJU_B 1FHL_A 1FOB_A 2GFT_A 1UR4_B 1UR0_A 1R8L_B 2CCR_A ....
Probab=20.72 E-value=56 Score=31.15 Aligned_cols=93 Identities=20% Similarity=0.344 Sum_probs=48.7
Q ss_pred CcceecCCCCCCCCcccCCCCCccEEEEeecccccccHHHHHhhhCCCCeEEEE-EEecCcc----ccccccccCCce--
Q 026044 88 DLELRPLWSTSSSRKKFGVYSNRNLLAIPAGIKQKDNVDAIVRKFLPENFTVIL-FHYDGDV----NAWRGLDWSNKA-- 160 (244)
Q Consensus 88 dl~lr~Lwg~p~~~~~~~~~~~k~Llam~VGikQK~~Vd~~V~KF~~~nF~vmL-FHYDG~v----d~W~dleWs~~a-- 160 (244)
|.-.-|+|-+|.. -|....+.|-++.|+--...+.||| |||-..- .++.-=.|.+..
T Consensus 39 N~vRlRvwv~P~~----------------~g~~~~~~~~~~akrak~~Gm~vlldfHYSD~WaDPg~Q~~P~aW~~~~~~ 102 (332)
T PF07745_consen 39 NAVRLRVWVNPYD----------------GGYNDLEDVIALAKRAKAAGMKVLLDFHYSDFWADPGKQNKPAAWANLSFD 102 (332)
T ss_dssp -EEEEEE-SS-TT----------------TTTTSHHHHHHHHHHHHHTT-EEEEEE-SSSS--BTTB-B--TTCTSSSHH
T ss_pred CeEEEEeccCCcc----------------cccCCHHHHHHHHHHHHHCCCeEEEeecccCCCCCCCCCCCCccCCCCCHH
Confidence 3344488998865 5888899999999998888899998 9994310 111112332210
Q ss_pred -EEEEEecccch---hhhcccCCCcccc---ccceEEEecccc
Q 026044 161 -IHIAAQNQTKW---WFAKRFLHPDVVS---NYDYIFLWDEDL 196 (244)
Q Consensus 161 -IHVsa~~QtKW---WfAKRFLHPdiVa---~YeYiFlWDEDL 196 (244)
+--++..=||- -+...=.-||+|+ +..+=|||++.-
T Consensus 103 ~l~~~v~~yT~~vl~~l~~~G~~pd~VQVGNEin~Gmlwp~g~ 145 (332)
T PF07745_consen 103 QLAKAVYDYTKDVLQALKAAGVTPDMVQVGNEINNGMLWPDGK 145 (332)
T ss_dssp HHHHHHHHHHHHHHHHHHHTT--ESEEEESSSGGGESTBTTTC
T ss_pred HHHHHHHHHHHHHHHHHHHCCCCccEEEeCccccccccCcCCC
Confidence 00000000111 0334456788886 677888887665
No 22
>PF12849 PBP_like_2: PBP superfamily domain; InterPro: IPR024370 This entry represents members of the periplasmic binding domain superfamily []. It is often associated with a helix-turn-helix domain.; PDB: 1QUL_A 1OIB_A 1A54_A 1IXH_A 1A40_A 1QUJ_A 1A55_A 1IXI_A 2ABH_A 1QUK_A ....
Probab=20.71 E-value=58 Score=27.96 Aligned_cols=25 Identities=16% Similarity=0.662 Sum_probs=18.5
Q ss_pred ecCcccccccc-ccCCceEEEEEecc
Q 026044 144 YDGDVNAWRGL-DWSNKAIHIAAQNQ 168 (244)
Q Consensus 144 YDG~vd~W~dl-eWs~~aIHVsa~~Q 168 (244)
|.|.++.|+|+ .|.++.|++..+..
T Consensus 121 ~~G~It~W~~~~~~~~~~I~~~~r~~ 146 (281)
T PF12849_consen 121 FSGEITNWSDLGGGPDRPIKVVGRSD 146 (281)
T ss_dssp HCTS--BGGGTTTCHSSB-EEEEESS
T ss_pred HhhhhhcccccccCCCCceEEEeCCC
Confidence 35779999998 89999999997754
No 23
>PRK05637 anthranilate synthase component II; Provisional
Probab=20.52 E-value=84 Score=27.41 Aligned_cols=49 Identities=20% Similarity=0.306 Sum_probs=32.9
Q ss_pred CCeEEEEEEecCcc---ccccccccCCc---eEEEEEecccchhhhcccCCCcccc
Q 026044 135 ENFTVILFHYDGDV---NAWRGLDWSNK---AIHIAAQNQTKWWFAKRFLHPDVVS 184 (244)
Q Consensus 135 ~nF~vmLFHYDG~v---d~W~dleWs~~---aIHVsa~~QtKWWfAKRFLHPdiVa 184 (244)
+.|.|..+|-|..+ ++..-+.||+. .+-.++.+..+..|+=.| ||+++-
T Consensus 136 ~~~~V~~~H~~~v~~lp~~~~vlA~s~~~~~~v~~a~~~~~~~~~GvQf-HPE~~~ 190 (208)
T PRK05637 136 RKVPIARYHSLGCVVAPDGMESLGTCSSEIGPVIMAAETTDGKAIGLQF-HPESVL 190 (208)
T ss_pred CceEEEEechhhhhcCCCCeEEEEEecCCCCCEEEEEEECCCCEEEEEe-CCccCc
Confidence 45888889988764 33444567654 244455666778888888 998764
No 24
>cd04186 GT_2_like_c Subfamily of Glycosyltransferase Family GT2 of unknown function. GT-2 includes diverse families of glycosyltransferases with a common GT-A type structural fold, which has two tightly associated beta/alpha/beta domains that tend to form a continuous central sheet of at least eight beta-strands. These are enzymes that catalyze the transfer of sugar moieties from activated donor molecules to specific acceptor molecules, forming glycosidic bonds. Glycosyltransferases have been classified into more than 90 distinct sequence based families.
Probab=20.41 E-value=1e+02 Score=22.86 Aligned_cols=25 Identities=28% Similarity=0.293 Sum_probs=18.4
Q ss_pred ccceEEEeccccccCCCChhHHHHH
Q 026044 185 NYDYIFLWDEDLGVENFDPRRYLEI 209 (244)
Q Consensus 185 ~YeYiFlWDEDLgve~F~~~rYl~I 209 (244)
.+|||++.|.|.-++.-..+++++.
T Consensus 74 ~~~~i~~~D~D~~~~~~~l~~~~~~ 98 (166)
T cd04186 74 KGDYVLLLNPDTVVEPGALLELLDA 98 (166)
T ss_pred CCCEEEEECCCcEECccHHHHHHHH
Confidence 7999999999987755444555553
No 25
>PF07976 Phe_hydrox_dim: Phenol hydroxylase, C-terminal dimerisation domain ; InterPro: IPR012941 Phenol hydroxylase is a homodimer which hydroxylates phenol to catechol, or similar products. The enzyme is comprised of three domains. The first two domains form the active site. The third domain, this domain, is involved in forming the dimerisation interface. The domain adopts a thioredoxin-like fold [].; PDB: 2DKH_A 2DKI_A 1PN0_A 1FOH_D.
Probab=20.11 E-value=2.1e+02 Score=24.16 Aligned_cols=74 Identities=19% Similarity=0.299 Sum_probs=42.0
Q ss_pred cCCCCCCceecCCCcceecCCCCCCCCcccCCCCCccEEEEeecccccc---cHH----------HHHhhhCCC------
Q 026044 75 LKSLPRGIVQARSDLELRPLWSTSSSRKKFGVYSNRNLLAIPAGIKQKD---NVD----------AIVRKFLPE------ 135 (244)
Q Consensus 75 ~e~LP~gIV~~~Sdl~lr~Lwg~p~~~~~~~~~~~k~Llam~VGikQK~---~Vd----------~~V~KF~~~------ 135 (244)
-++||+.-|.+-+|-....|-..=..+ -.=.++.++=-+.+-+ .++ .++++|...
T Consensus 33 G~Rlp~~~v~r~aD~~p~~l~~~l~sd------Grfri~vFagd~~~~~~~~~l~~l~~~L~~~~s~~~r~~~~~~~~~s 106 (169)
T PF07976_consen 33 GRRLPSAKVVRHADGNPVHLQDDLPSD------GRFRILVFAGDISLPEQLSRLSALADYLESPSSFLSRFTPKDRDPDS 106 (169)
T ss_dssp TCB----EEEETTTTEEEEGGGG--SS------S-EEEEEEEETTTTCHCCCHHHHHHHHHHSTTSHHHHHSBTTS-TTS
T ss_pred ccccCCceEEEEcCCCChhHhhhcccC------CCEEEEEEeCCCccchhHHHHHHHHHHHHhcchHHHhcCCCCCCCCC
Confidence 358999999999998888885421111 1225666665554433 333 355677653
Q ss_pred CeEEEEEEecCccccccccccCC
Q 026044 136 NFTVILFHYDGDVNAWRGLDWSN 158 (244)
Q Consensus 136 nF~vmLFHYDG~vd~W~dleWs~ 158 (244)
-|+++|+| =..++++||.+
T Consensus 107 ~~~~~~I~----~~~~~~~e~~d 125 (169)
T PF07976_consen 107 VFDVLLIH----SSPRDEVELFD 125 (169)
T ss_dssp SEEEEEEE----SS-CCCS-GGG
T ss_pred eeEEEEEe----cCCCCceeHHH
Confidence 39999999 45688888853
Done!