Query 015059
Match_columns 414
No_of_seqs 96 out of 98
Neff 2.6
Searched_HMMs 46136
Date Fri Mar 29 02:41:43 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/015059.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/015059hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 KOG1308 Hsp70-interacting prot 100.0 1.8E-41 4E-46 333.2 -3.3 310 62-413 51-366 (377)
2 smart00727 STI1 Heat shock cha 98.2 1.7E-06 3.6E-11 60.8 3.4 40 374-413 1-41 (41)
3 KOG3037 Cell membrane glycopro 96.2 0.013 2.7E-07 59.2 6.9 56 273-328 225-280 (330)
4 smart00727 STI1 Heat shock cha 95.6 0.0098 2.1E-07 41.8 2.4 37 282-322 4-41 (41)
5 KOG0010 Ubiquitin-like protein 93.2 0.1 2.2E-06 55.2 4.5 71 273-396 328-399 (493)
6 PRK01844 hypothetical protein; 92.7 0.31 6.7E-06 40.3 5.6 42 101-142 3-55 (72)
7 PRK00523 hypothetical protein; 92.6 0.32 6.9E-06 40.2 5.6 40 102-141 5-55 (72)
8 COG3763 Uncharacterized protei 90.9 0.57 1.2E-05 38.7 5.2 39 103-141 8-54 (71)
9 KOG0548 Molecular co-chaperone 89.8 0.32 6.9E-06 52.1 3.8 69 345-414 454-528 (539)
10 PF03672 UPF0154: Uncharacteri 87.5 1.6 3.5E-05 35.3 5.5 39 104-142 2-48 (64)
11 KOG0010 Ubiquitin-like protein 84.6 0.54 1.2E-05 50.0 1.9 25 278-304 155-179 (493)
12 KOG1924 RhoA GTPase effector D 70.7 8.5 0.00018 44.0 5.9 13 212-224 627-639 (1102)
13 KOG1308 Hsp70-interacting prot 58.0 0.99 2.1E-05 46.7 -3.7 43 364-408 275-317 (377)
14 PF13829 DUF4191: Domain of un 51.9 24 0.00053 34.5 4.7 38 98-142 49-87 (224)
15 KOG0548 Molecular co-chaperone 48.4 16 0.00034 39.8 3.1 48 365-412 127-174 (539)
16 PF13807 GNVR: G-rich domain o 48.1 33 0.00072 27.3 4.2 25 101-125 58-82 (82)
17 PF00042 Globin: Globin plant 47.0 39 0.00085 26.6 4.5 48 279-331 21-69 (110)
18 PF05957 DUF883: Bacterial pro 43.3 20 0.00044 29.2 2.4 22 100-121 72-93 (94)
19 PRK11677 hypothetical protein; 41.4 27 0.00059 31.6 3.1 21 102-122 3-23 (134)
20 PF07849 DUF1641: Protein of u 40.9 20 0.00044 26.3 1.8 18 376-393 15-32 (42)
21 PTZ00009 heat shock 70 kDa pro 39.4 58 0.0013 35.2 5.6 37 128-170 606-642 (653)
22 PF08370 PDR_assoc: Plant PDR 38.5 17 0.00038 29.3 1.3 15 99-113 25-39 (65)
23 PF11212 DUF2999: Protein of u 37.7 45 0.00098 28.4 3.6 26 340-365 5-44 (82)
24 KOG1924 RhoA GTPase effector D 37.4 46 0.001 38.5 4.6 8 154-161 577-584 (1102)
25 COG3877 Uncharacterized protei 36.6 70 0.0015 29.0 4.8 52 313-372 69-120 (122)
26 PF06305 DUF1049: Protein of u 36.6 1.1E+02 0.0024 23.1 5.3 34 100-133 18-52 (68)
27 PRK10132 hypothetical protein; 36.5 33 0.0007 29.9 2.7 22 100-121 85-106 (108)
28 COG3105 Uncharacterized protei 33.9 41 0.00088 31.1 3.0 23 101-123 7-29 (138)
29 PF04078 Rcd1: Cell differenti 33.8 30 0.00066 34.6 2.3 51 280-330 211-261 (262)
30 COG3763 Uncharacterized protei 33.3 84 0.0018 26.3 4.4 31 102-139 4-34 (71)
31 PRK10404 hypothetical protein; 31.3 44 0.00095 28.7 2.6 21 100-120 79-99 (101)
32 PF11075 DUF2780: Protein of u 30.9 42 0.00092 31.1 2.7 33 350-385 121-153 (163)
33 PF15050 SCIMP: SCIMP protein 30.0 3.4E+02 0.0074 25.1 8.1 37 128-169 55-92 (133)
34 PF06757 Ins_allergen_rp: Inse 29.6 70 0.0015 29.0 3.8 60 285-351 28-92 (179)
35 PHA00736 hypothetical protein 27.9 44 0.00096 28.1 2.0 17 101-118 56-72 (79)
36 PF03923 Lipoprotein_16: Uncha 27.2 44 0.00095 29.7 2.0 18 365-382 142-159 (159)
37 PF03960 ArsC: ArsC family; I 27.1 56 0.0012 27.0 2.5 57 313-378 29-87 (110)
38 KOG0011 Nucleotide excision re 27.0 3.9E+02 0.0084 28.1 8.9 41 280-324 214-259 (340)
39 PF10158 LOH1CR12: Tumour supp 25.8 20 0.00043 32.2 -0.4 63 295-357 2-72 (131)
40 PRK14581 hmsF outer membrane N 25.7 1.7E+02 0.0037 32.8 6.5 92 315-409 439-533 (672)
41 PF06295 DUF1043: Protein of u 25.5 59 0.0013 28.5 2.5 16 103-118 4-19 (128)
42 cd06199 SiR Cytochrome p450- l 24.8 60 0.0013 32.3 2.7 25 99-123 213-237 (360)
43 PF10474 DUF2451: Protein of u 24.8 44 0.00096 32.1 1.7 31 298-332 189-220 (234)
44 COG2838 Icd Monomeric isocitra 24.2 51 0.0011 36.5 2.2 60 352-411 283-342 (744)
45 PRK07983 exodeoxyribonuclease 24.1 1.1E+02 0.0024 28.9 4.2 49 278-327 157-218 (219)
46 KOG2366 Alpha-D-galactosidase 23.8 1.3E+02 0.0027 32.3 4.8 74 307-384 228-317 (414)
47 PF07301 DUF1453: Protein of u 23.6 82 0.0018 29.2 3.1 26 98-123 53-78 (148)
48 KOG3341 RNA polymerase II tran 23.4 65 0.0014 32.2 2.6 22 312-333 56-77 (249)
49 smart00845 GatB_Yqey GatB doma 23.2 91 0.002 27.4 3.2 68 339-411 47-120 (147)
50 COG2427 Uncharacterized conser 21.8 63 0.0014 29.1 2.0 82 303-395 50-137 (148)
51 PF01323 DSBA: DSBA-like thior 21.4 1.3E+02 0.0029 25.7 3.8 38 345-383 117-155 (193)
52 cd02977 ArsC_family Arsenate R 21.3 47 0.001 27.0 1.0 17 361-377 73-89 (105)
53 TIGR00601 rad23 UV excision re 21.2 1.6E+02 0.0034 30.6 4.9 32 281-326 247-278 (378)
54 PF07237 DUF1428: Protein of u 20.9 72 0.0016 28.0 2.1 20 279-298 79-100 (103)
55 KOG2629 Peroxisomal membrane a 20.7 2.1E+02 0.0044 29.6 5.5 56 72-127 53-109 (300)
56 PF02285 COX8: Cytochrome oxid 20.5 1.2E+02 0.0027 23.2 3.0 29 98-126 10-41 (44)
57 KOG2051 Nonsense-mediated mRNA 20.2 89 0.0019 36.9 3.2 101 310-412 571-689 (1128)
58 TIGR02384 RelB_DinJ addiction 20.1 1.7E+02 0.0036 24.2 3.9 28 383-410 53-80 (83)
No 1
>KOG1308 consensus Hsp70-interacting protein Hip/Transient component of progesterone receptor complexes and an Hsp70-binding protein [Posttranslational modification, protein turnover, chaperones; Signal transduction mechanisms]
Probab=100.00 E-value=1.8e-41 Score=333.20 Aligned_cols=310 Identities=25% Similarity=0.254 Sum_probs=223.0
Q ss_pred CCccccceeeeecCCCccccccccCCCCCCCCCCCCCCCchhhHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHhhcCCC
Q 015059 62 GQVGAGGFASLTSSGGQQTSSVGVNPNLPMPPPSSNVGSPLFWVGVGVGLSALFSFVASRLKQYAMQQALKGMMNQMNTQ 141 (414)
Q Consensus 62 ~~~~~~~fas~sss~~~~~~s~g~~p~~~~pp~~s~iGsPl~WiGvGVgLsalfs~V~~~VK~yaMQqamKsMM~Qmg~~ 141 (414)
.+-+++.|+++++++.- ++-...+++++|-.+.++.++||++.+|++..+++.|-.-.++|.+++= +.++
T Consensus 51 ~~e~~k~e~~~~~~~ee---~~~~~e~s~~~~~~~~d~egviepd~d~pq~MGds~~e~Tee~~eqa~e-------~k~~ 120 (377)
T KOG1308|consen 51 SEENTKAEASISKSVEE---SLKAPEVSSPESDLEIDGEGVIEPDTDAPQEMGDSNAEITEEMMDQAND-------KKVQ 120 (377)
T ss_pred ccccccccCCccccccc---ccccCCCCCCCcchhccCCCccccCCCcchhhchhhhhhhHHHHHHHHH-------HHHH
Confidence 66789999999887443 7777777775555568999999999999999999999999999998873 3343
Q ss_pred CCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCcccccCCccccccccccccccccccCcccccc---CCccccccc
Q 015059 142 NKPFGNAAFPQGSPFPFPNPPASGPTTPYPAASQPRFTMDIPATKVEAATATDVEGKKEVKGETEVKE---EPKKYAFVD 218 (414)
Q Consensus 142 ~~~fG~~pf~~gsPFpfp~Pp~~gp~~~~~~as~~~~tvD~~At~v~a~~~t~~~~~~e~~~~~e~~~---e~Kk~AFvD 218 (414)
+-.|-+.++ | ...-..++ .+++...++.--.+--.+. -.+.++ +-+.++|-+
T Consensus 121 A~eAln~G~-------~-----~~ai~~~t----~ai~lnp~~a~l~~kr~sv---------~lkl~kp~~airD~d~A~ 175 (377)
T KOG1308|consen 121 ASEALNDGE-------F-----DTAIELFT----SAIELNPPLAILYAKRASV---------FLKLKKPNAAIRDCDFAI 175 (377)
T ss_pred HHHHhcCcc-------h-----hhhhcccc----cccccCCchhhhcccccce---------eeeccCCchhhhhhhhhh
Confidence 333333333 0 00000000 1111111111111100010 112333 336788888
Q ss_pred CCchhhhhcccccccccccccCCCCCCCCC-CCCCCCCCCCCCCCCCC--CCCCccccCccccHHHHHHhhcChhhhhhh
Q 015059 219 VSPEETLQKSSFDNFEDVKETSSSKDAQPP-KDSQNGAAFNYNAGSPF--GGQSAKKEGRFLTVDTLEKLMEDPQVQKMV 295 (414)
Q Consensus 219 Vspee~~~k~~~~s~~d~~~~ss~k~~~~~-~~~~~g~~~~~g~g~~~--~~~~~~~~~~~~~v~~~e~M~~dP~~Qkm~ 295 (414)
....++-+..+|......-.+...+.+.-- .+.+++..-..++-... -..-....++...+...++++.++.+|+++
T Consensus 176 ein~Dsa~~ykfrg~A~rllg~~e~aa~dl~~a~kld~dE~~~a~lKeV~p~a~ki~e~~~k~er~~~e~~~~~r~er~r 255 (377)
T KOG1308|consen 176 EINPDSAKGYKFRGYAERLLGNWEEAAHDLALACKLDYDEANSATLKEVFPNAGKIEEHRRKYERAREEREIKERVERVR 255 (377)
T ss_pred ccCcccccccchhhHHHHHhhchHHHHHHHHHHHhccccHHHHHHHHHhccchhhhhhchhHHHHHHHHhcccccccccc
Confidence 888887777777766555444444222211 11111111100000000 001123446668899999999999999999
Q ss_pred ccCCCcccCChHHHHHHhhChHHHHHHHHHHHhccCCCCCchhhhhhcccCCCCcHHHHHHHHHcCCCcHHHHHHhhCCH
Q 015059 296 YPSLPEEMRNPASFKLMLQNPEYRKQLQEMLDGMCESGEFDGRVLDSLKNFDLNSAEVKQQFEQIGLTPEEVITKMMANP 375 (414)
Q Consensus 296 ypyLPe~MRNp~tfk~ml~NP~yr~QLe~ml~~mg~~~~~~~~m~D~lk~~d~~spev~~qF~q~GmtP~e~~skIm~DP 375 (414)
|+|+|++|||+++|+|+++|++||+++.+|...|+++..|+.+|.|.|++||+|++ +++||-|+| |+|||+||
T Consensus 256 ~~r~~~e~~~~e~~k~~~~~~~~~~~~g~~p~~M~g~~~~~~~m~~~m~~~~~n~~-~~~~p~~~g------i~ki~~dp 328 (377)
T KOG1308|consen 256 YAREPEEMANPEEFKRMLKNPQYRQFLGGFPGGMPGSFPGDKRMTDGMKGFDGNSP-VKQQPNQIG------ISKILSDP 328 (377)
T ss_pred cccchhhhcChhhhhhhhccCCCCcccCCCcccCCCCCCCccccccccccCCCCCc-cccCCCccc------HhhhcCch
Confidence 99999999999999999999999999999999999999999999999999999999 999999999 99999999
Q ss_pred HHHhhcCCHHHHHHHHHHhcChhhhhhhccChhhhccC
Q 015059 376 EIALGFQSPRVQAAIMECSQNPMNIIKYQNDKEVFSDF 413 (414)
Q Consensus 376 ELlaAfQDPEVmaA~qDi~sNPaNisKYqnnPKVmnl~ 413 (414)
||++||||||||+|||||++||+||+|||||||||+||
T Consensus 329 ev~aAfqdp~v~aal~d~~~np~n~~kyq~n~kv~~~i 366 (377)
T KOG1308|consen 329 EVAAAFQDPEVQAALMDVSQNPANMMKYQNNPKVMDVI 366 (377)
T ss_pred HHHHhhcChHHHhhhhhcccChHHHHHhccChHHHHHH
Confidence 99999999999999999999999999999999999998
No 2
>smart00727 STI1 Heat shock chaperonin-binding motif.
Probab=98.17 E-value=1.7e-06 Score=60.83 Aligned_cols=40 Identities=28% Similarity=0.512 Sum_probs=38.1
Q ss_pred CHHHHhhcCCHHHHHHHHHHhcChhhhhhhcc-ChhhhccC
Q 015059 374 NPEIALGFQSPRVQAAIMECSQNPMNIIKYQN-DKEVFSDF 413 (414)
Q Consensus 374 DPELlaAfQDPEVmaA~qDi~sNPaNisKYqn-nPKVmnl~ 413 (414)
|||+...++||+|+.+++++++||..+.+|.. ||++++.|
T Consensus 1 dP~~~~~l~~P~~~~~l~~~~~nP~~~~~~~~~nP~~~~~i 41 (41)
T smart00727 1 DPEMALRLQNPQVQSLLQDMQQNPDMLAQMLQENPQLLQLI 41 (41)
T ss_pred CHHHHHHHcCHHHHHHHHHHHHCHHHHHHHHHhCHHhHhhC
Confidence 79999999999999999999999999999999 99998865
No 3
>KOG3037 consensus Cell membrane glycoprotein [General function prediction only]
Probab=96.18 E-value=0.013 Score=59.20 Aligned_cols=56 Identities=20% Similarity=0.456 Sum_probs=48.1
Q ss_pred cCccccHHHHHHhhcChhhhhhhccCCCcccCChHHHHHHhhChHHHHHHHHHHHh
Q 015059 273 EGRFLTVDTLEKLMEDPQVQKMVYPSLPEEMRNPASFKLMLQNPEYRKQLQEMLDG 328 (414)
Q Consensus 273 ~~~~~~v~~~e~M~~dP~~Qkm~ypyLPe~MRNp~tfk~ml~NP~yr~QLe~ml~~ 328 (414)
-...|..|++..++.+|.+|+-++||||+---+.+-+.-++++|||||+|.-....
T Consensus 225 La~vL~~e~v~~vl~~~~v~erL~phlP~d~~~~~~i~e~l~spqF~qal~sfs~a 280 (330)
T KOG3037|consen 225 LATVLKPEAVAPVLANPGVQERLMPHLPSDHDRAEGILELLTSPQFRQALDSFSQA 280 (330)
T ss_pred hhhhcChHHHHHHhhCcchhhhhcccCCCCCcchHHHHHhhcCHHHHHHHHHHHHH
Confidence 34458899999999999999999999999877778888899999999999655443
No 4
>smart00727 STI1 Heat shock chaperonin-binding motif.
Probab=95.58 E-value=0.0098 Score=41.76 Aligned_cols=37 Identities=30% Similarity=0.542 Sum_probs=29.3
Q ss_pred HHHhhcChhhhhhhccCCCcccCChHHHHHHhh-ChHHHHHH
Q 015059 282 LEKLMEDPQVQKMVYPSLPEEMRNPASFKLMLQ-NPEYRKQL 322 (414)
Q Consensus 282 ~e~M~~dP~~Qkm~ypyLPe~MRNp~tfk~ml~-NP~yr~QL 322 (414)
+..+++||.+|.++= +-++||+.+..|++ ||+++..+
T Consensus 4 ~~~~l~~P~~~~~l~----~~~~nP~~~~~~~~~nP~~~~~i 41 (41)
T smart00727 4 MALRLQNPQVQSLLQ----DMQQNPDMLAQMLQENPQLLQLI 41 (41)
T ss_pred HHHHHcCHHHHHHHH----HHHHCHHHHHHHHHhCHHhHhhC
Confidence 445667998888754 56779999999999 99998753
No 5
>KOG0010 consensus Ubiquitin-like protein [Posttranslational modification, protein turnover, chaperones; General function prediction only]
Probab=93.22 E-value=0.1 Score=55.23 Aligned_cols=71 Identities=27% Similarity=0.404 Sum_probs=53.1
Q ss_pred cCccccHHHHHHhhcCh-hhhhhhccCCCcccCChHHHHHHhhChHHHHHHHHHHHhccCCCCCchhhhhhcccCCCCcH
Q 015059 273 EGRFLTVDTLEKLMEDP-QVQKMVYPSLPEEMRNPASFKLMLQNPEYRKQLQEMLDGMCESGEFDGRVLDSLKNFDLNSA 351 (414)
Q Consensus 273 ~~~~~~v~~~e~M~~dP-~~Qkm~ypyLPe~MRNp~tfk~ml~NP~yr~QLe~ml~~mg~~~~~~~~m~D~lk~~d~~sp 351 (414)
.++-..--.++.+.+|| -+|+|+-||.+ +.|.-+.+||.+..|
T Consensus 328 ~~~~~~~a~lq~i~~n~~~~~~l~s~~~~------~m~~~~s~~P~~a~~------------------------------ 371 (493)
T KOG0010|consen 328 LGSPGMQAGLQMITENPSLLQQLLSPYIR------SMFQSASQNPLQAAQ------------------------------ 371 (493)
T ss_pred cCCcchhhhhhccccChhhhhhccchhhH------HHHhhhccCchhhhc------------------------------
Confidence 34445567789999999 45666666654 456678899988777
Q ss_pred HHHHHHHHcCCCcHHHHHHhhCCHHHHhhcCCHHHHHHHHHHhcC
Q 015059 352 EVKQQFEQIGLTPEEVITKMMANPEIALGFQSPRVQAAIMECSQN 396 (414)
Q Consensus 352 ev~~qF~q~GmtP~e~~skIm~DPELlaAfQDPEVmaA~qDi~sN 396 (414)
|. . |.+|+++.+|.+|++|+||.-|-+-
T Consensus 372 -----~~-----------~-mq~p~~~~~~~np~a~~ai~qiqq~ 399 (493)
T KOG0010|consen 372 -----LR-----------Q-MQNPDVLRAMSNPRAMQAIRQIQQG 399 (493)
T ss_pred -----cc-----------c-ccCchHhhhhcChHHHHHHHHHHHH
Confidence 00 4 7799999999999999999987653
No 6
>PRK01844 hypothetical protein; Provisional
Probab=92.67 E-value=0.31 Score=40.26 Aligned_cols=42 Identities=29% Similarity=0.312 Sum_probs=26.5
Q ss_pred chhhHHHHH---HHHHHHHHHHH--HHHHHHH------HHHHHHHHhhcCCCC
Q 015059 101 PLFWVGVGV---GLSALFSFVAS--RLKQYAM------QQALKGMMNQMNTQN 142 (414)
Q Consensus 101 Pl~WiGvGV---gLsalfs~V~~--~VK~yaM------QqamKsMM~Qmg~~~ 142 (414)
-|+||+++| .+|++.++..+ ++++|.. |.|++.||.|||--+
T Consensus 3 ~~~~I~l~I~~li~G~~~Gff~ark~~~k~lk~NPpine~mir~Mm~QMGqkP 55 (72)
T PRK01844 3 IWLGILVGVVALVAGVALGFFIARKYMMNYLQKNPPINEQMLKMMMMQMGQKP 55 (72)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHCCCCCHHHHHHHHHHhCCCc
Confidence 367776543 33444444333 3566666 889999999999433
No 7
>PRK00523 hypothetical protein; Provisional
Probab=92.60 E-value=0.32 Score=40.20 Aligned_cols=40 Identities=15% Similarity=0.365 Sum_probs=25.6
Q ss_pred hhhHHHHHH---HHHHHHHHH--HHHHHHHH------HHHHHHHHhhcCCC
Q 015059 102 LFWVGVGVG---LSALFSFVA--SRLKQYAM------QQALKGMMNQMNTQ 141 (414)
Q Consensus 102 l~WiGvGVg---Lsalfs~V~--~~VK~yaM------QqamKsMM~Qmg~~ 141 (414)
++||+++|. .|++.++.. .++++|.. |.|++.||.|||--
T Consensus 5 ~l~I~l~i~~li~G~~~Gffiark~~~k~l~~NPpine~mir~M~~QMGqK 55 (72)
T PRK00523 5 GLALGLGIPLLIVGGIIGYFVSKKMFKKQIRENPPITENMIRAMYMQMGRK 55 (72)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHCcCCCHHHHHHHHHHhCCC
Confidence 566665433 333444433 33666766 89999999999943
No 8
>COG3763 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=90.87 E-value=0.57 Score=38.75 Aligned_cols=39 Identities=15% Similarity=0.265 Sum_probs=23.6
Q ss_pred hhHHHHHHHHHHHHHHHH--HHHHHHH------HHHHHHHHhhcCCC
Q 015059 103 FWVGVGVGLSALFSFVAS--RLKQYAM------QQALKGMMNQMNTQ 141 (414)
Q Consensus 103 ~WiGvGVgLsalfs~V~~--~VK~yaM------QqamKsMM~Qmg~~ 141 (414)
+||.+.+..|.+.++..+ ..|+|.+ ++|++.||.|||--
T Consensus 8 l~ivl~ll~G~~~G~fiark~~~k~lk~NPpine~~iR~M~~qmGqK 54 (71)
T COG3763 8 LLIVLALLAGLIGGFFIARKQMKKQLKDNPPINEEMIRMMMAQMGQK 54 (71)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHHhhCCCCCHHHHHHHHHHhCCC
Confidence 344444444444444433 2445555 89999999999943
No 9
>KOG0548 consensus Molecular co-chaperone STI1 [Posttranslational modification, protein turnover, chaperones]
Probab=89.82 E-value=0.32 Score=52.13 Aligned_cols=69 Identities=23% Similarity=0.376 Sum_probs=57.4
Q ss_pred cCCCCcHHHHHHHH------HcCCCcHHHHHHhhCCHHHHhhcCCHHHHHHHHHHhcChhhhhhhccChhhhccCC
Q 015059 345 NFDLNSAEVKQQFE------QIGLTPEEVITKMMANPEIALGFQSPRVQAAIMECSQNPMNIIKYQNDKEVFSDFV 414 (414)
Q Consensus 345 ~~d~~spev~~qF~------q~GmtP~e~~skIm~DPELlaAfQDPEVmaA~qDi~sNPaNisKYqnnPKVmnl~e 414 (414)
..|-++-|+-+.+. +.-.+|+++....|.|||+.+.+|||-....+.++-+|| +..|+..||.|++.|+
T Consensus 454 e~dp~~~e~~~~~~rc~~a~~~~~~~ee~~~r~~~dpev~~il~d~~m~~~l~q~q~~p-a~~~~~~n~~v~~ki~ 528 (539)
T KOG0548|consen 454 ELDPSNAEAIDGYRRCVEAQRGDETPEETKRRAMADPEVQAILQDPAMRQILEQMQENP-ALQEHLKNPMVMQKIE 528 (539)
T ss_pred hcCchhHHHHHHHHHHHHHhhcCCCHHHHHHhhccCHHHHHHHcCHHHHHHHHHHHhCH-HHHHHHhccHHHHHHH
Confidence 34555667766662 346799999999999999999999999999999999999 6669999999987653
No 10
>PF03672 UPF0154: Uncharacterised protein family (UPF0154); InterPro: IPR005359 The proteins in this entry are functionally uncharacterised.
Probab=87.52 E-value=1.6 Score=35.32 Aligned_cols=39 Identities=15% Similarity=0.290 Sum_probs=26.7
Q ss_pred hHHHHHHHHHHHHHHH--HHHHHHHH------HHHHHHHHhhcCCCC
Q 015059 104 WVGVGVGLSALFSFVA--SRLKQYAM------QQALKGMMNQMNTQN 142 (414)
Q Consensus 104 WiGvGVgLsalfs~V~--~~VK~yaM------QqamKsMM~Qmg~~~ 142 (414)
||-|++.+|++.+|.. -++++|.. +.|++.||.|||--+
T Consensus 2 ~iilali~G~~~Gff~ar~~~~k~l~~NPpine~mir~M~~QMG~kp 48 (64)
T PF03672_consen 2 LIILALIVGAVIGFFIARKYMEKQLKENPPINEKMIRAMMMQMGRKP 48 (64)
T ss_pred hHHHHHHHHHHHHHHHHHHHHHHHHHHCCCCCHHHHHHHHHHhCCCc
Confidence 4555555555555554 45667776 899999999999433
No 11
>KOG0010 consensus Ubiquitin-like protein [Posttranslational modification, protein turnover, chaperones; General function prediction only]
Probab=84.56 E-value=0.54 Score=50.03 Aligned_cols=25 Identities=36% Similarity=0.687 Sum_probs=19.2
Q ss_pred cHHHHHHhhcChhhhhhhccCCCcccC
Q 015059 278 TVDTLEKLMEDPQVQKMVYPSLPEEMR 304 (414)
Q Consensus 278 ~v~~~e~M~~dP~~Qkm~ypyLPe~MR 304 (414)
.-|.+-.||+||-+|.|+=. |+.||
T Consensus 155 npe~~~~~m~nP~vq~ll~N--pd~mr 179 (493)
T KOG0010|consen 155 NPEALRQMMENPIVQSLLNN--PDLMR 179 (493)
T ss_pred CHHHHHHhhhChHHHHHhcC--hHHHH
Confidence 46888999999999998654 44444
No 12
>KOG1924 consensus RhoA GTPase effector DIA/Diaphanous [Signal transduction mechanisms; Cytoskeleton]
Probab=70.73 E-value=8.5 Score=43.98 Aligned_cols=13 Identities=8% Similarity=0.636 Sum_probs=5.8
Q ss_pred cccccccCCchhh
Q 015059 212 KKYAFVDVSPEET 224 (414)
Q Consensus 212 Kk~AFvDVspee~ 224 (414)
|++-+--+.|.++
T Consensus 627 rr~nW~kI~p~d~ 639 (1102)
T KOG1924|consen 627 RRFNWSKIVPRDL 639 (1102)
T ss_pred ccCCccccCcccc
Confidence 3443444445544
No 13
>KOG1308 consensus Hsp70-interacting protein Hip/Transient component of progesterone receptor complexes and an Hsp70-binding protein [Posttranslational modification, protein turnover, chaperones; Signal transduction mechanisms]
Probab=57.95 E-value=0.99 Score=46.71 Aligned_cols=43 Identities=7% Similarity=-0.179 Sum_probs=35.4
Q ss_pred cHHHHHHhhCCHHHHhhcCCHHHHHHHHHHhcChhhhhhhccChh
Q 015059 364 PEEVITKMMANPEIALGFQSPRVQAAIMECSQNPMNIIKYQNDKE 408 (414)
Q Consensus 364 P~e~~skIm~DPELlaAfQDPEVmaA~qDi~sNPaNisKYqnnPK 408 (414)
..+..+.+.++|+.|.++++++++ +++||..||.|.. |..+|+
T Consensus 275 ~~~~~~~~g~~p~~M~g~~~~~~~-m~~~m~~~~~n~~-~~~~p~ 317 (377)
T KOG1308|consen 275 NPQYRQFLGGFPGGMPGSFPGDKR-MTDGMKGFDGNSP-VKQQPN 317 (377)
T ss_pred cCCCCcccCCCcccCCCCCCCccc-cccccccCCCCCc-cccCCC
Confidence 344577889999999999999999 9999999999977 444443
No 14
>PF13829 DUF4191: Domain of unknown function (DUF4191)
Probab=51.94 E-value=24 Score=34.48 Aligned_cols=38 Identities=26% Similarity=0.462 Sum_probs=25.9
Q ss_pred CCCchhhHHHHHHHHHHHHHH-HHHHHHHHHHHHHHHHHhhcCCCC
Q 015059 98 VGSPLFWVGVGVGLSALFSFV-ASRLKQYAMQQALKGMMNQMNTQN 142 (414)
Q Consensus 98 iGsPl~WiGvGVgLsalfs~V-~~~VK~yaMQqamKsMM~Qmg~~~ 142 (414)
+|++|+|+=+||.|+++++++ ++ ..+=|.+..|+-||+
T Consensus 49 ~~~~~~~~i~gi~~g~l~am~vl~-------rra~ra~Y~qieGqp 87 (224)
T PF13829_consen 49 FGSWWYWLIIGILLGLLAAMIVLS-------RRAQRAAYAQIEGQP 87 (224)
T ss_pred HccHHHHHHHHHHHHHHHHHHHHH-------HHHHHHHHHHhcCCC
Confidence 567899999999999888763 22 233455556666665
No 15
>KOG0548 consensus Molecular co-chaperone STI1 [Posttranslational modification, protein turnover, chaperones]
Probab=48.40 E-value=16 Score=39.82 Aligned_cols=48 Identities=17% Similarity=0.234 Sum_probs=44.9
Q ss_pred HHHHHHhhCCHHHHhhcCCHHHHHHHHHHhcChhhhhhhccChhhhcc
Q 015059 365 EEVITKMMANPEIALGFQSPRVQAAIMECSQNPMNIIKYQNDKEVFSD 412 (414)
Q Consensus 365 ~e~~skIm~DPELlaAfQDPEVmaA~qDi~sNPaNisKYqnnPKVmnl 412 (414)
..++.++-+||.....++||-++.-++.|-+||.++.-|-+||-+|..
T Consensus 127 p~~~~~l~~~p~t~~~~~~~~~~~~l~~~~~~p~~l~~~l~d~r~m~a 174 (539)
T KOG0548|consen 127 PYFHEKLANLPLTNYSLSDPAYVKILEIIQKNPTSLKLYLNDPRLMKA 174 (539)
T ss_pred cHHHHHhhcChhhhhhhccHHHHHHHHHhhcCcHhhhcccccHHHHHH
Confidence 348999999999999999999999999999999999999999999864
No 16
>PF13807 GNVR: G-rich domain on putative tyrosine kinase
Probab=48.10 E-value=33 Score=27.28 Aligned_cols=25 Identities=12% Similarity=0.220 Sum_probs=22.0
Q ss_pred chhhHHHHHHHHHHHHHHHHHHHHH
Q 015059 101 PLFWVGVGVGLSALFSFVASRLKQY 125 (414)
Q Consensus 101 Pl~WiGvGVgLsalfs~V~~~VK~y 125 (414)
.++++.+|+.+|.++|.+.-.+|++
T Consensus 58 ~~lil~l~~~~Gl~lgi~~~~~re~ 82 (82)
T PF13807_consen 58 RALILALGLFLGLILGIGLAFLREM 82 (82)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHhhC
Confidence 4899999999999999999888863
No 17
>PF00042 Globin: Globin plant globin signature erythrocruorin family signature alpha hemoglobin signature myoglobin signature thalassemia.; InterPro: IPR000971 Globins are haem-containing proteins involved in binding and/or transporting oxygen. They belong to a very large and well studied family that is widely distributed in many organisms []. Globins have evolved from a common ancestor and can be divided into three groups: single-domain globins, and two types of chimeric globins, flavohaemoglobins and globin-coupled sensors. Bacteria have all three types of globins, while archaea lack flavohaemoglobins, and eukaryotes lack globin-coupled sensors []. Several functionally different haemoglobins can coexist in the same species. The major types of globins include: Haemoglobin (Hb): trimer of two alpha and two beta chains, although embryonic and foetal forms can substitute the alpha or beta chain for ones with higher oxygen affinity, such as gamma, delta, epsilon or zeta chains. Hb transports oxygen from lungs to other tissues in vertebrates []. Hb proteins are also present in unicellular organisms where they act as enzymes or sensors []. Myoglobin (Mb): monomeric protein responsible for oxygen storage in vertebrate muscle []. Neuroglobin: a myoglobin-like haemprotein expressed in vertebrate brain and retina, where it is involved in neuroprotection from damage due to hypoxia or ischemia []. Neuroglobin belongs to a branch of the globin family that diverged early in evolution. Cytoglobin: an oxygen sensor expressed in multiple tissues. Related to neuroglobin []. Erythrocruorin: highly cooperative extracellular respiratory proteins found in annelids and arthropods that are assembled from as many as 180 subunit into hexagonal bilayers []. Leghaemoglobin (legHb or symbiotic Hb): occurs in the root nodules of leguminous plants, where it facilitates the diffusion of oxygen to symbiotic bacteriods in order to promote nitrogen fixation. Non-symbiotic haemoglobin (NsHb): occurs in non-leguminous plants, and can be over-expressed in stressed plants []. Flavohaemoglobins (FHb): chimeric, with an N-terminal globin domain and a C-terminal ferredoxin reductase-like NAD/FAD-binding domain. FHb provides protection against nitric oxide via its C-terminal domain, which transfers electrons to haem in the globin []. Globin-coupled sensors: chimeric, with an N-terminal myoglobin-like domain and a C-terminal domain that resembles the cytoplasmic signalling domain of bacterial chemoreceptors. They bind oxygen, and act to initiate an aerotactic response or regulate gene expression [, ]. Protoglobin: a single domain globin found in archaea that is related to the N-terminal domain of globin-coupled sensors []. Truncated 2/2 globin: lack the first helix, giving them a 2-over-2 instead of the canonical 3-over-3 alpha-helical sandwich fold. Can be divided into three main groups (I, II and II) based on structural features []. This entry covers most of the globin family of proteins, but it omits some bacterial globins and the protoglobins. More information about these proteins can be found at Protein of the Month: Haemoglobin [].; GO: 0005506 iron ion binding, 0020037 heme binding; PDB: 2WTH_A 2WTG_A 3A59_G 3FS4_C 3CY5_C 3D1A_B 2RI4_J 3EU1_B 1JEB_A 2Z6N_B ....
Probab=47.00 E-value=39 Score=26.63 Aligned_cols=48 Identities=19% Similarity=0.430 Sum_probs=35.9
Q ss_pred HHHHHHhhc-ChhhhhhhccCCCcccCChHHHHHHhhChHHHHHHHHHHHhccC
Q 015059 279 VDTLEKLME-DPQVQKMVYPSLPEEMRNPASFKLMLQNPEYRKQLQEMLDGMCE 331 (414)
Q Consensus 279 v~~~e~M~~-dP~~Qkm~ypyLPe~MRNp~tfk~ml~NP~yr~QLe~ml~~mg~ 331 (414)
.+...++++ +|++|++...+ ++-.+.+-+.+|+.++.|-..++.-.+.
T Consensus 21 ~~~f~~lF~~~P~~~~~F~~~-----~~~~~~~~l~~~~~~~~h~~~v~~~l~~ 69 (110)
T PF00042_consen 21 SEFFQRLFEEYPDYKKLFPKF-----KDIVPLEELKNNPEFKAHAQRVMEALDE 69 (110)
T ss_dssp HHHHHHHHHHSGGGGGGGTTG-----TTTSSHHHHTTSHHHHHHHHHHHHHHHH
T ss_pred HHHHHHHHHHCHHHHhhcccc-----cccchHHHHhccchHHHHHHHHHHHHHH
Confidence 455566665 99998876322 6666788999999999999888887543
No 18
>PF05957 DUF883: Bacterial protein of unknown function (DUF883); InterPro: IPR010279 This family consists of several bacterial proteins of unknown function that include the Escherichia coli genes for ElaB, YgaM and YqjD.
Probab=43.31 E-value=20 Score=29.20 Aligned_cols=22 Identities=27% Similarity=0.540 Sum_probs=18.9
Q ss_pred CchhhHHHHHHHHHHHHHHHHH
Q 015059 100 SPLFWVGVGVGLSALFSFVASR 121 (414)
Q Consensus 100 sPl~WiGvGVgLsalfs~V~~~ 121 (414)
.||-=|||.+|+|.|+|+.+.+
T Consensus 72 ~P~~svgiAagvG~llG~Ll~R 93 (94)
T PF05957_consen 72 NPWQSVGIAAGVGFLLGLLLRR 93 (94)
T ss_pred ChHHHHHHHHHHHHHHHHHHhC
Confidence 5888899999999999998753
No 19
>PRK11677 hypothetical protein; Provisional
Probab=41.37 E-value=27 Score=31.57 Aligned_cols=21 Identities=19% Similarity=0.284 Sum_probs=19.2
Q ss_pred hhhHHHHHHHHHHHHHHHHHH
Q 015059 102 LFWVGVGVGLSALFSFVASRL 122 (414)
Q Consensus 102 l~WiGvGVgLsalfs~V~~~V 122 (414)
|+++.||+++|+++|++..++
T Consensus 3 W~~a~i~livG~iiG~~~~R~ 23 (134)
T PRK11677 3 WEYALIGLVVGIIIGAVAMRF 23 (134)
T ss_pred HHHHHHHHHHHHHHHHHHHhh
Confidence 888899999999999999887
No 20
>PF07849 DUF1641: Protein of unknown function (DUF1641); InterPro: IPR012440 Archaeal and bacterial hypothetical proteins are found in this family, with the region in question being approximately 40 residues long.
Probab=40.89 E-value=20 Score=26.27 Aligned_cols=18 Identities=17% Similarity=0.193 Sum_probs=13.8
Q ss_pred HHHhhcCCHHHHHHHHHH
Q 015059 376 EIALGFQSPRVQAAIMEC 393 (414)
Q Consensus 376 ELlaAfQDPEVmaA~qDi 393 (414)
+|+.++.||+|+.++-=+
T Consensus 15 gl~~~l~DpdvqrgL~~l 32 (42)
T PF07849_consen 15 GLLRALRDPDVQRGLGFL 32 (42)
T ss_pred HHHHHHcCHHHHHHHHHH
Confidence 577888888888887543
No 21
>PTZ00009 heat shock 70 kDa protein; Provisional
Probab=39.39 E-value=58 Score=35.20 Aligned_cols=37 Identities=14% Similarity=0.097 Sum_probs=21.7
Q ss_pred HHHHHHHHhhcCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCCC
Q 015059 128 QQALKGMMNQMNTQNKPFGNAAFPQGSPFPFPNPPASGPTTPY 170 (414)
Q Consensus 128 QqamKsMM~Qmg~~~~~fG~~pf~~gsPFpfp~Pp~~gp~~~~ 170 (414)
+..+..|+.+++|+ +|.|...+||+.||++++|+.-.
T Consensus 606 ~pi~~r~~~~~~~~------~~~~~~~~~~~~~~~~~~~~~~~ 642 (653)
T PTZ00009 606 NPIMTKMYQAAGGG------MPGGMPGGMPGGMPGGAGPAGAG 642 (653)
T ss_pred HHHHHHHHhhccCC------CCCCCCCCCCCCCCCCCCCCCCC
Confidence 34445566666643 33334456777788777776544
No 22
>PF08370 PDR_assoc: Plant PDR ABC transporter associated; InterPro: IPR013581 ABC transporters belong to the ATP-Binding Cassette (ABC) superfamily, which uses the hydrolysis of ATP to energise diverse biological systems. ABC transporters minimally consist of two conserved regions: a highly conserved ATP binding cassette (ABC) and a less conserved transmembrane domain (TMD). These can be found on the same protein or on two different ones. Most ABC transporters function as a dimer and therefore are constituted of four domains, two ABC modules and two TMDs. ABC transporters are involved in the export or import of a wide variety of substrates ranging from small ions to macromolecules. The major function of ABC import systems is to provide essential nutrients to bacteria. They are found only in prokaryotes and their four constitutive domains are usually encoded by independent polypeptides (two ABC proteins and two TMD proteins). Prokaryotic importers require additional extracytoplasmic binding proteins (one or more per systems) for function. In contrast, export systems are involved in the extrusion of noxious substances, the export of extracellular toxins and the targeting of membrane components. They are found in all living organisms and in general the TMD is fused to the ABC module in a variety of combinations. Some eukaryotic exporters encode the four domains on the same polypeptide chain []. The ABC module (approximately two hundred amino acid residues) is known to bind and hydrolyse ATP, thereby coupling transport to ATP hydrolysis in a large number of biological processes. The cassette is duplicated in several subfamilies. Its primary sequence is highly conserved, displaying a typical phosphate-binding loop: Walker A, and a magnesium binding site: Walker B. Besides these two regions, three other conserved motifs are present in the ABC cassette: the switch region which contains a histidine loop, postulated to polarise the attaching water molecule for hydrolysis, the signature conserved motif (LSGGQ) specific to the ABC transporter, and the Q-motif (between Walker A and the signature), which interacts with the gamma phosphate through a water bond. The Walker A, Walker B, Q-loop and switch region form the nucleotide binding site [, , ]. The 3D structure of a monomeric ABC module adopts a stubby L-shape with two distinct arms. ArmI (mainly beta-strand) contains Walker A and Walker B. The important residues for ATP hydrolysis and/or binding are located in the P-loop. The ATP-binding pocket is located at the extremity of armI. The perpendicular armII contains mostly the alpha helical subdomain with the signature motif. It only seems to be required for structural integrity of the ABC module. ArmII is in direct contact with the TMD. The hinge between armI and armII contains both the histidine loop and the Q-loop, making contact with the gamma phosphate of the ATP molecule. ATP hydrolysis leads to a conformational change that could facilitate ADP release. In the dimer the two ABC cassettes contact each other through hydrophobic interactions at the antiparallel beta-sheet of armI by a two-fold axis [, , , , , ]. The ATP-Binding Cassette (ABC) superfamily forms one of the largest of all protein families with a diversity of physiological functions []. Several studies have shown that there is a correlation between the functional characterisation and the phylogenetic classification of the ABC cassette [, ]. More than 50 subfamilies have been described based on a phylogenetic and functional classification [, , ]; (for further information see http://www.tcdb.org/tcdb/index.php?tc=3.A.1). This domain is found on the C terminus of ABC-2 type transporter domains (IPR013525 from INTERPRO). It seems to be associated with the plant pleiotropic drug resistance (PDR) protein family of ABC transporters. Like in yeast, plant PDR ABC transporters may also play a role in the transport of antifungal agents [] (see also IPR010929 from INTERPRO). The PDR family is characterised by a configuration in which the ABC domain is nearer the N terminus of the protein than the transmembrane domain [].
Probab=38.52 E-value=17 Score=29.30 Aligned_cols=15 Identities=40% Similarity=0.764 Sum_probs=11.4
Q ss_pred CCchhhHHHHHHHHH
Q 015059 99 GSPLFWVGVGVGLSA 113 (414)
Q Consensus 99 GsPl~WiGvGVgLsa 113 (414)
..-|.|||||+-+|-
T Consensus 25 ~~~WyWIgvgaL~G~ 39 (65)
T PF08370_consen 25 ESYWYWIGVGALLGF 39 (65)
T ss_pred CCcEEeehHHHHHHH
Confidence 345999999987763
No 23
>PF11212 DUF2999: Protein of unknown function (DUF2999); InterPro: IPR021376 This family of proteins with unknown function appears to be restricted to Gammaproteobacteria.
Probab=37.73 E-value=45 Score=28.38 Aligned_cols=26 Identities=23% Similarity=0.602 Sum_probs=16.6
Q ss_pred hhhcccCCCCcHHHHHHH--------------HHcCCCcH
Q 015059 340 LDSLKNFDLNSAEVKQQF--------------EQIGLTPE 365 (414)
Q Consensus 340 ~D~lk~~d~~spev~~qF--------------~q~GmtP~ 365 (414)
...||+.+++-..+++-| .|+||+++
T Consensus 5 ia~LKehnvsd~qi~elFq~lT~NPl~AMa~i~qLGip~e 44 (82)
T PF11212_consen 5 IAILKEHNVSDEQINELFQALTQNPLAAMATIQQLGIPQE 44 (82)
T ss_pred HHHHHHcCCCHHHHHHHHHHHhhCHHHHHHHHHHcCCCHH
Confidence 345666666666666666 56777766
No 24
>KOG1924 consensus RhoA GTPase effector DIA/Diaphanous [Signal transduction mechanisms; Cytoskeleton]
Probab=37.36 E-value=46 Score=38.46 Aligned_cols=8 Identities=25% Similarity=0.048 Sum_probs=3.4
Q ss_pred CCCCCCCC
Q 015059 154 SPFPFPNP 161 (414)
Q Consensus 154 sPFpfp~P 161 (414)
++|+||+|
T Consensus 577 pg~~gppP 584 (1102)
T KOG1924|consen 577 PGGGGPPP 584 (1102)
T ss_pred CCCCCCCC
Confidence 34444443
No 25
>COG3877 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=36.63 E-value=70 Score=28.99 Aligned_cols=52 Identities=29% Similarity=0.463 Sum_probs=38.7
Q ss_pred hhChHHHHHHHHHHHhccCCCCCchhhhhhcccCCCCcHHHHHHHHHcCCCcHHHHHHhh
Q 015059 313 LQNPEYRKQLQEMLDGMCESGEFDGRVLDSLKNFDLNSAEVKQQFEQIGLTPEEVITKMM 372 (414)
Q Consensus 313 l~NP~yr~QLe~ml~~mg~~~~~~~~m~D~lk~~d~~spev~~qF~q~GmtP~e~~skIm 372 (414)
++-|.+|..|+++|..||.-+ +..|.. |+.-..+-+|+++=-|+|+|-+ ++|
T Consensus 69 ~sYptvR~kld~vlramgy~p--~~e~~~-----~i~~~~i~~qle~Gei~peeA~-~~L 120 (122)
T COG3877 69 ISYPTVRTKLDEVLRAMGYNP--DSENSV-----NIGKKKIIDQLEKGEISPEEAI-KML 120 (122)
T ss_pred CccHHHHHHHHHHHHHcCCCC--CCCChh-----hhhHHHHHHHHHcCCCCHHHHH-HHh
Confidence 567999999999999999855 222211 3445568888888899999887 444
No 26
>PF06305 DUF1049: Protein of unknown function (DUF1049); InterPro: IPR010445 This entry consists of several hypothetical bacterial proteins of unknown function.
Probab=36.59 E-value=1.1e+02 Score=23.09 Aligned_cols=34 Identities=15% Similarity=0.258 Sum_probs=23.9
Q ss_pred Cc-hhhHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Q 015059 100 SP-LFWVGVGVGLSALFSFVASRLKQYAMQQALKG 133 (414)
Q Consensus 100 sP-l~WiGvGVgLsalfs~V~~~VK~yaMQqamKs 133 (414)
-| .+||.+-.++|+++++.+...+.+-...-.|.
T Consensus 18 ~pl~l~il~~f~~G~llg~l~~~~~~~~~r~~~~~ 52 (68)
T PF06305_consen 18 LPLGLLILIAFLLGALLGWLLSLPSRLRLRRRIRR 52 (68)
T ss_pred chHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 45 67888888888888888887776554433333
No 27
>PRK10132 hypothetical protein; Provisional
Probab=36.47 E-value=33 Score=29.93 Aligned_cols=22 Identities=18% Similarity=0.256 Sum_probs=19.1
Q ss_pred CchhhHHHHHHHHHHHHHHHHH
Q 015059 100 SPLFWVGVGVGLSALFSFVASR 121 (414)
Q Consensus 100 sPl~WiGvGVgLsalfs~V~~~ 121 (414)
.||-=|||+.|+|.|+|+...+
T Consensus 85 ~Pw~svgiaagvG~llG~Ll~R 106 (108)
T PRK10132 85 RPWCSVGTAAAVGIFIGALLSL 106 (108)
T ss_pred CcHHHHHHHHHHHHHHHHHHhc
Confidence 6899999999999999988664
No 28
>COG3105 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=33.94 E-value=41 Score=31.11 Aligned_cols=23 Identities=13% Similarity=0.258 Sum_probs=20.0
Q ss_pred chhhHHHHHHHHHHHHHHHHHHH
Q 015059 101 PLFWVGVGVGLSALFSFVASRLK 123 (414)
Q Consensus 101 Pl~WiGvGVgLsalfs~V~~~VK 123 (414)
+|.++|||+..|+++|+++.++-
T Consensus 7 ~W~~a~igLvvGi~IG~li~Rlt 29 (138)
T COG3105 7 TWEYALIGLVVGIIIGALIARLT 29 (138)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHc
Confidence 58889999999999999988765
No 29
>PF04078 Rcd1: Cell differentiation family, Rcd1-like ; InterPro: IPR007216 Rcd1 (Required cell differentiation 1) -like proteins are found among a wide range of organisms []. Rcd1 was initially identified as an essential factor in nitrogen starvation-invoked differentiation in fission yeast. This results largely from a defect in nitrogen starvation-invoked induction of ste11+, a key transcriptional factor gene required for the onset of sexual development. It is one of the most conserved proteins in eukaryotes, and its mammalian homologue is expressed in a variety of differentiating tissues [, ]. The mammalian Rcd1 is a novel transcriptional cofactor and is critical for retinoic acid-induced differentiation of F9 mouse teratocarcinoma cells, at least in part, via forming complexes with retinoic acid receptor and activation transcription factor-2 (ATF-2) []. Two of the members in this family have been characterised as being involved in regulation of Ste11 regulated sex genes [, ].; PDB: 2FV2_B.
Probab=33.78 E-value=30 Score=34.56 Aligned_cols=51 Identities=18% Similarity=0.458 Sum_probs=40.6
Q ss_pred HHHHHhhcChhhhhhhccCCCcccCChHHHHHHhhChHHHHHHHHHHHhcc
Q 015059 280 DTLEKLMEDPQVQKMVYPSLPEEMRNPASFKLMLQNPEYRKQLQEMLDGMC 330 (414)
Q Consensus 280 ~~~e~M~~dP~~Qkm~ypyLPe~MRNp~tfk~ml~NP~yr~QLe~ml~~mg 330 (414)
.--..+-+||.-..++=-+||+.+||.+.-.-+-.||..|+-|++++.|.+
T Consensus 211 rCYlRLsdnprar~aL~~~LP~~Lrd~~f~~~l~~D~~~k~~l~qLl~nl~ 261 (262)
T PF04078_consen 211 RCYLRLSDNPRAREALRQCLPDQLRDGTFSNILKDDPSTKRWLQQLLSNLN 261 (262)
T ss_dssp HHHHHHTTSTTHHHHHHHHS-GGGTSSTTTTGGCS-HHHHHHHHHHHHHTT
T ss_pred HHHHHHccCHHHHHHHHHhCcHHHhcHHHHHHHhcCHHHHHHHHHHHHHhc
Confidence 344567789999999999999999998766666689999999999998864
No 30
>COG3763 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=33.27 E-value=84 Score=26.34 Aligned_cols=31 Identities=19% Similarity=0.209 Sum_probs=20.9
Q ss_pred hhhHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHhhcC
Q 015059 102 LFWVGVGVGLSALFSFVASRLKQYAMQQALKGMMNQMN 139 (414)
Q Consensus 102 l~WiGvGVgLsalfs~V~~~VK~yaMQqamKsMM~Qmg 139 (414)
|+|| ++++|+-+++....+. .+.|+|++|..
T Consensus 4 ~lai-l~ivl~ll~G~~~G~f------iark~~~k~lk 34 (71)
T COG3763 4 WLAI-LLIVLALLAGLIGGFF------IARKQMKKQLK 34 (71)
T ss_pred HHHH-HHHHHHHHHHHHHHHH------HHHHHHHHHHh
Confidence 6777 7777777777766632 45666666665
No 31
>PRK10404 hypothetical protein; Provisional
Probab=31.28 E-value=44 Score=28.73 Aligned_cols=21 Identities=19% Similarity=0.471 Sum_probs=17.8
Q ss_pred CchhhHHHHHHHHHHHHHHHH
Q 015059 100 SPLFWVGVGVGLSALFSFVAS 120 (414)
Q Consensus 100 sPl~WiGvGVgLsalfs~V~~ 120 (414)
.||-=|||+.|+|.|+++.+.
T Consensus 79 ~Pw~avGiaagvGlllG~Ll~ 99 (101)
T PRK10404 79 KPWQGIGVGAAVGLVLGLLLA 99 (101)
T ss_pred CcHHHHHHHHHHHHHHHHHHh
Confidence 688889999999999988754
No 32
>PF11075 DUF2780: Protein of unknown function VcgC/VcgE (DUF2780); InterPro: IPR021302 This is a bacterial family of uncharacterised proteins.
Probab=30.95 E-value=42 Score=31.13 Aligned_cols=33 Identities=33% Similarity=0.558 Sum_probs=22.7
Q ss_pred cHHHHHHHHHcCCCcHHHHHHhhCCHHHHhhcCCHH
Q 015059 350 SAEVKQQFEQIGLTPEEVITKMMANPEIALGFQSPR 385 (414)
Q Consensus 350 spev~~qF~q~GmtP~e~~skIm~DPELlaAfQDPE 385 (414)
-.+|+++|+++||+++ ++.++. |.|+...+..-
T Consensus 121 ~~~v~~~F~~LGld~~-mi~~f~--pii~~yL~~qG 153 (163)
T PF11075_consen 121 MADVNSAFSALGLDPS-MISQFV--PIILSYLQSQG 153 (163)
T ss_pred HHHHHHHHHHcCCCHH-HHHHHH--HHHHHHHHhhh
Confidence 5599999999999988 444443 45555555444
No 33
>PF15050 SCIMP: SCIMP protein
Probab=30.04 E-value=3.4e+02 Score=25.14 Aligned_cols=37 Identities=30% Similarity=0.588 Sum_probs=22.9
Q ss_pred HHHHHHHHhhcCCCCCCCCCCCCC-CCCCCCCCCCCCCCCCCC
Q 015059 128 QQALKGMMNQMNTQNKPFGNAAFP-QGSPFPFPNPPASGPTTP 169 (414)
Q Consensus 128 QqamKsMM~Qmg~~~~~fG~~pf~-~gsPFpfp~Pp~~gp~~~ 169 (414)
|.|.+-..+|...+= |+.+ -|.+|++-..|+-.|..|
T Consensus 55 EkmYENv~n~~~~~L-----PpLPPRg~~s~~~~spqetPs~p 92 (133)
T PF15050_consen 55 EKMYENVLNQSPVQL-----PPLPPRGSPSPEDSSPQETPSQP 92 (133)
T ss_pred HHHHHHhhcCCcCCC-----CCCCCCCCCCccccCcccCCCCC
Confidence 667777777777665 5665 347777766555554433
No 34
>PF06757 Ins_allergen_rp: Insect allergen related repeat, nitrile-specifier detoxification; InterPro: IPR010629 This entry represents several insect specific allergen repeats. These repeats are commonly found in various proteins from cockroaches, fruit flies and mosquitos. It has been suggested that the repeat sequences have evolved by duplication of an ancestral amino acid domain, which may have arisen from the mitochondrial energy transfer proteins []. This family exemplifies a case of novel gene evolution. The case in point is the arms-race between plants and their infective insective herbivores in the area of the glucosinolate-myrosinase system. Brassicas have developed the glucosinolate-myrosinase system as chemical defence mechanism against the insects, and consequently the insects have adapted to produce a detoxifying molecule, nitrile-specifier protein (NSP). NSP is present in the Pieris rapae (Cabbage white butterfly). NSP is structurally different from and has no amino acid homology to any known detoxifying enzymes, and it appears to have arisen by a process of domain and gene duplication of a sequence of unknown function that is widespread in insect species and referred to as insect-allergen-repeat protein. Thus this family is found either as a single domain or as a multiple repeat-domain [].
Probab=29.60 E-value=70 Score=29.01 Aligned_cols=60 Identities=23% Similarity=0.276 Sum_probs=39.9
Q ss_pred hhcChhhhhhhccCCCcccCChHHHH----HHhhChHHHHHHHHHHHhccCCC-CCchhhhhhcccCCCCcH
Q 015059 285 LMEDPQVQKMVYPSLPEEMRNPASFK----LMLQNPEYRKQLQEMLDGMCESG-EFDGRVLDSLKNFDLNSA 351 (414)
Q Consensus 285 M~~dP~~Qkm~ypyLPe~MRNp~tfk----~ml~NP~yr~QLe~ml~~mg~~~-~~~~~m~D~lk~~d~~sp 351 (414)
+..|+++|+.+ ..||.++ |+ .|.+.|+|| .|-+-|.+.|... .|=++.-+.|+-..++.+
T Consensus 28 ~~~D~efq~~~-----~yl~s~~-f~~l~~~l~~~pE~~-~l~~yL~~~gldv~~~i~~i~~~l~~~~~~p~ 92 (179)
T PF06757_consen 28 YLEDAEFQAAV-----RYLNSSE-FKQLWQQLEALPEVK-ALLDYLESAGLDVYYYINQINDLLGLPPLNPT 92 (179)
T ss_pred HHcCHHHHHHH-----HHHcChH-HHHHHHHHHcCHHHH-HHHHHHHHCCCCHHHHHHHHHHHHcCCcCCCC
Confidence 57899999976 3467765 55 467889998 5566667778876 445555565554444443
No 35
>PHA00736 hypothetical protein
Probab=27.87 E-value=44 Score=28.05 Aligned_cols=17 Identities=41% Similarity=0.937 Sum_probs=10.2
Q ss_pred chhhHHHHHHHHHHHHHH
Q 015059 101 PLFWVGVGVGLSALFSFV 118 (414)
Q Consensus 101 Pl~WiGvGVgLsalfs~V 118 (414)
|||| |++|.++-+-+.|
T Consensus 56 plfw-gi~vifgliag~v 72 (79)
T PHA00736 56 PLFW-GITVIFGLIAGLV 72 (79)
T ss_pred HHHH-HHHHHHHHHHHHh
Confidence 6788 5666555555444
No 36
>PF03923 Lipoprotein_16: Uncharacterized lipoprotein; InterPro: IPR005619 The function of this presumed lipoprotein is unknown. The family includes Escherichia coli YajG P36671 from SWISSPROT.
Probab=27.16 E-value=44 Score=29.74 Aligned_cols=18 Identities=22% Similarity=0.462 Sum_probs=14.5
Q ss_pred HHHHHHhhCCHHHHhhcC
Q 015059 365 EEVITKMMANPEIALGFQ 382 (414)
Q Consensus 365 ~e~~skIm~DPELlaAfQ 382 (414)
.++++.|++||||..++|
T Consensus 142 ~~~l~~i~~D~el~~~l~ 159 (159)
T PF03923_consen 142 SDVLNDIANDPELIQFLQ 159 (159)
T ss_pred HHHHHHHHcCHHHHHHhC
Confidence 457899999999888765
No 37
>PF03960 ArsC: ArsC family; InterPro: IPR006660 Several bacterial taxon have a chromosomal resistance system, encoded by the ars operon, for the detoxification of arsenate, arsenite, and antimonite []. This system transports arsenite and antimonite out of the cell. The pump is composed of two polypeptides, the products of the arsA and arsB genes. This two-subunit enzyme produces resistance to arsenite and antimonite. Arsenate, however, must first be reduced to arsenite before it is extruded. A third gene, arsC, expands the substrate specificity to allow for arsenate pumping and resistance. ArsC is an approximately 150-residue arsenate reductase that uses reduced glutathione (GSH) to convert arsenate to arsenite with a redox active cysteine residue in the active site. ArsC forms an active quaternary complex with GSH, arsenate, and glutaredoxin 1 (Grx1). The three ligands must be present simultaneously for reduction to occur []. The arsC family also comprises the Spx proteins which are GRAM-positive bacterial transcription factors that regulate the transcription of multiple genes in response to disulphide stress []. The arsC protein structure has been solved []. It belongs to the thioredoxin superfamily fold which is defined by a beta-sheet core surrounded by alpha-helices. The active cysteine residue of ArsC is located in the loop between the first beta-strand and the first helix, which is also conserved in the Spx protein and its homologues.; PDB: 2KOK_A 1SK1_A 1SK2_A 1JZW_A 1J9B_A 1S3C_A 1SD8_A 1SD9_A 1I9D_A 1SK0_A ....
Probab=27.11 E-value=56 Score=27.01 Aligned_cols=57 Identities=28% Similarity=0.411 Sum_probs=31.8
Q ss_pred hhChHHHHHHHHHHHhccCCCCCchhhhhhcccCCCCcHHHHHHH--HHcCCCcHHHHHHhhCCHHHH
Q 015059 313 LQNPEYRKQLQEMLDGMCESGEFDGRVLDSLKNFDLNSAEVKQQF--EQIGLTPEEVITKMMANPEIA 378 (414)
Q Consensus 313 l~NP~yr~QLe~ml~~mg~~~~~~~~m~D~lk~~d~~spev~~qF--~q~GmtP~e~~skIm~DPELl 378 (414)
+.+|-=+.+|.++|+..|.+. +.|-| .++...++.- +...++-++.+.-|+++|.|+
T Consensus 29 ~k~p~s~~el~~~l~~~~~~~-------~~lin--~~~~~~k~l~~~~~~~~s~~e~i~~l~~~p~Li 87 (110)
T PF03960_consen 29 KKEPLSREELRELLSKLGNGP-------DDLIN--TRSKTYKELGKLKKDDLSDEELIELLLENPKLI 87 (110)
T ss_dssp TTS---HHHHHHHHHHHTSSG-------GGGB---TTSHHHHHTTHHHCTTSBHHHHHHHHHHSGGGB
T ss_pred hhCCCCHHHHHHHHHHhcccH-------HHHhc--CccchHhhhhhhhhhhhhhHHHHHHHHhChhhe
Confidence 456667788888888877532 22333 4455444322 334577777777777777654
No 38
>KOG0011 consensus Nucleotide excision repair factor NEF2, RAD23 component [Replication, recombination and repair]
Probab=27.01 E-value=3.9e+02 Score=28.09 Aligned_cols=41 Identities=37% Similarity=0.506 Sum_probs=31.3
Q ss_pred HHHHHhhcChhhhhhhccCCCcccCChHHHHHHhh-----ChHHHHHHHH
Q 015059 280 DTLEKLMEDPQVQKMVYPSLPEEMRNPASFKLMLQ-----NPEYRKQLQE 324 (414)
Q Consensus 280 ~~~e~M~~dP~~Qkm~ypyLPe~MRNp~tfk~ml~-----NP~yr~QLe~ 324 (414)
+.|+-+..+|++|+|.-=. =.||+.++-||| ||+.+++|++
T Consensus 214 ~~l~fLr~~~qf~~lR~~i----qqNP~ll~~~Lqqlg~~nP~L~q~Iq~ 259 (340)
T KOG0011|consen 214 DPLEFLRNQPQFQQLRQMI----QQNPELLHPLLQQLGKQNPQLLQLIQE 259 (340)
T ss_pred CchhhhhccHHHHHHHHHH----hhCHHHHHHHHHHHhhhCHHHHHHHHH
Confidence 6788888999998763000 149999999996 8999999853
No 39
>PF10158 LOH1CR12: Tumour suppressor protein; InterPro: IPR018780 This entry represents a region of 130 amino acids that is the most conserved part of some hypothetical proteins involved in loss of heterozygosity, and thus, tumour suppression []. The exact function of these proteins is not known.
Probab=25.81 E-value=20 Score=32.16 Aligned_cols=63 Identities=27% Similarity=0.414 Sum_probs=43.7
Q ss_pred hccCCCc-----ccCChHHHHHHhhChHHH--HHHHHHHHhccCCCCCchh-hhhhcccCCCCcHHHHHHH
Q 015059 295 VYPSLPE-----EMRNPASFKLMLQNPEYR--KQLQEMLDGMCESGEFDGR-VLDSLKNFDLNSAEVKQQF 357 (414)
Q Consensus 295 ~ypyLPe-----~MRNp~tfk~ml~NP~yr--~QLe~ml~~mg~~~~~~~~-m~D~lk~~d~~spev~~qF 357 (414)
.||.|+. .+|+|++++.|=..|-.| ..+++-|+.-....+.++. ....+|++|.....+-+++
T Consensus 2 FlPilr~~l~~~~~rd~~~leklds~~~l~Lc~R~Q~HL~~cA~~Va~~Q~~L~~riKevd~~~~~l~~~~ 72 (131)
T PF10158_consen 2 FLPILRGSLNLPDSRDPEVLEKLDSRPVLRLCSRYQEHLNQCAEAVAFDQNALAKRIKEVDQEIAKLLQQM 72 (131)
T ss_pred CcccchhhcCCCCCCChHHHHccChHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 5788887 889999999998888888 8888888775444444442 3344566666555554444
No 40
>PRK14581 hmsF outer membrane N-deacetylase; Provisional
Probab=25.70 E-value=1.7e+02 Score=32.77 Aligned_cols=92 Identities=20% Similarity=0.394 Sum_probs=62.9
Q ss_pred ChHHHHHHHHHHHhccCCCCCchhh-hh--hcccCCCCcHHHHHHHHHcCCCcHHHHHHhhCCHHHHhhcCCHHHHHHHH
Q 015059 315 NPEYRKQLQEMLDGMCESGEFDGRV-LD--SLKNFDLNSAEVKQQFEQIGLTPEEVITKMMANPEIALGFQSPRVQAAIM 391 (414)
Q Consensus 315 NP~yr~QLe~ml~~mg~~~~~~~~m-~D--~lk~~d~~spev~~qF~q~GmtP~e~~skIm~DPELlaAfQDPEVmaA~q 391 (414)
+|+-|++|.++-+-.......|+=. -| .|.+|.=-||+-.++.++.|++. .+.+|-+||+++.....=+ -.+|.
T Consensus 439 ~~~~~~~i~~iy~DLa~~~~~~GilfhDd~~l~d~ed~sp~a~~~y~~~gl~~--~~~~~~~~~~~~~~w~~~k-~~~l~ 515 (672)
T PRK14581 439 NPEVRQRIIDIYRDMAYSAPIDGIIYHDDAVMSDFEDASPDAIRAYEKAGFPG--SITTIRQDPEMMQRWTRYK-SKYLI 515 (672)
T ss_pred CHHHHHHHHHHHHHHHhcCCCCeEEeccccccccccccCHHHHHHHHhcCCCc--cHHhHhcCHHHHHHHHHHH-HHHHH
Confidence 7899999999999887765565521 11 35666667887788889999874 4889999999886442222 24566
Q ss_pred HHhcChhhhhhhccChhh
Q 015059 392 ECSQNPMNIIKYQNDKEV 409 (414)
Q Consensus 392 Di~sNPaNisKYqnnPKV 409 (414)
|....-+.+.+.-..|+|
T Consensus 516 ~f~~~l~~~v~~~~~p~~ 533 (672)
T PRK14581 516 DFTNELTREVRDIRGPQV 533 (672)
T ss_pred HHHHHHHHHHHhhcCccc
Confidence 666666666655444443
No 41
>PF06295 DUF1043: Protein of unknown function (DUF1043); InterPro: IPR009386 This entry consists of several hypothetical bacterial proteins of unknown function.
Probab=25.47 E-value=59 Score=28.51 Aligned_cols=16 Identities=19% Similarity=0.162 Sum_probs=7.5
Q ss_pred hhHHHHHHHHHHHHHH
Q 015059 103 FWVGVGVGLSALFSFV 118 (414)
Q Consensus 103 ~WiGvGVgLsalfs~V 118 (414)
+|+-||+++|.+++..
T Consensus 4 i~lvvG~iiG~~~~r~ 19 (128)
T PF06295_consen 4 IGLVVGLIIGFLIGRL 19 (128)
T ss_pred HHHHHHHHHHHHHHHH
Confidence 4444455555444443
No 42
>cd06199 SiR Cytochrome p450- like alpha subunits of E. coli sulfite reductase (SiR) multimerize with beta subunits to catalyze the NADPH dependent reduction of sulfite to sulfide. Beta subunits have an Fe4S4 cluster and a siroheme, while the alpha subunits (cysJ gene) are of the cytochrome p450 (CyPor) family having FAD and FMN as prosthetic groups and utilizing NADPH. Cypor (including cyt -450 reductase, nitric oxide synthase, and methionine synthase reductase) are ferredoxin reductase (FNR)-like proteins with an additional N-terminal FMN domain and a connecting sub-domain inserted within the flavin binding portion of the FNR-like domain. The connecting domain orients the N-terminal FMN domain with the C-terminal FNR domain.
Probab=24.82 E-value=60 Score=32.29 Aligned_cols=25 Identities=28% Similarity=0.632 Sum_probs=21.9
Q ss_pred CCchhhHHHHHHHHHHHHHHHHHHH
Q 015059 99 GSPLFWVGVGVGLSALFSFVASRLK 123 (414)
Q Consensus 99 GsPl~WiGvGVgLsalfs~V~~~VK 123 (414)
+.|+++||-|.|++-+.|++..++.
T Consensus 213 ~~piImIa~GtGIAP~~s~l~~~~~ 237 (360)
T cd06199 213 DAPIIMVGPGTGIAPFRAFLQEREA 237 (360)
T ss_pred CCCEEEEecCcChHHHHHHHHHHHh
Confidence 5799999999999999999876653
No 43
>PF10474 DUF2451: Protein of unknown function C-terminus (DUF2451); InterPro: IPR019514 This protein is found in eukaryotes but its function is not known. The N-terminal domain of some members is PF10475 from PFAM (DUF2450).
Probab=24.77 E-value=44 Score=32.08 Aligned_cols=31 Identities=32% Similarity=0.447 Sum_probs=26.8
Q ss_pred CCCcccCChHHHHHHhhChHH-HHHHHHHHHhccCC
Q 015059 298 SLPEEMRNPASFKLMLQNPEY-RKQLQEMLDGMCES 332 (414)
Q Consensus 298 yLPe~MRNp~tfk~ml~NP~y-r~QLe~ml~~mg~~ 332 (414)
||||+ |.-+|+-+|++| ++||-.+++..+++
T Consensus 189 Yl~e~----e~~~W~~~h~eYs~~ql~~Lv~~~~~~ 220 (234)
T PF10474_consen 189 YLPEE----ELEEWIRTHTEYSKKQLVGLVNCAAAS 220 (234)
T ss_pred cCCHH----HHHHHHHhCcccCHHHHHHHHHHHHHh
Confidence 59985 899999999999 68999999986664
No 44
>COG2838 Icd Monomeric isocitrate dehydrogenase [Energy production and conversion]
Probab=24.18 E-value=51 Score=36.54 Aligned_cols=60 Identities=13% Similarity=0.218 Sum_probs=47.4
Q ss_pred HHHHHHHHcCCCcHHHHHHhhCCHHHHhhcCCHHHHHHHHHHhcChhhhhhhccChhhhc
Q 015059 352 EVKQQFEQIGLTPEEVITKMMANPEIALGFQSPRVQAAIMECSQNPMNIIKYQNDKEVFS 411 (414)
Q Consensus 352 ev~~qF~q~GmtP~e~~skIm~DPELlaAfQDPEVmaA~qDi~sNPaNisKYqnnPKVmn 411 (414)
+-.++|+++|+.+...++.+++.-|-+.+=|.-||.++|+.|...-..+.---+|.-|-|
T Consensus 283 k~~~~f~~lGvn~nNGl~~l~skiesl~~~~r~eI~~~~~~~~a~~p~laMVdS~kGItN 342 (744)
T COG2838 283 KHGDLFDALGVNVNNGLSDLYSKIESLPASQRAEIEADIHAVYAHRPDLAMVDSDKGITN 342 (744)
T ss_pred HHHHHHHHhCCCccccHHHHHHHHhcCChhhHHHHHHHHHHHHhcCCcceeeecccCccc
Confidence 446788999999999999999988888888899999999998776555554445554444
No 45
>PRK07983 exodeoxyribonuclease X; Provisional
Probab=24.14 E-value=1.1e+02 Score=28.88 Aligned_cols=49 Identities=20% Similarity=0.446 Sum_probs=30.0
Q ss_pred cHHHHHHhhcChhhhhhhccC--------CCcccCChHHHHHHhhC-----hHHHHHHHHHHH
Q 015059 278 TVDTLEKLMEDPQVQKMVYPS--------LPEEMRNPASFKLMLQN-----PEYRKQLQEMLD 327 (414)
Q Consensus 278 ~v~~~e~M~~dP~~Qkm~ypy--------LPe~MRNp~tfk~ml~N-----P~yr~QLe~ml~ 327 (414)
+.+.|-.+.+.|.+-+- +|+ =+-..|.++-++|||.| |.+|.-|+.-|+
T Consensus 157 ~~~~l~~~~~~~~~~~~-~~fGk~kg~~~~~~~~~~~~yl~wl~~~~~d~~~~l~~~~~~~l~ 218 (219)
T PRK07983 157 TAEEMADITGRPSLLTT-FTFGKYRGKAVSDVAERDPGYLRWLFNNLDDMSPELRLTLKHYLE 218 (219)
T ss_pred CHHHHHHHhcCCccCCC-ccccCccCcchhhhhhcchHHHHHHHhcccccCHHHHHHHHHHhh
Confidence 34555555555554432 222 11133678999999999 888887776654
No 46
>KOG2366 consensus Alpha-D-galactosidase (melibiase) [Carbohydrate transport and metabolism]
Probab=23.78 E-value=1.3e+02 Score=32.26 Aligned_cols=74 Identities=20% Similarity=0.226 Sum_probs=53.5
Q ss_pred HHHHHHhhChHHHHHHHHHHHhccCCCCCchhhhhhc--ccCCCCcHHHHHHHH--------------HcCCCcHHHHHH
Q 015059 307 ASFKLMLQNPEYRKQLQEMLDGMCESGEFDGRVLDSL--KNFDLNSAEVKQQFE--------------QIGLTPEEVITK 370 (414)
Q Consensus 307 ~tfk~ml~NP~yr~QLe~ml~~mg~~~~~~~~m~D~l--k~~d~~spev~~qF~--------------q~GmtP~e~~sk 370 (414)
||++-|.++--|---.++-++...+.+.|.. +||| -|+-+.-.+-+-||. +--|+++ +.+
T Consensus 228 dtW~Sv~~I~d~~~~nqd~~~~~agPg~WND--pDmL~iGN~G~s~e~y~~qf~lWai~kAPLlms~Dlr~is~~--~~~ 303 (414)
T KOG2366|consen 228 DTWKSVDSIIDYICWNQDRIAPLAGPGGWND--PDMLEIGNGGMSYEEYKGQFALWAILKAPLLMSNDLRLISKQ--TKE 303 (414)
T ss_pred hHHHHHHHHHHHHhhhhhhhccccCCCCCCC--hhHhhcCCCCccHHHHHHHHHHHHHhhchhhhccchhhcCHH--HHH
Confidence 6788888888777777888888888888866 4544 578888888888882 2234444 556
Q ss_pred hhCCHHHHhhcCCH
Q 015059 371 MMANPEIALGFQSP 384 (414)
Q Consensus 371 Im~DPELlaAfQDP 384 (414)
|++++|+.++=|||
T Consensus 304 il~nk~~IaiNQDp 317 (414)
T KOG2366|consen 304 ILQNKEVIAINQDP 317 (414)
T ss_pred HhcChhheeccCCc
Confidence 77777777777775
No 47
>PF07301 DUF1453: Protein of unknown function (DUF1453); InterPro: IPR009916 This family consists of several hypothetical bacterial proteins of around 150 residues in length. The function of this family is unknown. Members of this family seem to be found exclusively in the Order Bacillales.
Probab=23.60 E-value=82 Score=29.17 Aligned_cols=26 Identities=23% Similarity=0.220 Sum_probs=21.2
Q ss_pred CCCchhhHHHHHHHHHHHHHHHHHHH
Q 015059 98 VGSPLFWVGVGVGLSALFSFVASRLK 123 (414)
Q Consensus 98 iGsPl~WiGvGVgLsalfs~V~~~VK 123 (414)
..-||.|+..++.+|++||+..-+--
T Consensus 53 ~~~~~~~~l~A~~~G~lFs~~Li~ts 78 (148)
T PF07301_consen 53 FRPPWLEVLEAFLVGALFSYPLIKTS 78 (148)
T ss_pred ccchHHHHHHHHHHHHHHHHHHHHhc
Confidence 34578999999999999999876543
No 48
>KOG3341 consensus RNA polymerase II transcription factor complex subunit [Transcription]
Probab=23.42 E-value=65 Score=32.20 Aligned_cols=22 Identities=27% Similarity=0.451 Sum_probs=19.9
Q ss_pred HhhChHHHHHHHHHHHhccCCC
Q 015059 312 MLQNPEYRKQLQEMLDGMCESG 333 (414)
Q Consensus 312 ml~NP~yr~QLe~ml~~mg~~~ 333 (414)
+-+|||||.|.++|.+.-|.++
T Consensus 56 i~knsqFR~~Fq~Mca~IGvDP 77 (249)
T KOG3341|consen 56 IRKNSQFRNQFQEMCASIGVDP 77 (249)
T ss_pred HhhCHHHHHHHHHHHHHcCCCc
Confidence 4589999999999999998876
No 49
>smart00845 GatB_Yqey GatB domain. This domain is found in GatB and proteins related to bacterial Yqey. It is about 140 amino acid residues long. This domain is found at the C terminus of GatB which transamidates Glu-tRNA to Gln-tRNA. The function of this domain is uncertain. It does however suggest that Yqey and its relatives have a role in tRNA metabolism.
Probab=23.15 E-value=91 Score=27.40 Aligned_cols=68 Identities=19% Similarity=0.381 Sum_probs=45.2
Q ss_pred hhhhcccCCCCcHHHHHHHHHc---CCCcHHHHHHhhCCHHHHhhcCCH-HHHHHHHH-HhcChhhhhhhccC-hhhhc
Q 015059 339 VLDSLKNFDLNSAEVKQQFEQI---GLTPEEVITKMMANPEIALGFQSP-RVQAAIME-CSQNPMNIIKYQND-KEVFS 411 (414)
Q Consensus 339 m~D~lk~~d~~spev~~qF~q~---GmtP~e~~skIm~DPELlaAfQDP-EVmaA~qD-i~sNPaNisKYqnn-PKVmn 411 (414)
..+++.+=.++....++-|+.+ |-+|++++.+.. +..+.|. ++.+.+++ |.+||.-+.+|.+- .|++.
T Consensus 47 li~lv~~g~It~~~ak~vl~~~~~~~~~~~~ii~~~~-----l~~isd~~el~~~v~~vi~~~~~~v~~~~~g~~k~~~ 120 (147)
T smart00845 47 LLKLIEDGTISGKIAKEVLEELLESGKSPEEIVEEKG-----LKQISDEGELEAIVDEVIAENPKAVEDYRAGKKKALG 120 (147)
T ss_pred HHHHHHcCCCcHHHHHHHHHHHHHcCCCHHHHHHHcC-----CccCCCHHHHHHHHHHHHHHCHHHHHHHHCCHHHHHH
Confidence 4455556567777777777554 667777776652 3456776 67777777 66799999999754 34443
No 50
>COG2427 Uncharacterized conserved protein [Function unknown]
Probab=21.75 E-value=63 Score=29.07 Aligned_cols=82 Identities=17% Similarity=0.113 Sum_probs=41.5
Q ss_pred cCChH----HHHHHhhChHHHHHHHHHHHhccCCCCCch-hhhhhcccCCCCcHHHHHHHH-HcCCCcHHHHHHhhCCHH
Q 015059 303 MRNPA----SFKLMLQNPEYRKQLQEMLDGMCESGEFDG-RVLDSLKNFDLNSAEVKQQFE-QIGLTPEEVITKMMANPE 376 (414)
Q Consensus 303 MRNp~----tfk~ml~NP~yr~QLe~ml~~mg~~~~~~~-~m~D~lk~~d~~spev~~qF~-q~GmtP~e~~skIm~DPE 376 (414)
||... .++-+++++.+...|..++.-++.....+. |+.+++.+ +..-++ +-+.... +.++==+
T Consensus 50 ~~~~~~i~~~~~~~l~~e~~~~ll~~~~~~~~~l~~~~~e~~~~~~~~-------~~~a~~~~~~~~~~----~~vgl~~ 118 (148)
T COG2427 50 LRAKADIAKKLKDELAKELIENLLNNMLIMLGLLSLIDSERLSKLVEN-------LIKAIEAVKAEKNA----EPVGLLG 118 (148)
T ss_pred cccHHHHHHHHHHHHhHHHHHHHHHHHHHHHHHHHhccHHHHHHHHHH-------HHHHHHHHHhcccC----CCccHHH
Confidence 77764 445556666666665554444333222222 33333332 122222 1222221 2333447
Q ss_pred HHhhcCCHHHHHHHHHHhc
Q 015059 377 IALGFQSPRVQAAIMECSQ 395 (414)
Q Consensus 377 LlaAfQDPEVmaA~qDi~s 395 (414)
|+.+++||+|+.++-=+++
T Consensus 119 Llk~LkDPdvq~~Lg~lls 137 (148)
T COG2427 119 LLKALKDPDVQRGLGFLLS 137 (148)
T ss_pred HHHHcCCHHHHHHHHHHHH
Confidence 8999999999999876543
No 51
>PF01323 DSBA: DSBA-like thioredoxin domain; InterPro: IPR001853 DSBA is a sub-family of the Thioredoxin family []. The efficient and correct folding of bacterial disulphide bonded proteins in vivo is dependent upon a class of periplasmic oxidoreductase proteins called DsbA, after the Escherichia coli enzyme. The bacterial protein-folding factor DsbA is the most oxidizing of the thioredoxin family. DsbA catalyses disulphide-bond formation during the folding of secreted proteins. The extremely oxidizing nature of DsbA has been proposed to result from either domain motion or stabilising active-site interactions in the reduced form. DsbA's highly oxidizing nature is a result of hydrogen bond, electrostatic and helix-dipole interactions that favour the thiolate over the disulphide at the active site []. In the pathogenic bacterium Vibrio cholerae, the DsbA homologue (TcpG) is responsible for the folding, maturation and secretion of virulence factors. While the overall architecture of TcpG and DsbA is similar and the surface features are retained in TcpG, there are significant differences. For example, the kinked active site helix results from a three-residue loop in DsbA, but is caused by a proline in TcpG (making TcpG more similar to thioredoxin in this respect). Furthermore, the proposed peptide binding groove of TcpG is substantially shortened compared with that of DsbA due to a six-residue deletion. Also, the hydrophobic pocket of TcpG is more shallow and the acidic patch is much less extensive than that of E. coli DsbA [].; GO: 0015035 protein disulfide oxidoreductase activity; PDB: 3GL5_A 3DKS_D 3RPP_C 3RPN_B 1YZX_A 3L9V_C 2IMD_A 2IME_A 2IMF_A 2B3S_B ....
Probab=21.35 E-value=1.3e+02 Score=25.72 Aligned_cols=38 Identities=21% Similarity=0.504 Sum_probs=25.8
Q ss_pred cCCCCcHH-HHHHHHHcCCCcHHHHHHhhCCHHHHhhcCC
Q 015059 345 NFDLNSAE-VKQQFEQIGLTPEEVITKMMANPEIALGFQS 383 (414)
Q Consensus 345 ~~d~~spe-v~~qF~q~GmtP~e~~skIm~DPELlaAfQD 383 (414)
+-|++.++ +.+-+++.|+++++ +.+.++|+++.+..+.
T Consensus 117 ~~~i~~~~vl~~~~~~~Gld~~~-~~~~~~~~~~~~~~~~ 155 (193)
T PF01323_consen 117 GRDISDPDVLAEIAEEAGLDPDE-FDAALDSPEVKAALEE 155 (193)
T ss_dssp ST-TSSHHHHHHHHHHTT--HHH-HHHHHTSHHHHHHHHH
T ss_pred ccCCCCHHHHHHHHHHcCCcHHH-HHHHhcchHHHHHHHH
Confidence 45688885 88888999998875 5577777777665554
No 52
>cd02977 ArsC_family Arsenate Reductase (ArsC) family; composed of TRX-fold arsenic reductases and similar proteins including the transcriptional regulator, Spx. ArsC catalyzes the reduction of arsenate [As(V)] to arsenite [As(III)], using reducing equivalents derived from glutathione (GSH) via glutaredoxin (GRX), through a single catalytic cysteine. This family of predominantly bacterial enzymes is unrelated to two other families of arsenate reductases which show similarity to low-molecular-weight acid phosphatases and phosphotyrosyl phosphatases. Spx is a general regulator that exerts negative and positive control over transcription initiation by binding to the C-terminal domain of the alpha subunit of RNA polymerase.
Probab=21.27 E-value=47 Score=27.02 Aligned_cols=17 Identities=29% Similarity=0.556 Sum_probs=8.9
Q ss_pred CCCcHHHHHHhhCCHHH
Q 015059 361 GLTPEEVITKMMANPEI 377 (414)
Q Consensus 361 GmtP~e~~skIm~DPEL 377 (414)
+++-++.+..|.++|.|
T Consensus 73 ~ls~~e~~~~l~~~p~L 89 (105)
T cd02977 73 ELSDEEALELMAEHPKL 89 (105)
T ss_pred CCCHHHHHHHHHhCcCe
Confidence 34445555555555554
No 53
>TIGR00601 rad23 UV excision repair protein Rad23. All proteins in this family for which functions are known are components of a multiprotein complex used for targeting nucleotide excision repair to specific parts of the genome. In humans, Rad23 complexes with the XPC protein. This family is based on the phylogenomic analysis of JA Eisen (1999, Ph.D. Thesis, Stanford University).
Probab=21.17 E-value=1.6e+02 Score=30.62 Aligned_cols=32 Identities=34% Similarity=0.549 Sum_probs=18.9
Q ss_pred HHHHhhcChhhhhhhccCCCcccCChHHHHHHhhChHHHHHHHHHH
Q 015059 281 TLEKLMEDPQVQKMVYPSLPEEMRNPASFKLMLQNPEYRKQLQEML 326 (414)
Q Consensus 281 ~~e~M~~dP~~Qkm~ypyLPe~MRNp~tfk~ml~NP~yr~QLe~ml 326 (414)
.|+-+..+|++|+| | ..+-+||+.-.+|-+.|
T Consensus 247 ~l~~Lr~~pqf~~l---------R-----~~vq~NP~~L~~lLqql 278 (378)
T TIGR00601 247 PLEFLRNQPQFQQL---------R-----QVVQQNPQLLPPLLQQI 278 (378)
T ss_pred hHHHhhcCHHHHHH---------H-----HHHHHCHHHHHHHHHHH
Confidence 55666667777764 2 44567887654443333
No 54
>PF07237 DUF1428: Protein of unknown function (DUF1428); InterPro: IPR009874 This family consists of several hypothetical bacterial and one archaeal sequence of around 120 residues in length. The function of this family is unknown.; PDB: 2OKQ_A.
Probab=20.90 E-value=72 Score=28.00 Aligned_cols=20 Identities=35% Similarity=0.501 Sum_probs=15.7
Q ss_pred HHHHHHhhcChhhhhhh--ccC
Q 015059 279 VDTLEKLMEDPQVQKMV--YPS 298 (414)
Q Consensus 279 v~~~e~M~~dP~~Qkm~--ypy 298 (414)
-..+++||+||.||.|. +|+
T Consensus 79 D~~~~k~m~DPrm~~~~~~mPF 100 (103)
T PF07237_consen 79 DAANAKMMADPRMQEMDNEMPF 100 (103)
T ss_dssp HHHHHHHHCSHHHHHTTSS-SS
T ss_pred HHHHHHhhcCcCcCCCCCCCCC
Confidence 45689999999999975 554
No 55
>KOG2629 consensus Peroxisomal membrane anchor protein (peroxin) [Cell wall/membrane/envelope biogenesis; Posttranslational modification, protein turnover, chaperones; Intracellular trafficking, secretion, and vesicular transport]
Probab=20.72 E-value=2.1e+02 Score=29.59 Aligned_cols=56 Identities=20% Similarity=0.109 Sum_probs=38.5
Q ss_pred eecCCCccccccccCCCCCCCCCCCCCCCchhhHHHHHHHHHHHHHHH-HHHHHHHH
Q 015059 72 LTSSGGQQTSSVGVNPNLPMPPPSSNVGSPLFWVGVGVGLSALFSFVA-SRLKQYAM 127 (414)
Q Consensus 72 ~sss~~~~~~s~g~~p~~~~pp~~s~iGsPl~WiGvGVgLsalfs~V~-~~VK~yaM 127 (414)
.|+-.++.+.-+-..|++-+-.|....++-|=|.||--.+.+.|+|.+ .+||+|..
T Consensus 53 ~s~~~p~~~~~~~~~p~~~~~~P~~~~~~rwrdy~vmAvi~aGi~y~~y~~~K~YV~ 109 (300)
T KOG2629|consen 53 VSKQIPTANQVVSGGPPLLIIQPQQNVLRRWRDYFVMAVILAGIAYAAYRFVKSYVL 109 (300)
T ss_pred ccccCCCcccccCCCchhhhcCCCccchhhHHHHHHHHHHHhhHHHHHHHHHHHHHH
Confidence 455556666666666666666677888889999988666666666644 56777765
No 56
>PF02285 COX8: Cytochrome oxidase c subunit VIII; InterPro: IPR003205 Cytochrome c oxidase (1.9.3.1 from EC) is an oligomeric enzymatic complex which is a component of the respiratory chain complex and is involved in the transfer of electrons from cytochrome c to oxygen []. In eukaryotes this enzyme complex is located in the mitochondrial inner membrane; in aerobic prokaryotes it is found in the plasma membrane. In eukaryotes, in addition to the three large subunits, I, II and III, that form the catalytic centre of the enzyme complex, there are a variable number of small polypeptidic subunits.This family is composed of cytochrome c oxidase subunit VIII. ; GO: 0004129 cytochrome-c oxidase activity; PDB: 3AG3_Z 3ABM_M 1OCC_Z 3ASO_Z 3AG2_Z 3ABL_M 3AG4_M 3AG1_M 3ASN_M 1OCZ_M ....
Probab=20.55 E-value=1.2e+02 Score=23.16 Aligned_cols=29 Identities=21% Similarity=0.382 Sum_probs=19.5
Q ss_pred CCCchhhHHHHHHHHHHH---HHHHHHHHHHH
Q 015059 98 VGSPLFWVGVGVGLSALF---SFVASRLKQYA 126 (414)
Q Consensus 98 iGsPl~WiGvGVgLsalf---s~V~~~VK~ya 126 (414)
+|.-=--||+.|...++| +||++.+++|=
T Consensus 10 ~s~~e~aigltv~f~~~L~PagWVLshL~~YK 41 (44)
T PF02285_consen 10 LSPAEQAIGLTVCFVTFLGPAGWVLSHLESYK 41 (44)
T ss_dssp --HHHHHHHHHHHHHHHHHHHHHHHHTHHHHH
T ss_pred CCHHHHHHHHHHHHHHHHhhHHHHHHHHHHhh
Confidence 333334567777776665 79999999984
No 57
>KOG2051 consensus Nonsense-mediated mRNA decay 2 protein [RNA processing and modification]
Probab=20.25 E-value=89 Score=36.92 Aligned_cols=101 Identities=22% Similarity=0.443 Sum_probs=60.2
Q ss_pred HHHhhChHHHHHHHHHHHhcc---CCCCCchhhhhhccc--CCCCcHHHHHHHHHcCCCcHH-HHHHhhC--------CH
Q 015059 310 KLMLQNPEYRKQLQEMLDGMC---ESGEFDGRVLDSLKN--FDLNSAEVKQQFEQIGLTPEE-VITKMMA--------NP 375 (414)
Q Consensus 310 k~ml~NP~yr~QLe~ml~~mg---~~~~~~~~m~D~lk~--~d~~spev~~qF~q~GmtP~e-~~skIm~--------DP 375 (414)
..||.+|+++-+++.||.++- -.-..|.|+.-.+.| +=+++|+....... -.+|++ .+-+|+- |+
T Consensus 571 rfLlr~pEt~lrM~~~Le~i~rkK~a~~lDsr~~~~iENay~~~~PPe~~~~~~k-~r~p~~efiR~Li~~dL~k~tvd~ 649 (1128)
T KOG2051|consen 571 RFLLRSPETKLRMRVFLEQIKRKKRASALDSRQATLIENAYYLCNPPERSKRLSK-KRPPMQEFIRYLIRSDLSKDTVDR 649 (1128)
T ss_pred hhhhcChhHHHHHHHHHHHHHHHHHHhhhchHHHHHHHHhHHhccChhhcccccc-cCCcHHHHHHHHHHHHhccccHHH
Confidence 357899999999999988874 233456655444443 23567755441111 112222 2222222 12
Q ss_pred HH----HhhcCCHHHHHHHHHHhcChhhhhhhccChhhhcc
Q 015059 376 EI----ALGFQSPRVQAAIMECSQNPMNIIKYQNDKEVFSD 412 (414)
Q Consensus 376 EL----laAfQDPEVmaA~qDi~sNPaNisKYqnnPKVmnl 412 (414)
=| ..-.+||||.+-+-.|+.+|-+| ||++=+-|.++
T Consensus 650 ~lkllRkl~W~D~e~~~yli~~~~k~w~i-ky~~i~~lA~l 689 (1128)
T KOG2051|consen 650 VLKLLRKLDWSDPEVKQYLISCFSKPWKI-KYQNIHALASL 689 (1128)
T ss_pred HHHHHHhcccccHHHHHHHHHHhhhhhcc-ccccHHHHHHH
Confidence 11 12368999999999999999987 78775544443
No 58
>TIGR02384 RelB_DinJ addiction module antitoxin, RelB/DinJ family. Plasmids may be maintained stably in bacterial populations through the action of addiction modules, in which a toxin and antidote are encoded in a cassette on the plasmid. In any daughter cell that lacks the plasmid, the toxin persists and is lethal after the antidote protein is depleted. Toxin/antitoxin pairs are also found on main chromosomes, and likely represent selfish DNA. Sequences in the seed for this alignment all were found adjacent to toxin genes. The resulting model appears to describe a narrower set of proteins than Pfam model pfam04221, although many in the scope of this model are not obviously paired with toxin proteins. Several toxin/antitoxin pairs may occur in a single species.
Probab=20.12 E-value=1.7e+02 Score=24.18 Aligned_cols=28 Identities=18% Similarity=0.154 Sum_probs=16.1
Q ss_pred CHHHHHHHHHHhcChhhhhhhccChhhh
Q 015059 383 SPRVQAAIMECSQNPMNIIKYQNDKEVF 410 (414)
Q Consensus 383 DPEVmaA~qDi~sNPaNisKYqnnPKVm 410 (414)
+.|-.+|++|+-.-.-...+|.+-.+.|
T Consensus 53 n~et~~a~~e~~~~~~~~~~f~s~~el~ 80 (83)
T TIGR02384 53 NDETLAAIEEIKELRKLSHKFESVDDLL 80 (83)
T ss_pred CHHHHHHHHHHHHhcccCCCcCCHHHHH
Confidence 6777888888764222345565544433
Done!