Query 045282
Match_columns 130
No_of_seqs 106 out of 133
Neff 5.1
Searched_HMMs 46136
Date Fri Mar 29 11:41:21 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/045282.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/045282hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF09478 CBM49: Carbohydrate b 97.1 0.0044 9.6E-08 41.9 7.8 73 35-111 1-78 (80)
2 PLN02171 endoglucanase 92.7 0.43 9.2E-06 43.9 7.2 76 33-113 535-615 (629)
3 PF02933 CDC48_2: Cell divisio 67.4 5 0.00011 25.6 2.1 29 96-125 15-43 (64)
4 PF07172 GRP: Glycine rich pro 64.4 6.8 0.00015 27.7 2.5 13 1-13 1-13 (95)
5 PF10633 NPCBM_assoc: NPCBM-as 56.9 25 0.00054 22.9 4.1 50 48-111 5-57 (78)
6 PLN02340 endoglucanase 54.5 12 0.00027 34.5 3.1 78 31-111 518-600 (614)
7 PF07705 CARDB: CARDB; InterP 50.3 67 0.0015 20.7 6.1 68 33-114 2-71 (101)
8 PF07127 Nodulin_late: Late no 43.3 34 0.00074 21.3 3.0 15 1-16 1-16 (54)
9 COG3900 Predicted periplasmic 43.0 15 0.00032 30.6 1.5 19 42-60 206-224 (262)
10 PF03330 DPBB_1: Rare lipoprot 36.4 36 0.00078 22.2 2.4 37 50-86 38-75 (78)
11 PF06483 ChiC: Chitinase C; I 31.1 99 0.0021 24.6 4.4 49 64-118 35-86 (180)
12 PF06682 DUF1183: Protein of u 30.1 88 0.0019 26.8 4.3 25 62-88 91-115 (318)
13 PF03293 Pox_RNA_pol: Poxvirus 28.8 1E+02 0.0023 23.8 4.0 48 63-110 93-141 (160)
14 PF14016 DUF4232: Protein of u 27.5 1E+02 0.0022 22.0 3.7 73 30-114 1-82 (131)
15 PF03032 Brevenin: Brevenin/es 27.4 47 0.001 20.7 1.6 16 5-20 3-18 (46)
16 PLN00115 pollen allergen group 27.0 2.6E+02 0.0056 20.5 9.1 90 1-113 1-91 (118)
17 PF01345 DUF11: Domain of unkn 25.5 1.6E+02 0.0035 18.7 4.1 29 43-71 36-64 (76)
18 PF10731 Anophelin: Thrombin i 25.0 73 0.0016 21.3 2.3 21 1-21 1-21 (65)
19 KOG4063 Major epididymal secre 23.9 1.5E+02 0.0034 23.0 4.2 41 1-42 1-42 (158)
20 PRK15249 fimbrial chaperone pr 22.1 1.3E+02 0.0028 24.4 3.8 90 4-104 8-112 (253)
21 PF01456 Mucin: Mucin-like gly 21.2 83 0.0018 22.7 2.2 14 5-18 3-16 (143)
22 PF05753 TRAP_beta: Translocon 20.6 4.1E+02 0.0089 20.6 9.1 62 43-111 32-94 (181)
23 TIGR01451 B_ant_repeat conserv 20.0 2.1E+02 0.0047 17.5 3.7 25 46-70 10-34 (53)
No 1
>PF09478 CBM49: Carbohydrate binding domain CBM49; InterPro: IPR019028 A carbohydrate-binding module (CBM) is defined as a contiguous amino acid sequence within a carbohydrate-active enzyme with a discreet fold having carbohydrate-binding activity. A few exceptions are CBMs in cellulosomal scaffolding proteins and rare instances of independent putative CBMs. The requirement of CBMs existing as modules within larger enzymes sets this class of carbohydrate-binding protein apart from other non-catalytic sugar binding proteins such as lectins and sugar transport proteins. CBMs were previously classified as cellulose-binding domains (CBDs) based on the initial discovery of several modules that bound cellulose [, ]. However, additional modules in carbohydrate-active enzymes are continually being found that bind carbohydrates other than cellulose yet otherwise meet the CBM criteria, hence the need to reclassify these polypeptides using more inclusive terminology. Previous classification of cellulose-binding domains were based on amino acid similarity. Groupings of CBDs were called "Types" and numbered with roman numerals (e.g. Type I or Type II CBDs). In keeping with the glycoside hydrolase classification, these groupings are now called families and numbered with Arabic numerals. Families 1 to 13 are the same as Types I to XIII. For a detailed review on the structure and binding modes of CBMs see []. This domain is found at the C-terminal of cellulases and in vitro binding studies have shown it to binds to crystalline cellulose []. ; GO: 0030246 carbohydrate binding, 0005576 extracellular region
Probab=97.10 E-value=0.0044 Score=41.92 Aligned_cols=73 Identities=16% Similarity=0.172 Sum_probs=56.1
Q ss_pred ceEEEeeecCccCC---eeeEEEEEeeccccccceEEEecCCccc-cccCCCcceeEeCCeeEEeCC-cccCCCCeEEEE
Q 045282 35 LDISQKETGNIVQG---KKEFAVEVFNWCKCAQRNVTLDCDGFQT-VEKPDPVQMSISGFQCILLQG-RDIIPFSRVHFK 109 (130)
Q Consensus 35 i~V~Q~~tg~~v~G---~pe~~VtI~N~C~C~~~~V~l~C~gF~S-~~~VDP~ifr~~~~~CLVn~G-~pi~~g~~v~F~ 109 (130)
|+|.|..+..+..| ..+|.|+|+|.+.=+++++++.-+.+.+ .= .+-+..++..-+=+- .+|.+|++.+|-
T Consensus 1 i~i~q~~~~sW~~~g~~y~qy~v~I~N~~~~~I~~~~i~~~~l~~~iW----~l~~~~~~~y~lPs~~~~i~pg~s~~FG 76 (80)
T PF09478_consen 1 ITITQTLVNSWTENGQTYTQYDVTITNNGSKPIKSLKISIDNLYGSIW----GLDKVSGNTYTLPSYQPTIKPGQSFTFG 76 (80)
T ss_pred CEEEEEEEeEEEeCCEEEEEEEEEEEECCCCeEEEEEEEECccchhhe----eEEeccCCEEECCccccccCCCCEEEEE
Confidence 68899999888775 3579999999999999999999987651 11 222334577777454 399999999999
Q ss_pred Ec
Q 045282 110 YA 111 (130)
Q Consensus 110 YA 111 (130)
|-
T Consensus 77 YI 78 (80)
T PF09478_consen 77 YI 78 (80)
T ss_pred EE
Confidence 84
No 2
>PLN02171 endoglucanase
Probab=92.74 E-value=0.43 Score=43.95 Aligned_cols=76 Identities=14% Similarity=0.157 Sum_probs=55.0
Q ss_pred CcceEEEeeecCccC---CeeeEEEEEeeccccccceEEEecCCccc-cccCCCcceeEeCCeeEEeCCc-ccCCCCeEE
Q 045282 33 ETLDISQKETGNIVQ---GKKEFAVEVFNWCKCAQRNVTLDCDGFQT-VEKPDPVQMSISGFQCILLQGR-DIIPFSRVH 107 (130)
Q Consensus 33 ~di~V~Q~~tg~~v~---G~pe~~VtI~N~C~C~~~~V~l~C~gF~S-~~~VDP~ifr~~~~~CLVn~G~-pi~~g~~v~ 107 (130)
+.|+|.|..++.+.. +..+|+|+|+|++..+++++++.=..+-. .= .+. ..++...+=+-. -|.+|+..+
T Consensus 535 ~ei~i~q~v~~sW~~~g~~y~qy~v~I~N~s~~~ik~i~i~~~~~~~~iW----~v~-~~~ngytlPs~~~sL~aG~s~t 609 (629)
T PLN02171 535 SPIEIEQKATASWKAKGRTYYRYSTTVTNRSAKTLKELHLGISKLYGPLW----GLT-KAGYGYVLPSWMPSLPAGKSLE 609 (629)
T ss_pred ceeEEEEEEEEEEEcCCceEEEEEEEEEECCCCceeeeeeeeccccccch----hee-ecCCcccCchhhcccCCCCeeE
Confidence 368999999988885 47889999999999999999996544421 11 111 234445554443 788899999
Q ss_pred EEEccC
Q 045282 108 FKYAFD 113 (130)
Q Consensus 108 F~YAw~ 113 (130)
|-|=..
T Consensus 610 FgyI~~ 615 (629)
T PLN02171 610 FVYVHS 615 (629)
T ss_pred EEeecC
Confidence 999854
No 3
>PF02933 CDC48_2: Cell division protein 48 (CDC48), domain 2; InterPro: IPR004201 This domain has a double psi-beta barrel fold and includes VCP-like ATPase and N-ethylmaleimide sensitive fusion protein N-terminal domains. Both the VAT and NSF N-terminal functional domains consist of two structural domains of which this is at the C terminus. The VAT-N domain found in AAA ATPases (IPR003959 from INTERPRO) is a substrate 185-residue recognition domain [].; GO: 0005524 ATP binding; PDB: 1QDN_B 1QCS_A 1CR5_C 3QQ8_A 3HU2_A 3HU1_E 3HU3_A 3QWZ_A 3TIW_B 3QQ7_A ....
Probab=67.35 E-value=5 Score=25.60 Aligned_cols=29 Identities=28% Similarity=0.574 Sum_probs=23.7
Q ss_pred CCcccCCCCeEEEEEccCCccCeEeeeeec
Q 045282 96 QGRDIIPFSRVHFKYAFDDEFPFYVFSSAP 125 (130)
Q Consensus 96 ~G~pi~~g~~v~F~YAw~~~f~~~p~ss~~ 125 (130)
.|+|+..|+.|.|.|. ...++|.+.+.++
T Consensus 15 ~~~pv~~Gd~i~~~~~-~~~~~~~V~~~~P 43 (64)
T PF02933_consen 15 EGRPVTKGDTIVFPFF-GQALPFKVVSTEP 43 (64)
T ss_dssp TTEEEETT-EEEEEET-TEEEEEEEEEECS
T ss_pred cCCCccCCCEEEEEeC-CcEEEEEEEEEEc
Confidence 4689999999999996 6889999987654
No 4
>PF07172 GRP: Glycine rich protein family; InterPro: IPR010800 This family consists of glycine rich proteins. Some of them may be involved in resistance to environmental stress [].
Probab=64.36 E-value=6.8 Score=27.72 Aligned_cols=13 Identities=31% Similarity=0.087 Sum_probs=8.3
Q ss_pred ChhhhHHHHHHHH
Q 045282 1 MAVLVQKILYATL 13 (130)
Q Consensus 1 Ma~~~~k~l~~~l 13 (130)
||++.+-||.|+|
T Consensus 1 MaSK~~llL~l~L 13 (95)
T PF07172_consen 1 MASKAFLLLGLLL 13 (95)
T ss_pred CchhHHHHHHHHH
Confidence 8877766665444
No 5
>PF10633 NPCBM_assoc: NPCBM-associated, NEW3 domain of alpha-galactosidase; InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=56.95 E-value=25 Score=22.87 Aligned_cols=50 Identities=18% Similarity=0.296 Sum_probs=26.8
Q ss_pred CeeeEEEEEeeccccccceEEEec---CCccccccCCCcceeEeCCeeEEeCCcccCCCCeEEEEEc
Q 045282 48 GKKEFAVEVFNWCKCAQRNVTLDC---DGFQTVEKPDPVQMSISGFQCILLQGRDIIPFSRVHFKYA 111 (130)
Q Consensus 48 G~pe~~VtI~N~C~C~~~~V~l~C---~gF~S~~~VDP~ifr~~~~~CLVn~G~pi~~g~~v~F~YA 111 (130)
..-+++++|+|.+.-+..+++|+- .|+. ...+|.-+. .|.+|++.++++.
T Consensus 5 ~~~~~~~tv~N~g~~~~~~v~~~l~~P~GW~--~~~~~~~~~------------~l~pG~s~~~~~~ 57 (78)
T PF10633_consen 5 ETVTVTLTVTNTGTAPLTNVSLSLSLPEGWT--VSASPASVP------------SLPPGESVTVTFT 57 (78)
T ss_dssp EEEEEEEEEE--SSS-BSS-EEEEE--TTSE-----EEEEE--------------B-TTSEEEEEEE
T ss_pred CEEEEEEEEEECCCCceeeEEEEEeCCCCcc--ccCCccccc------------cCCCCCEEEEEEE
Confidence 345799999999987778888765 3544 223332221 6777877776653
No 6
>PLN02340 endoglucanase
Probab=54.45 E-value=12 Score=34.46 Aligned_cols=78 Identities=15% Similarity=0.139 Sum_probs=53.2
Q ss_pred CCCcceEEEeeecCccCC---eeeEEEEEeeccccccceEEEecCCcc-ccccCCCcceeEeCCeeEEeCC-cccCCCCe
Q 045282 31 PPETLDISQKETGNIVQG---KKEFAVEVFNWCKCAQRNVTLDCDGFQ-TVEKPDPVQMSISGFQCILLQG-RDIIPFSR 105 (130)
Q Consensus 31 ~~~di~V~Q~~tg~~v~G---~pe~~VtI~N~C~C~~~~V~l~C~gF~-S~~~VDP~ifr~~~~~CLVn~G-~pi~~g~~ 105 (130)
+..++.+.|.-|..+..+ .-+|+|+|+|+|.=+.+.+++.=..+- ..-.|.|++= .+++.+-+= ..|.+|+.
T Consensus 518 ~~~~~e~~~~~~~sw~~~g~~y~~~~v~i~N~s~~pi~~l~~~~~~l~g~lwgl~~~~~---~~~y~~p~~~~tl~~g~~ 594 (614)
T PLN02340 518 SGAPVEFVHSITNTWTAGGTTYYRHKVIIKNKSQKPITDLKLVIEDLSGPIWGLNPTKE---KNTYELPQWQKVLQPGSQ 594 (614)
T ss_pred CCCchhhhhhheeeeecCCceEEEEEEEEEeCCCCCchhhhhhhhhcccchhcceeccc---cCCccCchhhhccCCCCe
Confidence 355567778877776664 677999999999999999998774443 2222333211 244444443 47888999
Q ss_pred EEEEEc
Q 045282 106 VHFKYA 111 (130)
Q Consensus 106 v~F~YA 111 (130)
++|.|-
T Consensus 595 ~~f~yi 600 (614)
T PLN02340 595 LSFVYV 600 (614)
T ss_pred eEEEec
Confidence 999998
No 7
>PF07705 CARDB: CARDB; InterPro: IPR011635 The APHP (acidic peptide-dependent hydrolases/peptidase) domain is found in a variety of different proteins.; PDB: 2KUT_A 2L0D_A 3IDU_A 2KL6_A.
Probab=50.34 E-value=67 Score=20.72 Aligned_cols=68 Identities=12% Similarity=0.129 Sum_probs=32.8
Q ss_pred CcceE--EEeeecCccCCeeeEEEEEeeccccccceEEEecCCccccccCCCcceeEeCCeeEEeCCcccCCCCeEEEEE
Q 045282 33 ETLDI--SQKETGNIVQGKKEFAVEVFNWCKCAQRNVTLDCDGFQTVEKPDPVQMSISGFQCILLQGRDIIPFSRVHFKY 110 (130)
Q Consensus 33 ~di~V--~Q~~tg~~v~G~pe~~VtI~N~C~C~~~~V~l~C~gF~S~~~VDP~ifr~~~~~CLVn~G~pi~~g~~v~F~Y 110 (130)
-||.| ...+.-..++..-+..|+|.|.=.-...++.+.= +.+...+ +.-.| ..|.+|++..+++
T Consensus 2 pDL~v~~~~~~~~~~~g~~~~i~~~V~N~G~~~~~~~~v~~--~~~~~~~---------~~~~i---~~L~~g~~~~v~~ 67 (101)
T PF07705_consen 2 PDLTVSITVSPSNVVPGEPVTITVTVKNNGTADAENVTVRL--YLDGNSV---------STVTI---PSLAPGESETVTF 67 (101)
T ss_dssp --EEE-EEEC-SEEETTSEEEEEEEEEE-SSS-BEEEEEEE--EETTEEE---------EEEEE---SEB-TTEEEEEEE
T ss_pred CCEEEEEeeCCCcccCCCEEEEEEEEEECCCCCCCCEEEEE--EECCcee---------ccEEE---CCcCCCcEEEEEE
Confidence 36677 2222232334455699999998765566666651 1111111 11222 5788888766666
Q ss_pred ccCC
Q 045282 111 AFDD 114 (130)
Q Consensus 111 Aw~~ 114 (130)
.|..
T Consensus 68 ~~~~ 71 (101)
T PF07705_consen 68 TWTP 71 (101)
T ss_dssp EEE-
T ss_pred EEEe
Confidence 6554
No 8
>PF07127 Nodulin_late: Late nodulin protein; InterPro: IPR009810 This family consists of several plant specific late nodulin sequences which are homologous to the Pisum sativum (Garden pea) ENOD3 protein. ENOD3 is expressed in the late stages of root nodule formation and contains two pairs of cysteine residues toward the proteins C terminus which may be involved in metal-binding [].; GO: 0046872 metal ion binding, 0009878 nodule morphogenesis
Probab=43.35 E-value=34 Score=21.30 Aligned_cols=15 Identities=40% Similarity=0.744 Sum_probs=8.6
Q ss_pred ChhhhHHHHH-HHHHHH
Q 045282 1 MAVLVQKILY-ATLFLA 16 (130)
Q Consensus 1 Ma~~~~k~l~-~~lfl~ 16 (130)
|| +.+|++. +++||+
T Consensus 1 Ma-~ilKFvY~mIifls 16 (54)
T PF07127_consen 1 MA-KILKFVYAMIIFLS 16 (54)
T ss_pred Cc-cchhhHHHHHHHHH
Confidence 77 6777644 444444
No 9
>COG3900 Predicted periplasmic protein [Function unknown]
Probab=43.00 E-value=15 Score=30.57 Aligned_cols=19 Identities=32% Similarity=0.557 Sum_probs=17.0
Q ss_pred ecCccCCeeeEEEEEeecc
Q 045282 42 TGNIVQGKKEFAVEVFNWC 60 (130)
Q Consensus 42 tg~~v~G~pe~~VtI~N~C 60 (130)
|.+.+.|-|||+|++.|+=
T Consensus 206 Tsk~v~g~PqYtv~fsnwk 224 (262)
T COG3900 206 TSKDVPGEPQYTVVFSNWK 224 (262)
T ss_pred EecccCCCCcEEEEEcccc
Confidence 6778999999999999975
No 10
>PF03330 DPBB_1: Rare lipoprotein A (RlpA)-like double-psi beta-barrel; InterPro: IPR009009 Beta barrels are commonly observed in protein structures. They are classified in terms of two integral parameters: the number of strands in the sheet, n, and the shear number, S, a measure of the stagger of the strands in the beta-sheet. These two parameters have been shown to determine the major geometrical features of beta-barrels. Six-stranded beta-barrels with a pseudo-twofold axis are found in several proteins. One involving parallel strands forming two psi structures is known as the double-psi barrel. The first psi structure consists of the loop connecting strands beta1 and beta2 (a 'psi loop') and the strand beta5, whereas the second psi structure consists of the loop connecting strands beta4 and beta5 and the strand beta2. All the psi structures in double-psi barrels have a unique handedness, in that beta1 (beta4), beta2 (beta5) and the loop following beta5 (beta2) form a right-handed helix. The unique handedness may be related to the fact that the twisting angle between the parallel pair of strands is always larger than that between the antiparallel pair [].; PDB: 1N10_B 3D30_A 2BH0_A 2HCZ_X.
Probab=36.38 E-value=36 Score=22.18 Aligned_cols=37 Identities=24% Similarity=0.550 Sum_probs=26.2
Q ss_pred eeEEEEEeeccc-cccceEEEecCCccccccCCCccee
Q 045282 50 KEFAVEVFNWCK-CAQRNVTLDCDGFQTVEKPDPVQMS 86 (130)
Q Consensus 50 pe~~VtI~N~C~-C~~~~V~l~C~gF~S~~~VDP~ifr 86 (130)
..-.|+|+++|+ |...++-|+=..|..--..|..++.
T Consensus 38 ksV~v~V~D~Cp~~~~~~lDLS~~aF~~la~~~~G~i~ 75 (78)
T PF03330_consen 38 KSVTVTVVDRCPGCPPNHLDLSPAAFKALADPDAGVIP 75 (78)
T ss_dssp CEEEEEEEEE-TTSSSSEEEEEHHHHHHTBSTTCSSEE
T ss_pred CeEEEEEEccCCCCcCCEEEeCHHHHHHhCCCCceEEE
Confidence 567899999996 9999999887777654444444443
No 11
>PF06483 ChiC: Chitinase C; InterPro: IPR009470 This ~170 aa region is found at the C-terminal to the catalytic domain (IPR001223 from INTERPRO) found in members of glycoside hydrolase family 18.
Probab=31.12 E-value=99 Score=24.57 Aligned_cols=49 Identities=16% Similarity=0.172 Sum_probs=36.4
Q ss_pred cceEEEecCCcc---ccccCCCcceeEeCCeeEEeCCcccCCCCeEEEEEccCCccCe
Q 045282 64 QRNVTLDCDGFQ---TVEKPDPVQMSISGFQCILLQGRDIIPFSRVHFKYAFDDEFPF 118 (130)
Q Consensus 64 ~~~V~l~C~gF~---S~~~VDP~ifr~~~~~CLVn~G~pi~~g~~v~F~YAw~~~f~~ 118 (130)
.-||.+.=+||. +-=||+|++-= - =|.++.|+.|..|+|.|+-+.+-.+
T Consensus 35 ~ldv~v~~~gf~~GD~NYPI~Pkl~i-T-----Nns~~~iPGGt~~~FD~ptSa~~~~ 86 (180)
T PF06483_consen 35 ALDVSVSFTGFKLGDSNYPINPKLTI-T-----NNSGQTIPGGTEFEFDYPTSAPDNA 86 (180)
T ss_pred eEEEEEEeCCcccCCCCCCcCCcEEE-E-----cCCCcccCCccEEEEccccCCcccc
Confidence 447788888886 55789997532 1 2578899999999999997776543
No 12
>PF06682 DUF1183: Protein of unknown function (DUF1183); InterPro: IPR009567 This family consists of several eukaryotic proteins of around 360 residues in length. The function of this family is unknown.
Probab=30.09 E-value=88 Score=26.75 Aligned_cols=25 Identities=20% Similarity=0.478 Sum_probs=20.7
Q ss_pred cccceEEEecCCccccccCCCcceeEe
Q 045282 62 CAQRNVTLDCDGFQTVEKPDPVQMSIS 88 (130)
Q Consensus 62 C~~~~V~l~C~gF~S~~~VDP~ifr~~ 88 (130)
=....|.|.|.|+.+.+ ||=|||=+
T Consensus 91 ~klG~~~V~CEGY~~pd--DpyvLkGS 115 (318)
T PF06682_consen 91 YKLGSTDVSCEGYDYPD--DPYVLKGS 115 (318)
T ss_pred eeecceEEeeecccCCC--CceecCCc
Confidence 44567899999999977 89999944
No 13
>PF03293 Pox_RNA_pol: Poxvirus DNA-directed RNA polymerase, 18 kD subunit; InterPro: IPR004973 DNA-directed RNA polymerases 2.7.7.6 from EC (also known as DNA-dependent RNA polymerases) are responsible for the polymerisation of ribonucleotides into a sequence complementary to the template DNA. In eukaryotes, there are three different forms of DNA-directed RNA polymerases transcribing different sets of genes. Most RNA polymerases are multimeric enzymes and are composed of a variable number of subunits. The core RNA polymerase complex consists of five subunits (two alpha, one beta, one beta-prime and one omega) and is sufficient for transcription elongation and termination but is unable to initiate transcription. Transcription initiation from promoter elements requires a sixth, dissociable subunit called a sigma factor, which reversibly associates with the core RNA polymerase complex to form a holoenzyme []. The core RNA polymerase complex forms a "crab claw"-like structure with an internal channel running along the full length []. The key functional sites of the enzyme, as defined by mutational and cross-linking analysis, are located on the inner wall of this channel. RNA synthesis follows after the attachment of RNA polymerase to a specific site, the promoter, on the template DNA strand. The RNA synthesis process continues until a termination sequence is reached. The RNA product, which is synthesised in the 5' to 3'direction, is known as the primary transcript. Eukaryotic nuclei contain three distinct types of RNA polymerases that differ in the RNA they synthesise: RNA polymerase I: located in the nucleoli, synthesises precursors of most ribosomal RNAs. RNA polymerase II: occurs in the nucleoplasm, synthesises mRNA precursors. RNA polymerase III: also occurs in the nucleoplasm, synthesises the precursors of 5S ribosomal RNA, the tRNAs, and a variety of other small nuclear and cytosolic RNAs. Eukaryotic cells are also known to contain separate mitochondrial and chloroplast RNA polymerases. Eukaryotic RNA polymerases, whose molecular masses vary in size from 500 to 700 kDa, contain two non-identical large (>100 kDa) subunits and an array of up to 12 different small (less than 50 kDa) subunits. The Poxvirus DNA-directed RNA polymerase (2.7.7.6 from EC) catalyses DNA-template-directed extension of the 3'-end of an RNA strand by one nucleotide at a time. The enzyme consists of at least eight subunits, this is the 18 kDa subunit.; GO: 0003677 DNA binding, 0003899 DNA-directed RNA polymerase activity, 0019083 viral transcription
Probab=28.75 E-value=1e+02 Score=23.84 Aligned_cols=48 Identities=19% Similarity=0.217 Sum_probs=31.6
Q ss_pred ccceEEEecCCccccccCCCcceeEeC-CeeEEeCCcccCCCCeEEEEE
Q 045282 63 AQRNVTLDCDGFQTVEKPDPVQMSISG-FQCILLQGRDIIPFSRVHFKY 110 (130)
Q Consensus 63 ~~~~V~l~C~gF~S~~~VDP~ifr~~~-~~CLVn~G~pi~~g~~v~F~Y 110 (130)
..+||.+.|+..-=-..=|..-....+ .-|++.||..-..|+.|+-.-
T Consensus 93 dESni~V~CgDLiCkl~rdsGtVSf~dsKYCfirNg~vY~ngs~Vsv~L 141 (160)
T PF03293_consen 93 DESNITVQCGDLICKLSRDSGTVSFNDSKYCFIRNGVVYDNGSEVSVVL 141 (160)
T ss_pred ccCceEEEcCcEEEEeeccCCeEEecCceEEEEECCEEecCCCEEEEEe
Confidence 468899999875432222333333322 349999999999999887654
No 14
>PF14016 DUF4232: Protein of unknown function (DUF4232)
Probab=27.48 E-value=1e+02 Score=21.96 Aligned_cols=73 Identities=11% Similarity=0.145 Sum_probs=45.3
Q ss_pred CCCCcceEEEeeecCccCCeeeEEEEEeecc--ccccceEEEecCCccccccC-------CCcceeEeCCeeEEeCCccc
Q 045282 30 CPPETLDISQKETGNIVQGKKEFAVEVFNWC--KCAQRNVTLDCDGFQTVEKP-------DPVQMSISGFQCILLQGRDI 100 (130)
Q Consensus 30 C~~~di~V~Q~~tg~~v~G~pe~~VtI~N~C--~C~~~~V~l~C~gF~S~~~V-------DP~ifr~~~~~CLVn~G~pi 100 (130)
|...|++++-..... ..|...+.|+++|.= .|... ||..+..+ .+..-+.. + -..--.|
T Consensus 1 C~~~~L~~~~~~~~~-~~g~~~~~l~~tN~s~~~C~l~-------G~P~v~~~~~~g~~~~~~~~~~~-~---~~~~vtL 68 (131)
T PF14016_consen 1 CTAADLSVTVGPVDA-GAGQRHATLTFTNTSDTPCTLY-------GYPGVALVDADGAPLGVPAVREG-P---PPRPVTL 68 (131)
T ss_pred CCcccEEEEEecccC-CCCccEEEEEEEECCCCcEEec-------cCCcEEEECCCCCcCCccccccC-C---CCCcEEE
Confidence 888899998866533 468889999999976 37664 34333333 23222222 1 1222346
Q ss_pred CCCCeEEEEEccCC
Q 045282 101 IPFSRVHFKYAFDD 114 (130)
Q Consensus 101 ~~g~~v~F~YAw~~ 114 (130)
.+|++..|.-.|..
T Consensus 69 ~PG~sA~a~l~~~~ 82 (131)
T PF14016_consen 69 APGGSAYAGLRWSN 82 (131)
T ss_pred CCCCEEEEEEEEec
Confidence 78888888777754
No 15
>PF03032 Brevenin: Brevenin/esculentin/gaegurin/rugosin family; InterPro: IPR004275 In addition to the highly specific cell-mediated immune system, vertebrates possess an efficient host-defence mechanism against invading microorganisms which involves the synthesis of highly potent antimicrobial peptides with a large spectrum of activity. This entry represents a number of these defence peptides secreted from the skin of amphibians, including the opiate-like dermorphins and deltorphins, and the antimicrobial dermoseptins and temporins.; GO: 0006952 defense response, 0042742 defense response to bacterium, 0005576 extracellular region
Probab=27.44 E-value=47 Score=20.65 Aligned_cols=16 Identities=38% Similarity=0.455 Sum_probs=10.9
Q ss_pred hHHHHHHHHHHHHHhc
Q 045282 5 VQKILYATLFLALISE 20 (130)
Q Consensus 5 ~~k~l~~~lfl~lv~~ 20 (130)
..|-|.|++||-+|+=
T Consensus 3 lKKsllLlfflG~ISl 18 (46)
T PF03032_consen 3 LKKSLLLLFFLGTISL 18 (46)
T ss_pred chHHHHHHHHHHHccc
Confidence 3566777788877763
No 16
>PLN00115 pollen allergen group 3; Provisional
Probab=27.04 E-value=2.6e+02 Score=20.53 Aligned_cols=90 Identities=18% Similarity=0.165 Sum_probs=53.3
Q ss_pred ChhhhHHHHHHHHHHHHHhccCCCCCCCCCCCCcceEEEeeecCccCCeeeEEEEEeeccccccceEEEecCCccccccC
Q 045282 1 MAVLVQKILYATLFLALISETMSQPEQGPCPPETLDISQKETGNIVQGKKEFAVEVFNWCKCAQRNVTLDCDGFQTVEKP 80 (130)
Q Consensus 1 Ma~~~~k~l~~~lfl~lv~~g~~~~~~~~C~~~di~V~Q~~tg~~v~G~pe~~VtI~N~C~C~~~~V~l~C~gF~S~~~V 80 (130)
|++... +|++..+-.|..-| .|.. +|.++=.. | -.|.|-|-+.|. .+..|.+.-.| +.+-+
T Consensus 1 ~~~~~~-~~~~~~~a~l~~~~-------~~g~-~v~F~V~~-g----Snp~yL~ll~~~---dI~~V~Ik~~g--~~~W~ 61 (118)
T PLN00115 1 MSSLSF-LLLAVALAALFAVG-------SCAT-EVTFKVGK-G----SSSTSLELVTNV---AISEVEIKEKG--AKDWV 61 (118)
T ss_pred CchhHH-HHHHHHHHHHhhhh-------hcCC-ceEEEECC-C----CCcceEEEEEeC---CEEEEEEeecC--CCccc
Confidence 563333 55555555666666 5765 45444222 1 137888888865 57788887754 33334
Q ss_pred CCcceeEe-CCeeEEeCCcccCCCCeEEEEEccC
Q 045282 81 DPVQMSIS-GFQCILLQGRDIIPFSRVHFKYAFD 113 (130)
Q Consensus 81 DP~ifr~~-~~~CLVn~G~pi~~g~~v~F~YAw~ 113 (130)
|| +++. |..=-++.++|+. | +++|+...+
T Consensus 62 ~~--M~rswGavW~~~s~~pl~-G-PlS~R~t~~ 91 (118)
T PLN00115 62 DD--LKESSTNTWTLKSKAPLK-G-PFSVRFLVK 91 (118)
T ss_pred Cc--cccCccceeEecCCCCCC-C-ceEEEEEEe
Confidence 44 4565 5555566678876 4 788887654
No 17
>PF01345 DUF11: Domain of unknown function DUF11; InterPro: IPR001434 This group of sequences is represented by a conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis (Streptococcus faecalis), and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydia trachomatis outer membrane proteins. In C. trachomatis, three cysteine-rich proteins (also believed to be lipoproteins), MOMP, OMP6 and OMP3, make up the extracellular matrix of the outer membrane []. They are involved in the essential structural integrity of both the elementary body (EB) and recticulate body (RB) phase. They are thought to be involved in porin formation and, as these bacteria lack the peptidoglycan layer common to most Gram-negative microbes, such proteins are highly important in the pathogenicity of the organism.; GO: 0005727 extrachromosomal circular DNA
Probab=25.54 E-value=1.6e+02 Score=18.71 Aligned_cols=29 Identities=14% Similarity=0.062 Sum_probs=23.0
Q ss_pred cCccCCeeeEEEEEeeccccccceEEEec
Q 045282 43 GNIVQGKKEFAVEVFNWCKCAQRNVTLDC 71 (130)
Q Consensus 43 g~~v~G~pe~~VtI~N~C~C~~~~V~l~C 71 (130)
.-.++..-+|.++|+|.=.-+..||.|.-
T Consensus 36 ~~~~Gd~v~ytitvtN~G~~~a~nv~v~D 64 (76)
T PF01345_consen 36 TANPGDTVTYTITVTNTGPAPATNVVVTD 64 (76)
T ss_pred cccCCCEEEEEEEEEECCCCeeEeEEEEE
Confidence 33455677899999999988888888864
No 18
>PF10731 Anophelin: Thrombin inhibitor from mosquito; InterPro: IPR018932 Members of this family are all inhibitors of thrombin, the peptidase that is at the end of the blood coagulation cascade and which creates the clot by cleaving fibrinogen. The interaction between thrombin and fibrinogen involves two different areas of contact - via the thrombin active site and via a second substrate-binding site known as an exosite. The inhibitor acts by blocking the exosite, rather than by interacting with the active site. The inhibitors are from mosquitoes that feed on human blood and which, by inhibiting thrombin, prevent the blood from clotting and keep it flowing.
Probab=25.00 E-value=73 Score=21.28 Aligned_cols=21 Identities=24% Similarity=0.184 Sum_probs=12.2
Q ss_pred ChhhhHHHHHHHHHHHHHhcc
Q 045282 1 MAVLVQKILYATLFLALISET 21 (130)
Q Consensus 1 Ma~~~~k~l~~~lfl~lv~~g 21 (130)
||++++-+.+|.+.|+.+.|+
T Consensus 1 MA~Kl~vialLC~aLva~vQ~ 21 (65)
T PF10731_consen 1 MASKLIVIALLCVALVAIVQS 21 (65)
T ss_pred CcchhhHHHHHHHHHHHHHhc
Confidence 886655554444555556666
No 19
>KOG4063 consensus Major epididymal secretory protein HE1 [Function unknown]
Probab=23.94 E-value=1.5e+02 Score=23.04 Aligned_cols=41 Identities=12% Similarity=0.101 Sum_probs=21.9
Q ss_pred ChhhhHHHHHHHHHHHHHh-ccCCCCCCCCCCCCcceEEEeee
Q 045282 1 MAVLVQKILYATLFLALIS-ETMSQPEQGPCPPETLDISQKET 42 (130)
Q Consensus 1 Ma~~~~k~l~~~lfl~lv~-~g~~~~~~~~C~~~di~V~Q~~t 42 (130)
|+-++++.+++.++|.+-. |. -+-...+|.-+|-.+.+.+.
T Consensus 1 m~ms~~~~v~l~alls~a~aq~-~~t~~k~C~ss~g~~~~V~i 42 (158)
T KOG4063|consen 1 MMMSFLKTVILLALLSLAAAQA-ISTGVKQCGSSDGTPLEVKI 42 (158)
T ss_pred CchHHHHHHHHHHHHHHhhhcc-cCcccccccCCCCcceEEEe
Confidence 5545566555444444433 11 01124469888877777665
No 20
>PRK15249 fimbrial chaperone protein StbB; Provisional
Probab=22.08 E-value=1.3e+02 Score=24.40 Aligned_cols=90 Identities=14% Similarity=-0.007 Sum_probs=46.9
Q ss_pred hhHHHHHHHHHHHHHhccCCCCCCCCCCCCcceEEEeeecCccCCeeeEEEEEeeccccccceEEEecCCcc--------
Q 045282 4 LVQKILYATLFLALISETMSQPEQGPCPPETLDISQKETGNIVQGKKEFAVEVFNWCKCAQRNVTLDCDGFQ-------- 75 (130)
Q Consensus 4 ~~~k~l~~~lfl~lv~~g~~~~~~~~C~~~di~V~Q~~tg~~v~G~pe~~VtI~N~C~C~~~~V~l~C~gF~-------- 75 (130)
+.+++|.+++|+++-+. +....|.|..++.- ..++..+=.|+|.|.=.= ..-|..+=+.-.
T Consensus 8 ~~~~~~~~~~~~~~~~~---------~a~A~l~l~~TRvi-y~~~~~~~sl~l~N~~~~-p~LvQsWv~~~~~~~~p~~~ 76 (253)
T PRK15249 8 SALYYLIVFLFLALPAT---------ASWASVTILGSRII-YPSTASSVDVQLKNNDAI-PYIVQTWFDDGDMNTSPENS 76 (253)
T ss_pred hHHHHHHHHHHHHhhhH---------hheeEEEeCceEEE-EeCCCcceeEEEEcCCCC-cEEEEEEEeCCCCCCCcccc
Confidence 34666655444433222 23456888887763 445678888899886531 122221111111
Q ss_pred --ccccCCCcceeEeC-Cee---EEeCC-cccCCCC
Q 045282 76 --TVEKPDPVQMSISG-FQC---ILLQG-RDIIPFS 104 (130)
Q Consensus 76 --S~~~VDP~ifr~~~-~~C---LVn~G-~pi~~g~ 104 (130)
..-.|-|-+||... ..= ++..| .+++...
T Consensus 77 ~~~pFivtPPlfrl~p~~~q~lRI~~~~~~~lP~DR 112 (253)
T PRK15249 77 SAMPFIATPPVFRIQPKAGQVVRVIYNNTKKLPQDR 112 (253)
T ss_pred ccCcEEEcCCeEEecCCCceEEEEEEcCCCCCCCCc
Confidence 11347788999884 222 23344 3565543
No 21
>PF01456 Mucin: Mucin-like glycoprotein; InterPro: IPR000458 This family of trypanosomal proteins resemble vertebrate mucins. The protein consists of three regions. The N and C terminii are conserved between all members of the family, whereas the central region is not well conserved and contains a large number of threonine residues which can be glycosylated []. Indirect evidence suggested that these genes might encode the core protein of parasite mucins, glycoproteins that were proposed to be involved in the interaction with, and invasion of, mammalian host cells.
Probab=21.22 E-value=83 Score=22.68 Aligned_cols=14 Identities=43% Similarity=0.503 Sum_probs=11.4
Q ss_pred hHHHHHHHHHHHHH
Q 045282 5 VQKILYATLFLALI 18 (130)
Q Consensus 5 ~~k~l~~~lfl~lv 18 (130)
...||+.+|+|+|+
T Consensus 3 tcRLLCalLvlaLc 16 (143)
T PF01456_consen 3 TCRLLCALLVLALC 16 (143)
T ss_pred hHHHHHHHHHHHHH
Confidence 58899988888874
No 22
>PF05753 TRAP_beta: Translocon-associated protein beta (TRAPB); InterPro: IPR008856 This family consists of several eukaryotic translocon-associated protein beta (TRAPB) or signal sequence receptor beta subunit (SSR-beta) proteins. The normal translocation of nascent polypeptides into the lumen of the endoplasmic reticulum (ER) is thought to be aided in part by a translocon-associated protein (TRAP) complex consisting of 4 protein subunits. The association of mature proteins with the ER and Golgi, or other intracellular locales, such as lysosomes, depends on the initial targeting of the nascent polypeptide to the ER membrane. A similar scenario must also exist for proteins destined for secretion [].; GO: 0005783 endoplasmic reticulum, 0016021 integral to membrane
Probab=20.61 E-value=4.1e+02 Score=20.59 Aligned_cols=62 Identities=19% Similarity=0.249 Sum_probs=41.5
Q ss_pred cCccC-CeeeEEEEEeeccccccceEEEecCCccccccCCCcceeEeCCeeEEeCCcccCCCCeEEEEEc
Q 045282 43 GNIVQ-GKKEFAVEVFNWCKCAQRNVTLDCDGFQTVEKPDPVQMSISGFQCILLQGRDIIPFSRVHFKYA 111 (130)
Q Consensus 43 g~~v~-G~pe~~VtI~N~C~C~~~~V~l~C~gF~S~~~VDP~ifr~~~~~CLVn~G~pi~~g~~v~F~YA 111 (130)
.-++. -.-+.+++|.|.=.=+-.||+|.=++|. |.-|...++. +=-.=..|++|+.++..|.
T Consensus 32 ~~~v~g~~v~V~~~iyN~G~~~A~dV~l~D~~fp------~~~F~lvsG~-~s~~~~~i~pg~~vsh~~v 94 (181)
T PF05753_consen 32 KYLVEGEDVTVTYTIYNVGSSAAYDVKLTDDSFP------PEDFELVSGS-LSASWERIPPGENVSHSYV 94 (181)
T ss_pred ccccCCcEEEEEEEEEECCCCeEEEEEEECCCCC------ccccEeccCc-eEEEEEEECCCCeEEEEEE
Confidence 33443 3567999999999999999999987874 3455544321 1111137888888877775
No 23
>TIGR01451 B_ant_repeat conserved repeat domain. This model represents the conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis, and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydial outer membrane proteins.
Probab=20.03 E-value=2.1e+02 Score=17.48 Aligned_cols=25 Identities=16% Similarity=0.240 Sum_probs=18.7
Q ss_pred cCCeeeEEEEEeeccccccceEEEe
Q 045282 46 VQGKKEFAVEVFNWCKCAQRNVTLD 70 (130)
Q Consensus 46 v~G~pe~~VtI~N~C~C~~~~V~l~ 70 (130)
++-.-+|+++|.|.-.=+..+|.|.
T Consensus 10 ~Gd~v~Yti~v~N~g~~~a~~v~v~ 34 (53)
T TIGR01451 10 IGDTITYTITVTNNGNVPATNVVVT 34 (53)
T ss_pred CCCEEEEEEEEEECCCCceEeEEEE
Confidence 4556789999999886666666664
Done!