Query         045282
Match_columns 130
No_of_seqs    106 out of 133
Neff          5.1 
Searched_HMMs 46136
Date          Fri Mar 29 11:41:21 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/045282.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/045282hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF09478 CBM49:  Carbohydrate b  97.1  0.0044 9.6E-08   41.9   7.8   73   35-111     1-78  (80)
  2 PLN02171 endoglucanase          92.7    0.43 9.2E-06   43.9   7.2   76   33-113   535-615 (629)
  3 PF02933 CDC48_2:  Cell divisio  67.4       5 0.00011   25.6   2.1   29   96-125    15-43  (64)
  4 PF07172 GRP:  Glycine rich pro  64.4     6.8 0.00015   27.7   2.5   13    1-13      1-13  (95)
  5 PF10633 NPCBM_assoc:  NPCBM-as  56.9      25 0.00054   22.9   4.1   50   48-111     5-57  (78)
  6 PLN02340 endoglucanase          54.5      12 0.00027   34.5   3.1   78   31-111   518-600 (614)
  7 PF07705 CARDB:  CARDB;  InterP  50.3      67  0.0015   20.7   6.1   68   33-114     2-71  (101)
  8 PF07127 Nodulin_late:  Late no  43.3      34 0.00074   21.3   3.0   15    1-16      1-16  (54)
  9 COG3900 Predicted periplasmic   43.0      15 0.00032   30.6   1.5   19   42-60    206-224 (262)
 10 PF03330 DPBB_1:  Rare lipoprot  36.4      36 0.00078   22.2   2.4   37   50-86     38-75  (78)
 11 PF06483 ChiC:  Chitinase C;  I  31.1      99  0.0021   24.6   4.4   49   64-118    35-86  (180)
 12 PF06682 DUF1183:  Protein of u  30.1      88  0.0019   26.8   4.3   25   62-88     91-115 (318)
 13 PF03293 Pox_RNA_pol:  Poxvirus  28.8   1E+02  0.0023   23.8   4.0   48   63-110    93-141 (160)
 14 PF14016 DUF4232:  Protein of u  27.5   1E+02  0.0022   22.0   3.7   73   30-114     1-82  (131)
 15 PF03032 Brevenin:  Brevenin/es  27.4      47   0.001   20.7   1.6   16    5-20      3-18  (46)
 16 PLN00115 pollen allergen group  27.0 2.6E+02  0.0056   20.5   9.1   90    1-113     1-91  (118)
 17 PF01345 DUF11:  Domain of unkn  25.5 1.6E+02  0.0035   18.7   4.1   29   43-71     36-64  (76)
 18 PF10731 Anophelin:  Thrombin i  25.0      73  0.0016   21.3   2.3   21    1-21      1-21  (65)
 19 KOG4063 Major epididymal secre  23.9 1.5E+02  0.0034   23.0   4.2   41    1-42      1-42  (158)
 20 PRK15249 fimbrial chaperone pr  22.1 1.3E+02  0.0028   24.4   3.8   90    4-104     8-112 (253)
 21 PF01456 Mucin:  Mucin-like gly  21.2      83  0.0018   22.7   2.2   14    5-18      3-16  (143)
 22 PF05753 TRAP_beta:  Translocon  20.6 4.1E+02  0.0089   20.6   9.1   62   43-111    32-94  (181)
 23 TIGR01451 B_ant_repeat conserv  20.0 2.1E+02  0.0047   17.5   3.7   25   46-70     10-34  (53)

No 1  
>PF09478 CBM49:  Carbohydrate binding domain CBM49;  InterPro: IPR019028 A carbohydrate-binding module (CBM) is defined as a contiguous amino acid sequence within a carbohydrate-active enzyme with a discreet fold having carbohydrate-binding activity. A few exceptions are CBMs in cellulosomal scaffolding proteins and rare instances of independent putative CBMs. The requirement of CBMs existing as modules within larger enzymes sets this class of carbohydrate-binding protein apart from other non-catalytic sugar binding proteins such as lectins and sugar transport proteins. CBMs were previously classified as cellulose-binding domains (CBDs) based on the initial discovery of several modules that bound cellulose [, ]. However, additional modules in carbohydrate-active enzymes are continually being found that bind carbohydrates other than cellulose yet otherwise meet the CBM criteria, hence the need to reclassify these polypeptides using more inclusive terminology. Previous classification of cellulose-binding domains were based on amino acid similarity. Groupings of CBDs were called "Types" and numbered with roman numerals (e.g. Type I or Type II CBDs). In keeping with the glycoside hydrolase classification, these groupings are now called families and numbered with Arabic numerals. Families 1 to 13 are the same as Types I to XIII. For a detailed review on the structure and binding modes of CBMs see [].  This domain is found at the C-terminal of cellulases and in vitro binding studies have shown it to binds to crystalline cellulose []. ; GO: 0030246 carbohydrate binding, 0005576 extracellular region
Probab=97.10  E-value=0.0044  Score=41.92  Aligned_cols=73  Identities=16%  Similarity=0.172  Sum_probs=56.1

Q ss_pred             ceEEEeeecCccCC---eeeEEEEEeeccccccceEEEecCCccc-cccCCCcceeEeCCeeEEeCC-cccCCCCeEEEE
Q 045282           35 LDISQKETGNIVQG---KKEFAVEVFNWCKCAQRNVTLDCDGFQT-VEKPDPVQMSISGFQCILLQG-RDIIPFSRVHFK  109 (130)
Q Consensus        35 i~V~Q~~tg~~v~G---~pe~~VtI~N~C~C~~~~V~l~C~gF~S-~~~VDP~ifr~~~~~CLVn~G-~pi~~g~~v~F~  109 (130)
                      |+|.|..+..+..|   ..+|.|+|+|.+.=+++++++.-+.+.+ .=    .+-+..++..-+=+- .+|.+|++.+|-
T Consensus         1 i~i~q~~~~sW~~~g~~y~qy~v~I~N~~~~~I~~~~i~~~~l~~~iW----~l~~~~~~~y~lPs~~~~i~pg~s~~FG   76 (80)
T PF09478_consen    1 ITITQTLVNSWTENGQTYTQYDVTITNNGSKPIKSLKISIDNLYGSIW----GLDKVSGNTYTLPSYQPTIKPGQSFTFG   76 (80)
T ss_pred             CEEEEEEEeEEEeCCEEEEEEEEEEEECCCCeEEEEEEEECccchhhe----eEEeccCCEEECCccccccCCCCEEEEE
Confidence            68899999888775   3579999999999999999999987651 11    222334577777454 399999999999


Q ss_pred             Ec
Q 045282          110 YA  111 (130)
Q Consensus       110 YA  111 (130)
                      |-
T Consensus        77 YI   78 (80)
T PF09478_consen   77 YI   78 (80)
T ss_pred             EE
Confidence            84


No 2  
>PLN02171 endoglucanase
Probab=92.74  E-value=0.43  Score=43.95  Aligned_cols=76  Identities=14%  Similarity=0.157  Sum_probs=55.0

Q ss_pred             CcceEEEeeecCccC---CeeeEEEEEeeccccccceEEEecCCccc-cccCCCcceeEeCCeeEEeCCc-ccCCCCeEE
Q 045282           33 ETLDISQKETGNIVQ---GKKEFAVEVFNWCKCAQRNVTLDCDGFQT-VEKPDPVQMSISGFQCILLQGR-DIIPFSRVH  107 (130)
Q Consensus        33 ~di~V~Q~~tg~~v~---G~pe~~VtI~N~C~C~~~~V~l~C~gF~S-~~~VDP~ifr~~~~~CLVn~G~-pi~~g~~v~  107 (130)
                      +.|+|.|..++.+..   +..+|+|+|+|++..+++++++.=..+-. .=    .+. ..++...+=+-. -|.+|+..+
T Consensus       535 ~ei~i~q~v~~sW~~~g~~y~qy~v~I~N~s~~~ik~i~i~~~~~~~~iW----~v~-~~~ngytlPs~~~sL~aG~s~t  609 (629)
T PLN02171        535 SPIEIEQKATASWKAKGRTYYRYSTTVTNRSAKTLKELHLGISKLYGPLW----GLT-KAGYGYVLPSWMPSLPAGKSLE  609 (629)
T ss_pred             ceeEEEEEEEEEEEcCCceEEEEEEEEEECCCCceeeeeeeeccccccch----hee-ecCCcccCchhhcccCCCCeeE
Confidence            368999999988885   47889999999999999999996544421 11    111 234445554443 788899999


Q ss_pred             EEEccC
Q 045282          108 FKYAFD  113 (130)
Q Consensus       108 F~YAw~  113 (130)
                      |-|=..
T Consensus       610 FgyI~~  615 (629)
T PLN02171        610 FVYVHS  615 (629)
T ss_pred             EEeecC
Confidence            999854


No 3  
>PF02933 CDC48_2:  Cell division protein 48 (CDC48), domain 2;  InterPro: IPR004201 This domain has a double psi-beta barrel fold and includes VCP-like ATPase and N-ethylmaleimide sensitive fusion protein N-terminal domains. Both the VAT and NSF N-terminal functional domains consist of two structural domains of which this is at the C terminus. The VAT-N domain found in AAA ATPases (IPR003959 from INTERPRO) is a substrate 185-residue recognition domain [].; GO: 0005524 ATP binding; PDB: 1QDN_B 1QCS_A 1CR5_C 3QQ8_A 3HU2_A 3HU1_E 3HU3_A 3QWZ_A 3TIW_B 3QQ7_A ....
Probab=67.35  E-value=5  Score=25.60  Aligned_cols=29  Identities=28%  Similarity=0.574  Sum_probs=23.7

Q ss_pred             CCcccCCCCeEEEEEccCCccCeEeeeeec
Q 045282           96 QGRDIIPFSRVHFKYAFDDEFPFYVFSSAP  125 (130)
Q Consensus        96 ~G~pi~~g~~v~F~YAw~~~f~~~p~ss~~  125 (130)
                      .|+|+..|+.|.|.|. ...++|.+.+.++
T Consensus        15 ~~~pv~~Gd~i~~~~~-~~~~~~~V~~~~P   43 (64)
T PF02933_consen   15 EGRPVTKGDTIVFPFF-GQALPFKVVSTEP   43 (64)
T ss_dssp             TTEEEETT-EEEEEET-TEEEEEEEEEECS
T ss_pred             cCCCccCCCEEEEEeC-CcEEEEEEEEEEc
Confidence            4689999999999996 6889999987654


No 4  
>PF07172 GRP:  Glycine rich protein family;  InterPro: IPR010800 This family consists of glycine rich proteins. Some of them may be involved in resistance to environmental stress [].
Probab=64.36  E-value=6.8  Score=27.72  Aligned_cols=13  Identities=31%  Similarity=0.087  Sum_probs=8.3

Q ss_pred             ChhhhHHHHHHHH
Q 045282            1 MAVLVQKILYATL   13 (130)
Q Consensus         1 Ma~~~~k~l~~~l   13 (130)
                      ||++.+-||.|+|
T Consensus         1 MaSK~~llL~l~L   13 (95)
T PF07172_consen    1 MASKAFLLLGLLL   13 (95)
T ss_pred             CchhHHHHHHHHH
Confidence            8877766665444


No 5  
>PF10633 NPCBM_assoc:  NPCBM-associated, NEW3 domain of alpha-galactosidase;  InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=56.95  E-value=25  Score=22.87  Aligned_cols=50  Identities=18%  Similarity=0.296  Sum_probs=26.8

Q ss_pred             CeeeEEEEEeeccccccceEEEec---CCccccccCCCcceeEeCCeeEEeCCcccCCCCeEEEEEc
Q 045282           48 GKKEFAVEVFNWCKCAQRNVTLDC---DGFQTVEKPDPVQMSISGFQCILLQGRDIIPFSRVHFKYA  111 (130)
Q Consensus        48 G~pe~~VtI~N~C~C~~~~V~l~C---~gF~S~~~VDP~ifr~~~~~CLVn~G~pi~~g~~v~F~YA  111 (130)
                      ..-+++++|+|.+.-+..+++|+-   .|+.  ...+|.-+.            .|.+|++.++++.
T Consensus         5 ~~~~~~~tv~N~g~~~~~~v~~~l~~P~GW~--~~~~~~~~~------------~l~pG~s~~~~~~   57 (78)
T PF10633_consen    5 ETVTVTLTVTNTGTAPLTNVSLSLSLPEGWT--VSASPASVP------------SLPPGESVTVTFT   57 (78)
T ss_dssp             EEEEEEEEEE--SSS-BSS-EEEEE--TTSE-----EEEEE--------------B-TTSEEEEEEE
T ss_pred             CEEEEEEEEEECCCCceeeEEEEEeCCCCcc--ccCCccccc------------cCCCCCEEEEEEE
Confidence            345799999999987778888765   3544  223332221            6777877776653


No 6  
>PLN02340 endoglucanase
Probab=54.45  E-value=12  Score=34.46  Aligned_cols=78  Identities=15%  Similarity=0.139  Sum_probs=53.2

Q ss_pred             CCCcceEEEeeecCccCC---eeeEEEEEeeccccccceEEEecCCcc-ccccCCCcceeEeCCeeEEeCC-cccCCCCe
Q 045282           31 PPETLDISQKETGNIVQG---KKEFAVEVFNWCKCAQRNVTLDCDGFQ-TVEKPDPVQMSISGFQCILLQG-RDIIPFSR  105 (130)
Q Consensus        31 ~~~di~V~Q~~tg~~v~G---~pe~~VtI~N~C~C~~~~V~l~C~gF~-S~~~VDP~ifr~~~~~CLVn~G-~pi~~g~~  105 (130)
                      +..++.+.|.-|..+..+   .-+|+|+|+|+|.=+.+.+++.=..+- ..-.|.|++=   .+++.+-+= ..|.+|+.
T Consensus       518 ~~~~~e~~~~~~~sw~~~g~~y~~~~v~i~N~s~~pi~~l~~~~~~l~g~lwgl~~~~~---~~~y~~p~~~~tl~~g~~  594 (614)
T PLN02340        518 SGAPVEFVHSITNTWTAGGTTYYRHKVIIKNKSQKPITDLKLVIEDLSGPIWGLNPTKE---KNTYELPQWQKVLQPGSQ  594 (614)
T ss_pred             CCCchhhhhhheeeeecCCceEEEEEEEEEeCCCCCchhhhhhhhhcccchhcceeccc---cCCccCchhhhccCCCCe
Confidence            355567778877776664   677999999999999999998774443 2222333211   244444443 47888999


Q ss_pred             EEEEEc
Q 045282          106 VHFKYA  111 (130)
Q Consensus       106 v~F~YA  111 (130)
                      ++|.|-
T Consensus       595 ~~f~yi  600 (614)
T PLN02340        595 LSFVYV  600 (614)
T ss_pred             eEEEec
Confidence            999998


No 7  
>PF07705 CARDB:  CARDB;  InterPro: IPR011635 The APHP (acidic peptide-dependent hydrolases/peptidase) domain is found in a variety of different proteins.; PDB: 2KUT_A 2L0D_A 3IDU_A 2KL6_A.
Probab=50.34  E-value=67  Score=20.72  Aligned_cols=68  Identities=12%  Similarity=0.129  Sum_probs=32.8

Q ss_pred             CcceE--EEeeecCccCCeeeEEEEEeeccccccceEEEecCCccccccCCCcceeEeCCeeEEeCCcccCCCCeEEEEE
Q 045282           33 ETLDI--SQKETGNIVQGKKEFAVEVFNWCKCAQRNVTLDCDGFQTVEKPDPVQMSISGFQCILLQGRDIIPFSRVHFKY  110 (130)
Q Consensus        33 ~di~V--~Q~~tg~~v~G~pe~~VtI~N~C~C~~~~V~l~C~gF~S~~~VDP~ifr~~~~~CLVn~G~pi~~g~~v~F~Y  110 (130)
                      -||.|  ...+.-..++..-+..|+|.|.=.-...++.+.=  +.+...+         +.-.|   ..|.+|++..+++
T Consensus         2 pDL~v~~~~~~~~~~~g~~~~i~~~V~N~G~~~~~~~~v~~--~~~~~~~---------~~~~i---~~L~~g~~~~v~~   67 (101)
T PF07705_consen    2 PDLTVSITVSPSNVVPGEPVTITVTVKNNGTADAENVTVRL--YLDGNSV---------STVTI---PSLAPGESETVTF   67 (101)
T ss_dssp             --EEE-EEEC-SEEETTSEEEEEEEEEE-SSS-BEEEEEEE--EETTEEE---------EEEEE---SEB-TTEEEEEEE
T ss_pred             CCEEEEEeeCCCcccCCCEEEEEEEEEECCCCCCCCEEEEE--EECCcee---------ccEEE---CCcCCCcEEEEEE
Confidence            36677  2222232334455699999998765566666651  1111111         11222   5788888766666


Q ss_pred             ccCC
Q 045282          111 AFDD  114 (130)
Q Consensus       111 Aw~~  114 (130)
                      .|..
T Consensus        68 ~~~~   71 (101)
T PF07705_consen   68 TWTP   71 (101)
T ss_dssp             EEE-
T ss_pred             EEEe
Confidence            6554


No 8  
>PF07127 Nodulin_late:  Late nodulin protein;  InterPro: IPR009810 This family consists of several plant specific late nodulin sequences which are homologous to the Pisum sativum (Garden pea) ENOD3 protein. ENOD3 is expressed in the late stages of root nodule formation and contains two pairs of cysteine residues toward the proteins C terminus which may be involved in metal-binding [].; GO: 0046872 metal ion binding, 0009878 nodule morphogenesis
Probab=43.35  E-value=34  Score=21.30  Aligned_cols=15  Identities=40%  Similarity=0.744  Sum_probs=8.6

Q ss_pred             ChhhhHHHHH-HHHHHH
Q 045282            1 MAVLVQKILY-ATLFLA   16 (130)
Q Consensus         1 Ma~~~~k~l~-~~lfl~   16 (130)
                      || +.+|++. +++||+
T Consensus         1 Ma-~ilKFvY~mIifls   16 (54)
T PF07127_consen    1 MA-KILKFVYAMIIFLS   16 (54)
T ss_pred             Cc-cchhhHHHHHHHHH
Confidence            77 6777644 444444


No 9  
>COG3900 Predicted periplasmic protein [Function unknown]
Probab=43.00  E-value=15  Score=30.57  Aligned_cols=19  Identities=32%  Similarity=0.557  Sum_probs=17.0

Q ss_pred             ecCccCCeeeEEEEEeecc
Q 045282           42 TGNIVQGKKEFAVEVFNWC   60 (130)
Q Consensus        42 tg~~v~G~pe~~VtI~N~C   60 (130)
                      |.+.+.|-|||+|++.|+=
T Consensus       206 Tsk~v~g~PqYtv~fsnwk  224 (262)
T COG3900         206 TSKDVPGEPQYTVVFSNWK  224 (262)
T ss_pred             EecccCCCCcEEEEEcccc
Confidence            6778999999999999975


No 10 
>PF03330 DPBB_1:  Rare lipoprotein A (RlpA)-like double-psi beta-barrel;  InterPro: IPR009009  Beta barrels are commonly observed in protein structures. They are classified in terms of two integral parameters: the number of strands in the sheet, n, and the shear number, S, a measure of the stagger of the strands in the beta-sheet. These two parameters have been shown to determine the major geometrical features of beta-barrels. Six-stranded beta-barrels with a pseudo-twofold axis are found in several proteins. One involving parallel strands forming two psi structures is known as the double-psi barrel. The first psi structure consists of the loop connecting strands beta1 and beta2 (a 'psi loop') and the strand beta5, whereas the second psi structure consists of the loop connecting strands beta4 and beta5 and the strand beta2. All the psi structures in double-psi barrels have a unique handedness, in that beta1 (beta4), beta2 (beta5) and the loop following beta5 (beta2) form a right-handed helix. The unique handedness may be related to the fact that the twisting angle between the parallel pair of strands is always larger than that between the antiparallel pair [].; PDB: 1N10_B 3D30_A 2BH0_A 2HCZ_X.
Probab=36.38  E-value=36  Score=22.18  Aligned_cols=37  Identities=24%  Similarity=0.550  Sum_probs=26.2

Q ss_pred             eeEEEEEeeccc-cccceEEEecCCccccccCCCccee
Q 045282           50 KEFAVEVFNWCK-CAQRNVTLDCDGFQTVEKPDPVQMS   86 (130)
Q Consensus        50 pe~~VtI~N~C~-C~~~~V~l~C~gF~S~~~VDP~ifr   86 (130)
                      ..-.|+|+++|+ |...++-|+=..|..--..|..++.
T Consensus        38 ksV~v~V~D~Cp~~~~~~lDLS~~aF~~la~~~~G~i~   75 (78)
T PF03330_consen   38 KSVTVTVVDRCPGCPPNHLDLSPAAFKALADPDAGVIP   75 (78)
T ss_dssp             CEEEEEEEEE-TTSSSSEEEEEHHHHHHTBSTTCSSEE
T ss_pred             CeEEEEEEccCCCCcCCEEEeCHHHHHHhCCCCceEEE
Confidence            567899999996 9999999887777654444444443


No 11 
>PF06483 ChiC:  Chitinase C;  InterPro: IPR009470 This ~170 aa region is found at the C-terminal to the catalytic domain (IPR001223 from INTERPRO) found in members of glycoside hydrolase family 18.
Probab=31.12  E-value=99  Score=24.57  Aligned_cols=49  Identities=16%  Similarity=0.172  Sum_probs=36.4

Q ss_pred             cceEEEecCCcc---ccccCCCcceeEeCCeeEEeCCcccCCCCeEEEEEccCCccCe
Q 045282           64 QRNVTLDCDGFQ---TVEKPDPVQMSISGFQCILLQGRDIIPFSRVHFKYAFDDEFPF  118 (130)
Q Consensus        64 ~~~V~l~C~gF~---S~~~VDP~ifr~~~~~CLVn~G~pi~~g~~v~F~YAw~~~f~~  118 (130)
                      .-||.+.=+||.   +-=||+|++-= -     =|.++.|+.|..|+|.|+-+.+-.+
T Consensus        35 ~ldv~v~~~gf~~GD~NYPI~Pkl~i-T-----Nns~~~iPGGt~~~FD~ptSa~~~~   86 (180)
T PF06483_consen   35 ALDVSVSFTGFKLGDSNYPINPKLTI-T-----NNSGQTIPGGTEFEFDYPTSAPDNA   86 (180)
T ss_pred             eEEEEEEeCCcccCCCCCCcCCcEEE-E-----cCCCcccCCccEEEEccccCCcccc
Confidence            447788888886   55789997532 1     2578899999999999997776543


No 12 
>PF06682 DUF1183:  Protein of unknown function (DUF1183);  InterPro: IPR009567 This family consists of several eukaryotic proteins of around 360 residues in length. The function of this family is unknown.
Probab=30.09  E-value=88  Score=26.75  Aligned_cols=25  Identities=20%  Similarity=0.478  Sum_probs=20.7

Q ss_pred             cccceEEEecCCccccccCCCcceeEe
Q 045282           62 CAQRNVTLDCDGFQTVEKPDPVQMSIS   88 (130)
Q Consensus        62 C~~~~V~l~C~gF~S~~~VDP~ifr~~   88 (130)
                      =....|.|.|.|+.+.+  ||=|||=+
T Consensus        91 ~klG~~~V~CEGY~~pd--DpyvLkGS  115 (318)
T PF06682_consen   91 YKLGSTDVSCEGYDYPD--DPYVLKGS  115 (318)
T ss_pred             eeecceEEeeecccCCC--CceecCCc
Confidence            44567899999999977  89999944


No 13 
>PF03293 Pox_RNA_pol:  Poxvirus DNA-directed RNA polymerase, 18 kD subunit;  InterPro: IPR004973 DNA-directed RNA polymerases 2.7.7.6 from EC (also known as DNA-dependent RNA polymerases) are responsible for the polymerisation of ribonucleotides into a sequence complementary to the template DNA. In eukaryotes, there are three different forms of DNA-directed RNA polymerases transcribing different sets of genes. Most RNA polymerases are multimeric enzymes and are composed of a variable number of subunits. The core RNA polymerase complex consists of five subunits (two alpha, one beta, one beta-prime and one omega) and is sufficient for transcription elongation and termination but is unable to initiate transcription. Transcription initiation from promoter elements requires a sixth, dissociable subunit called a sigma factor, which reversibly associates with the core RNA polymerase complex to form a holoenzyme []. The core RNA polymerase complex forms a "crab claw"-like structure with an internal channel running along the full length []. The key functional sites of the enzyme, as defined by mutational and cross-linking analysis, are located on the inner wall of this channel. RNA synthesis follows after the attachment of RNA polymerase to a specific site, the promoter, on the template DNA strand. The RNA synthesis process continues until a termination sequence is reached. The RNA product, which is synthesised in the 5' to 3'direction, is known as the primary transcript. Eukaryotic nuclei contain three distinct types of RNA polymerases that differ in the RNA they synthesise:  RNA polymerase I: located in the nucleoli, synthesises precursors of most ribosomal RNAs. RNA polymerase II: occurs in the nucleoplasm, synthesises mRNA precursors.  RNA polymerase III: also occurs in the nucleoplasm, synthesises the precursors of 5S ribosomal RNA, the tRNAs, and a variety of other small nuclear and cytosolic RNAs.   Eukaryotic cells are also known to contain separate mitochondrial and chloroplast RNA polymerases. Eukaryotic RNA polymerases, whose molecular masses vary in size from 500 to 700 kDa, contain two non-identical large (>100 kDa) subunits and an array of up to 12 different small (less than 50 kDa) subunits. The Poxvirus DNA-directed RNA polymerase (2.7.7.6 from EC) catalyses DNA-template-directed extension of the 3'-end of an RNA strand by one nucleotide at a time. The enzyme consists of at least eight subunits, this is the 18 kDa subunit.; GO: 0003677 DNA binding, 0003899 DNA-directed RNA polymerase activity, 0019083 viral transcription
Probab=28.75  E-value=1e+02  Score=23.84  Aligned_cols=48  Identities=19%  Similarity=0.217  Sum_probs=31.6

Q ss_pred             ccceEEEecCCccccccCCCcceeEeC-CeeEEeCCcccCCCCeEEEEE
Q 045282           63 AQRNVTLDCDGFQTVEKPDPVQMSISG-FQCILLQGRDIIPFSRVHFKY  110 (130)
Q Consensus        63 ~~~~V~l~C~gF~S~~~VDP~ifr~~~-~~CLVn~G~pi~~g~~v~F~Y  110 (130)
                      ..+||.+.|+..-=-..=|..-....+ .-|++.||..-..|+.|+-.-
T Consensus        93 dESni~V~CgDLiCkl~rdsGtVSf~dsKYCfirNg~vY~ngs~Vsv~L  141 (160)
T PF03293_consen   93 DESNITVQCGDLICKLSRDSGTVSFNDSKYCFIRNGVVYDNGSEVSVVL  141 (160)
T ss_pred             ccCceEEEcCcEEEEeeccCCeEEecCceEEEEECCEEecCCCEEEEEe
Confidence            468899999875432222333333322 349999999999999887654


No 14 
>PF14016 DUF4232:  Protein of unknown function (DUF4232)
Probab=27.48  E-value=1e+02  Score=21.96  Aligned_cols=73  Identities=11%  Similarity=0.145  Sum_probs=45.3

Q ss_pred             CCCCcceEEEeeecCccCCeeeEEEEEeecc--ccccceEEEecCCccccccC-------CCcceeEeCCeeEEeCCccc
Q 045282           30 CPPETLDISQKETGNIVQGKKEFAVEVFNWC--KCAQRNVTLDCDGFQTVEKP-------DPVQMSISGFQCILLQGRDI  100 (130)
Q Consensus        30 C~~~di~V~Q~~tg~~v~G~pe~~VtI~N~C--~C~~~~V~l~C~gF~S~~~V-------DP~ifr~~~~~CLVn~G~pi  100 (130)
                      |...|++++-..... ..|...+.|+++|.=  .|...       ||..+..+       .+..-+.. +   -..--.|
T Consensus         1 C~~~~L~~~~~~~~~-~~g~~~~~l~~tN~s~~~C~l~-------G~P~v~~~~~~g~~~~~~~~~~~-~---~~~~vtL   68 (131)
T PF14016_consen    1 CTAADLSVTVGPVDA-GAGQRHATLTFTNTSDTPCTLY-------GYPGVALVDADGAPLGVPAVREG-P---PPRPVTL   68 (131)
T ss_pred             CCcccEEEEEecccC-CCCccEEEEEEEECCCCcEEec-------cCCcEEEECCCCCcCCccccccC-C---CCCcEEE
Confidence            888899998866533 468889999999976  37664       34333333       23222222 1   1222346


Q ss_pred             CCCCeEEEEEccCC
Q 045282          101 IPFSRVHFKYAFDD  114 (130)
Q Consensus       101 ~~g~~v~F~YAw~~  114 (130)
                      .+|++..|.-.|..
T Consensus        69 ~PG~sA~a~l~~~~   82 (131)
T PF14016_consen   69 APGGSAYAGLRWSN   82 (131)
T ss_pred             CCCCEEEEEEEEec
Confidence            78888888777754


No 15 
>PF03032 Brevenin:  Brevenin/esculentin/gaegurin/rugosin family;  InterPro: IPR004275 In addition to the highly specific cell-mediated immune system, vertebrates possess an efficient host-defence mechanism against invading microorganisms which involves the synthesis of highly potent antimicrobial peptides with a large spectrum of activity. This entry represents a number of these defence peptides secreted from the skin of amphibians, including the opiate-like dermorphins and deltorphins, and the antimicrobial dermoseptins and temporins.; GO: 0006952 defense response, 0042742 defense response to bacterium, 0005576 extracellular region
Probab=27.44  E-value=47  Score=20.65  Aligned_cols=16  Identities=38%  Similarity=0.455  Sum_probs=10.9

Q ss_pred             hHHHHHHHHHHHHHhc
Q 045282            5 VQKILYATLFLALISE   20 (130)
Q Consensus         5 ~~k~l~~~lfl~lv~~   20 (130)
                      ..|-|.|++||-+|+=
T Consensus         3 lKKsllLlfflG~ISl   18 (46)
T PF03032_consen    3 LKKSLLLLFFLGTISL   18 (46)
T ss_pred             chHHHHHHHHHHHccc
Confidence            3566777788877763


No 16 
>PLN00115 pollen allergen group 3; Provisional
Probab=27.04  E-value=2.6e+02  Score=20.53  Aligned_cols=90  Identities=18%  Similarity=0.165  Sum_probs=53.3

Q ss_pred             ChhhhHHHHHHHHHHHHHhccCCCCCCCCCCCCcceEEEeeecCccCCeeeEEEEEeeccccccceEEEecCCccccccC
Q 045282            1 MAVLVQKILYATLFLALISETMSQPEQGPCPPETLDISQKETGNIVQGKKEFAVEVFNWCKCAQRNVTLDCDGFQTVEKP   80 (130)
Q Consensus         1 Ma~~~~k~l~~~lfl~lv~~g~~~~~~~~C~~~di~V~Q~~tg~~v~G~pe~~VtI~N~C~C~~~~V~l~C~gF~S~~~V   80 (130)
                      |++... +|++..+-.|..-|       .|.. +|.++=.. |    -.|.|-|-+.|.   .+..|.+.-.|  +.+-+
T Consensus         1 ~~~~~~-~~~~~~~a~l~~~~-------~~g~-~v~F~V~~-g----Snp~yL~ll~~~---dI~~V~Ik~~g--~~~W~   61 (118)
T PLN00115          1 MSSLSF-LLLAVALAALFAVG-------SCAT-EVTFKVGK-G----SSSTSLELVTNV---AISEVEIKEKG--AKDWV   61 (118)
T ss_pred             CchhHH-HHHHHHHHHHhhhh-------hcCC-ceEEEECC-C----CCcceEEEEEeC---CEEEEEEeecC--CCccc
Confidence            563333 55555555666666       5765 45444222 1    137888888865   57788887754  33334


Q ss_pred             CCcceeEe-CCeeEEeCCcccCCCCeEEEEEccC
Q 045282           81 DPVQMSIS-GFQCILLQGRDIIPFSRVHFKYAFD  113 (130)
Q Consensus        81 DP~ifr~~-~~~CLVn~G~pi~~g~~v~F~YAw~  113 (130)
                      ||  +++. |..=-++.++|+. | +++|+...+
T Consensus        62 ~~--M~rswGavW~~~s~~pl~-G-PlS~R~t~~   91 (118)
T PLN00115         62 DD--LKESSTNTWTLKSKAPLK-G-PFSVRFLVK   91 (118)
T ss_pred             Cc--cccCccceeEecCCCCCC-C-ceEEEEEEe
Confidence            44  4565 5555566678876 4 788887654


No 17 
>PF01345 DUF11:  Domain of unknown function DUF11;  InterPro: IPR001434 This group of sequences is represented by a conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis (Streptococcus faecalis), and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydia trachomatis outer membrane proteins.  In C. trachomatis, three cysteine-rich proteins (also believed to be lipoproteins), MOMP, OMP6 and OMP3, make up the extracellular matrix of the outer membrane []. They are involved in the essential structural integrity of both the elementary body (EB) and recticulate body (RB) phase. They are thought to be involved in porin formation and, as these bacteria lack the peptidoglycan layer common to most Gram-negative microbes, such proteins are highly important in the pathogenicity of the organism.; GO: 0005727 extrachromosomal circular DNA
Probab=25.54  E-value=1.6e+02  Score=18.71  Aligned_cols=29  Identities=14%  Similarity=0.062  Sum_probs=23.0

Q ss_pred             cCccCCeeeEEEEEeeccccccceEEEec
Q 045282           43 GNIVQGKKEFAVEVFNWCKCAQRNVTLDC   71 (130)
Q Consensus        43 g~~v~G~pe~~VtI~N~C~C~~~~V~l~C   71 (130)
                      .-.++..-+|.++|+|.=.-+..||.|.-
T Consensus        36 ~~~~Gd~v~ytitvtN~G~~~a~nv~v~D   64 (76)
T PF01345_consen   36 TANPGDTVTYTITVTNTGPAPATNVVVTD   64 (76)
T ss_pred             cccCCCEEEEEEEEEECCCCeeEeEEEEE
Confidence            33455677899999999988888888864


No 18 
>PF10731 Anophelin:  Thrombin inhibitor from mosquito;  InterPro: IPR018932  Members of this family are all inhibitors of thrombin, the peptidase that is at the end of the blood coagulation cascade and which creates the clot by cleaving fibrinogen. The interaction between thrombin and fibrinogen involves two different areas of contact - via the thrombin active site and via a second substrate-binding site known as an exosite. The inhibitor acts by blocking the exosite, rather than by interacting with the active site. The inhibitors are from mosquitoes that feed on human blood and which, by inhibiting thrombin, prevent the blood from clotting and keep it flowing. 
Probab=25.00  E-value=73  Score=21.28  Aligned_cols=21  Identities=24%  Similarity=0.184  Sum_probs=12.2

Q ss_pred             ChhhhHHHHHHHHHHHHHhcc
Q 045282            1 MAVLVQKILYATLFLALISET   21 (130)
Q Consensus         1 Ma~~~~k~l~~~lfl~lv~~g   21 (130)
                      ||++++-+.+|.+.|+.+.|+
T Consensus         1 MA~Kl~vialLC~aLva~vQ~   21 (65)
T PF10731_consen    1 MASKLIVIALLCVALVAIVQS   21 (65)
T ss_pred             CcchhhHHHHHHHHHHHHHhc
Confidence            886655554444555556666


No 19 
>KOG4063 consensus Major epididymal secretory protein HE1 [Function unknown]
Probab=23.94  E-value=1.5e+02  Score=23.04  Aligned_cols=41  Identities=12%  Similarity=0.101  Sum_probs=21.9

Q ss_pred             ChhhhHHHHHHHHHHHHHh-ccCCCCCCCCCCCCcceEEEeee
Q 045282            1 MAVLVQKILYATLFLALIS-ETMSQPEQGPCPPETLDISQKET   42 (130)
Q Consensus         1 Ma~~~~k~l~~~lfl~lv~-~g~~~~~~~~C~~~di~V~Q~~t   42 (130)
                      |+-++++.+++.++|.+-. |. -+-...+|.-+|-.+.+.+.
T Consensus         1 m~ms~~~~v~l~alls~a~aq~-~~t~~k~C~ss~g~~~~V~i   42 (158)
T KOG4063|consen    1 MMMSFLKTVILLALLSLAAAQA-ISTGVKQCGSSDGTPLEVKI   42 (158)
T ss_pred             CchHHHHHHHHHHHHHHhhhcc-cCcccccccCCCCcceEEEe
Confidence            5545566555444444433 11 01124469888877777665


No 20 
>PRK15249 fimbrial chaperone protein StbB; Provisional
Probab=22.08  E-value=1.3e+02  Score=24.40  Aligned_cols=90  Identities=14%  Similarity=-0.007  Sum_probs=46.9

Q ss_pred             hhHHHHHHHHHHHHHhccCCCCCCCCCCCCcceEEEeeecCccCCeeeEEEEEeeccccccceEEEecCCcc--------
Q 045282            4 LVQKILYATLFLALISETMSQPEQGPCPPETLDISQKETGNIVQGKKEFAVEVFNWCKCAQRNVTLDCDGFQ--------   75 (130)
Q Consensus         4 ~~~k~l~~~lfl~lv~~g~~~~~~~~C~~~di~V~Q~~tg~~v~G~pe~~VtI~N~C~C~~~~V~l~C~gF~--------   75 (130)
                      +.+++|.+++|+++-+.         +....|.|..++.- ..++..+=.|+|.|.=.= ..-|..+=+.-.        
T Consensus         8 ~~~~~~~~~~~~~~~~~---------~a~A~l~l~~TRvi-y~~~~~~~sl~l~N~~~~-p~LvQsWv~~~~~~~~p~~~   76 (253)
T PRK15249          8 SALYYLIVFLFLALPAT---------ASWASVTILGSRII-YPSTASSVDVQLKNNDAI-PYIVQTWFDDGDMNTSPENS   76 (253)
T ss_pred             hHHHHHHHHHHHHhhhH---------hheeEEEeCceEEE-EeCCCcceeEEEEcCCCC-cEEEEEEEeCCCCCCCcccc
Confidence            34666655444433222         23456888887763 445678888899886531 122221111111        


Q ss_pred             --ccccCCCcceeEeC-Cee---EEeCC-cccCCCC
Q 045282           76 --TVEKPDPVQMSISG-FQC---ILLQG-RDIIPFS  104 (130)
Q Consensus        76 --S~~~VDP~ifr~~~-~~C---LVn~G-~pi~~g~  104 (130)
                        ..-.|-|-+||... ..=   ++..| .+++...
T Consensus        77 ~~~pFivtPPlfrl~p~~~q~lRI~~~~~~~lP~DR  112 (253)
T PRK15249         77 SAMPFIATPPVFRIQPKAGQVVRVIYNNTKKLPQDR  112 (253)
T ss_pred             ccCcEEEcCCeEEecCCCceEEEEEEcCCCCCCCCc
Confidence              11347788999884 222   23344 3565543


No 21 
>PF01456 Mucin:  Mucin-like glycoprotein;  InterPro: IPR000458 This family of trypanosomal proteins resemble vertebrate mucins. The protein consists of three regions. The N and C terminii are conserved between all members of the family, whereas the central region is not well conserved and contains a large number of threonine residues which can be glycosylated []. Indirect evidence suggested that these genes might encode the core protein of parasite mucins, glycoproteins that were proposed to be involved in the interaction with, and invasion of, mammalian host cells.
Probab=21.22  E-value=83  Score=22.68  Aligned_cols=14  Identities=43%  Similarity=0.503  Sum_probs=11.4

Q ss_pred             hHHHHHHHHHHHHH
Q 045282            5 VQKILYATLFLALI   18 (130)
Q Consensus         5 ~~k~l~~~lfl~lv   18 (130)
                      ...||+.+|+|+|+
T Consensus         3 tcRLLCalLvlaLc   16 (143)
T PF01456_consen    3 TCRLLCALLVLALC   16 (143)
T ss_pred             hHHHHHHHHHHHHH
Confidence            58899988888874


No 22 
>PF05753 TRAP_beta:  Translocon-associated protein beta (TRAPB);  InterPro: IPR008856 This family consists of several eukaryotic translocon-associated protein beta (TRAPB) or signal sequence receptor beta subunit (SSR-beta) proteins. The normal translocation of nascent polypeptides into the lumen of the endoplasmic reticulum (ER) is thought to be aided in part by a translocon-associated protein (TRAP) complex consisting of 4 protein subunits. The association of mature proteins with the ER and Golgi, or other intracellular locales, such as lysosomes, depends on the initial targeting of the nascent polypeptide to the ER membrane. A similar scenario must also exist for proteins destined for secretion [].; GO: 0005783 endoplasmic reticulum, 0016021 integral to membrane
Probab=20.61  E-value=4.1e+02  Score=20.59  Aligned_cols=62  Identities=19%  Similarity=0.249  Sum_probs=41.5

Q ss_pred             cCccC-CeeeEEEEEeeccccccceEEEecCCccccccCCCcceeEeCCeeEEeCCcccCCCCeEEEEEc
Q 045282           43 GNIVQ-GKKEFAVEVFNWCKCAQRNVTLDCDGFQTVEKPDPVQMSISGFQCILLQGRDIIPFSRVHFKYA  111 (130)
Q Consensus        43 g~~v~-G~pe~~VtI~N~C~C~~~~V~l~C~gF~S~~~VDP~ifr~~~~~CLVn~G~pi~~g~~v~F~YA  111 (130)
                      .-++. -.-+.+++|.|.=.=+-.||+|.=++|.      |.-|...++. +=-.=..|++|+.++..|.
T Consensus        32 ~~~v~g~~v~V~~~iyN~G~~~A~dV~l~D~~fp------~~~F~lvsG~-~s~~~~~i~pg~~vsh~~v   94 (181)
T PF05753_consen   32 KYLVEGEDVTVTYTIYNVGSSAAYDVKLTDDSFP------PEDFELVSGS-LSASWERIPPGENVSHSYV   94 (181)
T ss_pred             ccccCCcEEEEEEEEEECCCCeEEEEEEECCCCC------ccccEeccCc-eEEEEEEECCCCeEEEEEE
Confidence            33443 3567999999999999999999987874      3455544321 1111137888888877775


No 23 
>TIGR01451 B_ant_repeat conserved repeat domain. This model represents the conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis, and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydial outer membrane proteins.
Probab=20.03  E-value=2.1e+02  Score=17.48  Aligned_cols=25  Identities=16%  Similarity=0.240  Sum_probs=18.7

Q ss_pred             cCCeeeEEEEEeeccccccceEEEe
Q 045282           46 VQGKKEFAVEVFNWCKCAQRNVTLD   70 (130)
Q Consensus        46 v~G~pe~~VtI~N~C~C~~~~V~l~   70 (130)
                      ++-.-+|+++|.|.-.=+..+|.|.
T Consensus        10 ~Gd~v~Yti~v~N~g~~~a~~v~v~   34 (53)
T TIGR01451        10 IGDTITYTITVTNNGNVPATNVVVT   34 (53)
T ss_pred             CCCEEEEEEEEEECCCCceEeEEEE
Confidence            4556789999999886666666664


Done!