Query         011601
Match_columns 481
No_of_seqs    111 out of 141
Neff          3.4 
Searched_HMMs 46136
Date          Fri Mar 29 03:10:49 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/011601.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/011601hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF11891 DUF3411:  Domain of un 100.0 4.2E-66 9.1E-71  481.5  12.4  169  245-415     1-179 (180)
  2 PF06524 NOA36:  NOA36 protein;  88.5    0.44 9.5E-06   48.8   3.6   11  103-113   219-229 (314)
  3 PF02084 Bindin:  Bindin;  Inte  79.4     1.8 3.9E-05   43.5   3.2   30  181-210   104-133 (238)
  4 PF14812 PBP1_TM:  Transmembran  79.2    0.61 1.3E-05   40.0   0.0   14  148-161    35-48  (81)
  5 PF02979 NHase_alpha:  Nitrile   71.6     2.6 5.7E-05   41.2   2.1   48  205-255    17-68  (188)
  6 PLN03138 Protein TOC75; Provis  69.0     8.2 0.00018   44.8   5.6    9  407-415   458-466 (796)
  7 KOG0943 Predicted ubiquitin-pr  66.8     1.6 3.5E-05   52.6  -0.5   71  177-265  1796-1868(3015)
  8 COG4907 Predicted membrane pro  63.4     5.1 0.00011   44.2   2.5   22  113-134   562-583 (595)
  9 PF02957 TT_ORF2:  TT viral ORF  58.8     7.4 0.00016   34.5   2.3    9   63-71     30-38  (122)
 10 PLN00151 potassium transporter  54.2      16 0.00034   42.8   4.5   23  348-370   220-242 (852)
 11 KOG3540 Beta amyloid precursor  54.1      15 0.00032   40.9   4.0   21  171-192   255-275 (615)
 12 PF11705 RNA_pol_3_Rpc31:  DNA-  53.4     8.1 0.00018   37.7   1.8    7   66-72     90-96  (233)
 13 TIGR02877 spore_yhbH sporulati  50.9      16 0.00034   39.1   3.6    7  119-125    60-66  (371)
 14 PF09849 DUF2076:  Uncharacteri  49.5      19 0.00041   36.4   3.7   10   67-76    156-165 (247)
 15 TIGR01323 nitrile_alph nitrile  39.2      34 0.00074   33.6   3.6   46  205-253    11-60  (185)
 16 PF14812 PBP1_TM:  Transmembran  39.1      10 0.00022   32.8   0.0   12  150-161    33-44  (81)
 17 PF10446 DUF2457:  Protein of u  39.1      17 0.00036   39.9   1.6   14  321-334   294-307 (458)
 18 PLN03138 Protein TOC75; Provis  38.6      22 0.00048   41.4   2.6   12  199-210   189-200 (796)
 19 PF02957 TT_ORF2:  TT viral ORF  38.1      17 0.00037   32.2   1.3   11  118-128    75-85  (122)
 20 PHA03249 DNA packaging tegumen  35.0 1.2E+02  0.0026   34.9   7.3   23  158-180   171-195 (653)
 21 KOG0943 Predicted ubiquitin-pr  34.9      21 0.00045   43.9   1.6    7  441-447  2027-2033(3015)
 22 PF09026 CENP-B_dimeris:  Centr  34.8      13 0.00028   33.3   0.0   18  185-202    51-68  (101)
 23 PRK05325 hypothetical protein;  33.6      58  0.0013   35.3   4.5    9  119-127    48-56  (401)
 24 KOG3074 Transcriptional regula  32.7      26 0.00057   35.8   1.8   19  198-216    91-109 (263)
 25 PF11705 RNA_pol_3_Rpc31:  DNA-  32.0      28 0.00061   34.0   1.8    8  152-159   213-220 (233)
 26 KOG0921 Dosage compensation co  31.6      59  0.0013   39.1   4.5    9   27-35   1116-1124(1282)
 27 COG1512 Beta-propeller domains  31.2      59  0.0013   33.4   4.0   16   93-108   220-235 (271)
 28 PF04931 DNA_pol_phi:  DNA poly  30.5      48  0.0011   37.9   3.6    6   69-74    559-564 (784)
 29 PHA00458 single-stranded DNA-b  29.3      57  0.0012   33.1   3.4    8   69-76     59-66  (233)
 30 PF04858 TH1:  TH1 protein;  In  29.2 1.1E+02  0.0024   34.7   5.9   50  185-235    36-95  (584)
 31 COG2818 Tag 3-methyladenine DN  27.2      37 0.00081   33.4   1.7  125  167-300    50-177 (188)
 32 KOG2023 Nuclear transport rece  26.9      35 0.00075   39.6   1.6   39   89-127   302-343 (885)
 33 PF08595 RXT2_N:  RXT2-like, N-  24.1      41 0.00089   31.8   1.3    6  222-227   104-109 (149)
 34 PLN03083 E3 UFM1-protein ligas  23.4 1.2E+02  0.0027   35.6   5.1   35  176-210   484-522 (803)
 35 COG3685 Uncharacterized protei  23.3      54  0.0012   31.8   2.0   53  173-227    57-119 (167)
 36 PF04931 DNA_pol_phi:  DNA poly  22.7      48   0.001   38.0   1.8   10   65-74    559-568 (784)
 37 KOG3973 Uncharacterized conser  21.6      69  0.0015   34.7   2.5    9  102-110   340-348 (465)
 38 PRK05325 hypothetical protein;  21.4      71  0.0015   34.6   2.6    8  280-287   202-209 (401)
 39 KOG3241 Uncharacterized conser  20.9      70  0.0015   31.9   2.2    7  134-140   182-188 (227)
 40 PF06946 Phage_holin_5:  Phage   20.9 2.4E+02  0.0053   25.2   5.3   27  339-376    57-83  (93)
 41 PF04871 Uso1_p115_C:  Uso1 / p  20.7      71  0.0015   29.4   2.1   21  141-161   116-136 (136)
 42 KOG2023 Nuclear transport rece  20.5      99  0.0021   36.2   3.6   34  222-260   445-480 (885)

No 1  
>PF11891 DUF3411:  Domain of unknown function (DUF3411);  InterPro: IPR021825  This presumed domain is functionally uncharacterised. This domain is found in eukaryotes. This domain is typically between 168 to 186 amino acids in length. This domain has a conserved RYQ sequence motif. 
Probab=100.00  E-value=4.2e-66  Score=481.55  Aligned_cols=169  Identities=41%  Similarity=0.610  Sum_probs=162.3

Q ss_pred             hhhccCchhHHHHHHHHHHhhhhhhhhhhhhcchhhhhHHHHHHHHHHHHHHhhHHHhhhcccccccCCccc----cchH
Q 011601          245 GRMLADPSFLYKLILEQAATIGCTVLWELENRKERIKQEWDLALINVLTVTACNAFVVWSLAPCRSYGNTFR----FDLQ  320 (481)
Q Consensus       245 ~RlLADP~FLfKL~iE~~I~i~~~~~aE~~~Rge~F~~ElDfV~sdvv~g~v~nfaLVwLLAPt~s~G~~~~----~~lq  320 (481)
                      |||||||+|||||++||+||++|+++|||++|||+||+|||||+||+++++|+||+||||||||++++++++    +.+|
T Consensus         1 ~RllADP~Fl~Kl~~E~~i~i~~~~~~e~~~R~e~f~~E~d~v~~d~v~~~i~n~~lv~llAPt~s~~~~~~~~~~~~~~   80 (180)
T PF11891_consen    1 ERLLADPSFLFKLAIEEVIGIGCATAAEYAKRGERFWNELDFVFSDVVVGSIVNFALVWLLAPTRSFGSPAASSPGGGLQ   80 (180)
T ss_pred             CcccccchHHHHHHHHHHHHHHHHHHHHHHHcccchHHHHHHHHHHHHHHHHHHHHHHHhccchHhhCcccccccchHHH
Confidence            799999999999999999999999999999999999999999999999999999999999999999999988    6899


Q ss_pred             HHhccCCCccccccCCCCCcchhhhHHHHHhhchhhhhhHhHHHhhHHHHHHHHhh-c-cC----CCcccCCchhhhhhh
Q 011601          321 NTLQKLPNNIFERSYPFREFDLQKRIHSLFYKAAELCMVGLSAGAVQGSLSNYLAG-K-KD----RLSVTIPSVSTSALE  394 (481)
Q Consensus       321 k~l~~lP~N~Feks~pg~~fsl~qRiga~v~KGa~l~~VG~~aGlvG~glSN~L~~-k-K~----~~s~~vP~l~tsAl~  394 (481)
                      +++++||+||||+++||++||++||++||+|||++|++|||+||++|+++||+|++ | +.    ++++++|||++||+ 
T Consensus        81 ~~~~~~P~n~Fq~~~~g~~fsl~qR~~~~~~kg~~l~~VG~~ag~vg~~lsn~L~~~rk~~~~~~e~~~~~ppv~~ta~-  159 (180)
T PF11891_consen   81 KFLGSLPNNAFQKGYPGRSFSLAQRIGAFVYKGAKLAAVGFIAGLVGTGLSNALIAARKKVDPSFEPSVPVPPVLKTAL-  159 (180)
T ss_pred             HHHHhChHHHhccCCCCCcccHHHHHHHHHHcchHhhhhHHHHHHHHHHHHHHHHHHHHhcCccccCCCCCCCHHHHHH-
Confidence            99999999999999999999999999999999999999999999999999999999 4 33    35778899999999 


Q ss_pred             hhhhhhhcccccchhhhhHHH
Q 011601          395 RSRLAWLGVEADPLLQSDDLL  415 (481)
Q Consensus       395 ~~w~~fmGvSAN~RyQ~~nll  415 (481)
                       +|++|||+|||+|||..|++
T Consensus       160 -~~g~fmGvSsNlRYQil~Gi  179 (180)
T PF11891_consen  160 -GWGAFMGVSSNLRYQILNGI  179 (180)
T ss_pred             -HHHHHHhhhHhHHHHHHcCC
Confidence             69999999999999998876


No 2  
>PF06524 NOA36:  NOA36 protein;  InterPro: IPR010531 This family consists of several NOA36 proteins which contain 29 highly conserved cysteine residues. The function of this protein is unknown.; GO: 0008270 zinc ion binding, 0005634 nucleus
Probab=88.52  E-value=0.44  Score=48.84  Aligned_cols=11  Identities=27%  Similarity=0.286  Sum_probs=4.6

Q ss_pred             eeeccccccee
Q 011601          103 TLEKGKLDTTQ  113 (481)
Q Consensus       103 tleks~l~~~q  113 (481)
                      |-|-.-|.+|-
T Consensus       219 t~eTkdLSmSt  229 (314)
T PF06524_consen  219 TQETKDLSMST  229 (314)
T ss_pred             ccccccceeee
Confidence            33444444443


No 3  
>PF02084 Bindin:  Bindin;  InterPro: IPR000775 Bindin, the major protein component of the acrosome granule of sea urchin sperm, mediates species-specific adhesion of sperm to the egg surface during fertilisation [, ]. The protein coats the acrosomal process after externalisation by the acrosome reaction; it binds to sulphated, fucose-containing polysaccharides on the vitelline-layer receptor proteoglycans that cover the egg plasma membrane. Bindins from different genera show high levels of sequence similarity in both the mature bindin domain and in the probindin precursor region. The most highly conserved region is a 42-residue segment in the central portion of the mature bindin protein. This domain may be responsible for conserved functions of bindin, while the more highly divergent flanking regions may be responsible for its species-specific properties [].; GO: 0007342 fusion of sperm to egg plasma membrane
Probab=79.39  E-value=1.8  Score=43.50  Aligned_cols=30  Identities=27%  Similarity=0.554  Sum_probs=22.1

Q ss_pred             HHHHHHHHHHHhhccCchHHHHHHHHcCCc
Q 011601          181 KFVDAVLNEWMKTMMDLPAGFRQAYEMGLV  210 (481)
Q Consensus       181 ~~i~aVL~E~~rt~~sLPaDl~~A~e~Glv  210 (481)
                      |-.+.+.+-.+.|--+||-||-+=++.||+
T Consensus       104 Kvm~~ikavLgaTKiDLPVDINDPYDlGLL  133 (238)
T PF02084_consen  104 KVMEDIKAVLGATKIDLPVDINDPYDLGLL  133 (238)
T ss_pred             HHHHHHHHHhcccccccccccCChhhHHHH
Confidence            333444444578999999999999998864


No 4  
>PF14812 PBP1_TM:  Transmembrane domain of transglycosylase PBP1 at N-terminal; PDB: 3FWL_A 3VMA_A.
Probab=79.22  E-value=0.61  Score=40.04  Aligned_cols=14  Identities=71%  Similarity=0.902  Sum_probs=0.0

Q ss_pred             CCCCCCCCCCCCCC
Q 011601          148 DDDDYFDDFDDGDE  161 (481)
Q Consensus       148 ddddy~d~~ddgd~  161 (481)
                      ++|||+||++++||
T Consensus        35 ~ddd~~DDD~dDde   48 (81)
T PF14812_consen   35 YDDDYEDDDDDDDE   48 (81)
T ss_dssp             --------------
T ss_pred             cccccccccccchh
Confidence            34444444444333


No 5  
>PF02979 NHase_alpha:  Nitrile hydratase, alpha chain;  InterPro: IPR004232 Nitrile hydratases (4.2.1.84 from EC) are bacterial enzymes that catalyse the hydration of nitrile compounds to the corresponding amides. They are used as biocatalysts in acrylamide production, one of the few commercial scale bioprocesses, as well as in environmental remediation for the removal of nitriles from waste streams. Nitrile hydratases are composed of two subunits, alpha and beta, and are normally active as a tetramer, alpha(2)beta(2). Nitrile hydratases contain either a non-haem iron or a non-corrinoid cobalt centre, both types sharing a highly conserved peptide sequence in the alpha subunit (CXLCSC) that provides all the residues involved in coordinating the metal ion. Each type of nitrile hydratase specifically incorporated its metal with the help of activator proteins encoded by flanking regions of the nitrile hydratase genes that are necessary for metal insertion. The Fe-containing enzyme is photo-regulated: in the dark the enzyme is inactivated due to the association of nitric oxide (NO) to the iron, while in the light the enzyme is active by photo-dissociation of NO. The NO is held in place by a claw setting formed through specific oxygen atoms in two modified cysteines and a serine residue in the active site [, ]. The cobalt-containing enzyme is unaffected by NO, but was shown to undergo a similar effect with carbon monoxide [, ]. Fe- and cobalt-containing enzymes also display different inhibition patterns with nitrophenols. Thiocyanate hydrolase (SCNase) is a cobalt-containing metalloenzyme with a cysteine-sulphinic acid ligand that hydrolyses thiocyanate to carbonyl sulphide and ammonia []. The two enzymes, nitrile hydratase and SCNase, are homologous over regions corresponding to almost the entire coding regions of the genes: the beta and alpha subunits of thiocyanate hydrolase were homologous to the amino- and carboxyl-terminal halves of the beta subunit of nitrile hydratase, and the gamma subunit of thiocyanate hydrolase was homologous to the alpha subunit of nitrile hydratase [].  This entry represents the structural domain of the alpha subunit of both iron- and cobalt-containing nitrile hydratases; the alpha subunit is a duplication of two structural repeats, each consisting of 4 layers, alpha/beta/beta/alpha []. This structure is also found in the related protein, the gamma subunit of thiocyanate hydrolase (SCNase).; GO: 0003824 catalytic activity, 0046914 transition metal ion binding, 0006807 nitrogen compound metabolic process; PDB: 2DPP_A 3HHT_A 1V29_A 2ZZD_I 2DXC_F 2DXB_F 2DD5_C 2DD4_C 2ZPH_A 2CYZ_A ....
Probab=71.63  E-value=2.6  Score=41.18  Aligned_cols=48  Identities=27%  Similarity=0.479  Sum_probs=35.5

Q ss_pred             HHcCCcCHHHHHHHHHhccCc----chhHHHHhhcCcchhhhhhhhhccCchhHH
Q 011601          205 YEMGLVSSAQMVKFLAINARP----TTTRFISRSLPQGISRAFIGRMLADPSFLY  255 (481)
Q Consensus       205 ~e~Glvssa~L~rFl~L~~~P----~~~~~l~R~lp~~~srgfR~RlLADP~FLf  255 (481)
                      +|.|+|+++.+.++.+...+-    +-++.+.|.+-.   .+||.|||+||.=..
T Consensus        17 ~ekg~~~~~~~~~~~~~~~~~~~P~~GarvVArAW~D---p~FK~rLLaD~~aA~   68 (188)
T PF02979_consen   17 IEKGLITPAEVDRIIETYESRVGPRNGARVVARAWTD---PAFKARLLADPTAAI   68 (188)
T ss_dssp             HHTTSS-HHHHHHHHHHHHHTSSHHHHHHHHHHHHH----HHHHHHHHHSHHHHH
T ss_pred             HHcCCCCHHHHHHHHHHHHhccCccccceeehhhhCC---HHHHHHHHHCHHHHH
Confidence            679999999999988876543    236777887644   699999999997433


No 6  
>PLN03138 Protein TOC75; Provisional
Probab=69.01  E-value=8.2  Score=44.79  Aligned_cols=9  Identities=0%  Similarity=0.102  Sum_probs=4.5

Q ss_pred             chhhhhHHH
Q 011601          407 PLLQSDDLL  415 (481)
Q Consensus       407 ~RyQ~~nll  415 (481)
                      ..||-.||+
T Consensus       458 vs~~~~NL~  466 (796)
T PLN03138        458 VSFEHRNIQ  466 (796)
T ss_pred             EEEeccccc
Confidence            445555554


No 7  
>KOG0943 consensus Predicted ubiquitin-protein ligase/hyperplastic discs protein, HECT superfamily [Posttranslational modification, protein turnover, chaperones]
Probab=66.79  E-value=1.6  Score=52.60  Aligned_cols=71  Identities=27%  Similarity=0.152  Sum_probs=40.0

Q ss_pred             hhcHHHHHHHHHHHHhhccCchHHHHHHHHcCCcCHHHHHHHHHhccCcchhHHHHhhcCcchhhhhhh--hhccCchhH
Q 011601          177 LFDRKFVDAVLNEWMKTMMDLPAGFRQAYEMGLVSSAQMVKFLAINARPTTTRFISRSLPQGISRAFIG--RMLADPSFL  254 (481)
Q Consensus       177 ~fdR~~i~aVL~E~~rt~~sLPaDl~~A~e~Glvssa~L~rFl~L~~~P~~~~~l~R~lp~~~srgfR~--RlLADP~FL  254 (481)
                      .=||-+|.++..|.++--+.+|+-+.+|--.-.-+.      ++-+.+        -+  |  +.+.++  =|+|||.-+
T Consensus      1796 ~~dRieVq~ataeSE~aaepV~~afe~adpqp~dpd------~dDnSS--------ds--Q--e~~a~e~aFmiad~ahp 1857 (3015)
T KOG0943|consen 1796 SGDRIEVQAATAESEAAAEPVPAAFEEADPQPNDPD------DDDNSS--------DS--Q--EDDAEEEAFMIADPAHP 1857 (3015)
T ss_pred             ccceeEEeecccccccccCCcccccCcCCCCCCCCC------cccccc--------cc--c--cccccccceeecCCcch
Confidence            346677778888888888888876655532211111      000000        00  0  011111  189999999


Q ss_pred             HHHHHHHHHhh
Q 011601          255 YKLILEQAATI  265 (481)
Q Consensus       255 fKL~iE~~I~i  265 (481)
                      ..|+.|+-.-+
T Consensus      1858 leva~aqpaN~ 1868 (3015)
T KOG0943|consen 1858 LEVALAQPANS 1868 (3015)
T ss_pred             hhHhhcccCCC
Confidence            99999986544


No 8  
>COG4907 Predicted membrane protein [Function unknown]
Probab=63.45  E-value=5.1  Score=44.17  Aligned_cols=22  Identities=18%  Similarity=0.222  Sum_probs=11.2

Q ss_pred             ecccccCceeecCCCCCCCCCC
Q 011601          113 QQQSETTPELATGGGGGDIGKK  134 (481)
Q Consensus       113 q~~~~~~p~l~~GgGgG~~g~~  134 (481)
                      +.-+...|.-.+|-+||++|+-
T Consensus       562 raysa~a~S~~~~~~GGG~G~~  583 (595)
T COG4907         562 RAYSAIASSRRSSSSGGGGGFS  583 (595)
T ss_pred             hhhhcccccccCCCCCCCCCcC
Confidence            3345555655555555544443


No 9  
>PF02957 TT_ORF2:  TT viral ORF2;  InterPro: IPR004118 This entry represents the Gyroviral VP2 protein and TT viral ORF2.  Torque teno virus (TTV) is a nonenveloped and single-stranded DNA virus that was initially isolated from a Japanese patient with hepatitis of unknown aetiology, and which has since been found to infect both healthy and diseased individuals []. Numerous prevalence studies have raised questions about its role in unexplained hepatitis. ORF2 is a 150 residue protein of unknown function.  Gyroviruses are small circular single stranded viruses, such as the Chicken anaemia virus. The VP2 protein contains a set of conserved cysteine and histidine residues suggesting a zinc binding domain. VP2 may act as a scaffold protein in virion assembly and may also play a role in intracellular signaling during viral replication.
Probab=58.75  E-value=7.4  Score=34.46  Aligned_cols=9  Identities=11%  Similarity=0.106  Sum_probs=4.4

Q ss_pred             cchhhhhhh
Q 011601           63 DHSVVLLER   71 (481)
Q Consensus        63 ~~~~~~~er   71 (481)
                      ++-+.+|-+
T Consensus        30 ~dp~~HL~~   38 (122)
T PF02957_consen   30 GDPIAHLLA   38 (122)
T ss_pred             CCHHHHHHH
Confidence            355555444


No 10 
>PLN00151 potassium transporter; Provisional
Probab=54.23  E-value=16  Score=42.81  Aligned_cols=23  Identities=13%  Similarity=-0.096  Sum_probs=9.6

Q ss_pred             HHHhhchhhhhhHhHHHhhHHHH
Q 011601          348 SLFYKAAELCMVGLSAGAVQGSL  370 (481)
Q Consensus       348 a~v~KGa~l~~VG~~aGlvG~gl  370 (481)
                      .++-|-..+-.+=++.+++|+++
T Consensus       220 ~~lE~s~~~k~~ll~l~l~Gtam  242 (852)
T PLN00151        220 ERLETSSLLKKLLLLLVLAGTSM  242 (852)
T ss_pred             HHhhhhHHHHHHHHHHHHHhHHH
Confidence            33333333333334445555553


No 11 
>KOG3540 consensus Beta amyloid precursor protein [General function prediction only]
Probab=54.05  E-value=15  Score=40.89  Aligned_cols=21  Identities=43%  Similarity=0.664  Sum_probs=15.0

Q ss_pred             hhHHHHhhcHHHHHHHHHHHHh
Q 011601          171 RKILEELFDRKFVDAVLNEWMK  192 (481)
Q Consensus       171 r~~~~e~fdR~~i~aVL~E~~r  192 (481)
                      +.-++| =-|+-++.||+||+.
T Consensus       255 kmrlee-khr~rmd~VmkEW~~  275 (615)
T KOG3540|consen  255 KMRLEE-KHRKRMDKVMKEWEE  275 (615)
T ss_pred             HHHHHH-HHHHHHHHHHHHHHH
Confidence            333443 357889999999973


No 12 
>PF11705 RNA_pol_3_Rpc31:  DNA-directed RNA polymerase III subunit Rpc31;  InterPro: IPR024661 DNA-directed RNA polymerases 2.7.7.6 from EC (also known as DNA-dependent RNA polymerases) are responsible for the polymerisation of ribonucleotides into a sequence complementary to the template DNA. In eukaryotes, there are three different forms of DNA-directed RNA polymerases transcribing different sets of genes. Most RNA polymerases are multimeric enzymes and are composed of a variable number of subunits. The core RNA polymerase complex consists of five subunits (two alpha, one beta, one beta-prime and one omega) and is sufficient for transcription elongation and termination but is unable to initiate transcription. Transcription initiation from promoter elements requires a sixth, dissociable subunit called a sigma factor, which reversibly associates with the core RNA polymerase complex to form a holoenzyme []. The core RNA polymerase complex forms a "crab claw"-like structure with an internal channel running along the full length []. The key functional sites of the enzyme, as defined by mutational and cross-linking analysis, are located on the inner wall of this channel. RNA synthesis follows after the attachment of RNA polymerase to a specific site, the promoter, on the template DNA strand. The RNA synthesis process continues until a termination sequence is reached. The RNA product, which is synthesised in the 5' to 3'direction, is known as the primary transcript. Eukaryotic nuclei contain three distinct types of RNA polymerases that differ in the RNA they synthesise:  RNA polymerase I: located in the nucleoli, synthesises precursors of most ribosomal RNAs. RNA polymerase II: occurs in the nucleoplasm, synthesises mRNA precursors.  RNA polymerase III: also occurs in the nucleoplasm, synthesises the precursors of 5S ribosomal RNA, the tRNAs, and a variety of other small nuclear and cytosolic RNAs.   Eukaryotic cells are also known to contain separate mitochondrial and chloroplast RNA polymerases. Eukaryotic RNA polymerases, whose molecular masses vary in size from 500 to 700 kDa, contain two non-identical large (>100 kDa) subunits and an array of up to 12 different small (less than 50 kDa) subunits. RNA polymerase III contains seventeen subunits in yeasts and in human cells. Twelve of these are akin to RNA polymerase I or II and the other five are RNA polymerase III-specific, and form the functionally distinct groups: (i) Rpc31-Rpc34-Rpc82, and (ii) Rpc37-Rpc53. Rpc31, Rpc34 and Rpc82 form a cluster of enzyme-specific subunits that contribute to transcription initiation in Saccharomyces cerevisiae and Homo sapiens. There is evidence that these subunits are anchored at or near the N-terminal Zn-fold of Rpc1, itself prolonged by a highly conserved but RNA polymerase III-specific domain []. This entry represents the Rpc31 subunit.
Probab=53.39  E-value=8.1  Score=37.74  Aligned_cols=7  Identities=29%  Similarity=0.183  Sum_probs=5.0

Q ss_pred             hhhhhhh
Q 011601           66 VVLLERC   72 (481)
Q Consensus        66 ~~~~er~   72 (481)
                      ...+||+
T Consensus        90 ~~~ierY   96 (233)
T PF11705_consen   90 FDDIERY   96 (233)
T ss_pred             cccHHHH
Confidence            5568887


No 13 
>TIGR02877 spore_yhbH sporulation protein YhbH. This protein family, typified by YhbH in Bacillus subtilis, is found in nearly every endospore-forming bacterium and in no other genome (but note that the trusted cutoff score is set high to exclude a single high-scoring sequence from Nitrosococcus oceani ATCC 19707, which is classified in the Gammaproteobacteria). The gene in Bacillus subtilis was shown to be in the regulon of the sporulation sigma factor, sigma-E, and its mutation was shown to create a sporulation defect.
Probab=50.93  E-value=16  Score=39.14  Aligned_cols=7  Identities=0%  Similarity=-0.149  Sum_probs=3.6

Q ss_pred             CceeecC
Q 011601          119 TPELATG  125 (481)
Q Consensus       119 ~p~l~~G  125 (481)
                      +|....|
T Consensus        60 Ep~F~~g   66 (371)
T TIGR02877        60 EYRFRYD   66 (371)
T ss_pred             cceEEeC
Confidence            5555544


No 14 
>PF09849 DUF2076:  Uncharacterized protein conserved in bacteria (DUF2076);  InterPro: IPR018648  This family of hypothetical prokaryotic proteins has no known function but includes putative perimplasmic ligand-binding sensor proteins.
Probab=49.49  E-value=19  Score=36.41  Aligned_cols=10  Identities=10%  Similarity=0.052  Sum_probs=4.8

Q ss_pred             hhhhhhhcCC
Q 011601           67 VLLERCYKAP   76 (481)
Q Consensus        67 ~~~er~~~~~   76 (481)
                      .+||=-|...
T Consensus       156 n~i~~lF~~~  165 (247)
T PF09849_consen  156 NGIESLFGGH  165 (247)
T ss_pred             HHHHHHhcCC
Confidence            3455555443


No 15 
>TIGR01323 nitrile_alph nitrile hydratase, alpha subunit. This model describes both iron- and cobalt-containing nitrile hydratase alpha chains. It excludes the thiocyanate hydrolase gamma subunit of Thiobacillus thioparus, a sequence that appears to have evolved from within the family of nitrile hydratase alpha subunits but which differs by several indels and a more rapid accumulation of point mutations.
Probab=39.22  E-value=34  Score=33.60  Aligned_cols=46  Identities=13%  Similarity=0.323  Sum_probs=35.2

Q ss_pred             HHcCCcCHHHHHHHHHhccC---c-chhHHHHhhcCcchhhhhhhhhccCchh
Q 011601          205 YEMGLVSSAQMVKFLAINAR---P-TTTRFISRSLPQGISRAFIGRMLADPSF  253 (481)
Q Consensus       205 ~e~Glvssa~L~rFl~L~~~---P-~~~~~l~R~lp~~~srgfR~RlLADP~F  253 (481)
                      +|.|+|+++.+.+.++...+   | +-++.+.|.+-.   ..||.|||+|..=
T Consensus        11 ~eKGli~~~~id~~i~~~~~~~gP~nGA~vVArAW~D---p~fk~~Ll~d~~a   60 (185)
T TIGR01323        11 KSKGLIPEGAVDQLTSLYENEWGPENGAKVVAKAWVD---PEFRALLLKDATA   60 (185)
T ss_pred             HHcCCCCHHHHHHHHHHHHhccCCcchhhhhhHHhcC---HHHHHHHHhChHH
Confidence            67899999998888876543   2 236778888654   5999999999753


No 16 
>PF14812 PBP1_TM:  Transmembrane domain of transglycosylase PBP1 at N-terminal; PDB: 3FWL_A 3VMA_A.
Probab=39.15  E-value=10  Score=32.79  Aligned_cols=12  Identities=58%  Similarity=1.168  Sum_probs=0.0

Q ss_pred             CCCCCCCCCCCC
Q 011601          150 DDYFDDFDDGDE  161 (481)
Q Consensus       150 ddy~d~~ddgd~  161 (481)
                      |||+||.+|+|+
T Consensus        33 DD~ddd~~DDD~   44 (81)
T PF14812_consen   33 DDYDDDYEDDDD   44 (81)
T ss_dssp             ------------
T ss_pred             hccccccccccc
Confidence            344444334333


No 17 
>PF10446 DUF2457:  Protein of unknown function (DUF2457);  InterPro: IPR018853  This entry represents a family of uncharacterised proteins. 
Probab=39.06  E-value=17  Score=39.88  Aligned_cols=14  Identities=14%  Similarity=0.370  Sum_probs=7.3

Q ss_pred             HHhccCCCcccccc
Q 011601          321 NTLQKLPNNIFERS  334 (481)
Q Consensus       321 k~l~~lP~N~Feks  334 (481)
                      +.++.=|-..|...
T Consensus       294 k~~g~SPrrlf~sp  307 (458)
T PF10446_consen  294 KLRGRSPRRLFRSP  307 (458)
T ss_pred             hccCCCCcccccCC
Confidence            34444466666543


No 18 
>PLN03138 Protein TOC75; Provisional
Probab=38.56  E-value=22  Score=41.41  Aligned_cols=12  Identities=8%  Similarity=0.185  Sum_probs=4.8

Q ss_pred             HHHHHHHHcCCc
Q 011601          199 AGFRQAYEMGLV  210 (481)
Q Consensus       199 aDl~~A~e~Glv  210 (481)
                      +|+..-.+.|..
T Consensus       189 ~dv~~I~~tG~F  200 (796)
T PLN03138        189 KELETLASCGMF  200 (796)
T ss_pred             HHHHHHHhcCCc
Confidence            333333344444


No 19 
>PF02957 TT_ORF2:  TT viral ORF2;  InterPro: IPR004118 This entry represents the Gyroviral VP2 protein and TT viral ORF2.  Torque teno virus (TTV) is a nonenveloped and single-stranded DNA virus that was initially isolated from a Japanese patient with hepatitis of unknown aetiology, and which has since been found to infect both healthy and diseased individuals []. Numerous prevalence studies have raised questions about its role in unexplained hepatitis. ORF2 is a 150 residue protein of unknown function.  Gyroviruses are small circular single stranded viruses, such as the Chicken anaemia virus. The VP2 protein contains a set of conserved cysteine and histidine residues suggesting a zinc binding domain. VP2 may act as a scaffold protein in virion assembly and may also play a role in intracellular signaling during viral replication.
Probab=38.08  E-value=17  Score=32.22  Aligned_cols=11  Identities=36%  Similarity=0.622  Sum_probs=4.8

Q ss_pred             cCceeecCCCC
Q 011601          118 TTPELATGGGG  128 (481)
Q Consensus       118 ~~p~l~~GgGg  128 (481)
                      ..|-+.+||+.
T Consensus        75 ~~~~~~~gg~~   85 (122)
T PF02957_consen   75 RQPCPGGGGGE   85 (122)
T ss_pred             cccCCCCCCCC
Confidence            34455444433


No 20 
>PHA03249 DNA packaging tegument protein UL25; Provisional
Probab=35.03  E-value=1.2e+02  Score=34.87  Aligned_cols=23  Identities=30%  Similarity=0.398  Sum_probs=11.2

Q ss_pred             CCCCCCCCcchhh--hhHHHHhhcH
Q 011601          158 DGDEGDEGGLFRR--RKILEELFDR  180 (481)
Q Consensus       158 dgd~gd~~g~~r~--r~~~~e~fdR  180 (481)
                      +-+||+|+.|--+  ..+-+|||.|
T Consensus       171 ~~~~~~e~~~~~~~~~~f~~~~~~~  195 (653)
T PHA03249        171 DLAEGHEFSFCDSDIEDFEQECFER  195 (653)
T ss_pred             CcccccccccccccHHHHHHHHHHh
Confidence            3345566554433  3344556544


No 21 
>KOG0943 consensus Predicted ubiquitin-protein ligase/hyperplastic discs protein, HECT superfamily [Posttranslational modification, protein turnover, chaperones]
Probab=34.88  E-value=21  Score=43.94  Aligned_cols=7  Identities=29%  Similarity=0.444  Sum_probs=3.8

Q ss_pred             hhHhhhh
Q 011601          441 NAIVSGL  447 (481)
Q Consensus       441 ~~~~sgl  447 (481)
                      ++|++|+
T Consensus      2027 ~lmragc 2033 (3015)
T KOG0943|consen 2027 NLMRAGC 2033 (3015)
T ss_pred             HHHHHHH
Confidence            4555555


No 22 
>PF09026 CENP-B_dimeris:  Centromere protein B dimerisation domain;  InterPro: IPR015115 Centromere protein B (CENP-B) interacts with centromeric heterochromatin in chromosomes and binds to a specific subset of alphoid satellite DNA, called the CENP-B box. CENP-B may organise arrays of centromere satellite DNA into a higher order structure, which then directs centromere formation and kinetochore assembly in mammalian chromosomes. The CENP-B dimerisation domain is composed of two alpha-helices, which are folded into an antiparallel configuration. Dimerisation of CENP-B is mediated by this domain, in which monomers dimerise to form a symmetrical, antiparallel, four-helix bundle structure with a large hydrophobic patch in which 23 residues of one monomer form van der Waals contacts with the other monomer. This CENP-B dimer configuration may be suitable for capturing two distant CENP-B boxes during centromeric heterochromatin formation []. ; GO: 0003677 DNA binding, 0003682 chromatin binding, 0006355 regulation of transcription, DNA-dependent, 0000775 chromosome, centromeric region, 0005634 nucleus; PDB: 1UFI_A.
Probab=34.82  E-value=13  Score=33.34  Aligned_cols=18  Identities=11%  Similarity=0.157  Sum_probs=8.8

Q ss_pred             HHHHHHHhhccCchHHHH
Q 011601          185 AVLNEWMKTMMDLPAGFR  202 (481)
Q Consensus       185 aVL~E~~rt~~sLPaDl~  202 (481)
                      +-|..-.|-+-|+|-|=+
T Consensus        51 ~~~~~v~rYltSf~id~~   68 (101)
T PF09026_consen   51 AYFTMVKRYLTSFPIDDK   68 (101)
T ss_dssp             HHHHHHHHHHCTS---HH
T ss_pred             hhcchHhhhhhccchhHh
Confidence            445555566777776544


No 23 
>PRK05325 hypothetical protein; Provisional
Probab=33.57  E-value=58  Score=35.29  Aligned_cols=9  Identities=33%  Similarity=0.756  Sum_probs=5.6

Q ss_pred             CceeecCCC
Q 011601          119 TPELATGGG  127 (481)
Q Consensus       119 ~p~l~~GgG  127 (481)
                      +|....|-+
T Consensus        48 Ep~F~~g~~   56 (401)
T PRK05325         48 EPKFRYGRG   56 (401)
T ss_pred             cceEEeCCC
Confidence            677766544


No 24 
>KOG3074 consensus Transcriptional regulator of the PUR family, single-stranded-DNA-binding [Transcription]
Probab=32.70  E-value=26  Score=35.79  Aligned_cols=19  Identities=16%  Similarity=0.198  Sum_probs=8.3

Q ss_pred             hHHHHHHHHcCCcCHHHHH
Q 011601          198 PAGFRQAYEMGLVSSAQMV  216 (481)
Q Consensus       198 PaDl~~A~e~Glvssa~L~  216 (481)
                      |.+++...|...+-++.|+
T Consensus        91 ~~~~~~~~ee~~lkSe~L~  109 (263)
T KOG3074|consen   91 PPELAAPSEEHELKSEELQ  109 (263)
T ss_pred             CcccccchhhhhHHHHHHh
Confidence            3444444444433344443


No 25 
>PF11705 RNA_pol_3_Rpc31:  DNA-directed RNA polymerase III subunit Rpc31;  InterPro: IPR024661 DNA-directed RNA polymerases 2.7.7.6 from EC (also known as DNA-dependent RNA polymerases) are responsible for the polymerisation of ribonucleotides into a sequence complementary to the template DNA. In eukaryotes, there are three different forms of DNA-directed RNA polymerases transcribing different sets of genes. Most RNA polymerases are multimeric enzymes and are composed of a variable number of subunits. The core RNA polymerase complex consists of five subunits (two alpha, one beta, one beta-prime and one omega) and is sufficient for transcription elongation and termination but is unable to initiate transcription. Transcription initiation from promoter elements requires a sixth, dissociable subunit called a sigma factor, which reversibly associates with the core RNA polymerase complex to form a holoenzyme []. The core RNA polymerase complex forms a "crab claw"-like structure with an internal channel running along the full length []. The key functional sites of the enzyme, as defined by mutational and cross-linking analysis, are located on the inner wall of this channel. RNA synthesis follows after the attachment of RNA polymerase to a specific site, the promoter, on the template DNA strand. The RNA synthesis process continues until a termination sequence is reached. The RNA product, which is synthesised in the 5' to 3'direction, is known as the primary transcript. Eukaryotic nuclei contain three distinct types of RNA polymerases that differ in the RNA they synthesise:  RNA polymerase I: located in the nucleoli, synthesises precursors of most ribosomal RNAs. RNA polymerase II: occurs in the nucleoplasm, synthesises mRNA precursors.  RNA polymerase III: also occurs in the nucleoplasm, synthesises the precursors of 5S ribosomal RNA, the tRNAs, and a variety of other small nuclear and cytosolic RNAs.   Eukaryotic cells are also known to contain separate mitochondrial and chloroplast RNA polymerases. Eukaryotic RNA polymerases, whose molecular masses vary in size from 500 to 700 kDa, contain two non-identical large (>100 kDa) subunits and an array of up to 12 different small (less than 50 kDa) subunits. RNA polymerase III contains seventeen subunits in yeasts and in human cells. Twelve of these are akin to RNA polymerase I or II and the other five are RNA polymerase III-specific, and form the functionally distinct groups: (i) Rpc31-Rpc34-Rpc82, and (ii) Rpc37-Rpc53. Rpc31, Rpc34 and Rpc82 form a cluster of enzyme-specific subunits that contribute to transcription initiation in Saccharomyces cerevisiae and Homo sapiens. There is evidence that these subunits are anchored at or near the N-terminal Zn-fold of Rpc1, itself prolonged by a highly conserved but RNA polymerase III-specific domain []. This entry represents the Rpc31 subunit.
Probab=32.03  E-value=28  Score=34.04  Aligned_cols=8  Identities=50%  Similarity=0.999  Sum_probs=4.4

Q ss_pred             CCCCCCCC
Q 011601          152 YFDDFDDG  159 (481)
Q Consensus       152 y~d~~ddg  159 (481)
                      |||++||+
T Consensus       213 YFDnGedD  220 (233)
T PF11705_consen  213 YFDNGEDD  220 (233)
T ss_pred             cCCCCCcc
Confidence            36665543


No 26 
>KOG0921 consensus Dosage compensation complex, subunit MLE [Transcription]
Probab=31.55  E-value=59  Score=39.12  Aligned_cols=9  Identities=0%  Similarity=-0.005  Sum_probs=4.3

Q ss_pred             CCCCCCCCC
Q 011601           27 TLAAPRYST   35 (481)
Q Consensus        27 ~~~~~~~~~   35 (481)
                      |.-|-+|-+
T Consensus      1116 PaiIsqLdp 1124 (1282)
T KOG0921|consen 1116 PAIISQLDP 1124 (1282)
T ss_pred             hhHhhccCc
Confidence            334555554


No 27 
>COG1512 Beta-propeller domains of methanol dehydrogenase type [General function prediction only]
Probab=31.22  E-value=59  Score=33.40  Aligned_cols=16  Identities=19%  Similarity=0.167  Sum_probs=8.9

Q ss_pred             CcccCcccceeeeccc
Q 011601           93 GGQYGAFGAVTLEKGK  108 (481)
Q Consensus        93 g~~~g~~ga~tleks~  108 (481)
                      +.++..|+..++++..
T Consensus       220 ~~~~~~~~~~~~~~~~  235 (271)
T COG1512         220 GRQPDRWLNGVLGRRR  235 (271)
T ss_pred             ccccccccceeEEeee
Confidence            3355666666665554


No 28 
>PF04931 DNA_pol_phi:  DNA polymerase phi;  InterPro: IPR007015 Proteins of this family are predominantly nucleolar. The majority are described as transcription factor transactivators. The family also includes the fifth essential DNA polymerase (Pol5p) of Schizosaccharomyces pombe (Fission yeast) and Saccharomyces cerevisiae (Baker's yeast) (2.7.7.7 from EC). Pol5p is localized exclusively to the nucleolus and binds near or at the enhancer region of rRNA-encoding DNA repeating units.; GO: 0003677 DNA binding, 0003887 DNA-directed DNA polymerase activity, 0006351 transcription, DNA-dependent
Probab=30.49  E-value=48  Score=37.95  Aligned_cols=6  Identities=50%  Similarity=1.320  Sum_probs=2.7

Q ss_pred             hhhhhc
Q 011601           69 LERCYK   74 (481)
Q Consensus        69 ~er~~~   74 (481)
                      |.-|+.
T Consensus       559 l~~c~~  564 (784)
T PF04931_consen  559 LQICYE  564 (784)
T ss_pred             HHHHHH
Confidence            444543


No 29 
>PHA00458 single-stranded DNA-binding protein
Probab=29.31  E-value=57  Score=33.13  Aligned_cols=8  Identities=0%  Similarity=0.094  Sum_probs=3.5

Q ss_pred             hhhhhcCC
Q 011601           69 LERCYKAP   76 (481)
Q Consensus        69 ~er~~~~~   76 (481)
                      |++|+...
T Consensus        59 I~~~hee~   66 (233)
T PHA00458         59 IVKAHEEN   66 (233)
T ss_pred             HHHHHHHH
Confidence            44444433


No 30 
>PF04858 TH1:  TH1 protein;  InterPro: IPR006942 TH1 is a highly conserved but uncharacterised metazoan protein. No homologue has been identified in Caenorhabditis elegans []. TH1 binds specifically to A-Raf kinase [].; GO: 0045892 negative regulation of transcription, DNA-dependent, 0005634 nucleus
Probab=29.24  E-value=1.1e+02  Score=34.71  Aligned_cols=50  Identities=16%  Similarity=0.333  Sum_probs=31.4

Q ss_pred             HHHHHHHhhccCc-----h---HHHHHHHHcCCcCHHHHHHHHH--hccCcchhHHHHhhc
Q 011601          185 AVLNEWMKTMMDL-----P---AGFRQAYEMGLVSSAQMVKFLA--INARPTTTRFISRSL  235 (481)
Q Consensus       185 aVL~E~~rt~~sL-----P---aDl~~A~e~Glvssa~L~rFl~--L~~~P~~~~~l~R~l  235 (481)
                      .++.||.+.+..=     |   .+|.+=++.|.= +++++++|+  +.+.+.+..++..|+
T Consensus        36 ~~~~~~~~~l~~~D~Imep~i~~~i~~y~~~gG~-p~~vv~~Ls~~Y~g~aq~~~ll~~WL   95 (584)
T PF04858_consen   36 EVLEECLRRLSQPDAIMEPSIFDTIKRYFRAGGD-PEEVVELLSENYRGYAQMCNLLAEWL   95 (584)
T ss_pred             HHHHHHHHhcCCCCeeeCchHHHHHHHHHHCCCC-HHHHHHHHHHhcccHHHHHHHHHHHH
Confidence            4666776664332     2   345555667774 667777775  456677777777776


No 31 
>COG2818 Tag 3-methyladenine DNA glycosylase [DNA replication, recombination, and repair]
Probab=27.15  E-value=37  Score=33.43  Aligned_cols=125  Identities=10%  Similarity=0.028  Sum_probs=81.8

Q ss_pred             chhhhhHHHHhhcHHHHHHHHHHHHhhccCchHHHHHHHHcCCcCHH--HHHHHHHhccC-cchhHHHHhhcCcchhhhh
Q 011601          167 LFRRRKILEELFDRKFVDAVLNEWMKTMMDLPAGFRQAYEMGLVSSA--QMVKFLAINAR-PTTTRFISRSLPQGISRAF  243 (481)
Q Consensus       167 ~~r~r~~~~e~fdR~~i~aVL~E~~rt~~sLPaDl~~A~e~Glvssa--~L~rFl~L~~~-P~~~~~l~R~lp~~~srgf  243 (481)
                      ++++|...||.|..=+++.|.+.-..-++.|=+|-.=.--.|.|-+.  -=+.+++|+.. .....|+--+.+.   +.-
T Consensus        50 VL~KRe~freaF~~Fd~~kVA~~~~~dverLl~d~gIIR~r~KI~A~i~NA~~~l~l~~e~Gsf~~flWsf~~~---~~~  126 (188)
T COG2818          50 VLKKREAFREAFHGFDPEKVAAMTEEDVERLLADAGIIRNRGKIKATINNARAVLELQKEFGSFSEFLWSFVGG---KPS  126 (188)
T ss_pred             HHHhHHHHHHHHhcCCHHHHHcCCHHHHHHHHhCcchhhhHHHHHHHHHHHHHHHHHHHHcCCHHHHHHHhcCC---Ccc
Confidence            57888899999988777777765544444443332111113333121  23578888886 6666666555543   456


Q ss_pred             hhhhccCchhHHHHHHHHHHhhhhhhhhhhhhcchhhhhHHHHHHHHHHHHHHhhHH
Q 011601          244 IGRMLADPSFLYKLILEQAATIGCTVLWELENRKERIKQEWDLALINVLTVTACNAF  300 (481)
Q Consensus       244 R~RlLADP~FLfKL~iE~~I~i~~~~~aE~~~Rge~F~~ElDfV~sdvv~g~v~nfa  300 (481)
                      +.....=+.|+.+      ..++.++..++++||-.|.-.-=.+.-...+|.|.|-+
T Consensus       127 ~~~~~~~~~~pa~------t~~S~~mskaLKkrGf~fvGpTt~yafmqA~G~vndH~  177 (188)
T COG2818         127 RNQVNDGSEVPAS------TELSDAMSKALKKRGFKFVGPTTVYAFMQATGLVNDHA  177 (188)
T ss_pred             cccccchhhcccc------chhHHHHHHHHHHccCeecCcHHHHHHHHHHcchHHHH
Confidence            6666666777777      33466788899999999988887777777888887654


No 32 
>KOG2023 consensus Nuclear transport receptor Karyopherin-beta2/Transportin (importin beta superfamily) [Nuclear structure; Intracellular trafficking, secretion, and vesicular transport]
Probab=26.88  E-value=35  Score=39.64  Aligned_cols=39  Identities=15%  Similarity=0.347  Sum_probs=17.6

Q ss_pred             CCccCc-ccCcccceeeecccccceecc--cccCceeecCCC
Q 011601           89 PLMKGG-QYGAFGAVTLEKGKLDTTQQQ--SETTPELATGGG  127 (481)
Q Consensus        89 ~~~~g~-~~g~~ga~tleks~l~~~q~~--~~~~p~l~~GgG  127 (481)
                      ||.-++ .|.----+.|+.-.=|-++.-  .+.+|..+.+--
T Consensus       302 PvLl~~M~Ysd~D~~LL~~~eeD~~vpDreeDIkPRfhksk~  343 (885)
T KOG2023|consen  302 PVLLSGMVYSDDDIILLKNNEEDESVPDREEDIKPRFHKSKE  343 (885)
T ss_pred             HHHHccCccccccHHHhcCccccccCCchhhhccchhhhchh
Confidence            555554 444333333331222333333  355787776543


No 33 
>PF08595 RXT2_N:  RXT2-like, N-terminal;  InterPro: IPR013904  The entry represents the N-terminal region of RXT2-like proteins. In Saccharomyces cerevisiae (Baker's yeast), RXT2 has been demonstrated to be involved in conjugation with cellular fusion (mating) and invasive growth []. A high throughput localisation study has localised RXT2 to the nucleus []. 
Probab=24.13  E-value=41  Score=31.77  Aligned_cols=6  Identities=17%  Similarity=0.440  Sum_probs=2.7

Q ss_pred             ccCcch
Q 011601          222 NARPTT  227 (481)
Q Consensus       222 ~~~P~~  227 (481)
                      -.+|.+
T Consensus       104 ~tHPai  109 (149)
T PF08595_consen  104 PTHPAI  109 (149)
T ss_pred             cCCccc
Confidence            345544


No 34 
>PLN03083 E3 UFM1-protein ligase 1 homolog; Provisional
Probab=23.39  E-value=1.2e+02  Score=35.64  Aligned_cols=35  Identities=14%  Similarity=0.253  Sum_probs=23.3

Q ss_pred             HhhcHHHHHHHHHHHHhhccC----chHHHHHHHHcCCc
Q 011601          176 ELFDRKFVDAVLNEWMKTMMD----LPAGFRQAYEMGLV  210 (481)
Q Consensus       176 e~fdR~~i~aVL~E~~rt~~s----LPaDl~~A~e~Glv  210 (481)
                      .+|.=++|..++.+|...+++    -|..|...+..-+.
T Consensus       484 ~i~~~e~i~~~l~~~~~d~~e~~~~~~~~ll~~lA~~l~  522 (803)
T PLN03083        484 NIPPEEWVMKKILEWVPDLEEDGTEDPGSILKHLADHLR  522 (803)
T ss_pred             ccccHHHHHHHHHHHcccchhcccccHHHHHHHHHHHHH
Confidence            334667888899988887544    66677766554444


No 35 
>COG3685 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=23.32  E-value=54  Score=31.84  Aligned_cols=53  Identities=25%  Similarity=0.421  Sum_probs=45.6

Q ss_pred             HHHHhhcH----------HHHHHHHHHHHhhccCchHHHHHHHHcCCcCHHHHHHHHHhccCcch
Q 011601          173 ILEELFDR----------KFVDAVLNEWMKTMMDLPAGFRQAYEMGLVSSAQMVKFLAINARPTT  227 (481)
Q Consensus       173 ~~~e~fdR----------~~i~aVL~E~~rt~~sLPaDl~~A~e~Glvssa~L~rFl~L~~~P~~  227 (481)
                      .++|||+|          ..++-++.|...-++..|.+  ++++.|++.+++.+.-+++.-+.++
T Consensus        57 rLe~Vfe~~g~~~~~~~cda~~giiaegq~i~~~~~~~--evlda~L~~aaq~vEhyEIA~YgtL  119 (167)
T COG3685          57 RLEQVFERLGKKARRVTCDAMEGLIAEGQEIMEEFKSN--EVLDAGLIGAAQKVEHYEIACYGTL  119 (167)
T ss_pred             HHHHHHHHhCcccccchHHHHHHHHHHHHHHHHhcccc--HHHHHHHHHHHHHHHHHHHHHHHHH
Confidence            47899999          66778888888889999999  9999999999999999998876543


No 36 
>PF04931 DNA_pol_phi:  DNA polymerase phi;  InterPro: IPR007015 Proteins of this family are predominantly nucleolar. The majority are described as transcription factor transactivators. The family also includes the fifth essential DNA polymerase (Pol5p) of Schizosaccharomyces pombe (Fission yeast) and Saccharomyces cerevisiae (Baker's yeast) (2.7.7.7 from EC). Pol5p is localized exclusively to the nucleolus and binds near or at the enhancer region of rRNA-encoding DNA repeating units.; GO: 0003677 DNA binding, 0003887 DNA-directed DNA polymerase activity, 0006351 transcription, DNA-dependent
Probab=22.71  E-value=48  Score=37.98  Aligned_cols=10  Identities=10%  Similarity=0.295  Sum_probs=4.3

Q ss_pred             hhhhhhhhhc
Q 011601           65 SVVLLERCYK   74 (481)
Q Consensus        65 ~~~~~er~~~   74 (481)
                      .....+|.+.
T Consensus       559 l~~c~~~~~~  568 (784)
T PF04931_consen  559 LQICYEKAFG  568 (784)
T ss_pred             HHHHHHHHhc
Confidence            3444444444


No 37 
>KOG3973 consensus Uncharacterized conserved glycine-rich protein [Function unknown]
Probab=21.56  E-value=69  Score=34.73  Aligned_cols=9  Identities=22%  Similarity=0.272  Sum_probs=4.6

Q ss_pred             eeeeccccc
Q 011601          102 VTLEKGKLD  110 (481)
Q Consensus       102 ~tleks~l~  110 (481)
                      -|.+++|--
T Consensus       340 g~f~~~~~~  348 (465)
T KOG3973|consen  340 GTFDRPKTH  348 (465)
T ss_pred             CCCcCcccc
Confidence            455555543


No 38 
>PRK05325 hypothetical protein; Provisional
Probab=21.44  E-value=71  Score=34.64  Aligned_cols=8  Identities=25%  Similarity=0.459  Sum_probs=4.1

Q ss_pred             hhhHHHHH
Q 011601          280 IKQEWDLA  287 (481)
Q Consensus       280 F~~ElDfV  287 (481)
                      |++++|+-
T Consensus       202 f~d~~DlR  209 (401)
T PRK05325        202 FIDPFDLR  209 (401)
T ss_pred             CCCccccc
Confidence            55555543


No 39 
>KOG3241 consensus Uncharacterized conserved protein [Function unknown]
Probab=20.95  E-value=70  Score=31.87  Aligned_cols=7  Identities=43%  Similarity=0.572  Sum_probs=3.0

Q ss_pred             CCCCCCC
Q 011601          134 KINHGGG  140 (481)
Q Consensus       134 ~~~~GGG  140 (481)
                      -+++||-
T Consensus       182 ~i~~~~~  188 (227)
T KOG3241|consen  182 IIGHGSV  188 (227)
T ss_pred             hcccCCC
Confidence            3444443


No 40 
>PF06946 Phage_holin_5:  Phage holin;  InterPro: IPR009708 This entry represents the Bacteriophage A118, holin protein. The characteristics of the protein distribution suggest prophage matches in addition to the phage matches. This protein family represent one of a large number of mutually dissimilar families of phage holins. It is thought that the temporal precision of holin-mediated lysis may occur through the build-up of a holin oligomer which causes the lysis [].
Probab=20.93  E-value=2.4e+02  Score=25.15  Aligned_cols=27  Identities=19%  Similarity=0.122  Sum_probs=20.6

Q ss_pred             CcchhhhHHHHHhhchhhhhhHhHHHhhHHHHHHHHhh
Q 011601          339 EFDLQKRIHSLFYKAAELCMVGLSAGAVQGSLSNYLAG  376 (481)
Q Consensus       339 ~fsl~qRiga~v~KGa~l~~VG~~aGlvG~glSN~L~~  376 (481)
                      ++++.+++.+           |.+||+.+++|-...-+
T Consensus        57 ~~~l~~~~~a-----------G~laGlAaTGL~e~~t~   83 (93)
T PF06946_consen   57 DGNLALMAWA-----------GGLAGLAATGLFEQFTN   83 (93)
T ss_pred             CccHHHHHHH-----------HHHhhhhhhhHHHHHHh
Confidence            4566666544           88999999999888877


No 41 
>PF04871 Uso1_p115_C:  Uso1 / p115 like vesicle tethering protein, C terminal region;  InterPro: IPR006955 This domain identifies a group of proteins, which are described as: General vesicular transport factor, Transcytosis associate protein (TAP) and Vesicle docking protein. This myosin-shaped molecule consists of an N-terminal globular head region, a coiled-coil tail which mediates dimerisation, and a short C-terminal acidic region []. p115 tethers COP1 vesicles to the Golgi by binding the coiled coil proteins giantin (on the vesicles) and GM130 (on the Golgi), via its C-terminal acidic region. It is required for intercisternal transport in the Golgi stack. This domain is found in the acidic C-terminal region, which binds to the golgins giantin and GM130. p115 is thought to juxtapose two membranes by binding giantin with one acidic region, and GM130 with another [].; GO: 0008565 protein transporter activity, 0006886 intracellular protein transport, 0005737 cytoplasm, 0016020 membrane
Probab=20.74  E-value=71  Score=29.42  Aligned_cols=21  Identities=43%  Similarity=0.634  Sum_probs=0.0

Q ss_pred             CCCCCCCCCCCCCCCCCCCCC
Q 011601          141 DGGDDDGDDDDYFDDFDDGDE  161 (481)
Q Consensus       141 gGgdddGddddy~d~~ddgd~  161 (481)
                      ++.++++++||.+|+++|+|+
T Consensus       116 ddE~~~d~~dd~edd~~deee  136 (136)
T PF04871_consen  116 DDEDSEDDEDDDEDDDEDEEE  136 (136)
T ss_pred             CCccccccCCCCCCccCCCCC


No 42 
>KOG2023 consensus Nuclear transport receptor Karyopherin-beta2/Transportin (importin beta superfamily) [Nuclear structure; Intracellular trafficking, secretion, and vesicular transport]
Probab=20.50  E-value=99  Score=36.19  Aligned_cols=34  Identities=9%  Similarity=0.304  Sum_probs=15.3

Q ss_pred             ccCcchhHHHHhhcCcchhhhhhhhhccCch--hHHHHHHH
Q 011601          222 NARPTTTRFISRSLPQGISRAFIGRMLADPS--FLYKLILE  260 (481)
Q Consensus       222 ~~~P~~~~~l~R~lp~~~srgfR~RlLADP~--FLfKL~iE  260 (481)
                      ++.|.+....||.|..     |-.-+.-||.  |+-+|..+
T Consensus       445 DKkplVRsITCWTLsR-----ys~wv~~~~~~~~f~pvL~~  480 (885)
T KOG2023|consen  445 DKKPLVRSITCWTLSR-----YSKWVVQDSRDEYFKPVLEG  480 (885)
T ss_pred             cCccceeeeeeeeHhh-----hhhhHhcCChHhhhHHHHHH
Confidence            3446554444554432     2223444543  55555443


Done!