Query         030998
Match_columns 167
No_of_seqs    134 out of 860
Neff          5.8 
Searched_HMMs 46136
Date          Fri Mar 29 07:42:46 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/030998.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/030998hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF06280 DUF1034:  Fn3-like dom  95.3    0.25 5.3E-06   36.1   9.1   82   63-149    10-111 (112)
  2 PF10633 NPCBM_assoc:  NPCBM-as  95.0    0.32   7E-06   33.3   8.5   58   60-117     4-62  (78)
  3 PF11614 FixG_C:  IG-like fold   86.1     4.2 9.2E-05   29.7   6.9   54   64-118    34-87  (118)
  4 COG1470 Predicted membrane pro  76.8      25 0.00055   32.8   9.6  104   45-160   381-491 (513)
  5 smart00635 BID_2 Bacterial Ig-  69.5      22 0.00049   24.2   6.0   39   90-139     4-42  (81)
  6 PF14874 PapD-like:  Flagellar-  67.5      39 0.00085   23.5  12.1   79   62-151    21-99  (102)
  7 PF09244 DUF1964:  Domain of un  67.1      19 0.00041   24.7   4.9   44   97-144    15-58  (68)
  8 PF00345 PapD_N:  Pili and flag  59.3      65  0.0014   23.3   7.4   51   63-115    16-73  (122)
  9 TIGR02745 ccoG_rdxA_fixG cytoc  57.6      44 0.00095   30.6   7.1   53   64-117   349-401 (434)
 10 PF00347 Ribosomal_L6:  Ribosom  52.2      37  0.0008   22.6   4.5   26   86-111     2-27  (77)
 11 COG1470 Predicted membrane pro  51.5 1.4E+02   0.003   28.2   9.2  110    5-117   221-345 (513)
 12 PF00635 Motile_Sperm:  MSP (Ma  50.2      84  0.0018   21.8   8.2   53   62-117    19-71  (109)
 13 PF07718 Coatamer_beta_C:  Coat  44.3 1.2E+02  0.0027   23.7   6.8   50   81-138    90-139 (140)
 14 cd00237 p23 p23 binds heat sho  43.9      91   0.002   22.8   5.8   22   82-103    18-39  (106)
 15 PF13195 DUF4011:  Protein of u  36.9      53  0.0011   25.9   3.8   27  129-155   120-147 (176)
 16 PF14874 PapD-like:  Flagellar-  35.2 1.1E+02  0.0023   21.2   4.9   28   90-117     3-32  (102)
 17 PF12180 EABR:  TSG101 and ALIX  31.7     6.5 0.00014   23.8  -1.6   16    5-20     15-30  (35)
 18 PLN03080 Probable beta-xylosid  31.3 1.7E+02  0.0037   28.8   7.0   53   62-115   685-744 (779)
 19 PF12970 DUF3858:  Domain of Un  30.8      87  0.0019   23.8   3.9   36   76-111    36-71  (116)
 20 PF15496 DUF4646:  Domain of un  30.8      26 0.00056   26.5   1.1   16    7-22     45-60  (123)
 21 PF07610 DUF1573:  Protein of u  30.8 1.3E+02  0.0028   18.3   5.2   12  102-113    34-45  (45)
 22 PF00635 Motile_Sperm:  MSP (Ma  30.3 1.8E+02   0.004   20.1   5.5   53   90-152     1-55  (109)
 23 PF08310 LGFP:  LGFP repeat;  I  21.5      47   0.001   21.1   0.9   24  130-153    22-45  (54)

No 1  
>PF06280 DUF1034:  Fn3-like domain (DUF1034);  InterPro: IPR010435 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold:  Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases.   In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding.  Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This domain of unknown function is present in bacterial and plant peptidases belonging to MEROPS peptidase family S8 (subfamily S8A subtilisin, clan SB). It is C-terminal to and adjacent to the S8 peptidase domain and can be found in conjunction with the PA (Protease associated) domain (IPR003137 from INTERPRO) and additionally in Gram-positive bacteria with the surface protein anchor domain (IPR001899 from INTERPRO).; GO: 0004252 serine-type endopeptidase activity, 0005618 cell wall, 0016020 membrane; PDB: 3EIF_A 1XF1_B.
Probab=95.26  E-value=0.25  Score=36.14  Aligned_cols=82  Identities=23%  Similarity=0.336  Sum_probs=48.8

Q ss_pred             EEEEEEEEecCCCCcceEEEEEC--------CCC----------c-EEEEEcCEEEEeecCceEEEEEEEEEccccccCC
Q 030998           63 FTFKRVLTNVADTRSTYTAAVKA--------PVG----------M-TVTVEPATLSFAGKFSKAEFSLTVNINLGNAFSP  123 (167)
Q Consensus        63 ~tv~RTVTNVg~~~stY~a~V~~--------P~g----------v-~V~V~P~~L~F~~~gqk~sf~Vt~~~~~~~~~~~  123 (167)
                      .+++=+++|.|+..-+|+.+...        ..|          . .++..|..++. ++||++.++|+|...... ...
T Consensus        10 ~~~~itl~N~~~~~~ty~~~~~~~~t~~~~~~~~~~~~~~~~~~~~~~~~~~~~vTV-~ag~s~~v~vti~~p~~~-~~~   87 (112)
T PF06280_consen   10 FSFTITLHNYGDKPVTYTLSHVPVLTDKTDTEEGYSILVPPVPSISTVSFSPDTVTV-PAGQSKTVTVTITPPSGL-DAS   87 (112)
T ss_dssp             EEEEEEEEE-SSS-EEEEEEEE-EEEEEE--ETTEEEEEEEE----EEE---EEEEE--TTEEEEEEEEEE--GGG-HHT
T ss_pred             eEEEEEEEECCCCCEEEEEeeEEEEeeEeeccCCcccccccccceeeEEeCCCeEEE-CCCCEEEEEEEEEehhcC-Ccc
Confidence            56666777777777777665541        112          1 56777888888 789999999999984420 002


Q ss_pred             CCCccceEEEEEEEeCCCce-EEEceE
Q 030998          124 KSNFLGNFGYLTWYEVKRKH-TVRSPI  149 (167)
Q Consensus       124 ~~~~~~~fGsl~W~d~~g~h-~VRSPI  149 (167)
                      ...+  ..|.|..+++ ..+ .++.|.
T Consensus        88 ~~~~--~eG~I~~~~~-~~~~~lsIPy  111 (112)
T PF06280_consen   88 NGPF--YEGFITFKSS-DGEPDLSIPY  111 (112)
T ss_dssp             T-EE--EEEEEEEESS-TTSEEEEEEE
T ss_pred             cCCE--EEEEEEEEcC-CCCEEEEeee
Confidence            2345  8899999986 444 777774


No 2  
>PF10633 NPCBM_assoc:  NPCBM-associated, NEW3 domain of alpha-galactosidase;  InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=94.96  E-value=0.32  Score=33.27  Aligned_cols=58  Identities=19%  Similarity=0.260  Sum_probs=38.2

Q ss_pred             CeeEEEEEEEEecCCCC-cceEEEEECCCCcEEEEEcCEEEEeecCceEEEEEEEEEcc
Q 030998           60 TASFTFKRVLTNVADTR-STYTAAVKAPVGMTVTVEPATLSFAGKFSKAEFSLTVNINL  117 (167)
Q Consensus        60 ~~~~tv~RTVTNVg~~~-stY~a~V~~P~gv~V~V~P~~L~F~~~gqk~sf~Vt~~~~~  117 (167)
                      ....+++=+|+|-|... ...++++..|.|-.+...|..+.--++||+++++++|+...
T Consensus         4 G~~~~~~~tv~N~g~~~~~~v~~~l~~P~GW~~~~~~~~~~~l~pG~s~~~~~~V~vp~   62 (78)
T PF10633_consen    4 GETVTVTLTVTNTGTAPLTNVSLSLSLPEGWTVSASPASVPSLPPGESVTVTFTVTVPA   62 (78)
T ss_dssp             TEEEEEEEEEE--SSS-BSS-EEEEE--TTSE---EEEEE--B-TTSEEEEEEEEEE-T
T ss_pred             CCEEEEEEEEEECCCCceeeEEEEEeCCCCccccCCccccccCCCCCEEEEEEEEECCC
Confidence            35678888999999754 56889999999999999999887668999999999998764


No 3  
>PF11614 FixG_C:  IG-like fold at C-terminal of FixG, putative oxidoreductase; PDB: 2R39_A.
Probab=86.14  E-value=4.2  Score=29.72  Aligned_cols=54  Identities=20%  Similarity=0.197  Sum_probs=37.5

Q ss_pred             EEEEEEEecCCCCcceEEEEECCCCcEEEEEcCEEEEeecCceEEEEEEEEEccc
Q 030998           64 TFKRVLTNVADTRSTYTAAVKAPVGMTVTVEPATLSFAGKFSKAEFSLTVNINLG  118 (167)
Q Consensus        64 tv~RTVTNVg~~~stY~a~V~~P~gv~V~V~P~~L~F~~~gqk~sf~Vt~~~~~~  118 (167)
                      ..+=.++|-+...-+|+.++..++|+++......+.. ++|++..+.+.+.+...
T Consensus        34 ~Y~lkl~Nkt~~~~~~~i~~~g~~~~~l~~~~~~i~v-~~g~~~~~~v~v~~p~~   87 (118)
T PF11614_consen   34 QYTLKLTNKTNQPRTYTISVEGLPGAELQGPENTITV-PPGETREVPVFVTAPPD   87 (118)
T ss_dssp             EEEEEEEE-SSS-EEEEEEEES-SS-EE-ES--EEEE--TT-EEEEEEEEEE-GG
T ss_pred             EEEEEEEECCCCCEEEEEEEecCCCeEEECCCcceEE-CCCCEEEEEEEEEECHH
Confidence            3556789999999999999999999999654478888 78999999999987654


No 4  
>COG1470 Predicted membrane protein [Function unknown]
Probab=76.85  E-value=25  Score=32.85  Aligned_cols=104  Identities=12%  Similarity=0.144  Sum_probs=70.3

Q ss_pred             CCCcce--EEEEecCCCCeeEEEEEEEEecCCCC-cceEEEEECCCCcEEEEEcCEEEEeecCceEEEEEEEEEcccccc
Q 030998           45 DLNYPS--FIIILNNSNTASFTFKRVLTNVADTR-STYTAAVKAPVGMTVTVEPATLSFAGKFSKAEFSLTVNINLGNAF  121 (167)
Q Consensus        45 dLNYPS--i~v~~~~~~~~~~tv~RTVTNVg~~~-stY~a~V~~P~gv~V~V~P~~L~F~~~gqk~sf~Vt~~~~~~~~~  121 (167)
                      +|+.++  +.+...  .....++.=.++|-|+.+ .-=+.+|.+|.|-++.|.|+++---++|+.++..+++++...   
T Consensus       381 ~v~l~~g~~~lt~t--aGee~~i~i~I~NsGna~LtdIkl~v~~PqgWei~Vd~~~I~sL~pge~~tV~ltI~vP~~---  455 (513)
T COG1470         381 LVKLDNGPYRLTIT--AGEEKTIRISIENSGNAPLTDIKLTVNGPQGWEIEVDESTIPSLEPGESKTVSLTITVPED---  455 (513)
T ss_pred             eEEccCCcEEEEec--CCccceEEEEEEecCCCccceeeEEecCCccceEEECcccccccCCCCcceEEEEEEcCCC---
Confidence            566666  444432  233567777789999754 456899999999999999999887799999999999998764   


Q ss_pred             CCCCCccce----EEEEEEEeCCCceEEEceEEEEEeeCCCcc
Q 030998          122 SPKSNFLGN----FGYLTWYEVKRKHTVRSPIVAAFANNSRGV  160 (167)
Q Consensus       122 ~~~~~~~~~----fGsl~W~d~~g~h~VRSPIaV~~~~~~~~~  160 (167)
                      +..++|...    --..+|.|       +.-|.|...+.++.+
T Consensus       456 a~aGdY~i~i~~ksDq~s~e~-------tlrV~V~~sS~st~i  491 (513)
T COG1470         456 AGAGDYRITITAKSDQASSED-------TLRVVVGQSSTSTYI  491 (513)
T ss_pred             CCCCcEEEEEEEeeccccccc-------eEEEEEeccccchhh
Confidence            145555200    11345665       444566655555544


No 5  
>smart00635 BID_2 Bacterial Ig-like domain 2.
Probab=69.47  E-value=22  Score=24.21  Aligned_cols=39  Identities=33%  Similarity=0.364  Sum_probs=29.1

Q ss_pred             EEEEEcCEEEEeecCceEEEEEEEEEccccccCCCCCccceEEEEEEEeC
Q 030998           90 TVTVEPATLSFAGKFSKAEFSLTVNINLGNAFSPKSNFLGNFGYLTWYEV  139 (167)
Q Consensus        90 ~V~V~P~~L~F~~~gqk~sf~Vt~~~~~~~~~~~~~~~~~~fGsl~W~d~  139 (167)
                      .|++.|..+.+ ..|+++.|++++.....     .     ....++|+.+
T Consensus         4 ~i~i~p~~~~l-~~G~~~~l~a~~~~~~~-----~-----~~~~v~w~Ss   42 (81)
T smart00635        4 SVTVTPTTASV-KKGLTLQLTATVTPSSA-----K-----VTGKVTWTSS   42 (81)
T ss_pred             EEEEeCCeeEE-eCCCeEEEEEEEECCCC-----C-----ccceEEEEEC
Confidence            67889999998 58999999999764332     1     1366889874


No 6  
>PF14874 PapD-like:  Flagellar-associated PapD-like
Probab=67.52  E-value=39  Score=23.52  Aligned_cols=79  Identities=18%  Similarity=0.258  Sum_probs=50.7

Q ss_pred             eEEEEEEEEecCCCCcceEEEEECCCCcEEEEEcCEEEEeecCceEEEEEEEEEccccccCCCCCccceEEEEEEEeCCC
Q 030998           62 SFTFKRVLTNVADTRSTYTAAVKAPVGMTVTVEPATLSFAGKFSKAEFSLTVNINLGNAFSPKSNFLGNFGYLTWYEVKR  141 (167)
Q Consensus        62 ~~tv~RTVTNVg~~~stY~a~V~~P~gv~V~V~P~~L~F~~~gqk~sf~Vt~~~~~~~~~~~~~~~~~~fGsl~W~d~~g  141 (167)
                      ..+.+=+++|.|.....|++.......-..+|.|..= +-++|++..++|+|....     ..+.+   -+.|.-.-  .
T Consensus        21 ~~~~~v~l~N~s~~p~~f~v~~~~~~~~~~~v~~~~g-~l~PG~~~~~~V~~~~~~-----~~g~~---~~~l~i~~--e   89 (102)
T PF14874_consen   21 TYSRTVTLTNTSSIPARFRVRQPESLSSFFSVEPPSG-FLAPGESVELEVTFSPTK-----PLGDY---EGSLVITT--E   89 (102)
T ss_pred             EEEEEEEEEECCCCCEEEEEEeCCcCCCCEEEECCCC-EECCCCEEEEEEEEEeCC-----CCceE---EEEEEEEE--C
Confidence            3444556799999999999876432344566676653 337899999999999543     22333   46665555  3


Q ss_pred             ceEEEceEEE
Q 030998          142 KHTVRSPIVA  151 (167)
Q Consensus       142 ~h~VRSPIaV  151 (167)
                      ...+..|+-.
T Consensus        90 ~~~~~i~v~a   99 (102)
T PF14874_consen   90 GGSFEIPVKA   99 (102)
T ss_pred             CeEEEEEEEE
Confidence            4455555543


No 7  
>PF09244 DUF1964:  Domain of unknown function (DUF1964);  InterPro: IPR015325 This domain is C-terminal to the catalytic sucrose phosphorylase beta/alpha barrel domain. It adopts a beta-sandwich fold, with Greek-key topology and is functionally uncharacterised []. ; PDB: 1R7A_B 2GDU_A 2GDV_A.
Probab=67.07  E-value=19  Score=24.65  Aligned_cols=44  Identities=16%  Similarity=0.243  Sum_probs=25.6

Q ss_pred             EEEEeecCceEEEEEEEEEccccccCCCCCccceEEEEEEEeCCCceE
Q 030998           97 TLSFAGKFSKAEFSLTVNINLGNAFSPKSNFLGNFGYLTWYEVKRKHT  144 (167)
Q Consensus        97 ~L~F~~~gqk~sf~Vt~~~~~~~~~~~~~~~~~~fGsl~W~d~~g~h~  144 (167)
                      .|+|+=.|++-+=+++|+....-  -.....  .-.+|.|+|+.|.|.
T Consensus        15 Sitf~W~g~~t~atLtFePg~Gl--g~~n~~--pVatl~W~DsaG~H~   58 (68)
T PF09244_consen   15 SITFTWTGATTSATLTFEPGRGL--GVDNTT--PVATLAWTDSAGDHR   58 (68)
T ss_dssp             EEEEEEE-SS-EEEEEE-GGGC---STT--S----EEEEEEETTEEEE
T ss_pred             EEEEEEeccccEEEEEEccCccc--CccCCc--ceeEEEEeccCCCcc
Confidence            56788788888888999864320  011122  458999999877776


No 8  
>PF00345 PapD_N:  Pili and flagellar-assembly chaperone, PapD N-terminal domain;  InterPro: IPR016147 Most Gram-negative bacteria possess a supramolecular structure - the pili - on their surface, which mediates attachment to specific receptors. Many interactive subunits are required to assemble pili, but their assembly only takes place after translocation across the cytoplasmic membrane. Periplasmic chaperones assist pili assembly by binding to the subunits, thereby preventing premature aggregation [, ]. Pili chaperones are structurally, and possibly evolutionarily, related to the immunoglobulin superfamily [, ]: they contain two globular domains, with a topology identical to an immunoglobulin fold. This entry represents the N-terminal domain of pili assembly chaperone, and has a beta-sandwich fold consisting of seven strands in two sheets with a Greek key topology.; GO: 0007047 cellular cell wall organization, 0030288 outer membrane-bounded periplasmic space; PDB: 2CO6_B 2CO7_B 1L4I_B 3GFU_A 3F65_F 3F6L_A 3F6I_A 3GEW_B 3DSN_D 2OS7_B ....
Probab=59.26  E-value=65  Score=23.29  Aligned_cols=51  Identities=12%  Similarity=-0.033  Sum_probs=38.4

Q ss_pred             EEEEEEEEecCCCCcceEEEEEC---CCC----cEEEEEcCEEEEeecCceEEEEEEEEE
Q 030998           63 FTFKRVLTNVADTRSTYTAAVKA---PVG----MTVTVEPATLSFAGKFSKAEFSLTVNI  115 (167)
Q Consensus        63 ~tv~RTVTNVg~~~stY~a~V~~---P~g----v~V~V~P~~L~F~~~gqk~sf~Vt~~~  115 (167)
                      .+..=+|+|-|+..-.+.+.+.-   ..+    -.+-|.|..+.. ++|+++.+.| +..
T Consensus        16 ~~~~i~v~N~~~~~~~vq~~v~~~~~~~~~~~~~~~~vsPp~~~L-~pg~~q~vRv-~~~   73 (122)
T PF00345_consen   16 RSASITVTNNSDQPYLVQVWVYDQDDEDEDEPTDPFIVSPPIFRL-EPGESQTVRV-YRG   73 (122)
T ss_dssp             SEEEEEEEESSSSEEEEEEEEEETTSTTSSSSSSSEEEESSEEEE-ETTEEEEEEE-EEC
T ss_pred             CEEEEEEEcCCCCcEEEEEEEEcCCCcccccccccEEEeCCceEe-CCCCcEEEEE-Eec
Confidence            35567889988877777777763   111    257799999999 7899999999 663


No 9  
>TIGR02745 ccoG_rdxA_fixG cytochrome c oxidase accessory protein FixG. Member of this ferredoxin-like protein family are found exclusively in species with an operon encoding the cbb3 type of cytochrome c oxidase (cco-cbb3), and near the cco-cbb3 operon in about half the cases. The cco-cbb3 is found in a variety of proteobacteria and almost nowhere else, and is associated with oxygen use under microaerobic conditions. Some (but not all) of these proteobacteria are also nitrogen-fixing, hence the gene symbol fixG. FixG was shown essential for functional cco-cbb3 expression in Bradyrhizobium japonicum.
Probab=57.61  E-value=44  Score=30.61  Aligned_cols=53  Identities=13%  Similarity=0.149  Sum_probs=43.8

Q ss_pred             EEEEEEEecCCCCcceEEEEECCCCcEEEEEcCEEEEeecCceEEEEEEEEEcc
Q 030998           64 TFKRVLTNVADTRSTYTAAVKAPVGMTVTVEPATLSFAGKFSKAEFSLTVNINL  117 (167)
Q Consensus        64 tv~RTVTNVg~~~stY~a~V~~P~gv~V~V~P~~L~F~~~gqk~sf~Vt~~~~~  117 (167)
                      ..+=.+.|-+..+.+|+.+++.++|.++...++.+.. ++||+.++.|.+....
T Consensus       349 ~Y~~~i~Nk~~~~~~~~l~v~g~~~~~~~~~~~~i~v-~~g~~~~~~v~v~~~~  401 (434)
T TIGR02745       349 TYTLKILNKTEQPHEYYLSVLGLPGIKIEGPGAPIHV-KAGEKVKLPVFLRTPP  401 (434)
T ss_pred             EEEEEEEECCCCCEEEEEEEecCCCcEEEcCCceEEE-CCCCEEEEEEEEEech
Confidence            3555689999889999999999999998876557777 7899999999988754


No 10 
>PF00347 Ribosomal_L6:  Ribosomal protein L6;  InterPro: IPR020040 Ribosomes are the particles that catalyse mRNA-directed protein synthesis in all organisms. The codons of the mRNA are exposed on the ribosome to allow tRNA binding. This leads to the incorporation of amino acids into the growing polypeptide chain in accordance with the genetic information. Incoming amino acid monomers enter the ribosomal A site in the form of aminoacyl-tRNAs complexed with elongation factor Tu (EF-Tu) and GTP. The growing polypeptide chain, situated in the P site as peptidyl-tRNA, is then transferred to aminoacyl-tRNA and the new peptidyl-tRNA, extended by one residue, is translocated to the P site with the aid the elongation factor G (EF-G) and GTP as the deacylated tRNA is released from the ribosome through one or more exit sites [, ]. About 2/3 of the mass of the ribosome consists of RNA and 1/3 of protein. The proteins are named in accordance with the subunit of the ribosome which they belong to - the small (S1 to S31) and the large (L1 to L44). Usually they decorate the rRNA cores of the subunits.  Many ribosomal proteins, particularly those of the large subunit, are composed of a globular, surfaced-exposed domain with long finger-like projections that extend into the rRNA core to stabilise its structure. Most of the proteins interact with multiple RNA elements, often from different domains. In the large subunit, about 1/3 of the 23S rRNA nucleotides are at least in van der Waal's contact with protein, and L22 interacts with all six domains of the 23S rRNA. Proteins S4 and S7, which initiate assembly of the 16S rRNA, are located at junctions of five and four RNA helices, respectively. In this way proteins serve to organise and stabilise the rRNA tertiary structure. While the crucial activities of decoding and peptide transfer are RNA based, proteins play an active role in functions that may have evolved to streamline the process of protein synthesis. In addition to their function in the ribosome, many ribosomal proteins have some function 'outside' the ribosome [, ]. L6 is a protein from the large (50S) subunit. In Escherichia coli, it is located in the aminoacyl-tRNA binding site of the peptidyltransferase centre, and is known to bind directly to 23S rRNA. It belongs to a family of ribosomal proteins, including L6 from bacteria, cyanelles (structures that perform similar functions to chloroplasts, but have structural and biochemical characteristics of Cyanobacteria) and mitochondria; and L9 from mammals, Drosophila, plants and yeast. L6 contains two domains with almost identical folds, suggesting that is was derived by the duplication of an ancient RNA-binding protein gene. Analysis reveals several sites on the protein surface where interactions with other ribosome components may occur, the N terminus being involved in protein-protein interactions and the C terminus containing possible RNA-binding sites []. This entry represents the alpha-beta domain found duplicated in ribosomal L6 proteins. This domain consists of two beta-sheets and one alpha-helix packed around single core [].; GO: 0003735 structural constituent of ribosome, 0019843 rRNA binding, 0006412 translation, 0005840 ribosome; PDB: 2HGJ_H 2HGQ_H 2HGU_H 1S1I_H 3O5H_I 3O58_I 3J16_F 3IZS_F 2V47_H 2WDJ_H ....
Probab=52.24  E-value=37  Score=22.56  Aligned_cols=26  Identities=19%  Similarity=0.373  Sum_probs=19.2

Q ss_pred             CCCcEEEEEcCEEEEeecCceEEEEE
Q 030998           86 PVGMTVTVEPATLSFAGKFSKAEFSL  111 (167)
Q Consensus        86 P~gv~V~V~P~~L~F~~~gqk~sf~V  111 (167)
                      |.|++|+++...+.|..+.-++++.+
T Consensus         2 P~gV~v~~~~~~i~v~G~~g~l~~~~   27 (77)
T PF00347_consen    2 PEGVKVTIKGNIITVKGPKGELSRPI   27 (77)
T ss_dssp             STTCEEEEETTEEEEESSSSEEEEEE
T ss_pred             CCcEEEEEeCcEEEEECCCEeEEEEC
Confidence            78899999998888866555555543


No 11 
>COG1470 Predicted membrane protein [Function unknown]
Probab=51.53  E-value=1.4e+02  Score=28.17  Aligned_cols=110  Identities=17%  Similarity=0.232  Sum_probs=69.9

Q ss_pred             eeeeCChhhHHHhhhcCC-CCcceeeeeeccccccc-C------CCC--CCCCcceEEEEecCCCCeeEEEEEEEEecCC
Q 030998            5 LVYDIEIQDYLNYLCAMN-YTSQQIRVVTGTSDFTC-E------HGN--LDLNYPSFIIILNNSNTASFTFKRVLTNVAD   74 (167)
Q Consensus         5 LVYD~~~~DYi~fLC~lg-y~~~~i~~it~~~~~~C-~------~~~--~dLNYPSi~v~~~~~~~~~~tv~RTVTNVg~   74 (167)
                      +...+++.+|+--+-.-| |... .+.+......+- .      +..  ..||--++..-...  .....++=++-|-|.
T Consensus       221 ~~~e~t~g~y~~~i~~~g~ye~~-~~av~l~d~~t~dLkls~~~k~~~ftEl~~s~~~~~i~~--~~t~sf~V~IeN~g~  297 (513)
T COG1470         221 LEVEITPGKYVVLIAKKGIYEKK-KRAVKLNDGETKDLKLSVTEKKSYFTELNSSDIYLEISP--STTASFTVSIENRGK  297 (513)
T ss_pred             eeEEecCcceEEEeccccceecc-eEEEEcCCCcccceeEEEEeccceEEEeecccceeEEcc--CCceEEEEEEccCCC
Confidence            345678888888887778 5443 333332111111 0      111  46776666554431  234567777899999


Q ss_pred             CCcceEEEEE-CCCCcEEEEEcCEEEEe----ecCceEEEEEEEEEcc
Q 030998           75 TRSTYTAAVK-APVGMTVTVEPATLSFA----GKFSKAEFSLTVNINL  117 (167)
Q Consensus        75 ~~stY~a~V~-~P~gv~V~V~P~~L~F~----~~gqk~sf~Vt~~~~~  117 (167)
                      ..-.|.-++. .|+|....-.=..++.+    ++||++.++|.+....
T Consensus       298 ~~d~y~Le~~g~pe~w~~~Fteg~~~vt~vkL~~gE~kdvtleV~ps~  345 (513)
T COG1470         298 QDDEYALELSGLPEGWTAEFTEGELRVTSVKLKPGEEKDVTLEVYPSL  345 (513)
T ss_pred             CCceeEEEeccCCCCcceEEeeCceEEEEEEecCCCceEEEEEEecCC
Confidence            9999999998 89887765554433333    5799999999998754


No 12 
>PF00635 Motile_Sperm:  MSP (Major sperm protein) domain;  InterPro: IPR000535 Major sperm proteins (MSP) are central components in molecular interactions underlying sperm motility in Caenorhabditis elegans, whose sperm employ an amoebae-like crawling motion using a MSP-containing lamellipod, rather than the flagellar-based swimming motion associated with other sperm. These proteins oligomerise to form an extensive filament system that extends from sperm villipoda, along the leading edge of the pseudopod. About 30 MSP isoforms may exist in C. elegans. MSPs form a fibrous network, whereby MSP dimers form helical subfilaments that coil around one another to produce filaments, which in turn form supercoils to produce bundles. The crystal structure of MSP from C. elegans reveals an immunoglobulin (Ig)-like seven-stranded beta sandwich fold []. ; GO: 0005198 structural molecule activity; PDB: 1MSP_A 3MSP_B 2BVU_B 2MSP_C 1Z9O_F 1Z9L_A 3IKK_A 1WIC_A 2CRI_A 2RR3_A ....
Probab=50.22  E-value=84  Score=21.84  Aligned_cols=53  Identities=17%  Similarity=0.127  Sum_probs=38.7

Q ss_pred             eEEEEEEEEecCCCCcceEEEEECCCCcEEEEEcCEEEEeecCceEEEEEEEEEcc
Q 030998           62 SFTFKRVLTNVADTRSTYTAAVKAPVGMTVTVEPATLSFAGKFSKAEFSLTVNINL  117 (167)
Q Consensus        62 ~~tv~RTVTNVg~~~stY~a~V~~P~gv~V~V~P~~L~F~~~gqk~sf~Vt~~~~~  117 (167)
                      .....=+++|.++..-.|++.-..|..+.  |.|..=.. ++|++....|++....
T Consensus        19 ~~~~~l~l~N~s~~~i~fKiktt~~~~y~--v~P~~G~i-~p~~~~~i~I~~~~~~   71 (109)
T PF00635_consen   19 QQSCELTLTNPSDKPIAFKIKTTNPNRYR--VKPSYGII-EPGESVEITITFQPFD   71 (109)
T ss_dssp             -EEEEEEEEE-SSSEEEEEEEES-TTTEE--EESSEEEE--TTEEEEEEEEE-SSS
T ss_pred             eEEEEEEEECCCCCcEEEEEEcCCCceEE--ecCCCEEE-CCCCEEEEEEEEEecc
Confidence            34555689999999899999999998776  57997544 7899999999988643


No 13 
>PF07718 Coatamer_beta_C:  Coatomer beta C-terminal region;  InterPro: IPR011710 Proteins synthesised on the ribosome and processed in the endoplasmic reticulum are transported from the Golgi apparatus to the trans-Golgi network (TGN), and from there via small carrier vesicles to their final destination compartment. This traffic is bidirectional, to ensure that proteins required to form vesicles are recycled. Vesicles have specific coat proteins (such as clathrin or coatomer) that are important for cargo selection and direction of transfer []. While clathrin mediates endocytic protein transport, and transport from ER to Golgi, coatomers primarily mediate intra-Golgi transport, as well as the reverse Golgi to ER transport of dilysine-tagged proteins []. For example, the coatomer COP1 (coat protein complex 1) is responsible for reverse transport of recycled proteins from Golgi and pre-Golgi compartments back to the ER, while COPII buds vesicles from the ER to the Golgi []. Coatomers reversibly associate with Golgi (non-clathrin-coated) vesicles to mediate protein transport and for budding from Golgi membranes []. Activated small guanine triphosphatases (GTPases) attract coat proteins to specific membrane export sites, thereby linking coatomers to export cargos. As coat proteins polymerise, vesicles are formed and budded from membrane-bound organelles. Coatomer complexes also influence Golgi structural integrity, as well as the processing, activity, and endocytic recycling of LDL receptors. In mammals, coatomer complexes can only be recruited by membranes associated to ADP-ribosylation factors (ARFs), which are small GTP-binding proteins. Coatomer complexes are hetero-oligomers composed of at least an alpha, beta, beta', gamma, delta, epsilon and zeta subunits.  This entry represents the C-terminal domain of the beta subunit from coatomer proteins (Beta-coat proteins). The C-terminal domain probably adapts the function of the N-terminal IPR002553 from INTERPRO domain. Coatomer protein complex I (COPI)-coated vesicles are involved in transport between the endoplasmic reticulum and the Golgi but also participate in transport from early to late endosomes within the endocytic pathway [].  More information about these proteins can be found at Protein of the Month: Clathrin [].; GO: 0005198 structural molecule activity, 0006886 intracellular protein transport, 0016192 vesicle-mediated transport, 0030126 COPI vesicle coat
Probab=44.32  E-value=1.2e+02  Score=23.68  Aligned_cols=50  Identities=10%  Similarity=0.244  Sum_probs=37.8

Q ss_pred             EEEECCCCcEEEEEcCEEEEeecCceEEEEEEEEEccccccCCCCCccceEEEEEEEe
Q 030998           81 AAVKAPVGMTVTVEPATLSFAGKFSKAEFSLTVNINLGNAFSPKSNFLGNFGYLTWYE  138 (167)
Q Consensus        81 a~V~~P~gv~V~V~P~~L~F~~~gqk~sf~Vt~~~~~~~~~~~~~~~~~~fGsl~W~d  138 (167)
                      +....-.++++.=+|...+. .+++....+.+|+....     ..+.  .||.|+|..
T Consensus        90 vElat~gdLklve~p~~~tL-~P~~~~~i~~~iKVsSt-----etGv--IfG~I~Yd~  139 (140)
T PF07718_consen   90 VELATLGDLKLVERPQPITL-APHGFARIKATIKVSST-----ETGV--IFGNIVYDG  139 (140)
T ss_pred             EEEEecCCcEEccCCCceee-CCCcEEEEEEEEEEEec-----cCCE--EEEEEEEec
Confidence            33344456888889998887 78999999999988763     3456  899999853


No 14 
>cd00237 p23 p23 binds heat shock protein (Hsp)90 and participates in the folding of a number of Hsp90 clients, including the progesterone receptor. p23 also has a passive chaperoning activity and in addition may participate in prostaglandin synthesis.
Probab=43.94  E-value=91  Score=22.78  Aligned_cols=22  Identities=23%  Similarity=0.342  Sum_probs=16.6

Q ss_pred             EEECCCCcEEEEEcCEEEEeec
Q 030998           82 AVKAPVGMTVTVEPATLSFAGK  103 (167)
Q Consensus        82 ~V~~P~gv~V~V~P~~L~F~~~  103 (167)
                      .+.-+..+.|.++|..|.|+..
T Consensus        18 ~v~d~~d~~v~l~~~~l~f~~~   39 (106)
T cd00237          18 CVEDSKDVKVDFEKSKLTFSCL   39 (106)
T ss_pred             EeCCCCCcEEEEecCEEEEEEE
Confidence            3333578899999999999863


No 15 
>PF13195 DUF4011:  Protein of unknown function (DUF4011)
Probab=36.86  E-value=53  Score=25.87  Aligned_cols=27  Identities=30%  Similarity=0.682  Sum_probs=21.4

Q ss_pred             ceEEEEEEEeC-CCceEEEceEEEEEee
Q 030998          129 GNFGYLTWYEV-KRKHTVRSPIVAAFAN  155 (167)
Q Consensus       129 ~~fGsl~W~d~-~g~h~VRSPIaV~~~~  155 (167)
                      .+||.|.|.+. ......++|+...+..
T Consensus       120 La~G~L~W~~~~~~~~~~~APLlL~PV~  147 (176)
T PF13195_consen  120 LAFGFLEWYESDDSDKPRRAPLLLIPVE  147 (176)
T ss_pred             eeeeEEEeccCCCCCCEEECCEEEEeEE
Confidence            38999999974 4578999999987643


No 16 
>PF14874 PapD-like:  Flagellar-associated PapD-like
Probab=35.25  E-value=1.1e+02  Score=21.22  Aligned_cols=28  Identities=21%  Similarity=0.230  Sum_probs=21.3

Q ss_pred             EEEEEcCEEEEeec--CceEEEEEEEEEcc
Q 030998           90 TVTVEPATLSFAGK--FSKAEFSLTVNINL  117 (167)
Q Consensus        90 ~V~V~P~~L~F~~~--gqk~sf~Vt~~~~~  117 (167)
                      .++++|+.|.|...  |++.+.+|+++..+
T Consensus         3 ~l~v~P~~ldFG~v~~g~~~~~~v~l~N~s   32 (102)
T PF14874_consen    3 TLEVSPKELDFGNVFVGQTYSRTVTLTNTS   32 (102)
T ss_pred             EEEEeCCEEEeeEEccCCEEEEEEEEEECC
Confidence            67899999999854  67777777776654


No 17 
>PF12180 EABR:  TSG101 and ALIX binding domain of CEP55;  InterPro: IPR022008  This domain family is found in eukaryotes, and is approximately 40 amino acids in length. This domain is the active domain of CEP55. CEP55 is a protein involved in cytokinesis, specifically in abscission of the plasma membrane at the midbody. To perform this function, CEP55 complexes with ESCRT-I (by a Proline rich sequence in its TSG101 domain) and ALIX. This is the domain on CEP55 which binds to both TSG101 and ALIX. It also acts as a hinge between the N and C termini. This domain is called EABR. ; PDB: 3E1R_A.
Probab=31.69  E-value=6.5  Score=23.82  Aligned_cols=16  Identities=31%  Similarity=0.393  Sum_probs=14.5

Q ss_pred             eeeeCChhhHHHhhhc
Q 030998            5 LVYDIEIQDYLNYLCA   20 (167)
Q Consensus         5 LVYD~~~~DYi~fLC~   20 (167)
                      |.||.+-++|+.-||+
T Consensus        15 q~YD~qRE~YV~~L~~   30 (35)
T PF12180_consen   15 QKYDQQREAYVRGLLA   30 (35)
T ss_dssp             HHHHHHHHHHHHHHHH
T ss_pred             HHHHHhHHHHHHHHHH
Confidence            5799999999999997


No 18 
>PLN03080 Probable beta-xylosidase; Provisional
Probab=31.34  E-value=1.7e+02  Score=28.80  Aligned_cols=53  Identities=17%  Similarity=0.225  Sum_probs=33.6

Q ss_pred             eEEEEEEEEecCCCCcceEEE--EECCCCcEEEEEcCEEE-Ee----ecCceEEEEEEEEE
Q 030998           62 SFTFKRVLTNVADTRSTYTAA--VKAPVGMTVTVEPATLS-FA----GKFSKAEFSLTVNI  115 (167)
Q Consensus        62 ~~tv~RTVTNVg~~~stY~a~--V~~P~gv~V~V~P~~L~-F~----~~gqk~sf~Vt~~~  115 (167)
                      ..+|+=+|||.|+......+-  +..|.+- +..-+..|. |.    ++||++..++++..
T Consensus       685 ~~~v~v~VtNtG~~~G~evvQlYv~~p~~~-~~~P~k~L~gF~kv~L~~Ges~~V~~~l~~  744 (779)
T PLN03080        685 RFNVHISVSNVGEMDGSHVVMLFSRSPPVV-PGVPEKQLVGFDRVHTASGRSTETEIVVDP  744 (779)
T ss_pred             eEEEEEEEEECCcccCcEEEEEEEecCccC-CCCcchhccCcEeEeeCCCCEEEEEEEeCc
Confidence            477889999999866655544  4556431 111223443 33    67899988888865


No 19 
>PF12970 DUF3858:  Domain of Unknown Function with PDB structure (DUF3858);  InterPro: IPR024544 This domain of unknown function is structurally similar to part of neuropilin-2. The proteins it occurs in have not yet been functionally characterised.; PDB: 3KD4_A.
Probab=30.78  E-value=87  Score=23.85  Aligned_cols=36  Identities=25%  Similarity=0.432  Sum_probs=20.3

Q ss_pred             CcceEEEEECCCCcEEEEEcCEEEEeecCceEEEEE
Q 030998           76 RSTYTAAVKAPVGMTVTVEPATLSFAGKFSKAEFSL  111 (167)
Q Consensus        76 ~stY~a~V~~P~gv~V~V~P~~L~F~~~gqk~sf~V  111 (167)
                      .++|+-.|..|+|.++...|..-..+.+-.|.+++|
T Consensus        36 ~E~ytyti~~pegm~l~t~~~~K~I~N~~Gk~~isv   71 (116)
T PF12970_consen   36 DETYTYTIELPEGMKLVTPPMEKKIDNPVGKVSISV   71 (116)
T ss_dssp             EEEEEEEEEE-TT-EE-S--S-EEEEETTEEEEEEE
T ss_pred             CcceEEEEEcCCCCeeecCccceeccCCcceEEEEE
Confidence            357888888899999988887666655544544433


No 20 
>PF15496 DUF4646:  Domain of unknown function (DUF4646)
Probab=30.78  E-value=26  Score=26.47  Aligned_cols=16  Identities=19%  Similarity=0.486  Sum_probs=13.6

Q ss_pred             eeCChhhHHHhhhcCC
Q 030998            7 YDIEIQDYLNYLCAMN   22 (167)
Q Consensus         7 YD~~~~DYi~fLC~lg   22 (167)
                      ||.+++||..||-.+.
T Consensus        45 ~DVs~eDW~~F~~dl~   60 (123)
T PF15496_consen   45 HDVSEEDWTRFLNDLS   60 (123)
T ss_pred             cCCCHHHHHHHHHHHH
Confidence            7999999999986543


No 21 
>PF07610 DUF1573:  Protein of unknown function (DUF1573);  InterPro: IPR011467 These hypothetical proteins from bacteria, such as Rhodopirellula baltica, Bacteroides thetaiotaomicron and Porphyromonas gingivalis, share a region of conserved sequence towards their N termini.
Probab=30.76  E-value=1.3e+02  Score=18.33  Aligned_cols=12  Identities=8%  Similarity=0.030  Sum_probs=8.9

Q ss_pred             ecCceEEEEEEE
Q 030998          102 GKFSKAEFSLTV  113 (167)
Q Consensus       102 ~~gqk~sf~Vt~  113 (167)
                      ++||+...+|++
T Consensus        34 ~PGes~~i~v~y   45 (45)
T PF07610_consen   34 APGESGKIKVTY   45 (45)
T ss_pred             CCCCEEEEEEEC
Confidence            678888777764


No 22 
>PF00635 Motile_Sperm:  MSP (Major sperm protein) domain;  InterPro: IPR000535 Major sperm proteins (MSP) are central components in molecular interactions underlying sperm motility in Caenorhabditis elegans, whose sperm employ an amoebae-like crawling motion using a MSP-containing lamellipod, rather than the flagellar-based swimming motion associated with other sperm. These proteins oligomerise to form an extensive filament system that extends from sperm villipoda, along the leading edge of the pseudopod. About 30 MSP isoforms may exist in C. elegans. MSPs form a fibrous network, whereby MSP dimers form helical subfilaments that coil around one another to produce filaments, which in turn form supercoils to produce bundles. The crystal structure of MSP from C. elegans reveals an immunoglobulin (Ig)-like seven-stranded beta sandwich fold []. ; GO: 0005198 structural molecule activity; PDB: 1MSP_A 3MSP_B 2BVU_B 2MSP_C 1Z9O_F 1Z9L_A 3IKK_A 1WIC_A 2CRI_A 2RR3_A ....
Probab=30.34  E-value=1.8e+02  Score=20.06  Aligned_cols=53  Identities=19%  Similarity=0.237  Sum_probs=30.8

Q ss_pred             EEEEEcC-EEEEeecC-ceEEEEEEEEEccccccCCCCCccceEEEEEEEeCCCceEEEceEEEE
Q 030998           90 TVTVEPA-TLSFAGKF-SKAEFSLTVNINLGNAFSPKSNFLGNFGYLTWYEVKRKHTVRSPIVAA  152 (167)
Q Consensus        90 ~V~V~P~-~L~F~~~g-qk~sf~Vt~~~~~~~~~~~~~~~~~~fGsl~W~d~~g~h~VRSPIaV~  152 (167)
                      ++.|.|. .|.|.... +...-.+++.....      ...  +|---+...  ..+.|+-+.++-
T Consensus         1 ~l~v~P~~~i~F~~~~~~~~~~~l~l~N~s~------~~i--~fKiktt~~--~~y~v~P~~G~i   55 (109)
T PF00635_consen    1 DLSVEPSELIFFNAPFNKQQSCELTLTNPSD------KPI--AFKIKTTNP--NRYRVKPSYGII   55 (109)
T ss_dssp             -CEEESSSEEEEESSTSS-EEEEEEEEE-SS------SEE--EEEEEES-T--TTEEEESSEEEE
T ss_pred             CeEEeCCcceEEcCCCCceEEEEEEEECCCC------CcE--EEEEEcCCC--ceEEecCCCEEE
Confidence            3678999 88887754 55666667765542      233  554443333  567787776654


No 23 
>PF08310 LGFP:  LGFP repeat;  InterPro: IPR013207 This 54 amino acid repeat is found in many hypothetical proteins. Several hypothetical proteins from Corynebacterium glutamicum (Brevibacterium flavum) and Corynebacterium efficiens along with PS1 protein contain this repeat region. The N-terminal region of PS1 contains an esterase domain which transfers corynomycolic acid. The C-terminal region consists of 4 tandem LGFP repeats. It is hypothesised that the PS1 proteins in Corynebacterium, when associated with the cell wall, may be anchored via the LGFP tandem repeats that may be important for maintaining cell wall integrity. Deletion of Q01377 from SWISSPROT protein results in a 10-fold increase in the cell volume of the organism and infers the corresponding involvement of the protein in the cell shape formation []. The secondary structure of each repeat is predicted to comprise two beta-strands and one alpha-helix.
Probab=21.54  E-value=47  Score=21.14  Aligned_cols=24  Identities=21%  Similarity=0.470  Sum_probs=18.2

Q ss_pred             eEEEEEEEeCCCceEEEceEEEEE
Q 030998          130 NFGYLTWYEVKRKHTVRSPIVAAF  153 (167)
Q Consensus       130 ~fGsl~W~d~~g~h~VRSPIaV~~  153 (167)
                      .-|.|.|+...|.|.|+-+|.-++
T Consensus        22 ~~G~Iywsp~tGa~~v~G~I~~~w   45 (54)
T PF08310_consen   22 QNGTIYWSPATGAHAVHGAILDKW   45 (54)
T ss_pred             CCeEEEEeCCCCcEEECHHHHHHH
Confidence            569999998656799987765444


Done!