Query         026259
Match_columns 241
No_of_seqs    159 out of 281
Neff          3.7 
Searched_HMMs 46136
Date          Fri Mar 29 05:40:41 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/026259.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/026259hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF14379 Myb_CC_LHEQLE:  MYB-CC  99.9 3.8E-27 8.3E-32  167.5   7.1   51   71-121     1-51  (51)
  2 PLN03162 golden-2 like transcr  98.7 3.5E-09 7.6E-14  101.5   1.6   26    1-26    266-291 (526)
  3 TIGR01557 myb_SHAQKYF myb-like  98.0 2.1E-06 4.7E-11   62.1   1.4   23    2-24     34-56  (57)
  4 PF14379 Myb_CC_LHEQLE:  MYB-CC  94.1    0.13 2.8E-06   37.2   5.0   35   86-121     6-40  (51)
  5 PF15235 GRIN_C:  G protein-reg  80.3     1.3 2.7E-05   37.9   2.2   19   93-111    71-89  (137)
  6 cd07645 I-BAR_IMD_BAIAP2L1 Inv  51.1 1.5E+02  0.0032   27.4   9.3   70   71-143    63-141 (226)
  7 PF01519 DUF16:  Protein of unk  49.7      91   0.002   25.6   7.0   25   94-118    68-92  (102)
  8 cd07646 I-BAR_IMD_IRSp53 Inver  48.8 1.8E+02  0.0038   27.1   9.5   70   71-143    65-143 (232)
  9 PF06548 Kinesin-related:  Kine  35.5 5.1E+02   0.011   26.6  11.1   64   72-135   294-372 (488)
 10 PF00435 Spectrin:  Spectrin re  31.9 1.8E+02  0.0039   20.4   7.2   49   94-145    42-90  (105)
 11 PRK10803 tol-pal system protei  31.2 1.6E+02  0.0035   26.8   6.4   43   79-121    54-96  (263)
 12 KOG2620 Prohibitins and stomat  30.5 2.5E+02  0.0054   27.0   7.6   51   68-118   153-209 (301)
 13 KOG0994 Extracellular matrix g  29.6 3.6E+02  0.0079   31.1   9.6   73   74-147  1410-1483(1758)
 14 PF00752 XPG_N:  XPG N-terminal  28.2      24 0.00052   26.8   0.5   14    4-18      1-14  (101)
 15 KOG4466 Component of histone d  27.0 4.1E+02  0.0089   25.5   8.4   38   94-137    69-106 (291)
 16 KOG1916 Nuclear protein, conta  22.9 4.5E+02  0.0098   29.6   8.7   44   80-123   898-951 (1283)
 17 PF08898 DUF1843:  Domain of un  22.6 1.5E+02  0.0032   21.8   3.7   34  108-141    18-51  (53)
 18 COG5665 NOT5 CCR4-NOT transcri  22.6   7E+02   0.015   25.4   9.4   52   74-131   117-168 (548)
 19 KOG0804 Cytoplasmic Zn-finger   22.2   5E+02   0.011   26.7   8.4   46   73-119   321-366 (493)
 20 COG5420 Uncharacterized conser  22.0 3.8E+02  0.0082   20.7   6.5   47   83-141    10-68  (71)
 21 PF03816 LytR_cpsA_psr:  Cell e  21.4      78  0.0017   26.1   2.3   18  100-117   131-148 (149)
 22 PF01815 Rop:  Rop protein;  In  20.8      87  0.0019   23.5   2.2   21   85-105    38-59  (60)
 23 PF00517 GP41:  Retroviral enve  20.4 3.3E+02  0.0071   24.1   6.2   39   73-111    22-64  (204)
 24 PF07889 DUF1664:  Protein of u  20.4 5.2E+02   0.011   21.7   7.1   52   94-145    62-113 (126)

No 1  
>PF14379 Myb_CC_LHEQLE:  MYB-CC type transfactor, LHEQLE motif
Probab=99.94  E-value=3.8e-27  Score=167.51  Aligned_cols=51  Identities=88%  Similarity=1.125  Sum_probs=49.1

Q ss_pred             CccHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHhHHHHHHHHHHHHHhh
Q 026259           71 GYQVTEALRVQMEVQRRLHEQLEVQRRLQLRIEAQGKYLQSILEKACKALN  121 (241)
Q Consensus        71 ~~qI~EALrmQmEVQrRLhEQLEvQR~LQlRIEaQGkYLq~iLEKAqe~La  121 (241)
                      |++|+||||+||||||||||||||||+||+|||||||||++|||||+++++
T Consensus         1 g~~i~EALr~QmEvQrrLhEQLEvQr~Lqlrieaqgkyl~~ilek~~~~~s   51 (51)
T PF14379_consen    1 GMQITEALRMQMEVQRRLHEQLEVQRHLQLRIEAQGKYLQSILEKAQKALS   51 (51)
T ss_pred             CCcHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHhhHHHHHHHHHHHHhcC
Confidence            578999999999999999999999999999999999999999999999874


No 2  
>PLN03162 golden-2 like transcription factor; Provisional
Probab=98.72  E-value=3.5e-09  Score=101.50  Aligned_cols=26  Identities=46%  Similarity=0.676  Sum_probs=24.4

Q ss_pred             CcccCCCCchHHHHHHHhhhhhhccc
Q 026259            1 MRTMGVKGLTLYHLKSHLQKYRLGKQ   26 (241)
Q Consensus         1 lRLMGVkGLTLYHLKSHLQKYRLgK~   26 (241)
                      |++|+|+|||++||||||||||+.++
T Consensus       266 LelMnV~GLTRenVKSHLQKYRl~rk  291 (526)
T PLN03162        266 LELMGVQCLTRHNIASHLQKYRSHRR  291 (526)
T ss_pred             HHHcCCCCcCHHHHHHHHHHHHHhcc
Confidence            57999999999999999999999875


No 3  
>TIGR01557 myb_SHAQKYF myb-like DNA-binding domain, SHAQKYF class. This model describes a DNA-binding domain restricted to (but common in) plant proteins, many of which also contain a response regulator domain. The domain appears related to the Myb-like DNA-binding domain described by Pfam model pfam00249. It is distinguished in part by a well-conserved motif SH[AL]QKY[RF] at the C-terminal end of the motif.
Probab=98.01  E-value=2.1e-06  Score=62.07  Aligned_cols=23  Identities=57%  Similarity=0.699  Sum_probs=21.3

Q ss_pred             cccCCCCchHHHHHHHhhhhhhc
Q 026259            2 RTMGVKGLTLYHLKSHLQKYRLG   24 (241)
Q Consensus         2 RLMGVkGLTLYHLKSHLQKYRLg   24 (241)
                      .+|++++||..||+|||||||+-
T Consensus        34 ~~~~~~~lT~~qV~SH~QKy~~k   56 (57)
T TIGR01557        34 ELMVVDGLTRDQVASHLQKYRLK   56 (57)
T ss_pred             HHcCCCCCCHHHHHHHHHHHHcc
Confidence            57999999999999999999984


No 4  
>PF14379 Myb_CC_LHEQLE:  MYB-CC type transfactor, LHEQLE motif
Probab=94.12  E-value=0.13  Score=37.17  Aligned_cols=35  Identities=43%  Similarity=0.499  Sum_probs=27.3

Q ss_pred             HHHHHHHHHHHHHHHHHHHHhHHHHHHHHHHHHHhh
Q 026259           86 RRLHEQLEVQRRLQLRIEAQGKYLQSILEKACKALN  121 (241)
Q Consensus        86 rRLhEQLEvQR~LQlRIEaQGkYLq~iLEKAqe~La  121 (241)
                      --|..|+||||+|.=.+|.| |-|+.=+|..-+-|.
T Consensus         6 EALr~QmEvQrrLhEQLEvQ-r~Lqlrieaqgkyl~   40 (51)
T PF14379_consen    6 EALRMQMEVQRRLHEQLEVQ-RHLQLRIEAQGKYLQ   40 (51)
T ss_pred             HHHHHHHHHHHHHHHHHHHH-HHHHHHHHHhhHHHH
Confidence            45789999999999999999 677766666655543


No 5  
>PF15235 GRIN_C:  G protein-regulated inducer of neurite outgrowth C-terminus
Probab=80.28  E-value=1.3  Score=37.87  Aligned_cols=19  Identities=21%  Similarity=0.371  Sum_probs=16.6

Q ss_pred             HHHHHHHHHHHHHhHHHHH
Q 026259           93 EVQRRLQLRIEAQGKYLQS  111 (241)
Q Consensus        93 EvQR~LQlRIEaQGkYLq~  111 (241)
                      -||+||+++||+|++....
T Consensus        71 AIQkHLE~qi~e~~~q~~~   89 (137)
T PF15235_consen   71 AIQKHLERQIEEHERQRAP   89 (137)
T ss_pred             HHHHHHHHHHHHhhhcccc
Confidence            4899999999999988754


No 6  
>cd07645 I-BAR_IMD_BAIAP2L1 Inverse (I)-BAR, also known as the IRSp53/MIM homology Domain (IMD), of Brain-specific Angiogenesis Inhibitor 1-Associated Protein 2-Like 1. The IMD domain, also called Inverse-Bin/Amphiphysin/Rvs (I-BAR) domain, is a dimerization and lipid-binding module that bends membranes and induces membrane protrusions. BAIAP2L1 (Brain-specific Angiogenesis Inhibitor 1-Associated Protein 2-Like 1) is also known as IRTKS (Insulin Receptor Tyrosine Kinase Substrate). It is widely expressed, serves as a substrate for the insulin receptor, and binds the small GTPase Rac. It plays a role in regulating the actin cytoskeleton and colocalizes with F-actin, cortactin, VASP, and vinculin. BAIAP2L1 expression leads to the formation of short actin bundles, distinct from filopodia-like protrusions induced by the expression of the related protein IRSp53. It contains an N-terminal IMD, an SH3 domain, and a WASP homology 2 (WH2) actin-binding motif at the C-terminus. The IMD domain of 
Probab=51.05  E-value=1.5e+02  Score=27.44  Aligned_cols=70  Identities=19%  Similarity=0.293  Sum_probs=54.6

Q ss_pred             CccHHHHHHHHHHHHHHHHHHHHH---------HHHHHHHHHHHhHHHHHHHHHHHHHhhhhhhhhhchHHHHHHHHHHH
Q 026259           71 GYQVTEALRVQMEVQRRLHEQLEV---------QRRLQLRIEAQGKYLQSILEKACKALNDQAIVAAGLEAAREELSELA  141 (241)
Q Consensus        71 ~~qI~EALrmQmEVQrRLhEQLEv---------QR~LQlRIEaQGkYLq~iLEKAqe~La~~~~~~~glEaak~eLseL~  141 (241)
                      +..|.++|.-=-||+|+++.|||.         =..|.-.+|..-||+...+.+=+..   +-.-..+||-+.++|--+-
T Consensus        63 SkeLG~~L~qi~ev~r~i~~~le~~lK~Fh~Ell~~LE~k~elD~kyi~a~~Kkyq~E---~k~k~dsLeK~~seLKK~R  139 (226)
T cd07645          63 SKELGHVLMEISDVHKKLNDSLEENFKKFHREIIAELERKTDLDVKYMTATLKRYQTE---HKNKLDSLEKSQADLKKIR  139 (226)
T ss_pred             chHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH---HHHHHHHHHHHHHHHHHHH
Confidence            456778885445999999998873         3578999999999999988875443   4455678999999988887


Q ss_pred             HH
Q 026259          142 IK  143 (241)
Q Consensus       142 s~  143 (241)
                      -+
T Consensus       140 RK  141 (226)
T cd07645         140 RK  141 (226)
T ss_pred             hc
Confidence            66


No 7  
>PF01519 DUF16:  Protein of unknown function DUF16;  InterPro: IPR002862 Proteins that contain this domain are of unknown function. It appears to be confined to proteins from Mycoplasma pneumoniae [].; PDB: 2BA2_C.
Probab=49.68  E-value=91  Score=25.57  Aligned_cols=25  Identities=40%  Similarity=0.419  Sum_probs=20.2

Q ss_pred             HHHHHHHHHHHHhHHHHHHHHHHHH
Q 026259           94 VQRRLQLRIEAQGKYLQSILEKACK  118 (241)
Q Consensus        94 vQR~LQlRIEaQGkYLq~iLEKAqe  118 (241)
                      .=+.||.+|.+||+-|++|++.-+.
T Consensus        68 qIkel~~e~k~qgktL~~I~~~L~~   92 (102)
T PF01519_consen   68 QIKELQVEQKAQGKTLQLILKTLQS   92 (102)
T ss_dssp             HHHHHHHHHHHHHHHHHHHHHHHHH
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHHHH
Confidence            3378999999999999999875443


No 8  
>cd07646 I-BAR_IMD_IRSp53 Inverse (I)-BAR, also known as the IRSp53/MIM homology Domain (IMD), of Insulin Receptor tyrosine kinase Substrate p53. The IMD domain, also called Inverse-Bin/Amphiphysin/Rvs (I-BAR) domain, is a dimerization and lipid-binding module that bends membranes and induces membrane protrusions. IRSp53 (Insulin Receptor tyrosine kinase Substrate p53) is also known as BAIAP2 (Brain-specific Angiogenesis Inhibitor 1-Associated Protein 2). It is a scaffolding protein that takes part in many signaling pathways including Cdc42-induced filopodia formation, Rac-mediated lamellipodia extension, and spine morphogenesis. IRSp53 exists as multiple splicing variants that differ mainly at the C-termini. One variant (T-form) is expressed exclusively in human breast cancer cells. The gene encoding IRSp53 is a putative susceptibility gene for Gilles de la Tourette syndrome. IRSp53 contains an N-terminal IMD, a CRIB (Cdc42 and Rac interactive binding motif), an SH3 domain, and a WASP 
Probab=48.78  E-value=1.8e+02  Score=27.07  Aligned_cols=70  Identities=27%  Similarity=0.392  Sum_probs=53.4

Q ss_pred             CccHHHHHHHHHHHHHHHHHHHHHH---------HHHHHHHHHHhHHHHHHHHHHHHHhhhhhhhhhchHHHHHHHHHHH
Q 026259           71 GYQVTEALRVQMEVQRRLHEQLEVQ---------RRLQLRIEAQGKYLQSILEKACKALNDQAIVAAGLEAAREELSELA  141 (241)
Q Consensus        71 ~~qI~EALrmQmEVQrRLhEQLEvQ---------R~LQlRIEaQGkYLq~iLEKAqe~La~~~~~~~glEaak~eLseL~  141 (241)
                      +..|..||.-=-||+|.++.+||++         ..|+.++|..-|||...+.+=+-.   +-.-..++|-+++||-.|-
T Consensus        65 SkeLG~~L~~m~~~hr~i~~~le~~lk~Fh~eli~pLE~k~E~D~k~i~a~~Kky~~e---~k~k~~sleK~qseLKKlR  141 (232)
T cd07646          65 SKELGDVLFQMAEVHRQIQNQLEEMLKSFHNELLTQLEQKVELDSRYLTAALKKYQTE---HRSKGESLEKCQAELKKLR  141 (232)
T ss_pred             chHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH---HHHHHHHHHHHHHHHHHHH
Confidence            4567788855458888888888744         479999999999999876665443   4455678999999998877


Q ss_pred             HH
Q 026259          142 IK  143 (241)
Q Consensus       142 s~  143 (241)
                      -+
T Consensus       142 rK  143 (232)
T cd07646         142 KK  143 (232)
T ss_pred             Hh
Confidence            55


No 9  
>PF06548 Kinesin-related:  Kinesin-related;  InterPro: IPR010544 This entry represents a domain within kinesin-related proteins from higher plants. Many proteins containing this domain also contain the IPR001752 from INTERPRO domain. Kinesins are ATP-driven microtubule motor proteins that produce directed force []. Some family members are associated with the phragmoplast, a structure composed mainly of microtubules that executes cytokinesis in higher plants [].
Probab=35.46  E-value=5.1e+02  Score=26.59  Aligned_cols=64  Identities=31%  Similarity=0.475  Sum_probs=43.0

Q ss_pred             ccHHHHHHHHHHHHHHHHHHHH------------HHHHHHHHHHHHhHHHHHHH---HHHHHHhhhhhhhhhchHHHHH
Q 026259           72 YQVTEALRVQMEVQRRLHEQLE------------VQRRLQLRIEAQGKYLQSIL---EKACKALNDQAIVAAGLEAARE  135 (241)
Q Consensus        72 ~qI~EALrmQmEVQrRLhEQLE------------vQR~LQlRIEaQGkYLq~iL---EKAqe~La~~~~~~~glEaak~  135 (241)
                      +.++|-||+-+|..|.|-|-+|            ++--||.-|+-|.|.|.---   ||--.-++.|...-.||+-.|.
T Consensus       294 IsLteeLR~dle~~r~~aek~~~EL~~Ek~c~eEL~~al~~A~~GhaR~lEqYadLqEk~~~Ll~~Hr~i~egI~dVKk  372 (488)
T PF06548_consen  294 ISLTEELRVDLESSRSLAEKLEMELDSEKKCTEELDDALQRAMEGHARMLEQYADLQEKHNDLLARHRRIMEGIEDVKK  372 (488)
T ss_pred             hhhHHHHHHHHHHHHHHHHHHHHHHHHHHHhHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence            6688999999999999888765            55678888888877765432   2222334555555556654433


No 10 
>PF00435 Spectrin:  Spectrin repeat;  InterPro: IPR002017 Spectrin repeats [] are found in several proteins involved in cytoskeletal structure. These include spectrin alpha and beta subunits [, ], alpha-actinin [] and dystrophin. The spectrin repeat forms a three-helix bundle. The second helix is interrupted by proline in some sequences. The repeats are defined by a characteristic tryptophan (W) residue at position 17 in helix A and a leucine (L) at 2 residues from the carboxyl end of helix C.; GO: 0005515 protein binding; PDB: 1HCI_A 1QUU_A 3FB2_B 1S35_A 1U5P_A 1U4Q_A 1CUN_B 1YDI_B 3EDV_A 1AJ3_A ....
Probab=31.91  E-value=1.8e+02  Score=20.36  Aligned_cols=49  Identities=24%  Similarity=0.338  Sum_probs=30.0

Q ss_pred             HHHHHHHHHHHHhHHHHHHHHHHHHHhhhhhhhhhchHHHHHHHHHHHHHhh
Q 026259           94 VQRRLQLRIEAQGKYLQSILEKACKALNDQAIVAAGLEAAREELSELAIKVS  145 (241)
Q Consensus        94 vQR~LQlRIEaQGkYLq~iLEKAqe~La~~~~~~~glEaak~eLseL~s~v~  145 (241)
                      -.+.++-.|..+..-+..|.+.++.-....   +..-..-+..+.+|.....
T Consensus        42 ~~~~~~~ei~~~~~~l~~l~~~~~~L~~~~---~~~~~~i~~~~~~l~~~w~   90 (105)
T PF00435_consen   42 KHKELQEEIESRQERLESLNEQAQQLIDSG---PEDSDEIQEKLEELNQRWE   90 (105)
T ss_dssp             HHHHHHHHHHHHHHHHHHHHHHHHHHHHTT---HTTHHHHHHHHHHHHHHHH
T ss_pred             HHhhhhhHHHHHHHHHHHHHHHHHHHHHcC---CCcHHHHHHHHHHHHHHHH
Confidence            334455566777777888888877774433   3344555666666665543


No 11 
>PRK10803 tol-pal system protein YbgF; Provisional
Probab=31.20  E-value=1.6e+02  Score=26.85  Aligned_cols=43  Identities=14%  Similarity=0.202  Sum_probs=31.8

Q ss_pred             HHHHHHHHHHHHHHHHHHHHHHHHHHHhHHHHHHHHHHHHHhh
Q 026259           79 RVQMEVQRRLHEQLEVQRRLQLRIEAQGKYLQSILEKACKALN  121 (241)
Q Consensus        79 rmQmEVQrRLhEQLEvQR~LQlRIEaQGkYLq~iLEKAqe~La  121 (241)
                      ++|.|+|.+|.+.-.==+.|.=.||.+..-|+.|.+++.+--.
T Consensus        54 ~~~~~l~~ql~~lq~ev~~LrG~~E~~~~~l~~~~~rq~~~y~   96 (263)
T PRK10803         54 QLLTQLQQQLSDNQSDIDSLRGQIQENQYQLNQVVERQKQIYL   96 (263)
T ss_pred             HHHHHHHHHHHHHHHHHHHHhhHHHHHHHHHHHHHHHHHHHHH
Confidence            4567888888664333356788899999999999998777543


No 12 
>KOG2620 consensus Prohibitins and stomatins of the PID superfamily [Energy production and conversion]
Probab=30.53  E-value=2.5e+02  Score=27.00  Aligned_cols=51  Identities=24%  Similarity=0.158  Sum_probs=38.8

Q ss_pred             CCCCccHHHHHHHHHHHHHHHHHHH---HHHHHHHHHH---HHHhHHHHHHHHHHHH
Q 026259           68 PNDGYQVTEALRVQMEVQRRLHEQL---EVQRRLQLRI---EAQGKYLQSILEKACK  118 (241)
Q Consensus        68 ~~~~~qI~EALrmQmEVQrRLhEQL---EvQR~LQlRI---EaQGkYLq~iLEKAqe  118 (241)
                      ..-.-++.+|.+||-|.+|+=.-++   |--|.+|+.+   |++.|||.+.=.+++.
T Consensus       153 I~pp~~V~~AM~~q~~AeR~krAailesEger~~~InrAEGek~s~iL~seg~~~qr  209 (301)
T KOG2620|consen  153 IEPPPSVKRAMNMQNEAERMKRAAILESEGERIAQINRAEGEKESKILASEGIARQR  209 (301)
T ss_pred             cCCCHHHHHHHHHHHHHHHHHHHHHhhhhhhhHHhhhhhcchhhhHHhhhHHHHHHH
Confidence            3334578999999999999766553   4778888877   6889999887666554


No 13 
>KOG0994 consensus Extracellular matrix glycoprotein Laminin subunit beta [Extracellular structures]
Probab=29.60  E-value=3.6e+02  Score=31.10  Aligned_cols=73  Identities=22%  Similarity=0.244  Sum_probs=51.3

Q ss_pred             HHHHHHHHHHHHHHHHHHH-HHHHHHHHHHHHHhHHHHHHHHHHHHHhhhhhhhhhchHHHHHHHHHHHHHhhcC
Q 026259           74 VTEALRVQMEVQRRLHEQL-EVQRRLQLRIEAQGKYLQSILEKACKALNDQAIVAAGLEAAREELSELAIKVSND  147 (241)
Q Consensus        74 I~EALrmQmEVQrRLhEQL-EvQR~LQlRIEaQGkYLq~iLEKAqe~La~~~~~~~glEaak~eLseL~s~v~~~  147 (241)
                      -.+||.+=++++.+|.+-+ |+++-|++--||-- --...-++|+++|..-+.+..-.+.+.++|.+|...|.++
T Consensus      1410 A~~A~~~A~~~~~~l~~~~ae~eq~~~~v~ea~~-~aseA~~~Aq~~~~~a~as~~q~~~s~~el~~Li~~v~~F 1483 (1758)
T KOG0994|consen 1410 AGGALLMAGDADTQLRSKLAEAEQTLSMVREAKL-SASEAQQSAQRALEQANASRSQMEESNRELRNLIQQVRDF 1483 (1758)
T ss_pred             cchHHHHhhhHHHHHHHHHHHHHHHHHHHHHHHH-hhHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence            3478888888877776643 57777766544432 2224556777777777777777889999999999988863


No 14 
>PF00752 XPG_N:  XPG N-terminal domain;  InterPro: IPR006085 Xeroderma pigmentosum (XP) [] is a human autosomal recessive disease, characterised by a high incidence of sunlight-induced skin cancer. People's skin cells with this condition are hypersensitive to ultraviolet light, due to defects in the incision step of DNA excision repair. There are a minimum of seven genetic complementation groups involved in this pathway: XP-A to XP-G. XP-G is one of the most rare and phenotypically heterogeneous of XP, showing anything from slight to extreme dysfunction in DNA excision repair [, ]. XP-G can be corrected by a 133 Kd nuclear protein, XPGC []. XPGC is an acidic protein that confers normal UV resistance in expressing cells []. It is a magnesium-dependent, single-strand DNA endonuclease that makes structure-specific endonucleolytic incisions in a DNA substrate containing a duplex region and single-stranded arms [, ]. XPGC cleaves one strand of the duplex at the border with the single-stranded region []. XPG belongs to a family of proteins that includes RAD2 from Saccharomyces cerevisiae (Baker's yeast) and rad13 from Schizosaccharomyces pombe (Fission yeast), which are single-stranded DNA endonucleases [, ]; mouse and human FEN-1, a structure-specific endonuclease; RAD2 from fission yeast and RAD27 from budding yeast; fission yeast exo1, a 5'-3' double-stranded DNA exonuclease that may act in a pathway that corrects mismatched base pairs; yeast DHS1, and yeast DIN7. Sequence alignment of this family of proteins reveals that similarities are largely confined to two regions. The first is located at the N-terminal extremity (N-region) and corresponds to the first 95 to 105 amino acids. The second region is internal (I-region) and found towards the C terminus; it spans about 140 residues and contains a highly conserved core of 27 amino acids that includes a conserved pentapeptide (E-A-[DE]-A-[QS]). It is possible that the conserved acidic residues are involved in the catalytic mechanism of DNA excision repair in XPG. The amino acids linking the N- and I-regions are not conserved. This entry represents the N-terminal of XPG.; GO: 0004518 nuclease activity, 0006281 DNA repair; PDB: 1A77_A 1A76_A 1MC8_B 3QEB_Z 3QEA_Z 3QE9_Y 1UL1_Z 3Q8K_A 3Q8M_A 3Q8L_A ....
Probab=28.19  E-value=24  Score=26.75  Aligned_cols=14  Identities=50%  Similarity=0.674  Sum_probs=7.3

Q ss_pred             cCCCCchHHHHHHHh
Q 026259            4 MGVKGLTLYHLKSHL   18 (241)
Q Consensus         4 MGVkGLTLYHLKSHL   18 (241)
                      |||+||+-| +|.+.
T Consensus         1 MGI~gL~~~-l~~~~   14 (101)
T PF00752_consen    1 MGIKGLWQL-LKPAA   14 (101)
T ss_dssp             ---TTHHHH-CHHHE
T ss_pred             CCcccHHHH-HHhhc
Confidence            999999876 44443


No 15 
>KOG4466 consensus Component of histone deacetylase complex (breast carcinoma metastasis suppressor 1 protein in human) [Cell cycle control, cell division, chromosome partitioning; Transcription]
Probab=27.04  E-value=4.1e+02  Score=25.54  Aligned_cols=38  Identities=16%  Similarity=0.295  Sum_probs=28.6

Q ss_pred             HHHHHHHHHHHHhHHHHHHHHHHHHHhhhhhhhhhchHHHHHHH
Q 026259           94 VQRRLQLRIEAQGKYLQSILEKACKALNDQAIVAAGLEAAREEL  137 (241)
Q Consensus        94 vQR~LQlRIEaQGkYLq~iLEKAqe~La~~~~~~~glEaak~eL  137 (241)
                      +|+.++.||+--|.|.+-+++.++.-.-      .-++||++++
T Consensus        69 L~~~~kerl~~aely~e~~~e~v~~eYe------~E~~aAk~e~  106 (291)
T KOG4466|consen   69 LDESRKERLRVAELYREYCVERVEREYE------CEIKAAKKEY  106 (291)
T ss_pred             HHHHHHHHHHHHHHHHHHHHHHHHHHHH------HHHHHHHHHH
Confidence            8999999999999999999888776543      2345555553


No 16 
>KOG1916 consensus Nuclear protein, contains WD40 repeats [General function prediction only]
Probab=22.87  E-value=4.5e+02  Score=29.60  Aligned_cols=44  Identities=39%  Similarity=0.414  Sum_probs=30.6

Q ss_pred             HHHHHHHHHHHHHH----------HHHHHHHHHHHHhHHHHHHHHHHHHHhhhh
Q 026259           80 VQMEVQRRLHEQLE----------VQRRLQLRIEAQGKYLQSILEKACKALNDQ  123 (241)
Q Consensus        80 mQmEVQrRLhEQLE----------vQR~LQlRIEaQGkYLq~iLEKAqe~La~~  123 (241)
                      -|-|+|+||.-||+          |.|-|.-+-+|.-+-|+.-|-|-++++.++
T Consensus       898 sQ~el~~~l~~ql~g~le~~l~~~iEk~lks~~d~~~~rl~e~la~~e~~~r~~  951 (1283)
T KOG1916|consen  898 SQKELQRQLSNQLTGPLEVALGRMIEKSLKSNADALWARLQEELAKNEKALRDL  951 (1283)
T ss_pred             hHHHHHHHHHHhhcchHHHHHHHHHHHHHHhhHHHHHHHHHHHHHhhhhhhhHH
Confidence            36788888888876          456666777777777777776666655543


No 17 
>PF08898 DUF1843:  Domain of unknown function (DUF1843);  InterPro: IPR014994 This domain is found in functionally uncharacterised proteins. It can be found independently or at the C terminus of the protein. 
Probab=22.65  E-value=1.5e+02  Score=21.81  Aligned_cols=34  Identities=24%  Similarity=0.395  Sum_probs=28.0

Q ss_pred             HHHHHHHHHHHHhhhhhhhhhchHHHHHHHHHHH
Q 026259          108 YLQSILEKACKALNDQAIVAAGLEAAREELSELA  141 (241)
Q Consensus       108 YLq~iLEKAqe~La~~~~~~~glEaak~eLseL~  141 (241)
                      .|++++-.|.+.|+.+..-...++..++|...|.
T Consensus        18 ~MK~l~~~aeq~L~~~~~i~~al~~Lk~EIaklE   51 (53)
T PF08898_consen   18 QMKALAAQAEQQLAEAGDIAAALEKLKAEIAKLE   51 (53)
T ss_pred             HHHHHHHHHHHHHccchHHHHHHHHHHHHHHHHh
Confidence            4677888999999988877788888898887764


No 18 
>COG5665 NOT5 CCR4-NOT transcriptional regulation complex, NOT5 subunit [Transcription]
Probab=22.60  E-value=7e+02  Score=25.45  Aligned_cols=52  Identities=27%  Similarity=0.376  Sum_probs=35.8

Q ss_pred             HHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHhHHHHHHHHHHHHHhhhhhhhhhchH
Q 026259           74 VTEALRVQMEVQRRLHEQLEVQRRLQLRIEAQGKYLQSILEKACKALNDQAIVAAGLE  131 (241)
Q Consensus        74 I~EALrmQmEVQrRLhEQLEvQR~LQlRIEaQGkYLq~iLEKAqe~La~~~~~~~glE  131 (241)
                      |..||   -|+||+| ||+|.| .|.-|||.| ++-+.-||---+-|...-..+..++
T Consensus       117 i~~~~---~el~~q~-e~~ea~-e~e~~~erh-~~h~~~le~i~~~l~n~~~~pe~v~  168 (548)
T COG5665         117 IHDCL---DELQKQL-EQYEAQ-ENEEQTERH-EFHIANLENILKKLQNNEMDPEPVE  168 (548)
T ss_pred             HHHHH---HHHHHHH-HHHHHH-HhHHHHHHH-HHHHHHHHHHHHHHhccCCChhhHH
Confidence            55555   4788877 889998 788999988 6666667766666665444444443


No 19 
>KOG0804 consensus Cytoplasmic Zn-finger protein BRAP2 (BRCA1 associated protein) [General function prediction only]
Probab=22.19  E-value=5e+02  Score=26.66  Aligned_cols=46  Identities=26%  Similarity=0.359  Sum_probs=33.1

Q ss_pred             cHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHhHHHHHHHHHHHHH
Q 026259           73 QVTEALRVQMEVQRRLHEQLEVQRRLQLRIEAQGKYLQSILEKACKA  119 (241)
Q Consensus        73 qI~EALrmQmEVQrRLhEQLEvQR~LQlRIEaQGkYLq~iLEKAqe~  119 (241)
                      +-+--|--|+|-||+.-||. +=+-.|-.+|-|..|+...+.+|...
T Consensus       321 ~~s~ll~sqleSqr~y~e~~-~~e~~qsqlen~k~~~e~~~~e~~~l  366 (493)
T KOG0804|consen  321 EYSPLLTSQLESQRKYYEQI-MSEYEQSQLENQKQYYELLITEADSL  366 (493)
T ss_pred             ecchhhhhhhhHHHHHHHHH-HHHHHHHHHHhHHHHHHHHHHHHHhh
Confidence            33446667999999999953 33344568888999998888877663


No 20 
>COG5420 Uncharacterized conserved small protein containing a coiled-coil domain [Function unknown]
Probab=22.00  E-value=3.8e+02  Score=20.69  Aligned_cols=47  Identities=38%  Similarity=0.592  Sum_probs=26.8

Q ss_pred             HHHHHHHHHHHHHHHHHHHHHHH---------h---HHHHHHHHHHHHHhhhhhhhhhchHHHHHHHHHHH
Q 026259           83 EVQRRLHEQLEVQRRLQLRIEAQ---------G---KYLQSILEKACKALNDQAIVAAGLEAAREELSELA  141 (241)
Q Consensus        83 EVQrRLhEQLEvQR~LQlRIEaQ---------G---kYLq~iLEKAqe~La~~~~~~~glEaak~eLseL~  141 (241)
                      |+|+|.       |+||.|--+-         |   +|. .|.+-|-+++.-+    +-|+++|.||..|.
T Consensus        10 eiqkKv-------rkLqsrAg~akm~LhDLAEgLP~~wt-ei~~VA~kt~~~y----aeLD~~k~ELakle   68 (71)
T COG5420          10 EIQKKV-------RKLQSRAGQAKMELHDLAEGLPVKWT-EIMAVAEKTFEAY----AELDAAKRELAKLE   68 (71)
T ss_pred             HHHHHH-------HHHHHHHHHHHhhHHHHhccCCccHH-HHHHHHHHHHHHH----HHHHHHHHHHHHhh
Confidence            567776       6777764322         1   343 2444444444433    36788899887764


No 21 
>PF03816 LytR_cpsA_psr:  Cell envelope-related transcriptional attenuator domain;  InterPro: IPR004474 This entry describes a domain of unknown function that is found in the predicted extracellular domain of a number of putative membrane-bound proteins. One of these is protein psr, described as a penicillin binding protein 5 (PDP-5) synthesis repressor. Another is Bacillus subtilis LytR, described as a transcriptional attenuator of itself and the LytABC operon, where LytC is N-acetylmuramoyl-L-alanine amidase. A third is CpsA, a putative regulatory protein involved in exocellular polysaccharide biosynthesis. These proteins share the property of having a short putative N-terminal cytoplasmic domain and transmembrane domain forming a signal-anchor.; PDB: 3PE5_B 3QFI_A 3NRO_B 3OKZ_B 3OWQ_C 3MEJ_A 4DE9_A 3TEP_A 3TEL_A 3TFL_A ....
Probab=21.39  E-value=78  Score=26.07  Aligned_cols=18  Identities=39%  Similarity=0.434  Sum_probs=15.9

Q ss_pred             HHHHHHhHHHHHHHHHHH
Q 026259          100 LRIEAQGKYLQSILEKAC  117 (241)
Q Consensus       100 lRIEaQGkYLq~iLEKAq  117 (241)
                      -|++.|.+||.++++|+.
T Consensus       131 ~R~~rQ~~~l~al~~k~~  148 (149)
T PF03816_consen  131 GRIQRQQEVLKALLEKLK  148 (149)
T ss_dssp             HHHHHHHHHHHHHHHHHT
T ss_pred             HHHHHHHHHHHHHHHHhh
Confidence            489999999999999974


No 22 
>PF01815 Rop:  Rop protein;  InterPro: IPR000769 The Rop protein regulates plasmid DNA replication by modulating the initiation of transcription of the primer RNA precursor. Processing of the precursor, RNAII, is inhibited by hydrogen bonding of RNAII to its complementary sequence in RNAI. Rop increases the affinity of RNAI for RNAII and thus decreases the rate of replication initiation events. The 3D structure of Rop has been determined by X-ray crystallography and refined to 1.7A resolution. The 63 amino acid protein is a homodimer, each monomer consisting almost entirely of two alpha-helices, the whole molecule forming a highly regular four-alpha-helix bundle []. This can be approximated by a four-stranded rope, with radius 7.0 A, a left-handed helical twist, and pitch 172.5 A. A very compact packing of side chains in the helix interfaces of the Rop coiled-coil structure is presumed to account for its high stability []. The overall details of the structure have been confirmed by proton NMR [, ].; PDB: 1GTO_C 2IJH_A 2IJJ_B 1GMG_A 1ROP_A 1NKD_A 1QX8_A 1F4M_D 2GHY_B 3K79_A ....
Probab=20.77  E-value=87  Score=23.50  Aligned_cols=21  Identities=38%  Similarity=0.526  Sum_probs=16.2

Q ss_pred             HHHHHHHHH-HHHHHHHHHHHH
Q 026259           85 QRRLHEQLE-VQRRLQLRIEAQ  105 (241)
Q Consensus        85 QrRLhEQLE-vQR~LQlRIEaQ  105 (241)
                      =-||||+-| +.++|..|++..
T Consensus        38 CE~LHe~AE~L~~~l~~r~~~e   59 (60)
T PF01815_consen   38 CERLHELAEQLYRSLSARLGEE   59 (60)
T ss_dssp             HHHHHHHHHHHHHHHHHHHT-T
T ss_pred             HHHHHHHHHHHHHHHHHHhccC
Confidence            358999977 789999998764


No 23 
>PF00517 GP41:  Retroviral envelope protein;  InterPro: IPR000328 This entry represents envelope proteins from a variety of retroviruses. It includes the GP41 subunit of the envelope protein complex from Human immunodeficiency virus (HIV) and Simian-Human immunodeficiency virus (SIV), which mediate membrane fusion during viral entry []. It has a core composed of a six-helix bundle and is folded by its trimeric N- and C-terminal heptad-repeats (NHR and CHR) []. Derivatives of this protein prevent HIV-1 from entering cell lines and primary human CD4+ cells in vitro [], making it an attractive subject of gene therapy studies against HIV and related retroviruses. The entry also represents envelop proteins from Bovine immunodeficiency virus, Feline immunodeficiency virus and Equine infectious anemia virus (EIAV) [, ], as well as the Gp36 protein from Mouse mammary tumor virus (MMTV) and Human endogenous retrovirus (HERV).; GO: 0005198 structural molecule activity, 0019031 viral envelope; PDB: 2EZO_B 2EZQ_B 2EZR_A 2JNR_B 1F23_D 2EZP_A 1JEK_A 2Q7C_A 2Q5U_A 2Q3I_A ....
Probab=20.39  E-value=3.3e+02  Score=24.08  Aligned_cols=39  Identities=31%  Similarity=0.321  Sum_probs=27.2

Q ss_pred             cHHHHHHHHHHHHHHHHHHHHHHHH----HHHHHHHHhHHHHH
Q 026259           73 QVTEALRVQMEVQRRLHEQLEVQRR----LQLRIEAQGKYLQS  111 (241)
Q Consensus        73 qI~EALrmQmEVQrRLhEQLEvQR~----LQlRIEaQGkYLq~  111 (241)
                      +.+.+|+.|-..|..|.-++--=++    ||-|+.|=.+||+.
T Consensus        22 ~~~~ll~~~e~~~~lL~l~v~gik~~V~~L~aRV~alE~~l~d   64 (204)
T PF00517_consen   22 QQSNLLRAQEAQQHLLQLTVWGIKQGVKQLQARVLALERYLKD   64 (204)
T ss_dssp             HHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
T ss_pred             HHHHHHHHHHHHHHHhhhhhhhhhhhhhhhhhhHHHHHHHhhh
Confidence            3456777777777777766654444    88888888888743


No 24 
>PF07889 DUF1664:  Protein of unknown function (DUF1664);  InterPro: IPR012458 The members of this family are hypothetical plant proteins of unknown function. The region featured in this family is approximately 100 amino acids long. 
Probab=20.39  E-value=5.2e+02  Score=21.67  Aligned_cols=52  Identities=13%  Similarity=0.186  Sum_probs=32.5

Q ss_pred             HHHHHHHHHHHHhHHHHHHHHHHHHHhhhhhhhhhchHHHHHHHHHHHHHhh
Q 026259           94 VQRRLQLRIEAQGKYLQSILEKACKALNDQAIVAAGLEAAREELSELAIKVS  145 (241)
Q Consensus        94 vQR~LQlRIEaQGkYLq~iLEKAqe~La~~~~~~~glEaak~eLseL~s~v~  145 (241)
                      ..|||..||+.=++-|....|-++.+-.+=+.....++..+.++..+...|.
T Consensus        62 tKkhLsqRId~vd~klDe~~ei~~~i~~eV~~v~~dv~~i~~dv~~v~~~V~  113 (126)
T PF07889_consen   62 TKKHLSQRIDRVDDKLDEQKEISKQIKDEVTEVREDVSQIGDDVDSVQQMVE  113 (126)
T ss_pred             HHHHHHHHHHHHHhhHHHHHHHHHHHHHHHHHHHhhHHHHHHHHHHHHHHHH
Confidence            6789999999988888877776655544333334444444555444444443


Done!