Query         psy17223
Match_columns 323
No_of_seqs    271 out of 586
Neff          5.4 
Searched_HMMs 46136
Date          Fri Aug 16 16:30:26 2013
Command       hhsearch -i /work/01045/syshi/Psyhhblits/psy17223.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/17223hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 KOG3865|consensus              100.0 4.6E-62   1E-66  460.8  15.8  202    4-312    91-313 (402)
  2 KOG3780|consensus               99.7 9.7E-18 2.1E-22  165.3  13.2  125   16-150   101-230 (427)
  3 PF00339 Arrestin_N:  Arrestin   99.0 1.1E-09 2.4E-14   91.9   6.2   42   15-64     88-129 (149)
  4 PF02752 Arrestin_C:  Arrestin   98.6   9E-08   2E-12   78.7   5.3   45  105-149     2-46  (136)
  5 PF13002 LDB19:  Arrestin_N ter  98.0 4.1E-05   9E-10   69.7  10.5  121   13-148    43-175 (191)
  6 PF08737 Rgp1:  Rgp1;  InterPro  97.3  0.0044 9.5E-08   62.6  13.2   47  104-150   300-346 (415)
  7 KOG3865|consensus               97.1 0.00042 9.2E-09   67.4   3.6   32  120-151   207-238 (402)
  8 PF02752 Arrestin_C:  Arrestin   96.8  0.0026 5.7E-08   52.0   5.7   60  216-278     8-76  (136)
  9 PF03643 Vps26:  Vacuolar prote  96.4   0.043 9.2E-07   52.8  12.0  113   19-150    95-207 (275)
 10 KOG3780|consensus               92.8       1 2.2E-05   44.7  10.8   41  109-149     6-48  (427)
 11 KOG2717|consensus               89.0     2.2 4.9E-05   40.7   8.4  127   17-150    84-220 (313)
 12 PF01345 DUF11:  Domain of unkn  88.7     1.5 3.3E-05   33.3   5.9   44  104-147    22-65  (76)
 13 PF01835 A2M_N:  MG2 domain;  I  88.4       1 2.2E-05   35.7   5.0   26  110-135     2-27  (99)
 14 TIGR01451 B_ant_repeat conserv  85.8     1.5 3.2E-05   31.7   4.2   34  114-147     3-36  (53)
 15 PF09478 CBM49:  Carbohydrate b  85.6     2.1 4.5E-05   33.4   5.2   46  232-282    19-74  (80)
 16 PF00339 Arrestin_N:  Arrestin   85.0       1 2.2E-05   37.4   3.4   39  111-149     2-42  (149)
 17 PF00927 Transglut_C:  Transglu  83.9     1.9   4E-05   34.9   4.4   54  228-282    13-69  (107)
 18 PF07070 Spo0M:  SpoOM protein;  80.5       2 4.4E-05   40.1   3.9   38   15-64     80-118 (218)
 19 KOG3118|consensus               80.3    0.82 1.8E-05   47.1   1.3   32  168-199   106-138 (517)
 20 PF00207 A2M:  Alpha-2-macroglo  77.3     7.4 0.00016   30.7   5.7   39  105-145    53-91  (92)
 21 PF00927 Transglut_C:  Transglu  72.8     9.2  0.0002   30.8   5.3   38  109-147     2-39  (107)
 22 KOG3063|consensus               69.9     9.9 0.00022   36.4   5.5   60   17-87     95-158 (301)
 23 PF06159 DUF974:  Protein of un  69.0      12 0.00025   35.5   5.9   59  223-282     7-70  (249)
 24 PF14796 AP3B1_C:  Clathrin-ada  66.5       8 0.00017   34.0   3.9   54  227-286    82-138 (145)
 25 PF07070 Spo0M:  SpoOM protein;  66.3      15 0.00032   34.4   5.8   47  103-149     8-55  (218)
 26 PF10633 NPCBM_assoc:  NPCBM-as  64.5     5.9 0.00013   30.2   2.4   29  120-148     2-30  (78)
 27 PF07705 CARDB:  CARDB;  InterP  61.0      26 0.00057   26.8   5.6   41  108-148     4-44  (101)
 28 PF00963 Cohesin:  Cohesin doma  57.7      28 0.00061   29.2   5.7   38  110-148     1-38  (141)
 29 PLN02171 endoglucanase          54.1      17 0.00036   39.2   4.4   22  231-252   554-575 (629)
 30 PF07703 A2M_N_2:  Alpha-2-macr  53.8      16 0.00035   30.1   3.5   32  107-138    94-125 (136)
 31 KOG2625|consensus               52.4     8.4 0.00018   36.7   1.7   27  226-252    11-37  (348)
 32 PF05688 DUF824:  Salmonella re  49.7      18  0.0004   26.0   2.7   29  120-148    10-38  (47)
 33 PF11611 DUF4352:  Domain of un  49.5      30 0.00066   27.8   4.4   53  229-282    35-94  (123)
 34 PF04744 Monooxygenase_B:  Mono  48.9      20 0.00043   36.2   3.8   37  107-143   246-283 (381)
 35 COG2373 Large extracellular al  48.9      33 0.00071   40.9   6.0  131  107-252   393-535 (1621)
 36 PF09478 CBM49:  Carbohydrate b  48.2      50  0.0011   25.5   5.2   39  109-147     2-41  (80)
 37 PF07703 A2M_N_2:  Alpha-2-macr  46.5      24 0.00052   29.1   3.4   25  110-134     1-25  (136)
 38 smart00809 Alpha_adaptinC2 Ada  43.2      71  0.0015   25.1   5.6   41  103-147     2-42  (104)
 39 TIGR03079 CH4_NH3mon_ox_B meth  42.0      37  0.0008   34.4   4.4   36  105-140   263-299 (399)
 40 PF02014 Reeler:  Reeler domain  41.8      43 0.00092   28.1   4.3   35  112-148    23-57  (132)
 41 PF05753 TRAP_beta:  Translocon  40.5      57  0.0012   29.5   5.1   74  226-304    33-110 (181)
 42 PF09624 DUF2393:  Protein of u  40.3      68  0.0015   27.4   5.4   29  119-147    58-86  (149)
 43 TIGR01451 B_ant_repeat conserv  40.1      56  0.0012   23.5   4.1   29  226-254     8-36  (53)
 44 cd08544 Reeler Reeler, the N-t  38.1      56  0.0012   27.3   4.4   37  110-148    21-57  (135)
 45 PF01345 DUF11:  Domain of unkn  38.1      57  0.0012   24.4   4.1   30  224-253    35-64  (76)
 46 PF06030 DUF916:  Bacterial pro  36.9      37  0.0008   28.7   3.1   29  120-149    24-52  (121)
 47 PF14796 AP3B1_C:  Clathrin-ada  36.1      85  0.0018   27.7   5.3   46  103-148    64-110 (145)
 48 PF13199 Glyco_hydro_66:  Glyco  36.0      53  0.0011   35.0   4.7   25  113-137     1-25  (559)
 49 PF09624 DUF2393:  Protein of u  35.6      52  0.0011   28.2   3.9   34  216-252    51-84  (149)
 50 cd08548 Type_I_cohesin_like Ty  34.8      76  0.0016   27.1   4.7   40  110-149     1-40  (135)
 51 PF06159 DUF974:  Protein of un  34.0      51  0.0011   31.2   3.9   28  119-146    10-37  (249)
 52 COG4326 Spo0M Sporulation cont  34.0      57  0.0012   30.8   4.0   21   18-38    104-124 (270)
 53 PF02883 Alpha_adaptinC2:  Adap  34.0 1.4E+02   0.003   24.0   6.0   45  100-146     3-47  (115)
 54 COG2373 Large extracellular al  32.0      91   0.002   37.4   6.2   41  105-145   495-535 (1621)
 55 PF13595 DUF4138:  Domain of un  31.8      50  0.0011   31.4   3.4   21  224-244   136-156 (246)
 56 PF04425 Bul1_N:  Bul1 N termin  31.5      28  0.0006   36.0   1.7   22   12-33    252-273 (438)
 57 PF10633 NPCBM_assoc:  NPCBM-as  31.2      48   0.001   25.1   2.7   24  229-252     4-27  (78)
 58 cd00258 GM2-AP GM2 activator p  30.7      61  0.0013   29.2   3.6   35   17-63    109-145 (162)
 59 cd00917 PG-PI_TP The phosphati  30.7      87  0.0019   26.0   4.4   32   16-63     78-109 (122)
 60 PF11355 DUF3157:  Protein of u  29.9 1.1E+02  0.0024   28.4   5.2   37  110-147   100-136 (199)
 61 PF04314 DUF461:  Protein of un  28.3      87  0.0019   25.6   3.9   40  100-139    70-109 (110)
 62 smart00809 Alpha_adaptinC2 Ada  28.0 2.9E+02  0.0063   21.5   6.8   55  224-282    12-66  (104)
 63 PF07919 Gryzun:  Gryzun, putat  27.4 2.1E+02  0.0046   29.3   7.4   40  101-140   166-207 (554)
 64 PF11355 DUF3157:  Protein of u  26.8 1.3E+02  0.0029   27.9   5.1   39  216-258   101-139 (199)
 65 PF07705 CARDB:  CARDB;  InterP  26.8   2E+02  0.0044   21.7   5.6   34  226-262    15-48  (101)
 66 cd08546 cohesin_like Cohesin d  25.5 1.6E+02  0.0035   23.8   5.1   37  110-149     3-39  (135)
 67 PF11797 DUF3324:  Protein of u  24.6 1.5E+02  0.0032   25.4   4.8   66  231-299    43-110 (140)
 68 PF02019 WIF:  WIF domain;  Int  24.0 2.9E+02  0.0063   23.9   6.5   43  104-146    85-128 (132)
 69 PF10437 Lip_prot_lig_C:  Bacte  23.3 1.8E+02  0.0039   22.4   4.7   27  223-252     9-35  (86)
 70 PF04744 Monooxygenase_B:  Mono  23.1 1.1E+02  0.0024   31.0   4.2   67  214-282   248-328 (381)
 71 TIGR03780 Bac_Flav_CT_N Bacter  23.0      88  0.0019   30.6   3.4   20  224-243   175-194 (285)
 72 PF12389 Peptidase_M73:  Camely  23.0 1.3E+02  0.0029   27.9   4.5   32  117-148    59-90  (199)
 73 PF04442 CtaG_Cox11:  Cytochrom  21.8      97  0.0021   27.5   3.2   29  119-147    63-91  (152)
 74 KOG2540|consensus               21.1      82  0.0018   30.0   2.7   71  116-194   155-231 (269)
 75 PF04425 Bul1_N:  Bul1 N termin  20.9 1.3E+02  0.0028   31.2   4.2   44  105-148   131-191 (438)
 76 PF01050 MannoseP_isomer:  Mann  20.7 4.2E+02  0.0091   23.2   7.0   72   77-148    63-139 (151)
 77 PRK05089 cytochrome C oxidase   20.6   2E+02  0.0044   26.4   5.1   28  119-146    90-117 (188)
 78 PF04205 FMN_bind:  FMN-binding  20.4 1.7E+02  0.0037   21.9   4.0   22  229-252     3-24  (81)

No 1  
>KOG3865|consensus
Probab=100.00  E-value=4.6e-62  Score=460.85  Aligned_cols=202  Identities=62%  Similarity=0.899  Sum_probs=193.4

Q ss_pred             ccccCcHHHHHHcccCCCeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccCccccccceeeEEEeeeeee
Q psy17223          4 VIYKVTFGQERLMKKLGPNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGETAEDKIHKRNSVRLAIRKIM   83 (323)
Q Consensus         4 ~~~~lt~~Q~~l~~~~G~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~~~dk~~k~~tv~l~Ir~l~   83 (323)
                      ...|||++|++|++|+|.|+|||.|++|+++|+|+.+|+++++.||+|+|.|+||||++.+.+++++|+++|+|+||+++
T Consensus        91 ~~~plT~lQErLlkKLG~nAyPF~f~~pp~~P~SVtLQp~p~D~gKpcGVdyevkaF~~~s~edk~hKr~sVrL~IRKvq  170 (402)
T KOG3865|consen   91 DSRPLTRLQERLLKKLGSNAYPFTFEFPPNLPCSVTLQPGPEDTGKPCGVDYEVKAFVADSEEDKIHKRNSVRLVIRKVQ  170 (402)
T ss_pred             CCCcccHHHHHHHHHhCCCCCceEEeCCCCCCceEEeccCCccCCCcccceEEEEEEecCCcccccccccceeeeeeeee
Confidence            34679999999999999999999999999999999999999999999999999999999998999999999999999999


Q ss_pred             cCCCCCCCCCeEEEEEEeEecCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCCccchhhhhhhccCC
Q psy17223         84 YAPSKQGEQPSVEVSKEFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGAEDDQDLKDELADSD  163 (323)
Q Consensus        84 ~~P~~~~~~~~~e~~k~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~ed~~d~~~~~~~~d  163 (323)
                      ++|..+.++|+.++.+.|+|++|+++|+|+|||                                               
T Consensus       171 yAP~~~GpqP~~~v~k~FlmS~~~lhLevsLDk-----------------------------------------------  203 (402)
T KOG3865|consen  171 YAPLEPGPQPSAEVSKQFLMSDGPLHLEVSLDK-----------------------------------------------  203 (402)
T ss_pred             ecCCCCCCCchhHhhHhhccCCCceEEEEEecc-----------------------------------------------
Confidence            999998899999999999999999999999886                                               


Q ss_pred             CCCCccCCCccccccccccceeeecccccCCCCCCCCCCcchhhhhhccCcchhhccceeeEEeecceEEEEEEEecCCc
Q psy17223        164 IDGMEEDDLPNIKAWGKNKRMYYNTDYVDDDHGGIQSGTGFVWFEWLKKGSKEKSKKKYLFLYYHGESIAVNVHVANNSN  243 (323)
Q Consensus       164 i~~~~e~~lp~~~awg~~~~~~y~td~~~~~~~~~~~~~~~~~~e~~~~~~~~~s~~k~~~~yyhge~i~v~v~v~N~s~  243 (323)
                                                                                  ++|||||+|.|||+|+||||
T Consensus       204 ------------------------------------------------------------EiYyHGE~isvnV~V~NNsn  223 (402)
T KOG3865|consen  204 ------------------------------------------------------------EIYYHGEPISVNVHVTNNSN  223 (402)
T ss_pred             ------------------------------------------------------------hheecCCceeEEEEEecCCc
Confidence                                                                        57999999999999999999


Q ss_pred             ceEeeEEee-----eEEEeeCCeEeeeeccccCC--CccCCCCccc--------------ccceeecccccccccCCccc
Q psy17223        244 RTVKKIKVS-----DICLFSTAQYKCTVAETESD--CPIAPVSMFD--------------TEDLAMLRHGFKRMFGHAFS  302 (323)
Q Consensus       244 k~vkkikv~-----dv~l~s~~~y~~~Va~~e~~--~~i~p~~t~~--------------~~~~~~~~~~~~~~~~~a~~  302 (323)
                      |||||||++     ||||||++||+|+||.+|+.  |||+||+||+              |+||||||+|||+|+|||||
T Consensus       224 KtVKkIK~~V~Q~adi~Lfs~aqy~~~VA~~E~~eGc~v~Pgstl~Kvf~l~PllanN~dkrGlALDG~lKhEDtnLASS  303 (402)
T KOG3865|consen  224 KTVKKIKISVRQVADICLFSTAQYKKPVAMEETDEGCPVAPGSTLSKVFTLTPLLANNKDKRGLALDGKLKHEDTNLASS  303 (402)
T ss_pred             ceeeeeEEEeEeeceEEEEecccccceeeeeecccCCccCCCCeeeeeEEechhhhcCcccccccccccccccccccchh
Confidence            999999998     99999999999999999974  5999999999              99999999999999999999


Q ss_pred             ccccCccccc
Q psy17223        303 TSLAMPRDEL  312 (323)
Q Consensus       303 t~~~~~~~~~  312 (323)
                      |||+|+.||-
T Consensus       304 Tii~~~~~re  313 (402)
T KOG3865|consen  304 TIIREGADRE  313 (402)
T ss_pred             heecCCCCcc
Confidence            9999999874


No 2  
>KOG3780|consensus
Probab=99.75  E-value=9.7e-18  Score=165.35  Aligned_cols=125  Identities=29%  Similarity=0.380  Sum_probs=98.1

Q ss_pred             cccCCCeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccCccccccceeeEEEeeee---eecCCCCCCCC
Q psy17223         16 MKKLGPNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGETAEDKIHKRNSVRLAIRK---IMYAPSKQGEQ   92 (323)
Q Consensus        16 ~~~~G~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~~~dk~~k~~tv~l~Ir~---l~~~P~~~~~~   92 (323)
                      +.++|.|+|||.|+||.+||+||+     |.+|   +|||.|+|.|+|+  |+..+.....+.|..   ++..|....+.
T Consensus       101 ~l~~G~~~~pF~~~LP~~~P~Sfe-----g~~G---~irY~vk~~idr~--~~~~~~~~~~~~V~~~~~ln~~p~~~~~~  170 (427)
T KOG3780|consen  101 VLPPGNYEFPFSFTLPLNLPPSFE-----GKFG---HVRYFVKAEIDRP--WKLNKKNRKPFTVIETVDLNSSPSLLEPI  170 (427)
T ss_pred             ecCCCceEEeEeccCCCCCCCcee-----eCCc---eEEEEEEEEEecC--CCCCccceeeEEEecccccccCccccCcc
Confidence            578999999999999999999998     5566   9999999999995  455555444444322   22234333222


Q ss_pred             C--eEEEEEEeEecCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCCc
Q psy17223         93 P--SVEVSKEFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGAE  150 (323)
Q Consensus        93 ~--~~e~~k~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~e  150 (323)
                      .  ..+...++||..|+|.+++.|++.+|+|||.|.+++.|+|.|++.++.+++++.|..
T Consensus       171 ~~~~~k~~~~~~~~~g~v~~~~~ip~~~~~~ge~i~~~~~i~n~ss~~~~~~~~~l~q~~  230 (427)
T KOG3780|consen  171 ISKASKKLGCVCFSSGPVSLELTIPKTGYVPGETIPVTLEIENKSSRTIKKVKAKLIQKI  230 (427)
T ss_pred             hhhhhheeeEEEecCCcEEEEEEcccccCcCCccEEEEEEEecCCCCcceeeEEEEEEEE
Confidence            1  123334478999999999999999999999999999999999999999999998854


No 3  
>PF00339 Arrestin_N:  Arrestin (or S-antigen), N-terminal domain;  InterPro: IPR011021 G protein-coupled receptors are a large family of signalling molecules that respond to a wide variety of extracellular stimuli. The receptors relay the information encoded by the ligand through the activation of heterotrimeric G proteins and intracellular effector molecules. To ensure the appropriate regulation of the signalling cascade, it is vital to properly inactivate the receptor. This inactivation is achieved, in part, by the binding of a soluble protein, arrestin, which uncouples the receptor from the downstream G protein after the receptors are phosphorylated by G protein-coupled receptor kinases. In addition to the inactivation of G protein-coupled receptors, arrestins have also been implicated in the endocytosis of receptors and cross talk with other signalling pathways. Arrestin (retinal S-antigen) is a major protein of the retinal rod outer segments. It interacts with photo-activated phosphorylated rhodopsin, inhibiting or 'arresting' its ability to interact with transducin []. The protein binds calcium, and shows similarity in its C terminus to alpha-transducin and other purine nucleotide-binding proteins. In mammals, arrestin is associated with autoimmune uveitis. Arrestins comprise a family of closely-related proteins that includes beta-arrestin-1 and -2, which regulate the function of beta-adrenergic receptors by binding to their phosphorylated forms, impairing their capacity to activate G(S) proteins; Cone photoreceptors C-arrestin (arrestin-X) [], which could bind to phosphorylated red/green opsins; and Drosophila phosrestins I and II, which undergo light-induced phosphorylation, and probably play a role in photoreceptor transduction [, , ].  The crystal structure of bovine retinal arrestin comprises two domains of antiparallel beta-sheets connected through a hinge region and one short alpha-helix on the back of the amino-terminal fold []. The binding region for phosphorylated light-activated rhodopsin is located at the N-terminal domain, as indicated by the docking of the photoreceptor to the three-dimensional structure of arrestin.  The N-terminal domain consists of an immunoglobulin-like beta-sandwich structure. This entry represents proteins with immunoglobulin-like domains that are similar to those found in arrestin.; PDB: 1SUJ_A 3UGX_A 1CF1_B 1AYR_A 3UGU_A 3P2D_B 1ZSH_A 2WTR_B 3GC3_A 1G4R_A ....
Probab=98.96  E-value=1.1e-09  Score=91.90  Aligned_cols=42  Identities=33%  Similarity=0.572  Sum_probs=28.2

Q ss_pred             HcccCCCeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccC
Q psy17223         15 LMKKLGPNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGET   64 (323)
Q Consensus        15 l~~~~G~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~   64 (323)
                      ...++|.|.|||+|+||.+||+||+..     .|   +|+|.|+|.|+++
T Consensus        88 ~~l~~G~~~fpF~f~LP~~lP~S~~~~-----~g---~I~Y~l~a~l~~~  129 (149)
T PF00339_consen   88 NILPPGEYEFPFEFQLPSNLPSSFEGS-----HG---SIRYKLKATLDRP  129 (149)
T ss_dssp             ----C-TTEEEEEE---TTS--SEEEE------S---EEEEEEEEEESST
T ss_pred             ecccCCCEEEEEEEECCCCCCceEecc-----Cc---CEEEEEEEEEECC
Confidence            344599999999999999999999843     45   9999999999886


No 4  
>PF02752 Arrestin_C:  Arrestin (or S-antigen), C-terminal domain;  InterPro: IPR011022 G protein-coupled receptors are a large family of signalling molecules that respond to a wide variety of extracellular stimuli. The receptors relay the information encoded by the ligand through the activation of heterotrimeric G proteins and intracellular effector molecules. To ensure the appropriate regulation of the signalling cascade, it is vital to properly inactivate the receptor. This inactivation is achieved, in part, by the binding of a soluble protein, arrestin, which uncouples the receptor from the downstream G protein after the receptors are phosphorylated by G protein-coupled receptor kinases. In addition to the inactivation of G protein-coupled receptors, arrestins have also been implicated in the endocytosis of receptors and cross talk with other signalling pathways. Arrestin (retinal S-antigen) is a major protein of the retinal rod outer segments. It interacts with photo-activated phosphorylated rhodopsin, inhibiting or 'arresting' its ability to interact with transducin []. The protein binds calcium, and shows similarity in its C terminus to alpha-transducin and other purine nucleotide-binding proteins. In mammals, arrestin is associated with autoimmune uveitis. Arrestins comprise a family of closely-related proteins that includes beta-arrestin-1 and -2, which regulate the function of beta-adrenergic receptors by binding to their phosphorylated forms, impairing their capacity to activate G(S) proteins; Cone photoreceptors C-arrestin (arrestin-X) [], which could bind to phosphorylated red/green opsins; and Drosophila phosrestins I and II, which undergo light-induced phosphorylation, and probably play a role in photoreceptor transduction [, , ].  The crystal structure of bovine retinal arrestin comprises two domains of antiparallel beta-sheets connected through a hinge region and one short alpha-helix on the back of the amino-terminal fold []. The binding region for phosphorylated light-activated rhodopsin is located at the N-terminal domain, as indicated by the docking of the photoreceptor to the three-dimensional structure of arrestin.  The C-terminal domain consists of an immunoglobulin-like beta-sandwich structure. This entry represents proteins with immunoglobulin-like domains that are similar to those found in arrestin.; PDB: 1SUJ_A 3UGX_A 1CF1_B 1AYR_A 3UGU_A 3P2D_B 1ZSH_A 2WTR_B 3GC3_A 1G4R_A ....
Probab=98.56  E-value=9e-08  Score=78.73  Aligned_cols=45  Identities=40%  Similarity=0.573  Sum_probs=39.4

Q ss_pred             CCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCC
Q psy17223        105 PNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGA  149 (323)
Q Consensus       105 sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~  149 (323)
                      +|++++++++||++|.|||+|+|++.|+|.|++.|++|++.|.+.
T Consensus         2 ~g~i~~~~~i~~~~~~~Ge~i~v~v~i~n~s~~~i~~I~v~L~~~   46 (136)
T PF02752_consen    2 SGKISLSISIPRTAYVPGETIPVNVEIDNQSKKKIKKIKVSLVER   46 (136)
T ss_dssp             TEEEEEEEEES-SEEETT--EEEEEEEEE-SSSEEEEEEEEEEEE
T ss_pred             CCEEEEEEEECCCEECCCCEEEEEEEEEECCCCEEEEEEEEEEEE
Confidence            699999999999999999999999999999999999999999874


No 5  
>PF13002 LDB19:  Arrestin_N terminal like;  InterPro: IPR024391 This entry represents a predicted Ig-like beta sandwich domain found towards the N terminus of protein LDB19 []. It is also found in other sequences and is related to the arrestin N-terminal fold [].
Probab=98.03  E-value=4.1e-05  Score=69.67  Aligned_cols=121  Identities=22%  Similarity=0.337  Sum_probs=77.7

Q ss_pred             HHHcccCCCeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccC--c-ccccccee----eEEEee-eeeec
Q psy17223         13 ERLMKKLGPNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGET--A-EDKIHKRN----SVRLAI-RKIMY   84 (323)
Q Consensus        13 ~~l~~~~G~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~--~-~dk~~k~~----tv~l~I-r~l~~   84 (323)
                      ..+..+.|.|.|||++-||-+||+|..+-  .+-.   ..|.|+++|.....  . .....+..    ...+.| |-|. 
T Consensus        43 ~~t~l~~G~h~fPFS~LiPG~LPaS~~lg--s~~l---~~I~Yel~A~a~~~~~~~~~~~~~~~~~~~~~pl~V~Rsi~-  116 (191)
T PF13002_consen   43 HPTTLTKGSHAFPFSYLIPGHLPASMDLG--STPL---VSIKYELKAEATYKDPRRGSSSSKPRVLKLKRPLPVKRSIL-  116 (191)
T ss_pred             CccccCCCcccCCeeEECCCCCccccccC--CCCc---EEEEEEEEEEEEEccCccccCCCcceeEEEeeeEEEEEecC-
Confidence            34557899999999999999999999742  1223   48999999987541  0 00011111    112222 2221 


Q ss_pred             CCCCCCCCCeEEEEEEeEecCCceEEEEEeCccceecCCeEEEEEEEeccCcce----eeEEEeecCC
Q psy17223         85 APSKQGEQPSVEVSKEFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRT----VKKIKVSDSG  148 (323)
Q Consensus        85 ~P~~~~~~~~~e~~k~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~----Vk~Ikv~L~q  148 (323)
                       |     .+  .....-.|..=.|...|.||. +.+|--+.+|++.++|-++.+    ++++.=++.+
T Consensus       117 -~-----gp--d~~S~RvFPPT~l~a~a~lP~-VI~P~gtfpvel~LdGv~~~~~rWRlrKltWRiEE  175 (191)
T PF13002_consen  117 -P-----GP--DKNSLRVFPPTNLTASAVLPN-VIHPKGTFPVELRLDGVVSKDRRWRLRKLTWRIEE  175 (191)
T ss_pred             -C-----CC--CcccEEecCCCCcEEEEEcCC-eeCCCCcccEEEEEecccCCCCEEEEEeeeEEEee
Confidence             1     11  112334678888999999995 666888899999999987664    6666666544


No 6  
>PF08737 Rgp1:  Rgp1;  InterPro: IPR014848 Rgp1 forms heterodimer with Ric1 (IPR009771 from INTERPRO) which associates with Golgi membranes and functions as a guanyl-nucleotide exchange factor []. 
Probab=97.27  E-value=0.0044  Score=62.63  Aligned_cols=47  Identities=21%  Similarity=0.212  Sum_probs=41.0

Q ss_pred             cCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCCc
Q psy17223        104 SPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGAE  150 (323)
Q Consensus       104 ~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~e  150 (323)
                      .+|..-..++|.|..|.-||+|...+++++.....+..+.+.|.-.|
T Consensus       300 ~n~~~va~~~LsK~~yrlGE~I~g~idf~~~~~~~c~~v~~~LEs~E  346 (415)
T PF08737_consen  300 RNGQRVARLSLSKPAYRLGEDIVGTIDFNDASTIPCYQVSASLESEE  346 (415)
T ss_pred             ECCeEEEEEEecCCCcccCCeEEEEEEcCCCCcceeEEEEEEEEEEE
Confidence            37888889999999999999999999999998677888888886544


No 7  
>KOG3865|consensus
Probab=97.09  E-value=0.00042  Score=67.45  Aligned_cols=32  Identities=69%  Similarity=0.957  Sum_probs=29.6

Q ss_pred             ecCCeEEEEEEEeccCcceeeEEEeecCCCcc
Q psy17223        120 YHGESIAVNVHVANNSNRTVKKIKVSDSGAED  151 (323)
Q Consensus       120 ~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~ed  151 (323)
                      +|||+|.|+|.|.|+|+|.|++||+.+.|.-|
T Consensus       207 yHGE~isvnV~V~NNsnKtVKkIK~~V~Q~ad  238 (402)
T KOG3865|consen  207 YHGEPISVNVHVTNNSNKTVKKIKISVRQVAD  238 (402)
T ss_pred             ecCCceeEEEEEecCCcceeeeeEEEeEeece
Confidence            58999999999999999999999999998543


No 8  
>PF02752 Arrestin_C:  Arrestin (or S-antigen), C-terminal domain;  InterPro: IPR011022 G protein-coupled receptors are a large family of signalling molecules that respond to a wide variety of extracellular stimuli. The receptors relay the information encoded by the ligand through the activation of heterotrimeric G proteins and intracellular effector molecules. To ensure the appropriate regulation of the signalling cascade, it is vital to properly inactivate the receptor. This inactivation is achieved, in part, by the binding of a soluble protein, arrestin, which uncouples the receptor from the downstream G protein after the receptors are phosphorylated by G protein-coupled receptor kinases. In addition to the inactivation of G protein-coupled receptors, arrestins have also been implicated in the endocytosis of receptors and cross talk with other signalling pathways. Arrestin (retinal S-antigen) is a major protein of the retinal rod outer segments. It interacts with photo-activated phosphorylated rhodopsin, inhibiting or 'arresting' its ability to interact with transducin []. The protein binds calcium, and shows similarity in its C terminus to alpha-transducin and other purine nucleotide-binding proteins. In mammals, arrestin is associated with autoimmune uveitis. Arrestins comprise a family of closely-related proteins that includes beta-arrestin-1 and -2, which regulate the function of beta-adrenergic receptors by binding to their phosphorylated forms, impairing their capacity to activate G(S) proteins; Cone photoreceptors C-arrestin (arrestin-X) [], which could bind to phosphorylated red/green opsins; and Drosophila phosrestins I and II, which undergo light-induced phosphorylation, and probably play a role in photoreceptor transduction [, , ].  The crystal structure of bovine retinal arrestin comprises two domains of antiparallel beta-sheets connected through a hinge region and one short alpha-helix on the back of the amino-terminal fold []. The binding region for phosphorylated light-activated rhodopsin is located at the N-terminal domain, as indicated by the docking of the photoreceptor to the three-dimensional structure of arrestin.  The C-terminal domain consists of an immunoglobulin-like beta-sandwich structure. This entry represents proteins with immunoglobulin-like domains that are similar to those found in arrestin.; PDB: 1SUJ_A 3UGX_A 1CF1_B 1AYR_A 3UGU_A 3P2D_B 1ZSH_A 2WTR_B 3GC3_A 1G4R_A ....
Probab=96.82  E-value=0.0026  Score=52.02  Aligned_cols=60  Identities=37%  Similarity=0.504  Sum_probs=40.3

Q ss_pred             hhhccceeeEEeecceEEEEEEEecCCcceEeeEEee---eEEEeeC------CeEeeeeccccCCCccCCC
Q psy17223        216 EKSKKKYLFLYYHGESIAVNVHVANNSNRTVKKIKVS---DICLFST------AQYKCTVAETESDCPIAPV  278 (323)
Q Consensus       216 ~~s~~k~~~~yyhge~i~v~v~v~N~s~k~vkkikv~---dv~l~s~------~~y~~~Va~~e~~~~i~p~  278 (323)
                      ..++.|  ..|..||.|.|++.|+|.|++.|++|++.   .+..+..      .++.+.|+. ...+.+.++
T Consensus         8 ~~~i~~--~~~~~Ge~i~v~v~i~n~s~~~i~~I~v~L~~~~~~~~~~~~~~~~~~~~~v~~-~~~~~~~~~   76 (136)
T PF02752_consen    8 SISIPR--TAYVPGETIPVNVEIDNQSKKKIKKIKVSLVERITYKAKGGKDESKSEKRVVAK-SKNCGVDPG   76 (136)
T ss_dssp             EEEES---SEEETT--EEEEEEEEE-SSSEEEEEEEEEEEEEEE-SS----S-EEEEEEEEE-EECCEB-B-
T ss_pred             EEEECC--CEECCCCEEEEEEEEEECCCCEEEEEEEEEEEEEEEEEeeccccceEEEEEEEE-EecCCccCC
Confidence            566778  88999999999999999999999999998   5555544      346677777 334444333


No 9  
>PF03643 Vps26:  Vacuolar protein sorting-associated protein 26 ;  InterPro: IPR005377  The movement of lipid and protein components between intracellular organelles requires the regulated interactions of many molecules. Vacuolar protein sorting-associated protein (Vps)5 is a yeast protein that is a subunit of a large multimeric complex, termed the retromer complex, involved in retrograde transport of proteins from endosomes to the trans-Golgi network. Sorting nexin (SNX) 1 and SNX2 are its mammalian orthologs []. To carry out its biological functions, Vps5 forms the retromer complex with at least four other proteins: Vps17, Vps26, Vps29, and Vps35 []. This family of Vps26-proteins also contains Down syndrome critical region 3/A.; GO: 0007034 vacuolar transport, 0030904 retromer complex; PDB: 3LHA_A 3LH9_A 2R51_A 3LH8_B 2FAU_A.
Probab=96.45  E-value=0.043  Score=52.82  Aligned_cols=113  Identities=19%  Similarity=0.234  Sum_probs=64.7

Q ss_pred             CCCeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccCccccccceeeEEEeeeeeecCCCCCCCCCeEEEE
Q psy17223         19 LGPNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGETAEDKIHKRNSVRLAIRKIMYAPSKQGEQPSVEVS   98 (323)
Q Consensus        19 ~G~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~~~dk~~k~~tv~l~Ir~l~~~P~~~~~~~~~e~~   98 (323)
                      .|.. |||.|.+-.. |  ++     .|+|....|+|.|+|.+.|+.   ..-..++.|.|..+...|...     ..+ 
T Consensus        95 ~~~t-~pFeF~~~~k-~--yE-----TY~G~~v~i~Y~lrv~v~R~~---~~i~k~~ef~V~~~~~~p~~~-----~~i-  156 (275)
T PF03643_consen   95 EGKT-FPFEFPLVEK-P--YE-----TYHGVNVNIRYFLRVTVKRSY---KDISKEQEFWVQNFSITPESN-----QPI-  156 (275)
T ss_dssp             S-EE-EEEEE-SB------S-------EE-SSEEEEEEEEEEE--SS---S-EEEEEEEEEE-EB-------------E-
T ss_pred             CCcE-EeeEeCCCCC-C--Cc-----cEeeeEEEEEEEEEEEEEccC---CCcceEEEEEEEeccCCCCCC-----CCc-
Confidence            4445 9999987432 1  32     467888899999999998864   222345566776554444332     111 


Q ss_pred             EEeEecCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCCc
Q psy17223         99 KEFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGAE  150 (323)
Q Consensus        99 k~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~e  150 (323)
                      +.=.--.+-++++..++|..|.=.+.|.=.+.+.-.. ..|++|.++|.+-|
T Consensus       157 k~evgie~~lhief~~~k~~~~l~d~i~G~i~f~lv~-~kIk~~elqLiR~E  207 (275)
T PF03643_consen  157 KMEVGIEDCLHIEFEYDKSKYHLKDVITGKIYFLLVR-IKIKSMELQLIRVE  207 (275)
T ss_dssp             EEEECETTTEEEEEEES-SEEETT-EEEEEEEEEEES-S-EEEEEEEEEEEE
T ss_pred             ccccCCCccEEEEEEEcccceECCCCEEEEEEEEEEe-ecceEEEEEEEEEE
Confidence            1111135678999999999999999987776664333 67999999998844


No 10 
>KOG3780|consensus
Probab=92.84  E-value=1  Score=44.68  Aligned_cols=41  Identities=27%  Similarity=0.482  Sum_probs=35.6

Q ss_pred             EEEEEeCcc--ceecCCeEEEEEEEeccCcceeeEEEeecCCC
Q psy17223        109 HLEASLDKE--LYYHGESIAVNVHVANNSNRTVKKIKVSDSGA  149 (323)
Q Consensus       109 ~L~a~LdK~--~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~  149 (323)
                      .+.+.+|+.  +|.+||+|.=++.+.+.....++.|++++.+.
T Consensus         6 ~~~i~~d~~~~iy~~G~~vsG~v~l~~~~~~~~~~i~l~~~G~   48 (427)
T KOG3780|consen    6 SFEIVLDNPEAIYFPGEPVSGSVVLSTKEPIKVRAIKLQLKGR   48 (427)
T ss_pred             eEEEEeCCCccccCCCCeEEEEEEEEeCCccceeEEEEEEEEe
Confidence            345666666  59999999999999999999999999999874


No 11 
>KOG2717|consensus
Probab=89.03  E-value=2.2  Score=40.67  Aligned_cols=127  Identities=13%  Similarity=0.134  Sum_probs=74.4

Q ss_pred             ccCCCeeeeeEeeCCC-CCCCcEEeccCCCCCCCceeeEEEEEEEEccCccccccceeeEEEeeee----eecCCCCC--
Q psy17223         17 KKLGPNAFPFFFELPP-SCPASVTLQPAPGDTGKPCGVDYELKAFVGETAEDKIHKRNSVRLAIRK----IMYAPSKQ--   89 (323)
Q Consensus        17 ~~~G~h~FPFsFqLP~-~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~~~dk~~k~~tv~l~Ir~----l~~~P~~~--   89 (323)
                      .|+|.-+|||.|.|-. +=|--|.    ..++|-...|.|.+++.+.|.--.|+. ..++.|.|..    +...|...  
T Consensus        84 ~p~G~tEipFelpL~~kge~~~lY----ETyHGvfiNiqY~LtcdikR~~L~K~l-tkt~eFiv~s~pv~l~e~~p~iV~  158 (313)
T KOG2717|consen   84 IPPGTTEIPFELPLREKGEGEKLY----ETYHGVFINIQYLLTCDIKRGYLHKPL-TKTMEFIVESGPVDLPERPPEIVI  158 (313)
T ss_pred             CCCCceeeeeeeeeccCCCccEee----eeecceEEEEEEEEEEecccchhcCch-hhhheeeeccCCcccccCCCcceE
Confidence            4789999999988764 2232222    147888889999999999886333332 2345555531    11001000  


Q ss_pred             ---CCCCeEEEEEEeEecCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCCc
Q psy17223         90 ---GEQPSVEVSKEFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGAE  150 (323)
Q Consensus        90 ---~~~~~~e~~k~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~e  150 (323)
                         .|-..+...+. -..-|..-+.-.||+.-++--+++.=.+.|+ .|...|++|.++|.+-|
T Consensus       159 F~itpdtlq~~~ke-r~~~p~FlvtG~Ld~t~c~~t~PltGeltVe-~seaaI~Sie~qLvRVE  220 (313)
T KOG2717|consen  159 FYITPDTLQHPLKE-RIKTPGFLVTGKLDATQCSLTDPLTGELTVE-ASEAAITSIEIQLVRVE  220 (313)
T ss_pred             EEEChHHhhccchh-hccCCceEEEeeecceeeEecCCccceEEEE-eeccceeEEEEEEEEEE
Confidence               00000000000 1233455566778888888777777666665 45688999999998744


No 12 
>PF01345 DUF11:  Domain of unknown function DUF11;  InterPro: IPR001434 This group of sequences is represented by a conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis (Streptococcus faecalis), and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydia trachomatis outer membrane proteins.  In C. trachomatis, three cysteine-rich proteins (also believed to be lipoproteins), MOMP, OMP6 and OMP3, make up the extracellular matrix of the outer membrane []. They are involved in the essential structural integrity of both the elementary body (EB) and recticulate body (RB) phase. They are thought to be involved in porin formation and, as these bacteria lack the peptidoglycan layer common to most Gram-negative microbes, such proteins are highly important in the pathogenicity of the organism.; GO: 0005727 extrachromosomal circular DNA
Probab=88.68  E-value=1.5  Score=33.26  Aligned_cols=44  Identities=14%  Similarity=0.311  Sum_probs=39.4

Q ss_pred             cCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223        104 SPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDS  147 (323)
Q Consensus       104 ~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~  147 (323)
                      ....+.+.-+.++....+||.|..++.+.|..+.....+.+.+.
T Consensus        22 ~~~~~~~~k~~~~~~~~~Gd~v~ytitvtN~G~~~a~nv~v~D~   65 (76)
T PF01345_consen   22 AIPDLSITKTVNPSTANPGDTVTYTITVTNTGPAPATNVVVTDT   65 (76)
T ss_pred             CCCCEEEEEecCCCcccCCCEEEEEEEEEECCCCeeEeEEEEEc
Confidence            45678888889999999999999999999999999999988874


No 13 
>PF01835 A2M_N:  MG2 domain;  InterPro: IPR002890 The proteinase-binding alpha-macroglobulins (A2M) [] are large glycoproteins found in the plasma of vertebrates, in the hemolymph of some invertebrates and in reptilian and avian egg white. A2M-like proteins are able to inhibit all four classes of proteinases by a 'trapping' mechanism. They have a peptide stretch, called the 'bait region', which contains specific cleavage sites for different proteinases. When a proteinase cleaves the bait region, a conformational change is induced in the protein, thus trapping the proteinase. The entrapped enzyme remains active against low molecular weight substrates, whilst its activity toward larger substrates is greatly reduced, due to steric hindrance. Following cleavage in the bait region, a thiol ester bond, formed between the side chains of a cysteine and a glutamine, is cleaved and mediates the covalent binding of the A2M-like protein to the proteinase. This family includes the N-terminal region of the alpha-2-macroglobulin family. The inhibitor domains belong to MEROPS inhibitor family I39.; GO: 0004866 endopeptidase inhibitor activity; PDB: 2B39_B 3KLS_B 3PRX_C 3KM9_B 3PVM_C 3CU7_A 4E0S_A 4A5W_A 4ACQ_C 2P9R_B ....
Probab=88.38  E-value=1  Score=35.67  Aligned_cols=26  Identities=19%  Similarity=0.340  Sum_probs=21.8

Q ss_pred             EEEEeCccceecCCeEEEEEEEeccC
Q psy17223        110 LEASLDKELYYHGESIAVNVHVANNS  135 (323)
Q Consensus       110 L~a~LdK~~Y~PGE~I~V~v~IdN~S  135 (323)
                      +-+..||..|.|||+|.+.+-+.+..
T Consensus         2 ~~i~TDr~iYrPGetV~~~~~~~~~~   27 (99)
T PF01835_consen    2 IFIQTDRPIYRPGETVHFRAIVRDLD   27 (99)
T ss_dssp             EEEEESSSEE-TTSEEEEEEEEEEEC
T ss_pred             EEEECCccCcCCCCEEEEEEEEeccc
Confidence            45789999999999999999987776


No 14 
>TIGR01451 B_ant_repeat conserved repeat domain. This model represents the conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis, and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydial outer membrane proteins.
Probab=85.79  E-value=1.5  Score=31.75  Aligned_cols=34  Identities=29%  Similarity=0.460  Sum_probs=30.5

Q ss_pred             eCccceecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223        114 LDKELYYHGESIAVNVHVANNSNRTVKKIKVSDS  147 (323)
Q Consensus       114 LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~  147 (323)
                      .++....|||.|..++.|.|.....+..|.+.+.
T Consensus         3 ~d~~~~~~Gd~v~Yti~v~N~g~~~a~~v~v~D~   36 (53)
T TIGR01451         3 VDKTVATIGDTITYTITVTNNGNVPATNVVVTDI   36 (53)
T ss_pred             cCccccCCCCEEEEEEEEEECCCCceEeEEEEEc
Confidence            5778889999999999999999999998888763


No 15 
>PF09478 CBM49:  Carbohydrate binding domain CBM49;  InterPro: IPR019028 A carbohydrate-binding module (CBM) is defined as a contiguous amino acid sequence within a carbohydrate-active enzyme with a discreet fold having carbohydrate-binding activity. A few exceptions are CBMs in cellulosomal scaffolding proteins and rare instances of independent putative CBMs. The requirement of CBMs existing as modules within larger enzymes sets this class of carbohydrate-binding protein apart from other non-catalytic sugar binding proteins such as lectins and sugar transport proteins. CBMs were previously classified as cellulose-binding domains (CBDs) based on the initial discovery of several modules that bound cellulose [, ]. However, additional modules in carbohydrate-active enzymes are continually being found that bind carbohydrates other than cellulose yet otherwise meet the CBM criteria, hence the need to reclassify these polypeptides using more inclusive terminology. Previous classification of cellulose-binding domains were based on amino acid similarity. Groupings of CBDs were called "Types" and numbered with roman numerals (e.g. Type I or Type II CBDs). In keeping with the glycoside hydrolase classification, these groupings are now called families and numbered with Arabic numerals. Families 1 to 13 are the same as Types I to XIII. For a detailed review on the structure and binding modes of CBMs see [].  This domain is found at the C-terminal of cellulases and in vitro binding studies have shown it to binds to crystalline cellulose []. ; GO: 0030246 carbohydrate binding, 0005576 extracellular region
Probab=85.57  E-value=2.1  Score=33.36  Aligned_cols=46  Identities=28%  Similarity=0.432  Sum_probs=30.4

Q ss_pred             EEEEEEEecCCcceEeeEEee-e--------EEEeeCCeEeeeeccccC-CCccCCCCccc
Q psy17223        232 IAVNVHVANNSNRTVKKIKVS-D--------ICLFSTAQYKCTVAETES-DCPIAPVSMFD  282 (323)
Q Consensus       232 i~v~v~v~N~s~k~vkkikv~-d--------v~l~s~~~y~~~Va~~e~-~~~i~p~~t~~  282 (323)
                      ...+|.|.|+++++|+.+++. |        +..-+++.|.     +-. ..+|.||++++
T Consensus        19 ~qy~v~I~N~~~~~I~~~~i~~~~l~~~iW~l~~~~~~~y~-----lPs~~~~i~pg~s~~   74 (80)
T PF09478_consen   19 TQYDVTITNNGSKPIKSLKISIDNLYGSIWGLDKVSGNTYT-----LPSYQPTIKPGQSFT   74 (80)
T ss_pred             EEEEEEEEECCCCeEEEEEEEECccchhheeEEeccCCEEE-----CCccccccCCCCEEE
Confidence            457889999999999999997 4        2222233332     111 23889998864


No 16 
>PF00339 Arrestin_N:  Arrestin (or S-antigen), N-terminal domain;  InterPro: IPR011021 G protein-coupled receptors are a large family of signalling molecules that respond to a wide variety of extracellular stimuli. The receptors relay the information encoded by the ligand through the activation of heterotrimeric G proteins and intracellular effector molecules. To ensure the appropriate regulation of the signalling cascade, it is vital to properly inactivate the receptor. This inactivation is achieved, in part, by the binding of a soluble protein, arrestin, which uncouples the receptor from the downstream G protein after the receptors are phosphorylated by G protein-coupled receptor kinases. In addition to the inactivation of G protein-coupled receptors, arrestins have also been implicated in the endocytosis of receptors and cross talk with other signalling pathways. Arrestin (retinal S-antigen) is a major protein of the retinal rod outer segments. It interacts with photo-activated phosphorylated rhodopsin, inhibiting or 'arresting' its ability to interact with transducin []. The protein binds calcium, and shows similarity in its C terminus to alpha-transducin and other purine nucleotide-binding proteins. In mammals, arrestin is associated with autoimmune uveitis. Arrestins comprise a family of closely-related proteins that includes beta-arrestin-1 and -2, which regulate the function of beta-adrenergic receptors by binding to their phosphorylated forms, impairing their capacity to activate G(S) proteins; Cone photoreceptors C-arrestin (arrestin-X) [], which could bind to phosphorylated red/green opsins; and Drosophila phosrestins I and II, which undergo light-induced phosphorylation, and probably play a role in photoreceptor transduction [, , ].  The crystal structure of bovine retinal arrestin comprises two domains of antiparallel beta-sheets connected through a hinge region and one short alpha-helix on the back of the amino-terminal fold []. The binding region for phosphorylated light-activated rhodopsin is located at the N-terminal domain, as indicated by the docking of the photoreceptor to the three-dimensional structure of arrestin.  The N-terminal domain consists of an immunoglobulin-like beta-sandwich structure. This entry represents proteins with immunoglobulin-like domains that are similar to those found in arrestin.; PDB: 1SUJ_A 3UGX_A 1CF1_B 1AYR_A 3UGU_A 3P2D_B 1ZSH_A 2WTR_B 3GC3_A 1G4R_A ....
Probab=85.01  E-value=1  Score=37.40  Aligned_cols=39  Identities=36%  Similarity=0.486  Sum_probs=31.3

Q ss_pred             EEEeC--ccceecCCeEEEEEEEeccCcceeeEEEeecCCC
Q psy17223        111 EASLD--KELYYHGESIAVNVHVANNSNRTVKKIKVSDSGA  149 (323)
Q Consensus       111 ~a~Ld--K~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~  149 (323)
                      ++.||  +..|.|||.|.=+|.+.......++.|++++.+.
T Consensus         2 ~I~ld~~~~~y~~Ge~I~G~V~l~~~~~~~i~~i~v~l~G~   42 (149)
T PF00339_consen    2 EIELDNPKPVYFPGEVISGKVVLELSKPIKIKSIKVRLKGR   42 (149)
T ss_dssp             EEEES-SEEEEESS--EEEEEEECTTT-TTTSEEEEEEEEE
T ss_pred             EEEECCCCCEECCCCEEEEEEEEEECCccceeEEEEEEEEE
Confidence            34555  9999999999999999888888999999999874


No 17 
>PF00927 Transglut_C:  Transglutaminase family, C-terminal ig like domain;  InterPro: IPR008958 Synonym(s): Protein-glutamine gamma-glutamyltransferase, Fibrinoligase, TGase  Transglutaminases catalyse the post-translational modification of proteins at glutamine residues, with formation of isopeptide bonds. Members of the transglutaminase family usually have three domains: N-terminal (IPR001102 from INTERPRO), middle (IPR013808 from INTERPRO) and C-terminal. The middle domain is usually well conserved, but family members can display major differences in their N- and C-terminal domains, although their overall structure is conserved []. This entry represents the C-terminal domain found in transglutaminases, which consists of an immunoglobulin-like beta-sandwich consisting of seven strands in two sheets with a Greek key topology. The best known transglutaminase is blood coagulation factor XIII, a plasma tetrameric protein composed of two catalytic A subunits and two non-catalytic B subunits. Factor XIII is responsible for cross-linking fibrin chains, thus stabilising the fibrin clot. Protein-glutamine gamma-glutamyltransferases (2.3.2.13 from EC) are calcium-dependent enzymes that catalyse the cross-linking of proteins by promoting the formation of isopeptide bonds between the gamma-carboxyl group of a glutamine in one polypeptide chain and the epsilon-amino group of a lysine in a second polypeptide chain. TGases also catalyse the conjugation of polyamines to proteins [, ].; GO: 0003810 protein-glutamine gamma-glutamyltransferase activity, 0018149 peptide cross-linking; PDB: 2XZZ_A 1GGY_B 1FIE_B 1GGU_B 1GGT_B 1F13_A 1QRK_B 1EVU_A 1EX0_B 1L9N_B ....
Probab=83.86  E-value=1.9  Score=34.88  Aligned_cols=54  Identities=13%  Similarity=0.224  Sum_probs=38.8

Q ss_pred             ecceEEEEEEEecCCcceEeeEEee---eEEEeeCCeEeeeeccccCCCccCCCCccc
Q psy17223        228 HGESIAVNVHVANNSNRTVKKIKVS---DICLFSTAQYKCTVAETESDCPIAPVSMFD  282 (323)
Q Consensus       228 hge~i~v~v~v~N~s~k~vkkikv~---dv~l~s~~~y~~~Va~~e~~~~i~p~~t~~  282 (323)
                      -|+++.|.|.+.|.++..++.|++.   ..+.| +|-.+...-.....-.|.||.+..
T Consensus        13 vG~d~~v~v~~~N~~~~~l~~v~~~l~~~~v~y-tG~~~~~~~~~~~~~~l~p~~~~~   69 (107)
T PF00927_consen   13 VGQDFTVSVSFTNPSSEPLRNVSLNLCAFTVEY-TGLTRDQFKKEKFEVTLKPGETKS   69 (107)
T ss_dssp             TTSEEEEEEEEEE-SSS-EECEEEEEEEEEEEC-TTTEEEEEEEEEEEEEE-TTEEEE
T ss_pred             CCCCEEEEEEEEeCCcCccccceeEEEEEEEEE-CCcccccEeEEEcceeeCCCCEEE
Confidence            4999999999999999999999986   55556 676554444444455799998765


No 18 
>PF07070 Spo0M:  SpoOM protein;  InterPro: IPR009776 This family consists of several bacterial SpoOM proteins which are thought to control sporulation in Bacillus subtilis.Spo0M exerts certain negative effects on sporulation and its gene expression is controlled by sigmaH [].
Probab=80.49  E-value=2  Score=40.11  Aligned_cols=38  Identities=26%  Similarity=0.423  Sum_probs=29.3

Q ss_pred             HcccCC-CeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccC
Q psy17223         15 LMKKLG-PNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGET   64 (323)
Q Consensus        15 l~~~~G-~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~   64 (323)
                      +..++| .+.+||+|+||.++|-|.         |   +.+|+|+..++-.
T Consensus        80 f~I~~ge~~~iPF~~~lP~etPiT~---------~---~~~v~l~T~LdI~  118 (218)
T PF07070_consen   80 FTIEPGEEKEIPFSFPLPWETPITE---------G---GMRVWLRTGLDIA  118 (218)
T ss_pred             EEECCCCEEEEeEEEECCCCCCccC---------C---CcEEEEEEEEEeC
Confidence            333444 589999999999999976         2   5778888888764


No 19 
>KOG3118|consensus
Probab=80.34  E-value=0.82  Score=47.13  Aligned_cols=32  Identities=28%  Similarity=0.579  Sum_probs=26.6

Q ss_pred             ccCCCccccccccccceeeecc-cccCCCCCCC
Q psy17223        168 EEDDLPNIKAWGKNKRMYYNTD-YVDDDHGGIQ  199 (323)
Q Consensus       168 ~e~~lp~~~awg~~~~~~y~td-~~~~~~~~~~  199 (323)
                      ++.++.+..+||.++..||.+| |.++++++-+
T Consensus       106 ~e~e~ddn~~WG~~s~~yyg~dd~dddd~s~e~  138 (517)
T KOG3118|consen  106 KEEEEDDNSTWGGRSGLYYGGDDVDDDDLSSED  138 (517)
T ss_pred             cchhhhcccccccccccccCCccccchhhccch
Confidence            3457889999999999999997 8888887644


No 20 
>PF00207 A2M:  Alpha-2-macroglobulin family;  InterPro: IPR001599 This entry contains serum complement C3 and C4 precursors and alpha-macrogrobulins.  The alpha-macroglobulin (aM) family of proteins includes protease inhibitors [], typified by the human tetrameric a2-macroglobulin (a2M); they belong to the MEROPS proteinase inhibitor family I39, clan IL. These protease inhibitors share several defining properties, which include (i) the ability to inhibit proteases from all catalytic classes, (ii) the presence of a 'bait region' and a thiol ester, (iii) a similar protease inhibitory mechanism and (iv) the inactivation of the inhibitory capacity by reaction of the thiol ester with small primary amines. aM protease inhibitors inhibit by steric hindrance []. The mechanism involves protease cleavage of the bait region, a segment of the aM that is particularly susceptible to proteolytic cleavage, which initiates a conformational change such that the aM collapses about the protease. In the resulting aM-protease complex, the active site of the protease is sterically shielded, thus substantially decreasing access to protein substrates. Two additional events occur as a consequence of bait region cleavage, namely (i) the h-cysteinyl-g-glutamyl thiol ester becomes highly reactive and (ii) a major conformational change exposes a conserved COOH-terminal receptor binding domain [] (RBD). RBD exposure allows the aM protease complex to bind to clearance receptors and be removed from circulation []. Tetrameric, dimeric, and, more recently, monomeric aM protease inhibitors have been identified [, ].; GO: 0004866 endopeptidase inhibitor activity; PDB: 3KLS_B 3PRX_C 3KM9_B 3PVM_C 3CU7_A 4E0S_A 4A5W_A 2PN5_A 3FRP_G 3HRZ_B ....
Probab=77.32  E-value=7.4  Score=30.70  Aligned_cols=39  Identities=18%  Similarity=0.356  Sum_probs=27.4

Q ss_pred             CCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEee
Q psy17223        105 PNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVS  145 (323)
Q Consensus       105 sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~  145 (323)
                      .-++.++..||+ ....||.+.|.+.|-|+..+.+. ++|+
T Consensus        53 ~~p~~i~~~lP~-~l~~GD~~~i~v~v~N~~~~~~~-v~V~   91 (92)
T PF00207_consen   53 FKPFFIQLNLPR-SLRRGDQIQIPVTVFNYTDKDQE-VTVT   91 (92)
T ss_dssp             B-SEEEEEE--S-EEETTSEEEEEEEEEE-SSS-EE-EEEE
T ss_pred             EeeEEEEcCCCc-EEecCCEEEEEEEEEeCCCCCEE-EEEE
Confidence            448889999997 45799999999999999887764 4443


No 21 
>PF00927 Transglut_C:  Transglutaminase family, C-terminal ig like domain;  InterPro: IPR008958 Synonym(s): Protein-glutamine gamma-glutamyltransferase, Fibrinoligase, TGase  Transglutaminases catalyse the post-translational modification of proteins at glutamine residues, with formation of isopeptide bonds. Members of the transglutaminase family usually have three domains: N-terminal (IPR001102 from INTERPRO), middle (IPR013808 from INTERPRO) and C-terminal. The middle domain is usually well conserved, but family members can display major differences in their N- and C-terminal domains, although their overall structure is conserved []. This entry represents the C-terminal domain found in transglutaminases, which consists of an immunoglobulin-like beta-sandwich consisting of seven strands in two sheets with a Greek key topology. The best known transglutaminase is blood coagulation factor XIII, a plasma tetrameric protein composed of two catalytic A subunits and two non-catalytic B subunits. Factor XIII is responsible for cross-linking fibrin chains, thus stabilising the fibrin clot. Protein-glutamine gamma-glutamyltransferases (2.3.2.13 from EC) are calcium-dependent enzymes that catalyse the cross-linking of proteins by promoting the formation of isopeptide bonds between the gamma-carboxyl group of a glutamine in one polypeptide chain and the epsilon-amino group of a lysine in a second polypeptide chain. TGases also catalyse the conjugation of polyamines to proteins [, ].; GO: 0003810 protein-glutamine gamma-glutamyltransferase activity, 0018149 peptide cross-linking; PDB: 2XZZ_A 1GGY_B 1FIE_B 1GGU_B 1GGT_B 1F13_A 1QRK_B 1EVU_A 1EX0_B 1L9N_B ....
Probab=72.78  E-value=9.2  Score=30.80  Aligned_cols=38  Identities=16%  Similarity=0.314  Sum_probs=30.8

Q ss_pred             EEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223        109 HLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDS  147 (323)
Q Consensus       109 ~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~  147 (323)
                      ++++.++..+. -|+.+.+.+.+.|.++..++.|++.+.
T Consensus         2 ~~~i~~~~~~~-vG~d~~v~v~~~N~~~~~l~~v~~~l~   39 (107)
T PF00927_consen    2 EIKIKLPGDPV-VGQDFTVSVSFTNPSSEPLRNVSLNLC   39 (107)
T ss_dssp             EEEEEEESEEB-TTSEEEEEEEEEE-SSS-EECEEEEEE
T ss_pred             eEEEEECCCcc-CCCCEEEEEEEEeCCcCccccceeEEE
Confidence            56677766665 899999999999999999999998884


No 22 
>KOG3063|consensus
Probab=69.90  E-value=9.9  Score=36.43  Aligned_cols=60  Identities=23%  Similarity=0.361  Sum_probs=37.1

Q ss_pred             ccCCC----eeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccCccccccceeeEEEeeeeeecCCC
Q psy17223         17 KKLGP----NAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGETAEDKIHKRNSVRLAIRKIMYAPS   87 (323)
Q Consensus        17 ~~~G~----h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~~~dk~~k~~tv~l~Ir~l~~~P~   87 (323)
                      -+||.    ..|||-|.=   +---|+     .+.|+...+||.++|++.|...|-..   ...+.|.-+...|.
T Consensus        95 a~pGel~~~~~fpFeF~~---vekpyE-----sY~G~NV~lrY~lkvTv~Rr~~di~k---e~d~~V~~~~~~P~  158 (301)
T KOG3063|consen   95 ARPGELTQSQSFPFEFPH---VEKPYE-----SYIGKNVRLRYFLKVTVSRRLTDIVK---EKDLVVHNLSTYPE  158 (301)
T ss_pred             cCCcceeecccCCccccc---cccchh-----hhcCcceEEEEEEEEEEEechhhhhh---hhheeeEecccCCC
Confidence            35665    568887752   222233     57898889999999999887543222   23455655544443


No 23 
>PF06159 DUF974:  Protein of unknown function (DUF974);  InterPro: IPR010378 This is a family of uncharacterised eukaryotic proteins.
Probab=68.98  E-value=12  Score=35.45  Aligned_cols=59  Identities=22%  Similarity=0.411  Sum_probs=38.2

Q ss_pred             eeEEeecceEEEEEEEecCCcceEeeEEeeeEEEeeCCeE-eeeec-cccC---CCccCCCCccc
Q psy17223        223 LFLYYHGESIAVNVHVANNSNRTVKKIKVSDICLFSTAQY-KCTVA-ETES---DCPIAPVSMFD  282 (323)
Q Consensus       223 ~~~yyhge~i~v~v~v~N~s~k~vkkikv~dv~l~s~~~y-~~~Va-~~e~---~~~i~p~~t~~  282 (323)
                      .+--|=||+...-++|+|++++.|+.+.|. |-|-...+- +-... ..+.   ...+.||.++.
T Consensus         7 fG~iylGEtF~~~l~~~N~s~~~v~~v~ik-vemqT~s~~~r~~L~~~~~~~~~~~~L~p~~~l~   70 (249)
T PF06159_consen    7 FGSIYLGETFSCYLSVNNDSNKPVRNVRIK-VEMQTPSQSLRLPLSDNENSDSPVASLAPGESLD   70 (249)
T ss_pred             cCCEeecCCEEEEEEeecCCCCceEEeEEE-EEEeCCCCCccccCCCCccccccccccCCCCeEe
Confidence            455677999999999999999999999886 233322220 11111 1111   23688998887


No 24 
>PF14796 AP3B1_C:  Clathrin-adaptor complex-3 beta-1 subunit C-terminal
Probab=66.47  E-value=8  Score=34.04  Aligned_cols=54  Identities=15%  Similarity=0.289  Sum_probs=36.3

Q ss_pred             eecceEEEEEEEecCCcceEeeEEeeeEEEeeCCeEeeeeccccCC--CccCCCCccc-ccce
Q psy17223        227 YHGESIAVNVHVANNSNRTVKKIKVSDICLFSTAQYKCTVAETESD--CPIAPVSMFD-TEDL  286 (323)
Q Consensus       227 yhge~i~v~v~v~N~s~k~vkkikv~dv~l~s~~~y~~~Va~~e~~--~~i~p~~t~~-~~~~  286 (323)
                      |+.--+.|.+.++|+|...+++|+|.      +-+..+-+..-|+.  +.+.||++.+ ..||
T Consensus        82 ~s~~mvsIql~ftN~s~~~i~~I~i~------~k~l~~g~~i~~F~~I~~L~pg~s~t~~lgI  138 (145)
T PF14796_consen   82 YSPSMVSIQLTFTNNSDEPIKNIHIG------EKKLPAGMRIHEFPEIESLEPGASVTVSLGI  138 (145)
T ss_pred             CCCCcEEEEEEEEecCCCeecceEEC------CCCCCCCcEeeccCcccccCCCCeEEEEEEE
Confidence            45667899999999999999999997      22222222222332  2688888766 4444


No 25 
>PF07070 Spo0M:  SpoOM protein;  InterPro: IPR009776 This family consists of several bacterial SpoOM proteins which are thought to control sporulation in Bacillus subtilis.Spo0M exerts certain negative effects on sporulation and its gene expression is controlled by sigmaH [].
Probab=66.27  E-value=15  Score=34.44  Aligned_cols=47  Identities=17%  Similarity=0.288  Sum_probs=42.1

Q ss_pred             ecCCceEEEEEeCccceecCCeEEEEEEEeccC-cceeeEEEeecCCC
Q psy17223        103 MSPNKLHLEASLDKELYYHGESIAVNVHVANNS-NRTVKKIKVSDSGA  149 (323)
Q Consensus       103 f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~S-sk~Vk~Ikv~L~q~  149 (323)
                      +.-|-.++...|++..|.|||+|.=.|.|.-.+ ...|.+|.+.|.-.
T Consensus         8 ~GiG~akVDT~L~~~~~~pGe~v~G~V~i~GG~v~Q~I~~I~l~L~t~   55 (218)
T PF07070_consen    8 IGIGGAKVDTVLEKPSVRPGETVRGEVHIKGGSVDQEIDRIYLELVTR   55 (218)
T ss_pred             cCCCCceEEEEECCCCccCCCEEEEEEEEEeCCcceEEeEEEEEEEEE
Confidence            456889999999999999999999999999996 55899999999753


No 26 
>PF10633 NPCBM_assoc:  NPCBM-associated, NEW3 domain of alpha-galactosidase;  InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=64.50  E-value=5.9  Score=30.21  Aligned_cols=29  Identities=24%  Similarity=0.348  Sum_probs=21.1

Q ss_pred             ecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223        120 YHGESIAVNVHVANNSNRTVKKIKVSDSG  148 (323)
Q Consensus       120 ~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q  148 (323)
                      .|||++.+++.|.|.....+..+++.+.-
T Consensus         2 ~~G~~~~~~~tv~N~g~~~~~~v~~~l~~   30 (78)
T PF10633_consen    2 TPGETVTVTLTVTNTGTAPLTNVSLSLSL   30 (78)
T ss_dssp             -TTEEEEEEEEEE--SSS-BSS-EEEEE-
T ss_pred             CCCCEEEEEEEEEECCCCceeeEEEEEeC
Confidence            48999999999999998889888888853


No 27 
>PF07705 CARDB:  CARDB;  InterPro: IPR011635 The APHP (acidic peptide-dependent hydrolases/peptidase) domain is found in a variety of different proteins.; PDB: 2KUT_A 2L0D_A 3IDU_A 2KL6_A.
Probab=60.99  E-value=26  Score=26.77  Aligned_cols=41  Identities=20%  Similarity=0.263  Sum_probs=30.7

Q ss_pred             eEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223        108 LHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSG  148 (323)
Q Consensus       108 I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q  148 (323)
                      +.+.+........+|+.+.|++.|.|.-......+.+.+..
T Consensus         4 L~v~~~~~~~~~~~g~~~~i~~~V~N~G~~~~~~~~v~~~~   44 (101)
T PF07705_consen    4 LTVSITVSPSNVVPGEPVTITVTVKNNGTADAENVTVRLYL   44 (101)
T ss_dssp             EEE-EEEC-SEEETTSEEEEEEEEEE-SSS-BEEEEEEEEE
T ss_pred             EEEEEeeCCCcccCCCEEEEEEEEEECCCCCCCCEEEEEEE
Confidence            34455677778889999999999999988888888888764


No 28 
>PF00963 Cohesin:  Cohesin domain;  InterPro: IPR002102 Cohesin domains interact with a complementary domain, termed the dockerin domain (see IPR002105 from INTERPRO). The cohesin-dockerin interaction is the crucial interaction for complex formation in the cellulosome []. The scaffoldin component of the cellulolytic bacterium Clostridium thermocellum is a non-hydrolytic protein which organises the hydrolytic enzymes in a large complex, called the cellulosome. Scaffoldin comprises a series of functional domains, amongst which is a single cellulose-binding domain and nine cohesin domains which are responsible for integrating the individual enzymatic subunits into the complex.; GO: 0030246 carbohydrate binding, 0000272 polysaccharide catabolic process; PDB: 2BM3_A 3P0D_I 3KCP_A 2B59_A 3L8Q_B 3FNK_C 3GHP_B 2CCL_A 1ANU_A 1OHZ_A ....
Probab=57.66  E-value=28  Score=29.24  Aligned_cols=38  Identities=29%  Similarity=0.315  Sum_probs=31.6

Q ss_pred             EEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223        110 LEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSG  148 (323)
Q Consensus       110 L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q  148 (323)
                      |++.+++.--.|||++.|.+.++|-++. |..+.+.+.=
T Consensus         1 v~l~~~~~~a~~G~tv~V~V~v~~~~~~-i~~~~~~l~y   38 (141)
T PF00963_consen    1 VTLSVDSVSAKPGETVTVPVNVSNVSNS-IAGMQFTLSY   38 (141)
T ss_dssp             EEEEESECEE-TTSEEEEEEEEESCTTT-EEEEEEEEEE
T ss_pred             CEEEeCCceECCCCEEEEEEEEEcCCCc-EEEEEEEEEe
Confidence            5677888888999999999999999766 8888888864


No 29 
>PLN02171 endoglucanase
Probab=54.10  E-value=17  Score=39.17  Aligned_cols=22  Identities=23%  Similarity=0.354  Sum_probs=20.0

Q ss_pred             eEEEEEEEecCCcceEeeEEee
Q psy17223        231 SIAVNVHVANNSNRTVKKIKVS  252 (323)
Q Consensus       231 ~i~v~v~v~N~s~k~vkkikv~  252 (323)
                      -..+.|.|+|+|+|+||.|+|.
T Consensus       554 y~qy~v~I~N~s~~~ik~i~i~  575 (629)
T PLN02171        554 YYRYSTTVTNRSAKTLKELHLG  575 (629)
T ss_pred             EEEEEEEEEECCCCceeeeeee
Confidence            4678889999999999999997


No 30 
>PF07703 A2M_N_2:  Alpha-2-macroglobulin family N-terminal region;  InterPro: IPR011625 This is a domain of the alpha-2-macroglobulin family. The alpha-macroglobulin (aM) family of proteins includes protease inhibitors [], typified by the human tetrameric a2-macroglobulin (a2M); they belong to the MEROPS proteinase inhibitor family I39, clan IL. These protease inhibitors share several defining properties, which include (i) the ability to inhibit proteases from all catalytic classes, (ii) the presence of a 'bait region' and a thiol ester, (iii) a similar protease inhibitory mechanism and (iv) the inactivation of the inhibitory capacity by reaction of the thiol ester with small primary amines. aM protease inhibitors inhibit by steric hindrance []. The mechanism involves protease cleavage of the bait region, a segment of the aM that is particularly susceptible to proteolytic cleavage, which initiates a conformational change such that the aM collapses about the protease. In the resulting aM-protease complex, the active site of the protease is sterically shielded, thus substantially decreasing access to protein substrates. Two additional events occur as a consequence of bait region cleavage, namely (i) the h-cysteinyl-g-glutamyl thiol ester becomes highly reactive and (ii) a major conformational change exposes a conserved COOH-terminal receptor binding domain [] (RBD). RBD exposure allows the aM protease complex to bind to clearance receptors and be removed from circulation []. Tetrameric, dimeric, and, more recently, monomeric aM protease inhibitors have been identified [, ].; PDB: 2QKI_D 3L3O_D 3NMS_A 2ICF_A 2A73_A 2ICE_D 2HR0_A 2A74_A 2XWJ_G 3OHX_A ....
Probab=53.80  E-value=16  Score=30.07  Aligned_cols=32  Identities=22%  Similarity=0.356  Sum_probs=25.7

Q ss_pred             ceEEEEEeCccceecCCeEEEEEEEeccCcce
Q psy17223        107 KLHLEASLDKELYYHGESIAVNVHVANNSNRT  138 (323)
Q Consensus       107 ~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~  138 (323)
                      ..++++..++..|.|||++.+++.....|...
T Consensus        94 ~~~v~l~~~~~~~~Pg~~~~~~i~~~~~s~v~  125 (136)
T PF07703_consen   94 ELKVELTASPDEYKPGEEVTLRIKAPPNSLVG  125 (136)
T ss_dssp             SSSEEEEESSSSBTTTSEEEEEEEESTTEEEE
T ss_pred             cceEEEEEecceeCCCCEEEEEEEeCCCCEEE
Confidence            56788888999999999999999985554433


No 31 
>KOG2625|consensus
Probab=52.44  E-value=8.4  Score=36.74  Aligned_cols=27  Identities=33%  Similarity=0.582  Sum_probs=24.3

Q ss_pred             EeecceEEEEEEEecCCcceEeeEEee
Q psy17223        226 YYHGESIAVNVHVANNSNRTVKKIKVS  252 (323)
Q Consensus       226 yyhge~i~v~v~v~N~s~k~vkkikv~  252 (323)
                      -|-||+...-|+|.|.|+||||.|-+-
T Consensus        11 iflgetfs~yinv~nds~k~v~~i~lk   37 (348)
T KOG2625|consen   11 IFLGETFSFYINVHNDSEKTVKDILLK   37 (348)
T ss_pred             eeeccceEEEEEEecchhhhhhhheee
Confidence            356999999999999999999999884


No 32 
>PF05688 DUF824:  Salmonella repeat of unknown function (DUF824);  InterPro: IPR008542 This family consists of a series of repeated sequences (of around 180 residues) which are found in Salmonella typhimurium, Salmonella typhi and Escherichia coli. These repeats are almost always found with this entry. The repeats are associated with RatA and RatB, the coding sequences of which are found in the pathogeneicity island of Salmonella. The sequences may be determinants of pathogenicity [, ].
Probab=49.71  E-value=18  Score=25.96  Aligned_cols=29  Identities=21%  Similarity=0.280  Sum_probs=26.2

Q ss_pred             ecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223        120 YHGESIAVNVHVANNSNRTVKKIKVSDSG  148 (323)
Q Consensus       120 ~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q  148 (323)
                      .-||+|+++|.+.|.....|-..-+.|.+
T Consensus        10 K~Ge~I~ltVt~kda~G~pv~n~~f~l~r   38 (47)
T PF05688_consen   10 KVGETIPLTVTVKDANGNPVPNAPFTLTR   38 (47)
T ss_pred             ecCCeEEEEEEEECCCCCCcCCceEEEEe
Confidence            45999999999999999999999888876


No 33 
>PF11611 DUF4352:  Domain of unknown function (DUF4352);  InterPro: IPR021652 This entry is represented by Bacteriophage A118, Gp32. The characteristics of the protein distribution suggest prophage matches in addition to the phage matches. This entry represents a group of putative lipoproteins of unknown function.; PDB: 3CFU_A.
Probab=49.45  E-value=30  Score=27.78  Aligned_cols=53  Identities=21%  Similarity=0.359  Sum_probs=30.2

Q ss_pred             cceEEEEEEEecCCcceEeeEEeeeEEEeeC--CeEeeeeccccC-----CCccCCCCccc
Q psy17223        229 GESIAVNVHVANNSNRTVKKIKVSDICLFST--AQYKCTVAETES-----DCPIAPVSMFD  282 (323)
Q Consensus       229 ge~i~v~v~v~N~s~k~vkkikv~dv~l~s~--~~y~~~Va~~e~-----~~~i~p~~t~~  282 (323)
                      +..+.|+|.|.|++++.+- +-..++.|+..  .+|.........     ...|.||.+.+
T Consensus        35 ~~fv~v~v~v~N~~~~~~~-~~~~~f~l~d~~g~~~~~~~~~~~~~~~~~~~~i~pG~~~~   94 (123)
T PF11611_consen   35 NKFVVVDVTVKNNGDEPLD-FSPSDFKLYDSDGNKYDPDFSASSNDNDLFSETIKPGESVT   94 (123)
T ss_dssp             SEEEEEEEEEEE-SSS-EE-EEGGGEEEE-TT--B--EEE-CCCTTTB--EEEE-TT-EEE
T ss_pred             CEEEEEEEEEEECCCCcEE-ecccceEEEeCCCCEEcccccchhccccccccEECCCCEEE
Confidence            5679999999999888774 43348889833  345543333332     35899998866


No 34 
>PF04744 Monooxygenase_B:  Monooxygenase subunit B protein;  InterPro: IPR006833 Ammonia monooxygenase and the particulate methane monooxygenase are both integral membrane proteins, occurring in ammonia oxidisers and methanotrophs respectively, which are thought to be evolutionarily related []. These enzymes have a relatively wide substrate specificity and can catalyse the oxidation of a range of substrates including ammonia, methane, halogenated hydrocarbons and aromatic molecules []. These enzymes are composed of 3 subunits - A (IPR003393 from INTERPRO), B (IPR006833 from INTERPRO) and C (IPR006980 from INTERPRO) - and contain various metal centres, including copper. Particulate methane monooxygenase from Methylococcus capsulatus str. Bath is an ABC homotrimer, which contains mononuclear and dinuclear copper metal centres, and a third metal centre containing a metal ion whose identity in vivo is not certain[]. The soluble regions of these enzymes derive primarily from the B subunit. This subunit forms two antiparallel beta-barrel-like structures and contains the mono- and di- nuclear copper metal centres [].; PDB: 3CHX_E 3RFR_A 3RGB_A 1YEW_A.
Probab=48.91  E-value=20  Score=36.21  Aligned_cols=37  Identities=16%  Similarity=0.328  Sum_probs=29.7

Q ss_pred             ceEEEEEeCccce-ecCCeEEEEEEEeccCcceeeEEE
Q psy17223        107 KLHLEASLDKELY-YHGESIAVNVHVANNSNRTVKKIK  143 (323)
Q Consensus       107 ~I~L~a~LdK~~Y-~PGE~I~V~v~IdN~Ssk~Vk~Ik  143 (323)
                      +-.+++.+.+-.| .||-++.++++|+|+++..|+==+
T Consensus       246 ~~~V~~~v~~A~Y~vpgR~l~~~l~VtN~g~~pv~Lge  283 (381)
T PF04744_consen  246 PNSVKVKVTDATYRVPGRTLTMTLTVTNNGDSPVRLGE  283 (381)
T ss_dssp             -SSEEEEEEEEEEESSSSEEEEEEEEEEESSS-BEEEE
T ss_pred             CCceEEEEeccEEecCCcEEEEEEEEEcCCCCceEeee
Confidence            3338888888888 899999999999999999887544


No 35 
>COG2373 Large extracellular alpha-helical protein [General function prediction only]
Probab=48.87  E-value=33  Score=40.93  Aligned_cols=131  Identities=17%  Similarity=0.189  Sum_probs=72.6

Q ss_pred             ceEEEEEeCccceecCCeEEEEEEEeccCcc-eeeEEEeec--CCCccchhhhhhhccCCCCCCccC------CCccc-c
Q psy17223        107 KLHLEASLDKELYYHGESIAVNVHVANNSNR-TVKKIKVSD--SGAEDDQDLKDELADSDIDGMEED------DLPNI-K  176 (323)
Q Consensus       107 ~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk-~Vk~Ikv~L--~q~ed~~d~~~~~~~~di~~~~e~------~lp~~-~  176 (323)
                      .+++-++-||..|.|||++.+.+-....-.+ .+..+-+++  .+ +| +   ..++-..+..++++      .||.. +
T Consensus       393 ~~k~y~ftDRglYRpGE~v~~~~~~R~~~~~~a~~~~p~~l~v~~-Pd-G---~~~~~~~~~~~~~G~~~~~~~l~~na~  467 (1621)
T COG2373         393 GLKVYLFTDRGLYRPGETVHVNALLRDFDGKTALDNQPLKLRVLD-PD-G---SVLRTLTITLDEEGLYELSFPLPENAL  467 (1621)
T ss_pred             ceEEEEecCcccCCCCceeeeeeeehhhcccccccCCCeEEEEEC-CC-C---cEEEEEEEeccccCceEEeeeCCCCCC
Confidence            5788899999999999999999988776655 444443333  32 11 1   22333333332222      44433 1


Q ss_pred             ccccccceeeecccccCCCCCCCCCCcchhhhhhccCcc--hhhccceeeEEeecceEEEEEEEecCCcceEeeEEee
Q psy17223        177 AWGKNKRMYYNTDYVDDDHGGIQSGTGFVWFEWLKKGSK--EKSKKKYLFLYYHGESIAVNVHVANNSNRTVKKIKVS  252 (323)
Q Consensus       177 awg~~~~~~y~td~~~~~~~~~~~~~~~~~~e~~~~~~~--~~s~~k~~~~yyhge~i~v~v~v~N~s~k~vkkikv~  252 (323)
                      .=|=.-+++++..      . ...+..+--.+++.+ +.  ..+++|  ..|.||+++.++|...|-.-.=+.+-++.
T Consensus       468 tG~w~l~~~~~~~------~-~~~s~~f~V~df~p~-r~~i~l~~~k--~~~~~g~~v~~~v~~~yL~GaPa~g~~~~  535 (1621)
T COG2373         468 TGGYTLELYTGGK------S-AVISMSFRVEDFIPD-RFKINLTLDK--TEWVPGKDVKIKVDLRYLYGAPAAGLTVQ  535 (1621)
T ss_pred             cceEEEEEEeCCc------c-ceeeeeEEhhHhCCc-eEEEeccccc--ccccCCCcEEEEEEEEecCCCcccCceee
Confidence            1111122333111      0 222222222222322 33  456666  55999999999999988876666666654


No 36 
>PF09478 CBM49:  Carbohydrate binding domain CBM49;  InterPro: IPR019028 A carbohydrate-binding module (CBM) is defined as a contiguous amino acid sequence within a carbohydrate-active enzyme with a discreet fold having carbohydrate-binding activity. A few exceptions are CBMs in cellulosomal scaffolding proteins and rare instances of independent putative CBMs. The requirement of CBMs existing as modules within larger enzymes sets this class of carbohydrate-binding protein apart from other non-catalytic sugar binding proteins such as lectins and sugar transport proteins. CBMs were previously classified as cellulose-binding domains (CBDs) based on the initial discovery of several modules that bound cellulose [, ]. However, additional modules in carbohydrate-active enzymes are continually being found that bind carbohydrates other than cellulose yet otherwise meet the CBM criteria, hence the need to reclassify these polypeptides using more inclusive terminology. Previous classification of cellulose-binding domains were based on amino acid similarity. Groupings of CBDs were called "Types" and numbered with roman numerals (e.g. Type I or Type II CBDs). In keeping with the glycoside hydrolase classification, these groupings are now called families and numbered with Arabic numerals. Families 1 to 13 are the same as Types I to XIII. For a detailed review on the structure and binding modes of CBMs see [].  This domain is found at the C-terminal of cellulases and in vitro binding studies have shown it to binds to crystalline cellulose []. ; GO: 0030246 carbohydrate binding, 0005576 extracellular region
Probab=48.15  E-value=50  Score=25.53  Aligned_cols=39  Identities=21%  Similarity=0.354  Sum_probs=30.2

Q ss_pred             EEEEEeCccceecCC-eEEEEEEEeccCcceeeEEEeecC
Q psy17223        109 HLEASLDKELYYHGE-SIAVNVHVANNSNRTVKKIKVSDS  147 (323)
Q Consensus       109 ~L~a~LdK~~Y~PGE-~I~V~v~IdN~Ssk~Vk~Ikv~L~  147 (323)
                      +++-.+......-|. -....+.|.|+++++|+.+.+...
T Consensus         2 ~i~q~~~~sW~~~g~~y~qy~v~I~N~~~~~I~~~~i~~~   41 (80)
T PF09478_consen    2 TITQTLVNSWTENGQTYTQYDVTITNNGSKPIKSLKISID   41 (80)
T ss_pred             EEEEEEEeEEEeCCEEEEEEEEEEEECCCCeEEEEEEEEC
Confidence            455555666666665 357899999999999999999886


No 37 
>PF07703 A2M_N_2:  Alpha-2-macroglobulin family N-terminal region;  InterPro: IPR011625 This is a domain of the alpha-2-macroglobulin family. The alpha-macroglobulin (aM) family of proteins includes protease inhibitors [], typified by the human tetrameric a2-macroglobulin (a2M); they belong to the MEROPS proteinase inhibitor family I39, clan IL. These protease inhibitors share several defining properties, which include (i) the ability to inhibit proteases from all catalytic classes, (ii) the presence of a 'bait region' and a thiol ester, (iii) a similar protease inhibitory mechanism and (iv) the inactivation of the inhibitory capacity by reaction of the thiol ester with small primary amines. aM protease inhibitors inhibit by steric hindrance []. The mechanism involves protease cleavage of the bait region, a segment of the aM that is particularly susceptible to proteolytic cleavage, which initiates a conformational change such that the aM collapses about the protease. In the resulting aM-protease complex, the active site of the protease is sterically shielded, thus substantially decreasing access to protein substrates. Two additional events occur as a consequence of bait region cleavage, namely (i) the h-cysteinyl-g-glutamyl thiol ester becomes highly reactive and (ii) a major conformational change exposes a conserved COOH-terminal receptor binding domain [] (RBD). RBD exposure allows the aM protease complex to bind to clearance receptors and be removed from circulation []. Tetrameric, dimeric, and, more recently, monomeric aM protease inhibitors have been identified [, ].; PDB: 2QKI_D 3L3O_D 3NMS_A 2ICF_A 2A73_A 2ICE_D 2HR0_A 2A74_A 2XWJ_G 3OHX_A ....
Probab=46.46  E-value=24  Score=29.06  Aligned_cols=25  Identities=36%  Similarity=0.444  Sum_probs=20.2

Q ss_pred             EEEEeCccceecCCeEEEEEEEecc
Q psy17223        110 LEASLDKELYYHGESIAVNVHVANN  134 (323)
Q Consensus       110 L~a~LdK~~Y~PGE~I~V~v~IdN~  134 (323)
                      |++.+||..|.|||++.+.+.-.-.
T Consensus         1 l~i~~~~~~~~~Ge~~~v~v~~~~~   25 (136)
T PF07703_consen    1 LQISTDKDSYKPGETAKVTVQSPFP   25 (136)
T ss_dssp             EEEEE-SSSB-TTSEEEEEEEEESC
T ss_pred             CEEEcCCCCcCCCCEEEEEEEcCCC
Confidence            6789999999999999999887665


No 38 
>smart00809 Alpha_adaptinC2 Adaptin C-terminal domain. Adaptins are components of the adaptor complexes which link clathrin to receptors in coated vesicles. Clathrin-associated protein complexes are believed to interact with the cytoplasmic tails of membrane proteins, leading to their selection and concentration. Gamma-adaptin is a subunit of the golgi adaptor. Alpha adaptin is a heterotetramer that regulates clathrin-bud formation. The carboxyl-terminal appendage of the alpha subunit regulates translocation of endocytic accessory proteins to the bud site. This Ig-fold domain is found in alpha, beta and gamma adaptins and consists of a beta-sandwich containing 7 strands in 2 beta-sheets in a greek-key topology PUBMED:10430869, PUBMED:12176391. The adaptor appendage contains an additional N-terminal strand.
Probab=43.15  E-value=71  Score=25.12  Aligned_cols=41  Identities=12%  Similarity=0.259  Sum_probs=33.4

Q ss_pred             ecCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223        103 MSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDS  147 (323)
Q Consensus       103 f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~  147 (323)
                      +.+..+++.+.+.++    +..+.|.+.+.|.|..+|.++.+.+.
T Consensus         2 ~~~~~l~I~~~~~~~----~~~~~i~~~~~N~s~~~it~f~~~~a   42 (104)
T smart00809        2 YEKNGLQIGFKFERR----PGLIRITLTFTNKSPSPITNFSFQAA   42 (104)
T ss_pred             ccCCCEEEEEEEEcC----CCeEEEEEEEEeCCCCeeeeEEEEEE
Confidence            445668888888775    45689999999999999999998875


No 39 
>TIGR03079 CH4_NH3mon_ox_B methane monooxygenase/ammonia monooxygenase, subunit B. Both ammonia oxidizers such as Nitrosomonas europaea and methanotrophs (obligate methane oxidizers) such as Methylococcus capsulatus each can grow only on their own characteristic substrate. However, both groups have the ability to oxidize both substrates, and so the relevant enzymes must be named here according to their ability to oxidze both. The protein family represented here reflects subunit B of both the particulate methane monooxygenase of methylotrophs and the ammonia monooxygenase of nitrifying bacteria.
Probab=42.03  E-value=37  Score=34.41  Aligned_cols=36  Identities=17%  Similarity=0.350  Sum_probs=29.0

Q ss_pred             CCceEEEEEeCccce-ecCCeEEEEEEEeccCcceee
Q psy17223        105 PNKLHLEASLDKELY-YHGESIAVNVHVANNSNRTVK  140 (323)
Q Consensus       105 sG~I~L~a~LdK~~Y-~PGE~I~V~v~IdN~Ssk~Vk  140 (323)
                      -++-.+++.+.+..| +||-++.++++|+|+++..|+
T Consensus       263 ~~~~~V~~kv~~a~Y~VPGR~l~~~~~VTN~g~~~vr  299 (399)
T TIGR03079       263 VAPNPVSINVTKANYDVPGRALRVTMEITNNGDQVIS  299 (399)
T ss_pred             CCCCceEEEEeccEEecCCcEEEEEEEEEcCCCCceE
Confidence            344456666667666 899999999999999999886


No 40 
>PF02014 Reeler:  Reeler domain Schematic picture including Reeler domain;  InterPro: IPR002861 Extracellular matrix (ECM) proteins play an important role in early cortical development, specifically in the formation of neural connections and in controlling the cyto-architecture of the central nervous system. The product of the reeler gene in mouse is reelin,a large extracellular protein secreted by pioneer neurons that coordinates cell positioning during neurodevelopment []. F-spondin and mindin are a family of matrix-attached adhesion molecules that share structural similarities and overlapping domains of expression. Both F-spondin and mindin promote adhesion and outgrowth of hippocampal embryonic neurons and bind to a putative receptor(s) expressed on both hippocampal and sensory neurons []. This domain of unknown function is found at the N terminus of reelin and F-spondin.; PDB: 2ZOT_B 2ZOU_B 3COO_A.
Probab=41.79  E-value=43  Score=28.07  Aligned_cols=35  Identities=11%  Similarity=0.271  Sum_probs=25.3

Q ss_pred             EEeCccceecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223        112 ASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSG  148 (323)
Q Consensus       112 a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q  148 (323)
                      +.++...|.||+.+.|++  .+.++...+++-++...
T Consensus        23 i~~~~~~y~pg~~~~Vtl--~~~~~~~F~GFllqAr~   57 (132)
T PF02014_consen   23 ISVSPSSYEPGQTYTVTL--SSSGSSSFRGFLLQARD   57 (132)
T ss_dssp             EEET-SSB-TTBEEEEEE--EETTTEEBSEEEEEEEE
T ss_pred             EEeCCCeEcCCCEEEEEE--ECCCCCceeEEEEEEEe
Confidence            444599999999999998  66667788887776643


No 41 
>PF05753 TRAP_beta:  Translocon-associated protein beta (TRAPB);  InterPro: IPR008856 This family consists of several eukaryotic translocon-associated protein beta (TRAPB) or signal sequence receptor beta subunit (SSR-beta) proteins. The normal translocation of nascent polypeptides into the lumen of the endoplasmic reticulum (ER) is thought to be aided in part by a translocon-associated protein (TRAP) complex consisting of 4 protein subunits. The association of mature proteins with the ER and Golgi, or other intracellular locales, such as lysosomes, depends on the initial targeting of the nascent polypeptide to the ER membrane. A similar scenario must also exist for proteins destined for secretion [].; GO: 0005783 endoplasmic reticulum, 0016021 integral to membrane
Probab=40.51  E-value=57  Score=29.54  Aligned_cols=74  Identities=16%  Similarity=0.196  Sum_probs=45.7

Q ss_pred             Eee-cceEEEEEEEecCCcceEeeEEeeeEEEeeCCeEeeeecc-ccC-CCccCCCCccc-ccceeecccccccccCCcc
Q psy17223        226 YYH-GESIAVNVHVANNSNRTVKKIKVSDICLFSTAQYKCTVAE-TES-DCPIAPVSMFD-TEDLAMLRHGFKRMFGHAF  301 (323)
Q Consensus       226 yyh-ge~i~v~v~v~N~s~k~vkkikv~dv~l~s~~~y~~~Va~-~e~-~~~i~p~~t~~-~~~~~~~~~~~~~~~~~a~  301 (323)
                      |.. |+.+.|++.|-|.=+.+.-++++.| -=|..+.|. .|.. ... =+.|+||++.+ .+-|.+   .+...||+.|
T Consensus        33 ~~v~g~~v~V~~~iyN~G~~~A~dV~l~D-~~fp~~~F~-lvsG~~s~~~~~i~pg~~vsh~~vv~p---~~~G~f~~~~  107 (181)
T PF05753_consen   33 YLVEGEDVTVTYTIYNVGSSAAYDVKLTD-DSFPPEDFE-LVSGSLSASWERIPPGENVSHSYVVRP---KKSGYFNFTP  107 (181)
T ss_pred             cccCCcEEEEEEEEEECCCCeEEEEEEEC-CCCCccccE-eccCceEEEEEEECCCCeEEEEEEEee---eeeEEEEccC
Confidence            444 9999999999999999999999985 111112221 1111 011 14888998876 444443   4455666666


Q ss_pred             ccc
Q psy17223        302 STS  304 (323)
Q Consensus       302 ~t~  304 (323)
                      .++
T Consensus       108 a~V  110 (181)
T PF05753_consen  108 AVV  110 (181)
T ss_pred             EEE
Confidence            544


No 42 
>PF09624 DUF2393:  Protein of unknown function (DUF2393);  InterPro: IPR013417  The function of this protein is unknown. It is always found as part of a two-gene operon with IPR013416 from INTERPRO, a protein that appears to span the membrane seven times. It has so far been found in the bacteria Anabaena sp. (strain PCC 7120), Agrobacterium tumefaciens, Rhizobium meliloti, and Gloeobacter violaceus.
Probab=40.33  E-value=68  Score=27.40  Aligned_cols=29  Identities=31%  Similarity=0.292  Sum_probs=25.3

Q ss_pred             eecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223        119 YYHGESIAVNVHVANNSNRTVKKIKVSDS  147 (323)
Q Consensus       119 Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~  147 (323)
                      ..-+|.+-|...|.|.++++++++++.+.
T Consensus        58 l~~~~~~~v~g~V~N~g~~~i~~c~i~~~   86 (149)
T PF09624_consen   58 LQYSESFYVDGTVTNTGKFTIKKCKITVK   86 (149)
T ss_pred             eeeccEEEEEEEEEECCCCEeeEEEEEEE
Confidence            44689999999999999999998888774


No 43 
>TIGR01451 B_ant_repeat conserved repeat domain. This model represents the conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis, and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydial outer membrane proteins.
Probab=40.07  E-value=56  Score=23.46  Aligned_cols=29  Identities=31%  Similarity=0.475  Sum_probs=25.0

Q ss_pred             EeecceEEEEEEEecCCcceEeeEEeeeE
Q psy17223        226 YYHGESIAVNVHVANNSNRTVKKIKVSDI  254 (323)
Q Consensus       226 yyhge~i~v~v~v~N~s~k~vkkikv~dv  254 (323)
                      ..=|+.|...|.|.|+.......++|.|.
T Consensus         8 ~~~Gd~v~Yti~v~N~g~~~a~~v~v~D~   36 (53)
T TIGR01451         8 ATIGDTITYTITVTNNGNVPATNVVVTDI   36 (53)
T ss_pred             cCCCCEEEEEEEEEECCCCceEeEEEEEc
Confidence            34599999999999999999999988743


No 44 
>cd08544 Reeler Reeler, the N-terminal domain of reelin, F-spondin, and a variety of other proteins. This domain is found at the N-terminus of F-spondin, a protein attached to the extracellular matrix, which plays roles in neuronal development and vascular remodelling. The F-spondin reeler domain has been reported to bind heparin. The reeler domain is also found at the N-terminus of reelin, an extracellular glycoprotein involved in the development of the brain cortex, and in a variety of other eukaryotic proteins with different domain architectures, including the animal ferric-chelate reductase 1 or stromal cell-derived receptor 2, a member of the cytochrome B561 family, which reduces ferric iron before its transport from the endosome to the cytoplasm. Also included is the insect putative defense protein 1, which is expressed upon bacterial infection and appears to contain a single reeler domain.
Probab=38.12  E-value=56  Score=27.34  Aligned_cols=37  Identities=11%  Similarity=0.222  Sum_probs=27.1

Q ss_pred             EEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223        110 LEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSG  148 (323)
Q Consensus       110 L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q  148 (323)
                      ..+.++...|.|||.+.|++.-.|.  ..-+++-++...
T Consensus        21 y~i~~~~~~y~pG~~~~Vtl~~~~~--~~F~GF~lqAr~   57 (135)
T cd08544          21 YSITISGNSYVPGETYTVTLSGSSP--SPFRGFLLQARD   57 (135)
T ss_pred             EEEEeCCCEECCCCEEEEEEECCCC--CceeEEEEEEEc
Confidence            4556667799999999999988776  456666555543


No 45 
>PF01345 DUF11:  Domain of unknown function DUF11;  InterPro: IPR001434 This group of sequences is represented by a conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis (Streptococcus faecalis), and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydia trachomatis outer membrane proteins.  In C. trachomatis, three cysteine-rich proteins (also believed to be lipoproteins), MOMP, OMP6 and OMP3, make up the extracellular matrix of the outer membrane []. They are involved in the essential structural integrity of both the elementary body (EB) and recticulate body (RB) phase. They are thought to be involved in porin formation and, as these bacteria lack the peptidoglycan layer common to most Gram-negative microbes, such proteins are highly important in the pathogenicity of the organism.; GO: 0005727 extrachromosomal circular DNA
Probab=38.10  E-value=57  Score=24.45  Aligned_cols=30  Identities=17%  Similarity=0.304  Sum_probs=26.5

Q ss_pred             eEEeecceEEEEEEEecCCcceEeeEEeee
Q psy17223        224 FLYYHGESIAVNVHVANNSNRTVKKIKVSD  253 (323)
Q Consensus       224 ~~yyhge~i~v~v~v~N~s~k~vkkikv~d  253 (323)
                      ....=||.+...|.|+|..+....+++|.|
T Consensus        35 ~~~~~Gd~v~ytitvtN~G~~~a~nv~v~D   64 (76)
T PF01345_consen   35 STANPGDTVTYTITVTNTGPAPATNVVVTD   64 (76)
T ss_pred             CcccCCCEEEEEEEEEECCCCeeEeEEEEE
Confidence            445669999999999999999999999974


No 46 
>PF06030 DUF916:  Bacterial protein of unknown function (DUF916);  InterPro: IPR010317 This family consists of putative cell surface proteins, from Firmicutes, of unknown function. 
Probab=36.91  E-value=37  Score=28.67  Aligned_cols=29  Identities=31%  Similarity=0.528  Sum_probs=24.3

Q ss_pred             ecCCeEEEEEEEeccCcceeeEEEeecCCC
Q psy17223        120 YHGESIAVNVHVANNSNRTVKKIKVSDSGA  149 (323)
Q Consensus       120 ~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~  149 (323)
                      -||+...+++.|.|.|.+.++ +.+.+..+
T Consensus        24 ~P~q~~~l~v~i~N~s~~~~t-v~v~~~~A   52 (121)
T PF06030_consen   24 KPGQKQTLEVRITNNSDKEIT-VKVSANTA   52 (121)
T ss_pred             CCCCEEEEEEEEEeCCCCCEE-EEEEEeee
Confidence            489999999999999998876 77777654


No 47 
>PF14796 AP3B1_C:  Clathrin-adaptor complex-3 beta-1 subunit C-terminal
Probab=36.11  E-value=85  Score=27.66  Aligned_cols=46  Identities=15%  Similarity=0.340  Sum_probs=38.7

Q ss_pred             ecCCceEEEEEeCcccee-cCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223        103 MSPNKLHLEASLDKELYY-HGESIAVNVHVANNSNRTVKKIKVSDSG  148 (323)
Q Consensus       103 f~sG~I~L~a~LdK~~Y~-PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q  148 (323)
                      ...+=+.++-++.|.-+. ..-.+.|++.+.|+|...|++|.+....
T Consensus        64 v~G~GL~v~Y~F~RqP~~~s~~mvsIql~ftN~s~~~i~~I~i~~k~  110 (145)
T PF14796_consen   64 VNGKGLSVEYRFSRQPSLYSPSMVSIQLTFTNNSDEPIKNIHIGEKK  110 (145)
T ss_pred             cCCCceeEEEEEccCCcCCCCCcEEEEEEEEecCCCeecceEECCCC
Confidence            356778999999998884 4568899999999999999999987653


No 48 
>PF13199 Glyco_hydro_66:  Glycosyl hydrolase family 66; PDB: 3VMO_A 3VMN_A 3VMP_A.
Probab=36.04  E-value=53  Score=34.97  Aligned_cols=25  Identities=24%  Similarity=0.460  Sum_probs=16.2

Q ss_pred             EeCccceecCCeEEEEEEEeccCcc
Q psy17223        113 SLDKELYYHGESIAVNVHVANNSNR  137 (323)
Q Consensus       113 ~LdK~~Y~PGE~I~V~v~IdN~Ssk  137 (323)
                      +.||..|.|||+|.+++...|....
T Consensus         1 ~tDKA~Y~PGe~V~l~~~~~~~~~~   25 (559)
T PF13199_consen    1 TTDKARYRPGEKVTLTASLKNTTGS   25 (559)
T ss_dssp             EES-SSB-TTS-EEEE-EEE--SSS
T ss_pred             CCCcceeCCCCeEEEEEEeccCccc
Confidence            4689999999999999999998544


No 49 
>PF09624 DUF2393:  Protein of unknown function (DUF2393);  InterPro: IPR013417  The function of this protein is unknown. It is always found as part of a two-gene operon with IPR013416 from INTERPRO, a protein that appears to span the membrane seven times. It has so far been found in the bacteria Anabaena sp. (strain PCC 7120), Agrobacterium tumefaciens, Rhizobium meliloti, and Gloeobacter violaceus.
Probab=35.58  E-value=52  Score=28.17  Aligned_cols=34  Identities=35%  Similarity=0.428  Sum_probs=27.7

Q ss_pred             hhhccceeeEEeecceEEEEEEEecCCcceEeeEEee
Q psy17223        216 EKSKKKYLFLYYHGESIAVNVHVANNSNRTVKKIKVS  252 (323)
Q Consensus       216 ~~s~~k~~~~yyhge~i~v~v~v~N~s~k~vkkikv~  252 (323)
                      ....++++. |  +|.+-|...|+|.+++++++.+|.
T Consensus        51 ~~~~~~~l~-~--~~~~~v~g~V~N~g~~~i~~c~i~   84 (149)
T PF09624_consen   51 TLTSQKRLQ-Y--SESFYVDGTVTNTGKFTIKKCKIT   84 (149)
T ss_pred             EEeeeeeee-e--ccEEEEEEEEEECCCCEeeEEEEE
Confidence            344555533 3  899999999999999999999998


No 50 
>cd08548 Type_I_cohesin_like Type I cohesin domain, interaction partner of dockerin. Bacterial cohesin domains bind to a complementary protein domain named dockerin, and this interaction is required for the formation of the cellulosome, a cellulose-degrading complex. The cellulosome consists of scaffoldin, a noncatalytic scaffolding polypeptide, that comprises repeating cohesion modules and a single carbohydrate-binding module (CBM). Specific calcium-dependent interactions between cohesins and dockerins appear to be essential for cellulosome assembly. This subfamily represents type I cohesins; their interactions with dockerin mediate assembly of a range of dockerin-borne enzymes to the complex.
Probab=34.79  E-value=76  Score=27.07  Aligned_cols=40  Identities=15%  Similarity=0.160  Sum_probs=34.1

Q ss_pred             EEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCC
Q psy17223        110 LEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGA  149 (323)
Q Consensus       110 L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~  149 (323)
                      +++.+++---.||+++.|-|.++|-.+..|..+.+.+.-+
T Consensus         1 ~~v~ig~v~~~~G~tv~VpV~~~~v~~~~i~~~~f~l~yD   40 (135)
T cd08548           1 VEVKIGSVSGKPGDTVTVPVTLSNVPSKGIGACDFVLSYD   40 (135)
T ss_pred             CeEEeccEEecCCCEEEEEEEEecCCccCEEEEEEEEEeC
Confidence            3566777777899999999999999999999999988753


No 51 
>PF06159 DUF974:  Protein of unknown function (DUF974);  InterPro: IPR010378 This is a family of uncharacterised eukaryotic proteins.
Probab=34.01  E-value=51  Score=31.16  Aligned_cols=28  Identities=29%  Similarity=0.597  Sum_probs=25.3

Q ss_pred             eecCCeEEEEEEEeccCcceeeEEEeec
Q psy17223        119 YYHGESIAVNVHVANNSNRTVKKIKVSD  146 (323)
Q Consensus       119 Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L  146 (323)
                      .|-||+....++|.|.|+..|+.+.++.
T Consensus        10 iylGEtF~~~l~~~N~s~~~v~~v~ikv   37 (249)
T PF06159_consen   10 IYLGETFSCYLSVNNDSNKPVRNVRIKV   37 (249)
T ss_pred             EeecCCEEEEEEeecCCCCceEEeEEEE
Confidence            4679999999999999999999887776


No 52 
>COG4326 Spo0M Sporulation control protein [General function prediction only]
Probab=34.00  E-value=57  Score=30.85  Aligned_cols=21  Identities=29%  Similarity=0.523  Sum_probs=17.9

Q ss_pred             cCCCeeeeeEeeCCCCCCCcE
Q psy17223         18 KLGPNAFPFFFELPPSCPASV   38 (323)
Q Consensus        18 ~~G~h~FPFsFqLP~~LP~SF   38 (323)
                      |--.+.|||+|.||-+.|=+|
T Consensus       104 pgEe~~fpf~l~lP~~tPvT~  124 (270)
T COG4326         104 PGEERNFPFELSLPWNTPVTI  124 (270)
T ss_pred             CCceEeccEEEecCCCCceee
Confidence            334599999999999999887


No 53 
>PF02883 Alpha_adaptinC2:  Adaptin C-terminal domain;  InterPro: IPR008152 Proteins synthesized on the ribosome and processed in the endoplasmic reticulum are transported from the Golgi apparatus to the trans-Golgi network (TGN), and from there via small carrier vesicles to their final destination compartment. These vesicles have specific coat proteins (such as clathrin or coatomer) that are important for cargo selection and direction of transport []. Clathrin coats contain both clathrin (acts as a scaffold) and adaptor complexes that link clathrin to receptors in coated vesicles. Clathrin-associated protein complexes are believed to interact with the cytoplasmic tails of membrane proteins, leading to their selection and concentration. The two major types of clathrin adaptor complexes are the heterotetrameric adaptor protein (AP) complexes, and the monomeric GGA (Golgi-localising, Gamma-adaptin ear domain homology, ARF-binding proteins) adaptors [, ]. AP (adaptor protein) complexes are found in coated vesicles and clathrin-coated pits. AP complexes connect cargo proteins and lipids to clathrin at vesicle budding sites, as well as binding accessory proteins that regulate coat assembly and disassembly (such as AP180, epsins and auxilin). There are different AP complexes in mammals. AP1 is responsible for the transport of lysosomal hydrolases between the TGN and endosomes []. AP2 associates with the plasma membrane and is responsible for endocytosis []. AP3 is responsible for protein trafficking to lysosomes and other related organelles []. AP4 is less well characterised. AP complexes are heterotetramers composed of two large subunits (adaptins), a medium subunit (mu) and a small subunit (sigma). For example, in AP1 these subunits are gamma-1-adaptin, beta-1-adaptin, mu-1 and sigma-1, while in AP2 they are alpha-adaptin, beta-2-adaptin, mu-2 and sigma-2. Each subunit has a specific function. Adaptins recognise and bind to clathrin through their hinge region (clathrin box), and recruit accessory proteins that modulate AP function through their C-terminal ear (appendage) domains. Mu recognises tyrosine-based sorting signals within the cytoplasmic domains of transmembrane cargo proteins []. One function of clathrin and AP2 complex-mediated endocytosis is to regulate the number of GABA(A) receptors available at the cell surface [].  GGAs (Golgi-localising, Gamma-adaptin ear domain homology, ARF-binding proteins) are a family of monomeric clathrin adaptor proteins that are conserved from yeasts to humans. GGAs regulate clathrin-mediated the transport of proteins (such as mannose 6-phosphate receptors) from the TGN to endosomes and lysosomes through interactions with TGN-sorting receptors, sometimes in conjunction with AP-1 [, ]. GGAs bind cargo, membranes, clathrin and accessory factors. GGA1, GGA2 and GGA3 all contain a domain homologous to the ear domain of gamma-adaptin. GGAs are composed of a single polypeptide with four domains: an N-terminal VHS (Vps27p/Hrs/Stam) domain, a GAT (GGA and Tom1) domain, a hinge region, and a C-terminal GAE (gamma-adaptin ear) domain. The VHS domain is responsible for endocytosis and signal transduction, recognising transmembrane cargo through the ACLL sequence in the cytoplasmic domains of sorting receptors []. The GAT domain (also found in Tom1 proteins) interacts with ARF (ADP-ribosylation factor) to regulate membrane trafficking [], and with ubiquitin for receptor sorting []. The hinge region contains a clathrin box for recognition and binding to clathrin, similar to that found in AP adaptins. The GAE domain is similar to the AP gamma-adaptin ear domain, and is responsible for the recruitment of accessory proteins that regulate clathrin-mediated endocytosis [].  This entry represents a beta-sandwich structural motif found in the appendage (ear) domain of alpha-, beta- and gamma-adaptin from AP clathrin adaptor complexes, and the GAE (gamma-adaptin ear) domain of GGA adaptor proteins. These domains have an immunoglobulin-like beta-sandwich fold containing 7 or 8 strands in 2 beta-sheets in a Greek key topology [, ]. Although these domains share a similar fold, there is little sequence identity between the alpha/beta-adaptins and gamma-adaptin/GAE. More information about these proteins can be found at Protein of the Month: Clathrin [].; GO: 0006886 intracellular protein transport, 0016192 vesicle-mediated transport, 0030131 clathrin adaptor complex; PDB: 3MNM_B 3ZY7_B 1GYU_A 1GYW_B 2A7B_A 1GYV_A 2E9G_A 1E42_B 2G30_A 2IV9_B ....
Probab=33.97  E-value=1.4e+02  Score=24.02  Aligned_cols=45  Identities=11%  Similarity=0.179  Sum_probs=36.1

Q ss_pred             EeEecCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeec
Q psy17223        100 EFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSD  146 (323)
Q Consensus       100 ~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L  146 (323)
                      ..++.+..++|.+.+.+  -..+..+.|.+.+.|.+...|.++.+.+
T Consensus         3 ~~~ye~~~l~I~~~~~~--~~~~~~~~i~~~f~N~s~~~it~f~~q~   47 (115)
T PF02883_consen    3 GVLYEDNGLQIGFKSEK--SPNPNQGRIKLTFGNKSSQPITNFSFQA   47 (115)
T ss_dssp             EEEEEETTEEEEEEEEE--CCETTEEEEEEEEEE-SSS-BEEEEEEE
T ss_pred             EEEEeCCCEEEEEEEEe--cCCCCEEEEEEEEEECCCCCcceEEEEE
Confidence            34677788888888887  4567889999999999999999999987


No 54 
>COG2373 Large extracellular alpha-helical protein [General function prediction only]
Probab=32.00  E-value=91  Score=37.36  Aligned_cols=41  Identities=17%  Similarity=0.396  Sum_probs=35.9

Q ss_pred             CCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEee
Q psy17223        105 PNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVS  145 (323)
Q Consensus       105 sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~  145 (323)
                      .-.+.+.+++++..+.+|+.+.+++...|....++...++.
T Consensus       495 p~r~~i~l~~~k~~~~~g~~v~~~v~~~yL~GaPa~g~~~~  535 (1621)
T COG2373         495 PDRFKINLTLDKTEWVPGKDVKIKVDLRYLYGAPAAGLTVQ  535 (1621)
T ss_pred             CceEEEecccccccccCCCcEEEEEEEEecCCCcccCceee
Confidence            44567888999999999999999999999999888777766


No 55 
>PF13595 DUF4138:  Domain of unknown function (DUF4138)
Probab=31.80  E-value=50  Score=31.38  Aligned_cols=21  Identities=29%  Similarity=0.543  Sum_probs=19.9

Q ss_pred             eEEeecceEEEEEEEecCCcc
Q psy17223        224 FLYYHGESIAVNVHVANNSNR  244 (323)
Q Consensus       224 ~~yyhge~i~v~v~v~N~s~k  244 (323)
                      +||+||+-+-+.+.|.|+|+-
T Consensus       136 ~Iy~~~d~lyf~~~i~N~S~i  156 (246)
T PF13595_consen  136 NIYVHGDYLYFHLSIKNKSNI  156 (246)
T ss_pred             eEEEECCEEEEEEEEEcCCCC
Confidence            899999999999999999984


No 56 
>PF04425 Bul1_N:  Bul1 N terminus;  InterPro: IPR007519 This domain is the N terminus of Saccharomyces cerevisiae (Baker's yeast) Bul1. Bul1 binds the ubiquitin ligase Rsp5, via an N-terminal PPSY motif (157-160 in P48524 from SWISSPROT) []. The complex containing Bul1 and Rsp5 is involved in intracellular trafficking of the general amino acid permease Gap1 [], degradation of Rog1 in cooperation with Bul2 and GSK-3 [], and mitochondrial inheritance []. Bul1 may contain HEAT repeats. The C terminus is IPR007520 from INTERPRO.
Probab=31.47  E-value=28  Score=35.96  Aligned_cols=22  Identities=18%  Similarity=0.285  Sum_probs=18.1

Q ss_pred             HHHHcccCCCeeeeeEeeCCCC
Q psy17223         12 QERLMKKLGPNAFPFFFELPPS   33 (323)
Q Consensus        12 Q~~l~~~~G~h~FPFsFqLP~~   33 (323)
                      ..+.++|--.|.++|.|+||..
T Consensus       252 ~~r~l~p~~~Yk~fF~FkiP~~  273 (438)
T PF04425_consen  252 NKRILEPGVKYKKFFTFKIPEQ  273 (438)
T ss_pred             CCceecCCCeEeceeEEeCCch
Confidence            4566777778999999999984


No 57 
>PF10633 NPCBM_assoc:  NPCBM-associated, NEW3 domain of alpha-galactosidase;  InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=31.19  E-value=48  Score=25.09  Aligned_cols=24  Identities=25%  Similarity=0.548  Sum_probs=18.6

Q ss_pred             cceEEEEEEEecCCcceEeeEEee
Q psy17223        229 GESIAVNVHVANNSNRTVKKIKVS  252 (323)
Q Consensus       229 ge~i~v~v~v~N~s~k~vkkikv~  252 (323)
                      |+++.+.+.|.|+....+..++++
T Consensus         4 G~~~~~~~tv~N~g~~~~~~v~~~   27 (78)
T PF10633_consen    4 GETVTVTLTVTNTGTAPLTNVSLS   27 (78)
T ss_dssp             TEEEEEEEEEE--SSS-BSS-EEE
T ss_pred             CCEEEEEEEEEECCCCceeeEEEE
Confidence            899999999999999999988886


No 58 
>cd00258 GM2-AP GM2 activator protein (GM2-AP) is a non-enzymatic lysosomal protein that acts as cofactor in the sequential degradation of gangliosides. GM2A is an essential cofactor for beta-hexosaminidase A (Hex A) in the enzymatic hydrolysis of GM2 ganglioside to GM3. Mutation of the gene results in the AB variant of Tay-Sachs disease. GM2-AP and similar proteins belong to the ML domain family.
Probab=30.74  E-value=61  Score=29.18  Aligned_cols=35  Identities=29%  Similarity=0.473  Sum_probs=25.8

Q ss_pred             ccCCCeeeee-EeeCCC-CCCCcEEeccCCCCCCCceeeEEEEEEEEcc
Q psy17223         17 KKLGPNAFPF-FFELPP-SCPASVTLQPAPGDTGKPCGVDYELKAFVGE   63 (323)
Q Consensus        17 ~~~G~h~FPF-sFqLP~-~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r   63 (323)
                      .++|.|..|= +|.||. +||+...       .|+     |++++.+++
T Consensus       109 ~~~G~y~lp~s~f~lP~~~LPs~l~-------~G~-----Y~i~~~l~~  145 (162)
T cd00258         109 FKEGVYSLPDSTFTLPNVDLPSWLT-------NGN-----YRITGILMA  145 (162)
T ss_pred             CCCcceEccceeeecccccCCCccC-------CCc-----EEEEEEECC
Confidence            5688888854 558987 6887653       562     999999853


No 59 
>cd00917 PG-PI_TP The phosphatidylinositol/phosphatidylglycerol transfer protein (PG/PI-TP) has been shown to bind phosphatidylglycerol and phosphatidylinositol, but the biological significance of this is still obscure. These proteins belong to the ML domain family.
Probab=30.69  E-value=87  Score=26.05  Aligned_cols=32  Identities=19%  Similarity=0.225  Sum_probs=23.6

Q ss_pred             cccCCCeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEcc
Q psy17223         16 MKKLGPNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGE   63 (323)
Q Consensus        16 ~~~~G~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r   63 (323)
                      =.++|.+.+..+..||...|+                +.|.|++.+..
T Consensus        78 Pi~~G~~~~~~~~~ip~~~P~----------------g~y~v~~~l~d  109 (122)
T cd00917          78 PIEPGDKFLTKLVDLPGEIPP----------------GKYTVSARAYT  109 (122)
T ss_pred             CcCCCcEEEEEEeeCCCCCCC----------------ceEEEEEEEEC
Confidence            345788888888888876676                24888887743


No 60 
>PF11355 DUF3157:  Protein of unknown function (DUF3157);  InterPro: IPR021501  This family of proteins with unknown function appears to be restricted to Gammaproteobacteria. 
Probab=29.87  E-value=1.1e+02  Score=28.43  Aligned_cols=37  Identities=19%  Similarity=0.389  Sum_probs=30.6

Q ss_pred             EEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223        110 LEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDS  147 (323)
Q Consensus       110 L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~  147 (323)
                      +++.|...-|--| .+-+...|.|+|+..|..|++.+.
T Consensus       100 VdV~l~~~~y~~~-~L~l~~~ltnqSsqsVv~Vel~v~  136 (199)
T PF11355_consen  100 VDVSLGASQYEDG-QLGLPFSLTNQSSQSVVLVELEVT  136 (199)
T ss_pred             eeEEEeccceeCC-eEEEEEEEecCCCceEEEEEEEEE
Confidence            6777777777777 899999999999999988877663


No 61 
>PF04314 DUF461:  Protein of unknown function (DUF461);  InterPro: IPR007410 This entry represents a domain found in of proteins of unknown function, including DR1885 from Deinococcus radiodurans and CC3502 from Caulobacter crescentus (Caulobacter vibrioides), which share a potential metal binding motif H(M)X10MX21HXM. DR1885 was found to bind copper(I) through a histidine and three Mets in a cupredoxin-like fold []. The surface location of the copper-binding site as well as the type of coordination are well poised for metal transfer chemistry, suggesting that DR1885 might transfer copper, taking the role of Cox17 in bacteria (Cox17 being an accessory protein required for correct assembly of eukaryotic cyochrome c oxidase). ; PDB: 2K6W_A 2K6Z_A 2K6Y_A 2K70_A 1X9L_A 2JQA_A.
Probab=28.27  E-value=87  Score=25.65  Aligned_cols=40  Identities=18%  Similarity=0.241  Sum_probs=30.9

Q ss_pred             EeEecCCceEEEEEeCccceecCCeEEEEEEEeccCccee
Q psy17223        100 EFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTV  139 (323)
Q Consensus       100 ~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~V  139 (323)
                      ..-|..|-.+|-+.=.+....+||.+++++..+|.....|
T Consensus        70 ~v~l~pgg~HlmL~g~~~~l~~G~~v~ltL~f~~gg~v~v  109 (110)
T PF04314_consen   70 TVELKPGGYHLMLMGLKRPLKPGDTVPLTLTFEDGGKVTV  109 (110)
T ss_dssp             EEEE-CCCCEEEEECESS-B-TTEEEEEEEEETTTEEEEE
T ss_pred             eEEecCCCEEEEEeCCcccCCCCCEEEEEEEECCCCEEEe
Confidence            3457888888888777888999999999999999887665


No 62 
>smart00809 Alpha_adaptinC2 Adaptin C-terminal domain. Adaptins are components of the adaptor complexes which link clathrin to receptors in coated vesicles. Clathrin-associated protein complexes are believed to interact with the cytoplasmic tails of membrane proteins, leading to their selection and concentration. Gamma-adaptin is a subunit of the golgi adaptor. Alpha adaptin is a heterotetramer that regulates clathrin-bud formation. The carboxyl-terminal appendage of the alpha subunit regulates translocation of endocytic accessory proteins to the bud site. This Ig-fold domain is found in alpha, beta and gamma adaptins and consists of a beta-sandwich containing 7 strands in 2 beta-sheets in a greek-key topology PUBMED:10430869, PUBMED:12176391. The adaptor appendage contains an additional N-terminal strand.
Probab=27.99  E-value=2.9e+02  Score=21.53  Aligned_cols=55  Identities=9%  Similarity=0.059  Sum_probs=37.4

Q ss_pred             eEEeecceEEEEEEEecCCcceEeeEEeeeEEEeeCCeEeeeeccccCCCccCCCCccc
Q psy17223        224 FLYYHGESIAVNVHVANNSNRTVKKIKVSDICLFSTAQYKCTVAETESDCPIAPVSMFD  282 (323)
Q Consensus       224 ~~yyhge~i~v~v~v~N~s~k~vkkikv~dv~l~s~~~y~~~Va~~e~~~~i~p~~t~~  282 (323)
                      .+=+++..+.+.+...|++...+.++.+.    +..-+|-+.-..--++..|.||+..+
T Consensus        12 ~~~~~~~~~~i~~~~~N~s~~~it~f~~~----~avpk~~~l~l~~~s~~~l~p~~~i~   66 (104)
T smart00809       12 KFERRPGLIRITLTFTNKSPSPITNFSFQ----AAVPKSLKLQLQPPSSPTLPPGGQIT   66 (104)
T ss_pred             EEEcCCCeEEEEEEEEeCCCCeeeeEEEE----EEcccceEEEEcCCCCCccCCCCCEE
Confidence            44456778899999999999999888875    22344444444434566899987633


No 63 
>PF07919 Gryzun:  Gryzun, putative trafficking through Golgi;  InterPro: IPR012880 The proteins featured in this family are all hypothetical eukaryotic proteins of unknown function. The region in question is approximately 150 residues long. 
Probab=27.38  E-value=2.1e+02  Score=29.31  Aligned_cols=40  Identities=18%  Similarity=0.298  Sum_probs=32.0

Q ss_pred             eEecCCceEEEEEe--CccceecCCeEEEEEEEeccCcceee
Q psy17223        101 FMMSPNKLHLEASL--DKELYYHGESIAVNVHVANNSNRTVK  140 (323)
Q Consensus       101 f~f~sG~I~L~a~L--dK~~Y~PGE~I~V~v~IdN~Ssk~Vk  140 (323)
                      +.+..-+-+|++.+  .+.-|+-||.+.|.+.|.|.......
T Consensus       166 i~I~p~pp~v~I~~~~~~~~~l~gE~~~i~i~I~n~e~~~~~  207 (554)
T PF07919_consen  166 IRILPRPPKVSIKLPNHKPPALTGEFYPIPITISNNEDEEAS  207 (554)
T ss_pred             EEEECCCCCeEEEeCCCCCCeEcCCEEEEEEEEEcCCCccce
Confidence            34556677777777  78889999999999999999977544


No 64 
>PF11355 DUF3157:  Protein of unknown function (DUF3157);  InterPro: IPR021501  This family of proteins with unknown function appears to be restricted to Gammaproteobacteria. 
Probab=26.79  E-value=1.3e+02  Score=27.93  Aligned_cols=39  Identities=18%  Similarity=0.370  Sum_probs=30.3

Q ss_pred             hhhccceeeEEeecceEEEEEEEecCCcceEeeEEeeeEEEee
Q psy17223        216 EKSKKKYLFLYYHGESIAVNVHVANNSNRTVKKIKVSDICLFS  258 (323)
Q Consensus       216 ~~s~~k~~~~yyhge~i~v~v~v~N~s~k~vkkikv~dv~l~s  258 (323)
                      ..+++.   -+|.|....+...++|+|++.|.-|.+ +|.||.
T Consensus       101 dV~l~~---~~y~~~~L~l~~~ltnqSsqsVv~Vel-~v~l~d  139 (199)
T PF11355_consen  101 DVSLGA---SQYEDGQLGLPFSLTNQSSQSVVLVEL-EVTLFD  139 (199)
T ss_pred             eEEEec---cceeCCeEEEEEEEecCCCceEEEEEE-EEEEEc
Confidence            444543   456666999999999999999988876 588883


No 65 
>PF07705 CARDB:  CARDB;  InterPro: IPR011635 The APHP (acidic peptide-dependent hydrolases/peptidase) domain is found in a variety of different proteins.; PDB: 2KUT_A 2L0D_A 3IDU_A 2KL6_A.
Probab=26.76  E-value=2e+02  Score=21.66  Aligned_cols=34  Identities=24%  Similarity=0.427  Sum_probs=24.1

Q ss_pred             EeecceEEEEEEEecCCcceEeeEEeeeEEEeeCCeE
Q psy17223        226 YYHGESIAVNVHVANNSNRTVKKIKVSDICLFSTAQY  262 (323)
Q Consensus       226 yyhge~i~v~v~v~N~s~k~vkkikv~dv~l~s~~~y  262 (323)
                      .+=|+++.|.+.|.|.-......++|.   +|.++.-
T Consensus        15 ~~~g~~~~i~~~V~N~G~~~~~~~~v~---~~~~~~~   48 (101)
T PF07705_consen   15 VVPGEPVTITVTVKNNGTADAENVTVR---LYLDGNS   48 (101)
T ss_dssp             EETTSEEEEEEEEEE-SSS-BEEEEEE---EEETTEE
T ss_pred             ccCCCEEEEEEEEEECCCCCCCCEEEE---EEECCce
Confidence            345999999999999988877776665   5555554


No 66 
>cd08546 cohesin_like Cohesin domain, interaction parter of dockerin. Bacterial cohesin domains bind to a complementary protein domain named dockerin, and this interaction is required for the formation of the cellulosome, a cellulose-degrading complex. The cellulosome consists of scaffoldin, a noncatalytic scaffolding polypeptide, that comprises repeating cohesion modules and a single carbohydrate-binding module (CBM). Specific calcium-dependent interactions between cohesins and dockerins appear to be essential for cellulosome assembly. Cohesin modules are phylogenetically distributed into three groups:  type I cohesin-dockerin interactions mediate assembly of a range of dockerin-borne enzymes to the complex, while type-II interactions mediate attachment of the cellulosome complex to the bacterial cell wall. Recently discovered type-III cohesins, such as found in the anchoring scaffoldin ScaE, appears to contribute to increased stability of the elaborate cellulosome complex. While the p
Probab=25.50  E-value=1.6e+02  Score=23.78  Aligned_cols=37  Identities=22%  Similarity=0.215  Sum_probs=29.0

Q ss_pred             EEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCC
Q psy17223        110 LEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGA  149 (323)
Q Consensus       110 L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~  149 (323)
                      +.+..+... .+||++.|.+.++|.+  .+..+.+.+.=+
T Consensus         3 ~~~~~~~~~-~~G~~~~v~v~~~~~~--~~~~~~~~l~yD   39 (135)
T cd08546           3 VSLGAPSTV-KVGETVTVTVKVNNVP--NVAAADFTLSYD   39 (135)
T ss_pred             EEEeccccc-cCCCEEEEEEEEecCC--CeEEEEEEEEEC
Confidence            445555555 8999999999999998  888888877654


No 67 
>PF11797 DUF3324:  Protein of unknown function C-terminal (DUF3324);  InterPro: IPR021759  This family consists of several hypothetical bacterial proteins of unknown function. 
Probab=24.64  E-value=1.5e+02  Score=25.38  Aligned_cols=66  Identities=14%  Similarity=0.143  Sum_probs=43.8

Q ss_pred             eEEEEEEEecCCcceEeeEEeeeEEEeeCCeEeeeecccc-CCCccCCCCcccccceeec-ccccccccCC
Q psy17223        231 SIAVNVHVANNSNRTVKKIKVSDICLFSTAQYKCTVAETE-SDCPIAPVSMFDTEDLAML-RHGFKRMFGH  299 (323)
Q Consensus       231 ~i~v~v~v~N~s~k~vkkikv~dv~l~s~~~y~~~Va~~e-~~~~i~p~~t~~~~~~~~~-~~~~~~~~~~  299 (323)
                      .-.|.+.|.|.....++++++. ..++..++ .+++...+ .+-.++|+|.|. ..|.+. ..|+..+|.+
T Consensus        43 ~~~i~~~l~N~~~~~l~~~~v~-a~V~~~~~-~k~~~~~~~~~~~mAPNS~f~-~~i~~~~~~lk~G~Y~l  110 (140)
T PF11797_consen   43 RNVIQANLQNPQPAILKKLTVD-AKVTKKGS-KKVLYTFKKENMQMAPNSNFN-FPIPLGGKKLKPGKYTL  110 (140)
T ss_pred             eeEEEEEEECCCchhhcCcEEE-EEEEECCC-CeEEEEeeccCCEECCCCeEE-eEecCCCcCccCCEEEE
Confidence            4457778999999999999885 45554444 34555544 466999999986 344443 3555555543


No 68 
>PF02019 WIF:  WIF domain;  InterPro: IPR003306 Wnt proteins constitute a large family of secreted molecules that are involved in intercellular signalling during development. The name derives from the first 2 members of the family to be discovered: int-1 (mouse) and wingless (Drosophila) []. It is now recognised that Wnt signalling controls many cell fate decisions in a variety of different organisms, including mammals []. Wnt signalling has been implicated in tumourigenesis, early mesodermal patterning of the embryo, morphogenesis of the brain and kidneys, regulation of mammary gland proliferation and Alzheimer's disease [, ]. Wnt-mediated signalling is believed to proceed initially through binding to cell surface receptors of the frizzled family; the signal is subsequently transduced through several cytoplasmic components to B-catenin, which enters the nucleus and activates the transcription of several genes important in development []. Several non-canonical Wnt signalling pathways have also been elucidated that act independently of B-catenin. Canonical and noncanonical Wnt signaling branches are highly interconnected, and cross-regulate each other []. Members of the Wnt gene family are defined by their sequence similarity to mouse Wnt-1 and Wingless in Drosophila. They encode proteins of ~350-400 residues in length, with orthologues identified in several, mostly vertebrate, species. Very little is known about the structure of Wnts as they are notoriously insoluble, but they share the following features characteristics of secretory proteins: a signal peptide, several potential N-glycosylation sites and 22 conserved cysteines [] that are probably involved in disulphide bonds. The Wnt proteins seem to adhere to the plasma membrane of the secreting cells and are therefore likely to signal over only few cell diameters. Fifteen major Wnt gene families have been identified in vertebrates, with multiple subtypes within some classes. This entry represents the WIF domain, and is found in the RYK tyrosine kinase receptors and WIF the Wnt-inhibitory-factor. The domain is extracellular and contains two conserved cysteines that may form a disulphide bridge. This domain is Wnt binding in WIF, and it has been suggested that RYK may also bind to Wnt [].; GO: 0004713 protein tyrosine kinase activity; PDB: 2YGP_A 2YGO_A 2YGN_A 2D3J_A 2YGQ_A.
Probab=23.99  E-value=2.9e+02  Score=23.89  Aligned_cols=43  Identities=9%  Similarity=0.159  Sum_probs=25.5

Q ss_pred             cCCceEEEEEeCccceecCC-eEEEEEEEeccCcceeeEEEeec
Q psy17223        104 SPNKLHLEASLDKELYYHGE-SIAVNVHVANNSNRTVKKIKVSD  146 (323)
Q Consensus       104 ~sG~I~L~a~LdK~~Y~PGE-~I~V~v~IdN~Ssk~Vk~Ikv~L  146 (323)
                      ...+-..++.|+-.|...|| ++.|+++|.+.+++..+.+.++.
T Consensus        85 P~~~~~F~V~L~CtG~~~g~a~v~i~lni~~~~~~n~T~L~~kr  128 (132)
T PF02019_consen   85 PHSPQVFSVELPCTGKRSGEATVTIQLNITLPSSKNGTPLRFKR  128 (132)
T ss_dssp             -SS-EEEEEE--B-SSS-EEEEEEEEEEEEETT--S-EEEE--T
T ss_pred             cCCCEEEEEEEEecCccceEEEEEEEEEEEeCCCCCceEEEEee
Confidence            34466778889999999998 78999999999987666555543


No 69 
>PF10437 Lip_prot_lig_C:  Bacterial lipoate protein ligase C-terminus;  InterPro: IPR019491  This is the C-terminal domain of a bacterial lipoate protein ligase. There is no conservation between this C terminus and that of vertebrate lipoate protein ligase C-termini, but both are associated with IPR004143 from INTERPRO, further upstream. This C-terminal domain is more stable than IPR004143 from INTERPRO and the hypothesis is that the C-terminal domain has a role in recognising the lipoyl domain and/or transferring the lipoyl group onto it from the lipoyl-AMP intermediate. C-terminal fragments of length 172 to 193 amino acid residues are observed in the eubacterial enzymes whereas in their archaeal counterparts the C-terminal segment is significantly smaller, ranging in size from 87 to 107 amino acid residues. ; PDB: 1X2G_A 3A7R_A 3A7A_A 1X2H_C 1VQZ_A 3R07_C.
Probab=23.31  E-value=1.8e+02  Score=22.41  Aligned_cols=27  Identities=22%  Similarity=0.288  Sum_probs=21.9

Q ss_pred             eeEEeecceEEEEEEEecCCcceEeeEEee
Q psy17223        223 LFLYYHGESIAVNVHVANNSNRTVKKIKVS  252 (323)
Q Consensus       223 ~~~yyhge~i~v~v~v~N~s~k~vkkikv~  252 (323)
                      +..++.|-.|.|+++|.|.   .|+.|++.
T Consensus         9 ~~~rf~~G~v~v~~~V~~G---~I~~i~i~   35 (86)
T PF10437_consen    9 KERRFPWGTVEVHLNVKNG---IIKDIKIY   35 (86)
T ss_dssp             EEEEETTEEEEEEEEEETT---EEEEEEEE
T ss_pred             eeeEcCCceEEEEEEEECC---EEEEEEEE
Confidence            3678999999999999888   66666665


No 70 
>PF04744 Monooxygenase_B:  Monooxygenase subunit B protein;  InterPro: IPR006833 Ammonia monooxygenase and the particulate methane monooxygenase are both integral membrane proteins, occurring in ammonia oxidisers and methanotrophs respectively, which are thought to be evolutionarily related []. These enzymes have a relatively wide substrate specificity and can catalyse the oxidation of a range of substrates including ammonia, methane, halogenated hydrocarbons and aromatic molecules []. These enzymes are composed of 3 subunits - A (IPR003393 from INTERPRO), B (IPR006833 from INTERPRO) and C (IPR006980 from INTERPRO) - and contain various metal centres, including copper. Particulate methane monooxygenase from Methylococcus capsulatus str. Bath is an ABC homotrimer, which contains mononuclear and dinuclear copper metal centres, and a third metal centre containing a metal ion whose identity in vivo is not certain[]. The soluble regions of these enzymes derive primarily from the B subunit. This subunit forms two antiparallel beta-barrel-like structures and contains the mono- and di- nuclear copper metal centres [].; PDB: 3CHX_E 3RFR_A 3RGB_A 1YEW_A.
Probab=23.11  E-value=1.1e+02  Score=31.04  Aligned_cols=67  Identities=21%  Similarity=0.342  Sum_probs=34.7

Q ss_pred             cchhhccceeeEEe-ecceEEEEEEEecCCcceEeeEEee--eEEEeeCC------eEee-eecc----ccCCCccCCCC
Q psy17223        214 SKEKSKKKYLFLYY-HGESIAVNVHVANNSNRTVKKIKVS--DICLFSTA------QYKC-TVAE----TESDCPIAPVS  279 (323)
Q Consensus       214 ~~~~s~~k~~~~yy-hge~i~v~v~v~N~s~k~vkkikv~--dv~l~s~~------~y~~-~Va~----~e~~~~i~p~~  279 (323)
                      .++.-..+  |.|. -|..+.+++.|+||+++-|+==...  +|..-.-+      .|-. .+|.    +....||+||.
T Consensus       248 ~V~~~v~~--A~Y~vpgR~l~~~l~VtN~g~~pv~LgeF~tA~vrFln~~v~~~~~~~P~~l~A~~gL~vs~~~pI~PGE  325 (381)
T PF04744_consen  248 SVKVKVTD--ATYRVPGRTLTMTLTVTNNGDSPVRLGEFNTANVRFLNPDVPTDDPDYPDELLAERGLSVSDNSPIAPGE  325 (381)
T ss_dssp             SEEEEEEE--EEEESSSSEEEEEEEEEEESSS-BEEEEEESSS-EEE-TTT-SS-S---TTTEETT-EEES--S-B-TT-
T ss_pred             ceEEEEec--cEEecCCcEEEEEEEEEcCCCCceEeeeEEeccEEEeCcccccCCCCCchhhhccCcceeCCCCCcCCCc
Confidence            34444555  6665 6899999999999999987644443  33332111      1111 1332    23345999998


Q ss_pred             ccc
Q psy17223        280 MFD  282 (323)
Q Consensus       280 t~~  282 (323)
                      |-.
T Consensus       326 Trt  328 (381)
T PF04744_consen  326 TRT  328 (381)
T ss_dssp             EEE
T ss_pred             eEE
Confidence            744


No 71 
>TIGR03780 Bac_Flav_CT_N Bacteroides conjugative transposon TraN protein. Members of this family are the TraN protein encoded by transfer region genes of conjugative transposons of Bacteroides. The family is related to conjugative transfer proteins VirB9 and TrbG of Agrobacterium Ti plasmids.
Probab=23.04  E-value=88  Score=30.57  Aligned_cols=20  Identities=20%  Similarity=0.399  Sum_probs=19.4

Q ss_pred             eEEeecceEEEEEEEecCCc
Q psy17223        224 FLYYHGESIAVNVHVANNSN  243 (323)
Q Consensus       224 ~~yyhge~i~v~v~v~N~s~  243 (323)
                      +||.||+-+-+++.+.|+||
T Consensus       175 ~Iy~~~d~lyf~~~l~N~Sn  194 (285)
T TIGR03780       175 GIYTHNDLLYFHTSLENKTN  194 (285)
T ss_pred             eEEEECCEEEEEEEEEcCCC
Confidence            89999999999999999987


No 72 
>PF12389 Peptidase_M73:  Camelysin metallo-endopeptidase;  InterPro: IPR022121 Camelysin is a novel surface metallopeptidase from Bacillus cereus []. Camelysin prefers cleavage sites in front of aliphatic and hydrophilic amino acid residues (-OH, -SO3H, amido group), and requires zinc for activity [, ].
Probab=23.03  E-value=1.3e+02  Score=27.85  Aligned_cols=32  Identities=9%  Similarity=0.188  Sum_probs=27.3

Q ss_pred             cceecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223        117 ELYYHGESIAVNVHVANNSNRTVKKIKVSDSG  148 (323)
Q Consensus       117 ~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q  148 (323)
                      .-..||+++.-.+.|.|.-+-+|+.|.+...-
T Consensus        59 ~nlkPGD~v~k~f~l~N~Gtldi~~v~l~~~y   90 (199)
T PF12389_consen   59 SNLKPGDTVEKEFTLKNSGTLDIKDVLLKTDY   90 (199)
T ss_pred             ccCCCCCeEEEEEEEEeCCeeeeeeEEEEEEE
Confidence            35789999999999999999999888777643


No 73 
>PF04442 CtaG_Cox11:  Cytochrome c oxidase assembly protein CtaG/Cox11;  InterPro: IPR007533 Cytochrome c oxidase assembly protein is essential for the assembly of functional cytochrome oxidase protein. In eukaryotes it is an integral protein of the mitochondrial inner membrane. Cox11 is essential for the insertion of Cu(I) ions to form the CuB site. This is essential for the stability of other structures in subunit I, for example haems a and a3, and the magnesium/manganese centre. Cox11 is probably only required in sub-stoichiometric amounts relative to the structural units []. The C-terminal region of the protein is known to form a dimer. Each monomer coordinates one Cu(I) ion via three conserved cysteine residues (111, 208 and 210) in Saccharomyces cerevisiae (P19516 from SWISSPROT). Met 224 is also thought to play a role in copper transfer or stabilising the copper site [].; GO: 0005507 copper ion binding; PDB: 1SO9_A 1SP0_A.
Probab=21.83  E-value=97  Score=27.53  Aligned_cols=29  Identities=17%  Similarity=0.234  Sum_probs=18.5

Q ss_pred             eecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223        119 YYHGESIAVNVHVANNSNRTVKKIKVSDS  147 (323)
Q Consensus       119 Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~  147 (323)
                      ..|||...+...+.|.|++.|..+-+-=.
T Consensus        63 V~pGe~~~~~y~a~N~s~~~i~g~A~~nV   91 (152)
T PF04442_consen   63 VHPGETALVFYEATNPSDKPITGQAIPNV   91 (152)
T ss_dssp             EETT--EEEEEEEEE-SSS-EE---EEEE
T ss_pred             eCCCCEEEEEEEEECCCCCcEEEEEeeeE
Confidence            36999999999999999999987765443


No 74 
>KOG2540|consensus
Probab=21.06  E-value=82  Score=29.96  Aligned_cols=71  Identities=15%  Similarity=0.252  Sum_probs=42.8

Q ss_pred             ccce-ecCCeEEEEEEEeccCcceeeEEEeecCCCccchh----hhhhhccCCCCCC-ccCCCccccccccccceeeecc
Q psy17223        116 KELY-YHGESIAVNVHVANNSNRTVKKIKVSDSGAEDDQD----LKDELADSDIDGM-EEDDLPNIKAWGKNKRMYYNTD  189 (323)
Q Consensus       116 K~~Y-~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~ed~~d----~~~~~~~~di~~~-~e~~lp~~~awg~~~~~~y~td  189 (323)
                      +.+| .|||+...-....|.|.++|.+|.---+--.+..-    .+=.+.+.-.-.. |+-|||        .=-|.|.|
T Consensus       155 rEiyV~PGEtALaFYta~N~sdkpIiGvstYni~P~~Aa~YFnKiqCFCFEEQ~L~pgE~vDmP--------VFFyIDPe  226 (269)
T KOG2540|consen  155 REIYVLPGETALAFYTAENPSDKPIIGVSTYNITPGQAAVYFNKIQCFCFEEQKLNPGEQVDMP--------VFFYIDPE  226 (269)
T ss_pred             eEEEEcCCcceeeeEeccCCCCCCceeeEeeccCccHhhhheeceeEEeehhhccCCCcccCcc--------eEEEeCcc
Confidence            4556 69999999999999999999888754332111110    1112222222222 445788        55677777


Q ss_pred             cccCC
Q psy17223        190 YVDDD  194 (323)
Q Consensus       190 ~~~~~  194 (323)
                      |+++.
T Consensus       227 fa~DP  231 (269)
T KOG2540|consen  227 FATDP  231 (269)
T ss_pred             cccCc
Confidence            76443


No 75 
>PF04425 Bul1_N:  Bul1 N terminus;  InterPro: IPR007519 This domain is the N terminus of Saccharomyces cerevisiae (Baker's yeast) Bul1. Bul1 binds the ubiquitin ligase Rsp5, via an N-terminal PPSY motif (157-160 in P48524 from SWISSPROT) []. The complex containing Bul1 and Rsp5 is involved in intracellular trafficking of the general amino acid permease Gap1 [], degradation of Rog1 in cooperation with Bul2 and GSK-3 [], and mitochondrial inheritance []. Bul1 may contain HEAT repeats. The C terminus is IPR007520 from INTERPRO.
Probab=20.87  E-value=1.3e+02  Score=31.23  Aligned_cols=44  Identities=27%  Similarity=0.381  Sum_probs=32.7

Q ss_pred             CCceEEEEEeCc---------------cceecCCeEEEEEEEeccCcceee--EEEeecCC
Q psy17223        105 PNKLHLEASLDK---------------ELYYHGESIAVNVHVANNSNRTVK--KIKVSDSG  148 (323)
Q Consensus       105 sG~I~L~a~LdK---------------~~Y~PGE~I~V~v~IdN~Ssk~Vk--~Ikv~L~q  148 (323)
                      .-+|.+++.+-|               .-|.+|+.|.=-|.|+|.|++.|.  =+.|.|.+
T Consensus       131 s~~l~I~I~~Tk~v~~~g~p~~id~~l~Ey~qGD~I~GyvtI~N~S~~pIpFdMFyV~lEG  191 (438)
T PF04425_consen  131 SSPLEIEIYVTKDVGKPGKPPEIDPSLKEYTQGDIIHGYVTIENTSSKPIPFDMFYVSLEG  191 (438)
T ss_pred             CCceEEEEEEeccCCCCCCCcccCcccccccCCCEEEEEEEEEECCCCCcccceEEEEEEE
Confidence            456777777766               468889999999999999999875  34444443


No 76 
>PF01050 MannoseP_isomer:  Mannose-6-phosphate isomerase;  InterPro: IPR001538 Mannose-6-phosphate isomerase or phosphomannose isomerase (5.3.1.8 from EC) (PMI) is the enzyme that catalyses the interconversion of mannose-6-phosphate and fructose-6-phosphate. In eukaryotes PMI is involved in the synthesis of GDP-mannose, a constituent of N- and O-linked glycans and GPI anchors and in prokaryotes it participates in a variety of pathways, including capsular polysaccharide biosynthesis and D-mannose metabolism. PMI's belong to the cupin superfamily whose functions range from isomerase and epimerase activities involved in the modification of cell wall carbohydrates in bacteria and plants, to non-enzymatic storage proteins in plant seeds, and transcription factors linked to congenital baldness in mammals []. Three classes of PMI have been defined []. The type II phosphomannose isomerases are bifunctional enzymes 5.3.1.8 from EC. This entry covers the isomerase region of the protein []. The guanosine diphospho-D-mannose pyrophosphorylase region is described in another InterPro entry (see IPR005836 from INTERPRO).; GO: 0016779 nucleotidyltransferase activity, 0005976 polysaccharide metabolic process
Probab=20.66  E-value=4.2e+02  Score=23.19  Aligned_cols=72  Identities=14%  Similarity=0.221  Sum_probs=45.3

Q ss_pred             EeeeeeecCCCCC-CCCCeEEEEEEeEecCCceEEEEEeCccceecCCeEEEEE----EEeccCcceeeEEEeecCC
Q psy17223         77 LAIRKIMYAPSKQ-GEQPSVEVSKEFMMSPNKLHLEASLDKELYYHGESIAVNV----HVANNSNRTVKKIKVSDSG  148 (323)
Q Consensus        77 l~Ir~l~~~P~~~-~~~~~~e~~k~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v----~IdN~Ssk~Vk~Ikv~L~q  148 (323)
                      +.++.|...|-.. ..+....-...+.+-+|.-.+.+-=....+.+||.|.|-.    .|.|.++.++.=|.|+.-.
T Consensus        63 ~~vkri~V~pG~~lSlq~H~~R~E~W~Vv~G~a~v~~~~~~~~~~~g~sv~Ip~g~~H~i~n~g~~~L~~IEVq~G~  139 (151)
T PF01050_consen   63 YKVKRITVNPGKRLSLQYHHHRSEHWTVVSGTAEVTLDDEEFTLKEGDSVYIPRGAKHRIENPGKTPLEIIEVQTGE  139 (151)
T ss_pred             EEEEEEEEcCCCccceeeecccccEEEEEeCeEEEEECCEEEEEcCCCEEEECCCCEEEEECCCCcCcEEEEEecCC
Confidence            3445555555432 2232223334566777877777655555668899886643    5889988888888888854


No 77 
>PRK05089 cytochrome C oxidase assembly protein; Provisional
Probab=20.63  E-value=2e+02  Score=26.44  Aligned_cols=28  Identities=21%  Similarity=0.138  Sum_probs=24.2

Q ss_pred             eecCCeEEEEEEEeccCcceeeEEEeec
Q psy17223        119 YYHGESIAVNVHVANNSNRTVKKIKVSD  146 (323)
Q Consensus       119 Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L  146 (323)
                      ..|||+..+...+.|.|.+.|...-+-=
T Consensus        90 V~pGE~~~~~y~a~N~sd~~i~g~A~~n  117 (188)
T PRK05089         90 VHPGELNLVFYEAENLSDRPIVGQAIPS  117 (188)
T ss_pred             EcCCCeEEEEEEEECCCCCcEEEEEecc
Confidence            3599999999999999999998776644


No 78 
>PF04205 FMN_bind:  FMN-binding domain;  InterPro: IPR007329 This conserved region includes the FMN-binding site of the NqrC protein [] as well as the NosR and NirI regulatory proteins.; GO: 0010181 FMN binding, 0016020 membrane; PDB: 3LWX_A 2KZX_A 3DCZ_A 3O6U_D.
Probab=20.36  E-value=1.7e+02  Score=21.89  Aligned_cols=22  Identities=23%  Similarity=0.511  Sum_probs=17.0

Q ss_pred             cceEEEEEEEecCCcceEeeEEee
Q psy17223        229 GESIAVNVHVANNSNRTVKKIKVS  252 (323)
Q Consensus       229 ge~i~v~v~v~N~s~k~vkkikv~  252 (323)
                      |.+|.|.|.|+++  ..|..|++.
T Consensus         3 ~g~i~v~v~i~~d--g~I~~v~~~   24 (81)
T PF04205_consen    3 GGPITVTVTIDKD--GKITDVKIL   24 (81)
T ss_dssp             EEEEEEEEEEETT--TEEEEEEEE
T ss_pred             CceEEEEEEEeCC--CEEEEEEEe
Confidence            4489999999886  567777776


Done!