Query         027502
Match_columns 222
No_of_seqs    196 out of 1134
Neff          6.7 
Searched_HMMs 46136
Date          Fri Mar 29 10:53:20 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/027502.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/027502hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF07430 PP1:  Phloem filament   99.9 1.8E-21   4E-26  160.1  16.1  179   31-218     6-193 (202)
  2 cd00042 CY Substituted updates  99.9 5.2E-21 1.1E-25  144.8  11.8   85   33-119     1-105 (105)
  3 smart00043 CY Cystatin-like do  99.8 4.1E-21   9E-26  145.9   8.9   88   31-120     2-107 (107)
  4 PF00031 Cystatin:  Cystatin do  99.8 3.1E-18 6.8E-23  126.9  11.6   74   33-108     1-94  (94)
  5 PF00031 Cystatin:  Cystatin do  99.7 2.7E-16 5.9E-21  116.5   9.5   61  140-201     1-61  (94)
  6 cd00042 CY Substituted updates  99.7 4.9E-16 1.1E-20  117.5  10.1   75  140-216     1-94  (105)
  7 smart00043 CY Cystatin-like do  99.6 7.5E-16 1.6E-20  116.9   7.7   78  138-216     2-95  (107)
  8 PF07430 PP1:  Phloem filament   99.4 7.1E-13 1.5E-17  109.6   8.4   85   29-115   111-199 (202)
  9 TIGR01638 Atha_cystat_rel Arab  98.3 1.8E-06 3.9E-11   64.1   6.9   64   42-105    10-76  (92)
 10 TIGR01638 Atha_cystat_rel Arab  98.0 3.1E-05 6.7E-10   57.6   6.7   66  144-212     7-75  (92)
 11 PF06907 Latexin:  Latexin;  In  96.4    0.55 1.2E-05   40.2  19.5  167   42-210     4-194 (220)
 12 TIGR01572 A_thl_para_3677 Arab  93.0       6 0.00013   34.9  15.6   73   47-120    42-117 (265)
 13 PF00666 Cathelicidins:  Cathel  89.6    0.88 1.9E-05   32.0   4.7   52  152-203     4-57  (67)
 14 PF06907 Latexin:  Latexin;  In  83.1      15 0.00033   31.6   9.7   67  145-212     2-74  (220)
 15 TIGR01572 A_thl_para_3677 Arab  72.8     8.1 0.00018   34.1   5.2   67  152-221    42-111 (265)
 16 PF01789 PsbP:  PsbP;  InterPro  70.3      29 0.00062   28.3   7.8   58  154-214    84-142 (175)
 17 PRK13859 type IV secretion sys  56.7     9.9 0.00022   25.3   2.1   29   11-39      4-40  (55)
 18 PF05679 CHGN:  Chondroitin N-a  53.2      43 0.00093   32.2   6.7   49   44-92    160-210 (499)
 19 PF00666 Cathelicidins:  Cathel  53.0      17 0.00037   25.5   2.9   46   47-93      4-53  (67)
 20 PLN00042 photosystem II oxygen  52.2      40 0.00086   29.8   5.7   24  176-199   185-208 (260)
 21 PLN00067 PsbP domain-containin  48.8      39 0.00084   29.9   5.1   28  173-200   190-217 (263)
 22 PF15240 Pro-rich:  Proline-ric  43.5      16 0.00034   30.5   1.8   16    9-24      3-18  (179)
 23 CHL00132 psaF photosystem I su  37.7      71  0.0015   26.7   4.7   27   31-60     25-51  (185)
 24 PF08294 TIM21:  TIM21;  InterP  36.3      40 0.00087   27.0   3.1   73  147-219    51-124 (145)
 25 PLN00059 PsbP domain-containin  36.2 2.2E+02  0.0047   25.5   7.8   44  154-200   172-221 (286)
 26 COG3360 Uncharacterized conser  32.9 1.8E+02  0.0039   20.6   6.4   45   46-91     18-64  (71)
 27 PTZ00444 hypothetical protein;  32.2      38 0.00083   28.3   2.4   23    1-23      1-23  (184)
 28 PF05679 CHGN:  Chondroitin N-a  32.0 1.8E+02  0.0039   27.9   7.3   48  150-200   161-210 (499)
 29 PRK10081 entericidin B membran  31.7      32 0.00069   22.6   1.5   22    1-22      1-22  (48)
 30 COG3360 Uncharacterized conser  31.2 1.8E+02  0.0038   20.6   5.1   45  152-199    19-64  (71)
 31 smart00773 WGR Proposed nuclei  31.1   1E+02  0.0022   21.8   4.3   30  192-221     6-37  (84)
 32 TIGR02105 III_needle type III   28.6      66  0.0014   22.8   2.8   23  145-167    29-51  (72)
 33 PF12276 DUF3617:  Protein of u  28.6      57  0.0012   25.8   2.8   35    1-38      1-35  (162)
 34 PF12274 DUF3615:  Protein of u  26.9 2.5E+02  0.0054   20.4   6.2   50  163-212     1-58  (96)
 35 PRK15344 type III secretion sy  26.3      83  0.0018   22.3   2.9   22  145-166    28-49  (71)
 36 KOG4306 Glycosylphosphatidylin  26.1      42 0.00091   30.4   1.7   27  179-205    70-96  (306)
 37 PF02995 DUF229:  Protein of un  25.3      74  0.0016   30.5   3.4   32  138-169   450-481 (497)
 38 PRK13883 conjugal transfer pro  24.7 1.5E+02  0.0034   24.0   4.6   20   42-61     28-47  (151)
 39 PF10828 DUF2570:  Protein of u  24.7      65  0.0014   24.3   2.4   22    1-22      1-22  (110)
 40 PF09049 SNN_transmemb:  Stanni  24.6      33 0.00072   20.2   0.5   17    5-21     11-27  (33)
 41 PF07311 Dodecin:  Dodecin;  In  24.5 2.5E+02  0.0054   19.5   7.5   47   44-91     13-61  (66)
 42 PF15418 DUF4625:  Domain of un  24.1 1.5E+02  0.0033   23.2   4.4   35  184-219    31-65  (132)
 43 PF01456 Mucin:  Mucin-like gly  24.0      40 0.00087   26.3   1.1   10   12-21      6-15  (143)
 44 smart00557 IG_FLMN Filamin-typ  23.8 1.2E+02  0.0025   21.8   3.5   28  193-221    33-60  (93)
 45 PF05399 EVI2A:  Ectropic viral  22.9      65  0.0014   27.7   2.2   17    5-21    135-151 (227)
 46 PLN03207 stomagen; Provisional  22.5      92   0.002   23.6   2.7   13    5-17     11-23  (113)
 47 PF13956 Ibs_toxin:  Toxin Ibs,  22.4      40 0.00086   17.6   0.5   16    1-16      1-16  (19)
 48 MTH00261 ATP8 ATP synthase F0   22.2      79  0.0017   21.5   2.1   18    2-19     11-28  (68)
 49 PF13721 SecD-TM1:  SecD export  20.5      89  0.0019   23.4   2.3   14    1-14      1-14  (101)
 50 PF12984 DUF3868:  Domain of un  20.2      62  0.0013   24.8   1.4   15    1-15      1-15  (115)

No 1  
>PF07430 PP1:  Phloem filament protein PP1;  InterPro: IPR009994 This domain represents a conserved region approximately 200 residues long, four copies of which are found within the plant phloem filament protein PP1. This is one of the constituents of the proteinaceous filaments found in the sieve elements of Cucurbita phloem [].
Probab=99.88  E-value=1.8e-21  Score=160.14  Aligned_cols=179  Identities=16%  Similarity=0.201  Sum_probs=148.1

Q ss_pred             CCceeeeCCCCCCCHHHHHHHHHHHHHHHhhcCCceeEEEEEEEE--EEeeccEEEEEEEEEEe-CCcceEEEEEEEEec
Q 027502           31 RPGGVYDYGGNQNSAEIEGLARFAVQEHNKKENALLQFARVLKAK--EQVVAGKLYYLTLEVID-AGKNKIYEAKIWVKP  107 (222)
Q Consensus        31 l~GG~~~i~~~~~d~~v~~~a~fAv~~~N~~sn~~~~~~kV~~a~--~QVVaG~nY~l~v~v~~-~~~~~~c~~~V~~~P  107 (222)
                      ..|||.+++ |+.+|.+|++++||+.+++.+-++.++|..|.+.+  .|.+.++.|+|.+++.| -++...|++.|+++-
T Consensus         6 ~~~~w~~ip-~v~~~~~q~v~~~~veq~k~~~~~~l~~~~v~egwy~el~~~~~~yrlhv~a~d~l~r~l~~e~ii~e~~   84 (202)
T PF07430_consen    6 FSPKWIKIP-DVKEPCLQEVAKFAVEQFKIQYGDSLKFRSVVEGWYFELCPNSLKYRLHVKAIDFLGRSLKYEAIIIEEK   84 (202)
T ss_pred             cCcccccCC-cccchHHHHHHHHHHHHHhhhcccceeeeeeeeceeecccccceeEEEeehhhhhhccccceeeeeeehh
Confidence            479999997 58999999999999999999998889999999998  88899999999999988 588899999999995


Q ss_pred             --CCCceeeEEEeeCCCCCCcccccccccccCCCCcceec-cCCCHHHHHHHHHHHHHHHhhccCccceeEEEEEEeeee
Q 027502          108 --WINFKQLQEFKHAEHGPFSALSDLNLKRGCHGQEWLAV-STNDLEVKNAANHAVKSMQRKSNSLFLYELLEILQAKAK  184 (222)
Q Consensus       108 --W~~~~~l~s~~c~~~~~~~~~~~~~~k~~~~~gg~~~i-~~~d~~v~e~a~fAv~~~N~~sn~~~~~~~~kV~~a~~Q  184 (222)
                        |++.++|.|+--.-..-....+    ...+....|.++ |+++|.||+++.|||.||| ..++.  .+|.+|.++..|
T Consensus        85 ~~~~~~~kl~s~l~~v~~~~~v~p----v~~p~~~~Wi~I~nin~p~VQeLgkFAV~EhN-K~gd~--LkF~KV~eGw~q  157 (202)
T PF07430_consen   85 PQLTRIRKLASILAIVSGGAVVYP----VATPQSKKWIPIPNINNPFVQELGKFAVIEHN-KAGDK--LKFEKVYEGWYQ  157 (202)
T ss_pred             hhhhhhhhhheeeEEEeeceeecc----ccCcccCCCEECCCCCcHHHHHHHHHHHHHHh-hcCCc--eEEEEEeeEEEE
Confidence              9999998887543211000000    001123569888 4999999999999999999 55665  499999999999


Q ss_pred             eec--ceeEEEEEEEEccC-CccceEEEEEEeCCCCc
Q 027502          185 VIE--DYAKFELHLKLRRG-SEEEKHWVEIIKNSEGK  218 (222)
Q Consensus       185 VVa--G~~yy~l~l~~~~g-~~~~~~~~~v~~~~~~~  218 (222)
                      =++  |++ |+|+|.+++| ++...|+|.|++++--|
T Consensus       158 ~l~~d~ik-YrLhI~AkDg~G~~~~YeAvV~~k~~~s  193 (202)
T PF07430_consen  158 DLGNDGIK-YRLHIVAKDGLGRLGNYEAVVWEKQFLS  193 (202)
T ss_pred             eccCCCce-EEEEEEeecCCCCcCceEEEEEEeccCc
Confidence            775  688 6899999999 67889999999987544


No 2  
>cd00042 CY Substituted updates: Jan 30, 2002
Probab=99.86  E-value=5.2e-21  Score=144.82  Aligned_cols=85  Identities=54%  Similarity=0.822  Sum_probs=80.5

Q ss_pred             ceeeeCCCCCCCHHHHHHHHHHHHHHHhhcCCc-eeEEEEEEEEEEeeccEEEEEEEEEEeC------------------
Q 027502           33 GGVYDYGGNQNSAEIEGLARFAVQEHNKKENAL-LQFARVLKAKEQVVAGKLYYLTLEVIDA------------------   93 (222)
Q Consensus        33 GG~~~i~~~~~d~~v~~~a~fAv~~~N~~sn~~-~~~~kV~~a~~QVVaG~nY~l~v~v~~~------------------   93 (222)
                      |||.+++  .+||+++++++||+.+||+.+++. |.+.+|++|++|||+|++|+|++++.++                  
T Consensus         1 gg~~~~~--~~d~~~~~~~~~a~~~~N~~~~~~~~~~~~i~~~~~QvvaG~~y~i~~~~~~t~C~k~~~~~~~~~c~~~~   78 (105)
T cd00042           1 GGPSDIP--ANDPEVQELADFAVAEYNKKSNDKYLEFFKVLSAKSQVVAGTNYYITVEAGDTNCKKSSVPLDCPDCKLLE   78 (105)
T ss_pred             CCCccCC--CCCHHHHHHHHHHHHHHHhhcCccceeEEEEEEEEEEEEeeeEEEEEEEEecccccccCcccccccccccc
Confidence            8999887  799999999999999999999998 8889999999999999999999999975                  


Q ss_pred             -CcceEEEEEEEEecCCCceeeEEEee
Q 027502           94 -GKNKIYEAKIWVKPWINFKQLQEFKH  119 (222)
Q Consensus        94 -~~~~~c~~~V~~~PW~~~~~l~s~~c  119 (222)
                       +....|++.||.+||.+..++++++|
T Consensus        79 ~~~~~~C~~~V~~~pw~~~~~l~~~~C  105 (105)
T cd00042          79 EGKKKFCTAKVWEKPWENFKELLSFKC  105 (105)
T ss_pred             cCCCEEEEEEEEecCCCCceeeeeccC
Confidence             57899999999999999999999988


No 3  
>smart00043 CY Cystatin-like domain. Cystatins are a family of cysteine protease inhibitors that occur mainly as single domain proteins. However some extracellular proteins such as  kininogen, His-rich glycoprotein and fetuin also contain these domains.
Probab=99.85  E-value=4.1e-21  Score=145.93  Aligned_cols=88  Identities=44%  Similarity=0.622  Sum_probs=80.5

Q ss_pred             CCceeeeCCCCCCCHHHHHHHHHHHHHHHhhcCCcee--EEEEEEEEEEeeccEEEEEEEEEEeCCcc------------
Q 027502           31 RPGGVYDYGGNQNSAEIEGLARFAVQEHNKKENALLQ--FARVLKAKEQVVAGKLYYLTLEVIDAGKN------------   96 (222)
Q Consensus        31 l~GG~~~i~~~~~d~~v~~~a~fAv~~~N~~sn~~~~--~~kV~~a~~QVVaG~nY~l~v~v~~~~~~------------   96 (222)
                      ++|||.+++  .+||+++++|+||+.+||+++++.|.  +.+|++|++|||+|++|+|++++.++.-.            
T Consensus         2 ~~Gg~~~~~--~~d~~~~~~~~~a~~~~N~~~~~~~~~~~~~v~~a~~QvvaG~~y~l~~~v~~t~C~k~~~~~~~C~~~   79 (107)
T smart00043        2 CLGGPSDVP--PNDPEVQEAADFAVAEYNKKSNDKYELRVIKVVSAKSQVVAGTNYYLKVEVGETNCKKLSVDLENCPFL   79 (107)
T ss_pred             CCCCCccCC--CCCHHHHHHHHHHHHHHHHhcccchhhhhhhhheeeeeeecceEEEEEEEEEeceeccCCcccccCCCC
Confidence            689999997  68999999999999999999998876  79999999999999999999999985322            


Q ss_pred             ----eEEEEEEEEecCCCceeeEEEeeC
Q 027502           97 ----KIYEAKIWVKPWINFKQLQEFKHA  120 (222)
Q Consensus        97 ----~~c~~~V~~~PW~~~~~l~s~~c~  120 (222)
                          ..|.++||.+||.++.++++++|.
T Consensus        80 ~~~~~~C~~~V~~~pw~~~~~~~~~~C~  107 (107)
T smart00043       80 DQGEKFCTAKVWEKPWENKIKLVEFKCT  107 (107)
T ss_pred             CCCccEEEEEEEecCCCCccCccceecC
Confidence                489999999999999999999984


No 4  
>PF00031 Cystatin:  Cystatin domain;  InterPro: IPR000010 Peptide proteinase inhibitors can be found as single domain proteins or as single or multiple domains within proteins; these are referred to as either simple or compound inhibitors, respectively. In many cases they are synthesised as part of a larger precursor protein, either as a prepropeptide or as an N-terminal domain associated with an inactive peptidase or zymogen. This domain prevents access of the substrate to the active site. Removal of the N-terminal inhibitor domain either by interaction with a second peptidase or by autocatalytic cleavage activates the zymogen. Other inhibitors interact direct with proteinases using a simple noncovalent lock and key mechanism; while yet others use a conformational change-based trapping mechanism that depends on their structural and thermodynamic properties.  The cystatins are cysteine proteinase inhibitors belonging to MEROPS inhibitor family I25, clan IH [, , ]. They mainly inhibit peptidases belonging to peptidase families C1 (papain family) and C13 (legumain family). The cystatin family includes:   The Type 1 cystatins, which are intracellular cystatins that are present in the cytosol of many cell types, but can also appear in body fluids at significant concentrations. They are single-chain polypeptides of about 100 residues, which have neither disulphide bonds nor carbohydrate side chains.  The Type 2 cystatins, which are mainly extracellular secreted polypeptides synthesised with a 19-28 residue signal peptide. They are broadly distributed and found in most body fluids.  The Type 3 cystatins, which are multidomain proteins. The mammalian representatives of this group are the kininogens. There are three different kininogens in mammals: H- (high molecular mass, IPR002395 from INTERPRO) and L- (low molecular mass) kininogen which are found in a number of species, and T-kininogen that is found only in rat.  Unclassified cystatins. These are cystatin-like proteins found in a range of organisms: plant phytocystatins, fetuin in mammals, insect cystatins and a puff adder venom cystatin which inhibits metalloproteases of the MEROPS peptidase family M12 (astacin/adamalysin). Also a number of the cystatins-like proteins have been shown to be devoid of inhibitory activity.   All true cystatins inhibit cysteine peptidases of the papain family (MEROPS peptidase family C1), and some also inhibit legumain family enzymes (MEROPS peptidase family C13). These peptidases play key roles in physiological processes, such as intracellular protein degradation (cathepsins B, H and L), are pivotal in the remodelling of bone (cathepsin K), and may be important in the control of antigen presentation (cathepsin S, mammalian legumain). Moreover, the activities of such peptidases are increased in pathophysiological conditions, such as cancer metastasis and inflammation. Additionally, such peptidases are essential for several pathogenic parasites and bacteria. Thus in animals cystatins not only have capacity to regulate normal body processes and perhaps cause disease when down-regulated, but in other organisms may also participate in defence against biotic and abiotic stress. ; GO: 0004869 cysteine-type endopeptidase inhibitor activity; PDB: 3L0R_B 2W9P_K 2W9Q_A 3S67_A 3QRD_D 1R4C_G 3GAX_A 1TIJ_B 1G96_A 3NX0_A ....
Probab=99.78  E-value=3.1e-18  Score=126.94  Aligned_cols=74  Identities=32%  Similarity=0.614  Sum_probs=68.3

Q ss_pred             ceeeeCCCCCCCHHHHHHHHHHHHHHHhhcCCc--eeEEEEEEEEEEeeccEEEEEEEEEEeC-----------------
Q 027502           33 GGVYDYGGNQNSAEIEGLARFAVQEHNKKENAL--LQFARVLKAKEQVVAGKLYYLTLEVIDA-----------------   93 (222)
Q Consensus        33 GG~~~i~~~~~d~~v~~~a~fAv~~~N~~sn~~--~~~~kV~~a~~QVVaG~nY~l~v~v~~~-----------------   93 (222)
                      |||.+++  .+||+++++|+||+.+||+++++.  |.+.+|++|++|||+|++|+|++++.++                 
T Consensus         1 Gg~~~~~--~~dp~v~~~~~~al~~~N~~~~~~~~~~~~~v~~a~~QvV~G~~Y~i~~~~~~t~C~k~~~~~~~C~~~~~   78 (94)
T PF00031_consen    1 GGPSPVD--PNDPEVQEAAEFALDKFNEQSNSGYKFKLVKVISATTQVVAGINYYIEFEVGETNCKKSSKDFENCPFQEE   78 (94)
T ss_dssp             SSEEEEC--TTSHHHHHHHHHHHHHHHHHSTTSEEEEEEEEEEEEEEESSSEEEEEEEEEEEEEEETCEEEEEECEBEST
T ss_pred             CCCccCC--CCCHHHHHHHHHHHHHHHHhCcccCcceeeeeeEEEEeecCCceEEEEEEEEcccccccccccccCCcccc
Confidence            8999998  599999999999999999999766  6789999999999999999999999873                 


Q ss_pred             -CcceEEEEEEEEecC
Q 027502           94 -GKNKIYEAKIWVKPW  108 (222)
Q Consensus        94 -~~~~~c~~~V~~~PW  108 (222)
                       .....|.++||.+||
T Consensus        79 ~~~~~~C~~~v~~~pW   94 (94)
T PF00031_consen   79 QPWTKFCKFTVWERPW   94 (94)
T ss_dssp             TSSEEEEEEEEEEECG
T ss_pred             CCceeeEEEEEEECCC
Confidence             457899999999999


No 5  
>PF00031 Cystatin:  Cystatin domain;  InterPro: IPR000010 Peptide proteinase inhibitors can be found as single domain proteins or as single or multiple domains within proteins; these are referred to as either simple or compound inhibitors, respectively. In many cases they are synthesised as part of a larger precursor protein, either as a prepropeptide or as an N-terminal domain associated with an inactive peptidase or zymogen. This domain prevents access of the substrate to the active site. Removal of the N-terminal inhibitor domain either by interaction with a second peptidase or by autocatalytic cleavage activates the zymogen. Other inhibitors interact direct with proteinases using a simple noncovalent lock and key mechanism; while yet others use a conformational change-based trapping mechanism that depends on their structural and thermodynamic properties.  The cystatins are cysteine proteinase inhibitors belonging to MEROPS inhibitor family I25, clan IH [, , ]. They mainly inhibit peptidases belonging to peptidase families C1 (papain family) and C13 (legumain family). The cystatin family includes:   The Type 1 cystatins, which are intracellular cystatins that are present in the cytosol of many cell types, but can also appear in body fluids at significant concentrations. They are single-chain polypeptides of about 100 residues, which have neither disulphide bonds nor carbohydrate side chains.  The Type 2 cystatins, which are mainly extracellular secreted polypeptides synthesised with a 19-28 residue signal peptide. They are broadly distributed and found in most body fluids.  The Type 3 cystatins, which are multidomain proteins. The mammalian representatives of this group are the kininogens. There are three different kininogens in mammals: H- (high molecular mass, IPR002395 from INTERPRO) and L- (low molecular mass) kininogen which are found in a number of species, and T-kininogen that is found only in rat.  Unclassified cystatins. These are cystatin-like proteins found in a range of organisms: plant phytocystatins, fetuin in mammals, insect cystatins and a puff adder venom cystatin which inhibits metalloproteases of the MEROPS peptidase family M12 (astacin/adamalysin). Also a number of the cystatins-like proteins have been shown to be devoid of inhibitory activity.   All true cystatins inhibit cysteine peptidases of the papain family (MEROPS peptidase family C1), and some also inhibit legumain family enzymes (MEROPS peptidase family C13). These peptidases play key roles in physiological processes, such as intracellular protein degradation (cathepsins B, H and L), are pivotal in the remodelling of bone (cathepsin K), and may be important in the control of antigen presentation (cathepsin S, mammalian legumain). Moreover, the activities of such peptidases are increased in pathophysiological conditions, such as cancer metastasis and inflammation. Additionally, such peptidases are essential for several pathogenic parasites and bacteria. Thus in animals cystatins not only have capacity to regulate normal body processes and perhaps cause disease when down-regulated, but in other organisms may also participate in defence against biotic and abiotic stress. ; GO: 0004869 cysteine-type endopeptidase inhibitor activity; PDB: 3L0R_B 2W9P_K 2W9Q_A 3S67_A 3QRD_D 1R4C_G 3GAX_A 1TIJ_B 1G96_A 3NX0_A ....
Probab=99.68  E-value=2.7e-16  Score=116.51  Aligned_cols=61  Identities=23%  Similarity=0.349  Sum_probs=58.6

Q ss_pred             CcceeccCCCHHHHHHHHHHHHHHHhhccCccceeEEEEEEeeeeeecceeEEEEEEEEccC
Q 027502          140 QEWLAVSTNDLEVKNAANHAVKSMQRKSNSLFLYELLEILQAKAKVIEDYAKFELHLKLRRG  201 (222)
Q Consensus       140 gg~~~i~~~d~~v~e~a~fAv~~~N~~sn~~~~~~~~kV~~a~~QVVaG~~yy~l~l~~~~g  201 (222)
                      |||.++++|||+++++|+||+.+||+++|+.+.|++.+|++|++|||+|++| +|++++.++
T Consensus         1 Gg~~~~~~~dp~v~~~~~~al~~~N~~~~~~~~~~~~~v~~a~~QvV~G~~Y-~i~~~~~~t   61 (94)
T PF00031_consen    1 GGPSPVDPNDPEVQEAAEFALDKFNEQSNSGYKFKLVKVISATTQVVAGINY-YIEFEVGET   61 (94)
T ss_dssp             SSEEEECTTSHHHHHHHHHHHHHHHHHSTTSEEEEEEEEEEEEEEESSSEEE-EEEEEEEEE
T ss_pred             CCCccCCCCCHHHHHHHHHHHHHHHHhCcccCcceeeeeeEEEEeecCCceE-EEEEEEEcc
Confidence            7999999999999999999999999999999999999999999999999995 699999885


No 6  
>cd00042 CY Substituted updates: Jan 30, 2002
Probab=99.67  E-value=4.9e-16  Score=117.47  Aligned_cols=75  Identities=20%  Similarity=0.260  Sum_probs=68.4

Q ss_pred             CcceeccCCCHHHHHHHHHHHHHHHhhccCccceeEEEEEEeeeeeecceeEEEEEEEEccCC-----------------
Q 027502          140 QEWLAVSTNDLEVKNAANHAVKSMQRKSNSLFLYELLEILQAKAKVIEDYAKFELHLKLRRGS-----------------  202 (222)
Q Consensus       140 gg~~~i~~~d~~v~e~a~fAv~~~N~~sn~~~~~~~~kV~~a~~QVVaG~~yy~l~l~~~~g~-----------------  202 (222)
                      |||.+++++||+++++++||+++||+.+|+.| |++.+|++|++|||+|++| .|+++++++.                 
T Consensus         1 gg~~~~~~~d~~~~~~~~~a~~~~N~~~~~~~-~~~~~i~~~~~QvvaG~~y-~i~~~~~~t~C~k~~~~~~~~~c~~~~   78 (105)
T cd00042           1 GGPSDIPANDPEVQELADFAVAEYNKKSNDKY-LEFFKVLSAKSQVVAGTNY-YITVEAGDTNCKKSSVPLDCPDCKLLE   78 (105)
T ss_pred             CCCccCCCCCHHHHHHHHHHHHHHHhhcCccc-eeEEEEEEEEEEEEeeeEE-EEEEEEecccccccCcccccccccccc
Confidence            78999999999999999999999999999999 9999999999999999995 7999999752                 


Q ss_pred             --ccceEEEEEEeCCC
Q 027502          203 --EEEKHWVEIIKNSE  216 (222)
Q Consensus       203 --~~~~~~~~v~~~~~  216 (222)
                        ....+.+.||+.|-
T Consensus        79 ~~~~~~C~~~V~~~pw   94 (105)
T cd00042          79 EGKKKFCTAKVWEKPW   94 (105)
T ss_pred             cCCCEEEEEEEEecCC
Confidence              35578999999885


No 7  
>smart00043 CY Cystatin-like domain. Cystatins are a family of cysteine protease inhibitors that occur mainly as single domain proteins. However some extracellular proteins such as  kininogen, His-rich glycoprotein and fetuin also contain these domains.
Probab=99.63  E-value=7.5e-16  Score=116.92  Aligned_cols=78  Identities=21%  Similarity=0.216  Sum_probs=69.1

Q ss_pred             CCCcceeccCCCHHHHHHHHHHHHHHHhhccCccceeEEEEEEeeeeeecceeEEEEEEEEccCC--cc-----------
Q 027502          138 HGQEWLAVSTNDLEVKNAANHAVKSMQRKSNSLFLYELLEILQAKAKVIEDYAKFELHLKLRRGS--EE-----------  204 (222)
Q Consensus       138 ~~gg~~~i~~~d~~v~e~a~fAv~~~N~~sn~~~~~~~~kV~~a~~QVVaG~~yy~l~l~~~~g~--~~-----------  204 (222)
                      .+|||.+++++||+++++|+||+.+||+++|+.|.+++.+|++|++|||+|++| .|++++.++.  +.           
T Consensus         2 ~~Gg~~~~~~~d~~~~~~~~~a~~~~N~~~~~~~~~~~~~v~~a~~QvvaG~~y-~l~~~v~~t~C~k~~~~~~~C~~~~   80 (107)
T smart00043        2 CLGGPSDVPPNDPEVQEAADFAVAEYNKKSNDKYELRVIKVVSAKSQVVAGTNY-YLKVEVGETNCKKLSVDLENCPFLD   80 (107)
T ss_pred             CCCCCccCCCCCHHHHHHHHHHHHHHHHhcccchhhhhhhhheeeeeeecceEE-EEEEEEEeceeccCCcccccCCCCC
Confidence            579999999999999999999999999999999988899999999999999995 6999988643  32           


Q ss_pred             ---ceEEEEEEeCCC
Q 027502          205 ---EKHWVEIIKNSE  216 (222)
Q Consensus       205 ---~~~~~~v~~~~~  216 (222)
                         ..+.++||..|-
T Consensus        81 ~~~~~C~~~V~~~pw   95 (107)
T smart00043       81 QGEKFCTAKVWEKPW   95 (107)
T ss_pred             CCccEEEEEEEecCC
Confidence               268899998874


No 8  
>PF07430 PP1:  Phloem filament protein PP1;  InterPro: IPR009994 This domain represents a conserved region approximately 200 residues long, four copies of which are found within the plant phloem filament protein PP1. This is one of the constituents of the proteinaceous filaments found in the sieve elements of Cucurbita phloem [].
Probab=99.41  E-value=7.1e-13  Score=109.55  Aligned_cols=85  Identities=28%  Similarity=0.346  Sum_probs=75.9

Q ss_pred             ccCCceeeeCCCCCCCHHHHHHHHHHHHHHHhhcCCceeEEEEEEEEEEeec--cEEEEEEEEEEeC-CcceEEEEEEEE
Q 027502           29 QMRPGGVYDYGGNQNSAEIEGLARFAVQEHNKKENALLQFARVLKAKEQVVA--GKLYYLTLEVIDA-GKNKIYEAKIWV  105 (222)
Q Consensus        29 ~~l~GG~~~i~~~~~d~~v~~~a~fAv~~~N~~sn~~~~~~kV~~a~~QVVa--G~nY~l~v~v~~~-~~~~~c~~~V~~  105 (222)
                      ++....|.+++ |+++|.+|++++|||.+|| +.++.++|.+|.+++.|-++  |++|+|++.+.++ |+...|+|.||+
T Consensus       111 ~p~~~~Wi~I~-nin~p~VQeLgkFAV~EhN-K~gd~LkF~KV~eGw~q~l~~d~ikYrLhI~AkDg~G~~~~YeAvV~~  188 (202)
T PF07430_consen  111 TPQSKKWIPIP-NINNPFVQELGKFAVIEHN-KAGDKLKFEKVYEGWYQDLGNDGIKYRLHIVAKDGLGRLGNYEAVVWE  188 (202)
T ss_pred             CcccCCCEECC-CCCcHHHHHHHHHHHHHHh-hcCCceEEEEEeeEEEEeccCCCceEEEEEEeecCCCCcCceEEEEEE
Confidence            55578999997 5899999999999999999 67889999999999999996  6999999999997 999999999999


Q ss_pred             e-cCCCceeeE
Q 027502          106 K-PWINFKQLQ  115 (222)
Q Consensus       106 ~-PW~~~~~l~  115 (222)
                      + +|.+..+++
T Consensus       189 k~~~sk~i~i~  199 (202)
T PF07430_consen  189 KQFLSKKIKIL  199 (202)
T ss_pred             eccCcceEEEE
Confidence            9 577666654


No 9  
>TIGR01638 Atha_cystat_rel Arabidopsis thaliana cystatin-related protein. This model represents a family similar in sequence and probably homologous to a large family of cysteine proteinase inhibitors, or cystatins, as described by pfam model pfam00031. Cystatins may help plants resist attack by insects.
Probab=98.34  E-value=1.8e-06  Score=64.15  Aligned_cols=64  Identities=23%  Similarity=0.241  Sum_probs=55.1

Q ss_pred             CCCHHHHHHHHHHHHHHHhhcCCceeEEEEEEEEEEeeccEEEEEEEEEEeC--C-cceEEEEEEEE
Q 027502           42 QNSAEIEGLARFAVQEHNKKENALLQFARVLKAKEQVVAGKLYYLTLEVIDA--G-KNKIYEAKIWV  105 (222)
Q Consensus        42 ~~d~~v~~~a~fAv~~~N~~sn~~~~~~kV~~a~~QVVaG~nY~l~v~v~~~--~-~~~~c~~~V~~  105 (222)
                      .+..-+..++++|+++||...+..+.|++|++|..|..+|+.|+||+.+.+.  + ....+++.||.
T Consensus        10 T~rd~~~~la~~al~k~N~~~~t~lEfV~vVrAn~~~~~g~~~yITF~Ard~~d~p~~e~~q~~v~~   76 (92)
T TIGR01638        10 TNRDLLERLSYVASKKYNDTKFLNLELVEVVRANYRGGAKSKSYITFEARDKPDGPLGEYQQAAVVY   76 (92)
T ss_pred             CHHHHHHHHHHHHHHHhhhhcCceEEEEEEEEEEeeccceEEEEEEEEEecCCCCCHHHhhheeeEe
Confidence            4667889999999999999999999999999999999999999999999983  3 44555666665


No 10 
>TIGR01638 Atha_cystat_rel Arabidopsis thaliana cystatin-related protein. This model represents a family similar in sequence and probably homologous to a large family of cysteine proteinase inhibitors, or cystatins, as described by pfam model pfam00031. Cystatins may help plants resist attack by insects.
Probab=97.96  E-value=3.1e-05  Score=57.58  Aligned_cols=66  Identities=15%  Similarity=0.148  Sum_probs=53.6

Q ss_pred             eccCCCHHHHHHHHHHHHHHHhhccCccceeEEEEEEeeeeeecceeEEEEEEEEccCCc---cceEEEEEE
Q 027502          144 AVSTNDLEVKNAANHAVKSMQRKSNSLFLYELLEILQAKAKVIEDYAKFELHLKLRRGSE---EEKHWVEII  212 (222)
Q Consensus       144 ~i~~~d~~v~e~a~fAv~~~N~~sn~~~~~~~~kV~~a~~QVVaG~~yy~l~l~~~~g~~---~~~~~~~v~  212 (222)
                      +...+..-+.+++++|+++||+.-+..+  +|++|++|.-|+.+|..| +||+++.+...   .+.|.+.|+
T Consensus         7 p~~T~rd~~~~la~~al~k~N~~~~t~l--EfV~vVrAn~~~~~g~~~-yITF~Ard~~d~p~~e~~q~~v~   75 (92)
T TIGR01638         7 PIETNRDLLERLSYVASKKYNDTKFLNL--ELVEVVRANYRGGAKSKS-YITFEARDKPDGPLGEYQQAAVV   75 (92)
T ss_pred             cccCHHHHHHHHHHHHHHHhhhhcCceE--EEEEEEEEEeeccceEEE-EEEEEEecCCCCCHHHhhheeeE
Confidence            4456777899999999999999987665  999999999999999995 69999997543   344555554


No 11 
>PF06907 Latexin:  Latexin;  InterPro: IPR009684 This family consists of several animal specific latexin and proteins related to latexin that belong to MEROPS proteinase inhibitor family I47, clan I- [].  Latexin, a protein possessing inhibitory activity against rat carboxypeptidase A1 (CPA1) and CPA2 (MEROPS peptidase family M14A), is expressed in a neuronal subset in the cerebral cortex and cells in other neural and non-neural tissues of rat [, ]. OCX-32, the 32 kDa eggshell matrix protein, is present at high levels in the uterine fluid during the terminal phase of eggshell formation, and is localised predominantly in the outer eggshell. The timing of OCX-32 secretion into the uterine fluid suggests that it may play a role in the termination of mineral deposition []. OCX-32 protein possesses limited identity (32%) to two unrelated proteins: latexin and to a skin protein that is encoded by a retinoic acid receptor-responsive gene, TIG1. Tazarotene Induced Gene 1 (TIG1) is a putative 228 transmembrane protein with a small N-terminal intracellular region, a single membrane-spanning hydrophobic region, and a large C-terminal extracellular region containing a glycosylation signal. TIG1 is up-regulated by retinoic acid receptor but not by retinoid X receptor-specific synthetic retinoids []. TIG1 may be a tumour suppressor gene whose diminished expression is involved in the malignant progression of prostate cancer [].; PDB: 1WNH_A 2BO9_B.
Probab=96.38  E-value=0.55  Score=40.24  Aligned_cols=167  Identities=16%  Similarity=0.090  Sum_probs=91.8

Q ss_pred             CCCHHHHHHHHHHHHHHHhhcCCc---eeEEEEEEEEEEeec--cEEEEEEEEEEe---CCcceEEEEEEEEecCCCcee
Q 027502           42 QNSAEIEGLARFAVQEHNKKENAL---LQFARVLKAKEQVVA--GKLYYLTLEVID---AGKNKIYEAKIWVKPWINFKQ  113 (222)
Q Consensus        42 ~~d~~v~~~a~fAv~~~N~~sn~~---~~~~kV~~a~~QVVa--G~nY~l~v~v~~---~~~~~~c~~~V~~~PW~~~~~  113 (222)
                      ++.-..+.+|+-|..-+|-..+++   +.+.+|.+|...++.  |-+|+|++.+.+   ++....|.|+|+-. -.+..-
T Consensus         4 p~h~~a~rAA~va~hy~N~~~GSP~~l~~l~~V~~a~~e~ip~~G~Ky~L~FSte~~~~~e~~g~CsA~V~f~-~qkp~P   82 (220)
T PF06907_consen    4 PSHRPAQRAARVAQHYINYRAGSPSRLFVLQQVQKARAEDIPGEGCKYDLVFSTEEYIEGEHLGNCSAEVFFK-NQKPRP   82 (220)
T ss_dssp             TTSHHHHHHHHHHHHHHHHHH-BTTB-EEEEEEEEEEEEEETTTEEEEEEEEEEEETTT---EEEEEEEEEET-T-----
T ss_pred             CcchHHHHHHHHHHHHhccccCCCceeeehhhhhhhhheeccCCCCEEEEEEEhHHhhcCCceeEeEEEEEec-CCCCCC
Confidence            355678899999999999998887   456999999999985  799999999997   45788999999982 222334


Q ss_pred             eEEEeeCCCCCCccc--cccc----ccccCCCCcceeccCC----CHHHHHHHHHHH--HHHHhh--ccCccceeEEEEE
Q 027502          114 LQEFKHAEHGPFSAL--SDLN----LKRGCHGQEWLAVSTN----DLEVKNAANHAV--KSMQRK--SNSLFLYELLEIL  179 (222)
Q Consensus       114 l~s~~c~~~~~~~~~--~~~~----~k~~~~~gg~~~i~~~----d~~v~e~a~fAv--~~~N~~--sn~~~~~~~~kV~  179 (222)
                      -.++.|.......+.  .|..    .|....+---.+||.+    +|+..-+=..|.  ..|=.-  |...-.|....|.
T Consensus        83 ~V~vtc~~~~~k~~~qeeD~~fY~~~k~~~~pl~a~~IPDs~G~i~p~m~P~w~La~v~ssyVmwq~STe~t~Y~maQi~  162 (220)
T PF06907_consen   83 AVNVTCTGLIEKNKRQEEDYAFYQQMKSLKKPLSAQSIPDSHGNIEPEMEPVWHLAIVASSYVMWQKSTENTLYNMAQIK  162 (220)
T ss_dssp             EEEEEECS-------HHHHHHHHHHHHC-SS--EEEEES-TTS---HHHHHHHHHHHHHHHHHHHHH--TT--EEEEEEE
T ss_pred             cEEEEEEeccccCcchhHHHHHHHHHHhhcCccccccCCCCcCCcCccccchhhhhhhheeeEEEeccccceeheeeeec
Confidence            568899876532221  1100    0000000111223221    444444322222  233222  2223356788888


Q ss_pred             Eeeeeee--cceeEEEEEEEEccCCccceEEEE
Q 027502          180 QAKAKVI--EDYAKFELHLKLRRGSEEEKHWVE  210 (222)
Q Consensus       180 ~a~~QVV--aG~~yy~l~l~~~~g~~~~~~~~~  210 (222)
                      +++++--  +-+. |+.+|-+.+-...++--|.
T Consensus       163 sVkQ~kr~DD~i~-FdytVLLHe~~sQEIipc~  194 (220)
T PF06907_consen  163 SVKQWKRNDDFIE-FDYTVLLHEMSSQEIIPCQ  194 (220)
T ss_dssp             EEEEE--SSS-EE-EEEEEEEEETTTTEEEEEE
T ss_pred             ceeeeeeccceee-eceEEEEeecccccceeeE
Confidence            8875533  2334 7888888887777764443


No 12 
>TIGR01572 A_thl_para_3677 Arabidopsis paralogous family TIGR01572. This model describes a paralogous family of hypothetical proteins in Arabidopsis thaliana. No homologs are detected from other species. Length heterogeneity within the family is attributable partly to a 21-residue repeat present in from zero to three tandem copies. The central region of the repeat resembles the pattern [VIF][FY][QK]GX[LM]P[DEK]XXXDDAL.
Probab=93.01  E-value=6  Score=34.90  Aligned_cols=73  Identities=19%  Similarity=0.204  Sum_probs=59.5

Q ss_pred             HHHHHHHHHHHHHhhcCCceeEEEEEEEEEEeeccEEEEEEEEEEeC---CcceEEEEEEEEecCCCceeeEEEeeC
Q 027502           47 IEGLARFAVQEHNKKENALLQFARVLKAKEQVVAGKLYYLTLEVIDA---GKNKIYEAKIWVKPWINFKQLQEFKHA  120 (222)
Q Consensus        47 v~~~a~fAv~~~N~~sn~~~~~~kV~~a~~QVVaG~nY~l~v~v~~~---~~~~~c~~~V~~~PW~~~~~l~s~~c~  120 (222)
                      ++-.|+.++.-||-..+..+.|..|.+.-.+..+-+.|+||+++-+-   +..+.|+..|.++- .+...|+..-|-
T Consensus        42 vklyAr~GLH~YN~~~GTNlel~~v~K~N~~~~~~~syyITL~A~DP~s~~s~qTFQtrV~e~~-~~~L~ltt~iaR  117 (265)
T TIGR01572        42 VKIYARVGLHRYNFLEGTNLELDHVDKFNKRMCALSSYYITLLAVDPDSRFLQQTFQVRVDEQK-LETLDLTVEIAR  117 (265)
T ss_pred             HHHHHHhhhhhhhhccCccceehhhhhhccchhhheeeeEEEEEecCCccccceEEEEEEEecc-CCcEEEEEEEEe
Confidence            58899999999999998899999999999999999999999999984   46778888887753 234455544443


No 13 
>PF00666 Cathelicidins:  Cathelicidin;  InterPro: IPR001894 The precursor sequences of a number of antimicrobial peptides secreted by neutrophils (polymorphonuclear leukocytes) upon activation have been found to be evolutionarily related and are collectively known as cathelicidins []. Structurally, these proteins consist of three domains: a signal sequence, a conserved region of about 100 residues that contains four cysteines involved in two disulphide bonds, and a highly divergent C-terminal section of variable size. It is in this C-terminal section that the antibacterial peptides are found; they are proteolytically processed from their precursor by enzymes such as elastase. This structure is shown in the following schematic representation:  +---+--------------------------------+--------------------+ |Sig| Propeptide C C C C | Antibacterial pep. | +---+----------------|--|--|--|------+--------------------+ | | | | +--+ +--+ 'C': conserved cysteine involved in a disulphide bond. ; GO: 0006952 defense response, 0005576 extracellular region; PDB: 1KWI_A 1PFP_A 1LXE_A 1N5P_A 1N5H_A.
Probab=89.58  E-value=0.88  Score=31.96  Aligned_cols=52  Identities=17%  Similarity=0.127  Sum_probs=34.7

Q ss_pred             HHHHHHHHHHHHHhhccCccceeEEEEEEeeeeeec-ce-eEEEEEEEEccCCc
Q 027502          152 VKNAANHAVKSMQRKSNSLFLYELLEILQAKAKVIE-DY-AKFELHLKLRRGSE  203 (222)
Q Consensus       152 v~e~a~fAv~~~N~~sn~~~~~~~~kV~~a~~QVVa-G~-~yy~l~l~~~~g~~  203 (222)
                      ++||+..||..||+++.+.+.|++....---.+... ++ +...++|+-+.|.+
T Consensus         4 Y~eav~~Av~~yN~~s~~~nlfRLLe~~p~P~~~~~~~~~~pl~FtIkETVC~~   57 (67)
T PF00666_consen    4 YEEAVLRAVDFYNQGSSGENLFRLLELDPPPGWDEDPSTPKPLNFTIKETVCPK   57 (67)
T ss_dssp             CHHHHHHHHHHHHHCS-SSEEEEEEEE---SSSSSSSSS-EEEEEEEEEEEEES
T ss_pred             HHHHHHHHHHHHhcCCCccCceeeeeccCCCCCCCCcCcceeeEEEEeeccCCC
Confidence            578999999999999999998888777644444333 23 33456777666554


No 14 
>PF06907 Latexin:  Latexin;  InterPro: IPR009684 This family consists of several animal specific latexin and proteins related to latexin that belong to MEROPS proteinase inhibitor family I47, clan I- [].  Latexin, a protein possessing inhibitory activity against rat carboxypeptidase A1 (CPA1) and CPA2 (MEROPS peptidase family M14A), is expressed in a neuronal subset in the cerebral cortex and cells in other neural and non-neural tissues of rat [, ]. OCX-32, the 32 kDa eggshell matrix protein, is present at high levels in the uterine fluid during the terminal phase of eggshell formation, and is localised predominantly in the outer eggshell. The timing of OCX-32 secretion into the uterine fluid suggests that it may play a role in the termination of mineral deposition []. OCX-32 protein possesses limited identity (32%) to two unrelated proteins: latexin and to a skin protein that is encoded by a retinoic acid receptor-responsive gene, TIG1. Tazarotene Induced Gene 1 (TIG1) is a putative 228 transmembrane protein with a small N-terminal intracellular region, a single membrane-spanning hydrophobic region, and a large C-terminal extracellular region containing a glycosylation signal. TIG1 is up-regulated by retinoic acid receptor but not by retinoid X receptor-specific synthetic retinoids []. TIG1 may be a tumour suppressor gene whose diminished expression is involved in the malignant progression of prostate cancer [].; PDB: 1WNH_A 2BO9_B.
Probab=83.11  E-value=15  Score=31.56  Aligned_cols=67  Identities=16%  Similarity=0.186  Sum_probs=48.0

Q ss_pred             ccCCCHHHHHHHHHHHHHHHhhccCcc-ceeEEEEEEeeeeeecce--eEEEEEEEEccCC---ccceEEEEEE
Q 027502          145 VSTNDLEVKNAANHAVKSMQRKSNSLF-LYELLEILQAKAKVIEDY--AKFELHLKLRRGS---EEEKHWVEII  212 (222)
Q Consensus       145 i~~~d~~v~e~a~fAv~~~N~~sn~~~-~~~~~kV~~a~~QVVaG~--~yy~l~l~~~~g~---~~~~~~~~v~  212 (222)
                      ++|+.--.++||+-|+-=+|-+..+.+ .|.+.+|.+|+..++.|.  + |+|.+.+.+-.   ...+=.|+|.
T Consensus         2 ~~p~h~~a~rAA~va~hy~N~~~GSP~~l~~l~~V~~a~~e~ip~~G~K-y~L~FSte~~~~~e~~g~CsA~V~   74 (220)
T PF06907_consen    2 INPSHRPAQRAARVAQHYINYRAGSPSRLFVLQQVQKARAEDIPGEGCK-YDLVFSTEEYIEGEHLGNCSAEVF   74 (220)
T ss_dssp             --TTSHHHHHHHHHHHHHHHHHH-BTTB-EEEEEEEEEEEEEETTTEEE-EEEEEEEEETTT---EEEEEEEEE
T ss_pred             CCCcchHHHHHHHHHHHHhccccCCCceeeehhhhhhhhheeccCCCCE-EEEEEEhHHhhcCCceeEeEEEEE
Confidence            456666789999999999999988776 578899999999999765  6 67998888732   2233445554


No 15 
>TIGR01572 A_thl_para_3677 Arabidopsis paralogous family TIGR01572. This model describes a paralogous family of hypothetical proteins in Arabidopsis thaliana. No homologs are detected from other species. Length heterogeneity within the family is attributable partly to a 21-residue repeat present in from zero to three tandem copies. The central region of the repeat resembles the pattern [VIF][FY][QK]GX[LM]P[DEK]XXXDDAL.
Probab=72.79  E-value=8.1  Score=34.10  Aligned_cols=67  Identities=12%  Similarity=0.003  Sum_probs=56.1

Q ss_pred             HHHHHHHHHHHHHhhccCccceeEEEEEEeeeeeecceeEEEEEEEEccCC---ccceEEEEEEeCCCCceec
Q 027502          152 VKNAANHAVKSMQRKSNSLFLYELLEILQAKAKVIEDYAKFELHLKLRRGS---EEEKHWVEIIKNSEGKFYL  221 (222)
Q Consensus       152 v~e~a~fAv~~~N~~sn~~~~~~~~kV~~a~~QVVaG~~yy~l~l~~~~g~---~~~~~~~~v~~~~~~~~~~  221 (222)
                      |+-.|+.++-.||......+  ++.+|.+.-++..+-+.| +||+++-+..   ..+.|...|.|...|+..|
T Consensus        42 vklyAr~GLH~YN~~~GTNl--el~~v~K~N~~~~~~~sy-yITL~A~DP~s~~s~qTFQtrV~e~~~~~L~l  111 (265)
T TIGR01572        42 VKIYARVGLHRYNFLEGTNL--ELDHVDKFNKRMCALSSY-YITLLAVDPDSRFLQQTFQVRVDEQKLETLDL  111 (265)
T ss_pred             HHHHHHhhhhhhhhccCccc--eehhhhhhccchhhheee-eEEEEEecCCccccceEEEEEEEeccCCcEEE
Confidence            68899999999999976665  999999999999988884 6999999743   4667999999988777543


No 16 
>PF01789 PsbP:  PsbP;  InterPro: IPR002683 Oxygenic photosynthesis uses two multi-subunit photosystems (I and II) located in the cell membranes of cyanobacteria and in the thylakoid membranes of chloroplasts in plants and algae. Photosystem II (PSII) has a P680 reaction centre containing chlorophyll 'a' that uses light energy to carry out the oxidation (splitting) of water molecules, and to produce ATP via a proton pump. Photosystem I (PSI) has a P700 reaction centre containing chlorophyll that takes the electron and associated hydrogen donated from PSII to reduce NADP+ to NADPH. Both ATP and NADPH are subsequently used in the light-independent reactions to convert carbon dioxide to glucose using the hydrogen atom extracted from water by PSII, releasing oxygen as a by-product. PSII is a multisubunit protein-pigment complex containing polypeptides both intrinsic and extrinsic to the photosynthetic membrane [, ]. Within the core of the complex, the chlorophyll and beta-carotene pigments are mainly bound to the antenna proteins CP43 (PsbC) and CP47 (PsbB), which pass the excitation energy on to the reaction centre proteins D1 (Qb, PsbA) and D2 (Qa, PsbD) that bind all the redox-active cofactors involved in the energy conversion process. The PSII oxygen-evolving complex (OEC) oxidises water to provide protons for use by PSI, and consists of OEE1 (PsbO), OEE2 (PsbP) and OEE3 (PsbQ). The remaining subunits in PSII are of low molecular weight (less than 10 kDa), and are involved in PSII assembly, stabilisation, dimerisation, and photo-protection [].  In PSII, the oxygen-evolving complex (OEC) is responsible for catalysing the splitting of water to O(2) and 4H+. The OEC is composed of a cluster of manganese, calcium and chloride ions bound to extrinsic proteins. In cyanobacteria there are five extrinsic proteins in OEC (PsbO, PsbP-like, PsbQ-like, PsbU and PsbV), while in plants there are only three (PsbO, PsbP and PsbQ), PsbU and PsbV having been lost during the evolution of green plants []. This family represents the PSII OEC protein PsbP. Both PsbP and PsbQ (IPR008797 from INTERPRO) are regulators that are necessary for the biogenesis of optically active PSII. PsbP increases the affinity of the water oxidation site for chloride ions and provides the conditions required for high affinity binding of calcium ions [, ]. The crystal structure of PsbP from Nicotiana tabacum (Common tobacco) revealed a two-domain structure, where domain 1 may play a role in the ion retention activity in PSII, the N-terminal residues being essential for calcium and chloride ion retention activity []. PsbP is encoded in the nuclear genome in plants.; GO: 0005509 calcium ion binding, 0015979 photosynthesis, 0009523 photosystem II, 0009654 oxygen evolving complex, 0019898 extrinsic to membrane; PDB: 2VU4_A 1V2B_A 2LNJ_A 2XB3_A.
Probab=70.27  E-value=29  Score=28.26  Aligned_cols=58  Identities=9%  Similarity=0.146  Sum_probs=36.8

Q ss_pred             HHHHHHHHHHHhhccCccceeEEEEEEeeeeeecceeEEEEEEEEccCC-ccceEEEEEEeC
Q 027502          154 NAANHAVKSMQRKSNSLFLYELLEILQAKAKVIEDYAKFELHLKLRRGS-EEEKHWVEIIKN  214 (222)
Q Consensus       154 e~a~fAv~~~N~~sn~~~~~~~~kV~~a~~QVVaG~~yy~l~l~~~~g~-~~~~~~~~v~~~  214 (222)
                      +++...+........+.   +..+++++.+....|..||.+...+..+. ....+.+.+..+
T Consensus        84 ~va~~l~~~~~~~~~~~---~~a~li~a~~~~~~g~~yY~~Ey~~~~~~~~~rh~l~~~tv~  142 (175)
T PF01789_consen   84 EVAERLLNGELASPGSG---REAELISASEREVDGKTYYEYEYTVQSPNEGRRHNLAVVTVK  142 (175)
T ss_dssp             HHHHHHHHHCCCHCTSS---EEEEEEEEEEEEETTEEEEEEEEEEEETTEEEEEEEEEEEEE
T ss_pred             HHHHHHhhhhcccccCC---cceEEEEeeeeecCCccEEEEEEEeccCCCcccEEEEEEEEE
Confidence            34444444433333333   88899999999999999988888777655 333334444333


No 17 
>PRK13859 type IV secretion system lipoprotein VirB7; Provisional
Probab=56.69  E-value=9.9  Score=25.29  Aligned_cols=29  Identities=17%  Similarity=0.273  Sum_probs=19.7

Q ss_pred             HHHHHHHHhhccccCCCccc--------CCceeeeCC
Q 027502           11 VLVCGFIELGLCREGNFIQM--------RPGGVYDYG   39 (222)
Q Consensus        11 ~~~~~~~~~~~~~~~~~~~~--------l~GG~~~i~   39 (222)
                      ++||+++++++|+..+...+        -+|-|.|.+
T Consensus         4 ~lL~l~l~La~CqT~D~lAtckGpiFpLNVgrWqptp   40 (55)
T PRK13859          4 CLLCLALALAGCQTNDTLASCKGPIFPLNVGRWQPTP   40 (55)
T ss_pred             hHHHHHHHHHhccccCccccccCCccccccccccCCh
Confidence            57788888998886655433        356676664


No 18 
>PF05679 CHGN:  Chondroitin N-acetylgalactosaminyltransferase;  InterPro: IPR008428 This family represents Chondroitin N-acetylgalactosaminyltransferase. Proteins have a type II transmembrane topology. The enzyme is involved in the biosynthetic initiation and elongation of chondroitin sulphate and is the key enzyme responsible for the selective chain assembly of chondroitin/dermatan sulphate on the linkage region tetrasaccharide common to various proteoglycans containing chondroitin/dermatan sulphate or heparin/heparan sulphate chains. ; GO: 0016758 transferase activity, transferring hexosyl groups, 0032580 Golgi cisterna membrane
Probab=53.15  E-value=43  Score=32.17  Aligned_cols=49  Identities=20%  Similarity=0.363  Sum_probs=43.4

Q ss_pred             CHHHHHHHHHHHHHHHhhcCCceeEEEEEEEEEEe--eccEEEEEEEEEEe
Q 027502           44 SAEIEGLARFAVQEHNKKENALLQFARVLKAKEQV--VAGKLYYLTLEVID   92 (222)
Q Consensus        44 d~~v~~~a~fAv~~~N~~sn~~~~~~kV~~a~~QV--VaG~nY~l~v~v~~   92 (222)
                      -.++.++.+.|++.+|+.+...+.|.+++.+.+.+  .-|+-|.|++.+..
T Consensus       160 ~~dl~~vi~~a~~~ln~~~~~~~~~~~l~~GY~R~dp~rG~~Y~Ldl~l~~  210 (499)
T PF05679_consen  160 REDLDDVIEQAMEELNRKSRRVLEFRDLINGYRRFDPTRGMDYILDLLLKY  210 (499)
T ss_pred             HHHHHHHHHHHHHHHhccccccEEeeeeeeEEEEecCCCCceEEEEEEEee
Confidence            47899999999999999888779999999998777  46999999998875


No 19 
>PF00666 Cathelicidins:  Cathelicidin;  InterPro: IPR001894 The precursor sequences of a number of antimicrobial peptides secreted by neutrophils (polymorphonuclear leukocytes) upon activation have been found to be evolutionarily related and are collectively known as cathelicidins []. Structurally, these proteins consist of three domains: a signal sequence, a conserved region of about 100 residues that contains four cysteines involved in two disulphide bonds, and a highly divergent C-terminal section of variable size. It is in this C-terminal section that the antibacterial peptides are found; they are proteolytically processed from their precursor by enzymes such as elastase. This structure is shown in the following schematic representation:  +---+--------------------------------+--------------------+ |Sig| Propeptide C C C C | Antibacterial pep. | +---+----------------|--|--|--|------+--------------------+ | | | | +--+ +--+ 'C': conserved cysteine involved in a disulphide bond. ; GO: 0006952 defense response, 0005576 extracellular region; PDB: 1KWI_A 1PFP_A 1LXE_A 1N5P_A 1N5H_A.
Probab=52.98  E-value=17  Score=25.47  Aligned_cols=46  Identities=15%  Similarity=0.066  Sum_probs=26.2

Q ss_pred             HHHHHHHHHHHHHhhcCCceeEEEEEEEEEEee----ccEEEEEEEEEEeC
Q 027502           47 IEGLARFAVQEHNKKENALLQFARVLKAKEQVV----AGKLYYLTLEVIDA   93 (222)
Q Consensus        47 v~~~a~fAv~~~N~~sn~~~~~~kV~~a~~QVV----aG~nY~l~v~v~~~   93 (222)
                      ++++...||+.||+.+.+. .+.+++.+.-|--    .++.--+.+.|.++
T Consensus         4 Y~eav~~Av~~yN~~s~~~-nlfRLLe~~p~P~~~~~~~~~~pl~FtIkET   53 (67)
T PF00666_consen    4 YEEAVLRAVDFYNQGSSGE-NLFRLLELDPPPGWDEDPSTPKPLNFTIKET   53 (67)
T ss_dssp             CHHHHHHHHHHHHHCS-SS-EEEEEEEE---SSSSSSSSS-EEEEEEEEEE
T ss_pred             HHHHHHHHHHHHhcCCCcc-CceeeeeccCCCCCCCCcCcceeeEEEEeec
Confidence            5789999999999998764 3344555544432    22344555555553


No 20 
>PLN00042 photosystem II oxygen-evolving enhancer protein 2; Provisional
Probab=52.23  E-value=40  Score=29.82  Aligned_cols=24  Identities=13%  Similarity=0.235  Sum_probs=20.3

Q ss_pred             EEEEEeeeeeecceeEEEEEEEEc
Q 027502          176 LEILQAKAKVIEDYAKFELHLKLR  199 (222)
Q Consensus       176 ~kV~~a~~QVVaG~~yy~l~l~~~  199 (222)
                      .+|+++++..+.|..||.|.+.+.
T Consensus       185 a~Lleas~re~dGk~YY~lE~~~~  208 (260)
T PLN00042        185 AAVLESSTQEVGGKPYYYLSVLTR  208 (260)
T ss_pred             eeEEEeeeEEeCCeEEEEEEEEEe
Confidence            478999999999999987777754


No 21 
>PLN00067 PsbP domain-containing protein 6; Provisional
Probab=48.85  E-value=39  Score=29.91  Aligned_cols=28  Identities=21%  Similarity=0.304  Sum_probs=23.4

Q ss_pred             eeEEEEEEeeeeeecceeEEEEEEEEcc
Q 027502          173 YELLEILQAKAKVIEDYAKFELHLKLRR  200 (222)
Q Consensus       173 ~~~~kV~~a~~QVVaG~~yy~l~l~~~~  200 (222)
                      +...+|++|++..+.|..||.++++..-
T Consensus       190 ~~~~eLLeAs~re~dGktYY~~E~~tp~  217 (263)
T PLN00067        190 YDPDELLETSVEKIGDQTYYKYVLETPF  217 (263)
T ss_pred             CCCcceEEeeeEeeCCeEEEEEEEEecC
Confidence            4566899999999999999988887653


No 22 
>PF15240 Pro-rich:  Proline-rich
Probab=43.48  E-value=16  Score=30.54  Aligned_cols=16  Identities=19%  Similarity=0.368  Sum_probs=12.0

Q ss_pred             HHHHHHHHHHhhcccc
Q 027502            9 LSVLVCGFIELGLCRE   24 (222)
Q Consensus         9 ~~~~~~~~~~~~~~~~   24 (222)
                      |++|.++||||++|++
T Consensus         3 lVLLSvALLALSSAQ~   18 (179)
T PF15240_consen    3 LVLLSVALLALSSAQS   18 (179)
T ss_pred             hHHHHHHHHHhhhccc
Confidence            4556688999999973


No 23 
>CHL00132 psaF photosystem I subunit III; Validated
Probab=37.71  E-value=71  Score=26.73  Aligned_cols=27  Identities=11%  Similarity=0.124  Sum_probs=21.1

Q ss_pred             CCceeeeCCCCCCCHHHHHHHHHHHHHHHh
Q 027502           31 RPGGVYDYGGNQNSAEIEGLARFAVQEHNK   60 (222)
Q Consensus        31 l~GG~~~i~~~~~d~~v~~~a~fAv~~~N~   60 (222)
                      -.+|.+|=+   ++|.+++-++-++.++.+
T Consensus        25 d~agLtpCs---es~aF~kR~~~~~k~Le~   51 (185)
T CHL00132         25 DVAGLTPCS---ESPAFQKRLNNSVKKLEN   51 (185)
T ss_pred             cccCCccCc---cCHHHHHHHHHHHHHHHh
Confidence            478888887   789999988888866443


No 24 
>PF08294 TIM21:  TIM21;  InterPro: IPR013261 TIM21 interacts with the outer mitochondrial TOM complex and promotes the insertion of proteins into the inner mitochondrial membrane [].; PDB: 2CIU_A.
Probab=36.26  E-value=40  Score=27.01  Aligned_cols=73  Identities=10%  Similarity=0.167  Sum_probs=35.5

Q ss_pred             CCCHHHHHHHHHHHHHHHhhccCccceeEEEEEEeeeeeecceeEEEEEEEEccCCccceEEEEEEeCCC-Cce
Q 027502          147 TNDLEVKNAANHAVKSMQRKSNSLFLYELLEILQAKAKVIEDYAKFELHLKLRRGSEEEKHWVEIIKNSE-GKF  219 (222)
Q Consensus       147 ~~d~~v~e~a~fAv~~~N~~sn~~~~~~~~kV~~a~~QVVaG~~yy~l~l~~~~g~~~~~~~~~v~~~~~-~~~  219 (222)
                      .+||++++++.--+..|...+....--+=..+++-+.---.|..+..+..-+....+.-.=.+|+.++++ ++|
T Consensus        51 ~~d~~v~~~LG~~ikayGe~~~~~Rw~R~R~~~s~~~~d~~G~eh~~m~F~V~G~~~~G~V~~e~~k~~~~~~~  124 (145)
T PF08294_consen   51 KKDPRVQDLLGEPIKAYGEETGRNRWRRNRPIVSHREYDKDGREHMRMKFYVEGPRGKGVVHLEMVKDDGSGEY  124 (145)
T ss_dssp             HH-HHHHHHT----EEEE-EEE-SS-EEE----EEEEE-TTS-EEEEEEEEEE-SS-EEEEEEEEE--SS-SS-
T ss_pred             hcCHHHHHHhCCCeEEecCCCCCCcccccCCccceEEEcCCCCEEEEEEEEEEeCCCeEEEEEEEEECCCCCCe
Confidence            4799999999888888887765221111223444444456888875666666654455556788887775 665


No 25 
>PLN00059 PsbP domain-containing protein 1; Provisional
Probab=36.25  E-value=2.2e+02  Score=25.47  Aligned_cols=44  Identities=7%  Similarity=0.222  Sum_probs=30.4

Q ss_pred             HHHHHHHHHHHhh-----ccCccceeEEEEEEeeeeee-cceeEEEEEEEEcc
Q 027502          154 NAANHAVKSMQRK-----SNSLFLYELLEILQAKAKVI-EDYAKFELHLKLRR  200 (222)
Q Consensus       154 e~a~fAv~~~N~~-----sn~~~~~~~~kV~~a~~QVV-aG~~yy~l~l~~~~  200 (222)
                      +++..-++++...     .++.   +-.++++|++... +|..||.|...+.-
T Consensus       172 eVgerLlkqvLa~f~str~Gsg---ReaeLVsA~~Re~~DGktYY~lEY~Vks  221 (286)
T PLN00059        172 EVGKRVLRQYLTEFMSTRLGVK---REANILSTSSRVADDGKLYYQVEVNIKS  221 (286)
T ss_pred             HHHHHHHHHHhcccccccCCCC---cceEEEEeeeEEccCCcEEEEEEEEEEc
Confidence            4556666666543     1222   4678999998866 89999988888765


No 26 
>COG3360 Uncharacterized conserved protein [Function unknown]
Probab=32.89  E-value=1.8e+02  Score=20.56  Aligned_cols=45  Identities=24%  Similarity=0.316  Sum_probs=31.7

Q ss_pred             HHHHHHHHHHHHHHhhcCCceeEEEEEEEEEEeecc--EEEEEEEEEE
Q 027502           46 EIEGLARFAVQEHNKKENALLQFARVLKAKEQVVAG--KLYYLTLEVI   91 (222)
Q Consensus        46 ~v~~~a~fAv~~~N~~sn~~~~~~kV~~a~~QVVaG--~nY~l~v~v~   91 (222)
                      .+.++++-|+..-. ++-+.+.+.+|++-+-+|+.|  ..|.++++++
T Consensus        18 S~d~Ai~~Ai~RA~-~t~~~l~wfeV~~~rg~v~~g~v~hyqv~lkVg   64 (71)
T COG3360          18 SIDAAIANAIARAA-DTLDNLDWFEVVETRGHVVDGAVAHYQVTLKVG   64 (71)
T ss_pred             cHHHHHHHHHHHHH-hhhhcceEEEEEeecccEeecceEEEEEEEEEE
Confidence            34456666665432 234568889999999999988  5688888776


No 27 
>PTZ00444 hypothetical protein; Provisional
Probab=32.17  E-value=38  Score=28.32  Aligned_cols=23  Identities=30%  Similarity=0.615  Sum_probs=21.3

Q ss_pred             CCchhhHHHHHHHHHHHHhhccc
Q 027502            1 MNRYSVIVLSVLVCGFIELGLCR   23 (222)
Q Consensus         1 ~~~~~~~~~~~~~~~~~~~~~~~   23 (222)
                      |++++++++++|+.+....++|-
T Consensus         1 m~~~~~~~~~~l~~~~~~~~a~l   23 (184)
T PTZ00444          1 MRQRSLLFLLLLVFSYINFSACL   23 (184)
T ss_pred             CchHHHHHHHHHHHHHHHHHHHh
Confidence            89999999999999999999885


No 28 
>PF05679 CHGN:  Chondroitin N-acetylgalactosaminyltransferase;  InterPro: IPR008428 This family represents Chondroitin N-acetylgalactosaminyltransferase. Proteins have a type II transmembrane topology. The enzyme is involved in the biosynthetic initiation and elongation of chondroitin sulphate and is the key enzyme responsible for the selective chain assembly of chondroitin/dermatan sulphate on the linkage region tetrasaccharide common to various proteoglycans containing chondroitin/dermatan sulphate or heparin/heparan sulphate chains. ; GO: 0016758 transferase activity, transferring hexosyl groups, 0032580 Golgi cisterna membrane
Probab=32.02  E-value=1.8e+02  Score=27.94  Aligned_cols=48  Identities=17%  Similarity=0.317  Sum_probs=38.9

Q ss_pred             HHHHHHHHHHHHHHHhhccCccceeEEEEEEeeeee--ecceeEEEEEEEEcc
Q 027502          150 LEVKNAANHAVKSMQRKSNSLFLYELLEILQAKAKV--IEDYAKFELHLKLRR  200 (222)
Q Consensus       150 ~~v~e~a~fAv~~~N~~sn~~~~~~~~kV~~a~~QV--VaG~~yy~l~l~~~~  200 (222)
                      .++.++.+.|++..|+.+..  .+++.+++.+-..+  .-|+.| .|++.+..
T Consensus       161 ~dl~~vi~~a~~~ln~~~~~--~~~~~~l~~GY~R~dp~rG~~Y-~Ldl~l~~  210 (499)
T PF05679_consen  161 EDLDDVIEQAMEELNRKSRR--VLEFRDLINGYRRFDPTRGMDY-ILDLLLKY  210 (499)
T ss_pred             HHHHHHHHHHHHHHhccccc--cEEeeeeeeEEEEecCCCCceE-EEEEEEee
Confidence            68999999999999998863  44789999998775  579994 68877653


No 29 
>PRK10081 entericidin B membrane lipoprotein; Provisional
Probab=31.71  E-value=32  Score=22.56  Aligned_cols=22  Identities=14%  Similarity=0.194  Sum_probs=12.4

Q ss_pred             CCchhhHHHHHHHHHHHHhhcc
Q 027502            1 MNRYSVIVLSVLVCGFIELGLC   22 (222)
Q Consensus         1 ~~~~~~~~~~~~~~~~~~~~~~   22 (222)
                      |-+-.+..++++++++++++.|
T Consensus         1 MmKk~i~~i~~~l~~~~~l~~C   22 (48)
T PRK10081          1 MVKKTIAAIFSVLVLSTVLTAC   22 (48)
T ss_pred             ChHHHHHHHHHHHHHHHHHhhh
Confidence            3344445555556666667766


No 30 
>COG3360 Uncharacterized conserved protein [Function unknown]
Probab=31.20  E-value=1.8e+02  Score=20.60  Aligned_cols=45  Identities=20%  Similarity=0.457  Sum_probs=31.1

Q ss_pred             HHHHHHHHHHHHHhhccCccceeEEEEEEeeeeeecce-eEEEEEEEEc
Q 027502          152 VKNAANHAVKSMQRKSNSLFLYELLEILQAKAKVIEDY-AKFELHLKLR  199 (222)
Q Consensus       152 v~e~a~fAv~~~N~~sn~~~~~~~~kV~~a~~QVVaG~-~yy~l~l~~~  199 (222)
                      +.+|++-|+.   +.+.+.-....-+|+.-+-+|+.|. .+|.++++++
T Consensus        19 ~d~Ai~~Ai~---RA~~t~~~l~wfeV~~~rg~v~~g~v~hyqv~lkVg   64 (71)
T COG3360          19 IDAAIANAIA---RAADTLDNLDWFEVVETRGHVVDGAVAHYQVTLKVG   64 (71)
T ss_pred             HHHHHHHHHH---HHHhhhhcceEEEEEeecccEeecceEEEEEEEEEE
Confidence            4445554443   3344555668889999999999887 5578888876


No 31 
>smart00773 WGR Proposed nucleic acid binding domain. This domain is named after its most conserved central motif. It is found in a variety of polyA polymerases as well as in molybdate metabolism regulators (e.g. in E.coli) and other proteins of unknown function. The domain is found in isolation in some proteins and is between 70 and 80 residues in length. It is proposed that it may be a nucleic acid binding domain.
Probab=31.10  E-value=1e+02  Score=21.78  Aligned_cols=30  Identities=7%  Similarity=0.256  Sum_probs=19.0

Q ss_pred             EEEEEEEccCCc--cceEEEEEEeCCCCceec
Q 027502          192 FELHLKLRRGSE--EEKHWVEIIKNSEGKFYL  221 (222)
Q Consensus       192 y~l~l~~~~g~~--~~~~~~~v~~~~~~~~~~  221 (222)
                      |..+|+..+.+.  .+-|..+++++..|.|.|
T Consensus         6 ~~~~L~~~d~~~n~nkfy~iql~~~~~~~~~v   37 (84)
T smart00773        6 YDVYLNQTDLASNNNKFYRIQLLEDDFGGYSV   37 (84)
T ss_pred             eEEEEEccccccCCeeEEEEEEEEcCCCCEEE
Confidence            466666665444  455777777777766654


No 32 
>TIGR02105 III_needle type III secretion apparatus needle protein. Type III secretion systems translocate proteins, usually virulence factors, out across both inner and outer membranes of certain Gram-negative bacteria and further across the plasma membrane and into the cytoplasm of the host cell. This protein, termed YscF in Yersinia, and EscF, PscF, EprI, etc. in other systems, forms the needle of the injection apparatus.
Probab=28.60  E-value=66  Score=22.79  Aligned_cols=23  Identities=17%  Similarity=0.075  Sum_probs=19.8

Q ss_pred             ccCCCHHHHHHHHHHHHHHHhhc
Q 027502          145 VSTNDLEVKNAANHAVKSMQRKS  167 (222)
Q Consensus       145 i~~~d~~v~e~a~fAv~~~N~~s  167 (222)
                      -.|+||+..--.+|++.+||--.
T Consensus        29 ~~~~nP~~La~~Q~~~~qYs~~~   51 (72)
T TIGR02105        29 DLPNDPELMAELQFALNQYSAYY   51 (72)
T ss_pred             CCCCCHHHHHHHHHHHHHHHHHH
Confidence            34799999999999999999753


No 33 
>PF12276 DUF3617:  Protein of unknown function (DUF3617);  InterPro: IPR022061  This family of proteins is found in bacteria. Proteins in this family are typically between 155 and 179 amino acids in length. There is a single completely conserved residue C that may be functionally important. 
Probab=28.59  E-value=57  Score=25.80  Aligned_cols=35  Identities=14%  Similarity=0.166  Sum_probs=19.0

Q ss_pred             CCchhhHHHHHHHHHHHHhhccccCCCcccCCceeeeC
Q 027502            1 MNRYSVIVLSVLVCGFIELGLCREGNFIQMRPGGVYDY   38 (222)
Q Consensus         1 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~l~GG~~~i   38 (222)
                      |.+..+++++++++++.+.++++   .....+|=|.--
T Consensus         1 M~~~~~~~~~~~~~~~~~~~~a~---~~~~kpGlWe~t   35 (162)
T PF12276_consen    1 MKRRLLLALALALLALAAAAAAA---APDIKPGLWEVT   35 (162)
T ss_pred             CchHHHHHHHHHHHHhhcccccc---cCCCCCcccEEE
Confidence            77776666666665543333332   223447777544


No 34 
>PF12274 DUF3615:  Protein of unknown function (DUF3615);  InterPro: IPR022059  This domain family is found in bacteria and eukaryotes, and is typically between 86 and 97 amino acids in length. There is a conserved FAE sequence motif. There is a single completely conserved residue F that may be functionally important. 
Probab=26.92  E-value=2.5e+02  Score=20.37  Aligned_cols=50  Identities=10%  Similarity=0.147  Sum_probs=31.4

Q ss_pred             HHhhcc-CccceeEEEEEEeeeeeeccee-EEEEEEEEcc------CCccceEEEEEE
Q 027502          163 MQRKSN-SLFLYELLEILQAKAKVIEDYA-KFELHLKLRR------GSEEEKHWVEII  212 (222)
Q Consensus       163 ~N~~sn-~~~~~~~~kV~~a~~QVVaG~~-yy~l~l~~~~------g~~~~~~~~~v~  212 (222)
                      ||++.+ ....|++.+|+.+..=.-.|.. |+++-..+..      .+....|=|||.
T Consensus         1 Yn~~~~~~~~~yeL~~v~~~~~~~e~~~~~y~HvNF~A~~~~~~~~~~~~~LFFAE~~   58 (96)
T PF12274_consen    1 YNEDHPLLGLEYELVDVLHSCFIFERGGWNYYHVNFTAKTKGPDSDDGSPTLFFAEVS   58 (96)
T ss_pred             CcccCCCCCcCEEEeEEEeeeeeEeCCCcEEEeEEEEEEcCCccCCCCCceEEEEEEe
Confidence            455542 1445799999887644444443 6655555553      235778999999


No 35 
>PRK15344 type III secretion system needle protein SsaG; Provisional
Probab=26.28  E-value=83  Score=22.32  Aligned_cols=22  Identities=14%  Similarity=-0.038  Sum_probs=19.4

Q ss_pred             ccCCCHHHHHHHHHHHHHHHhh
Q 027502          145 VSTNDLEVKNAANHAVKSMQRK  166 (222)
Q Consensus       145 i~~~d~~v~e~a~fAv~~~N~~  166 (222)
                      -+++||+..--++|++.+|+.-
T Consensus        28 ~~~~nP~~ml~lQf~i~QyS~~   49 (71)
T PRK15344         28 NDLLNPESMIKAQFALQQYSTF   49 (71)
T ss_pred             CCCCCHHHHHHHHHHHHHHHHH
Confidence            4588999999999999999874


No 36 
>KOG4306 consensus Glycosylphosphatidylinositol-specific phospholipase C [Signal transduction mechanisms]
Probab=26.09  E-value=42  Score=30.36  Aligned_cols=27  Identities=7%  Similarity=0.256  Sum_probs=20.5

Q ss_pred             EEeeeeeecceeEEEEEEEEccCCccc
Q 027502          179 LQAKAKVIEDYAKFELHLKLRRGSEEE  205 (222)
Q Consensus       179 ~~a~~QVVaG~~yy~l~l~~~~g~~~~  205 (222)
                      +....|+++|+.|++|.|.....+.+.
T Consensus        70 l~i~~QL~~GvRylDlRi~~~~~~~D~   96 (306)
T KOG4306|consen   70 LDIREQLVAGVRYLDLRIGYKLMDPDR   96 (306)
T ss_pred             cchHHHHhhcceEEEEEeeeccCCCCc
Confidence            345679999999888888877666544


No 37 
>PF02995 DUF229:  Protein of unknown function (DUF229);  InterPro: IPR004245 Members of this family are uncharacterised with a long conserved region that may contain several domains.
Probab=25.34  E-value=74  Score=30.50  Aligned_cols=32  Identities=22%  Similarity=0.328  Sum_probs=28.4

Q ss_pred             CCCcceeccCCCHHHHHHHHHHHHHHHhhccC
Q 027502          138 HGQEWLAVSTNDLEVKNAANHAVKSMQRKSNS  169 (222)
Q Consensus       138 ~~gg~~~i~~~d~~v~e~a~fAv~~~N~~sn~  169 (222)
                      .+.+|.+++.+++.++.+|+++|..+|+.-..
T Consensus       450 ~C~~~~~~~~~~~~~~~~a~~~v~~iN~~l~~  481 (497)
T PF02995_consen  450 TCEGWKTIPTNDSLVQRIAKFLVDHINEYLKE  481 (497)
T ss_pred             cCcCccccccCcHHHHHHHHHHHHHHHHHHhc
Confidence            56789999999999999999999999997544


No 38 
>PRK13883 conjugal transfer protein TrbH; Provisional
Probab=24.75  E-value=1.5e+02  Score=23.99  Aligned_cols=20  Identities=30%  Similarity=0.224  Sum_probs=16.1

Q ss_pred             CCCHHHHHHHHHHHHHHHhh
Q 027502           42 QNSAEIEGLARFAVQEHNKK   61 (222)
Q Consensus        42 ~~d~~v~~~a~fAv~~~N~~   61 (222)
                      +..+.-+.+|.-++.++-+.
T Consensus        28 ~s~~~a~~iA~D~v~qL~~~   47 (151)
T PRK13883         28 ASAADQQKLATDAVQQLATL   47 (151)
T ss_pred             cCHHHHHHHHHHHHHHHHHh
Confidence            46788889999999888665


No 39 
>PF10828 DUF2570:  Protein of unknown function (DUF2570);  InterPro: IPR022538 This entry is represented by Bacteriophage IME08, pseT.3. The characteristics of the protein distribution suggest prophage matches in addition to the phage matches.  This is a family of proteins with unknown function. 
Probab=24.68  E-value=65  Score=24.35  Aligned_cols=22  Identities=36%  Similarity=0.428  Sum_probs=15.5

Q ss_pred             CCchhhHHHHHHHHHHHHhhcc
Q 027502            1 MNRYSVIVLSVLVCGFIELGLC   22 (222)
Q Consensus         1 ~~~~~~~~~~~~~~~~~~~~~~   22 (222)
                      |.+|..++|.++++++.+....
T Consensus         1 ~~~~~~~~l~~lvl~L~~~l~~   22 (110)
T PF10828_consen    1 MKKYIYIALAVLVLGLGGWLWY   22 (110)
T ss_pred             ChHHHHHHHHHHHHHHHHHHHH
Confidence            8999888877776666555443


No 40 
>PF09049 SNN_transmemb:  Stannin transmembrane;  InterPro: IPR015135 This region consists of a single highly hydrophobic transmembrane helix that transverses the lipid bilayer at a 20 degree angle with respect to the membrane normal. It contains a conserved cysteine residue (Cys32) that, together with Cys34 found in the stannin unstructured linker domain, constitutes the putative trimethyltin-binding site that resides at the end of the transmembrane domain close to the lipid/solvent interface []. ; PDB: 1ZZA_A.
Probab=24.58  E-value=33  Score=20.25  Aligned_cols=17  Identities=24%  Similarity=0.472  Sum_probs=11.7

Q ss_pred             hhHHHHHHHHHHHHhhc
Q 027502            5 SVIVLSVLVCGFIELGL   21 (222)
Q Consensus         5 ~~~~~~~~~~~~~~~~~   21 (222)
                      -+.+..++|.+..+++.
T Consensus        11 gvvti~viliavaalg~   27 (33)
T PF09049_consen   11 GVVTIIVILIAVAALGA   27 (33)
T ss_dssp             HHHHHHHHHHHHHHHHH
T ss_pred             cEEEehhHHHHHHHHhh
Confidence            46677777777777663


No 41 
>PF07311 Dodecin:  Dodecin;  InterPro: IPR009923 This entry represents proteins with a Dodecin-like topology. Dodecin flavoprotein is a small dodecameric flavin-binding protein from Halobacterium salinarium (Halobacterium halobium) that contains two flavins stacked in a single binding pocket between two tryptophan residues to form an aromatic tetrade []. Dodecin binds riboflavin, although it appears to have a broad specificity for flavins. Lumichrome, a molecule associated with flavin metabolism, appears to be a ligand of dodecin, which could act as a waste-trapping device. ; PDB: 2VYX_L 2DEG_F 2V18_K 2V19_D 2UX9_B 2CZ8_E 2V21_F 2CC8_A 2CCB_A 2VX9_A ....
Probab=24.52  E-value=2.5e+02  Score=19.52  Aligned_cols=47  Identities=21%  Similarity=0.241  Sum_probs=36.7

Q ss_pred             CHHHHHHHHHHHHHHHhhcCCceeEEEEEEEEEEeecc--EEEEEEEEEE
Q 027502           44 SAEIEGLARFAVQEHNKKENALLQFARVLKAKEQVVAG--KLYYLTLEVI   91 (222)
Q Consensus        44 d~~v~~~a~fAv~~~N~~sn~~~~~~kV~~a~~QVVaG--~nY~l~v~v~   91 (222)
                      .....++++-|+.+-++ +=..++..+|..-+-.|..|  ..|+.+++++
T Consensus        13 ~~S~edAv~~Av~~A~k-Tl~ni~~~eV~e~~~~v~dg~i~~y~v~lkv~   61 (66)
T PF07311_consen   13 PKSWEDAVQNAVARASK-TLRNIRWFEVKEQRGHVEDGKITEYQVNLKVS   61 (66)
T ss_dssp             SSHHHHHHHHHHHHHHH-HSSSEEEEEEEEEEEEEETTCEEEEEEEEEEE
T ss_pred             CCCHHHHHHHHHHHHhh-chhCcEEEEEEEEEEEEeCCcEEEEEEEEEEE
Confidence            45678888888888765 33567889999999999888  6788888775


No 42 
>PF15418 DUF4625:  Domain of unknown function (DUF4625)
Probab=24.07  E-value=1.5e+02  Score=23.25  Aligned_cols=35  Identities=14%  Similarity=0.063  Sum_probs=28.8

Q ss_pred             eeecceeEEEEEEEEccCCccceEEEEEEeCCCCce
Q 027502          184 KVIEDYAKFELHLKLRRGSEEEKHWVEIIKNSEGKF  219 (222)
Q Consensus       184 QVVaG~~yy~l~l~~~~g~~~~~~~~~v~~~~~~~~  219 (222)
                      ..+.|.. +.+..+++...+-..|++++|.|-|||-
T Consensus        31 ~~~~G~~-ihfe~~i~d~~~i~si~VeIH~nfd~H~   65 (132)
T PF15418_consen   31 VATRGDD-IHFEADISDNSAIKSIKVEIHNNFDHHT   65 (132)
T ss_pred             EEecCCc-EEEEEEEEcccceeEEEEEEecCcCccc
Confidence            3678888 6788899988888889999988877774


No 43 
>PF01456 Mucin:  Mucin-like glycoprotein;  InterPro: IPR000458 This family of trypanosomal proteins resemble vertebrate mucins. The protein consists of three regions. The N and C terminii are conserved between all members of the family, whereas the central region is not well conserved and contains a large number of threonine residues which can be glycosylated []. Indirect evidence suggested that these genes might encode the core protein of parasite mucins, glycoproteins that were proposed to be involved in the interaction with, and invasion of, mammalian host cells.
Probab=24.01  E-value=40  Score=26.29  Aligned_cols=10  Identities=40%  Similarity=0.972  Sum_probs=5.9

Q ss_pred             HHHHHHHhhc
Q 027502           12 LVCGFIELGL   21 (222)
Q Consensus        12 ~~~~~~~~~~   21 (222)
                      |||+||+|+.
T Consensus         6 LLCalLvlaL   15 (143)
T PF01456_consen    6 LLCALLVLAL   15 (143)
T ss_pred             HHHHHHHHHH
Confidence            4566666555


No 44 
>smart00557 IG_FLMN Filamin-type immunoglobulin domains. These form a rod-like structure in the actin-binding cytoskeleton protein, filamin. The C-terminal repeats of filamin bind beta1-integrin (CD29).
Probab=23.76  E-value=1.2e+02  Score=21.79  Aligned_cols=28  Identities=29%  Similarity=0.428  Sum_probs=19.4

Q ss_pred             EEEEEEccCCccceEEEEEEeCCCCceec
Q 027502          193 ELHLKLRRGSEEEKHWVEIIKNSEGKFYL  221 (222)
Q Consensus       193 ~l~l~~~~g~~~~~~~~~v~~~~~~~~~~  221 (222)
                      .|.+.+...+. +...+.|.++.||+|.+
T Consensus        33 ~~~v~i~~p~g-~~~~~~v~d~~dGty~v   60 (93)
T smart00557       33 ELEVEVTGPSG-KKVPVEVKDNGDGTYTV   60 (93)
T ss_pred             cEEEEEECCCC-CeeEeEEEeCCCCEEEE
Confidence            35666665443 34688889999998865


No 45 
>PF05399 EVI2A:  Ectropic viral integration site 2A protein (EVI2A);  InterPro: IPR008608 This family contains several mammalian ectropic viral integration site 2A (EVI2A) proteins. The function of this protein is unknown although it is thought to be a membrane protein and may function as an oncogene in retrovirus induced myeloid tumours [, ].; GO: 0016021 integral to membrane
Probab=22.86  E-value=65  Score=27.73  Aligned_cols=17  Identities=24%  Similarity=0.524  Sum_probs=13.4

Q ss_pred             hhHHHHHHHHHHHHhhc
Q 027502            5 SVIVLSVLVCGFIELGL   21 (222)
Q Consensus         5 ~~~~~~~~~~~~~~~~~   21 (222)
                      .+|++|||+|.||-++-
T Consensus       135 IIIAVLfLICT~LfLST  151 (227)
T PF05399_consen  135 IIIAVLFLICTLLFLST  151 (227)
T ss_pred             HHHHHHHHHHHHHHHHH
Confidence            57888999998887663


No 46 
>PLN03207 stomagen; Provisional
Probab=22.55  E-value=92  Score=23.63  Aligned_cols=13  Identities=15%  Similarity=0.220  Sum_probs=7.1

Q ss_pred             hhHHHHHHHHHHH
Q 027502            5 SVIVLSVLVCGFI   17 (222)
Q Consensus         5 ~~~~~~~~~~~~~   17 (222)
                      ..+.|++|||.|+
T Consensus        11 ~~~~lffLl~~ll   23 (113)
T PLN03207         11 RCLTLFFLLFFLL   23 (113)
T ss_pred             hhHHHHHHHHHHH
Confidence            4555555555555


No 47 
>PF13956 Ibs_toxin:  Toxin Ibs, type I toxin-antitoxin system
Probab=22.37  E-value=40  Score=17.62  Aligned_cols=16  Identities=25%  Similarity=0.571  Sum_probs=6.9

Q ss_pred             CCchhhHHHHHHHHHH
Q 027502            1 MNRYSVIVLSVLVCGF   16 (222)
Q Consensus         1 ~~~~~~~~~~~~~~~~   16 (222)
                      |-|..++...+|+.+|
T Consensus         1 MMk~vIIlvvLLliSf   16 (19)
T PF13956_consen    1 MMKLVIILVVLLLISF   16 (19)
T ss_pred             CceehHHHHHHHhccc
Confidence            4444444344444444


No 48 
>MTH00261 ATP8 ATP synthase F0 subunit 8; Provisional
Probab=22.16  E-value=79  Score=21.50  Aligned_cols=18  Identities=22%  Similarity=0.325  Sum_probs=12.4

Q ss_pred             CchhhHHHHHHHHHHHHh
Q 027502            2 NRYSVIVLSVLVCGFIEL   19 (222)
Q Consensus         2 ~~~~~~~~~~~~~~~~~~   19 (222)
                      ++|+++.|+++++.++..
T Consensus        11 nhyfvllllf~iliilis   28 (68)
T MTH00261         11 NHYFVLLLLFFILIILIS   28 (68)
T ss_pred             HHHHHHHHHHHHHHHHHH
Confidence            578888877777665543


No 49 
>PF13721 SecD-TM1:  SecD export protein N-terminal TM region
Probab=20.50  E-value=89  Score=23.36  Aligned_cols=14  Identities=29%  Similarity=0.562  Sum_probs=7.7

Q ss_pred             CCchhhHHHHHHHH
Q 027502            1 MNRYSVIVLSVLVC   14 (222)
Q Consensus         1 ~~~~~~~~~~~~~~   14 (222)
                      |+||+.|=-+++++
T Consensus         1 mN~yp~WKyllil~   14 (101)
T PF13721_consen    1 MNRYPLWKYLLILV   14 (101)
T ss_pred             CCCcchHHHHHHHH
Confidence            78886554333333


No 50 
>PF12984 DUF3868:  Domain of unknown function, B. Theta Gene description (DUF3868);  InterPro: IPR024480 This domain of unknown function is found in a number of bacterial proteins. The function of the proteins is not known, but the Bacteroides thetaiotaomicron gene appears to be upregulated in the presence of host or other bacterial species compared to pure culture [, ].
Probab=20.25  E-value=62  Score=24.77  Aligned_cols=15  Identities=13%  Similarity=0.210  Sum_probs=8.9

Q ss_pred             CCchhhHHHHHHHHH
Q 027502            1 MNRYSVIVLSVLVCG   15 (222)
Q Consensus         1 ~~~~~~~~~~~~~~~   15 (222)
                      |.+|+++++++++|.
T Consensus         1 ~~~~~i~~~Ll~~~~   15 (115)
T PF12984_consen    1 KKIYFILFFLLLCSL   15 (115)
T ss_pred             CcEEEHHHHHHHHhh
Confidence            566777665555544


Done!