Query         047082
Match_columns 263
No_of_seqs    119 out of 1284
Neff          7.6 
Searched_HMMs 46136
Date          Fri Mar 29 07:26:30 2013
Command       hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/047082.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/047082hhsearch_cdd -cpu 12 -v 0 

 No Hit                             Prob E-value P-value  Score    SS Cols Query HMM  Template HMM
  1 PF01453 B_lectin:  D-mannose b 100.0   3E-29 6.4E-34  195.8   6.4   87    5-91      1-93  (114)
  2 PF00954 S_locus_glycop:  S-loc  99.8 4.2E-21 9.1E-26  148.8   9.6   99  120-218     1-108 (110)
  3 cd00028 B_lectin Bulb-type man  99.8   8E-21 1.7E-25  148.5   9.8   75    6-80     41-116 (116)
  4 smart00108 B_lectin Bulb-type   99.8 8.8E-20 1.9E-24  142.2   9.7   74    6-79     40-114 (114)
  5 smart00108 B_lectin Bulb-type   98.4 1.1E-06 2.4E-11   68.1   7.3   55   22-76     22-79  (114)
  6 cd00028 B_lectin Bulb-type man  98.4 5.6E-07 1.2E-11   70.1   5.4   56   23-78     23-82  (116)
  7 PF08276 PAN_2:  PAN-like domai  98.4 2.8E-07   6E-12   64.7   2.9   25  235-259    25-49  (66)
  8 PF01453 B_lectin:  D-mannose b  98.1 2.6E-05 5.6E-10   60.7   8.8   73    6-80     38-113 (114)
  9 cd01098 PAN_AP_plant Plant PAN  98.1   3E-06 6.4E-11   61.5   3.0   27  235-261    30-56  (84)
 10 cd00129 PAN_APPLE PAN/APPLE-li  97.8 2.1E-05 4.6E-10   57.3   3.2   25  236-260    24-51  (80)
 11 smart00473 PAN_AP divergent su  96.5  0.0029 6.3E-08   44.5   3.6   25  235-259    23-48  (78)
 12 cd01100 APPLE_Factor_XI_like S  96.3  0.0045 9.7E-08   43.9   3.3   26  235-260    23-48  (73)
 13 PF00024 PAN_1:  PAN domain Thi  77.0     2.8   6E-05   29.0   2.9   26  236-261    22-48  (79)
 14 smart00223 APPLE APPLE domain.  76.2     3.2   7E-05   29.9   3.1   32  231-262    16-47  (79)
 15 PF07354 Sp38:  Zona-pellucida-  75.9     4.2 9.2E-05   36.1   4.3   36    2-37      9-44  (271)
 16 PF01436 NHL:  NHL repeat;  Int  74.8     5.9 0.00013   22.4   3.4   21   23-43      6-26  (28)
 17 PF14295 PAN_4:  PAN domain; PD  73.1       4 8.6E-05   25.9   2.7   25  235-259    14-38  (51)
 18 cd05845 Ig2_L1-CAM_like Second  67.0      10 0.00023   28.3   4.1   33    3-35     31-63  (95)
 19 cd01099 PAN_AP_HGF Subfamily o  66.5     7.5 0.00016   27.8   3.2   25  236-260    24-50  (80)
 20 PF13360 PQQ_2:  PQQ-like domai  65.3      14  0.0003   31.1   5.1   73    4-76     11-102 (238)
 21 cd00053 EGF Epidermal growth f  62.4     7.9 0.00017   21.9   2.2   29  190-218     2-31  (36)
 22 TIGR03066 Gem_osc_para_1 Gemma  61.5      27 0.00058   27.0   5.5   55   18-72     32-104 (111)
 23 PRK11138 outer membrane biogen  60.9      35 0.00076   31.6   7.4   56   21-76    120-186 (394)
 24 PRK11138 outer membrane biogen  58.9      23 0.00051   32.8   5.9   19   58-76    342-361 (394)
 25 PF13360 PQQ_2:  PQQ-like domai  50.7      63  0.0014   26.9   6.8   73    4-76     54-148 (238)
 26 KOG0640 mRNA cleavage stimulat  45.7 1.1E+02  0.0024   28.2   7.6   71    5-75    247-333 (430)
 27 PLN00033 photosystem II stabil  44.7      73  0.0016   30.1   6.7   52   23-74    252-305 (398)
 28 PF07974 EGF_2:  EGF-like domai  43.7      19  0.0004   21.3   1.7   22  194-216     6-27  (32)
 29 TIGR03300 assembly_YfgL outer   43.1 1.4E+02   0.003   27.2   8.3   72    4-75     83-170 (377)
 30 PF09064 Tme5_EGF_like:  Thromb  41.5      16 0.00035   22.0   1.1   16  201-216    11-26  (34)
 31 PF08277 PAN_3:  PAN-like domai  41.3      36 0.00078   23.2   3.2   24  235-258    18-41  (71)
 32 KOG4649 PQQ (pyrrolo-quinoline  39.8      50  0.0011   29.7   4.4   43    6-48    169-217 (354)
 33 TIGR03300 assembly_YfgL outer   39.8      68  0.0015   29.3   5.7   17   58-74    327-344 (377)
 34 cd05852 Ig5_Contactin-1 Fifth   39.0      35 0.00076   23.6   2.8   33    3-36     13-45  (73)
 35 smart00605 CW CW domain.        39.0      37 0.00081   24.8   3.1   23  236-258    21-43  (94)
 36 PF07645 EGF_CA:  Calcium-bindi  35.7     9.4  0.0002   23.7  -0.5   28  190-217     5-34  (42)
 37 KOG0291 WD40-repeat-containing  35.3 5.3E+02   0.011   26.8  16.7  108    5-120   372-505 (893)
 38 smart00179 EGF_CA Calcium-bind  34.9      36 0.00079   19.7   2.1   28  190-217     5-33  (39)
 39 cd00216 PQQ_DH Dehydrogenases   34.0 1.4E+02   0.003   28.8   7.0   72    4-75     37-135 (488)
 40 PF13570 PQQ_3:  PQQ-like domai  31.2      40 0.00088   20.2   1.9    8   40-47      2-9   (40)
 41 PF05935 Arylsulfotrans:  Aryls  30.3      49  0.0011   31.9   3.2   52   29-81    127-186 (477)
 42 PF01683 EB:  EB module;  Inter  28.4      43 0.00093   21.5   1.7   27  189-218    21-47  (52)
 43 smart00564 PQQ beta-propeller   27.4 1.2E+02  0.0026   16.9   3.4   17   28-44     14-31  (33)
 44 PF06006 DUF905:  Bacterial pro  24.8      80  0.0017   22.2   2.5   17   63-79     35-51  (70)
 45 cd00054 EGF_CA Calcium-binding  24.5      70  0.0015   18.0   2.1   28  190-217     5-33  (38)
 46 cd00216 PQQ_DH Dehydrogenases   24.2 1.7E+02  0.0037   28.1   5.7   16   60-75    415-431 (488)
 47 KOG3881 Uncharacterized conser  23.8 2.9E+02  0.0064   26.1   6.7   61   22-83    218-280 (412)
 48 PF10636 hemP:  Hemin uptake pr  23.4      90   0.002   19.3   2.3   13   23-35     25-37  (38)
 49 PF02035 Coagulin:  Coagulin;    22.1      76  0.0016   25.2   2.2   34  172-205   101-142 (174)
 50 cd05764 Ig_2 Subgroup of the i  21.8 1.6E+02  0.0035   19.6   3.8   33    3-35     13-45  (74)
 51 PF05833 FbpA:  Fibronectin-bin  21.8      75  0.0016   30.2   2.7   39   53-91    115-159 (455)
 52 KOG3848 Extracellular protein   21.4 3.3E+02  0.0072   26.1   6.6   51   11-67    194-250 (516)
 53 TIGR03075 PQQ_enz_alc_DH PQQ-d  20.7 4.1E+02  0.0089   26.0   7.6   75    1-75     84-196 (527)
 54 COG1520 FOG: WD40-like repeat   20.5 4.3E+02  0.0093   24.1   7.4   42    5-46    130-180 (370)
 55 COG3236 Uncharacterized protei  20.3      55  0.0012   26.5   1.2   21   55-75    116-137 (162)

No 1  
>PF01453 B_lectin:  D-mannose binding lectin;  InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]:  Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein   This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity.  Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=99.95  E-value=3e-29  Score=195.75  Aligned_cols=87  Identities=47%  Similarity=0.679  Sum_probs=67.0

Q ss_pred             CCcEEEEcCCCCCCC---CCCEEEEecCCcEEEEcCCCcEEEcC-CCCCCC--eeEEEEeeCCCeeEEcCCCceEeeecc
Q 047082            5 KNSCSLNSIFASPVC---GKPTFSLGSDGNLVLAEADGTVVCQS-NTANKG--VVGFKLLPNGNMVLHDSKGNFIWQSFD   78 (263)
Q Consensus         5 ~~~~vWvANr~~Pv~---~~~~l~l~~~G~L~l~~~~g~~~Wst-~~~~~~--~~~a~Lld~GNlvl~~~~~~~~WqSFd   78 (263)
                      ++++||+|||+.||.   +..+|.|+.||+|+|++..++.+|++ .+.+.+  ...|+|+|+|||||++..+.+||||||
T Consensus         1 ~~tvvW~an~~~p~~~~s~~~~L~l~~dGnLvl~~~~~~~iWss~~t~~~~~~~~~~~L~~~GNlvl~d~~~~~lW~Sf~   80 (114)
T PF01453_consen    1 PRTVVWVANRNSPLTSSSGNYTLILQSDGNLVLYDSNGSVIWSSNNTSGRGNSGCYLVLQDDGNLVLYDSSGNVLWQSFD   80 (114)
T ss_dssp             ---------TTEEEEECETTEEEEEETTSEEEEEETTTEEEEE--S-TTSS-SSEEEEEETTSEEEEEETTSEEEEESTT
T ss_pred             CcccccccccccccccccccccceECCCCeEEEEcCCCCEEEEecccCCccccCeEEEEeCCCCEEEEeecceEEEeecC
Confidence            468999999999994   34899999999999999998999999 666544  689999999999999988999999999


Q ss_pred             CCCceeccCcccC
Q 047082           79 CPTDTLLVGQSLL   91 (263)
Q Consensus        79 ~PTDTlLpGq~l~   91 (263)
                      |||||+||||+|+
T Consensus        81 ~ptdt~L~~q~l~   93 (114)
T PF01453_consen   81 YPTDTLLPGQKLG   93 (114)
T ss_dssp             SSS-EEEEEET--
T ss_pred             CCccEEEeccCcc
Confidence            9999999999986


No 2  
>PF00954 S_locus_glycop:  S-locus glycoprotein family;  InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=99.85  E-value=4.2e-21  Score=148.79  Aligned_cols=99  Identities=15%  Similarity=0.192  Sum_probs=85.0

Q ss_pred             eecCCCCceeeecccCCCccceeEEEEEecCCCceEEEeecCCCceEEEEEeecCCEEEEEecCCC-CCCC--------C
Q 047082          120 LYDPSVQNLTFNSRPETDEAFAFKLTLDISDSGSDILARPKYNIRSSFLRLGMHGNLKIYTHYDKV-DSQP--------T  190 (263)
Q Consensus       120 ~~sg~w~~~~f~~~p~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~r~~Ld~dG~lr~y~~~~~~-~w~~--------C  190 (263)
                      |++|+|+|..|++.|+|.....+.+.|+.++.+.++.+.....+.++|++||.+|+|++|.|.+.. .|..        |
T Consensus         1 wrsG~WnG~~f~g~p~~~~~~~~~~~fv~~~~e~~~t~~~~~~s~~~r~~ld~~G~l~~~~w~~~~~~W~~~~~~p~d~C   80 (110)
T PF00954_consen    1 WRSGPWNGQRFSGIPEMSSNSLYNYSFVSNNEEVYYTYSLSNSSVLSRLVLDSDGQLQRYIWNESTQSWSVFWSAPKDQC   80 (110)
T ss_pred             CCccccCCeEECCcccccccceeEEEEEECCCeEEEEEecCCCceEEEEEEeeeeEEEEEEEecCCCcEEEEEEecccCC
Confidence            568999999999999987656677778777778888888777778999999999999999998654 4542        9


Q ss_pred             CCCCCCCCCcccCCCCCcCCCCCCCCcc
Q 047082          191 QLPERCSKLGVCDDNQCVACPTEKGLLG  218 (263)
Q Consensus       191 ~~~~~CG~~g~C~~~~~~~C~c~~g~~~  218 (263)
                      |+|++||+||+|+.+..+.|.|++||..
T Consensus        81 d~y~~CG~~g~C~~~~~~~C~Cl~GF~P  108 (110)
T PF00954_consen   81 DVYGFCGPNGICNSNNSPKCSCLPGFEP  108 (110)
T ss_pred             CCccccCCccEeCCCCCCceECCCCcCC
Confidence            9999999999999888889999999964


No 3  
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=99.84  E-value=8e-21  Score=148.53  Aligned_cols=75  Identities=45%  Similarity=0.678  Sum_probs=68.7

Q ss_pred             CcEEEEcCCCCCCCCCCEEEEecCCcEEEEcCCCcEEEcCCCCC-CCeeEEEEeeCCCeeEEcCCCceEeeeccCC
Q 047082            6 NSCSLNSIFASPVCGKPTFSLGSDGNLVLAEADGTVVCQSNTAN-KGVVGFKLLPNGNMVLHDSKGNFIWQSFDCP   80 (263)
Q Consensus         6 ~~~vWvANr~~Pv~~~~~l~l~~~G~L~l~~~~g~~~Wst~~~~-~~~~~a~Lld~GNlvl~~~~~~~~WqSFd~P   80 (263)
                      .++||+|||+.|....+.|.|+.+|+|+|.|.+|.++|++++.+ .+...|+|+|+|||||++.++++||||||||
T Consensus        41 ~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S~~~~~~~~~~~~L~ddGnlvl~~~~~~~~W~Sf~~P  116 (116)
T cd00028          41 RTVVWVANRDNPSGSSCTLTLQSDGNLVIYDGSGTVVWSSNTTRVNGNYVLVLLDDGNLVLYDSDGNFLWQSFDYP  116 (116)
T ss_pred             CeEEEECCCCCCCCCCEEEEEecCCCeEEEcCCCcEEEEecccCCCCceEEEEeCCCCEEEECCCCCEEEcCCCCC
Confidence            57999999999966678999999999999999999999999876 5567899999999999998899999999999


No 4  
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=99.81  E-value=8.8e-20  Score=142.16  Aligned_cols=74  Identities=46%  Similarity=0.685  Sum_probs=67.8

Q ss_pred             CcEEEEcCCCCCCCCCCEEEEecCCcEEEEcCCCcEEEcCCCC-CCCeeEEEEeeCCCeeEEcCCCceEeeeccC
Q 047082            6 NSCSLNSIFASPVCGKPTFSLGSDGNLVLAEADGTVVCQSNTA-NKGVVGFKLLPNGNMVLHDSKGNFIWQSFDC   79 (263)
Q Consensus         6 ~~~vWvANr~~Pv~~~~~l~l~~~G~L~l~~~~g~~~Wst~~~-~~~~~~a~Lld~GNlvl~~~~~~~~WqSFd~   79 (263)
                      .++||+|||+.|+.+++.|.|+++|+|+|.+.+|.++|++++. +.+...|+|+|+|||||++..+++|||||||
T Consensus        40 ~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S~t~~~~~~~~~~L~ddGnlvl~~~~~~~~W~Sf~~  114 (114)
T smart00108       40 RTVVWVANRDNPVSDSCTLTLQSDGNLVLYDGDGRVVWSSNTTGANGNYVLVLLDDGNLVIYDSDGNFLWQSFDY  114 (114)
T ss_pred             CcEEEECCCCCCCCCCEEEEEeCCCCEEEEeCCCCEEEEecccCCCCceEEEEeCCCCEEEECCCCCEEeCCCCC
Confidence            5799999999999877899999999999999999999999986 4556789999999999999988999999997


No 5  
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=98.39  E-value=1.1e-06  Score=68.12  Aligned_cols=55  Identities=35%  Similarity=0.528  Sum_probs=44.9

Q ss_pred             CEEEEecCCcEEEEcCC-CcEEEcCCCCCC--CeeEEEEeeCCCeeEEcCCCceEeee
Q 047082           22 PTFSLGSDGNLVLAEAD-GTVVCQSNTANK--GVVGFKLLPNGNMVLHDSKGNFIWQS   76 (263)
Q Consensus        22 ~~l~l~~~G~L~l~~~~-g~~~Wst~~~~~--~~~~a~Lld~GNlvl~~~~~~~~WqS   76 (263)
                      -.+.++.+|+||+.... +.++|++++...  ....+.|.++|||||++.++.++|+|
T Consensus        22 ~~~~~q~dgnlV~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S   79 (114)
T smart00108       22 FTLIMQNDYNLILYKSSSRTVVWVANRDNPVSDSCTLTLQSDGNLVLYDGDGRVVWSS   79 (114)
T ss_pred             cccCCCCCEEEEEEECCCCcEEEECCCCCCCCCCEEEEEeCCCCEEEEeCCCCEEEEe
Confidence            35667789999999765 479999998532  23678999999999999888999997


No 6  
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=98.38  E-value=5.6e-07  Score=70.06  Aligned_cols=56  Identities=32%  Similarity=0.503  Sum_probs=45.1

Q ss_pred             EEEEec-CCcEEEEcCC-CcEEEcCCCCC--CCeeEEEEeeCCCeeEEcCCCceEeeecc
Q 047082           23 TFSLGS-DGNLVLAEAD-GTVVCQSNTAN--KGVVGFKLLPNGNMVLHDSKGNFIWQSFD   78 (263)
Q Consensus        23 ~l~l~~-~G~L~l~~~~-g~~~Wst~~~~--~~~~~a~Lld~GNlvl~~~~~~~~WqSFd   78 (263)
                      .+.++. +|+|++++.. +.++|++++..  .....+.|.++|||||.+.++.++|+|--
T Consensus        23 ~~~~q~~dgnlv~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S~~   82 (116)
T cd00028          23 KLIMQSRDYNLILYKGSSRTVVWVANRDNPSGSSCTLTLQSDGNLVIYDGSGTVVWSSNT   82 (116)
T ss_pred             cCCCCCCeEEEEEEeCCCCeEEEECCCCCCCCCCEEEEEecCCCeEEEcCCCcEEEEecc
Confidence            455666 9999999764 47999999854  24567899999999999988899999643


No 7  
>PF08276 PAN_2:  PAN-like domain;  InterPro: IPR013227 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs
Probab=98.36  E-value=2.8e-07  Score=64.74  Aligned_cols=25  Identities=32%  Similarity=1.071  Sum_probs=23.1

Q ss_pred             CCCCHHHHHHHhhccCCeEEEEecc
Q 047082          235 TAIKVEDCGRKCTSDCKCSGYFYHQ  259 (263)
Q Consensus       235 ~~~s~~~C~~~Cl~nCsC~a~~y~~  259 (263)
                      ...++++|+++||+||||+||+|.+
T Consensus        25 ~~~s~~~C~~~Cl~nCsC~Ayay~~   49 (66)
T PF08276_consen   25 SSVSLEECEKACLSNCSCTAYAYSN   49 (66)
T ss_pred             cCCCHHHHHhhcCCCCCEeeEEeec
Confidence            4589999999999999999999985


No 8  
>PF01453 B_lectin:  D-mannose binding lectin;  InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]:  Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein   This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity.  Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=98.08  E-value=2.6e-05  Score=60.65  Aligned_cols=73  Identities=25%  Similarity=0.202  Sum_probs=50.1

Q ss_pred             CcEEEEc-CCCCCCCCCCEEEEecCCcEEEEcCCCcEEEcCCCCCCCeeEEEEee--CCCeeEEcCCCceEeeeccCC
Q 047082            6 NSCSLNS-IFASPVCGKPTFSLGSDGNLVLAEADGTVVCQSNTANKGVVGFKLLP--NGNMVLHDSKGNFIWQSFDCP   80 (263)
Q Consensus         6 ~~~vWvA-Nr~~Pv~~~~~l~l~~~G~L~l~~~~g~~~Wst~~~~~~~~~a~Lld--~GNlvl~~~~~~~~WqSFd~P   80 (263)
                      .++||.. +........+.+.|..+|||||.|..+.++|++... ..-+.+.+++  .||++ ......++|.|=..|
T Consensus        38 ~~~iWss~~t~~~~~~~~~~~L~~~GNlvl~d~~~~~lW~Sf~~-ptdt~L~~q~l~~~~~~-~~~~~~~sw~s~~dp  113 (114)
T PF01453_consen   38 GSVIWSSNNTSGRGNSGCYLVLQDDGNLVLYDSSGNVLWQSFDY-PTDTLLPGQKLGDGNVT-GKNDSLTSWSSNTDP  113 (114)
T ss_dssp             TEEEEE--S-TTSS-SSEEEEEETTSEEEEEETTSEEEEESTTS-SS-EEEEEET--TSEEE-EESTSSEEEESS---
T ss_pred             CCEEEEecccCCccccCeEEEEeCCCCEEEEeecceEEEeecCC-CccEEEeccCcccCCCc-cccceEEeECCCCCC
Confidence            4679999 434333346889999999999999999999999542 2335566677  88888 654567899876665


No 9  
>cd01098 PAN_AP_plant Plant PAN/APPLE-like domain; present in plant S-receptor protein kinases and secreted glycoproteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions. S-receptor protein kinases and S-locus glycoproteins are involved in sporophytic self-incompatibility response in Brassica, one of probably many molecular mechanisms, by which hermaphrodite flowering plants avoid self-fertilization.
Probab=98.06  E-value=3e-06  Score=61.50  Aligned_cols=27  Identities=41%  Similarity=1.026  Sum_probs=23.9

Q ss_pred             CCCCHHHHHHHhhccCCeEEEEeccCC
Q 047082          235 TAIKVEDCGRKCTSDCKCSGYFYHQET  261 (263)
Q Consensus       235 ~~~s~~~C~~~Cl~nCsC~a~~y~~~~  261 (263)
                      ...++++|+++||+||+|+||+|.+++
T Consensus        30 ~~~s~~~C~~~Cl~nCsC~a~~~~~~~   56 (84)
T cd01098          30 TAISLEECREACLSNCSCTAYAYNNGS   56 (84)
T ss_pred             ccCCHHHHHHHHhcCCCcceeeecCCC
Confidence            457999999999999999999998643


No 10 
>cd00129 PAN_APPLE PAN/APPLE-like domain; present in N-terminal (N) domains of plasminogen/ hepatocyte growth factor proteins,  plasma prekallikrein/coagulation factor XI and microneme antigen proteins, plant receptor-like protein kinases, and various nematode and leech anti-platelet proteins. Common structural features include two disulfide bonds that link the alpha-helix to the central region of the protein. PAN domains have significant functional versatility, fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=97.77  E-value=2.1e-05  Score=57.32  Aligned_cols=25  Identities=16%  Similarity=0.712  Sum_probs=22.4

Q ss_pred             CCCHHHHHHHhhc---cCCeEEEEeccC
Q 047082          236 AIKVEDCGRKCTS---DCKCSGYFYHQE  260 (263)
Q Consensus       236 ~~s~~~C~~~Cl~---nCsC~a~~y~~~  260 (263)
                      ..++++|+++|++   ||||.||+|.+.
T Consensus        24 ~~s~~eC~~~Cl~~~~nCsC~Aya~~~~   51 (80)
T cd00129          24 ANTADECANRCEKNGLPFSCKAFVFAKA   51 (80)
T ss_pred             ccCHHHHHHHHhcCCCCCCceeeeccCC
Confidence            3789999999999   999999999653


No 11 
>smart00473 PAN_AP divergent subfamily of APPLE domains. Apple-like domains present in Plasminogen, C. elegans hypothetical ORFs and the extracellular portion of plant receptor-like protein kinases. Predicted to possess protein- and/or carbohydrate-binding functions.
Probab=96.55  E-value=0.0029  Score=44.50  Aligned_cols=25  Identities=28%  Similarity=0.975  Sum_probs=22.7

Q ss_pred             CCCCHHHHHHHhhc-cCCeEEEEecc
Q 047082          235 TAIKVEDCGRKCTS-DCKCSGYFYHQ  259 (263)
Q Consensus       235 ~~~s~~~C~~~Cl~-nCsC~a~~y~~  259 (263)
                      ...++++|++.|++ +|+|.||.|..
T Consensus        23 ~~~s~~~C~~~C~~~~~~C~s~~y~~   48 (78)
T smart00473       23 SVASLEECASKCLNSNCSCRSFTYNN   48 (78)
T ss_pred             cCCCHHHHHHHhCCCCCceEEEEEcC
Confidence            35799999999999 99999999975


No 12 
>cd01100 APPLE_Factor_XI_like Subfamily of PAN/APPLE-like domains; present in plasma prekallikrein/coagulation factor XI, microneme antigen proteins, and a few prokaryotic proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=96.28  E-value=0.0045  Score=43.94  Aligned_cols=26  Identities=31%  Similarity=0.657  Sum_probs=23.0

Q ss_pred             CCCCHHHHHHHhhccCCeEEEEeccC
Q 047082          235 TAIKVEDCGRKCTSDCKCSGYFYHQE  260 (263)
Q Consensus       235 ~~~s~~~C~~~Cl~nCsC~a~~y~~~  260 (263)
                      ...+.++|++.|+.+|+|.||.|..+
T Consensus        23 ~~~s~~~Cq~~C~~~~~C~afT~~~~   48 (73)
T cd01100          23 FASSAEQCQAACTADPGCLAFTYNTK   48 (73)
T ss_pred             ecCCHHHHHHHcCCCCCceEEEEECC
Confidence            34689999999999999999999754


No 13 
>PF00024 PAN_1:  PAN domain This Prosite entry concerns apple domains, a subset of PAN domains;  InterPro: IPR003014 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs It has been shown that, the N-terminal N domains of members of the plasminogen/hepatocyte growth factor family, the apple domains of the plasma prekallikrein/coagulation factor XI family, and domains of various nematode proteins belong to the same module superfamily, the PAN module []. PAN contains a conserved core of three disulphide bridges. In some members of the family there is an additional fourth disulphide bridge that links the N and C termini of the domain.; PDB: 1GP9_C 2QJ2_B 1GMO_H 1NK1_B 3MKP_B 1BHT_B 3HN4_A 1GMN_A 3HMS_A 3HMT_B ....
Probab=76.96  E-value=2.8  Score=29.04  Aligned_cols=26  Identities=19%  Similarity=0.695  Sum_probs=23.0

Q ss_pred             CCCHHHHHHHhhccCC-eEEEEeccCC
Q 047082          236 AIKVEDCGRKCTSDCK-CSGYFYHQET  261 (263)
Q Consensus       236 ~~s~~~C~~~Cl~nCs-C~a~~y~~~~  261 (263)
                      ..++++|.+.|+.+=. |.+|.|...+
T Consensus        22 v~s~~~C~~~C~~~~~~C~s~~y~~~~   48 (79)
T PF00024_consen   22 VPSLEECAQLCLNEPRRCKSFNYDPSS   48 (79)
T ss_dssp             ESSHHHHHHHHHHSTT-ESEEEEETTT
T ss_pred             CCCHHHHHhhcCcCcccCCeEEEECCC
Confidence            4599999999999999 9999997653


No 14 
>smart00223 APPLE APPLE domain. Four-fold repeat in plasma kallikrein and coagulation factor XI. Factor XI apple 3 mediates binding to platelets. Factor XI apple 1 binds high-molecular-mass kininogen. Apple 4 in factor XI mediates dimer formation and binds to factor XIIa. Mutations in apple 4 cause factor XI deficiency, an inherited bleeding disorder.
Probab=76.24  E-value=3.2  Score=29.92  Aligned_cols=32  Identities=16%  Similarity=0.396  Sum_probs=26.0

Q ss_pred             cccCCCCCHHHHHHHhhccCCeEEEEeccCCC
Q 047082          231 YTSGTAIKVEDCGRKCTSDCKCSGYFYHQETS  262 (263)
Q Consensus       231 ~~~~~~~s~~~C~~~Cl~nCsC~a~~y~~~~~  262 (263)
                      .......+.++|++.|..+=.|.+|.|...+.
T Consensus        16 l~~~~~~~~~~Cq~~Ct~~~~C~~FTf~~~~~   47 (79)
T smart00223       16 INTVYVPSAQVCQKRCTSHPRCLFFTFSTNEP   47 (79)
T ss_pred             eeeeecCCHHHHHHhhcCCCCccEEEeeCCCC
Confidence            33344579999999999999999999976654


No 15 
>PF07354 Sp38:  Zona-pellucida-binding protein (Sp38);  InterPro: IPR010857 This family contains a number of zona-pellucida-binding proteins that seem to be restricted to mammals. These are sperm proteins that bind to the 90 kDa family of zona pellucida glycoproteins in a calcium-dependent manner []. These represent some of the specific molecules that mediate the first steps of gamete interaction, allowing fertilisation to occur [].; GO: 0007339 binding of sperm to zona pellucida, 0005576 extracellular region
Probab=75.89  E-value=4.2  Score=36.06  Aligned_cols=36  Identities=14%  Similarity=0.243  Sum_probs=33.0

Q ss_pred             CCCCCcEEEEcCCCCCCCCCCEEEEecCCcEEEEcC
Q 047082            2 EYPKNSCSLNSIFASPVCGKPTFSLGSDGNLVLAEA   37 (263)
Q Consensus         2 ~~~~~~~vWvANr~~Pv~~~~~l~l~~~G~L~l~~~   37 (263)
                      |+...+..|+--.+.++++++.+.|++.|.|++.+-
T Consensus         9 E~iDP~y~W~GP~g~~l~gn~~~nIT~TG~L~~~~F   44 (271)
T PF07354_consen    9 ELIDPTYLWTGPNGKPLSGNSYVNITETGKLMFKNF   44 (271)
T ss_pred             ccCCCceEEECCCCcccCCCCeEEEccCceEEeecc
Confidence            677889999999999999999999999999999764


No 16 
>PF01436 NHL:  NHL repeat;  InterPro: IPR001258 The NHL repeat, named after NCL-1, HT2A and Lin-41, is found largely in a large number of eukaryotic and prokaryotic proteins. For example, the repeat is found in a variety of enzymes of the copper type II, ascorbate-dependent monooxygenase family which catalyse the C terminus alpha-amidation of biological peptides []. In many it occurs in tandem arrays, for example in the ringfinger beta-box, coiled-coil (RBCC) eukaryotic growth regulators []. The 'Brain Tumor' protein (Brat) is one such growth regulator that contains a 6-bladed NHL-repeat beta-propeller [, ].  The NHL repeats are also found in serine/threonine protein kinase (STPK) in diverse range of pathogenic bacteria. These STPK are transmembrane receptors with a intracellular N-terminal kinase domain and extracellular C-terminal sensor domain. In the STPK, PknD, from Mycobacterium tuberculosis, the sensor domain forms a rigid, six-bladed b-propeller composed of NHL repeats with a flexible tether to the transmembrane domain.; GO: 0005515 protein binding; PDB: 3FVZ_A 3FW0_A 1RWL_A 1RWI_A 1Q7F_A.
Probab=74.81  E-value=5.9  Score=22.36  Aligned_cols=21  Identities=29%  Similarity=0.449  Sum_probs=16.6

Q ss_pred             EEEEecCCcEEEEcCCCcEEE
Q 047082           23 TFSLGSDGNLVLAEADGTVVC   43 (263)
Q Consensus        23 ~l~l~~~G~L~l~~~~g~~~W   43 (263)
                      -+.++.+|+|++.|..+.-||
T Consensus         6 gvav~~~g~i~VaD~~n~rV~   26 (28)
T PF01436_consen    6 GVAVDSDGNIYVADSGNHRVQ   26 (28)
T ss_dssp             EEEEETTSEEEEEECCCTEEE
T ss_pred             EEEEeCCCCEEEEECCCCEEE
Confidence            477888899999988776665


No 17 
>PF14295 PAN_4:  PAN domain; PDB: 2YIL_E 2YIP_C 2YIO_A.
Probab=73.12  E-value=4  Score=25.88  Aligned_cols=25  Identities=28%  Similarity=0.695  Sum_probs=17.9

Q ss_pred             CCCCHHHHHHHhhccCCeEEEEecc
Q 047082          235 TAIKVEDCGRKCTSDCKCSGYFYHQ  259 (263)
Q Consensus       235 ~~~s~~~C~~~Cl~nCsC~a~~y~~  259 (263)
                      ...+.++|.++|..+=.|.+|.|..
T Consensus        14 ~~~s~~~C~~~C~~~~~C~~~~~~~   38 (51)
T PF14295_consen   14 TASSPEECQAACAADPGCQAFTFNP   38 (51)
T ss_dssp             ----HHHHHHHHHTSTT--EEEEET
T ss_pred             cCCCHHHHHHHccCCCCCCEEEEEC
Confidence            4568999999999999999999976


No 18 
>cd05845 Ig2_L1-CAM_like Second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM) and similar proteins. Ig2_L1-CAM_like: domain similar to the second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM). L1 belongs to the L1 subfamily of cell adhesion molecules (CAMs) and is comprised of an extracellular region having six Ig-like domains, five fibronectin type III domains, a transmembrane region and an intracellular domain. L1 is primarily expressed in the nervous system and is involved in its development and function. L1 is associated with an X-linked recessive disorder, X-linked hydrocephalus, MASA syndrome, or spastic paraplegia type 1, that involves abnormalities of axonal growth.
Probab=66.98  E-value=10  Score=28.25  Aligned_cols=33  Identities=18%  Similarity=0.100  Sum_probs=22.3

Q ss_pred             CCCCcEEEEcCCCCCCCCCCEEEEecCCcEEEE
Q 047082            3 YPKNSCSLNSIFASPVCGKPTFSLGSDGNLVLA   35 (263)
Q Consensus         3 ~~~~~~vWvANr~~Pv~~~~~l~l~~~G~L~l~   35 (263)
                      +|+.++.|+-+....+.....+.++.+|+|.+.
T Consensus        31 ~P~P~i~W~~~~~~~i~~~~Ri~~~~~GnL~fs   63 (95)
T cd05845          31 AVPLRIYWMNSDLLHITQDERVSMGQNGNLYFA   63 (95)
T ss_pred             CCCCEEEEECCCCccccccccEEECCCceEEEE
Confidence            567788888544444554567777777888774


No 19 
>cd01099 PAN_AP_HGF Subfamily of PAN/APPLE-like domains; present in N-terminal (N) domains of plasminogen/hepatocyte growth factor proteins, and various proteins found in Bilateria, such as leech anti-platelet proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=66.50  E-value=7.5  Score=27.78  Aligned_cols=25  Identities=28%  Similarity=0.761  Sum_probs=21.9

Q ss_pred             CCCHHHHHHHhhc--cCCeEEEEeccC
Q 047082          236 AIKVEDCGRKCTS--DCKCSGYFYHQE  260 (263)
Q Consensus       236 ~~s~~~C~~~Cl~--nCsC~a~~y~~~  260 (263)
                      ..++++|.++|++  +=.|.+|.|...
T Consensus        24 ~~s~~~C~~~C~~~~~f~CrSf~y~~~   50 (80)
T cd01099          24 VASLEECLRKCLEETEFTCRSFNYNYK   50 (80)
T ss_pred             cCCHHHHHHHhCCCCCceEeEEEEEcC
Confidence            4799999999999  899999988654


No 20 
>PF13360 PQQ_2:  PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=65.30  E-value=14  Score=31.07  Aligned_cols=73  Identities=18%  Similarity=0.236  Sum_probs=43.9

Q ss_pred             CCCcEEEEcCC----CCCC----CCCCEEEE-ecCCcEEEEcC-CCcEEEcCCCCCCCe-------eEEEEe-eCCCeeE
Q 047082            4 PKNSCSLNSIF----ASPV----CGKPTFSL-GSDGNLVLAEA-DGTVVCQSNTANKGV-------VGFKLL-PNGNMVL   65 (263)
Q Consensus         4 ~~~~~vWvANr----~~Pv----~~~~~l~l-~~~G~L~l~~~-~g~~~Wst~~~~~~~-------~~a~Ll-d~GNlvl   65 (263)
                      .....+|..+-    ..++    .....|.+ +.+|.|+..|. .|..+|+........       ..+.+. .+|.|+.
T Consensus        11 ~tG~~~W~~~~~~~~~~~~~~~~~~~~~v~~~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~   90 (238)
T PF13360_consen   11 RTGKELWSYDLGPGIGGPVATAVPDGGRVYVASGDGNLYALDAKTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYA   90 (238)
T ss_dssp             TTTEEEEEEECSSSCSSEEETEEEETTEEEEEETTSEEEEEETTTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEE
T ss_pred             CCCCEEEEEECCCCCCCccceEEEeCCEEEEEcCCCEEEEEECCCCCEEEEeeccccccceeeecccccccccceeeeEe
Confidence            45678888753    2222    12343444 48899999996 899999987632210       111222 2344666


Q ss_pred             Ec-CCCceEeee
Q 047082           66 HD-SKGNFIWQS   76 (263)
Q Consensus        66 ~~-~~~~~~WqS   76 (263)
                      .| .+++++|+.
T Consensus        91 ~d~~tG~~~W~~  102 (238)
T PF13360_consen   91 LDAKTGKVLWSI  102 (238)
T ss_dssp             EETTTSCEEEEE
T ss_pred             cccCCcceeeee
Confidence            66 578999995


No 21 
>cd00053 EGF Epidermal growth factor domain, found in epidermal growth factor (EGF) presents in a large number of proteins, mostly animal; the list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied; the functional significance of EGF-like domains in what appear to be unrelated proteins is not yet clear; a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase); the domain includes six cysteine residues which have been shown to be involved in disulfide bonds; the main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet; Subdomains between the conserved cysteines vary in length; the region between the 5th and 6th cysteine contains two conserved glycines of which at  least  one  is  present  in  most EGF-like domains; a subset of these bind calcium.
Probab=62.40  E-value=7.9  Score=21.94  Aligned_cols=29  Identities=24%  Similarity=0.505  Sum_probs=20.6

Q ss_pred             CCCCCCCCCCcccCCC-CCcCCCCCCCCcc
Q 047082          190 TQLPERCSKLGVCDDN-QCVACPTEKGLLG  218 (263)
Q Consensus       190 C~~~~~CG~~g~C~~~-~~~~C~c~~g~~~  218 (263)
                      |.....|...++|... ....|.|+.||..
T Consensus         2 C~~~~~C~~~~~C~~~~~~~~C~C~~g~~g   31 (36)
T cd00053           2 CAASNPCSNGGTCVNTPGSYRCVCPPGYTG   31 (36)
T ss_pred             CCCCCCCCCCCEEecCCCCeEeECCCCCcc
Confidence            4435678888999643 4568999998854


No 22 
>TIGR03066 Gem_osc_para_1 Gemmata obscuriglobus paralogous family TIGR03066. This model represents an uncharacterized paralogous family in Gemmata obscuriglobus UQM 2246, a member of the Planctomycetes. This family shows sequence similarity to TIGR03067, which is also found in Gemmata obscuriglobus as well as in a few other species.
Probab=61.49  E-value=27  Score=26.97  Aligned_cols=55  Identities=18%  Similarity=0.272  Sum_probs=32.8

Q ss_pred             CCCCCEEEEecCCcEEEEcCCCcE------EEcCC---------CCCC---CeeEEEEeeCCCeeEEcCCCce
Q 047082           18 VCGKPTFSLGSDGNLVLAEADGTV------VCQSN---------TANK---GVVGFKLLPNGNMVLHDSKGNF   72 (263)
Q Consensus        18 v~~~~~l~l~~~G~L~l~~~~g~~------~Wst~---------~~~~---~~~~a~Lld~GNlvl~~~~~~~   72 (263)
                      +.....|.|..+|.|+|+.++++-      -|+-.         ..+.   .-..-.-+++|.|||.|++++.
T Consensus        32 ~~~~~~leF~~dGKL~v~~gnng~~~~~~Gty~L~G~kLtL~~~p~g~t~k~~Vtv~~l~~~~Lvl~d~dg~~  104 (111)
T TIGR03066        32 TKDDVVIEFAKDGKLVVTIGEKGKEVKADGTYKLDGNKLTLTLKAGGKEKKETLTVKKLTDDELVGKDPDGKK  104 (111)
T ss_pred             eCCceEEEEcCCCeEEEecCCCCcEeccCceEEEECCEEEEEEcCCCccccceEEEEEecCCeEEEEcCCCCE
Confidence            334678999999999998775442      12211         0011   1011123688899999887653


No 23 
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=60.93  E-value=35  Score=31.64  Aligned_cols=56  Identities=23%  Similarity=0.370  Sum_probs=34.3

Q ss_pred             CCEEEEe-cCCcEEEEcC-CCcEEEcCCCCCC----Ce----eEEEEeeCCCeeEEcC-CCceEeee
Q 047082           21 KPTFSLG-SDGNLVLAEA-DGTVVCQSNTANK----GV----VGFKLLPNGNMVLHDS-KGNFIWQS   76 (263)
Q Consensus        21 ~~~l~l~-~~G~L~l~~~-~g~~~Wst~~~~~----~~----~~a~Lld~GNlvl~~~-~~~~~WqS   76 (263)
                      ...|.+. .+|.|+-+|. +|..+|+......    ++    ....-..+|.|+-.|. +++.+|+-
T Consensus       120 ~~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~ssP~v~~~~v~v~~~~g~l~ald~~tG~~~W~~  186 (394)
T PRK11138        120 GGKVYIGSEKGQVYALNAEDGEVAWQTKVAGEALSRPVVSDGLVLVHTSNGMLQALNESDGAVKWTV  186 (394)
T ss_pred             CCEEEEEcCCCEEEEEECCCCCCcccccCCCceecCCEEECCEEEEECCCCEEEEEEccCCCEeeee
Confidence            3444444 6688887775 6899999876431    11    1112234566777775 68899964


No 24 
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=58.89  E-value=23  Score=32.80  Aligned_cols=19  Identities=21%  Similarity=0.338  Sum_probs=10.1

Q ss_pred             eeCCCeeEEcC-CCceEeee
Q 047082           58 LPNGNMVLHDS-KGNFIWQS   76 (263)
Q Consensus        58 ld~GNlvl~~~-~~~~~WqS   76 (263)
                      -++|.|...|. +++++|+-
T Consensus       342 ~~~G~l~~ld~~tG~~~~~~  361 (394)
T PRK11138        342 DSEGYLHWINREDGRFVAQQ  361 (394)
T ss_pred             eCCCEEEEEECCCCCEEEEE
Confidence            34555555553 45666653


No 25 
>PF13360 PQQ_2:  PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=50.66  E-value=63  Score=26.90  Aligned_cols=73  Identities=18%  Similarity=0.321  Sum_probs=44.5

Q ss_pred             CCCcEEEEcCCCCCCC-----CCC-EEEEecCCcEEEEc-CCCcEEEcC-CCC----C---CCe-e----EEEE-eeCCC
Q 047082            4 PKNSCSLNSIFASPVC-----GKP-TFSLGSDGNLVLAE-ADGTVVCQS-NTA----N---KGV-V----GFKL-LPNGN   62 (263)
Q Consensus         4 ~~~~~vWvANr~~Pv~-----~~~-~l~l~~~G~L~l~~-~~g~~~Wst-~~~----~---~~~-~----~a~L-ld~GN   62 (263)
                      ..-+.+|....+.++.     ... .+..+.+|.|...| .+|.++|.. ...    .   ... +    .+.+ ..+|.
T Consensus        54 ~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~g~  133 (238)
T PF13360_consen   54 KTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYALDAKTGKVLWSIYLTSSPPAGVRSSSSPAVDGDRLYVGTSSGK  133 (238)
T ss_dssp             TTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEEEETTTSCEEEEEEE-SSCTCSTB--SEEEEETTEEEEEETCSE
T ss_pred             CCCCEEEEeeccccccceeeecccccccccceeeeEecccCCcceeeeeccccccccccccccCceEecCEEEEEeccCc
Confidence            4567788877655532     233 44455678888888 789999994 321    1   010 0    1222 33788


Q ss_pred             eeEEcC-CCceEeee
Q 047082           63 MVLHDS-KGNFIWQS   76 (263)
Q Consensus        63 lvl~~~-~~~~~WqS   76 (263)
                      |+..|. +++.+|+-
T Consensus       134 l~~~d~~tG~~~w~~  148 (238)
T PF13360_consen  134 LVALDPKTGKLLWKY  148 (238)
T ss_dssp             EEEEETTTTEEEEEE
T ss_pred             EEEEecCCCcEEEEe
Confidence            888884 68899964


No 26 
>KOG0640 consensus mRNA cleavage stimulating factor complex; subunit 1 [RNA processing and modification]
Probab=45.73  E-value=1.1e+02  Score=28.19  Aligned_cols=71  Identities=18%  Similarity=0.350  Sum_probs=49.7

Q ss_pred             CCcEEEEcCCCCCCCCC-CEEEEecCCcEEEEcC-CCc-EEEcCCCC-----------CCCeeEEEEeeCCCeeEEcCCC
Q 047082            5 KNSCSLNSIFASPVCGK-PTFSLGSDGNLVLAEA-DGT-VVCQSNTA-----------NKGVVGFKLLPNGNMVLHDSKG   70 (263)
Q Consensus         5 ~~~~vWvANr~~Pv~~~-~~l~l~~~G~L~l~~~-~g~-~~Wst~~~-----------~~~~~~a~Lld~GNlvl~~~~~   70 (263)
                      ..-+-=.||.+..+++. ..+.-++.|+|+++.+ +|. -+|.-...           +..+.+|++-.+|.++|....+
T Consensus       247 T~QcfvsanPd~qht~ai~~V~Ys~t~~lYvTaSkDG~IklwDGVS~rCv~t~~~AH~gsevcSa~Ftkn~kyiLsSG~D  326 (430)
T KOG0640|consen  247 TYQCFVSANPDDQHTGAITQVRYSSTGSLYVTASKDGAIKLWDGVSNRCVRTIGNAHGGSEVCSAVFTKNGKYILSSGKD  326 (430)
T ss_pred             ceeEeeecCcccccccceeEEEecCCccEEEEeccCCcEEeeccccHHHHHHHHhhcCCceeeeEEEccCCeEEeecCCc
Confidence            33445568888777775 6788999999999855 344 47874321           2336789999999999986533


Q ss_pred             c--eEee
Q 047082           71 N--FIWQ   75 (263)
Q Consensus        71 ~--~~Wq   75 (263)
                      .  -||+
T Consensus       327 S~vkLWE  333 (430)
T KOG0640|consen  327 STVKLWE  333 (430)
T ss_pred             ceeeeee
Confidence            3  4786


No 27 
>PLN00033 photosystem II stability/assembly factor; Provisional
Probab=44.70  E-value=73  Score=30.14  Aligned_cols=52  Identities=17%  Similarity=0.248  Sum_probs=31.7

Q ss_pred             EEEEecCCcEEEEcCCCcEEEcCCCCC--CCeeEEEEeeCCCeeEEcCCCceEe
Q 047082           23 TFSLGSDGNLVLAEADGTVVCQSNTAN--KGVVGFKLLPNGNMVLHDSKGNFIW   74 (263)
Q Consensus        23 ~l~l~~~G~L~l~~~~g~~~Wst~~~~--~~~~~a~Lld~GNlvl~~~~~~~~W   74 (263)
                      .+.+...|++++.+.+|...|......  ...+.+...++|.|+|....+.++|
T Consensus       252 ~~~vg~~G~~~~s~d~G~~~W~~~~~~~~~~l~~v~~~~dg~l~l~g~~G~l~~  305 (398)
T PLN00033        252 YVAVSSRGNFYLTWEPGQPYWQPHNRASARRIQNMGWRADGGLWLLTRGGGLYV  305 (398)
T ss_pred             EEEEECCccEEEecCCCCcceEEecCCCccceeeeeEcCCCCEEEEeCCceEEE
Confidence            444445556555555566667744322  2334556688999999887666555


No 28 
>PF07974 EGF_2:  EGF-like domain;  InterPro: IPR013111 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length. This entry contains EGF domains found in a variety of extracellular and membrane proteins
Probab=43.71  E-value=19  Score=21.25  Aligned_cols=22  Identities=32%  Similarity=0.761  Sum_probs=17.3

Q ss_pred             CCCCCCcccCCCCCcCCCCCCCC
Q 047082          194 ERCSKLGVCDDNQCVACPTEKGL  216 (263)
Q Consensus       194 ~~CG~~g~C~~~~~~~C~c~~g~  216 (263)
                      .+|...|+|+.. ...|.|.+||
T Consensus         6 ~~C~~~G~C~~~-~g~C~C~~g~   27 (32)
T PF07974_consen    6 NICSGHGTCVSP-CGRCVCDSGY   27 (32)
T ss_pred             CccCCCCEEeCC-CCEEECCCCC
Confidence            479999999765 3379999886


No 29 
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=43.14  E-value=1.4e+02  Score=27.21  Aligned_cols=72  Identities=11%  Similarity=0.197  Sum_probs=42.3

Q ss_pred             CCCcEEEEcCCCC-----CCCCCCEEEE-ecCCcEEEEcC-CCcEEEcCCCCCC----Ce----eEEEEeeCCCeeEEcC
Q 047082            4 PKNSCSLNSIFAS-----PVCGKPTFSL-GSDGNLVLAEA-DGTVVCQSNTANK----GV----VGFKLLPNGNMVLHDS   68 (263)
Q Consensus         4 ~~~~~vWvANr~~-----Pv~~~~~l~l-~~~G~L~l~~~-~g~~~Wst~~~~~----~~----~~a~Lld~GNlvl~~~   68 (263)
                      ..-..+|.-+-..     |+.+...+.+ +.+|.|+.+|. +|..+|+......    ..    ....-..+|.|+..|.
T Consensus        83 ~tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~a~d~  162 (377)
T TIGR03300        83 ETGKRLWRVDLDERLSGGVGADGGLVFVGTEKGEVIALDAEDGKELWRAKLSSEVLSPPLVANGLVVVRTNDGRLTALDA  162 (377)
T ss_pred             cCCcEeeeecCCCCcccceEEcCCEEEEEcCCCEEEEEECCCCcEeeeeccCceeecCCEEECCEEEEECCCCeEEEEEc
Confidence            3456788655433     3333444444 46788888886 6889998765321    11    1111234567777775


Q ss_pred             -CCceEee
Q 047082           69 -KGNFIWQ   75 (263)
Q Consensus        69 -~~~~~Wq   75 (263)
                       +++.+|+
T Consensus       163 ~tG~~~W~  170 (377)
T TIGR03300       163 ATGERLWT  170 (377)
T ss_pred             CCCceeeE
Confidence             6788996


No 30 
>PF09064 Tme5_EGF_like:  Thrombomodulin like fifth domain, EGF-like;  InterPro: IPR015149 This domain adopts a fold similar to other EGF domains, with a flat major and a twisted minor beta sheet. Disulphide pairing, however, is not of the usual 1-3, 2-4, 5-6 type; rather 1-2, 3-4, 5-6 pairing is found. Its extended major sheet (strands beta-2 and beta-3 and the connecting loop) projects into thrombin's active site groove. This domain is required for interaction of thrombomodulin with thrombin, and subsequent activation of protein-C []. ; GO: 0004888 transmembrane signaling receptor activity, 0016021 integral to membrane
Probab=41.53  E-value=16  Score=21.99  Aligned_cols=16  Identities=31%  Similarity=0.511  Sum_probs=11.7

Q ss_pred             ccCCCCCcCCCCCCCC
Q 047082          201 VCDDNQCVACPTEKGL  216 (263)
Q Consensus       201 ~C~~~~~~~C~c~~g~  216 (263)
                      .|+.+...+|.||.||
T Consensus        11 ~CDpn~~~~C~CPeGy   26 (34)
T PF09064_consen   11 DCDPNSPGQCFCPEGY   26 (34)
T ss_pred             ccCCCCCCceeCCCce
Confidence            4665555589999987


No 31 
>PF08277 PAN_3:  PAN-like domain;  InterPro: IPR006583 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs The PAN-3 or CW is a domain associated with a number of Caenorhabditis elegans hypothetical proteins.
Probab=41.30  E-value=36  Score=23.17  Aligned_cols=24  Identities=29%  Similarity=0.652  Sum_probs=21.3

Q ss_pred             CCCCHHHHHHHhhccCCeEEEEec
Q 047082          235 TAIKVEDCGRKCTSDCKCSGYFYH  258 (263)
Q Consensus       235 ~~~s~~~C~~~Cl~nCsC~a~~y~  258 (263)
                      ...+.++|-..|..+=.|+.+.+.
T Consensus        18 ~~~sw~~Cv~~C~~~~~C~la~~~   41 (71)
T PF08277_consen   18 TNTSWDDCVQKCYNDENCVLAYFD   41 (71)
T ss_pred             cCCCHHHHhHHhCCCCEEEEEEeC
Confidence            347889999999999999998876


No 32 
>KOG4649 consensus PQQ (pyrrolo-quinoline quinone) repeat protein [Secondary metabolites biosynthesis, transport and catabolism]
Probab=39.79  E-value=50  Score=29.72  Aligned_cols=43  Identities=16%  Similarity=0.198  Sum_probs=32.9

Q ss_pred             CcEEEEcCCCCCCCCC-----CEEEE-ecCCcEEEEcCCCcEEEcCCCC
Q 047082            6 NSCSLNSIFASPVCGK-----PTFSL-GSDGNLVLAEADGTVVCQSNTA   48 (263)
Q Consensus         6 ~~~vWvANr~~Pv~~~-----~~l~l-~~~G~L~l~~~~g~~~Wst~~~   48 (263)
                      .+.+|.|.|..||-.+     ..+.+ +-||+|.-.|+.|+.||.-.+.
T Consensus       169 ~~~~w~~~~~~PiF~splcv~~sv~i~~VdG~l~~f~~sG~qvwr~~t~  217 (354)
T KOG4649|consen  169 STEFWAATRFGPIFASPLCVGSSVIITTVDGVLTSFDESGRQVWRPATK  217 (354)
T ss_pred             cceehhhhcCCccccCceeccceEEEEEeccEEEEEcCCCcEEEeecCC
Confidence            4678999999998763     33444 4789999889989999976554


No 33 
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=39.78  E-value=68  Score=29.30  Aligned_cols=17  Identities=18%  Similarity=0.223  Sum_probs=9.9

Q ss_pred             eeCCCeeEEcC-CCceEe
Q 047082           58 LPNGNMVLHDS-KGNFIW   74 (263)
Q Consensus        58 ld~GNlvl~~~-~~~~~W   74 (263)
                      -.+|.|.+.+. +++.+|
T Consensus       327 ~~~G~l~~~d~~tG~~~~  344 (377)
T TIGR03300       327 DFEGYLHWLSREDGSFVA  344 (377)
T ss_pred             eCCCEEEEEECCCCCEEE
Confidence            34566666654 456666


No 34 
>cd05852 Ig5_Contactin-1 Fifth Ig domain of contactin-1. Ig5_Contactin-1: fifth Ig domain of the neural cell adhesion molecule contactin-1. Contactins are comprised of six Ig domains followed by four fibronectin type III (FnIII) domains anchored to the membrane by glycosylphosphatidylinositol. Contactin-1 is differentially expressed in tumor tissues and may through a RhoA mechanism, facilitate invasion and metastasis of human lung adenocarcinoma.
Probab=39.05  E-value=35  Score=23.57  Aligned_cols=33  Identities=21%  Similarity=0.236  Sum_probs=22.3

Q ss_pred             CCCCcEEEEcCCCCCCCCCCEEEEecCCcEEEEc
Q 047082            3 YPKNSCSLNSIFASPVCGKPTFSLGSDGNLVLAE   36 (263)
Q Consensus         3 ~~~~~~vWvANr~~Pv~~~~~l~l~~~G~L~l~~   36 (263)
                      .|..++.|.=+. .++..+..+.+..+|.|+|.+
T Consensus        13 ~P~p~v~W~k~~-~~l~~~~r~~~~~~g~L~I~~   45 (73)
T cd05852          13 APKPKFSWSKGT-ELLVNNSRISIWDDGSLEILN   45 (73)
T ss_pred             eCCCEEEEEeCC-EecccCCCEEEcCCCEEEECc
Confidence            467788998654 355555567777778887754


No 35 
>smart00605 CW CW domain.
Probab=39.00  E-value=37  Score=24.79  Aligned_cols=23  Identities=22%  Similarity=0.574  Sum_probs=19.7

Q ss_pred             CCCHHHHHHHhhccCCeEEEEec
Q 047082          236 AIKVEDCGRKCTSDCKCSGYFYH  258 (263)
Q Consensus       236 ~~s~~~C~~~Cl~nCsC~a~~y~  258 (263)
                      ..+.++|...|..+..|+.+...
T Consensus        21 ~~sw~~Ci~~C~~~~~Cvlay~~   43 (94)
T smart00605       21 TLSWDECIQKCYEDSNCVLAYGN   43 (94)
T ss_pred             CCCHHHHHHHHhCCCceEEEecC
Confidence            46889999999999999987544


No 36 
>PF07645 EGF_CA:  Calcium-binding EGF domain;  InterPro: IPR001881 A sequence of about forty amino-acid residues found in epidermal growth factor (EGF) has been shown [, , , , , ] to be present in a large number of membrane-bound and extracellular, mostly animal, proteins. Many of these proteins require calcium for their biological function and a calcium-binding site has been found at the N terminus of some EGF-like domains []. Calcium-binding may be crucial for numerous protein-protein interactions. For human coagulation factor IX it has been shown [] that the calcium-ligands form a pentagonal bipyramid. The first, third and fourth conserved negatively charged or polar residues are side chain ligands. The latter is possibly hydroxylated (see aspartic acid and asparagine hydroxylation site) []. A conserved aromatic residue, as well as the second conserved negative residue, are thought to be involved in stabilising the calcium-binding site. As in non-calcium binding EGF-like domains, there are six conserved cysteines and the structure of both types is very similar as calcium-binding induces only strictly local structural changes [].  +------------------+ +---------+ | | | | nxnnC-x(3,14)-C-x(3,7)-CxxbxxxxaxC-x(1,6)-C-x(8,13)-Cx | | +------------------+ 'n': negatively charged or polar residue [DEQN] 'b': possibly beta-hydroxylated residue [DN] 'a': aromatic amino acid 'C': cysteine, involved in disulphide bond 'x': any amino acid. ; GO: 0005509 calcium ion binding; PDB: 2VJ3_A 1TOZ_A 1LMJ_A 1UZQ_A 1UZK_A 1UZJ_B 1UZP_A 1EMO_A 1EMN_A 2RR0_A ....
Probab=35.69  E-value=9.4  Score=23.67  Aligned_cols=28  Identities=18%  Similarity=0.418  Sum_probs=20.8

Q ss_pred             CCCC-CCCCCCcccC-CCCCcCCCCCCCCc
Q 047082          190 TQLP-ERCSKLGVCD-DNQCVACPTEKGLL  217 (263)
Q Consensus       190 C~~~-~~CG~~g~C~-~~~~~~C~c~~g~~  217 (263)
                      |... ..|..++.|. ......|.|++||.
T Consensus         5 C~~~~~~C~~~~~C~N~~Gsy~C~C~~Gy~   34 (42)
T PF07645_consen    5 CAEGPHNCPENGTCVNTEGSYSCSCPPGYE   34 (42)
T ss_dssp             TTTTSSSSSTTSEEEEETTEEEEEESTTEE
T ss_pred             cCCCCCcCCCCCEEEcCCCCEEeeCCCCcE
Confidence            5553 5798899995 34566899999985


No 37 
>KOG0291 consensus WD40-repeat-containing subunit of the 18S rRNA processing complex [RNA processing and modification]
Probab=35.28  E-value=5.3e+02  Score=26.75  Aligned_cols=108  Identities=19%  Similarity=0.246  Sum_probs=62.0

Q ss_pred             CCcEEEEcCCCCCC-------CCCCEEEEecCCcEEEEcC-CCcE-EEcCCC--------CCCCeeEEEEee--CCCeeE
Q 047082            5 KNSCSLNSIFASPV-------CGKPTFSLGSDGNLVLAEA-DGTV-VCQSNT--------ANKGVVGFKLLP--NGNMVL   65 (263)
Q Consensus         5 ~~~~vWvANr~~Pv-------~~~~~l~l~~~G~L~l~~~-~g~~-~Wst~~--------~~~~~~~a~Lld--~GNlvl   65 (263)
                      .+.-||-..+..-+       ++-..+.++..|+.+|... +|++ +|--..        ...+..-..|..  +|.||.
T Consensus       372 gKVKvWn~~SgfC~vTFteHts~Vt~v~f~~~g~~llssSLDGtVRAwDlkRYrNfRTft~P~p~QfscvavD~sGelV~  451 (893)
T KOG0291|consen  372 GKVKVWNTQSGFCFVTFTEHTSGVTAVQFTARGNVLLSSSLDGTVRAWDLKRYRNFRTFTSPEPIQFSCVAVDPSGELVC  451 (893)
T ss_pred             CcEEEEeccCceEEEEeccCCCceEEEEEEecCCEEEEeecCCeEEeeeecccceeeeecCCCceeeeEEEEcCCCCEEE
Confidence            34678888875532       2225789999999888644 6776 787652        222332333444  499999


Q ss_pred             EcCCCc---eEeeeccCCCceeccCcccC---CCCCCe-EEEEecCCceEEEecCCCCCcee
Q 047082           66 HDSKGN---FIWQSFDCPTDTLLVGQSLL---SVKENV-SFVMEPKRFTLYYKGSNSPQPVL  120 (263)
Q Consensus        66 ~~~~~~---~~WqSFd~PTDTlLpGq~l~---S~~dps-sl~l~~~~~~~~~~~~~~~~~~~  120 (263)
                      ...-+.   .+| |+.       -||-|-   .+.-|. .|.+.+.+-.+....|+.+.+.|
T Consensus       452 AG~~d~F~IfvW-S~q-------TGqllDiLsGHEgPVs~l~f~~~~~~LaS~SWDkTVRiW  505 (893)
T KOG0291|consen  452 AGAQDSFEIFVW-SVQ-------TGQLLDILSGHEGPVSGLSFSPDGSLLASGSWDKTVRIW  505 (893)
T ss_pred             eeccceEEEEEE-Eee-------cCeeeehhcCCCCcceeeEEccccCeEEeccccceEEEE
Confidence            865433   477 332       355443   344454 67777766555544444443333


No 38 
>smart00179 EGF_CA Calcium-binding EGF-like domain.
Probab=34.90  E-value=36  Score=19.73  Aligned_cols=28  Identities=18%  Similarity=0.402  Sum_probs=19.0

Q ss_pred             CCCCCCCCCCcccCCC-CCcCCCCCCCCc
Q 047082          190 TQLPERCSKLGVCDDN-QCVACPTEKGLL  217 (263)
Q Consensus       190 C~~~~~CG~~g~C~~~-~~~~C~c~~g~~  217 (263)
                      |.....|...+.|... ....|.|++||.
T Consensus         5 C~~~~~C~~~~~C~~~~g~~~C~C~~g~~   33 (39)
T smart00179        5 CASGNPCQNGGTCVNTVGSYRCECPPGYT   33 (39)
T ss_pred             CcCCCCcCCCCEeECCCCCeEeECCCCCc
Confidence            5443568878889642 345799999875


No 39 
>cd00216 PQQ_DH Dehydrogenases with pyrrolo-quinoline quinone (PQQ) as cofactor, like ethanol, methanol, and membrane bound glucose dehydrogenases. The alignment model contains an 8-bladed beta-propeller.
Probab=34.00  E-value=1.4e+02  Score=28.75  Aligned_cols=72  Identities=19%  Similarity=0.290  Sum_probs=42.3

Q ss_pred             CCCcEEEEcCCC-------CCCCCCCEEEEe-cCCcEEEEcC-CCcEEEcCCCCCC-----------Cee----EEEE--
Q 047082            4 PKNSCSLNSIFA-------SPVCGKPTFSLG-SDGNLVLAEA-DGTVVCQSNTANK-----------GVV----GFKL--   57 (263)
Q Consensus         4 ~~~~~vWvANr~-------~Pv~~~~~l~l~-~~G~L~l~~~-~g~~~Wst~~~~~-----------~~~----~a~L--   57 (263)
                      ++-.++|...-.       .|+-....+-+. .+|.|+-+|. .|.++|+......           +++    ..++  
T Consensus        37 ~~~~~~W~~~~~~~~~~~~sPvv~~g~vy~~~~~g~l~AlD~~tG~~~W~~~~~~~~~~~~~~~~~~g~~~~~~~~V~v~  116 (488)
T cd00216          37 KKLKVAWTFSTGDERGQEGTPLVVDGDMYFTTSHSALFALDAATGKVLWRYDPKLPADRGCCDVVNRGVAYWDPRKVFFG  116 (488)
T ss_pred             hcceeeEEEECCCCCCcccCCEEECCEEEEeCCCCcEEEEECCCChhhceeCCCCCccccccccccCCcEEccCCeEEEe
Confidence            345678887654       354444444444 5788887775 5889999764211           100    0111  


Q ss_pred             eeCCCeeEEcC-CCceEee
Q 047082           58 LPNGNMVLHDS-KGNFIWQ   75 (263)
Q Consensus        58 ld~GNlvl~~~-~~~~~Wq   75 (263)
                      -.+|.|+-.|. +++.+|+
T Consensus       117 ~~~g~v~AlD~~TG~~~W~  135 (488)
T cd00216         117 TFDGRLVALDAETGKQVWK  135 (488)
T ss_pred             cCCCeEEEEECCCCCEeee
Confidence            13566666665 6899997


No 40 
>PF13570 PQQ_3:  PQQ-like domain; PDB: 3HXJ_B 3Q54_A.
Probab=31.24  E-value=40  Score=20.21  Aligned_cols=8  Identities=25%  Similarity=0.305  Sum_probs=2.8

Q ss_pred             cEEEcCCC
Q 047082           40 TVVCQSNT   47 (263)
Q Consensus        40 ~~~Wst~~   47 (263)
                      .++|+..+
T Consensus         2 ~~~W~~~~    9 (40)
T PF13570_consen    2 KVLWSYDT    9 (40)
T ss_dssp             -EEEEEE-
T ss_pred             ceeEEEEC
Confidence            34444443


No 41 
>PF05935 Arylsulfotrans:  Arylsulfotransferase (ASST);  InterPro: IPR010262 This family consists of several bacterial arylsulphotransferase proteins. Arylsulphotransferase (ASST) transfers a sulphate group from phenolic sulphate esters to a phenolic acceptor substrate [].; PDB: 3ETT_B 3ELQ_A 3ETS_A.
Probab=30.30  E-value=49  Score=31.94  Aligned_cols=52  Identities=29%  Similarity=0.514  Sum_probs=30.6

Q ss_pred             CCcEEEEcCCCcEEEcCCCCCCCeeEEEEeeCCCeeEEc--------CCCceEeeeccCCC
Q 047082           29 DGNLVLAEADGTVVCQSNTANKGVVGFKLLPNGNMVLHD--------SKGNFIWQSFDCPT   81 (263)
Q Consensus        29 ~G~L~l~~~~g~~~Wst~~~~~~~~~a~Lld~GNlvl~~--------~~~~~~WqSFd~PT   81 (263)
                      .+..++.|.+|.++|..............+++|+|....        -.|+++|+ ++.|.
T Consensus       127 ~~~~~~iD~~G~Vrw~~~~~~~~~~~~~~l~nG~ll~~~~~~~~e~D~~G~v~~~-~~l~~  186 (477)
T PF05935_consen  127 SSYTYLIDNNGDVRWYLPLDSGSDNSFKQLPNGNLLIGSGNRLYEIDLLGKVIWE-YDLPG  186 (477)
T ss_dssp             EEEEEEEETTS-EEEEE-GGGT--SSEEE-TTS-EEEEEBTEEEEE-TT--EEEE-EE--T
T ss_pred             CceEEEECCCccEEEEEccCccccceeeEcCCCCEEEecCCceEEEcCCCCEEEe-eecCC
Confidence            467889999999999987543222226789999987653        35789998 77776


No 42 
>PF01683 EB:  EB module;  InterPro: IPR006149  The EB domain has no known function. It is found in several Caenorhabditis sp. and Drosophila sp. proteins. The domain contains 8 conserved cysteines that probably form four disulphide bridges and is found associated with kunitz domains IPR002223 from INTERPRO 
Probab=28.40  E-value=43  Score=21.51  Aligned_cols=27  Identities=22%  Similarity=0.492  Sum_probs=21.1

Q ss_pred             CCCCCCCCCCCcccCCCCCcCCCCCCCCcc
Q 047082          189 PTQLPERCSKLGVCDDNQCVACPTEKGLLG  218 (263)
Q Consensus       189 ~C~~~~~CG~~g~C~~~~~~~C~c~~g~~~  218 (263)
                      .|.....|-.+++|..+   .|.|++||..
T Consensus        21 ~C~~~~qC~~~s~C~~g---~C~C~~g~~~   47 (52)
T PF01683_consen   21 SCESDEQCIGGSVCVNG---RCQCPPGYVE   47 (52)
T ss_pred             CCCCcCCCCCcCEEcCC---EeECCCCCEe
Confidence            39988999999999432   5889998743


No 43 
>smart00564 PQQ beta-propeller repeat. Beta-propeller repeat occurring in enzymes with pyrrolo-quinoline quinone (PQQ) as cofactor, in Ire1p-like Ser/Thr kinases, and in prokaryotic dehydrogenases.
Probab=27.36  E-value=1.2e+02  Score=16.86  Aligned_cols=17  Identities=29%  Similarity=0.575  Sum_probs=8.8

Q ss_pred             cCCcEEEEcC-CCcEEEc
Q 047082           28 SDGNLVLAEA-DGTVVCQ   44 (263)
Q Consensus        28 ~~G~L~l~~~-~g~~~Ws   44 (263)
                      .+|.|+-.|. +|..+|.
T Consensus        14 ~~g~l~a~d~~~G~~~W~   31 (33)
T smart00564       14 TDGTLYALDAKTGEILWT   31 (33)
T ss_pred             CCCEEEEEEcccCcEEEE
Confidence            3455555544 4555664


No 44 
>PF06006 DUF905:  Bacterial protein of unknown function (DUF905);  InterPro: IPR009253 This family consists of several short hypothetical proteobacterial proteins of unknown function.; PDB: 2HJJ_A.
Probab=24.77  E-value=80  Score=22.20  Aligned_cols=17  Identities=24%  Similarity=0.976  Sum_probs=10.1

Q ss_pred             eeEEcCCCceEeeeccC
Q 047082           63 MVLHDSKGNFIWQSFDC   79 (263)
Q Consensus        63 lvl~~~~~~~~WqSFd~   79 (263)
                      ||+|+.++.-+|..|.+
T Consensus        35 lvvRd~~g~mvWRaWNF   51 (70)
T PF06006_consen   35 LVVRDTEGQMVWRAWNF   51 (70)
T ss_dssp             EEEE-SS--EEEEEESS
T ss_pred             EEEEcCCCcEEEEeecc
Confidence            67777777778877654


No 45 
>cd00054 EGF_CA Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.
Probab=24.51  E-value=70  Score=18.03  Aligned_cols=28  Identities=18%  Similarity=0.413  Sum_probs=18.0

Q ss_pred             CCCCCCCCCCcccCCC-CCcCCCCCCCCc
Q 047082          190 TQLPERCSKLGVCDDN-QCVACPTEKGLL  217 (263)
Q Consensus       190 C~~~~~CG~~g~C~~~-~~~~C~c~~g~~  217 (263)
                      |.....|...+.|... ....|.|+.||.
T Consensus         5 C~~~~~C~~~~~C~~~~~~~~C~C~~g~~   33 (38)
T cd00054           5 CASGNPCQNGGTCVNTVGSYRCSCPPGYT   33 (38)
T ss_pred             CCCCCCcCCCCEeECCCCCeEeECCCCCc
Confidence            4433467777788543 344799998874


No 46 
>cd00216 PQQ_DH Dehydrogenases with pyrrolo-quinoline quinone (PQQ) as cofactor, like ethanol, methanol, and membrane bound glucose dehydrogenases. The alignment model contains an 8-bladed beta-propeller.
Probab=24.19  E-value=1.7e+02  Score=28.14  Aligned_cols=16  Identities=25%  Similarity=0.700  Sum_probs=6.9

Q ss_pred             CCCeeEEcC-CCceEee
Q 047082           60 NGNMVLHDS-KGNFIWQ   75 (263)
Q Consensus        60 ~GNlvl~~~-~~~~~Wq   75 (263)
                      +|.|.-.|. +++.+|+
T Consensus       415 dG~l~ald~~tG~~lW~  431 (488)
T cd00216         415 DGYFRAFDATTGKELWK  431 (488)
T ss_pred             CCeEEEEECCCCceeeE
Confidence            344444432 3445553


No 47 
>KOG3881 consensus Uncharacterized conserved protein [Function unknown]
Probab=23.82  E-value=2.9e+02  Score=26.10  Aligned_cols=61  Identities=15%  Similarity=0.160  Sum_probs=41.9

Q ss_pred             CEEEEecCCcEEEEcCCC--cEEEcCCCCCCCeeEEEEeeCCCeeEEcCCCceEeeeccCCCce
Q 047082           22 PTFSLGSDGNLVLAEADG--TVVCQSNTANKGVVGFKLLPNGNMVLHDSKGNFIWQSFDCPTDT   83 (263)
Q Consensus        22 ~~l~l~~~G~L~l~~~~g--~~~Wst~~~~~~~~~a~Lld~GNlvl~~~~~~~~WqSFd~PTDT   83 (263)
                      --++++.-|.+.|+|...  +||=+..-...+.+...|.-+||+|+....... --+||+-+--
T Consensus       218 ~fat~T~~hqvR~YDt~~qRRPV~~fd~~E~~is~~~l~p~gn~Iy~gn~~g~-l~~FD~r~~k  280 (412)
T KOG3881|consen  218 KFATITRYHQVRLYDTRHQRRPVAQFDFLENPISSTGLTPSGNFIYTGNTKGQ-LAKFDLRGGK  280 (412)
T ss_pred             eEEEEecceeEEEecCcccCcceeEeccccCcceeeeecCCCcEEEEecccch-hheecccCce
Confidence            467888889999998642  466555544455677889999999988543222 3589976543


No 48 
>PF10636 hemP:  Hemin uptake protein hemP;  InterPro: IPR019600  This entry represents bacterial proteins that are involved in the uptake of the iron source hemin []. ; PDB: 2JRA_B 2LOJ_A.
Probab=23.44  E-value=90  Score=19.29  Aligned_cols=13  Identities=23%  Similarity=0.593  Sum_probs=8.5

Q ss_pred             EEEEecCCcEEEE
Q 047082           23 TFSLGSDGNLVLA   35 (263)
Q Consensus        23 ~l~l~~~G~L~l~   35 (263)
                      .|.++..|.|+|+
T Consensus        25 ~LR~Tr~gKLILT   37 (38)
T PF10636_consen   25 RLRITRQGKLILT   37 (38)
T ss_dssp             EEEEETTTEEEEE
T ss_pred             EeeEccCCcEEEc
Confidence            5666666666664


No 49 
>PF02035 Coagulin:  Coagulin;  InterPro: IPR000275 Coagulogen is a gel-forming protein of hemolymph that hinders the spread of invaders by immobilising them [, ]. The protein contains a single 175- residue polypeptide chain; this is cleaved after Arg-18 and Arg-46 by a clotting enzyme contained in the hemocyte and activated by a bacterial endotoxin (lipopolysaccharide). Cleavage releases two chains of coagulin, A and B, linked by two disulphide bonds, together with the peptide C [, ]. Gel formation results from interlinking of coagulin molecules. Secondary structure prediction suggests the C peptide forms an alpha- helix, which is released during the proteolytic conversion of coagulogen to coagulin gel []. The beta-sheet structure and 16 half-cystines found in the molecule appear to yield a compact protein stable to acid and heat. Mammalian blood coagulation is based on the proteolytically induced polymerisation of fibrinogens. Initially, fibrin monomers noncovalently interact with each other. The resulting homopolymers are further stabilised when the plasma transglutaminase (TGase) intermolecularly cross-links epsilon-(gamma-glutamyl)lysine bonds. In crustaceans, hemolymph coagulation depends on the TGase-mediated cross-linking of specific plasma-clotting proteins, but without the proteolytic cascade. In horseshoe crabs, the proteolytic coagulation cascade triggered by lipopolysaccharides and beta-1,3-glucans leads to the conversion of coagulogen into coagulin, resulting in noncovalent coagulin homopolymers through head-to-tail interaction. Horseshoe crab TGase, however, does not cross-link coagulins intermolecularly. Recently, we found that coagulins are cross-linked on hemocyte cell surface proteins called proxins. This indicates that a cross-linking reaction at the final stage of hemolymph coagulation is an important innate immune system of horseshoe crabs [].; GO: 0042381 hemolymph coagulation, 0005576 extracellular region; PDB: 1AOC_A.
Probab=22.08  E-value=76  Score=25.21  Aligned_cols=34  Identities=15%  Similarity=0.304  Sum_probs=15.0

Q ss_pred             ecCCEEEEEecCCC-----CCCC-CCCCCC--CCCCcccCCC
Q 047082          172 MHGNLKIYTHYDKV-----DSQP-TQLPER--CSKLGVCDDN  205 (263)
Q Consensus       172 ~dG~lr~y~~~~~~-----~w~~-C~~~~~--CG~~g~C~~~  205 (263)
                      ..|.+|+..-.+..     .|+. |..||.  ||.+|-|+..
T Consensus       101 ~a~efrvivqapragfrqcvwqhkcraygsn~c~~~grctqq  142 (174)
T PF02035_consen  101 VAGEFRVIVQAPRAGFRQCVWQHKCRAYGSNNCGFNGRCTQQ  142 (174)
T ss_dssp             TTS-EEEE--BCCCTB-B---EEEET-TSSSB-SSS-EE--E
T ss_pred             ecceEEEEEeCchhhHHHHHHHhhhccccccccCcCceeccc
Confidence            34556655533322     3665 988754  9999999753


No 50 
>cd05764 Ig_2 Subgroup of the immunoglobulin (Ig) superfamily. Ig_2: subgroup of the immunoglobulin (Ig) domain found in the Ig superfamily. The Ig superfamily is a heterogenous group of proteins, built on a common fold comprised of a sandwich of two beta sheets. Members of the Ig superfamily are components of immunoglobulin, neuroglia, cell surface glycoproteins, such as T-cell receptors, CD2, CD4, CD8, and membrane glycoproteins, such as butyrophilin and chondroitin sulfate proteoglycan core protein. A predominant feature of most Ig domains is a disulfide bridge connecting the two beta-sheets with a tryptophan residue packed against the disulfide bond.
Probab=21.81  E-value=1.6e+02  Score=19.61  Aligned_cols=33  Identities=12%  Similarity=0.050  Sum_probs=20.3

Q ss_pred             CCCCcEEEEcCCCCCCCCCCEEEEecCCcEEEE
Q 047082            3 YPKNSCSLNSIFASPVCGKPTFSLGSDGNLVLA   35 (263)
Q Consensus         3 ~~~~~~vWvANr~~Pv~~~~~l~l~~~G~L~l~   35 (263)
                      .|...+.|.-+.+.++.......+..+|.|.|.
T Consensus        13 ~P~p~v~W~~~~~~~~~~~~~~~~~~~~~L~i~   45 (74)
T cd05764          13 DPEPAIHWISPDGKLISNSSRTLVYDNGTLDIL   45 (74)
T ss_pred             cCCCEEEEEeCCCEEecCCCeEEEecCCEEEEE
Confidence            366788888655556554444445556666664


No 51 
>PF05833 FbpA:  Fibronectin-binding protein A N-terminus (FbpA);  InterPro: IPR008616 This family consists of the N-terminal region of the prokaryotic fibronectin-binding protein, the C-terminal region is IPR008532 from INTERPRO. Fibronectin binding is considered to be an important virulence factor in streptococcal infections. Fibronectin is a dimeric glycoprotein that is present in a soluble form in plasma and extracellular fluids; it is also present in a fibrillar form on cell surfaces. Both the soluble and cellular forms of fibronectin may be incorporated into the extracellular tissue matrix. While fibronectin has critical roles in eukaryotic cellular processes, such as adhesion, migration and differentiation, it is also a substrate for the attachment of bacteria. The binding of pathogenic Streptococcus pyogenes and Staphylococcus aureus to epithelial cells via fibronectin facilitates their internalisation and systemic spread within the host [].; PDB: 3DOA_A 2ZBK_F 2HKJ_A 1Z5B_A 1Z5C_B 1MX0_F 1Z5A_A 1MU5_A 1Z59_A.
Probab=21.80  E-value=75  Score=30.18  Aligned_cols=39  Identities=18%  Similarity=0.336  Sum_probs=22.2

Q ss_pred             eEEEEeeC-CCeeEEcCCCceEeeeccCCCc-----eeccCcccC
Q 047082           53 VGFKLLPN-GNMVLHDSKGNFIWQSFDCPTD-----TLLVGQSLL   91 (263)
Q Consensus        53 ~~a~Lld~-GNlvl~~~~~~~~WqSFd~PTD-----TlLpGq~l~   91 (263)
                      -.++|... ||++|.|+++.+|+---.++.+     +++||+...
T Consensus       115 Li~El~g~~~NiiL~d~~~~Il~a~~~~~~~~~~~R~i~~G~~Y~  159 (455)
T PF05833_consen  115 LIIELMGRHSNIILTDEDGKILDALRRVSFSQSRDREILPGEPYI  159 (455)
T ss_dssp             EEEE--GGG-EEEEEETT-BEEEESS-B---------BSTTSB--
T ss_pred             EEEEEcCCcccEEEEcCCCeEEeehhhcCcccccceeeccCcccc
Confidence            45677777 9999999888877754444554     899999976


No 52 
>KOG3848 consensus Extracellular protein TEM7, contains PSI domain (tumor endothelial marker in humans) [Extracellular structures]
Probab=21.42  E-value=3.3e+02  Score=26.06  Aligned_cols=51  Identities=14%  Similarity=0.136  Sum_probs=35.8

Q ss_pred             EcCCCCCCCCCCEEEEecCCcEEEEcCCCcEEEcCCCC------CCCeeEEEEeeCCCeeEEc
Q 047082           11 NSIFASPVCGKPTFSLGSDGNLVLAEADGTVVCQSNTA------NKGVVGFKLLPNGNMVLHD   67 (263)
Q Consensus        11 vANr~~Pv~~~~~l~l~~~G~L~l~~~~g~~~Wst~~~------~~~~~~a~Lld~GNlvl~~   67 (263)
                      .||-+...++++.+..-.+|.+++      +.|.....      ++-.-.|.|+.+|.+|..-
T Consensus       194 MANFdts~snnS~V~y~DnGtafv------vqWdnV~Lqd~~d~gsFTFqatL~~dGdIVFaY  250 (516)
T KOG3848|consen  194 MANFDTSYSNNSTVVYFDNGTAFV------VQWDNVQLQDDKDEGSFTFQATLHKDGDIVFAY  250 (516)
T ss_pred             hhcCCccccCCceEEEecCCeEEE------EEeeeEEeccCCCCCcEEEEEEeccCCcEEEEE
Confidence            488888788888888888998776      34554321      2223467888889888764


No 53 
>TIGR03075 PQQ_enz_alc_DH PQQ-dependent dehydrogenase, methanol/ethanol family. This protein family has a phylogenetic distribution very similar to that coenzyme PQQ biosynthesis enzymes, as shown by partial phylogenetic profiling. Genes in this family often are found adjacent to the PQQ biosynthesis genes themselves. An unusual, strained disulfide bond between adjacent Cys residues contributes to PQQ-binding, as does a Trp residue that is part of a PQQ enzyme repeat (see pfam01011). Characterized members include the dehydrogenase subunit of a membrane-anchored, three subunit alcohol (ethanol) dehydrogenase of Gluconobacter suboxydans, a homodimeric ethanol dehydrogenase in Pseudomonas aeruginosa, and the large subunit of an alpha2/beta2 heterotetrameric methanol dehydrogenase in Methylobacterium extorquens.
Probab=20.70  E-value=4.1e+02  Score=26.00  Aligned_cols=75  Identities=19%  Similarity=0.208  Sum_probs=0.0

Q ss_pred             CCCCCCcEEEEcCCCCCCCCCC-----------------EEEEecCCcEEEEcC-CCcEEEcCCCC-----CCCeeEEEE
Q 047082            1 MEYPKNSCSLNSIFASPVCGKP-----------------TFSLGSDGNLVLAEA-DGTVVCQSNTA-----NKGVVGFKL   57 (263)
Q Consensus         1 ~~~~~~~~vWvANr~~Pv~~~~-----------------~l~l~~~G~L~l~~~-~g~~~Wst~~~-----~~~~~~a~L   57 (263)
                      ++..+-..+|.-+...|.....                 .+.-+.+|.|+-+|. .|.++|+....     ....+...+
T Consensus        84 lDa~TGk~lW~~~~~~~~~~~~~~~~~~~~rg~av~~~~v~v~t~dg~l~ALDa~TGk~~W~~~~~~~~~~~~~tssP~v  163 (527)
T TIGR03075        84 LDAKTGKELWKYDPKLPDDVIPVMCCDVVNRGVALYDGKVFFGTLDARLVALDAKTGKVVWSKKNGDYKAGYTITAAPLV  163 (527)
T ss_pred             EECCCCceeeEecCCCCcccccccccccccccceEECCEEEEEcCCCEEEEEECCCCCEEeecccccccccccccCCcEE


Q ss_pred             ee--------------CCCeeEEcC-CCceEee
Q 047082           58 LP--------------NGNMVLHDS-KGNFIWQ   75 (263)
Q Consensus        58 ld--------------~GNlvl~~~-~~~~~Wq   75 (263)
                      .+              .|.|+-.|. +++.+|+
T Consensus       164 ~~g~Vivg~~~~~~~~~G~v~AlD~~TG~~lW~  196 (527)
T TIGR03075       164 VKGKVITGISGGEFGVRGYVTAYDAKTGKLVWR  196 (527)
T ss_pred             ECCEEEEeecccccCCCcEEEEEECCCCceeEe


No 54 
>COG1520 FOG: WD40-like repeat [Function unknown]
Probab=20.48  E-value=4.3e+02  Score=24.07  Aligned_cols=42  Identities=29%  Similarity=0.414  Sum_probs=29.1

Q ss_pred             CCcEEEEcCCCC-------CCCCCCEEEEe-cCCcEEEEcCC-CcEEEcCC
Q 047082            5 KNSCSLNSIFAS-------PVCGKPTFSLG-SDGNLVLAEAD-GTVVCQSN   46 (263)
Q Consensus         5 ~~~~vWvANr~~-------Pv~~~~~l~l~-~~G~L~l~~~~-g~~~Wst~   46 (263)
                      .-+.+|..+...       |+...+.+-+. .+|.|+-.+.+ |..+|...
T Consensus       130 ~G~~~W~~~~~~~~~~~~~~v~~~~~v~~~s~~g~~~al~~~tG~~~W~~~  180 (370)
T COG1520         130 TGTLVWSRNVGGSPYYASPPVVGDGTVYVGTDDGHLYALNADTGTLKWTYE  180 (370)
T ss_pred             CCcEEEEEecCCCeEEecCcEEcCcEEEEecCCCeEEEEEccCCcEEEEEe
Confidence            456788877666       23334556666 57999888877 89999944


No 55 
>COG3236 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=20.29  E-value=55  Score=26.52  Aligned_cols=21  Identities=33%  Similarity=0.588  Sum_probs=16.1

Q ss_pred             EEEeeCCCeeEEcC-CCceEee
Q 047082           55 FKLLPNGNMVLHDS-KGNFIWQ   75 (263)
Q Consensus        55 a~Lld~GNlvl~~~-~~~~~Wq   75 (263)
                      ..||+||+.||... .+..+|-
T Consensus       116 e~LL~Tgd~vLVE~s~~D~~WG  137 (162)
T COG3236         116 ELLLATGDAVLVEASPNDAIWG  137 (162)
T ss_pred             HHHHhcCCeeEEecCCCcceee
Confidence            45899999999954 4567884


Done!