Query 003257
Match_columns 836
No_of_seqs 127 out of 151
Neff 3.8
Searched_HMMs 46136
Date Thu Mar 28 20:02:56 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/003257.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/003257hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF12036 DUF3522: Protein of u 100.0 1.6E-47 3.6E-52 380.9 16.6 179 576-783 2-186 (186)
2 PF05875 Ceramidase: Ceramidas 97.9 0.00018 3.8E-09 75.5 13.0 52 585-644 29-88 (262)
3 TIGR01065 hlyIII channel prote 96.6 0.038 8.3E-07 56.6 13.5 51 604-657 35-89 (204)
4 PF04080 Per1: Per1-like ; In 96.5 0.064 1.4E-06 58.0 14.8 44 604-655 89-132 (267)
5 PF12955 DUF3844: Domain of un 96.2 0.0048 1E-07 58.2 3.8 64 528-597 4-83 (103)
6 PF07974 EGF_2: EGF-like domai 96.1 0.005 1.1E-07 46.7 2.8 26 535-568 7-32 (32)
7 PF03006 HlyIII: Haemolysin-II 96.0 0.1 2.2E-06 52.5 12.3 44 609-655 49-94 (222)
8 PRK15087 hemolysin; Provisiona 95.9 0.25 5.4E-06 51.6 15.1 46 608-657 54-103 (219)
9 KOG2329 Alkaline ceramidase [L 93.5 0.3 6.4E-06 53.2 8.5 43 584-626 36-86 (276)
10 COG1272 Predicted membrane pro 93.0 0.86 1.9E-05 48.4 10.9 39 752-791 176-214 (226)
11 PF13965 SID-1_RNA_chan: dsRNA 92.0 2.9 6.2E-05 49.9 14.6 19 773-791 529-547 (570)
12 KOG2970 Predicted membrane pro 91.4 1.5 3.2E-05 48.6 10.5 141 604-784 141-294 (319)
13 cd00053 EGF Epidermal growth f 89.3 0.4 8.6E-06 34.1 2.9 30 534-569 6-36 (36)
14 PF00008 EGF: EGF-like domain 88.1 0.4 8.7E-06 36.0 2.3 29 533-566 3-31 (32)
15 KOG1225 Teneurin-1 and related 87.6 0.32 7E-06 57.1 2.3 32 530-571 312-343 (525)
16 PHA02887 EGF-like protein; Pro 87.1 0.39 8.6E-06 46.7 2.1 45 521-570 75-123 (126)
17 cd00054 EGF_CA Calcium-binding 86.1 0.72 1.6E-05 33.4 2.7 34 530-569 3-38 (38)
18 PF04863 EGF_alliinase: Alliin 85.8 0.3 6.5E-06 41.8 0.6 35 534-571 17-52 (56)
19 PF12036 DUF3522: Protein of u 84.0 7 0.00015 40.2 9.6 44 632-676 59-102 (186)
20 KOG4289 Cadherin EGF LAG seven 83.7 0.77 1.7E-05 58.8 2.9 36 530-571 1240-1276(2531)
21 PF12661 hEGF: Human growth fa 83.6 0.5 1.1E-05 29.7 0.7 13 556-568 1-13 (13)
22 smart00179 EGF_CA Calcium-bind 83.2 1.2 2.7E-05 32.8 2.9 34 530-569 3-39 (39)
23 KOG1225 Teneurin-1 and related 82.9 0.85 1.8E-05 53.7 2.8 58 501-570 217-280 (525)
24 PF04151 PPC: Bacterial pre-pe 82.1 9.2 0.0002 32.6 8.1 66 429-514 4-69 (70)
25 smart00181 EGF Epidermal growt 78.0 2.2 4.7E-05 31.3 2.6 28 534-568 6-34 (35)
26 PHA03099 epidermal growth fact 76.4 2 4.4E-05 42.5 2.6 37 529-570 42-82 (139)
27 KOG1226 Integrin beta subunit 76.4 1.7 3.6E-05 53.0 2.5 34 527-570 544-581 (783)
28 KOG3607 Meltrins, fertilins an 74.9 1.7 3.7E-05 53.0 2.1 35 528-571 624-658 (716)
29 COG5237 PER1 Predicted membran 72.1 6 0.00013 43.3 5.1 51 603-661 135-190 (319)
30 KOG4260 Uncharacterized conser 67.5 3.1 6.7E-05 45.8 1.8 40 530-572 142-185 (350)
31 KOG4243 Macrophage maturation- 66.0 23 0.00049 38.6 7.7 24 767-790 255-278 (298)
32 PF07645 EGF_CA: Calcium-bindi 55.9 9 0.0002 30.1 2.1 25 534-564 10-34 (42)
33 smart00051 DSL delta serrate l 54.3 8.9 0.00019 33.4 2.0 26 534-568 38-63 (63)
34 PF00954 S_locus_glycop: S-loc 54.0 28 0.00061 32.3 5.4 34 524-564 72-107 (110)
35 PF00053 Laminin_EGF: Laminin 51.5 9.4 0.0002 30.7 1.6 28 535-570 2-33 (49)
36 KOG3879 Predicted membrane pro 50.0 2.3E+02 0.0049 31.2 11.8 24 810-833 212-235 (267)
37 cd00055 EGF_Lam Laminin-type e 49.5 14 0.00031 30.1 2.4 28 535-570 3-34 (50)
38 PF12947 EGF_3: EGF domain; I 46.3 11 0.00023 29.5 1.2 28 534-567 6-33 (36)
39 KOG1219 Uncharacterized conser 32.8 33 0.00072 47.3 3.0 41 525-571 3899-3940(4289)
40 PF12658 Ten1: Telomere cappin 30.9 54 0.0012 32.1 3.5 48 125-172 36-97 (124)
41 PF12662 cEGF: Complement Clr- 30.7 32 0.00069 25.2 1.4 15 556-570 3-21 (24)
42 KOG1219 Uncharacterized conser 30.2 39 0.00084 46.7 3.0 38 529-572 3942-3980(4289)
43 PF04151 PPC: Bacterial pre-pe 23.7 3.3E+02 0.0072 23.1 6.7 64 256-334 3-68 (70)
44 PRK05420 aquaporin Z; Provisio 22.6 4.1E+02 0.0088 28.4 8.5 21 746-766 203-223 (231)
45 KOG4812 Golgi-associated prote 22.4 1.1E+02 0.0023 33.7 4.2 31 670-706 176-207 (262)
No 1
>PF12036 DUF3522: Protein of unknown function (DUF3522); InterPro: IPR021910 This family of proteins is functionally uncharacterised. This protein is found in eukaryotes. Proteins in this family are typically between 220 to 787 amino acids in length.
Probab=100.00 E-value=1.6e-47 Score=380.89 Aligned_cols=179 Identities=31% Similarity=0.432 Sum_probs=163.9
Q ss_pred hHHHHHHHHHHHhhhhhHHHHHHHHHHhHHHHHHHHHHHHHHhhhhhccccc----ceeeccchhHhHhHhHHHHHHHHH
Q 003257 576 RGHVQQSVALIASNAAALLPAYQALRQKAFAEWVLFTASGISSGLYHACDVG----TWCALSFNVLQFMDFWLSFMAVVS 651 (836)
Q Consensus 576 ~~~~~q~lLLtLSNLaFlP~I~vA~kRr~~~Ea~Vy~fTMffS~fYHACD~g----~~Cim~ydvLQf~DF~gSimSiwv 651 (836)
.+...|+++||+||++|+|+|++|+|||+++|++||+|||++|+||||||++ .+|++++++||++||+++++++|+
T Consensus 2 ~~~~~~~l~l~lSnl~~lP~i~~a~rr~~~~Ea~v~~~tm~~S~~YHacd~~~~~~~lc~~~~~~L~~~~~~~s~~~~~v 81 (186)
T PF12036_consen 2 FEQLLQFLLLTLSNLAFLPTIYVAVRRRYHFEAFVYTFTMFFSTFYHACDSGPGEIFLCIMDWHRLQNIDFIGSFLSIWV 81 (186)
T ss_pred hhhHHHHHHHHHHHHHHHHHHHHHHHHhhHHHHHHHHHHHHHHHhcccccCCCCceEEeechHHHHHHHHHHHHHHHHHH
Confidence 4568899999999999999999999999999999999999999999999964 499999999999999999999999
Q ss_pred HHHhhccchhHHHhhhhhhhHHHHHHHHHhhccCC--ccchhhHHHHHHHHHHHHHhhhcccccceeeeccccccccchh
Q 003257 652 TFIYLTTIDEALKRTIHTVVAILTAMMAITKATRS--SNIILVISIGAAGLLIGLLVELSTKFRSFSLRFGFCMNMVDRQ 729 (836)
Q Consensus 652 T~I~MA~~~e~lk~~~~~~~~IL~Al~~~~q~~R~--wn~iiPI~i~~lgili~Wl~~~~t~~R~~~~s~~~~~~yP~~~ 729 (836)
|+++||++++++|+.+++++++++++. .|.||+ ||+++|+++++++++++|++|+++ |+.+ ||+++
T Consensus 82 tl~~~a~~~~~~~~~l~~~~~~~~ai~--~~~~~~~~~~~~~Pi~~~~~i~~~~w~~r~~~--~~~~--------~~~~~ 149 (186)
T PF12036_consen 82 TLCAMARLDEPLKSVLHYFGALVIAIF--QQKDRWSLWNTIGPILIGLLILLVSWLYRCRR--RRRC--------YPPSW 149 (186)
T ss_pred HHHHhccCCHHHHHHHHHHHHHHHHHH--HhhCcccchhhHHHHHHHHHHHHHHHheeccc--CCcc--------CChHH
Confidence 999999999999999999999998877 455555 699999999999999999998653 3334 77876
Q ss_pred HHHHHHHHHhHHhhhhcccchhhhHHHHHHHHHHhhhhcccCcceeEehhHHHH
Q 003257 730 QTIMEWLRNFMKTILRRFRWGFVLVGFAALAMAAISWKLETSQSYWIWHSIWHV 783 (836)
Q Consensus 730 ~~i~~w~~~~~~~l~rrfRw~f~L~Ggi~la~~aI~~flET~dnY~y~HSiWHi 783 (836)
+ ||++++.||+++++.|+.+|+||+|||||+||+||+
T Consensus 150 ~-----------------~~~~~l~~g~~~~~~Gl~~f~et~dnY~~~HSlWHi 186 (186)
T PF12036_consen 150 R-----------------RWLFYLLPGIIFFILGLDLFLETNDNYRIVHSLWHI 186 (186)
T ss_pred H-----------------HHHHHHHHHHHHHHHHHhHhhcCCCcEEEEeeeeeC
Confidence 5 799999999999999999999999999999999996
No 2
>PF05875 Ceramidase: Ceramidase; InterPro: IPR008901 This entry consists of several ceramidases. Ceramidases are enzymes involved in regulating cellular levels of ceramides, sphingoid bases, and their phosphates.; GO: 0016811 hydrolase activity, acting on carbon-nitrogen (but not peptide) bonds, in linear amides, 0006672 ceramide metabolic process, 0016021 integral to membrane
Probab=97.86 E-value=0.00018 Score=75.54 Aligned_cols=52 Identities=29% Similarity=0.281 Sum_probs=35.6
Q ss_pred HHHhhhhhHHHHHHHHH---H-----hHHHHHHHHHHHHHHhhhhhcccccceeeccchhHhHhHhHH
Q 003257 585 LIASNAAALLPAYQALR---Q-----KAFAEWVLFTASGISSGLYHACDVGTWCALSFNVLQFMDFWL 644 (836)
Q Consensus 585 LtLSNLaFlP~I~vA~k---R-----r~~~Ea~Vy~fTMffS~fYHACD~g~~Cim~ydvLQf~DF~g 644 (836)
=|+||++|+......++ | ++..-.+...+-++.|+.||+= +++ ..|.+|=+-
T Consensus 29 NtlSNl~fi~~al~gl~~~~~~~~~~~~~l~~~~l~~VGiGS~~FHaT-------l~~-~~ql~DelP 88 (262)
T PF05875_consen 29 NTLSNLAFIVAALYGLYLARRRGLERRFALLYLGLALVGIGSFLFHAT-------LSY-WTQLLDELP 88 (262)
T ss_pred HHHHHHHHHHHHHHHHHHHhhccccchhHHHHHHHHHHHHhHHHHHhC-------hhh-hHHHhhhhh
Confidence 37999999887654433 2 3455555566778999999984 454 367788654
No 3
>TIGR01065 hlyIII channel protein, hemolysin III family. This family includes proteins from pathogenic and non-pathogenic bacteria, Homo sapiens and Drosophila. In Bacillus cereus, a pathogen, it has been show to function as a channel-forming cytolysin. The human protein is expressed preferentially in mature macrophages, consistent with a role cytolytic role.
Probab=96.61 E-value=0.038 Score=56.65 Aligned_cols=51 Identities=24% Similarity=0.342 Sum_probs=35.3
Q ss_pred HHHHHHHHHHHHH----HhhhhhcccccceeeccchhHhHhHhHHHHHHHHHHHHhhc
Q 003257 604 AFAEWVLFTASGI----SSGLYHACDVGTWCALSFNVLQFMDFWLSFMAVVSTFIYLT 657 (836)
Q Consensus 604 ~~~Ea~Vy~fTMf----fS~fYHACD~g~~Cim~ydvLQf~DF~gSimSiwvT~I~MA 657 (836)
......+|.+++. .|++||.=.... -..+.|+++|-.+=.+.|+.|++-..
T Consensus 35 ~~~~~~vy~~~~~~~~~~St~yH~~~~s~---~~~~~~~rlD~~gI~~lIaGsytP~~ 89 (204)
T TIGR01065 35 AVLGFSIYGISLILLFLVSTLYHSIPKGS---KAKNWLRKIDHSMIYVLIAGTYTPFL 89 (204)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHCCcCch---hHHHHHHHccHHHHHHHHHHhhHHHH
Confidence 3455667766654 599999765211 24568999999998888888765543
No 4
>PF04080 Per1: Per1-like ; InterPro: IPR007217 A member of this family has been implemented in protein processing in the endoplasmic reticulum [].
Probab=96.50 E-value=0.064 Score=58.00 Aligned_cols=44 Identities=14% Similarity=0.119 Sum_probs=35.0
Q ss_pred HHHHHHHHHHHHHHhhhhhcccccceeeccchhHhHhHhHHHHHHHHHHHHh
Q 003257 604 AFAEWVLFTASGISSGLYHACDVGTWCALSFNVLQFMDFWLSFMAVVSTFIY 655 (836)
Q Consensus 604 ~~~Ea~Vy~fTMffS~fYHACD~g~~Cim~ydvLQf~DF~gSimSiwvT~I~ 655 (836)
+..-+++...+=++|+.+|+.|.. .=+.+|-+++.+.|...+.+
T Consensus 89 ~~~~~~v~~naW~wStvFH~RD~~--------~TE~lDYf~A~a~vl~~l~~ 132 (267)
T PF04080_consen 89 YIIYAIVSMNAWIWSTVFHTRDTP--------LTEKLDYFSAGATVLFGLYA 132 (267)
T ss_pred eehHHHHHHHHHHHHHHHHHhccc--------HhhHhHHhhhHHHHHHHHHH
Confidence 567889999999999999999974 12368999988777776654
No 5
>PF12955 DUF3844: Domain of unknown function (DUF3844); InterPro: IPR024382 This presumed domain is found in fungal species. It contains 8 largely conserved cysteine residues. This domain is found in proteins thought to be located in the endoplasmic reticulum.
Probab=96.18 E-value=0.0048 Score=58.20 Aligned_cols=64 Identities=25% Similarity=0.438 Sum_probs=40.7
Q ss_pred Eecccc---cCCCCCceeeeeeccCCceEEeeeeeCC-------------CCCCcCCCccccchhHHHHHHHHHHHhhhh
Q 003257 528 SLERCP---KRCSSHGQCRNAFDASGLTLYSFCACDR-------------DHGGFDCSVELVSHRGHVQQSVALIASNAA 591 (836)
Q Consensus 528 sls~C~---~~Cg~~G~C~~l~~~sG~~~ys~C~C~~-------------Gy~GwdCtd~svs~~~~~~q~lLLtLSNLa 591 (836)
+.+.|. ++|++||+|......+++ .-=.|.|.+ .|+|.+|...-++ .+..|++.+-++
T Consensus 4 S~~aC~~~Tn~CsgHG~C~~~~~~~~~-~C~~C~C~~T~~~~~~~~~ktt~W~G~aCqKkDvS-----~~F~L~~~~ti~ 77 (103)
T PF12955_consen 4 SNDACENATNNCSGHGSCVKKYGSGGG-DCFACKCKPTVVKTGSGKGKTTHWGGPACQKKDVS-----VPFWLFAGFTIA 77 (103)
T ss_pred CHHHHHHhccCCCCCceEeeccCCCcc-ceEEEEeeccccccccccCceeeeccccccccccc-----chhhHHHHHHHH
Confidence 446675 799999999987543321 222799999 7999999865333 234444444444
Q ss_pred hHHHHH
Q 003257 592 ALLPAY 597 (836)
Q Consensus 592 FlP~I~ 597 (836)
++..+.
T Consensus 78 lv~~~~ 83 (103)
T PF12955_consen 78 LVVLVA 83 (103)
T ss_pred HHHHHH
Confidence 444433
No 6
>PF07974 EGF_2: EGF-like domain; InterPro: IPR013111 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length. This entry contains EGF domains found in a variety of extracellular and membrane proteins
Probab=96.09 E-value=0.005 Score=46.72 Aligned_cols=26 Identities=46% Similarity=1.049 Sum_probs=22.7
Q ss_pred CCCCCceeeeeeccCCceEEeeeeeCCCCCCcCC
Q 003257 535 RCSSHGQCRNAFDASGLTLYSFCACDRDHGGFDC 568 (836)
Q Consensus 535 ~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdC 568 (836)
.|++||+|... .| .|.|++||.|.+|
T Consensus 7 ~C~~~G~C~~~---~g-----~C~C~~g~~G~~C 32 (32)
T PF07974_consen 7 ICSGHGTCVSP---CG-----RCVCDSGYTGPDC 32 (32)
T ss_pred ccCCCCEEeCC---CC-----EEECCCCCcCCCC
Confidence 59999999964 24 8999999999988
No 7
>PF03006 HlyIII: Haemolysin-III related; InterPro: IPR004254 Members of this family are integral membrane proteins. This family includes proteins that are hemolysin-III homologs.; GO: 0016021 integral to membrane
Probab=95.96 E-value=0.1 Score=52.52 Aligned_cols=44 Identities=16% Similarity=0.305 Sum_probs=30.2
Q ss_pred HHHHHHHHHhhhhhc--ccccceeeccchhHhHhHhHHHHHHHHHHHHh
Q 003257 609 VLFTASGISSGLYHA--CDVGTWCALSFNVLQFMDFWLSFMAVVSTFIY 655 (836)
Q Consensus 609 ~Vy~fTMffS~fYHA--CD~g~~Cim~ydvLQf~DF~gSimSiwvT~I~ 655 (836)
+-....+++|++||. |-+... .+..|+++|-.|-.+.+..+.+.
T Consensus 49 ~~~~~~~~~St~yH~f~~~s~~~---~~~~~~~lD~~gI~l~i~gs~~p 94 (222)
T PF03006_consen 49 LSAILCFLCSTLYHLFSCHSEGK---VYHIFLRLDYAGIFLLIAGSYTP 94 (222)
T ss_pred HHHHHHHHhHHHhhCCCcCCcHH---HHHHHHhcchhhhhHhHhhhhhh
Confidence 334455778999999 533211 57899999999976666665443
No 8
>PRK15087 hemolysin; Provisional
Probab=95.89 E-value=0.25 Score=51.61 Aligned_cols=46 Identities=22% Similarity=0.303 Sum_probs=32.8
Q ss_pred HHHHHHH----HHHhhhhhcccccceeeccchhHhHhHhHHHHHHHHHHHHhhc
Q 003257 608 WVLFTAS----GISSGLYHACDVGTWCALSFNVLQFMDFWLSFMAVVSTFIYLT 657 (836)
Q Consensus 608 a~Vy~fT----MffS~fYHACD~g~~Cim~ydvLQf~DF~gSimSiwvT~I~MA 657 (836)
..+|..+ +.+|++||.-... -..+.|+++|=.+=.+.|..|+.-++
T Consensus 54 ~~vy~~s~~~l~~~StlYH~~~~~----~~~~~~~rlDh~~I~llIaGsytP~~ 103 (219)
T PRK15087 54 YSLYGGSMILLFLASTLYHAIPHQ----RAKRWLKKFDHCAIYLLIAGTYTPFL 103 (219)
T ss_pred HHHHHHHHHHHHHHHHHHHCCCch----HHHHHHHHccHHHHHHHHHHhhHHHH
Confidence 3455554 4579999987632 23569999999998888888776543
No 9
>KOG2329 consensus Alkaline ceramidase [Lipid transport and metabolism]
Probab=93.49 E-value=0.3 Score=53.22 Aligned_cols=43 Identities=40% Similarity=0.478 Sum_probs=32.2
Q ss_pred HHHHhhhhhHHHHH----HHHHH----hHHHHHHHHHHHHHHhhhhhcccc
Q 003257 584 ALIASNAAALLPAY----QALRQ----KAFAEWVLFTASGISSGLYHACDV 626 (836)
Q Consensus 584 LLtLSNLaFlP~I~----vA~kR----r~~~Ea~Vy~fTMffS~fYHACD~ 626 (836)
.=|.||+.|+.++. -++|+ |++.-.+.+++-+++|..|||-=+
T Consensus 36 ~NT~sN~~fil~~~~~l~~~y~~~~e~~~~l~~v~~~ivgl~S~~fH~TL~ 86 (276)
T KOG2329|consen 36 ANTESNSPFILLAFIGLHCAYRQKLEKRAYLICVLFTIVGLGSMYFHMTLV 86 (276)
T ss_pred HHHhhcchHHHHHHHHHHHHHHHHhhhhHHHHHHHHHHHHHHHhhhhhhHH
Confidence 34778888874433 44443 578899999999999999999754
No 10
>COG1272 Predicted membrane protein, hemolysin III homolog [General function prediction only]
Probab=92.98 E-value=0.86 Score=48.45 Aligned_cols=39 Identities=21% Similarity=0.312 Sum_probs=27.6
Q ss_pred hhHHHHHHHHHHhhhhcccCcceeEehhHHHHHHhhheeE
Q 003257 752 VLVGFAALAMAAISWKLETSQSYWIWHSIWHVSIYTSSFF 791 (836)
Q Consensus 752 ~L~Ggi~la~~aI~~flET~dnY~y~HSiWHi~Ia~S~~F 791 (836)
+..||++..++++++..+- |-..+.|-+||+++-+++++
T Consensus 176 l~~GGv~YsvG~ifY~~~~-~~~~~~H~iwH~fVv~ga~~ 214 (226)
T COG1272 176 LALGGVLYSVGAIFYVLRI-DRIPYSHAIWHLFVVGGAAC 214 (226)
T ss_pred HHHHhHHheeeeEEEEEee-ccCCchHHHHHHHHHHHHHH
Confidence 4666666666666553332 77889999999998776654
No 11
>PF13965 SID-1_RNA_chan: dsRNA-gated channel SID-1
Probab=91.99 E-value=2.9 Score=49.90 Aligned_cols=19 Identities=37% Similarity=0.992 Sum_probs=17.2
Q ss_pred ceeEehhHHHHHHhhheeE
Q 003257 773 SYWIWHSIWHVSIYTSSFF 791 (836)
Q Consensus 773 nY~y~HSiWHi~Ia~S~~F 791 (836)
+++=+|-+||++-|++.||
T Consensus 529 ~f~D~HDiwH~~SA~alff 547 (570)
T PF13965_consen 529 GFFDWHDIWHFLSAIALFF 547 (570)
T ss_pred CccccHHHHHHHHHHHHHH
Confidence 6788999999999999887
No 12
>KOG2970 consensus Predicted membrane protein [Function unknown]
Probab=91.38 E-value=1.5 Score=48.62 Aligned_cols=141 Identities=23% Similarity=0.281 Sum_probs=73.1
Q ss_pred HHHHHHHHHHHHHHhhhhhcccccceeeccchhHhHhHhHHHHH----HHHHHHHhhccchhH-HHhhhhhhhHHHHHHH
Q 003257 604 AFAEWVLFTASGISSGLYHACDVGTWCALSFNVLQFMDFWLSFM----AVVSTFIYLTTIDEA-LKRTIHTVVAILTAMM 678 (836)
Q Consensus 604 ~~~Ea~Vy~fTMffS~fYHACD~g~~Cim~ydvLQf~DF~gSim----SiwvT~I~MA~~~e~-lk~~~~~~~~IL~Al~ 678 (836)
.+.-|.+...+-+.|+.+|.=|.. .=+.||-.++.+ +.-++++-|-+++.. ..+- ++.++..|..
T Consensus 141 ~~I~a~i~mnawiwSsvFH~rD~~--------lTEklDYf~A~~~vlf~ly~a~ir~~~i~~~~~~~~--~ita~fla~y 210 (319)
T KOG2970|consen 141 WLIYAYIGMNAWIWSSVFHIRDVP--------LTEKLDYFSAYLTVLFGLYVALIRMLSIQSLPALRG--MITAIFLAFY 210 (319)
T ss_pred hhhHHHHHHHHHHHHHhhhhcCCc--------hHhhhhHHHHHHHHHHHHHHHHHHHHHHhcchhhhH--HHHHHHHHHH
Confidence 467788888999999999999873 123567666543 333444444444433 2222 2233333333
Q ss_pred H--Hhh-----ccCCccchhhHHHHHHHHHHHHHhhhcccccceeeeccccccccchhHHHHHHHHHhHHhhhhcccchh
Q 003257 679 A--ITK-----ATRSSNIILVISIGAAGLLIGLLVELSTKFRSFSLRFGFCMNMVDRQQTIMEWLRNFMKTILRRFRWGF 751 (836)
Q Consensus 679 ~--~~q-----~~R~wn~iiPI~i~~lgili~Wl~~~~t~~R~~~~s~~~~~~yP~~~~~i~~w~~~~~~~l~rrfRw~f 751 (836)
+ +.+ .|=.+|+..=+++|.+. ++.|++-. -|+|+ .|..|+ +|.+
T Consensus 211 a~Hi~yls~~~fdYgyNm~~~v~~g~iq-~vlw~~~~-~~~~~----------~~s~~~-----------------i~~~ 261 (319)
T KOG2970|consen 211 ANHILYLSFYNFDYGYNMIVCVAIGVIQ-LVLWLVWS-FKKRN----------LPSFWR-----------------IWPI 261 (319)
T ss_pred HHHHHHHhheecccccceeeehhhHHHH-HHHHHHHH-HHhhc----------Ccchhh-----------------hhHH
Confidence 2 112 23334766545555433 34554432 12343 344332 6777
Q ss_pred hhHHHHHHHHHHhhhhc-ccCcceeEehhHHHHH
Q 003257 752 VLVGFAALAMAAISWKL-ETSQSYWIWHSIWHVS 784 (836)
Q Consensus 752 ~L~Ggi~la~~aI~~fl-ET~dnY~y~HSiWHi~ 784 (836)
++.....+|++ +=.++ ..=..|.=-|++||.+
T Consensus 262 ~i~~~~~LA~s-LEi~DFpPy~~~iDAHALWHla 294 (319)
T KOG2970|consen 262 LIVIFFFLAMS-LEIFDFPPYAWLIDAHALWHLA 294 (319)
T ss_pred HHHHHHHHHHH-HHhhcCCchhhhcchHHHHHhh
Confidence 77765554442 22233 2233444469999975
No 13
>cd00053 EGF Epidermal growth factor domain, found in epidermal growth factor (EGF) presents in a large number of proteins, mostly animal; the list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied; the functional significance of EGF-like domains in what appear to be unrelated proteins is not yet clear; a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase); the domain includes six cysteine residues which have been shown to be involved in disulfide bonds; the main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet; Subdomains between the conserved cysteines vary in length; the region between the 5th and 6th cysteine contains two conserved glycines of which at least one is present in most EGF-like domains; a subset of these bind calcium.
Probab=89.28 E-value=0.4 Score=34.12 Aligned_cols=30 Identities=33% Similarity=0.687 Sum_probs=24.0
Q ss_pred cCCCCCceeeeeeccCCceEEeeeeeCCCCCCc-CCC
Q 003257 534 KRCSSHGQCRNAFDASGLTLYSFCACDRDHGGF-DCS 569 (836)
Q Consensus 534 ~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~Gw-dCt 569 (836)
..|.++|+|.... |.+ .|.|..||.|+ .|.
T Consensus 6 ~~C~~~~~C~~~~---~~~---~C~C~~g~~g~~~C~ 36 (36)
T cd00053 6 NPCSNGGTCVNTP---GSY---RCVCPPGYTGDRSCE 36 (36)
T ss_pred CCCCCCCEEecCC---CCe---EeECCCCCcccCCcC
Confidence 6788899999753 323 79999999999 773
No 14
>PF00008 EGF: EGF-like domain This is a sub-family of the Pfam entry This is a sub-family of the Pfam entry; InterPro: IPR006209 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length.; GO: 0005515 protein binding; PDB: 1WHE_A 1CCF_A 1APO_A 1WHF_A 2VJ3_A 1TOZ_A 4D90_B 3CFW_A 1EDM_B 1IXA_A ....
Probab=88.15 E-value=0.4 Score=35.99 Aligned_cols=29 Identities=24% Similarity=0.511 Sum_probs=23.3
Q ss_pred ccCCCCCceeeeeeccCCceEEeeeeeCCCCCCc
Q 003257 533 PKRCSSHGQCRNAFDASGLTLYSFCACDRDHGGF 566 (836)
Q Consensus 533 ~~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~Gw 566 (836)
++.|.++|.|.... .+.| .|.|.+||.|.
T Consensus 3 ~~~C~n~g~C~~~~--~~~y---~C~C~~G~~G~ 31 (32)
T PF00008_consen 3 SNPCQNGGTCIDLP--GGGY---TCECPPGYTGK 31 (32)
T ss_dssp TTSSTTTEEEEEES--TSEE---EEEEBTTEEST
T ss_pred CCcCCCCeEEEeCC--CCCE---EeECCCCCccC
Confidence 36899999999976 2324 99999999984
No 15
>KOG1225 consensus Teneurin-1 and related extracellular matrix proteins, contain EGF-like repeats [Signal transduction mechanisms; Extracellular structures]
Probab=87.58 E-value=0.32 Score=57.05 Aligned_cols=32 Identities=44% Similarity=1.006 Sum_probs=27.7
Q ss_pred cccccCCCCCceeeeeeccCCceEEeeeeeCCCCCCcCCCcc
Q 003257 530 ERCPKRCSSHGQCRNAFDASGLTLYSFCACDRDHGGFDCSVE 571 (836)
Q Consensus 530 s~C~~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdCtd~ 571 (836)
..||.||.+||+|.. | .|.|++||.|.+|+..
T Consensus 312 ~~cpadC~g~G~Ci~-----G-----~C~C~~Gy~G~~C~~~ 343 (525)
T KOG1225|consen 312 RRCPADCSGHGKCID-----G-----ECLCDEGYTGELCIQR 343 (525)
T ss_pred ccCCccCCCCCcccC-----C-----ceEeCCCCcCCccccc
Confidence 449999999999993 5 8999999999999873
No 16
>PHA02887 EGF-like protein; Provisional
Probab=87.05 E-value=0.39 Score=46.67 Aligned_cols=45 Identities=29% Similarity=0.700 Sum_probs=35.2
Q ss_pred ceEEEEEEecccc----cCCCCCceeeeeeccCCceEEeeeeeCCCCCCcCCCc
Q 003257 521 SETVMSVSLERCP----KRCSSHGQCRNAFDASGLTLYSFCACDRDHGGFDCSV 570 (836)
Q Consensus 521 ~~v~~sisls~C~----~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdCtd 570 (836)
.+...+..-.||+ +=|= ||+|.++.+.. -.+|.|..||.|.-|..
T Consensus 75 ~~rk~~~hf~pC~~eyk~YCi-HG~C~yI~dL~----epsCrC~~GYtG~RCE~ 123 (126)
T PHA02887 75 FKRKNSMFFEKCKNDFNDFCI-NGECMNIIDLD----EKFCICNKGYTGIRCDE 123 (126)
T ss_pred hhhccccCccccChHhhCEee-CCEEEccccCC----CceeECCCCcccCCCCc
Confidence 3445566678998 4684 99999998765 35999999999999964
No 17
>cd00054 EGF_CA Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.
Probab=86.12 E-value=0.72 Score=33.38 Aligned_cols=34 Identities=29% Similarity=0.810 Sum_probs=25.5
Q ss_pred cccc--cCCCCCceeeeeeccCCceEEeeeeeCCCCCCcCCC
Q 003257 530 ERCP--KRCSSHGQCRNAFDASGLTLYSFCACDRDHGGFDCS 569 (836)
Q Consensus 530 s~C~--~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdCt 569 (836)
..|. ..|..+|.|.... |.| .|.|..||.|..|.
T Consensus 3 ~~C~~~~~C~~~~~C~~~~---~~~---~C~C~~g~~g~~C~ 38 (38)
T cd00054 3 DECASGNPCQNGGTCVNTV---GSY---RCSCPPGYTGRNCE 38 (38)
T ss_pred ccCCCCCCcCCCCEeECCC---CCe---EeECCCCCcCCcCC
Confidence 4565 4788889998642 333 79999999998883
No 18
>PF04863 EGF_alliinase: Alliinase EGF-like domain; InterPro: IPR006947 Allicin is a thiosulphinate that gives rise to dithiines, allyl sulphides and ajoenes, the three groups of active compounds in Allium species. Allicin is synthesised from sulphoxide cysteine derivatives by alliinase, whose C-S lyase activity cleaves C(beta)-S(gamma) bonds. It is thought that this enzyme forms part of a primitive plant defence system [].; GO: 0016846 carbon-sulfur lyase activity; PDB: 1LK9_B 2HOX_C 2HOR_A.
Probab=85.82 E-value=0.3 Score=41.83 Aligned_cols=35 Identities=34% Similarity=0.574 Sum_probs=19.0
Q ss_pred cCCCCCceeeeeecc-CCceEEeeeeeCCCCCCcCCCcc
Q 003257 534 KRCSSHGQCRNAFDA-SGLTLYSFCACDRDHGGFDCSVE 571 (836)
Q Consensus 534 ~~Cg~~G~C~~l~~~-sG~~~ys~C~C~~Gy~GwdCtd~ 571 (836)
-.|++||+..+-.-. .| ...|.|...|+|.||+.-
T Consensus 17 i~CSGHGr~flDg~~~dG---~p~CECn~Cy~GpdCS~~ 52 (56)
T PF04863_consen 17 ISCSGHGRAFLDGLIADG---SPVCECNSCYGGPDCSTL 52 (56)
T ss_dssp S--TTSEE--TTS-EETT---EE--EE-TTEESTTS-EE
T ss_pred CCcCCCCeeeeccccccC---CccccccCCcCCCCcccC
Confidence 379999999863211 22 278999999999999853
No 19
>PF12036 DUF3522: Protein of unknown function (DUF3522); InterPro: IPR021910 This family of proteins is functionally uncharacterised. This protein is found in eukaryotes. Proteins in this family are typically between 220 to 787 amino acids in length.
Probab=84.01 E-value=7 Score=40.24 Aligned_cols=44 Identities=14% Similarity=0.152 Sum_probs=33.0
Q ss_pred ccchhHhHhHhHHHHHHHHHHHHhhccchhHHHhhhhhhhHHHHH
Q 003257 632 LSFNVLQFMDFWLSFMAVVSTFIYLTTIDEALKRTIHTVVAILTA 676 (836)
Q Consensus 632 m~ydvLQf~DF~gSimSiwvT~I~MA~~~e~lk~~~~~~~~IL~A 676 (836)
|....||++||...+.++....+.|+++.+ +++........+++
T Consensus 59 lc~~~~~~L~~~~~~~s~~~~~vtl~~~a~-~~~~~~~~l~~~~~ 102 (186)
T PF12036_consen 59 LCIMDWHRLQNIDFIGSFLSIWVTLCAMAR-LDEPLKSVLHYFGA 102 (186)
T ss_pred EeechHHHHHHHHHHHHHHHHHHHHHHhcc-CCHHHHHHHHHHHH
Confidence 789999999999999999999999998776 44433333333333
No 20
>KOG4289 consensus Cadherin EGF LAG seven-pass G-type receptor [Signal transduction mechanisms]
Probab=83.66 E-value=0.77 Score=58.80 Aligned_cols=36 Identities=31% Similarity=0.798 Sum_probs=29.1
Q ss_pred ccc-ccCCCCCceeeeeeccCCceEEeeeeeCCCCCCcCCCcc
Q 003257 530 ERC-PKRCSSHGQCRNAFDASGLTLYSFCACDRDHGGFDCSVE 571 (836)
Q Consensus 530 s~C-~~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdCtd~ 571 (836)
.-| -+.||+||.|..- + |+| +|.|++||.|.+|-.+
T Consensus 1240 DlCYs~pC~nng~C~sr--E-ggY---tCeCrpg~tGehCEvs 1276 (2531)
T KOG4289|consen 1240 DLCYSGPCGNNGRCRSR--E-GGY---TCECRPGFTGEHCEVS 1276 (2531)
T ss_pred HhhhcCCCCCCCceEEe--c-Cce---eEEecCCccccceeee
Confidence 456 3799999999974 3 457 9999999999999643
No 21
>PF12661 hEGF: Human growth factor-like EGF; PDB: 2YGQ_A 2E26_A 3A7Q_A 2YGP_A 2YGO_A 1HRE_A 1HAE_A 1HAF_A 1HRF_A.
Probab=83.56 E-value=0.5 Score=29.72 Aligned_cols=13 Identities=31% Similarity=0.864 Sum_probs=10.9
Q ss_pred eeeeCCCCCCcCC
Q 003257 556 FCACDRDHGGFDC 568 (836)
Q Consensus 556 ~C~C~~Gy~GwdC 568 (836)
.|.|.+||.|..|
T Consensus 1 ~C~C~~G~~G~~C 13 (13)
T PF12661_consen 1 TCQCPPGWTGPNC 13 (13)
T ss_dssp EEEE-TTEETTTT
T ss_pred CccCcCCCcCCCC
Confidence 4999999999988
No 22
>smart00179 EGF_CA Calcium-binding EGF-like domain.
Probab=83.16 E-value=1.2 Score=32.79 Aligned_cols=34 Identities=29% Similarity=0.788 Sum_probs=25.7
Q ss_pred cccc--cCCCCCceeeeeeccCCceEEeeeeeCCCCC-CcCCC
Q 003257 530 ERCP--KRCSSHGQCRNAFDASGLTLYSFCACDRDHG-GFDCS 569 (836)
Q Consensus 530 s~C~--~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~-GwdCt 569 (836)
..|. +.|..+|.|... .|.| .|.|..||. |..|.
T Consensus 3 ~~C~~~~~C~~~~~C~~~---~g~~---~C~C~~g~~~g~~C~ 39 (39)
T smart00179 3 DECASGNPCQNGGTCVNT---VGSY---RCECPPGYTDGRNCE 39 (39)
T ss_pred ccCcCCCCcCCCCEeECC---CCCe---EeECCCCCccCCcCC
Confidence 4565 479888999854 3434 699999999 98883
No 23
>KOG1225 consensus Teneurin-1 and related extracellular matrix proteins, contain EGF-like repeats [Signal transduction mechanisms; Extracellular structures]
Probab=82.88 E-value=0.85 Score=53.69 Aligned_cols=58 Identities=28% Similarity=0.371 Sum_probs=40.5
Q ss_pred EeeccCCcEEEEEEeeeCCC------ceEEEEEEecccccCCCCCceeeeeeccCCceEEeeeeeCCCCCCcCCCc
Q 003257 501 ILYVREGTWGFGIRHVNTSK------SETVMSVSLERCPKRCSSHGQCRNAFDASGLTLYSFCACDRDHGGFDCSV 570 (836)
Q Consensus 501 IpYPqtGtWYLsL~~~n~~~------~~v~~sisls~C~~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdCtd 570 (836)
--|-++|.|+... ..... ...-...+...||.+|.++|+|.. | .|.|+.||.|.||+.
T Consensus 217 ~~r~~~~~~~~~~--~~~~~ic~c~~~~~g~~c~~~~C~~~c~~~g~c~~-----G-----~CIC~~Gf~G~dC~e 280 (525)
T KOG1225|consen 217 TGRCREGRCFCTA--GFFDGICECPEGYFGPLCSTIYCPGGCTGRGQCVE-----G-----RCICPPGFTGDDCDE 280 (525)
T ss_pred ccccccCcccccc--cccCceeecCCceeCCccccccCCCCCcccceEeC-----C-----eEeCCCCCcCCCCCc
Confidence 3456778888754 22111 111222335689999999999997 5 899999999999986
No 24
>PF04151 PPC: Bacterial pre-peptidase C-terminal domain; InterPro: IPR007280 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. This domain is normally found at the C terminus of secreted archaeal and bacterial peptidases, the majority of which belong to MEROPS peptidase families M4 (vibriolysin, IPR001570 from INTERPRO), M9A amd M9B (microbial collangenase, IPR002169 from INTERPRO), M28 (aminopeptidase Ap1, IPR007484 from INTERPRO) and S8 (subtilisin family peptidases, IPR000209 from INTERPRO).; GO: 0008233 peptidase activity, 0006508 proteolysis; PDB: 4DY5_B 4DXZ_A 4DY3_B 3JQW_A 3JQX_C 1NQJ_B 1NQD_A 2O8O_A 1WMF_A 1WME_A ....
Probab=82.11 E-value=9.2 Score=32.59 Aligned_cols=66 Identities=23% Similarity=0.425 Sum_probs=38.2
Q ss_pred EEEeecCCCCCCceEEEEEeecceeeEEEEEeecCCCCCCcccccccccccccccccccccccCCccceeEEEeeccCCc
Q 003257 429 YFLLDIPRGAAGGSIHIQLTSDTKIKHEIYAKSGGLPSLQSWDYYYANRTNNSVGSMFFKLYNSSEEKVDFYILYVREGT 508 (836)
Q Consensus 429 ~f~l~Lp~gdSGG~L~v~L~~nks~~~~Vyar~g~~Ptlt~~D~~~~~~ts~s~~s~f~~~~nsS~~~a~L~IpYPqtGt 508 (836)
+|.+++| +|+.|+|+|..... +..++.....-+++.++|. .+ .....+..+.+.-|++|+
T Consensus 4 ~y~f~v~---ag~~l~i~l~~~~~-d~dl~l~~~~g~~~~~~d~-----~~-----------~~~~~~~~i~~~~~~~Gt 63 (70)
T PF04151_consen 4 YYSFTVP---AGGTLTIDLSGGSG-DADLYLYDSNGNSLASYDD-----SS-----------QSGGNDESITFTAPAAGT 63 (70)
T ss_dssp EEEEEES---TTEEEEEEECETTS-SEEEEEEETTSSSCEECCC-----CT-----------CETTSEEEEEEEESSSEE
T ss_pred EEEEEEc---CCCEEEEEEcCCCC-CeEEEEEcCCCCchhhhee-----cC-----------CCCCCccEEEEEcCCCEE
Confidence 6777777 67789999865552 3334433332355544431 00 001223445566799999
Q ss_pred EEEEEE
Q 003257 509 WGFGIR 514 (836)
Q Consensus 509 WYLsL~ 514 (836)
||+.++
T Consensus 64 Yyi~V~ 69 (70)
T PF04151_consen 64 YYIRVY 69 (70)
T ss_dssp EEEEEE
T ss_pred EEEEEE
Confidence 999874
No 25
>smart00181 EGF Epidermal growth factor-like domain.
Probab=77.98 E-value=2.2 Score=31.27 Aligned_cols=28 Identities=32% Similarity=0.706 Sum_probs=21.3
Q ss_pred cCCCCCceeeeeeccCCceEEeeeeeCCCCCC-cCC
Q 003257 534 KRCSSHGQCRNAFDASGLTLYSFCACDRDHGG-FDC 568 (836)
Q Consensus 534 ~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~G-wdC 568 (836)
+.|..+ .|... .|.+ .|.|..||.| ..|
T Consensus 6 ~~C~~~-~C~~~---~~~~---~C~C~~g~~g~~~C 34 (35)
T smart00181 6 GPCSNG-TCINT---PGSY---TCSCPPGYTGDKRC 34 (35)
T ss_pred CCCCCC-EEECC---CCCe---EeECCCCCccCCcc
Confidence 368777 89865 2333 8999999999 887
No 26
>PHA03099 epidermal growth factor-like protein (EGF-like protein); Provisional
Probab=76.42 E-value=2 Score=42.54 Aligned_cols=37 Identities=32% Similarity=0.832 Sum_probs=29.6
Q ss_pred ecccc----cCCCCCceeeeeeccCCceEEeeeeeCCCCCCcCCCc
Q 003257 529 LERCP----KRCSSHGQCRNAFDASGLTLYSFCACDRDHGGFDCSV 570 (836)
Q Consensus 529 ls~C~----~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdCtd 570 (836)
.++|+ +=| =||+|..+.+.+. .+|.|..||.|.-|-.
T Consensus 42 i~~Cp~ey~~YC-lHG~C~yI~dl~~----~~CrC~~GYtGeRCEh 82 (139)
T PHA03099 42 IRLCGPEGDGYC-LHGDCIHARDIDG----MYCRCSHGYTGIRCQH 82 (139)
T ss_pred cccCChhhCCEe-ECCEEEeeccCCC----ceeECCCCcccccccc
Confidence 46887 346 5899999987654 4899999999999964
No 27
>KOG1226 consensus Integrin beta subunit (N-terminal portion of extracellular region) [Signal transduction mechanisms; Extracellular structures]
Probab=76.35 E-value=1.7 Score=53.01 Aligned_cols=34 Identities=29% Similarity=0.844 Sum_probs=28.3
Q ss_pred EEecccccC----CCCCceeeeeeccCCceEEeeeeeCCCCCCcCCCc
Q 003257 527 VSLERCPKR----CSSHGQCRNAFDASGLTLYSFCACDRDHGGFDCSV 570 (836)
Q Consensus 527 isls~C~~~----Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdCtd 570 (836)
...-.|+.. ||+||+|.- | .|.|++||.|-.|.=
T Consensus 544 CDnfsC~r~~g~lC~g~G~C~C-----G-----~CvC~~GwtG~~C~C 581 (783)
T KOG1226|consen 544 CDNFSCERHKGVLCGGHGRCEC-----G-----RCVCNPGWTGSACNC 581 (783)
T ss_pred ccCcccccccCcccCCCCeEeC-----C-----cEEcCCCCccCCCCC
Confidence 344578877 999999997 5 899999999999863
No 28
>KOG3607 consensus Meltrins, fertilins and related Zn-dependent metalloproteinases of the ADAMs family [Posttranslational modification, protein turnover, chaperones]
Probab=74.86 E-value=1.7 Score=53.00 Aligned_cols=35 Identities=29% Similarity=0.663 Sum_probs=29.7
Q ss_pred EecccccCCCCCceeeeeeccCCceEEeeeeeCCCCCCcCCCcc
Q 003257 528 SLERCPKRCSSHGQCRNAFDASGLTLYSFCACDRDHGGFDCSVE 571 (836)
Q Consensus 528 sls~C~~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdCtd~ 571 (836)
..+-||.+|++||.|-.- + .|+|.+||.+.+|...
T Consensus 624 ~~~~~~~~C~g~GVCnn~----~-----~ChC~~gwapp~C~~~ 658 (716)
T KOG3607|consen 624 NSSCCPTTCNGHGVCNNE----L-----NCHCEPGWAPPFCFIF 658 (716)
T ss_pred cccccccccCCCcccCCC----c-----ceeeCCCCCCCccccc
Confidence 346789999999999863 2 8999999999999864
No 29
>COG5237 PER1 Predicted membrane protein [Function unknown]
Probab=72.11 E-value=6 Score=43.25 Aligned_cols=51 Identities=18% Similarity=0.333 Sum_probs=30.7
Q ss_pred hHHHHH-HHHHHHHHHhhhhhcccccceeeccchhHhHhHhHHHHH----HHHHHHHhhccchh
Q 003257 603 KAFAEW-VLFTASGISSGLYHACDVGTWCALSFNVLQFMDFWLSFM----AVVSTFIYLTTIDE 661 (836)
Q Consensus 603 r~~~Ea-~Vy~fTMffS~fYHACD~g~~Cim~ydvLQf~DF~gSim----SiwvT~I~MA~~~e 661 (836)
+++..+ ++.-.+-..|+.+|.=|.- .=+-||-+.+.+ .+-++++-|-.+..
T Consensus 135 ~~~l~wv~igmlAwi~SsvFHird~~--------iTeklDYF~AgltVLfGfy~~lvrm~~~~~ 190 (319)
T COG5237 135 LYYLQWVYIGMLAWISSSVFHIRDNT--------ITEKLDYFLAGLTVLFGFYMALVRMILIVS 190 (319)
T ss_pred eEEeeHHHHHHHHHHHHhheeeeccc--------hhhhHHHHHhhHHHHHHHHHHHHHHHHhhc
Confidence 355666 6677778899999999862 111355555433 34445555555543
No 30
>KOG4260 consensus Uncharacterized conserved protein [Function unknown]
Probab=67.52 E-value=3.1 Score=45.83 Aligned_cols=40 Identities=25% Similarity=0.576 Sum_probs=29.8
Q ss_pred cccc----cCCCCCceeeeeeccCCceEEeeeeeCCCCCCcCCCccc
Q 003257 530 ERCP----KRCSSHGQCRNAFDASGLTLYSFCACDRDHGGFDCSVEL 572 (836)
Q Consensus 530 s~C~----~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdCtd~s 572 (836)
.+|| +.|+++|+|+=-.+..| ...|.|..||+|.-|.+=-
T Consensus 142 l~Cpggser~C~GnG~C~GdGsR~G---sGkCkC~~GY~Gp~C~~Cg 185 (350)
T KOG4260|consen 142 LQCPGGSERPCFGNGSCHGDGSREG---SGKCKCETGYTGPLCRYCG 185 (350)
T ss_pred ccCCCCCcCCcCCCCcccCCCCCCC---CCcccccCCCCCccccccc
Confidence 3565 68999999985433322 2399999999999998643
No 31
>KOG4243 consensus Macrophage maturation-associated protein [Defense mechanisms]
Probab=66.02 E-value=23 Score=38.64 Aligned_cols=24 Identities=17% Similarity=0.247 Sum_probs=17.9
Q ss_pred hcccCcceeEehhHHHHHHhhhee
Q 003257 767 KLETSQSYWIWHSIWHVSIYTSSF 790 (836)
Q Consensus 767 flET~dnY~y~HSiWHi~Ia~S~~ 790 (836)
|+..+.---+-|-|||.++++++.
T Consensus 255 FFK~DG~ipfAHAIWHLFV~l~A~ 278 (298)
T KOG4243|consen 255 FFKSDGIIPFAHAIWHLFVALAAG 278 (298)
T ss_pred EEecCCceehHHHHHHHHHHHHcc
Confidence 344555566789999999998763
No 32
>PF07645 EGF_CA: Calcium-binding EGF domain; InterPro: IPR001881 A sequence of about forty amino-acid residues found in epidermal growth factor (EGF) has been shown [, , , , , ] to be present in a large number of membrane-bound and extracellular, mostly animal, proteins. Many of these proteins require calcium for their biological function and a calcium-binding site has been found at the N terminus of some EGF-like domains []. Calcium-binding may be crucial for numerous protein-protein interactions. For human coagulation factor IX it has been shown [] that the calcium-ligands form a pentagonal bipyramid. The first, third and fourth conserved negatively charged or polar residues are side chain ligands. The latter is possibly hydroxylated (see aspartic acid and asparagine hydroxylation site) []. A conserved aromatic residue, as well as the second conserved negative residue, are thought to be involved in stabilising the calcium-binding site. As in non-calcium binding EGF-like domains, there are six conserved cysteines and the structure of both types is very similar as calcium-binding induces only strictly local structural changes []. +------------------+ +---------+ | | | | nxnnC-x(3,14)-C-x(3,7)-CxxbxxxxaxC-x(1,6)-C-x(8,13)-Cx | | +------------------+ 'n': negatively charged or polar residue [DEQN] 'b': possibly beta-hydroxylated residue [DN] 'a': aromatic amino acid 'C': cysteine, involved in disulphide bond 'x': any amino acid. ; GO: 0005509 calcium ion binding; PDB: 2VJ3_A 1TOZ_A 1LMJ_A 1UZQ_A 1UZK_A 1UZJ_B 1UZP_A 1EMO_A 1EMN_A 2RR0_A ....
Probab=55.88 E-value=9 Score=30.13 Aligned_cols=25 Identities=28% Similarity=0.670 Sum_probs=21.3
Q ss_pred cCCCCCceeeeeeccCCceEEeeeeeCCCCC
Q 003257 534 KRCSSHGQCRNAFDASGLTLYSFCACDRDHG 564 (836)
Q Consensus 534 ~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~ 564 (836)
+.|..++.|.-.. |.| .|.|++||.
T Consensus 10 ~~C~~~~~C~N~~---Gsy---~C~C~~Gy~ 34 (42)
T PF07645_consen 10 HNCPENGTCVNTE---GSY---SCSCPPGYE 34 (42)
T ss_dssp SSSSTTSEEEEET---TEE---EEEESTTEE
T ss_pred CcCCCCCEEEcCC---CCE---EeeCCCCcE
Confidence 5798899999864 656 899999998
No 33
>smart00051 DSL delta serrate ligand.
Probab=54.25 E-value=8.9 Score=33.36 Aligned_cols=26 Identities=23% Similarity=0.365 Sum_probs=21.3
Q ss_pred cCCCCCceeeeeeccCCceEEeeeeeCCCCCCcCC
Q 003257 534 KRCSSHGQCRNAFDASGLTLYSFCACDRDHGGFDC 568 (836)
Q Consensus 534 ~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdC 568 (836)
+++.+|..|.. .| .|.|.+||.|..|
T Consensus 38 ~d~~~~~~Cd~----~G-----~~~C~~Gw~G~~C 63 (63)
T smart00051 38 DDFFGHYTCDE----NG-----NKGCLEGWMGPYC 63 (63)
T ss_pred ccccCCccCCc----CC-----CEecCCCCcCCCC
Confidence 56778899964 35 7999999999988
No 34
>PF00954 S_locus_glycop: S-locus glycoprotein family; InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=53.96 E-value=28 Score=32.26 Aligned_cols=34 Identities=21% Similarity=0.533 Sum_probs=25.0
Q ss_pred EEEEEecccc--cCCCCCceeeeeeccCCceEEeeeeeCCCCC
Q 003257 524 VMSVSLERCP--KRCSSHGQCRNAFDASGLTLYSFCACDRDHG 564 (836)
Q Consensus 524 ~~sisls~C~--~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~ 564 (836)
..+.-.+.|- +.||.+|.|... . -..|.|-+||.
T Consensus 72 ~~~~p~d~Cd~y~~CG~~g~C~~~--~-----~~~C~Cl~GF~ 107 (110)
T PF00954_consen 72 FWSAPKDQCDVYGFCGPNGICNSN--N-----SPKCSCLPGFE 107 (110)
T ss_pred EEEecccCCCCccccCCccEeCCC--C-----CCceECCCCcC
Confidence 4455557895 899999999642 1 22799999985
No 35
>PF00053 Laminin_EGF: Laminin EGF-like (Domains III and V); InterPro: IPR002049 Laminins [] are the major noncollagenous components of basement membranes that mediate cell adhesion, growth migration, and differentiation. They are composed of distinct but related alpha, beta and gamma chains. The three chains form a cross-shaped molecule that consist of a long arm and three short globular arms. The long arm consist of a coiled coil structure contributed by all three chains and cross-linked by interchain disulphide bonds. Beside different types of globular domains each subunit contains, in its first half, consecutive repeats of about 60 amino acids in length that include eight conserved cysteines []. The tertiary structure [, ] of this domain is remotely similar in its N-terminal to that of the EGF-like module (see PDOC00021 from PROSITEDOC). It is known as a 'LE' or 'laminin-type EGF-like' domain. The number of copies of the LE domain in the different forms of laminins is highly variable; from 3 up to 22 copies have been found. A schematic representation of the topology of the four disulphide bonds in the LE domain is shown below. +-------------------+ +-|-----------+ | +--------+ +-----------------+ | | | | | | | | xxCxCxxxxxxxxxxxCxxxxxxxCxxCxxxxxGxxCxxCxxgaagxxxxxxxxxxxCxx sssssssssssssssssssssssssssssssssss 'C': conserved cysteine involved in a disulphide bond 'a': conserved aromatic residue 'G': conserved glycine (lower case = less conserved) 's': region similar to the EGF-like domain In mouse laminin gamma-1 chain, the seventh LE domain has been shown to be the only one that binds with a high affinity to nidogen []. The binding-sites are located on the surface within the loops C1-C3 and C5-C6 [, ]. Long consecutive arrays of LE domains in laminins form rod-like elements of limited flexibility [], which determine the spacing in the formation of laminin networks of basement membranes [].; PDB: 3TBD_A 3ZYG_B 3ZYI_B 2Y38_A 1KLO_A 1NPE_B 3ZYJ_B 1TLE_A.
Probab=51.52 E-value=9.4 Score=30.74 Aligned_cols=28 Identities=32% Similarity=0.909 Sum_probs=21.4
Q ss_pred CCCCCc----eeeeeeccCCceEEeeeeeCCCCCCcCCCc
Q 003257 535 RCSSHG----QCRNAFDASGLTLYSFCACDRDHGGFDCSV 570 (836)
Q Consensus 535 ~Cg~~G----~C~~l~~~sG~~~ys~C~C~~Gy~GwdCtd 570 (836)
+|.++| .|.. ..| .|.|++||.|..|..
T Consensus 2 ~C~~~~~~~~~C~~---~~G-----~C~C~~~~~G~~C~~ 33 (49)
T PF00053_consen 2 DCNPHGSSSQTCDP---STG-----QCVCKPGTTGPRCDQ 33 (49)
T ss_dssp SSTTCCBCCSSEEE---TCE-----EESBSTTEESTTS-E
T ss_pred cCcCCCCCCCcccC---CCC-----EEeccccccCCcCcC
Confidence 466666 8887 234 999999999999974
No 36
>KOG3879 consensus Predicted membrane protein [Function unknown]
Probab=49.98 E-value=2.3e+02 Score=31.24 Aligned_cols=24 Identities=25% Similarity=0.536 Sum_probs=17.9
Q ss_pred CCeeeEeecCCCCCCCCCCCCCcc
Q 003257 810 GTYELTRQDSMPRGDSEGRERPEV 833 (836)
Q Consensus 810 ~~y~~t~~d~~~r~~~~~~~~~~~ 833 (836)
=+|...|.|.-+-.|.|-+|.|.+
T Consensus 212 isy~th~~d~e~~ee~~~~~~~~i 235 (267)
T KOG3879|consen 212 ISYDTHHEDNEPEEETEVPEEPKI 235 (267)
T ss_pred cceecccccCCCCcccCCCCCcch
Confidence 357777888888888777777765
No 37
>cd00055 EGF_Lam Laminin-type epidermal growth factor-like domain; laminins are the major noncollagenous components of basement membranes that mediate cell adhesion, growth migration, and differentiation; the laminin-type epidermal growth factor-like module occurs in tandem arrays; the domain contains 4 disulfide bonds (loops a-d) the first three resemble epidermal growth factor (EGF); the number of copies of this domain in the different forms of laminins is highly variable ranging from 3 up to 22 copies
Probab=49.48 E-value=14 Score=30.06 Aligned_cols=28 Identities=32% Similarity=0.929 Sum_probs=21.2
Q ss_pred CCCCCce----eeeeeccCCceEEeeeeeCCCCCCcCCCc
Q 003257 535 RCSSHGQ----CRNAFDASGLTLYSFCACDRDHGGFDCSV 570 (836)
Q Consensus 535 ~Cg~~G~----C~~l~~~sG~~~ys~C~C~~Gy~GwdCtd 570 (836)
+|.++|. |.. .+| .|.|+.|+.|..|..
T Consensus 3 ~C~~~g~~~~~C~~---~~G-----~C~C~~~~~G~~C~~ 34 (50)
T cd00055 3 DCNGHGSLSGQCDP---GTG-----QCECKPNTTGRRCDR 34 (50)
T ss_pred cCcCCCCCCccccC---CCC-----EEeCCCcCCCCCCCC
Confidence 4666665 865 245 899999999999963
No 38
>PF12947 EGF_3: EGF domain; InterPro: IPR024731 This entry represents an EGF domain found in the the C terminus of malarial parasite merozoite surface protein 1 [], as well as other proteins.; PDB: 2NPR_A 1N1I_C 1B9W_A 1YO8_A 2RHP_A.
Probab=46.28 E-value=11 Score=29.51 Aligned_cols=28 Identities=21% Similarity=0.456 Sum_probs=20.0
Q ss_pred cCCCCCceeeeeeccCCceEEeeeeeCCCCCCcC
Q 003257 534 KRCSSHGQCRNAFDASGLTLYSFCACDRDHGGFD 567 (836)
Q Consensus 534 ~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~Gwd 567 (836)
..|+.+-+|..... .+ .|.|++||.|-+
T Consensus 6 ~~C~~nA~C~~~~~---~~---~C~C~~Gy~GdG 33 (36)
T PF12947_consen 6 GGCHPNATCTNTGG---SY---TCTCKPGYEGDG 33 (36)
T ss_dssp GGS-TTCEEEE-TT---SE---EEEE-CEEECCS
T ss_pred CCCCCCcEeecCCC---CE---EeECCCCCccCC
Confidence 58999999998643 23 999999999864
No 39
>KOG1219 consensus Uncharacterized conserved protein, contains laminin, cadherin and EGF domains [Signal transduction mechanisms]
Probab=32.79 E-value=33 Score=47.29 Aligned_cols=41 Identities=27% Similarity=0.673 Sum_probs=32.8
Q ss_pred EEEEeccc-ccCCCCCceeeeeeccCCceEEeeeeeCCCCCCcCCCcc
Q 003257 525 MSVSLERC-PKRCSSHGQCRNAFDASGLTLYSFCACDRDHGGFDCSVE 571 (836)
Q Consensus 525 ~sisls~C-~~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdCtd~ 571 (836)
-++.++|| ++.|-.-|+|... ++| | -|.|+.||.|--|-.+
T Consensus 3899 CEi~~epC~snPC~~GgtCip~--~n~-f---~CnC~~gyTG~~Ce~~ 3940 (4289)
T KOG1219|consen 3899 CEIDLEPCASNPCLTGGTCIPF--YNG-F---LCNCPNGYTGKRCEAR 3940 (4289)
T ss_pred cccccccccCCCCCCCCEEEec--CCC-e---eEeCCCCccCceeecc
Confidence 45777899 5899999999985 333 4 7999999999999643
No 40
>PF12658 Ten1: Telomere capping, CST complex subunit; InterPro: IPR024222 Stn1 and Ten1 are DNA-binding proteins with specificity for telomeric DNA substrates and both protect chromosome termini from unregulated resection and regulate telomere length. Stn1 complexes with Ten1 and Cdc13 to function as a telomere-specific replication protein A (RPA)-like complex []. These three interacting proteins associate with the telomeric overhang in budding yeast, whereas a single protein known as Pot1 (protection of telomeres-1) performs this function in fission yeast, and a two-subunit complex consisting of POT1 and TPP1 associates with telomeric ssDNA in humans. S.pombe has Stn1- and Ten1-like proteins that are essential for chromosome end protection. Stn1 orthologues exist in all species that have Pot1, whereas Ten1-like proteins can be found in all fungi. Fission yeast Stn1 and Ten1 localise at telomeres in a manner that correlates with the length of the ssDNA overhang, suggesting that they specifically associate with the telomeric ssDNA. Two separate protein complexes are required for chromosome end protection in fission yeast. Protection of telomeres by multiple proteins with OB-fold domains is conserved in eukaryotic evolution [].; PDB: 3KF8_D 3KF6_B 3K0X_A.
Probab=30.89 E-value=54 Score=32.12 Aligned_cols=48 Identities=25% Similarity=0.474 Sum_probs=27.0
Q ss_pred CCccccccccccccccc---ccceeee---------eeccccCCcceE--EEeecCCCccce
Q 003257 125 SSNELEDIQNEEQCYPM---QKNISVK---------LTNEQISPGAWY--LGFFNGVGAIRT 172 (836)
Q Consensus 125 ~~~~~~~~~~~~qc~p~---~~~~~~~---------l~~~qi~~g~wy--~g~f~~~~~~r~ 172 (836)
++....+....+.+||- ..+.++. ++.+++..|.|+ +|+++|-.+..+
T Consensus 36 ~Y~~~~~~L~l~h~~p~~~~~~~~~v~VdI~~vL~tv~~~~~rvG~WvNV~Gy~~~~~~~~~ 97 (124)
T PF12658_consen 36 SYDTSTGTLTLEHNYPRENDSQPSSVSVDINLVLETVSSEELRVGEWVNVVGYIRGEKPSQT 97 (124)
T ss_dssp EEECCCTEEEEEETCCC---S----EEEE-TTTTTTS-GGGGSTT-EEEEEEEEECTT----
T ss_pred EEecCccEEEEeecCCCCcCCCCceEEEEHHHHhhhcCccceecceEEEEEEEecccccccc
Confidence 34445556666677777 2221121 367789999998 899999997763
No 41
>PF12662 cEGF: Complement Clr-like EGF-like
Probab=30.71 E-value=32 Score=25.15 Aligned_cols=15 Identities=27% Similarity=0.665 Sum_probs=12.5
Q ss_pred eeeeCCCCC----CcCCCc
Q 003257 556 FCACDRDHG----GFDCSV 570 (836)
Q Consensus 556 ~C~C~~Gy~----GwdCtd 570 (836)
.|.|.+||. |-.|.|
T Consensus 3 ~C~C~~Gy~l~~d~~~C~D 21 (24)
T PF12662_consen 3 TCSCPPGYQLSPDGRSCED 21 (24)
T ss_pred EeeCCCCCcCCCCCCcccc
Confidence 799999997 667775
No 42
>KOG1219 consensus Uncharacterized conserved protein, contains laminin, cadherin and EGF domains [Signal transduction mechanisms]
Probab=30.18 E-value=39 Score=46.71 Aligned_cols=38 Identities=29% Similarity=0.654 Sum_probs=30.6
Q ss_pred ecccc-cCCCCCceeeeeeccCCceEEeeeeeCCCCCCcCCCccc
Q 003257 529 LERCP-KRCSSHGQCRNAFDASGLTLYSFCACDRDHGGFDCSVEL 572 (836)
Q Consensus 529 ls~C~-~~Cg~~G~C~~l~~~sG~~~ys~C~C~~Gy~GwdCtd~s 572 (836)
.+.|- +-|+.-|+|.....+ | .|.|.+||.|-.|-++.
T Consensus 3942 i~eCs~n~C~~gg~C~n~~gs---f---~CncT~g~~gr~c~~~~ 3980 (4289)
T KOG1219|consen 3942 ISECSKNVCGTGGQCINIPGS---F---HCNCTPGILGRTCCAEK 3980 (4289)
T ss_pred ccccccccccCCceeeccCCc---e---EeccChhHhcccCcccc
Confidence 45576 789999999987533 3 89999999999997654
No 43
>PF04151 PPC: Bacterial pre-peptidase C-terminal domain; InterPro: IPR007280 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. This domain is normally found at the C terminus of secreted archaeal and bacterial peptidases, the majority of which belong to MEROPS peptidase families M4 (vibriolysin, IPR001570 from INTERPRO), M9A amd M9B (microbial collangenase, IPR002169 from INTERPRO), M28 (aminopeptidase Ap1, IPR007484 from INTERPRO) and S8 (subtilisin family peptidases, IPR000209 from INTERPRO).; GO: 0008233 peptidase activity, 0006508 proteolysis; PDB: 4DY5_B 4DXZ_A 4DY3_B 3JQW_A 3JQX_C 1NQJ_B 1NQD_A 2O8O_A 1WMF_A 1WME_A ....
Probab=23.65 E-value=3.3e+02 Score=23.10 Aligned_cols=64 Identities=20% Similarity=0.302 Sum_probs=41.7
Q ss_pred eEEEEeccCchheeeEEeceeeeeccccCCcCCCCCceEEEEeecCCCCCccccccCC--CCCcceeeecCCCCcceEEE
Q 003257 256 KVFFLDVLGIAEQLIIMAMNVTFSMTQSNNTLNAGGANIVCFARHGAMPSEILHDYSG--DISNGPLIVDSPKVGRWYIT 333 (836)
Q Consensus 256 ~~y~ldV~~~a~~l~i~a~n~~~~~~~s~~~~~~~~~~l~~~~r~~a~P~~~~~d~sg--~~~~c~L~l~sPpwgrW~~v 333 (836)
.+|++++|.-.. ++|++.+-. .+.-|-++...| +.....|+++ ....-.+.+..|.-|+||+.
T Consensus 3 D~y~f~v~ag~~-l~i~l~~~~------------~d~dl~l~~~~g--~~~~~~d~~~~~~~~~~~i~~~~~~~GtYyi~ 67 (70)
T PF04151_consen 3 DYYSFTVPAGGT-LTIDLSGGS------------GDADLYLYDSNG--NSLASYDDSSQSGGNDESITFTAPAAGTYYIR 67 (70)
T ss_dssp EEEEEEESTTEE-EEEEECETT------------SSEEEEEEETTS--SSCEECCCCTCETTSEEEEEEEESSSEEEEEE
T ss_pred EEEEEEEcCCCE-EEEEEcCCC------------CCeEEEEEcCCC--CchhhheecCCCCCCccEEEEEcCCCEEEEEE
Confidence 579999998776 888874432 145566666665 4444445444 23446677788999999887
Q ss_pred E
Q 003257 334 I 334 (836)
Q Consensus 334 i 334 (836)
+
T Consensus 68 V 68 (70)
T PF04151_consen 68 V 68 (70)
T ss_dssp E
T ss_pred E
Confidence 4
No 44
>PRK05420 aquaporin Z; Provisional
Probab=22.62 E-value=4.1e+02 Score=28.38 Aligned_cols=21 Identities=10% Similarity=0.292 Sum_probs=17.0
Q ss_pred cccchhhhHHHHHHHHHHhhh
Q 003257 746 RFRWGFVLVGFAALAMAAISW 766 (836)
Q Consensus 746 rfRw~f~L~Ggi~la~~aI~~ 766 (836)
.+.|+|.+.|++-..++++.+
T Consensus 203 ~~~wvy~vgP~~Ga~laa~~y 223 (231)
T PRK05420 203 EQLWLFWVAPIVGAIIGGLIY 223 (231)
T ss_pred cceEEeehHHHHHHHHHHHHH
Confidence 358999999999877777765
No 45
>KOG4812 consensus Golgi-associated protein/Nedd4 WW domain-binding protein [General function prediction only]
Probab=22.39 E-value=1.1e+02 Score=33.73 Aligned_cols=31 Identities=35% Similarity=0.449 Sum_probs=17.5
Q ss_pred hhHHHHHHHHHhhccCCccchhhHHHHHHHH-HHHHHh
Q 003257 670 VVAILTAMMAITKATRSSNIILVISIGAAGL-LIGLLV 706 (836)
Q Consensus 670 ~~~IL~Al~~~~q~~R~wn~iiPI~i~~lgi-li~Wl~ 706 (836)
++|+++.++..+.+-|.+-+. |+ |+ +|+|+.
T Consensus 176 IGFlltycl~tT~agRYGA~~-----Gf-GLsLikwil 207 (262)
T KOG4812|consen 176 IGFLLTYCLTTTHAGRYGAIS-----GF-GLSLIKWIL 207 (262)
T ss_pred HHHHHHHHHHhhHhhhhhhhh-----cc-chhhheeeE
Confidence 466666666555667776322 22 43 567764
Done!