Query 000135
Match_columns 2087
No_of_seqs 280 out of 1176
Neff 2.9
Searched_HMMs 46136
Date Thu Mar 28 20:03:45 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/000135.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/000135hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 smart00230 CysPc Calpain-like 100.0 1E-70 2.2E-75 621.3 28.9 302 1694-2014 4-317 (318)
2 cd00044 CysPc Calpains, domain 100.0 5.3E-68 1.2E-72 595.5 27.9 304 1695-2005 3-315 (315)
3 KOG0045 Cytosolic Ca2+-depende 100.0 1.6E-66 3.4E-71 627.6 29.0 364 1693-2070 15-406 (612)
4 PF00648 Peptidase_C2: Calpain 100.0 1.4E-66 3.1E-71 576.7 20.0 283 1705-2006 1-297 (298)
5 KOG0045 Cytosolic Ca2+-depende 100.0 2E-31 4.3E-36 323.9 -11.3 598 874-1874 13-610 (612)
6 smart00720 calpain_III calpain 98.4 6.2E-07 1.4E-11 92.2 6.5 51 2013-2064 4-56 (143)
7 cd00214 Calpain_III Calpain, s 98.2 3E-06 6.4E-11 88.7 6.3 52 2013-2064 6-60 (150)
8 PF01067 Calpain_III: Calpain 98.1 4E-06 8.6E-11 85.4 5.3 51 2013-2063 5-58 (147)
9 cd00152 PTX Pentraxins are pla 98.0 5.3E-05 1.1E-09 82.5 11.9 162 1434-1618 32-195 (201)
10 smart00159 PTX Pentraxin / C-r 97.9 0.00011 2.5E-09 80.5 12.1 160 1434-1618 32-195 (206)
11 PF13385 Laminin_G_3: Concanav 97.4 0.00065 1.4E-08 66.6 8.7 80 1498-1594 78-157 (157)
12 PF00354 Pentaxin: Pentaxin fa 97.3 0.00075 1.6E-08 74.3 8.9 159 1434-1618 26-188 (195)
13 cd00110 LamG Laminin G domain; 94.5 0.23 5E-06 50.1 9.7 110 1432-1559 19-129 (151)
14 smart00210 TSPN Thrombospondin 94.3 0.25 5.4E-06 53.9 10.1 104 1434-1554 53-163 (184)
15 smart00282 LamG Laminin G doma 93.5 0.55 1.2E-05 47.5 10.0 109 1435-1560 3-112 (135)
16 smart00560 LamGL LamG-like jel 91.5 0.72 1.6E-05 47.7 8.0 85 1434-1532 2-88 (133)
17 cd02619 Peptidase_C1 C1 Peptid 85.6 1.5 3.2E-05 47.2 5.7 49 1924-1998 168-218 (223)
18 KOG1029 Endocytic adaptor prot 79.1 3.2 6.9E-05 54.4 6.0 33 1011-1043 72-104 (1118)
19 PF02210 Laminin_G_2: Laminin 78.6 4.4 9.6E-05 39.6 5.7 62 1497-1562 46-107 (128)
20 cd02248 Peptidase_C1A Peptidas 68.9 14 0.00031 40.2 7.2 43 1925-1993 156-198 (210)
21 PRK12438 hypothetical protein; 66.5 8.3 0.00018 52.5 5.7 43 894-936 115-174 (991)
22 PF03699 UPF0182: Uncharacteri 62.5 13 0.00028 49.8 6.2 62 865-926 62-154 (774)
23 PF02057 Glyco_hydro_59: Glyco 60.0 25 0.00055 46.4 8.1 91 1431-1534 542-638 (669)
24 KOG4326 Mitochondrial F1F0-ATP 58.0 15 0.00032 36.9 4.2 17 1242-1258 13-29 (81)
25 TIGR00805 oat sodium-independe 57.3 26 0.00056 45.4 7.6 94 919-1014 328-436 (633)
26 PF00112 Peptidase_C1: Papain 54.5 35 0.00075 37.0 6.9 44 1925-1994 163-206 (219)
27 PF09323 DUF1980: Domain of un 46.5 55 0.0012 36.5 7.0 57 64-124 4-60 (182)
28 PTZ00334 trans-sialidase; Prov 45.4 30 0.00064 46.6 5.5 77 1504-1598 642-724 (780)
29 PF07946 DUF1682: Protein of u 44.1 33 0.00072 41.4 5.3 10 1332-1341 305-314 (321)
30 COG1390 NtpE Archaeal/vacuolar 43.9 2.2E+02 0.0047 32.9 11.2 113 1255-1384 17-133 (194)
31 PF04156 IncA: IncA protein; 41.5 19 0.00042 39.5 2.6 22 925-946 11-32 (191)
32 PF11770 GAPT: GRB2-binding ad 41.5 18 0.00038 40.4 2.2 14 958-971 21-34 (158)
33 PF00054 Laminin_G_1: Laminin 41.5 34 0.00074 35.6 4.2 51 1472-1530 26-76 (131)
34 cd08045 TAF4 TATA Binding Prot 41.0 15 0.00032 41.9 1.6 44 1353-1420 166-209 (212)
35 PRK00068 hypothetical protein; 41.0 29 0.00062 47.6 4.5 40 895-934 114-170 (970)
36 PTZ00266 NIMA-related protein 40.9 45 0.00097 46.2 6.2 16 823-838 227-242 (1021)
37 KOG1029 Endocytic adaptor prot 40.2 37 0.00081 45.3 5.1 43 1290-1332 356-398 (1118)
38 PLN02316 synthase/transferase 40.2 45 0.00097 46.2 6.1 45 1309-1353 270-329 (1036)
39 cd02620 Peptidase_C1A_Cathepsi 38.7 69 0.0015 36.7 6.5 27 1927-1953 184-210 (236)
40 PF05297 Herpes_LMP1: Herpesvi 37.7 11 0.00024 45.4 0.0 52 949-1002 107-158 (381)
41 PF09472 MtrF: Tetrahydrometha 36.7 12 0.00025 36.7 0.0 47 796-842 17-64 (64)
42 cd02698 Peptidase_C1A_Cathepsi 36.6 83 0.0018 36.2 6.6 42 1927-1994 178-220 (239)
43 KOG1144 Translation initiation 35.1 92 0.002 42.1 7.3 17 1521-1537 397-413 (1064)
44 PF09323 DUF1980: Domain of un 34.9 70 0.0015 35.7 5.6 65 956-1020 4-83 (182)
45 PF05875 Ceramidase: Ceramidas 32.9 50 0.0011 38.4 4.2 143 806-969 14-159 (262)
46 COG4870 Cysteine protease [Pos 31.9 45 0.00098 41.6 3.8 49 1923-1997 260-318 (372)
47 PF14023 DUF4239: Protein of u 31.9 1.2E+02 0.0027 33.9 6.9 31 979-1010 167-197 (209)
48 PF09586 YfhO: Bacterial membr 31.0 1.3E+02 0.0027 40.1 7.9 24 846-869 214-238 (843)
49 PF09991 DUF2232: Predicted me 30.9 51 0.0011 37.6 3.9 87 915-1002 199-288 (290)
50 COG0815 Lnt Apolipoprotein N-a 30.7 1.2E+02 0.0025 39.4 7.2 77 99-180 97-185 (518)
51 KOG2341 TATA box binding prote 30.1 54 0.0012 42.8 4.2 27 1055-1081 189-215 (563)
52 PF12065 DUF3545: Protein of u 29.1 24 0.00052 34.3 0.7 10 1333-1342 23-32 (59)
53 TIGR00570 cdk7 CDK-activating 29.0 85 0.0018 38.6 5.4 104 1204-1329 57-164 (309)
54 PF04405 ScdA_N: Domain of Unk 28.6 38 0.00083 32.1 2.0 33 511-543 11-47 (56)
55 cd06899 lectin_legume_LecRK_Ar 28.4 2E+02 0.0043 33.4 7.9 37 1492-1528 150-186 (236)
56 PF05154 TM2: TM2 domain; Int 28.0 20 0.00043 32.9 0.0 33 290-326 3-38 (51)
57 PF11877 DUF3397: Protein of u 26.9 90 0.0019 32.9 4.5 96 895-994 9-109 (116)
58 PF14402 7TM_transglut: 7 tran 26.3 88 0.0019 38.5 4.8 54 943-1001 147-207 (313)
59 PF06439 DUF1080: Domain of Un 26.3 2.2E+02 0.0049 30.4 7.4 102 1415-1529 38-149 (185)
60 TIGR00917 2A060601 Niemann-Pic 26.0 35 0.00077 47.8 1.8 79 956-1035 640-743 (1204)
61 PF02460 Patched: Patched fami 25.3 92 0.002 41.6 5.2 53 955-1007 282-348 (798)
62 TIGR02916 PEP_his_kin putative 25.0 39 0.00085 43.8 1.8 36 886-922 58-93 (679)
63 PF13801 Metal_resist: Heavy-m 24.3 2.6E+02 0.0057 27.5 7.0 20 1287-1306 43-62 (125)
64 PF11911 DUF3429: Protein of u 23.9 87 0.0019 34.0 3.9 44 852-895 35-78 (142)
65 cd01951 lectin_L-type legume l 23.6 3.2E+02 0.007 30.9 8.3 50 1505-1555 154-203 (223)
66 KOG3011 Ubiquitin-conjugating 23.2 2.1E+02 0.0045 34.7 6.8 115 852-984 83-225 (293)
67 PRK15097 cytochrome d terminal 22.9 2.4E+02 0.0051 37.1 7.9 91 948-1074 393-491 (522)
68 PLN00122 serine/threonine prot 22.8 93 0.002 35.4 3.9 22 1323-1344 142-163 (170)
69 PF15412 Nse4-Nse3_bdg: Bindin 22.4 62 0.0013 30.5 2.1 28 182-209 18-45 (56)
70 KOG3583 Uncharacterized conser 22.3 1.6E+02 0.0034 35.1 5.6 122 1233-1363 38-185 (279)
71 PF02387 IncFII_repA: IncFII R 22.3 98 0.0021 37.5 4.2 87 1247-1349 159-251 (281)
72 PRK10263 DNA translocase FtsK; 21.9 58 0.0013 46.1 2.6 30 772-805 23-52 (1355)
73 PF04123 DUF373: Domain of unk 21.9 49 0.0011 40.9 1.8 138 923-1074 161-320 (344)
74 KOG4661 Hsp27-ERE-TATA-binding 21.8 1.2E+02 0.0026 39.7 5.0 30 1313-1342 626-655 (940)
75 PRK11588 hypothetical protein; 21.5 2.6E+02 0.0056 36.6 7.8 46 890-952 172-217 (506)
76 PTZ00358 hypothetical protein; 20.4 67 0.0015 39.9 2.4 93 811-911 247-343 (367)
77 KOG2751 Beclin-like protein [S 20.2 2E+02 0.0043 37.0 6.2 63 1283-1353 185-247 (447)
78 PRK09776 putative diguanylate 20.1 3.3E+02 0.0072 36.9 8.8 191 821-1022 47-256 (1092)
No 1
>smart00230 CysPc Calpain-like thiol protease family. Calpain-like thiol protease family (peptidase family C2). Calcium activated neutral protease (large subunit).
Probab=100.00 E-value=1e-70 Score=621.27 Aligned_cols=302 Identities=41% Similarity=0.812 Sum_probs=267.3
Q ss_pred HHHHHHHcCCCceecCCCCCCCCCcccCCCCCCcccccccccccccccccccccCCCceeecCCCCCCCcccCCCCCchH
Q 000135 1694 VKEALSARGERQFTDHEFPPDDQSLYVDPGNPPSKLQVVAEWMRPSEIVKESRLDCQPCLFSGAVNPSDVCQGRLGDCWF 1773 (2087)
Q Consensus 1694 IKE~cl~rGeklFeDPEFPPndsSLy~Dp~~PpsKlq~vIeWKRPsEI~~e~k~~snP~LF~dgISP~DIkQGsLGDCWF 1773 (2087)
+.+.|.+++ .+|+|++|||++.||+.++..+ ..++|+||+|+++ +|.+|.++++|.||+||.+|||||
T Consensus 4 i~~~c~~~~-~~f~D~~Fpp~~~sl~~~~~~~-----~~~~W~Rp~e~~~------~~~~~~~~i~~~di~QG~lgDC~~ 71 (318)
T smart00230 4 LRQYCKESG-TLFEDPLFPANNGSLFFSQRQR-----KFVVWKRPHEIFE------NPPFIVGGASRTDICQGVLGDCWL 71 (318)
T ss_pred HHHHHHHcC-CCccCCCCCCCcCccccCCCCC-----CCcEEECcHHHcC------CCEEEeCCCChhhccCcccccHHH
Confidence 455677665 6999999999999998765422 2479999999986 478898999999999999999999
Q ss_pred HHHHHHHhccccccccccccc----cCCCCcEEEEEeeCCEEEEEEEeccccCCCCCceEEeecCCCCchhHHHHHHHHH
Q 000135 1774 LSAVAVLTEVSQISEVIITPE----YNEEGIYTVRFCIQGEWVPVVVDDWIPCESPGKPAFATSKKGHELWVSILEKAYA 1849 (2087)
Q Consensus 1774 LAALAALAE~PrLle~fI~Pe----yNe~GIY~VRL~iNGeWReVVVDDrLPc~~nGKPLFArSsd~nELWpSLLEKAYA 1849 (2087)
+|||++|+++|.+++.++++. .|+.|+|+||||+||+|+.|+|||+||+.. |+++|+++.+++|+|++|||||||
T Consensus 72 lsal~~la~~~~~i~~if~~~~~~~~~~~G~y~vrl~~~G~w~~V~VDd~lP~~~-~~~~~~~~~~~~e~W~~LLEKAyA 150 (318)
T smart00230 72 LAALASLTLREKLLDRVIPHDQEFSENYAGIFHFRFWRFGKWVDVVIDDRLPTYN-GELVFMHSNSRNEFWSALLEKAYA 150 (318)
T ss_pred HHHHHHHHhCHHHHhheEeCCcccccccCCEEEEEEEECCEEEEEEecCCCeeeC-CceEEEEeCCCCcchhHHHHHHHH
Confidence 999999999998888777532 468999999999999999999999999964 569999999999999999999999
Q ss_pred HhcCCcccccCCChhhhhhhcCCCcceEEeCCchhhhhccchhHHHHHHHHHhcCCCEEEecCCCCC---CccccccCcc
Q 000135 1850 KLHGSYEALEGGLVQDALVDLTGGAGEEIDMRSAQAQIDLASGRLWSQLLRFKQEGFLLGAGSPSGS---DVHISSSGIV 1926 (2087)
Q Consensus 1850 KLhGSYEALeGGnpsEALqDLTGGP~E~IDL~sa~aq~DldsdeLWk~Llkalk~G~LMgcSTPsgS---Deeves~GLV 1926 (2087)
|+||||++|.||++.+||++|||++++.+++++.. .+.+++|+.|.++.++|++|+|+++..+ +...++.||+
T Consensus 151 K~~GsY~~i~gg~~~~al~~LTG~~~~~i~l~~~~----~~~~~~w~~l~~~~~~g~lv~~~t~~~~~~~~~~~~~~GLv 226 (318)
T smart00230 151 KLNGCYEALKGGSTTEALEDLTGGVAESIDLKEAS----KDPDNLFEDLFKAFERGSLMGCSIGAGTAVEEEEQKDCGLV 226 (318)
T ss_pred HHcCCCcccCCCCHHHHHHHhcCCCeEEEEccccc----CCHHHHHHHHHHHHhCCCeEEEEcCCCCcchhhhhhhcCcc
Confidence 99999999999999999999999999999988643 2467899999999999999999987553 3445689999
Q ss_pred cCceeEEEEEEEECCEE--EEEEecCCCCCccccCCCCCCCcccc---hHHhhhhcCCCCCCCCeEEEehhhhhhcccce
Q 000135 1927 QGHAYSILQVREVDGHK--LVQIRNPWANEVEWNGPWSDSSPEWT---DRMKHKLKHVPQSKDGIFWMSWQDFQIHFRSI 2001 (2087)
Q Consensus 1927 sGHAYSVLDVrEVdG~R--LVRLRNPWG~~~EWKGdWSD~S~eWT---eeLKkkL~~~~~sDDGeFWMSfEDFLkyFssL 2001 (2087)
++|||+|++++++++++ ||+|||||| ..||+|+|||+|++|+ +++++++++. ..+||+|||+|+||++||+++
T Consensus 227 ~~HaYsVl~v~~~~~~~~~Ll~lrNPWg-~~eW~G~wsd~s~~W~~~~~~~~~~l~~~-~~~dG~FWM~~~df~~~F~~~ 304 (318)
T smart00230 227 KGHAYSVTDVREVQGRRQELLRLRNPWG-QVEWNGPWSDDSPEWRSVSASEKKNLGLT-FDDDGEFWMSFEDFLRHFDKV 304 (318)
T ss_pred cCccEEEEEEEEEecCCeEEEEEECCCC-CCCcCCCCCCCCccccccCHHHHHHhCCC-CCCCCEEEEEhHHHHhhCCeE
Confidence 99999999999998766 999999999 5899999999999999 6678888764 469999999999999999999
Q ss_pred eEeeEcCCCCcee
Q 000135 2002 YVCRVYPSEMRYS 2014 (2087)
Q Consensus 2002 yICrL~Pds~ryr 2014 (2087)
+||++.|+.+.|+
T Consensus 305 ~vc~~~~~~~~~r 317 (318)
T smart00230 305 EICNLNPDSLEER 317 (318)
T ss_pred EEeccCCcccccc
Confidence 9999999987664
No 2
>cd00044 CysPc Calpains, domains IIa, IIb; calcium-dependent cytoplasmic cysteine proteinases, papain-like. Functions in cytoskeletal remodeling processes, cell differentiation, apoptosis and signal transduction.
Probab=100.00 E-value=5.3e-68 Score=595.47 Aligned_cols=304 Identities=47% Similarity=0.846 Sum_probs=261.3
Q ss_pred HHHHHHcCCCceecCCCCCCCCCcccCCCCCCcccccccccccccccccccccCCCceeecCCCCCCCcccCCCCCchHH
Q 000135 1695 KEALSARGERQFTDHEFPPDDQSLYVDPGNPPSKLQVVAEWMRPSEIVKESRLDCQPCLFSGAVNPSDVCQGRLGDCWFL 1774 (2087)
Q Consensus 1695 KE~cl~rGeklFeDPEFPPndsSLy~Dp~~PpsKlq~vIeWKRPsEI~~e~k~~snP~LF~dgISP~DIkQGsLGDCWFL 1774 (2087)
.+.|.+.+ .+|+|++|||+++|++.++..+..+....++|+||+|+++.... .+|.+|.++++|.||+||.+|||||+
T Consensus 3 ~~~c~~~~-~~f~D~~Fpp~~~s~~~~~~~~~~~~~~~~~W~Rp~~~~~~~~~-~~~~~~~~~~~~~dI~QG~lgDC~~l 80 (315)
T cd00044 3 LQICLLSG-VLFEDPDFPPNDSSLGFDDSLSNGQPKKVIEWKRPSEIFADDGN-SNPRLFVNGASPSDVCQGILGDCWFL 80 (315)
T ss_pred HHHHHHcC-CCccCCCCCCCccccccccccccccCcCcceEECcHHHhCcccC-CCCEEEeCCCChhhcccCcccchHHH
Confidence 45566665 69999999999999987543333334556799999999975322 46899999999999999999999999
Q ss_pred HHHHHHhccccccccccccc-c---CCCCcEEEEEeeCCEEEEEEEeccccCCCCCceEEeecCCCCchhHHHHHHHHHH
Q 000135 1775 SAVAVLTEVSQISEVIITPE-Y---NEEGIYTVRFCIQGEWVPVVVDDWIPCESPGKPAFATSKKGHELWVSILEKAYAK 1850 (2087)
Q Consensus 1775 AALAALAE~PrLle~fI~Pe-y---Ne~GIY~VRL~iNGeWReVVVDDrLPc~~nGKPLFArSsd~nELWpSLLEKAYAK 1850 (2087)
|||++|+++|.+++.++++. . ++.|+|+||||+||+|+.|+|||+||+..++ |+|+++.+.+|+|++||||||||
T Consensus 81 saL~~la~~~~~i~~lf~~~~~~~~~~~G~y~v~l~~~G~w~~V~VDD~lP~~~~~-~~~~~s~~~~e~W~~LlEKAyAK 159 (315)
T cd00044 81 AALAALAERPELLKRVIPPDQSFEENYAGIYHFRFWKNGEWVEVVIDDRLPTSNGG-LLFMHSRDRNELWVALLEKAYAK 159 (315)
T ss_pred HHHHHHHcCHHHHhheEcCCcccccCcCcEEEEEEEECCEEEEEEecCCCeecCCc-eEEEEECCCCeEcHHHHHHHHHh
Confidence 99999999998777766543 3 6899999999999999999999999997655 99999988899999999999999
Q ss_pred hcCCcccccCCChhhhhhhcCCCcceEEeCCchhhhhccchhHHHHHHHHHhcCCCEEEecCCCCCCcc-ccccCcccCc
Q 000135 1851 LHGSYEALEGGLVQDALVDLTGGAGEEIDMRSAQAQIDLASGRLWSQLLRFKQEGFLLGAGSPSGSDVH-ISSSGIVQGH 1929 (2087)
Q Consensus 1851 LhGSYEALeGGnpsEALqDLTGGP~E~IDL~sa~aq~DldsdeLWk~Llkalk~G~LMgcSTPsgSDee-ves~GLVsGH 1929 (2087)
+||||++|.||++.+||++|||++++.+++++.... ...+++|+.|.++.+++++|+|+|+...+.. .+..||+.+|
T Consensus 160 ~~GsY~~i~gg~~~~al~~LTG~~~~~i~~~~~~~~--~~~~~~~~~l~~~~~~~~lv~~~t~~~~~~~~~~~~Gl~~~H 237 (315)
T cd00044 160 LHGSYEALVGGNTAEALEDLTGGPTERIDLKSADAS--SGDNDLFALLLSFLQGGSLIGCSTGSRSEEEARTANGLVKGH 237 (315)
T ss_pred hcCCccccCCCCHHHHHHHhhCCCcEEEEccccccc--cCHHHHHHHHHHHhhCCCEEEEEcCCCCcchhhccCCcccCc
Confidence 999999999999999999999999999998865321 2467899999999999999999998654432 5689999999
Q ss_pred eeEEEEEEEEC--CEEEEEEecCCCCCccccCCCCCCCcccch--HHhhhhcCCCCCCCCeEEEehhhhhhcccceeEee
Q 000135 1930 AYSILQVREVD--GHKLVQIRNPWANEVEWNGPWSDSSPEWTD--RMKHKLKHVPQSKDGIFWMSWQDFQIHFRSIYVCR 2005 (2087)
Q Consensus 1930 AYSVLDVrEVd--G~RLVRLRNPWG~~~EWKGdWSD~S~eWTe--eLKkkL~~~~~sDDGeFWMSfEDFLkyFssLyICr 2005 (2087)
||+|+++++++ |+|||+||||||. .||+|+|||+|++|+. ..++.+. ....+||+|||+|+||++||+++++|+
T Consensus 238 aY~Vl~~~~~~~~~~~lv~lrNPWg~-~~w~G~ws~~~~~w~~~~~~~~~~~-~~~~~dG~Fwm~~~df~~~F~~~~vc~ 315 (315)
T cd00044 238 AYSVLDVREVQEEGLRLLRLRNPWGV-GEWWGGWSDDSSEWWVIDAERKKLL-LSGKDDGEFWMSFEDFLRNFDGLYVCN 315 (315)
T ss_pred ceEEeEEEEEccCceEEEEecCCccC-CCccCCCCCCCchhccChHHHHHhc-CCCCCCCEEEEEhHHhheeeCeEEEeC
Confidence 99999999998 8999999999996 7999999999999963 2333333 346799999999999999999999994
No 3
>KOG0045 consensus Cytosolic Ca2+-dependent cysteine protease (calpain), large subunit (EF-Hand protein superfamily) [Posttranslational modification, protein turnover, chaperones; Signal transduction mechanisms]
Probab=100.00 E-value=1.6e-66 Score=627.61 Aligned_cols=364 Identities=40% Similarity=0.698 Sum_probs=305.3
Q ss_pred HHHHHHHHcCCCceecCCCCCCCCCcccCCCCCCcccccccccccccccccccccCCCceeecCCCCCCCcccCCCCCch
Q 000135 1693 AVKEALSARGERQFTDHEFPPDDQSLYVDPGNPPSKLQVVAEWMRPSEIVKESRLDCQPCLFSGAVNPSDVCQGRLGDCW 1772 (2087)
Q Consensus 1693 aIKE~cl~rGeklFeDPEFPPndsSLy~Dp~~PpsKlq~vIeWKRPsEI~~e~k~~snP~LF~dgISP~DIkQGsLGDCW 1772 (2087)
.+++.|...+ ..|+|++|||+++|++.+...|..+. ..+.|+||+|++. +|++|.+++++.||+||.+||||
T Consensus 15 ~~~~~cl~~~-~~F~D~~FP~~~~Sl~~~~~~p~~~~-~~i~W~RP~ei~~------~p~~i~~~~~~~di~Qg~lgdCw 86 (612)
T KOG0045|consen 15 RLRRDCLPAK-SLFVDALFPAADSSLFYKLSTPLAQF-SDIVWKRPQEICA------NPRLIVDGPSRFDVKQGLLGDCW 86 (612)
T ss_pred HHHHHHhhcC-CcccccCCCCCCccccccccCCCccc-ccceecCcccccC------CCCeecCCCCcceeEEeeecchH
Confidence 3556677665 58999999999999998766555332 4579999999875 58899999999999999999999
Q ss_pred HHHHHHHHhcccccccccccc----ccCCCCcEEEEEeeCCEEEEEEEeccccCCCCCceEEeecCCCCchhHHHHHHHH
Q 000135 1773 FLSAVAVLTEVSQISEVIITP----EYNEEGIYTVRFCIQGEWVPVVVDDWIPCESPGKPAFATSKKGHELWVSILEKAY 1848 (2087)
Q Consensus 1773 FLAALAALAE~PrLle~fI~P----eyNe~GIY~VRL~iNGeWReVVVDDrLPc~~nGKPLFArSsd~nELWpSLLEKAY 1848 (2087)
||||+|+||.++.++.+++++ .+++.|+|+||||++|+|+.|+|||+|||. +|+..|+++..++|+|++||||||
T Consensus 87 ~laA~a~la~~~~ll~~vip~~~~~~~~yaGif~f~~w~~G~W~~VvIDD~LP~~-~~~~~~~~s~~~~efW~aLlEKAy 165 (612)
T KOG0045|consen 87 FLAACAALALRPELLDKVIPQDQSFQENYAGIFHFRFWQNGEWVEVVIDDRLPTS-NGGLLFSHSSGKNEFWAALLEKAY 165 (612)
T ss_pred HHHHHHHhhcCHHHHHhccCCCcccccccceEEEEEEEeCCeEEEEEeeeecceE-cCCEEEEeecCCceeHHHHHHHHH
Confidence 999999999999999888873 268999999999999999999999999996 577889999888999999999999
Q ss_pred HHhcCCcccccCCChhhhhhhcCCCcceEEeCCchhhhhccchhHHHHHHHHHhcCCCEEEecCCC-C-C-C--cccccc
Q 000135 1849 AKLHGSYEALEGGLVQDALVDLTGGAGEEIDMRSAQAQIDLASGRLWSQLLRFKQEGFLLGAGSPS-G-S-D--VHISSS 1923 (2087)
Q Consensus 1849 AKLhGSYEALeGGnpsEALqDLTGGP~E~IDL~sa~aq~DldsdeLWk~Llkalk~G~LMgcSTPs-g-S-D--eeves~ 1923 (2087)
||++|||+++.||...+|+++|||+++|.+++++.... +.+ +++..+.+..++|.+++|++.. + . + +....+
T Consensus 166 aKl~GsY~~l~gg~~~~a~~~lTG~~~e~~~l~~~~~~-~~~--~l~~~~~~~~~~~~~l~c~~~~~~~~~~~~~~~~~~ 242 (612)
T KOG0045|consen 166 AKLLGSYEALHGGSTIDALVDLTGGVTEPFDLNKTPKS-FKN--NLVWALLKSAHRGSLLLCSIESKDPTEEEEEAKLRN 242 (612)
T ss_pred HHHhCcccCCCCCchhhHHHhccCCccceeEcccCcch-hHH--HHHHHHHHhhhccCceeeeccccccchhHHHHHhhc
Confidence 99999999999999999999999999999999875421 111 4455555666666666666532 2 1 2 235789
Q ss_pred CcccCceeEEEEEEEECC----EEEEEEecCCCCCccccCCCCCCCcccchHHhhhhcCCC--CCCCCeEEEehhhhhhc
Q 000135 1924 GIVQGHAYSILQVREVDG----HKLVQIRNPWANEVEWNGPWSDSSPEWTDRMKHKLKHVP--QSKDGIFWMSWQDFQIH 1997 (2087)
Q Consensus 1924 GLVsGHAYSVLDVrEVdG----~RLVRLRNPWG~~~EWKGdWSD~S~eWTeeLKkkL~~~~--~sDDGeFWMSfEDFLky 1997 (2087)
||+++|||+|++++++++ ++|+||||||| +.||||+|||++++|...++.++.... ..+||+|||+++||+++
T Consensus 243 gL~~~HaYsit~~~~~~~~~~~~~lirlrNPwg-~~~W~G~wsd~~~~W~~v~~~~~~~~~~~~~~dGeFWms~~dF~~~ 321 (612)
T KOG0045|consen 243 GLVKGHAYAITDVREVQGRGGKHRLIRLRNPWG-ESEWNGPWSDGSEEWHLVDKSKLSELGRQPLDDGEFWMSFDDFLRE 321 (612)
T ss_pred CccccccEEEEEEEEeecccccceeEEecCCcC-CceeccccccCCcchhhhCHHHHhhcccccccCCCeeeeHHHHHhh
Confidence 999999999999999998 99999999999 589999999999999987765544221 26899999999999999
Q ss_pred ccceeEeeEcCCCC---------ceeeccee--e-ccCCCCCCCC-CCCCcCCeEEEEecCCCCCCCEEEEEEeccCCcc
Q 000135 1998 FRSIYVCRVYPSEM---------RYSVHGQW--R-GYSAGGCQDY-ASWNQNPQFRLRASGSDASFPIHVFITLTQSRFY 2064 (2087)
Q Consensus 1998 FssLyICrL~Pds~---------ryrVhGeW--r-GsSAGGc~n~-~SF~~NPQF~LeVtssD~sep~eVlISLsQkr~Y 2064 (2087)
|..++||++.++.. ....+|.| . +.++|||.++ ++|++||||.+.+..++. ..+.+++.++|+...
T Consensus 322 F~~~~vC~~~~~~~~~~~~~~~~~~~~~~~w~~~~~~t~ggc~~~~~tF~~npq~~~~~~~~~~-~~~~~v~~~~q~~~~ 400 (612)
T KOG0045|consen 322 FDSLTVCRLRPDWLESRNQLQWVKLSLDGEWELARGVTAGGCRNSVDTFDRNPQYILAVRKPTK-SLCAVVLALFQKTRR 400 (612)
T ss_pred CCeEeecCCCcchhhhhheeeeeeeecCCccceeecccCCCCccCcccccCCceEEEEecCCCc-cceEEEEEeeccccc
Confidence 99999999988854 13578999 3 6789999998 799999999999986553 568999999998876
Q ss_pred cceeec
Q 000135 2065 DVLYWD 2070 (2087)
Q Consensus 2065 s~L~~~ 2070 (2087)
...+..
T Consensus 401 ~~~~~~ 406 (612)
T KOG0045|consen 401 GERSFG 406 (612)
T ss_pred cccccc
Confidence 555444
No 4
>PF00648 Peptidase_C2: Calpain family cysteine protease This is family C2 in the peptidase classification. ; InterPro: IPR001300 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This group of cysteine peptidases belong to the MEROPS peptidase family C2 (calpain family, clan CA). A type example is calpain, which is an intracellular protease involved in many important cellular functions that are regulated by calcium []. The protein is a complex of 2 polypeptide chains (light and heavy), with three known forms in mammals [, ]: a highly calcium-sensitive (i.e., micro-molar range) form known as mu-calpain, mu-CANP or calpain I; a form sensitive to calcium in the milli-molar range, known as m-calpain, m-CANP or calpain II; and a third form, known as p94, which is found in skeletal muscle only []. All forms have identical light but different heavy chains. Both mu- and m-calpain are heterodimers containing an identical 28kDa subunit and an 80kDa subunit that shares 55-65% sequence homology between the two proteases [, ]. The crystallographic structure of m-calpain reveals six "domains" in the 80kDa subunit: A 19-amino acid NH2-terminal sequence; Active site domain IIa; Active site domain IIb. Domain 2 shows low levels of sequence similarity to papain; although the catalytic His has not been located by biochemical means, it is likely that calpain and papain are related []. Domain III; An 18-amino acid extended sequence linking domain III to domain IV; Domain IV, which resembles the penta EF-hand family of polypeptides, binds calcium and regulates activity []. />]. Ca2+-binding causes a rearrangement of the protein backbone, the net effect of which is that a Trp side chain, which acts as a wedge between catalytic domains IIa and IIb in the apo state, moves away from the active site cleft allowing for the proper formation of the catalytic triad []. Calpain-like mRNAs have been identified in other organisms including bacteria, but the molecules encoded by these mRNAs have not been isolated, so little is known about their properties. How calpain activity is regulated in these organisms cells is still unclear In metazoans, the activity of calpain is controlled by a single proteinase inhibitor, calpastatin (IPR001259 from INTERPRO). The calpastatin gene can produce eight or more calpastatin polypeptides ranging from 17 to 85 kDa by use of different promoters and alternative splicing events. The physiological significance of these different calpastatins is unclear, although all bind to three different places on the calpain molecule; binding to at least two of the sites is Ca2+ dependent. The calpains ostensibly participate in a variety of cellular processes including remodelling of cytoskeletal/membrane attachments, different signal transduction pathways, and apoptosis. Deregulated calpain activity following loss of Ca2+ homeostasis results in tissue damage in response to events such as myocardial infarcts, stroke, and brain trauma []. Calpains are a family of cytosolic cysteine proteinases (see PDOC00126 from PROSITEDOC). Members of the calpain family are believed to function in various biological processes, including integrin-mediated cell migration, cytoskeletal remodeling, cell differentiation and apoptosis [, ]. The calpain family includes numerous members from C. elegans to mammals and with homologues in yeast and bacteria. The best characterised members are the m- and mu-calpains, both proteins are heterodimer composed of a large catalytic subunit and a small regulatory subunit. The large subunit comprises four domains (dI-dIV) while the small subunit has two domains (dV-dVI). Domain dI is a short region cleaved by autolysis, dII is the catalytic core, dIII is a C2-like domain, dIV consists of five calcium binding EF-hand motifs []. The crystal structure of calpain has been solved [, ]. The catalytic region consists of two distinct structural domains (dIIa and dIIb). dIIa contains a central helix flanked on three faces by a cluster of alpha-helices and is entirely unrelated to the corresponding domain in the typical thiol proteinases. The fold of dIIb is similar to the corresponding domain in other cysteine proteinases and contains two three-stranded anti-parallel beta-sheets. The catalytic triad residues (C,H,N) are located in dIIa and dIIb. The activation of the domain is dependent on the binding of two calcium atoms in two non EF-hand calcium binding sites located in the catalytic core, one close to the Cys active site in dIIa and one at the end of dIIb. Calcium-binding induced conformational changes in the catalytic domain which align the active site [][]. The profile covers the whole catalytic domain.; GO: 0004198 calcium-dependent cysteine-type endopeptidase activity, 0006508 proteolysis, 0005622 intracellular; PDB: 2NQA_A 1KFU_L 1KFX_L 1QXP_B 2R9C_A 1TL9_A 2G8E_A 1KXR_B 2G8J_A 2NQG_A ....
Probab=100.00 E-value=1.4e-66 Score=576.66 Aligned_cols=283 Identities=51% Similarity=0.960 Sum_probs=228.1
Q ss_pred ceecCCCCCCCCCcccCCCCCCcccccccccccccccccccccCCCceeecCCCCCCCcccCCCCCchHHHHHHHHhccc
Q 000135 1705 QFTDHEFPPDDQSLYVDPGNPPSKLQVVAEWMRPSEIVKESRLDCQPCLFSGAVNPSDVCQGRLGDCWFLSAVAVLTEVS 1784 (2087)
Q Consensus 1705 lFeDPEFPPndsSLy~Dp~~PpsKlq~vIeWKRPsEI~~e~k~~snP~LF~dgISP~DIkQGsLGDCWFLAALAALAE~P 1784 (2087)
+|+||+|||+++||+.++..+ ..++|+||+|+++ +|++|.+++.+.||+||.+|||||+|||++|+++|
T Consensus 1 ~f~D~~Fpp~~~Sl~~~~~~~-----~~~~W~R~~e~~~------~~~~~~~~~~~~di~QG~lgDc~llaaL~~la~~~ 69 (298)
T PF00648_consen 1 LFEDPEFPPNDSSLGFDDQKP-----KNVEWKRPSEICE------NPQFFIDGISPSDIRQGSLGDCWLLAALAALAEHP 69 (298)
T ss_dssp ----TTS-SSHHHHTSSTTST-----TT-EEE-HHHHSS------S-BSSSSSSSGGGEBE-SSSSHHHHHHHHHHTTSH
T ss_pred CccCCCCccCccccccCCCCC-----CcceeEechhcCC------CCeEEECCCccccccccccCChhHHHHHHHHHhcc
Confidence 599999999999998765433 3469999999985 47788899999999999999999999999999999
Q ss_pred cccccccc--ccc--CCCCcEEEEEeeCCEEEEEEEeccccCCCCCceEEeecCCCCchhHHHHHHHHHHhcCCcccccC
Q 000135 1785 QISEVIIT--PEY--NEEGIYTVRFCIQGEWVPVVVDDWIPCESPGKPAFATSKKGHELWVSILEKAYAKLHGSYEALEG 1860 (2087)
Q Consensus 1785 rLle~fI~--Pey--Ne~GIY~VRL~iNGeWReVVVDDrLPc~~nGKPLFArSsd~nELWpSLLEKAYAKLhGSYEALeG 1860 (2087)
.+++.+++ +.. ++.|+|+||||++|+|++|+|||+||+ .+|+|+|++|.+++|+|++||||||||+||||++|.|
T Consensus 70 ~~i~~i~~~~~~~~~~~~G~y~v~l~~~G~w~~V~VDd~lP~-~~g~~~f~~s~~~~elW~~LlEKAyAKl~GsY~~l~g 148 (298)
T PF00648_consen 70 DLIKKIFPVNQSFNENYNGIYTVRLFKNGEWREVTVDDRLPC-KNGKPLFARSSDPNELWPSLLEKAYAKLHGSYSALEG 148 (298)
T ss_dssp HHHHHHS-SS--SSTT-SSEEEEEEEETTEEEEEEEES-EEE-ETTEESSSBESSTTB-HHHHHHHHHHHHTTSSGGGSS
T ss_pred cccccccccccccccccCceeeEeeccCCeeeeeccchhhhc-cccceeeeccCCcccchhhhhhchhhhccccccccCC
Confidence 88777763 222 346999999999999999999999999 6899999999899999999999999999999999999
Q ss_pred CChhhhhhhcCCCcceEEeCCchhhhhccchhHHHHHHHHHhcCCCEEEecCCCC---CCccccccCcccCceeEEEEEE
Q 000135 1861 GLVQDALVDLTGGAGEEIDMRSAQAQIDLASGRLWSQLLRFKQEGFLLGAGSPSG---SDVHISSSGIVQGHAYSILQVR 1937 (2087)
Q Consensus 1861 GnpsEALqDLTGGP~E~IDL~sa~aq~DldsdeLWk~Llkalk~G~LMgcSTPsg---SDeeves~GLVsGHAYSVLDVr 1937 (2087)
|++.++|++|||++++.+++++.. ..+++|+.+.+..+++.++++.+... .....+..||+++|||+|++++
T Consensus 149 g~~~~al~~LTG~~~~~~~l~~~~-----~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~gl~~~HaY~Vl~~~ 223 (298)
T PF00648_consen 149 GNPSEALQDLTGGPPESIDLRDDS-----SDDELWELWKKLLKSGSLVGCSTGSSTPFDSEEYEKNGLVPGHAYAVLDVR 223 (298)
T ss_dssp BSHHHHHHHHHSSEEEEEEGGG-------T--THHHHHHHHHHCT-EEEEE--SSSGGGTTSBCTTSBBTTS-EEEEEEE
T ss_pred CChhhhhHhhcCCcceeeeccccc-----hhhhHHHHHHHHHHhccccccccccccccccccccccCcccceeEEEEEEE
Confidence 999999999999999999987543 13468888888899999888776432 1233568999999999999999
Q ss_pred EECC----EEEEEEecCCCCCccccCCCCCCCcccc---hHHhhhhcCCCCCCCCeEEEehhhhhhcccceeEeeE
Q 000135 1938 EVDG----HKLVQIRNPWANEVEWNGPWSDSSPEWT---DRMKHKLKHVPQSKDGIFWMSWQDFQIHFRSIYVCRV 2006 (2087)
Q Consensus 1938 EVdG----~RLVRLRNPWG~~~EWKGdWSD~S~eWT---eeLKkkL~~~~~sDDGeFWMSfEDFLkyFssLyICrL 2006 (2087)
++++ +|||||||||| ..||+|+|||+|++|+ +..++.++. ...+||+|||+|+||++||+.++||++
T Consensus 224 ~~~~~~~~~~lv~LrNPwg-~~~w~G~ws~~s~~W~~~~~~~~~~~~~-~~~~dg~FWM~~~df~~~F~~i~vc~~ 297 (298)
T PF00648_consen 224 EVNGNGEGHRLVKLRNPWG-STEWKGDWSDDSPEWTEIHPSLRKRLNQ-SSSDDGTFWMSFEDFLKYFSSIYVCRL 297 (298)
T ss_dssp EEEETTEEEEEEEEE-TTS-S---SSTTSTTSGGGGGS-HHHHHHHTT-TSSSSSEEEEEHHHHHHHSEEEEEEES
T ss_pred eeccccceeEEEEEcCCCc-cccccccccccccccccCCHHHHhhccc-ccccCccHhHhHHHHHhhCCceEEEee
Confidence 9975 89999999999 5799999999999999 456777765 346899999999999999999999986
No 5
>KOG0045 consensus Cytosolic Ca2+-dependent cysteine protease (calpain), large subunit (EF-Hand protein superfamily) [Posttranslational modification, protein turnover, chaperones; Signal transduction mechanisms]
Probab=99.95 E-value=2e-31 Score=323.90 Aligned_cols=598 Identities=23% Similarity=0.181 Sum_probs=482.6
Q ss_pred EEeccCCCCCChhhHHHhhhhhhhHHHHHHhhcccceeecCccccccceeeeehhHHHHHHhhhhheeeeechhHHHHHH
Q 000135 874 VVKSREDQVPTKGDFLAALLPLVCIPALLSLCSGLLKWKDDDWKLSRGVYVFITIGLVLLLGAISAVIVVITPWTIGVAF 953 (2087)
Q Consensus 874 v~~sr~~~~p~~~dfl~allpl~~ipa~~~l~~gl~kw~dd~w~~s~~~y~f~~~gl~ll~~aisa~~~~~~pw~~gvaf 953 (2087)
..+.|++..|++..|..+.+|..+.+.++.+++..-+|++-.|++..- .+.+||+|..-.
T Consensus 13 ~~~~~~~cl~~~~~F~D~~FP~~~~Sl~~~~~~p~~~~~~i~W~RP~e--------------------i~~~p~~i~~~~ 72 (612)
T KOG0045|consen 13 FERLRRDCLPAKSLFVDALFPAADSSLFYKLSTPLAQFSDIVWKRPQE--------------------ICANPRLIVDGP 72 (612)
T ss_pred HHHHHHHHhhcCCcccccCCCCCCccccccccCCCcccccceecCccc--------------------ccCCCCeecCCC
Confidence 346789999999999999999999999999999998887777777665 247999976655
Q ss_pred HHHHHHHHHHHhhhhcccccceeeehhhHHHHHHHHHHHHHHHHHhhhcCCCCccccchhHHHHHHHhhccceeeeccCC
Q 000135 954 LLLLLLIVLAIGVIHHWASNNFYLTRTQMFFVCFLAFLLGLAAFLVGWFDDKPFVGASVGYFTFLFLLAGRALTVLLSPP 1033 (2087)
Q Consensus 954 ll~~~~~v~~igvih~wasnnfyl~r~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1033 (2087)
..+-+.... .+|.++..-...+.+.-.++.-... +|+-|-..+.|+|.|-|...|+..+|.
T Consensus 73 ~~~di~Qg~---------lgdCw~laA~a~la~~~~ll~~vip------~~~~~~~~yaGif~f~~w~~G~W~~Vv---- 133 (612)
T KOG0045|consen 73 SRFDVKQGL---------LGDCWFLAACAALALRPELLDKVIP------QDQSFQENYAGIFHFRFWQNGEWVEVV---- 133 (612)
T ss_pred CcceeEEee---------ecchHHHHHHHHhhcCHHHHHhccC------CCcccccccceEEEEEEEeCCeEEEEE----
Confidence 444332221 4566655554444444444444333 899999999999999999999988764
Q ss_pred EEEecCceeeEEEeecccccCCCchhhHHHHHHHHhhhccceeEEEEEEcCCCcccchhhhhheeeeccccccccchhhc
Q 000135 1034 IVVYSPRVLPVYVYDAHADCGKNVSVAFLVLYGVALAIEGWGVVASLKIYPPFAGAAVSAITLVVAFGFAVSRPCLTLKT 1113 (2087)
Q Consensus 1034 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1113 (2087)
| --.||+|+++.| .+.|... ..+.+||...+| ...+-.|+++.|..++.+. ++|+.+++.|+...|+
T Consensus 134 --I--DD~LP~~~~~~~----~~~s~~~-~efW~aLlEKAy--aKl~GsY~~l~gg~~~~a~--~~lTG~~~e~~~l~~~ 200 (612)
T KOG0045|consen 134 --I--DDRLPTSNGGLL----FSHSSGK-NEFWAALLEKAY--AKLLGSYEALHGGSTIDAL--VDLTGGVTEPFDLNKT 200 (612)
T ss_pred --e--eeecceEcCCEE----EEeecCC-ceeHHHHHHHHH--HHHhCcccCCCCCchhhHH--HhccCCccceeEcccC
Confidence 2 568999999998 6677777 788999999999 6678899999999887766 9999999999999999
Q ss_pred hHHHhhhcchhhHHHHHhhhccccccccccccccccccccccceeccCCccccccCCCccccchhhhHHHHhhccccccc
Q 000135 1114 MEDAVHFLSKDTVVQAISRSATKTRNALSGTYSAPQRSASSTALLVGDPNATRDKQGNLMLPRDDVVKLRDRLKNEEFVA 1193 (2087)
Q Consensus 1114 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1193 (2087)
+++... ++++++-++++|..+++..+++ ..+.+++ +.+++|+.|++..-.+
T Consensus 201 ~~~~~~-----~l~~~~~~~~~~~~~l~c~~~~----------------~~~~~~~--------~~~~~~~gL~~~HaYs 251 (612)
T KOG0045|consen 201 PKSFKN-----NLVWALLKSAHRGSLLLCSIES----------------KDPTEEE--------EEAKLRNGLVKGHAYA 251 (612)
T ss_pred cchhHH-----HHHHHHHHhhhccCceeeeccc----------------cccchhH--------HHHHhhcCccccccEE
Confidence 998876 7899999999999999988876 1222222 7999999999999999
Q ss_pred ccccccccccccccCCCCCchhhHhhhhhhhhhhhhhhcccceeeeeccchhhhHhhhccchhhhhhhhhhhhhhhhccc
Q 000135 1194 GSFFCRMKYKRFRHELSSDYDYRREMCTHARILALEEAIDTEWVYMWDKFGGYLLLLLGLTAKAERVQDEVRLRLFLDSI 1273 (2087)
Q Consensus 1194 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1273 (2087)
.+-.+.++. |+.|+.|.||..-.. ++||.++|++.+.+...+.....+..++|.
T Consensus 252 it~~~~~~~-------------~~~~~~lirlrNPwg--~~~W~G~wsd~~~~W~~v~~~~~~~~~~~~----------- 305 (612)
T KOG0045|consen 252 ITDVREVQG-------------RGGKHRLIRLRNPWG--ESEWNGPWSDGSEEWHLVDKSKLSELGRQP----------- 305 (612)
T ss_pred EEEEEEeec-------------ccccceeEEecCCcC--CceeccccccCCcchhhhCHHHHhhccccc-----------
Confidence 998888875 999999999999988 999999999999999999988888777775
Q ss_pred CCCcCChhhhhccCchhhhhHHHHHHhhhhhhhhHHHHHHHHHhhhcccHHHHHHHHHHHHhhHHhhhhhhcccCCCCCc
Q 000135 1274 GFSDLSAKKIKKWMPEDRRQFEIIQESYIREKEMEEEILMQRREEEGRGKERRKALLEKEERKWKEIEASLISSIPNAGN 1353 (2087)
Q Consensus 1274 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1353 (2087)
++|...||++|..+.+..+..+-+++++.+|..+| +..+.++++++.+
T Consensus 306 ------~~dGeFWms~~dF~~~F~~~~vC~~~~~~~~~~~~--------------------------~~~~~~~~~~~w~ 353 (612)
T KOG0045|consen 306 ------LDDGEFWMSFDDFLREFDSLTVCRLRPDWLESRNQ--------------------------LQWVKLSLDGEWE 353 (612)
T ss_pred ------ccCCCeeeeHHHHHhhCCeEeecCCCcchhhhhhe--------------------------eeeeeeecCCccc
Confidence 67889999999999999999999999999988877 5567788999988
Q ss_pred hHHHHHHHHHHHhcCCccccchhhhHHHHHHHHHHHHHHHHHHHHhcCCcceEEeeCCCCCccCccccccccccccccee
Q 000135 1354 REAAAMAAAVRAVGGDSVLEDSFARERVSSIARRIRTAQLARRALQTGITGAICVLDDEPTTSGRHCGQIDASICQSQKV 1433 (2087)
Q Consensus 1354 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1433 (2087)
.++....||.....++|.+.....|+++....
T Consensus 354 ------~~~~~t~ggc~~~~~tF~~npq~~~~~~~~~~------------------------------------------ 385 (612)
T KOG0045|consen 354 ------LARGVTAGGCRNSVDTFDRNPQYILAVRKPTK------------------------------------------ 385 (612)
T ss_pred ------eeecccCCCCccCcccccCCceEEEEecCCCc------------------------------------------
Confidence 66778899999999999987766665544333
Q ss_pred EEEEEEEeecCCCceeeecccccchhhhheeeccccccccccceeEEEEEecCCceeeeeeeccccceecCCceEEEEEE
Q 000135 1434 SFSIAVMIQPESGPVCLLGTEFQKKVCWEILVAGSEQGIEAGQVGLRLITKGDRQTTVAKDWSISATSIADGRWHIVTMT 1513 (2087)
Q Consensus 1434 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1513 (2087)
+++.+..+..|++..++|.+|+ .+.+..||+++++
T Consensus 386 --------------------------------------------~~~~~v~~~~q~~~~~~~~~~~-~~~~ig~~i~~v~ 420 (612)
T KOG0045|consen 386 --------------------------------------------SLCAVVLALFQKTRRGERSFGA-NILDIGFHIYEVP 420 (612)
T ss_pred --------------------------------------------cceEEEEEeecccccccccccc-eeeecceEEEEec
Confidence 8999999999999999999999 9999999999988
Q ss_pred EeccccceeeeecccccccccccccccccccccCCceEEeecCCCCccccccCCCccccccchhhheehhhcccCChHHH
Q 000135 1514 IDADIGEATCYLDGGFDGYQTGLALSAGNSIWEEGAEVWVGVRPPTDMDVFGRSDSEGAESKMHIMDVFLWGRCLTEDEI 1593 (2087)
Q Consensus 1514 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~cLtedEi 1593 (2087)
.+ ++.|. ++++.+.....-||.++..+|.. +||.+...|+ |+.|..|+.+|+|+||-++.|.+|+
T Consensus 421 ~~----------~~~~~-~~~~~~~~~~~~i~~r~v~~~~~-~P~~~y~~~p-st~~~~~~~~f~lrvfs~~~~~~~~-- 485 (612)
T KOG0045|consen 421 LE----------GKYFV-LDNAPIASSSSFINNREVSVRFR-LPPGTYVIVP-STFEPGEEGEFLLRVFSNVKVKSEE-- 485 (612)
T ss_pred CC----------CCceE-ecccchhcccccccceeEEEEec-CCCcceeecc-cCCCCCCCccEEEEEeecccccCcc--
Confidence 76 66777 99999999999999999999999 9999999999 9999999999999999999999998
Q ss_pred HHHhhcccccccccccCCCCCcccCCCCcccccCCCCCcceeeccccccccccccccCccccCcccceeeccchhhhhhc
Q 000135 1594 ASLYSAICSAELNMNEFPEDNWQWADSPPRVDEWDSDPADVDLYDRDDIDWDGQYSSGRKRRADRDGIVVNVDSFARKFR 1673 (2087)
Q Consensus 1594 ~~~~~~~~~aey~~~d~~dd~WQ~~dsp~R~~~~~~d~a~v~ly~rE~v~~~~q~ssGrk~~~~rd~i~ldmDsf~RKlr 1673 (2087)
+..+..+...|+..
T Consensus 486 ------------------------------------------------------------------~~~i~~~~~~~~~~ 499 (612)
T KOG0045|consen 486 ------------------------------------------------------------------DMEISLDETKRSTN 499 (612)
T ss_pred ------------------------------------------------------------------ceEEeeccccccee
Confidence 01111111111111
Q ss_pred CCCcCcHHHHHHHHHHHHHHHHHHHHHcCCCceecCCCCCCCCCcccCCCCCCcccccccccccccccccccccCCCcee
Q 000135 1674 KPRMETQEEIYQRMLSVELAVKEALSARGERQFTDHEFPPDDQSLYVDPGNPPSKLQVVAEWMRPSEIVKESRLDCQPCL 1753 (2087)
Q Consensus 1674 kpr~Et~EEI~QrL~svE~aIKE~cl~rGeklFeDPEFPPndsSLy~Dp~~PpsKlq~vIeWKRPsEI~~e~k~~snP~L 1753 (2087)
....
T Consensus 500 ~~~~---------------------------------------------------------------------------- 503 (612)
T KOG0045|consen 500 IIVM---------------------------------------------------------------------------- 503 (612)
T ss_pred eeee----------------------------------------------------------------------------
Confidence 1111
Q ss_pred ecCCCCCCCcccCCCCCchHHHHHHHHhccccccccccccccCCCCcEEEEEeeCCEEEEEEEeccccCCCCCceEEeec
Q 000135 1754 FSGAVNPSDVCQGRLGDCWFLSAVAVLTEVSQISEVIITPEYNEEGIYTVRFCIQGEWVPVVVDDWIPCESPGKPAFATS 1833 (2087)
Q Consensus 1754 F~dgISP~DIkQGsLGDCWFLAALAALAE~PrLle~fI~PeyNe~GIY~VRL~iNGeWReVVVDDrLPc~~nGKPLFArS 1833 (2087)
.+..++..+|+|.+.......++++....+...+.++. | ..++++..++ |..+++...|...+...
T Consensus 504 -------~~~~~~~~~~~~~~~~~~~~~k~s~~~~~~~~~~~~~~--~----~~~~~~~~~~-~~~~~~~~~~~~~~~~~ 569 (612)
T KOG0045|consen 504 -------KGFSLGECGDKWKLSSTLVNTKVSRSSEFILTVEVVSP--L----DIEGESTLVV-DIPIAIESKGSGDVAPL 569 (612)
T ss_pred -------cceehhhhchhhhccccccccccchhhceeeeeccccc--E----EEeccccccc-cccceeeccCCcccccc
Confidence 03345556666666555555555444333333333333 2 6788888888 99899877777777776
Q ss_pred CCCCchhHHHHHHHHHHhcCCcccccCCChhhhhhhcCCCc
Q 000135 1834 KKGHELWVSILEKAYAKLHGSYEALEGGLVQDALVDLTGGA 1874 (2087)
Q Consensus 1834 sd~nELWpSLLEKAYAKLhGSYEALeGGnpsEALqDLTGGP 1874 (2087)
.+..+.|....|++|++.+..+...+++...+.+.++++..
T Consensus 570 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 610 (612)
T KOG0045|consen 570 LNVIRLRIADPEIAYSFDSTSCCATEGPLVLDELFDLSSKK 610 (612)
T ss_pred eeeeeeeccChhheeeccccccccccCcchhhhhhcCCCCC
Confidence 66789999999999999999999999999999999888754
No 6
>smart00720 calpain_III calpain_III.
Probab=98.36 E-value=6.2e-07 Score=92.17 Aligned_cols=51 Identities=41% Similarity=0.724 Sum_probs=43.9
Q ss_pred eeecceee-ccCCCCCCCC-CCCCcCCeEEEEecCCCCCCCEEEEEEeccCCcc
Q 000135 2013 YSVHGQWR-GYSAGGCQDY-ASWNQNPQFRLRASGSDASFPIHVFITLTQSRFY 2064 (2087)
Q Consensus 2013 yrVhGeWr-GsSAGGc~n~-~SF~~NPQF~LeVtssD~sep~eVlISLsQkr~Y 2064 (2087)
..++|+|. +.+||||.++ .+|++||||.|++.+++. ..|+|+|+|+|+...
T Consensus 4 ~~~~G~W~~~~tAGG~~~~~~tf~~NPqy~l~v~~~~~-~~~~v~i~L~q~~~r 56 (143)
T smart00720 4 KSVQGSWTRGQTAGGCRNYPATFWTNPQFRITLEEPDD-DDCTVLIALMQKNRR 56 (143)
T ss_pred EEEeCeEECCCccCCccccccccccCCeEEEEecCCCC-CceEEEEEecccCcc
Confidence 46899997 8999999999 899999999999986653 348999999998643
No 7
>cd00214 Calpain_III Calpain, subdomain III. Calpains are calcium-activated cytoplasmic cysteine proteinases, participate in cytoskeletal remodeling processes, cell differentiation, apoptosis and signal transduction. Catalytic domain and the two calmodulin-like domains are separated by C2-like domain III. Domain III plays an important role in calcium-induced activation of calpain involving electrostatic interactions with subdomain II. Proposed to mediate calpain's interaction with phospholipids and translocation to cytoplasmic/nuclear membranes. CD includes subdomain III of typical and atypical calpains.
Probab=98.15 E-value=3e-06 Score=88.73 Aligned_cols=52 Identities=37% Similarity=0.643 Sum_probs=42.6
Q ss_pred eeecceeec-cCCCCCCCC-CCCCcCCeEEEEecCCCC-CCCEEEEEEeccCCcc
Q 000135 2013 YSVHGQWRG-YSAGGCQDY-ASWNQNPQFRLRASGSDA-SFPIHVFITLTQSRFY 2064 (2087)
Q Consensus 2013 yrVhGeWrG-sSAGGc~n~-~SF~~NPQF~LeVtssD~-sep~eVlISLsQkr~Y 2064 (2087)
..++|+|+. .+||||.++ .+|++||||.|+++++|. ...++|+|+|+|++..
T Consensus 6 ~~~~G~W~~g~tAGGc~~~~~tf~~NPQf~l~v~~~~~~~~~~~v~i~L~q~~~r 60 (150)
T cd00214 6 KSFNGEWRRGQTAGGCRNNPDTFWTNPQFRIRVPEPDDDEGKCTVLIALMQKNRR 60 (150)
T ss_pred EEEeCeEeCCcccCCCCCcccccccCceEEEEecCCCCCCCccEEEEEeccCCcc
Confidence 468999975 999999665 799999999999987642 2348999999998643
No 8
>PF01067 Calpain_III: Calpain large subunit, domain III; InterPro: IPR022682 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This group of cysteine peptidases belong to the MEROPS peptidase family C2 (calpain family, clan CA). A type example is calpain, which is an intracellular protease involved in many important cellular functions that are regulated by calcium []. The protein is a complex of 2 polypeptide chains (light and heavy), with three known forms in mammals [, ]: a highly calcium-sensitive (i.e., micro-molar range) form known as mu-calpain, mu-CANP or calpain I; a form sensitive to calcium in the milli-molar range, known as m-calpain, m-CANP or calpain II; and a third form, known as p94, which is found in skeletal muscle only []. All forms have identical light but different heavy chains. Both mu- and m-calpain are heterodimers containing an identical 28kDa subunit and an 80kDa subunit that shares 55-65% sequence homology between the two proteases [, ]. The crystallographic structure of m-calpain reveals six "domains" in the 80kDa subunit: A 19-amino acid NH2-terminal sequence; Active site domain IIa; Active site domain IIb. Domain 2 shows low levels of sequence similarity to papain; although the catalytic His has not been located by biochemical means, it is likely that calpain and papain are related []. Domain III; An 18-amino acid extended sequence linking domain III to domain IV; Domain IV, which resembles the penta EF-hand family of polypeptides, binds calcium and regulates activity []. />]. Ca2+-binding causes a rearrangement of the protein backbone, the net effect of which is that a Trp side chain, which acts as a wedge between catalytic domains IIa and IIb in the apo state, moves away from the active site cleft allowing for the proper formation of the catalytic triad []. Calpain-like mRNAs have been identified in other organisms including bacteria, but the molecules encoded by these mRNAs have not been isolated, so little is known about their properties. How calpain activity is regulated in these organisms cells is still unclear In metazoans, the activity of calpain is controlled by a single proteinase inhibitor, calpastatin (IPR001259 from INTERPRO). The calpastatin gene can produce eight or more calpastatin polypeptides ranging from 17 to 85 kDa by use of different promoters and alternative splicing events. The physiological significance of these different calpastatins is unclear, although all bind to three different places on the calpain molecule; binding to at least two of the sites is Ca2+ dependent. The calpains ostensibly participate in a variety of cellular processes including remodelling of cytoskeletal/membrane attachments, different signal transduction pathways, and apoptosis. Deregulated calpain activity following loss of Ca2+ homeostasis results in tissue damage in response to events such as myocardial infarcts, stroke, and brain trauma []. This entry represents domain III. It is found in association with PF00648 from PFAM. The function of the domain III and I are currently unknown. Domain II is a cysteine protease and domain IV is a calcium binding domain. Calpains are believed to participate in intracellular signaling pathways mediated by calcium ions. ; PDB: 1QXP_B 2QFE_A 1DF0_A 1U5I_A 3DF0_A 3BOW_A 1KFU_L 1KFX_L.
Probab=98.08 E-value=4e-06 Score=85.37 Aligned_cols=51 Identities=39% Similarity=0.737 Sum_probs=39.1
Q ss_pred eeeccee-eccCCCCCCCCC-CCCcCCeEEEEecCCCC-CCCEEEEEEeccCCc
Q 000135 2013 YSVHGQW-RGYSAGGCQDYA-SWNQNPQFRLRASGSDA-SFPIHVFITLTQSRF 2063 (2087)
Q Consensus 2013 yrVhGeW-rGsSAGGc~n~~-SF~~NPQF~LeVtssD~-sep~eVlISLsQkr~ 2063 (2087)
..++|+| ++.+||||.++. +|++||||.|+++.++. +.+++|+|+|+|++.
T Consensus 5 ~~~~G~W~~~~taGG~~~~~~s~~~NPQy~l~v~~~~~~~~~~~v~i~L~q~~~ 58 (147)
T PF01067_consen 5 VTIEGEWVTGNTAGGCPNNPYSWWNNPQYRLTVSEPTEESNKCTVVISLMQKDR 58 (147)
T ss_dssp EEEEEEE-TTTS---STT-TTTGGGS-EEEEEESSGCCCSSBEEEEEEEEECSG
T ss_pred EEEeCEEeCCCcCCCCcccccccccCcEEEEEEcCCCCCcceeEEEEEEEecCc
Confidence 4689999 899999999998 99999999999987653 236899999999664
No 9
>cd00152 PTX Pentraxins are plasma proteins characterized by their pentameric discoid assembly and their Ca2+ dependent ligand binding, such as Serum amyloid P component (SAP) and C-reactive Protein (CRP), which are cytokine-inducible acute-phase proteins implicated in innate immunity. CRP binds to ligands containing phosphocholine, SAP binds to amyloid fibrils, DNA, chromatin, fibronectin, C4-binding proteins and glycosaminoglycans. "Long" pentraxins have N-terminal extensions to the common pentraxin domain; one group, the neuronal pentraxins, may be involved in synapse formation and remodeling, and they may also be able to form heteromultimers.
Probab=97.98 E-value=5.3e-05 Score=82.51 Aligned_cols=162 Identities=20% Similarity=0.328 Sum_probs=100.1
Q ss_pred EEEEEEEeecC--CCceeeecccccchhhhheeeccccccccccceeEEEEEecCCceeeeeeeccccceecCCceEEEE
Q 000135 1434 SFSIAVMIQPE--SGPVCLLGTEFQKKVCWEILVAGSEQGIEAGQVGLRLITKGDRQTTVAKDWSISATSIADGRWHIVT 1511 (2087)
Q Consensus 1434 ~~~~~~~~~~~--~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1511 (2087)
+|++++-++.+ +++..+|..--.++ =-|+++.+..+ |+ +.+-..|...++ . ....||+||.|+
T Consensus 32 ~fTv~~Wv~~~~~~~~~~ifSy~~~~~-~~~~~l~~~~~----g~--~~~~i~~~~~~~-------~-~~~~~g~W~hv~ 96 (201)
T cd00152 32 AFTLCLWVYTDLSTREYSLFSYATKGQ-DNELLLYKEKD----GG--YSLYIGGKEVTF-------K-VPESDGAWHHIC 96 (201)
T ss_pred hEEEEEEEEecCCCCCeEEEEEeCCCC-CCeEEEEEcCC----Ce--EEEEEcCEEEEE-------e-ccCCCCCEEEEE
Confidence 57788888776 47777774333211 22777664432 33 333333332221 2 234899999999
Q ss_pred EEEeccccceeeeecccccccccccccccccccccCCceEEeecCCCCccccccCCCccccccchhhheehhhcccCChH
Q 000135 1512 MTIDADIGEATCYLDGGFDGYQTGLALSAGNSIWEEGAEVWVGVRPPTDMDVFGRSDSEGAESKMHIMDVFLWGRCLTED 1591 (2087)
Q Consensus 1512 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~cLted 1591 (2087)
+|-|..+|+.+-|+||.-.+.++ + . .+..+..+....+|-++ |..|..-...-.=+=+|=|+-+|.|-||.+
T Consensus 97 ~t~d~~~g~~~lyvnG~~~~~~~-~--~-~~~~~~~~g~l~lG~~q----~~~gg~~~~~~~f~G~I~~v~iw~~~Ls~~ 168 (201)
T cd00152 97 VTWESTSGIAELWVNGKLSVRKS-L--K-KGYTVGPGGSIILGQEQ----DSYGGGFDATQSFVGEISDVNMWDSVLSPE 168 (201)
T ss_pred EEEECCCCcEEEEECCEEecccc-c--c-CCCEECCCCeEEEeecc----cCCCCCCCCCcceEEEEceeEEEcccCCHH
Confidence 99999999999999998776554 1 1 12344556667777654 233322111111233567888999999999
Q ss_pred HHHHHhhcccccccccccCCCCCcccC
Q 000135 1592 EIASLYSAICSAELNMNEFPEDNWQWA 1618 (2087)
Q Consensus 1592 Ei~~~~~~~~~aey~~~d~~dd~WQ~~ 1618 (2087)
||..+++.-+...=++++-.++.|+.+
T Consensus 169 eI~~l~~~~~~~~Gnv~~W~~~~~~~~ 195 (201)
T cd00152 169 EIKNVYSEGGTLSGNILNWRALNYEIN 195 (201)
T ss_pred HHHHHHhcCCCCCCCEEechhhEEEEe
Confidence 999998744444555555555555554
No 10
>smart00159 PTX Pentraxin / C-reactive protein / pentaxin family. This family form a doscoid pentameric structure. Human serum amyloid P demonstrates calcium-mediated ligand-binding.
Probab=97.87 E-value=0.00011 Score=80.49 Aligned_cols=160 Identities=21% Similarity=0.322 Sum_probs=100.6
Q ss_pred EEEEEEEeecCC--Cceeee--cccccchhhhheeeccccccccccceeEEEEEecCCceeeeeeeccccceecCCceEE
Q 000135 1434 SFSIAVMIQPES--GPVCLL--GTEFQKKVCWEILVAGSEQGIEAGQVGLRLITKGDRQTTVAKDWSISATSIADGRWHI 1509 (2087)
Q Consensus 1434 ~~~~~~~~~~~~--~~~~~~--~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1509 (2087)
+|++++-++++. ++-.|| .+..|. -|+++... ++.++.+...|...+ ....+.||+||.
T Consensus 32 ~fTvc~W~k~~~~~~~~~ifSy~~~~~~---ne~~~~~~------~~~~~~l~i~g~~~~--------~~~~~~~g~W~h 94 (206)
T smart00159 32 AFTVCLWFYSDLSPRGYSLFSYATKGQD---NELLLYKE------KQGEYSLYIGGKKVQ--------FPVPESDGKWHH 94 (206)
T ss_pred HEEEEEEEEecCCCCceEEEEEeCCCCC---CeEEEEEc------CCcEEEEEEcCeEEE--------ecccccCCceEE
Confidence 567777777653 444454 665554 36766533 233466666664211 123578999999
Q ss_pred EEEEEeccccceeeeecccccccccccccccccccccCCceEEeecCCCCccccccCCCccccccchhhheehhhcccCC
Q 000135 1510 VTMTIDADIGEATCYLDGGFDGYQTGLALSAGNSIWEEGAEVWVGVRPPTDMDVFGRSDSEGAESKMHIMDVFLWGRCLT 1589 (2087)
Q Consensus 1510 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~cLt 1589 (2087)
|++|-|..+|+++-|+||... .+.++ .. +..+..+-.+.+|-++ |..|-.-.+...=+=.|=|+=||.|-||
T Consensus 95 vc~tw~~~~g~~~lyvnG~~~-~~~~~--~~-g~~i~~~G~lvlGq~q----d~~gg~f~~~~~f~G~i~~v~iw~~~Ls 166 (206)
T smart00159 95 ICTTWESSSGIAELWVDGKPG-VRKGL--AK-GYTVKPGGSIILGQEQ----DSYGGGFDATQSFVGEIGDLNMWDSVLS 166 (206)
T ss_pred EEEEEECCCCcEEEEECCEEc-ccccc--cC-CcEECCCCEEEEEecc----cCCCCCCCCCcceeEEEeeeEEecccCC
Confidence 999999999999999999875 33322 11 2344566677888764 3333221111112335668889999999
Q ss_pred hHHHHHHhhcccccccccccCCCCCcccC
Q 000135 1590 EDEIASLYSAICSAELNMNEFPEDNWQWA 1618 (2087)
Q Consensus 1590 edEi~~~~~~~~~aey~~~d~~dd~WQ~~ 1618 (2087)
++||..+++.-...+=++.+-.++.|+.+
T Consensus 167 ~~eI~~l~~~~~~~~Gnv~~W~~~~~~~~ 195 (206)
T smart00159 167 PEEIKSVYKGSTFSIGNILNWRALNYEVH 195 (206)
T ss_pred HHHHHHHHcCCCCCCCCEEeccccEEEEe
Confidence 99999998743333345666666666665
No 11
>PF13385 Laminin_G_3: Concanavalin A-like lectin/glucanases superfamily; PDB: 4DQA_A 1N1Y_A 1MZ6_A 1MZ5_A 1N1S_A 2A75_A 1WCS_A 1N1T_A 1N1V_A 2FHR_A ....
Probab=97.40 E-value=0.00065 Score=66.55 Aligned_cols=80 Identities=24% Similarity=0.379 Sum_probs=49.6
Q ss_pred ccceecCCceEEEEEEEeccccceeeeecccccccccccccccccccccCCceEEeecCCCCccccccCCCccccccchh
Q 000135 1498 SATSIADGRWHIVTMTIDADIGEATCYLDGGFDGYQTGLALSAGNSIWEEGAEVWVGVRPPTDMDVFGRSDSEGAESKMH 1577 (2087)
Q Consensus 1498 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1577 (2087)
....+.+++||.+++|+| .++.+.|+||-..+....-.. ..+......-+|-.+ ....--+..
T Consensus 78 ~~~~~~~~~W~~l~~~~~--~~~~~lyvnG~~~~~~~~~~~----~~~~~~~~~~iG~~~-----------~~~~~~~g~ 140 (157)
T PF13385_consen 78 SDSNLPDNKWHHLALTYD--GSTVTLYVNGELVGSSTIPSN----ISLNSNGPLFIGGSG-----------GGSSPFNGY 140 (157)
T ss_dssp -BS---TT-EEEEEEEEE--TTEEEEEETTEEETTCTEESS----SSTTSCCEEEESS-S-----------TT--B-EEE
T ss_pred cCcccCCCCEEEEEEEEE--CCeEEEEECCEEEEeEeccCC----cCCCCcceEEEeecC-----------CCCCceEEE
Confidence 556788999999999999 556999999999986543222 123444455555433 223344678
Q ss_pred hheehhhcccCChHHHH
Q 000135 1578 IMDVFLWGRCLTEDEIA 1594 (2087)
Q Consensus 1578 ~~~~~~~~~cLtedEi~ 1594 (2087)
|-|+-+|.|+||++||+
T Consensus 141 i~~~~i~~~aLt~~eI~ 157 (157)
T PF13385_consen 141 IDDLRIYNRALTAEEIQ 157 (157)
T ss_dssp EEEEEEESS---HHHHH
T ss_pred EEEEEEECccCCHHHcC
Confidence 88999999999999996
No 12
>PF00354 Pentaxin: Pentaxin family; InterPro: IPR001759 Pentaxins (or pentraxins) [, ] are a family of proteins which show, under electron microscopy, a discoid arrangement of five noncovalently bound subunits. Proteins of the pentaxin family are involved in acute immunological responses []. Three of the principal members of the pentaxin family are serum proteins: namely, C-reactive protein (CRP) [], serum amyloid P component protein (SAP) [], and female protein (FP) []. CRP is expressed during acute phase response to tissue injury or inflammation in mammals. The protein resembles antibody and performs several functions associated with host defence: it promotes agglutination, bacterial capsular swelling and phagocytosis, and activates the classical complement pathway through its calcium-dependent binding to phosphocholine. CRPs have also been sequenced in an invertebrate, Limulus polyphemus (Atlantic horseshoe crab), where they are a normal constituent of the hemolymph. SAP is a vertebrate protein that is a precursor of amyloid component P. It is found in all types of amyloid deposits, in glomerular basement menbrane and in elastic fibres in blood vessels. SAP binds to various lipoprotein ligands in a calcium-dependent manner, and it has been suggested that, in mammals, this may have important implications in atherosclerosis and amyloidosis. FP is a SAP homologue found in Mesocricetus auratus (Golden hamster). The concentration of this plasma protein is altered by sex steroids and stimuli that elicit an acute phase response. Pentaxin proteins expressed in the nervous system are neural pentaxin I (NPI) and II (NPII) []. NPI and NPII are homologous and can exist within one species. It is suggested that both proteins mediate the uptake of synaptic macromolecules and play a role in synaptic plasticity. Apexin, a sperm acrosomal protein, is a homologue of NPII found in Cavia porcellus (Guinea pig) []. PTX3 (or TSG-14) protein is a cytokine-induced protein that is homologous to CRPs and SAPs, but its function is not yet known.; PDB: 2A3W_F 3KQR_C 3D5O_D 2A3X_G 1SAC_D 2W08_B 1GYK_B 1LGN_A 2A3Y_A 1B09_D ....
Probab=97.32 E-value=0.00075 Score=74.29 Aligned_cols=159 Identities=29% Similarity=0.496 Sum_probs=94.4
Q ss_pred EEEEEEEeecCC--Cceeee--cccccchhhhheeeccccccccccceeEEEEEecCCceeeeeeeccccceecCCceEE
Q 000135 1434 SFSIAVMIQPES--GPVCLL--GTEFQKKVCWEILVAGSEQGIEAGQVGLRLITKGDRQTTVAKDWSISATSIADGRWHI 1509 (2087)
Q Consensus 1434 ~~~~~~~~~~~~--~~~~~~--~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1509 (2087)
+|++.+-++++. ..-+|| .|+.|. -|+++.+..++ +++|.-.|.... + ...+.||+||.
T Consensus 26 ~fTvC~w~k~~~~~~~~tifSYat~~~~---nell~~~~~~~------~~~l~i~~~~~~-------~-~~~~~~~~Whh 88 (195)
T PF00354_consen 26 AFTVCFWVKTDDSSNDGTIFSYATSSQD---NELLLFGSSSG------SLRLYINGSSVS-------F-SGPIRDGQWHH 88 (195)
T ss_dssp EEEEEEEEEESGSGS-EEEEEEEETTEE---EEEEEEEETTT------EEEEEETTEEEE-------E-EECS-TSS-EE
T ss_pred cEEEEEEEEeccCCCceEEEEEccCCCC---ccEEEEEeCCc------eEEEEECCeEeE-------e-ccccCCCCcEE
Confidence 355555555533 355555 444443 37888765442 566776666221 1 13578999999
Q ss_pred EEEEEeccccceeeeecccccccccccccccccccccCCceEEeecCCCCccccccCCCccccccchhhheehhhcccCC
Q 000135 1510 VTMTIDADIGEATCYLDGGFDGYQTGLALSAGNSIWEEGAEVWVGVRPPTDMDVFGRSDSEGAESKMHIMDVFLWGRCLT 1589 (2087)
Q Consensus 1510 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~cLt 1589 (2087)
+.+|-|..+|+..-|+||-. ....+ +..+..|-..|+ +=+|-+. |.+|-.-.+...=.=.|-|+-||.|-||
T Consensus 89 ~C~tW~s~~G~~~ly~dG~~-~~~~~--~~~g~~i~~gG~-~vlGQeQ----d~~gG~fd~~q~F~G~i~~~~iWd~vLs 160 (195)
T PF00354_consen 89 ICVTWDSSTGRWQLYVDGVR-LSSTG--LATGHSIPGGGT-LVLGQEQ----DSYGGGFDESQAFVGEISDFNIWDRVLS 160 (195)
T ss_dssp EEEEEETTTTEEEEEETTEE-EEEEE--SSTT--B-SSEE-EEESS-B----SBTTBTCSGGGB--EEEEEEEEESS---
T ss_pred EEEEEecCCcEEEEEECCEe-ccccc--ccCCceECCCCE-EEECccc----cccCCCcCCccEeeEEEeceEEEeeeCC
Confidence 99999999999999999983 22233 345556655555 4477654 6666544443333446889999999999
Q ss_pred hHHHHHHhhcccccccccccCCCCCcccC
Q 000135 1590 EDEIASLYSAICSAELNMNEFPEDNWQWA 1618 (2087)
Q Consensus 1590 edEi~~~~~~~~~aey~~~d~~dd~WQ~~ 1618 (2087)
++||+.++.. +..+=++++-.+..|+..
T Consensus 161 ~~eI~~l~~~-~~~~Gnvi~W~~~~~~~~ 188 (195)
T PF00354_consen 161 PEEIRALASC-CCYKGNVISWDDLRWSIS 188 (195)
T ss_dssp HHHHHHHHHT--S---SSEEGGGBEEEEE
T ss_pred HHHHHHHHhC-CCCCCCEEccccCeEEee
Confidence 9999999986 555566666666666544
No 13
>cd00110 LamG Laminin G domain; Laminin G-like domains are usually Ca++ mediated receptors that can have binding sites for steroids, beta1 integrins, heparin, sulfatides, fibulin-1, and alpha-dystroglycans. Proteins that contain LamG domains serve a variety of purposes including signal transduction via cell-surface steroid receptors, adhesion, migration and differentiation through mediation of cell adhesion molecules.
Probab=94.54 E-value=0.23 Score=50.06 Aligned_cols=110 Identities=22% Similarity=0.327 Sum_probs=63.8
Q ss_pred eeEEEEEEEeecCCCceeeecccccchhhhheeeccccccccccceeEEEEEecCCceeeeeeeccccc-eecCCceEEE
Q 000135 1432 KVSFSIAVMIQPESGPVCLLGTEFQKKVCWEILVAGSEQGIEAGQVGLRLITKGDRQTTVAKDWSISAT-SIADGRWHIV 1510 (2087)
Q Consensus 1432 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~-~~~~~~~~~~ 1510 (2087)
.-.+++.+.+.|.+..-.||-...+.. -+.+.+ .++.|++-+++-.. .+.. .+... .+.||+||.|
T Consensus 19 ~~~~~i~~~frt~~~~g~l~~~~~~~~--~~~~~l----~l~~g~l~~~~~~g-~~~~------~~~~~~~v~dg~Wh~v 85 (151)
T cd00110 19 RTRLSISFSFRTTSPNGLLLYAGSQNG--GDFLAL----ELEDGRLVLRYDLG-SGSL------VLSSKTPLNDGQWHSV 85 (151)
T ss_pred cceeEEEEEEEeCCCCeEEEEecCCCC--CCEEEE----EEECCEEEEEEcCC-cccE------EEEccCccCCCCEEEE
Confidence 446777788888765555554444321 112221 24566766654433 2222 22222 6999999999
Q ss_pred EEEEeccccceeeeecccccccccccccccccccccCCceEEeecCCCC
Q 000135 1511 TMTIDADIGEATCYLDGGFDGYQTGLALSAGNSIWEEGAEVWVGVRPPT 1559 (2087)
Q Consensus 1511 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1559 (2087)
+++.+. ++++-|+||.--. +... +.+.-.=...+.+++|-.|..
T Consensus 86 ~i~~~~--~~~~l~VD~~~~~-~~~~--~~~~~~~~~~~~~~iGg~~~~ 129 (151)
T cd00110 86 SVERNG--RSVTLSVDGERVV-ESGS--PGGSALLNLDGPLYLGGLPED 129 (151)
T ss_pred EEEECC--CEEEEEECCccEE-eeeC--CCCceeecCCCCeEEcCCCCc
Confidence 999987 7899999997111 1111 111112246778899988864
No 14
>smart00210 TSPN Thrombospondin N-terminal -like domains. Heparin-binding and cell adhesion domain of thrombospondin
Probab=94.35 E-value=0.25 Score=53.89 Aligned_cols=104 Identities=23% Similarity=0.299 Sum_probs=58.2
Q ss_pred EEEEEEEeecC-CCceeeecccc-cchhhhheeeccccccccccceeEEEEEe---cCCceeeeeeeccccceecCCceE
Q 000135 1434 SFSIAVMIQPE-SGPVCLLGTEF-QKKVCWEILVAGSEQGIEAGQVGLRLITK---GDRQTTVAKDWSISATSIADGRWH 1508 (2087)
Q Consensus 1434 ~~~~~~~~~~~-~~~~~~~~~~~-~~~~~~~~~~~~~~~~~~~~~~~~~~~~~---~~~~~~~~~~~~~~~~~~~~~~~~ 1508 (2087)
.||+.+.++|. ..+--||...- |++.=+++.+-| ++.-+.+.++ |+.++.+- ....++|||||
T Consensus 53 ~fsi~~~~r~~~~~~g~L~si~~~~~~~~l~v~l~g-------~~~~~~~~~~~~~g~~~~~~f-----~~~~l~dg~WH 120 (184)
T smart00210 53 DFSLLTTFRQTPKSRGVLFAIYDAQNVRQFGLEVDG-------RANTLLLRYQGVDGKQHTVSF-----RNLPLADGQWH 120 (184)
T ss_pred CeEEEEEEEeCCCCCeEEEEEEcCCCcEEEEEEEeC-------CccEEEEEECCCCCcEEEEee-----cCCccccCCce
Confidence 46666667665 34444554432 444334443332 2344555542 32232221 12469999999
Q ss_pred EEEEEEeccccceeeeeccccccccccccccccc--ccccCCceEEee
Q 000135 1509 IVTMTIDADIGEATCYLDGGFDGYQTGLALSAGN--SIWEEGAEVWVG 1554 (2087)
Q Consensus 1509 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~--~~~~~~~~~~~~ 1554 (2087)
.++++|+.+ .++-|+|+..-+-+ +|+... .+=-.|..++.+
T Consensus 121 ~lal~V~~~--~v~LyvDC~~~~~~---~l~~~~~~~~~~~g~~~~g~ 163 (184)
T smart00210 121 KLALSVSGS--SATLYVDCNEIDSR---PLDRPGQPPIDTDGIEVRGA 163 (184)
T ss_pred EEEEEEeCC--EEEEEECCccccce---ecCCcccccccccceEEEee
Confidence 999999887 69999999876544 333333 333345544443
No 15
>smart00282 LamG Laminin G domain.
Probab=93.53 E-value=0.55 Score=47.47 Aligned_cols=109 Identities=20% Similarity=0.215 Sum_probs=64.8
Q ss_pred EEEEEEeecCCCceeeecccc-cchhhhheeeccccccccccceeEEEEEecCCceeeeeeeccccceecCCceEEEEEE
Q 000135 1435 FSIAVMIQPESGPVCLLGTEF-QKKVCWEILVAGSEQGIEAGQVGLRLITKGDRQTTVAKDWSISATSIADGRWHIVTMT 1513 (2087)
Q Consensus 1435 ~~~~~~~~~~~~~~~~~~~~~-~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1513 (2087)
+++.+.++|.+..=.||-+.. +.+. ++.+ .++.|++-++.-..+ +... -......+.||+||.|.++
T Consensus 3 ~~i~~~frt~~~~g~l~~~~~~~~~~---~l~l----~l~~g~l~~~~~~g~-~~~~----~~~~~~~~~dg~WH~v~i~ 70 (135)
T smart00282 3 LSISFSFRTTSPNGLLLYAGSKNGGD---YLAL----ELRDGRLVLRYDLGS-GPAR----LTSDPTPLNDGQWHRVAVE 70 (135)
T ss_pred eEEEEEEEeCCCCEEEEEeCCCCCCC---EEEE----EEECCEEEEEEECCC-CCEE----EEECCeEeCCCCEEEEEEE
Confidence 566777777765445554433 1111 1221 235688777666533 2211 1224478999999999999
Q ss_pred EeccccceeeeecccccccccccccccccccccCCceEEeecCCCCc
Q 000135 1514 IDADIGEATCYLDGGFDGYQTGLALSAGNSIWEEGAEVWVGVRPPTD 1560 (2087)
Q Consensus 1514 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1560 (2087)
.+ .++.+-++||...-... .+.....-+..+.+++|-.|+..
T Consensus 71 ~~--~~~~~l~VD~~~~~~~~---~~~~~~~l~~~~~l~iGG~p~~~ 112 (135)
T smart00282 71 RN--GRRVTLSVDGENPVSGE---SPGGLTILNLDGPLYLGGLPEDL 112 (135)
T ss_pred Ee--CCEEEEEECCCccccEE---CCCCceEEecCCCcEEccCCchh
Confidence 87 46788999996432221 12222344556789999888753
No 16
>smart00560 LamGL LamG-like jellyroll fold domain.
Probab=91.47 E-value=0.72 Score=47.68 Aligned_cols=85 Identities=19% Similarity=0.180 Sum_probs=49.5
Q ss_pred EEEEEEEeecCCCce--eeecccccchhhhheeeccccccccccceeEEEEEecCCceeeeeeeccccceecCCceEEEE
Q 000135 1434 SFSIAVMIQPESGPV--CLLGTEFQKKVCWEILVAGSEQGIEAGQVGLRLITKGDRQTTVAKDWSISATSIADGRWHIVT 1511 (2087)
Q Consensus 1434 ~~~~~~~~~~~~~~~--~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1511 (2087)
+|++++.|.|++.|- .+++ .- +.+-. +-..++.++-|-...+.+.... .-+.+....|+||-|+
T Consensus 2 ~fTv~aWv~~~~~~~~~~~~~---------~~-v~~~~-~~~~~~~~f~l~~~~~~~w~~~---~~~~~~~~~~~W~hva 67 (133)
T smart00560 2 SFTLEAWVKLESAGGSQPIIT---------GA-AVAQP-TISEKALTFFLRAKSVQGWQTA---RTGATADWIGVWVHLA 67 (133)
T ss_pred cEEEEEEEeecccCcccceee---------eE-EEEcc-CCCCCceEEEEEeeccCCEEEe---ccccCCCCCCCEEEEE
Confidence 699999999997642 1110 01 11111 2233556655544432221111 1111222239999999
Q ss_pred EEEeccccceeeeeccccccc
Q 000135 1512 MTIDADIGEATCYLDGGFDGY 1532 (2087)
Q Consensus 1512 ~~~~~~~~~~~~~~~~~~~~~ 1532 (2087)
++.|.+.|+.+.|+||-..+-
T Consensus 68 ~v~d~~~g~~~lYvnG~~~~~ 88 (133)
T smart00560 68 GVYDGGAGKLSLYVNGVEVAT 88 (133)
T ss_pred EEEECCCCeEEEEECCEEccc
Confidence 999999999999999976653
No 17
>cd02619 Peptidase_C1 C1 Peptidase family (MEROPS database nomenclature), also referred to as the papain family; composed of two subfamilies of cysteine peptidases (CPs), C1A (papain) and C1B (bleomycin hydrolase). Papain-like enzymes are mostly endopeptidases with some exceptions like cathepsins B, C, H and X, which are exopeptidases. Papain-like CPs have different functions in various organisms. Plant CPs are used to mobilize storage proteins in seeds while mammalian CPs are primarily lysosomal enzymes responsible for protein degradation in the lysosome. Papain-like CPs are synthesized as inactive proenzymes with N-terminal propeptide regions, which are removed upon activation. Bleomycin hydrolase (BH) is a CP that detoxifies bleomycin by hydrolysis of an amide group. It acts as a carboxypeptidase on its C-terminus to convert itself into an aminopeptidase and peptide ligase. BH is found in all tissues in mammals as well as in many other eukaryotes. It forms a hexameric ring barrel str
Probab=85.58 E-value=1.5 Score=47.23 Aligned_cols=49 Identities=24% Similarity=0.493 Sum_probs=39.3
Q ss_pred CcccCceeEEEEEEEEC--CEEEEEEecCCCCCccccCCCCCCCcccchHHhhhhcCCCCCCCCeEEEehhhhhhcc
Q 000135 1924 GIVQGHAYSILQVREVD--GHKLVQIRNPWANEVEWNGPWSDSSPEWTDRMKHKLKHVPQSKDGIFWMSWQDFQIHF 1998 (2087)
Q Consensus 1924 GLVsGHAYSVLDVrEVd--G~RLVRLRNPWG~~~EWKGdWSD~S~eWTeeLKkkL~~~~~sDDGeFWMSfEDFLkyF 1998 (2087)
.-..+||-.|++...-. +.....+||-||. .| .++|-|||+++++..++
T Consensus 168 ~~~~~Hav~ivGy~~~~~~~~~~~i~~NSwG~--~w------------------------g~~Gy~~i~~~~~~~~~ 218 (223)
T cd02619 168 GDLGGHAVVIVGYDDNYVEGKGAFIVKNSWGT--DW------------------------GDNGYGRISYEDVYEMT 218 (223)
T ss_pred CccCCeEEEEEeecCCCCCCCCEEEEEeCCCC--cc------------------------ccCCEEEEehhhhhhhh
Confidence 44579999999998654 6788999999994 44 35799999999998554
No 18
>KOG1029 consensus Endocytic adaptor protein intersectin [Signal transduction mechanisms; Intracellular trafficking, secretion, and vesicular transport]
Probab=79.07 E-value=3.2 Score=54.36 Aligned_cols=33 Identities=27% Similarity=0.427 Sum_probs=19.4
Q ss_pred chhHHHHHHHhhccceeeeccCCEEEecCceee
Q 000135 1011 SVGYFTFLFLLAGRALTVLLSPPIVVYSPRVLP 1043 (2087)
Q Consensus 1011 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1043 (2087)
|+..=....-|.|--+-+.|-|-+.+--||-.|
T Consensus 72 SIAmkLi~lkLqG~~lP~~LPPsll~~~~~~~p 104 (1118)
T KOG1029|consen 72 SIAMKLIKLKLQGIQLPPVLPPSLLKQPPRNAP 104 (1118)
T ss_pred HHHHHHHHHHhcCCcCCCCCChHHhccCCcCCC
Confidence 455555556677777777665546655555444
No 19
>PF02210 Laminin_G_2: Laminin G domain; InterPro: IPR012680 Laminins are large heterotrimeric glycoproteins involved in basement membrane function []. The laminin globular (G) domain can be found in one to several copies in various laminin family members, including a large number of extracellular proteins. The C terminus of the laminin alpha chain contains a tandem repeat of five laminin G domains, which are critical for heparin-binding and cell attachment activity []. Laminin alpha4 is distributed in a variety of tissues including peripheral nerves, dorsal root ganglion, skeletal muscle and capillaries; in the neuromuscular junction, it is required for synaptic specialisation []. The structure of the laminin-G domain has been predicted to resemble that of pentraxin []. Laminin G domains can vary in their function, and a variety of binding functions have been ascribed to different LamG modules. For example, the laminin alpha1 and alpha2 chains each have five C-teminal laminin G domains, where only domains LG4 and LG5 contain binding sites for heparin, sulphatides and the cell surface receptor dystroglycan []. Laminin G-containing proteins appear to have a wide variety of roles in cell adhesion, signalling, migration, assembly and differentiation. This entry represents one subtype of laminin G domains, which is sometimes found in association with thrombospondin-type laminin G domains (IPR012679 from INTERPRO).; PDB: 3POY_A 3QCW_B 3R05_B 3ASI_A 3MW4_B 3MW3_A 1QU0_D 1DYK_A 1OKQ_A 3SH4_A ....
Probab=78.57 E-value=4.4 Score=39.58 Aligned_cols=62 Identities=21% Similarity=0.394 Sum_probs=39.9
Q ss_pred cccceecCCceEEEEEEEeccccceeeeecccccccccccccccccccccCCceEEeecCCCCccc
Q 000135 1497 ISATSIADGRWHIVTMTIDADIGEATCYLDGGFDGYQTGLALSAGNSIWEEGAEVWVGVRPPTDMD 1562 (2087)
Q Consensus 1497 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1562 (2087)
.....++||+||.|+++.+... ++-++|+.-.-.+.-.... ...=+....+++|-.|+....
T Consensus 46 ~~~~~~~dg~wh~v~i~~~~~~--~~l~Vd~~~~~~~~~~~~~--~~~~~~~~~l~iGg~~~~~~~ 107 (128)
T PF02210_consen 46 FSNSNLNDGQWHKVSISRDGNR--VTLTVDGQSVSSESLPSSS--SDSLDPDGSLYIGGLPESNQP 107 (128)
T ss_dssp ECSSSSTSSSEEEEEEEEETTE--EEEEETTSEEEEEESSSTT--HHCBESEEEEEESSTTTTCTC
T ss_pred ccCccccccceeEEEEEEeeee--EEEEecCccceEEeccccc--eecccCCCCEEEecccCcccc
Confidence 3455699999999999887765 7788887643322111111 013345667999999886543
No 20
>cd02248 Peptidase_C1A Peptidase C1A subfamily (MEROPS database nomenclature); composed of cysteine peptidases (CPs) similar to papain, including the mammalian CPs (cathepsins B, C, F, H, L, K, O, S, V, X and W). Papain is an endopeptidase with specific substrate preferences, primarily for bulky hydrophobic or aromatic residues at the S2 subsite, a hydrophobic pocket in papain that accommodates the P2 sidechain of the substrate (the second residue away from the scissile bond). Most members of the papain subfamily are endopeptidases. Some exceptions to this rule can be explained by specific details of the catalytic domains like the occluding loop in cathepsin B which confers an additional carboxydipeptidyl activity and the mini-chain of cathepsin H resulting in an N-terminal exopeptidase activity. Papain-like CPs have different functions in various organisms. Plant CPs are used to mobilize storage proteins in seeds. Parasitic CPs act extracellularly to help invade tissues and cells, to h
Probab=68.94 E-value=14 Score=40.15 Aligned_cols=43 Identities=16% Similarity=0.404 Sum_probs=35.6
Q ss_pred cccCceeEEEEEEEECCEEEEEEecCCCCCccccCCCCCCCcccchHHhhhhcCCCCCCCCeEEEehhh
Q 000135 1925 IVQGHAYSILQVREVDGHKLVQIRNPWANEVEWNGPWSDSSPEWTDRMKHKLKHVPQSKDGIFWMSWQD 1993 (2087)
Q Consensus 1925 LVsGHAYSVLDVrEVdG~RLVRLRNPWG~~~EWKGdWSD~S~eWTeeLKkkL~~~~~sDDGeFWMSfED 1993 (2087)
...+|+=.|++..+-.+.+...+||-||. +| .++|-|||+.++
T Consensus 156 ~~~~Hav~iVGy~~~~~~~ywiv~NSWG~--~W------------------------G~~Gy~~i~~~~ 198 (210)
T cd02248 156 TNLNHAVLLVGYGTENGVDYWIVKNSWGT--SW------------------------GEKGYIRIARGS 198 (210)
T ss_pred CcCCEEEEEEEEeecCCceEEEEEcCCCC--cc------------------------ccCcEEEEEcCC
Confidence 44689999999988767889999999994 44 356999999887
No 21
>PRK12438 hypothetical protein; Provisional
Probab=66.49 E-value=8.3 Score=52.46 Aligned_cols=43 Identities=26% Similarity=0.318 Sum_probs=33.7
Q ss_pred hhhhHHHHHHhhcccc---ee--------------ecCccccccceeeeehhHHHHHHhh
Q 000135 894 PLVCIPALLSLCSGLL---KW--------------KDDDWKLSRGVYVFITIGLVLLLGA 936 (2087)
Q Consensus 894 pl~~ipa~~~l~~gl~---kw--------------~dd~w~~s~~~y~f~~~gl~ll~~a 936 (2087)
.++.+|++++|..|+. .| +|--....-|-|+|.-=-+-+|++.
T Consensus 115 ~~~~v~~~~gl~~g~~~~~~W~~~Llfln~~~FG~~DP~Fg~DigFYvF~LPf~~~l~~~ 174 (991)
T PRK12438 115 FGWGIAVTLGVVCGLIAQFDWVTVQLFVHGGTFGIVDPEFGYDIGFYVFDLPFYRSVLNW 174 (991)
T ss_pred HHHHHHHHHHHHHHHHHHhHHHHHHHHhCCCCCCCCCCCCCCcceEEEEecHHHHHHHHH
Confidence 3677899999999975 57 8888999999999986665555543
No 22
>PF03699 UPF0182: Uncharacterised protein family (UPF0182); InterPro: IPR005372 This family contains uncharacterised integral membrane proteins.; GO: 0016021 integral to membrane
Probab=62.48 E-value=13 Score=49.78 Aligned_cols=62 Identities=23% Similarity=0.379 Sum_probs=37.6
Q ss_pred hhcCceEEEEEeccCCCCCChh-----h----HHHhh-----hhhhhHHHHHHhhcccc---ee--------------ec
Q 000135 865 AFCGASYLEVVKSREDQVPTKG-----D----FLAAL-----LPLVCIPALLSLCSGLL---KW--------------KD 913 (2087)
Q Consensus 865 ~fc~~sy~~v~~sr~~~~p~~~-----d----fl~al-----lpl~~ipa~~~l~~gl~---kw--------------~d 913 (2087)
.+...+.+-..+.|....|... + +.... +-++.++++++++.|+. .| +|
T Consensus 62 ~~~~~~~~~a~r~~~~~~~~~~~~~~~~~l~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~W~~~L~f~n~~~Fg~~D 141 (774)
T PF03699_consen 62 LFVFLNLWLAYRSRPKFRPPSPEQQRSDPLERYRELIEPRRRWVIIGVSLVLGLFAGLSASSQWETILLFLNGTPFGITD 141 (774)
T ss_pred HHHHHHHHHHHhcccccccccccccccchHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHhhHHHHHHHhCCCCCCCCC
Confidence 4455555556666665444322 1 22222 22456777888887764 35 78
Q ss_pred Cccccccceeeee
Q 000135 914 DDWKLSRGVYVFI 926 (2087)
Q Consensus 914 d~w~~s~~~y~f~ 926 (2087)
--....-|-|+|.
T Consensus 142 P~Fg~Di~FYvF~ 154 (774)
T PF03699_consen 142 PIFGKDISFYVFS 154 (774)
T ss_pred CCCCCCceeeeeh
Confidence 8888889999985
No 23
>PF02057 Glyco_hydro_59: Glycosyl hydrolase family 59; InterPro: IPR001286 O-Glycosyl hydrolases 3.2.1. from EC are a widespread group of enzymes that hydrolyse the glycosidic bond between two or more carbohydrates, or between a carbohydrate and a non-carbohydrate moiety. A classification system for glycosyl hydrolases, based on sequence similarity, has led to the definition of 85 different families [, ]. This classification is available on the CAZy (CArbohydrate-Active EnZymes) web site. Glycoside hydrolase family 59 GH59 from CAZY comprises enzymes with only one known activity; galactocerebrosidase (3.2.1.46 from EC). Globoid cell leukodystrophy (Krabbe disease) is a severe, autosomal recessive disorder that results from deficiency of galactocerebrosidase (GALC) activity [, , ]. GALC is responsible for the lysosomal catabolism of certain galactolipids, including galactosylceramide and psychosine [].; GO: 0004336 galactosylceramidase activity, 0006683 galactosylceramide catabolic process; PDB: 3ZR6_A 3ZR5_A.
Probab=60.02 E-value=25 Score=46.40 Aligned_cols=91 Identities=29% Similarity=0.418 Sum_probs=49.8
Q ss_pred ceeEEEEEEEee-cCCCceeeecccccchhhhheeeccccccc-----cccceeEEEEEecCCceeeeeeeccccceecC
Q 000135 1431 QKVSFSIAVMIQ-PESGPVCLLGTEFQKKVCWEILVAGSEQGI-----EAGQVGLRLITKGDRQTTVAKDWSISATSIAD 1504 (2087)
Q Consensus 1431 ~~~~~~~~~~~~-~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~-----~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1504 (2087)
+.++.|.-||+. |++|-|+|.|.--+.- | ...+.+|+ +.|.-- ||+.-..+++-+... +.+.-
T Consensus 542 ~NytVs~DV~ie~~~~ggv~lagRv~~~g-~----~~~~~~G~~f~v~~~G~w~---vt~d~~~~~~l~~G~---~~~~~ 610 (669)
T PF02057_consen 542 SNYTVSCDVYIETPDTGGVFLAGRVNKGG-C----DVRSARGYFFWVYANGTWS---VTSDLAGTTTLASGT---ADIGA 610 (669)
T ss_dssp -EEEEEEEEEE-STTT-EEEEEEEE---G-G----GGGG-EEEEEEEETTTEEE---EEEETTS-SEEEEEE----S--T
T ss_pred eEEEEEEEEEeccCCcCcEEEEEeecccc-c----ccCCCCeEEEEEEcCCcEE---EeccCCCcEEEeeee---ecccC
Confidence 346777888887 5899999987654332 1 12223332 222221 333333333434433 45777
Q ss_pred CceEEEEEEEeccccceeeeeccccccccc
Q 000135 1505 GRWHIVTMTIDADIGEATCYLDGGFDGYQT 1534 (2087)
Q Consensus 1505 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1534 (2087)
||||+++++|+-++ +++++||..-++-.
T Consensus 611 ~~WhtltL~~~g~~--~ta~lng~~l~~~~ 638 (669)
T PF02057_consen 611 GKWHTLTLTISGST--ATAMLNGTVLWTDV 638 (669)
T ss_dssp T-EEEEEEEEETTE--EEEEETTEEEEEEE
T ss_pred CeEEEEEEEEECCE--EEEEECCEEeEEec
Confidence 99999999998887 99999998766543
No 24
>KOG4326 consensus Mitochondrial F1F0-ATP synthase, subunit e [Energy production and conversion]
Probab=58.00 E-value=15 Score=36.86 Aligned_cols=17 Identities=47% Similarity=0.597 Sum_probs=14.4
Q ss_pred cchhhhHhhhccchhhh
Q 000135 1242 KFGGYLLLLLGLTAKAE 1258 (2087)
Q Consensus 1242 ~~~~~~~~~~~~~~~~~ 1258 (2087)
|||.|-+|+||.+--|-
T Consensus 13 kfGRysaL~lGvaYGa~ 29 (81)
T KOG4326|consen 13 KFGRYSALSLGVAYGAF 29 (81)
T ss_pred HhhHHHHHHHHHHHhHH
Confidence 89999999999876554
No 25
>TIGR00805 oat sodium-independent organic anion transporter. Proteins of the OAT family catalyze the Na+-independent facilitated transport of organic anions such as bromosulfobromophthalein and prostaglandins as well as conjugated and unconjugated bile acids (taurocholate and cholate, respectively). These transporters have been characterized in mammals, but homologues are present in C. elegans and A. thaliana. Some of the mammalian proteins exhibit a high degree of tissue specificity. For example, the rat OAT is found at high levels in liver and kidney and at lower levels in other tissues. These proteins possess 10-12 putative a-helical transmembrane spanners. They may catalyze electrogenic anion uniport or anion exchange.
Probab=57.32 E-value=26 Score=45.45 Aligned_cols=94 Identities=16% Similarity=0.277 Sum_probs=54.1
Q ss_pred ccceeeeehhHHHHHHhhhhheeeeechhH----------HHHHHHHHHHHHHHHHh-hhhcccccceeeehhhHHHHHH
Q 000135 919 SRGVYVFITIGLVLLLGAISAVIVVITPWT----------IGVAFLLLLLLIVLAIG-VIHHWASNNFYLTRTQMFFVCF 987 (2087)
Q Consensus 919 s~~~y~f~~~gl~ll~~aisa~~~~~~pw~----------~gvafll~~~~~v~~ig-vih~wasnnfyl~r~~~~~~~~ 987 (2087)
+...|++..++..+..++..++...+..+. .|..+.+..+.. ..+| .+.-|.++.+-+..++++..|+
T Consensus 328 ~n~~f~~~~l~~~~~~~~~~~~~~~lP~yl~~~~g~s~~~ag~l~~~~~i~~-~~vG~~l~G~l~~r~~~~~~~~~~~~~ 406 (633)
T TIGR00805 328 CNPIYMLVILAQVIDSLAFNGYITFLPKYLENQYGISSAEANFLIGVVNLPA-AGLGYLIGGFIMKKFKLNVKKAAYFAI 406 (633)
T ss_pred cCcHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHcCCcHHHHHHHhhhhhhhH-HHHHHhhhhheeeeecccHHHHHHHHH
Confidence 344566666666666655555444333332 333322222211 2233 3566777777777777777666
Q ss_pred HHHHHHHH----HHHhhhcCCCCccccchhH
Q 000135 988 LAFLLGLA----AFLVGWFDDKPFVGASVGY 1014 (2087)
Q Consensus 988 ~~~~~~~~----~~~~~~~~~~~~~~~~~~~ 1014 (2087)
+..+++++ .|++| -++-|+.|..+.|
T Consensus 407 ~~~~~~~~~~~~~~~~~-C~~~~~agv~~~y 436 (633)
T TIGR00805 407 CLSTLSYLLCSPLFLIG-CESAPVAGVNNPS 436 (633)
T ss_pred HHHHHHHHHHHHHHeec-CCCCccceeeccC
Confidence 65555543 45555 5889999999987
No 26
>PF00112 Peptidase_C1: Papain family cysteine protease This is family C1 in the peptidase classification. ; InterPro: IPR000668 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Cysteine peptidases have characteristic molecular topologies, which can be seen not only in their three-dimensional structures, but commonly also in the two-dimensional structures. These are peptidases in which the nucleophile is the sulphydryl group of a cysteine residue. Cysteine proteases are divided into clans (proteins which are evolutionary related), and further sub-divided into families, on the basis of the architecture of their catalytic dyad or triad []. This group of proteins belong to the peptidase family C1, sub-family C1A (papain family, clan CA). It includes proteins classed as non-peptidase homologs. These are have either been shown experimentally to lack peptidase activity or lack one or more of the active site residues. The papain family has a wide variety of activities, including broad-range (papain) and narrow-range endo-peptidases, aminopeptidases, dipeptidyl peptidases and enzymes with both exo- and endo-peptidase activity []. Members of the papain family are widespread, found in baculovirus [], eubacteria, yeast, and practically all protozoa, plants and mammals []. The proteins are typically lysosomal or secreted, and proteolytic cleavage of the propeptide is required for enzyme activation, although bleomycin hydrolase is cytosolic in fungi and mammals []. Papain-like cysteine proteinases are essentially synthesised as inactive proenzymes (zymogens) with N-terminal propeptide regions. The activation process of these enzymes includes the removal of propeptide regions. The propeptide regions serve a variety of functions in vivo and in vitro. The pro-region is required for the proper folding of the newly synthesised enzyme, the inactivation of the peptidase domain and stabilisation of the enzyme against denaturing at neutral to alkaline pH conditions. Amino acid residues within the pro-region mediate their membrane association, and play a role in the transport of the proenzyme to lysosomes. Among the most notable features of propeptides is their ability to inhibit the activity of their cognate enzymes and that certain propeptides exhibit high selectivity for inhibition of the peptidases from which they originate []. The catalytic residues of papain are Cys-25 and His-159, other important residues being Gln-19, which helps form the 'oxyanion hole', and Asn-175, which orientates the imidazole ring of His-159. ; GO: 0008234 cysteine-type peptidase activity, 0006508 proteolysis; PDB: 3MOR_B 3HHI_B 1S4V_A 3F75_A 1MEG_A 1PCI_C 1PPO_A 3HD3_B 1F29_A 1EWL_A ....
Probab=54.55 E-value=35 Score=36.96 Aligned_cols=44 Identities=25% Similarity=0.572 Sum_probs=36.1
Q ss_pred cccCceeEEEEEEEECCEEEEEEecCCCCCccccCCCCCCCcccchHHhhhhcCCCCCCCCeEEEehhhh
Q 000135 1925 IVQGHAYSILQVREVDGHKLVQIRNPWANEVEWNGPWSDSSPEWTDRMKHKLKHVPQSKDGIFWMSWQDF 1994 (2087)
Q Consensus 1925 LVsGHAYSVLDVrEVdG~RLVRLRNPWG~~~EWKGdWSD~S~eWTeeLKkkL~~~~~sDDGeFWMSfEDF 1994 (2087)
-..+|+-.|++..+-.+.....+||-||. .| .++|.|||+.++.
T Consensus 163 ~~~~Hav~iVGy~~~~~~~~wiv~NSWG~--~W------------------------G~~Gy~~i~~~~~ 206 (219)
T PF00112_consen 163 ESGGHAVLIVGYDDENGKGYWIVKNSWGT--DW------------------------GDNGYFRISYDYN 206 (219)
T ss_dssp SSEEEEEEEEEEEEETTEEEEEEE-SBTT--TS------------------------TBTTEEEEESSSS
T ss_pred ccccccccccccccccceeeEeeehhhCC--cc------------------------CCCeEEEEeeCCC
Confidence 46699999999999888899999999994 34 3579999999865
No 27
>PF09323 DUF1980: Domain of unknown function (DUF1980); InterPro: IPR015402 Members of this occur in gene pairs with members of PF03773 from PFAM. The N-terminal region contains several predicted transmembrane helix regions while the few invariant residues (G, CxxD, and W) occur in the C-terminal region. Members of this family are found in a set of prokaryotic hypothetical proteins. Their exact function has not, as yet, been defined.
Probab=46.46 E-value=55 Score=36.46 Aligned_cols=57 Identities=19% Similarity=0.386 Sum_probs=43.9
Q ss_pred HHHHHHHHHHHhHHHHHHHhhhhhheeeccceehhhHHHHHHHHHHHHHHHHHHhhhhccc
Q 000135 64 FLALSAWMVVISPVAVLIMWGSWLIVILGRDIIGLAIIMAGTALLLAFYSIMLWWRTQWQS 124 (2087)
Q Consensus 64 ~l~l~a~~~v~sp~~~l~~wg~~~~~~~~~~~~gla~~m~g~~~~la~y~i~~w~~tqwqs 124 (2087)
+|.|++|.+.+ +-+.+-.-+...+.++.+.++++.+.+.++||.+.++.|+|.+=++
T Consensus 4 ~liL~~~~~l~----~~l~~sG~i~~YI~P~~~~~~~~a~i~l~ilai~q~~~~~~~~~~~ 60 (182)
T PF09323_consen 4 FLILLGFGILL----FYLILSGKILLYIHPRYIPLLYFAAILLLILAIVQLWRWFRPKRRK 60 (182)
T ss_pred HHHHHHHHHHH----HHHHHhCcHHHHhCccHHHHHHHHHHHHHHHHHHHHHHHHhccccc
Confidence 34455554432 2344556677788999999999999999999999999999988774
No 28
>PTZ00334 trans-sialidase; Provisional
Probab=45.43 E-value=30 Score=46.57 Aligned_cols=77 Identities=23% Similarity=0.403 Sum_probs=50.6
Q ss_pred CCceEEEEEEEeccccceeeeeccccccc-ccccccccccccccCCceEEeecCCCCccccc--cC-CCcccc--ccchh
Q 000135 1504 DGRWHIVTMTIDADIGEATCYLDGGFDGY-QTGLALSAGNSIWEEGAEVWVGVRPPTDMDVF--GR-SDSEGA--ESKMH 1577 (2087)
Q Consensus 1504 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~-~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~--~~-~~~~~~--~~~~~ 1577 (2087)
-|+-|.|.++++-. .+.+.|+||--=|- ++ ++.. +.|.++--| |- ..+.+. ++++-
T Consensus 642 ~~k~yqVal~L~~G-~~gsvYVDG~~vg~~~~--~l~~---------------~~~~~IshFyiGgdg~~~~~~~~~~VT 703 (780)
T PTZ00334 642 PETTHQVAIVLRNG-KQGSAYVDGQRVGDASC--ELKN---------------TDSKGISHFYIGGDGGSAGSKEDVPVT 703 (780)
T ss_pred CCCeEEEEEEEeCC-CeEEEEECCEEecCccc--ccCC---------------CCCcccceEEECCCccccccCCCCCEE
Confidence 36779999999542 26899999976552 22 2221 124444444 11 111111 46788
Q ss_pred hheehhhcccCChHHHHHHhh
Q 000135 1578 IMDVFLWGRCLTEDEIASLYS 1598 (2087)
Q Consensus 1578 ~~~~~~~~~cLtedEi~~~~~ 1598 (2087)
...|||.-|+|+++||.+|..
T Consensus 704 V~NVlLYNRpL~~~Ei~~l~~ 724 (780)
T PTZ00334 704 ATNVLLYNRPLDDNEIRVLNA 724 (780)
T ss_pred EeEeEEeCCCCCHHHHHhhhc
Confidence 999999999999999999975
No 29
>PF07946 DUF1682: Protein of unknown function (DUF1682); InterPro: IPR012879 The members of this family are all hypothetical eukaryotic proteins of unknown function. One member (Q920S6 from SWISSPROT) is described as being an adipocyte-specific protein, but no evidence of this was found.
Probab=44.09 E-value=33 Score=41.36 Aligned_cols=10 Identities=50% Similarity=0.826 Sum_probs=6.1
Q ss_pred HHHhhHHhhh
Q 000135 1332 KEERKWKEIE 1341 (2087)
Q Consensus 1332 ~~~~~~~~~~ 1341 (2087)
.|.|||.|-|
T Consensus 305 eeQrK~eeKe 314 (321)
T PF07946_consen 305 EEQRKYEEKE 314 (321)
T ss_pred HHHHHHHHHH
Confidence 5666666655
No 30
>COG1390 NtpE Archaeal/vacuolar-type H+-ATPase subunit E [Energy production and conversion]
Probab=43.88 E-value=2.2e+02 Score=32.93 Aligned_cols=113 Identities=25% Similarity=0.262 Sum_probs=73.3
Q ss_pred hhhhhhhhhhhhhhhhcccCCCcCChhhhhccCchhhhhHHHHHHhhhhhhhhHHHHHHHHHhhhcccHHHHHHHHHHHH
Q 000135 1255 AKAERVQDEVRLRLFLDSIGFSDLSAKKIKKWMPEDRRQFEIIQESYIREKEMEEEILMQRREEEGRGKERRKALLEKEE 1334 (2087)
Q Consensus 1255 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1334 (2087)
.||+++.+|.+-+ .++=..|-++.-+-.++.+.+.++-|.+...||=-..-+-.-||+.|-.+||
T Consensus 17 eeak~I~~eA~~e---------------ae~i~~ea~~~~~~~~~~~~~~~~~ea~~~~~~iis~A~le~r~~~Le~~ee 81 (194)
T COG1390 17 EEAEEILEEAREE---------------AEKIKEEAKREAEEAIEEILRKAEKEAERERQRIISSALLEARRKLLEAKEE 81 (194)
T ss_pred HHHHHHHHHHHHH---------------HHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 5677777776543 3333456777778888899988887777777665555455555555555544
Q ss_pred --hhHHhhhhhhcccCCCCCchHH--HHHHHHHHHhcCCccccchhhhHHHHHH
Q 000135 1335 --RKWKEIEASLISSIPNAGNREA--AAMAAAVRAVGGDSVLEDSFARERVSSI 1384 (2087)
Q Consensus 1335 --~~~~~~~~~~~~~~~~~~~~~~--~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1384 (2087)
..|-+..-.-|..+++...-++ +-|.+++....|+.+. -+.+++.+.+
T Consensus 82 ~l~~~~~~~~e~L~~i~~~~~~~~l~~ll~~~~~~~~~~~~i--V~~~e~d~~~ 133 (194)
T COG1390 82 ILESVFEAVEEKLRNIASDPEYESLQELLIEALEKLLGGELV--VYLNEKDKAL 133 (194)
T ss_pred HHHHHHHHHHHHHHcCcCCcchHHHHHHHHHHHHhcCCCCeE--EEeCcccHHH
Confidence 2344455556667777666666 6688888888777766 4555555555
No 31
>PF04156 IncA: IncA protein; InterPro: IPR007285 Chlamydia trachomatis is an obligate intracellular bacterium that develops within a parasitophorous vacuole termed an inclusion. The inclusion is nonfusogenic with lysosomes but intercepts lipids from a host cell exocytic pathway. Initiation of chlamydial development is concurrent with modification of the inclusion membrane by a set of C. trachomatis-encoded proteins collectively designated Incs. One of these Incs, IncA (Inclusion membrane protein A), is functionally associated with the homotypic fusion of inclusions [].
Probab=41.50 E-value=19 Score=39.47 Aligned_cols=22 Identities=27% Similarity=0.668 Sum_probs=13.9
Q ss_pred eehhHHHHHHhhhhheeeeech
Q 000135 925 FITIGLVLLLGAISAVIVVITP 946 (2087)
Q Consensus 925 f~~~gl~ll~~aisa~~~~~~p 946 (2087)
++.+|++|+.++|.+++.++.+
T Consensus 11 ~iilgilli~~gI~~Lv~~~~~ 32 (191)
T PF04156_consen 11 LIILGILLIASGIAALVLFISG 32 (191)
T ss_pred HHHHHHHHHHHHHHHHHHHHhh
Confidence 4556777777777775554444
No 32
>PF11770 GAPT: GRB2-binding adapter (GAPT); InterPro: IPR021082 This entry represents a family of transmembrane proteins which bind the growth factor receptor-bound protein 2 (GRB2) in B cells []. In contrast to other transmembrane adaptor proteins, GAPT, which this entry represents, is not phosphorylated upon BCR ligation. It associates with GRB2 constitutively through its proline-rich region [].
Probab=41.50 E-value=18 Score=40.40 Aligned_cols=14 Identities=43% Similarity=1.054 Sum_probs=6.6
Q ss_pred HHHHHHHhhhhccc
Q 000135 958 LLIVLAIGVIHHWA 971 (2087)
Q Consensus 958 ~~~v~~igvih~wa 971 (2087)
||++.|||++-||-
T Consensus 21 lLl~cgiGcvwhwk 34 (158)
T PF11770_consen 21 LLLLCGIGCVWHWK 34 (158)
T ss_pred HHHHHhcceEEEee
Confidence 44444555544443
No 33
>PF00054 Laminin_G_1: Laminin G domain; InterPro: IPR012679 Laminins are large heterotrimeric glycoproteins involved in basement membrane function []. The laminin globular (G) domain can be found in one to several copies in various laminin family members, which includes a large number of extracellular proteins. The C terminus of laminin alpha chain contains a tandem repeat of five laminin G domains, which are critical for heparin-binding and cell attachment activity []. Laminin alpha4 is distributed in a variety of tissues including peripheral nerves, dorsal root ganglion, skeletal muscle and capillaries; in the neuromuscular junction, it is required for synaptic specialisation []. The structure of the laminin-G domain has been predicted to resemble that of pentraxin []. Laminin G domains can vary in their function, and a variety of binding functions has been ascribed to different LamG modules. For example, the laminin alpha1 and alpha2 chains each has five C-teminal laminin G domains, where only domains LG4 and LG5 contain binding sites for heparin, sulphatides and the cell surface receptor dystroglycan []. Laminin G-containing proteins appear to have a wide variety of roles in cell adhesion, signalling, migration, assembly and differentiation. This entry represents one subtype of laminin G domains, which is sometimes found in association with thrombospondin-type laminin G domains (IPR012680 from INTERPRO).; PDB: 1OKQ_A 1DYK_A 2C5D_A 1H30_A 1LHW_A 1KDK_A 1LHU_A 1KDM_A 1LHO_A 1D2S_A ....
Probab=41.45 E-value=34 Score=35.57 Aligned_cols=51 Identities=22% Similarity=0.473 Sum_probs=33.9
Q ss_pred ccccceeEEEEEecCCceeeeeeeccccceecCCceEEEEEEEeccccceeeeeccccc
Q 000135 1472 IEAGQVGLRLITKGDRQTTVAKDWSISATSIADGRWHIVTMTIDADIGEATCYLDGGFD 1530 (2087)
Q Consensus 1472 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1530 (2087)
|..|++=+|.- -|.+..++ ..+.+ |.||+||.|++..... +++-.+||...
T Consensus 26 L~~G~l~~~~~-~G~~~~~~----~~~~~-i~dg~wh~v~~~r~~~--~~~L~Vd~~~~ 76 (131)
T PF00054_consen 26 LRDGRLEFRYN-LGSGPASL----RSPQK-INDGKWHTVSVSRNGR--NGSLSVDGEEV 76 (131)
T ss_dssp EETTEEEEEEE-SSSEEEEE----EESSE-TTSSSEEEEEEEEETT--EEEEEETTSEE
T ss_pred EECCEEEEEEe-CCCcccee----cCCCc-cCCCcceEEEEEEcCc--EEEEEECCccc
Confidence 66788777763 33333333 12334 9999999999988754 55667888765
No 34
>cd08045 TAF4 TATA Binding Protein (TBP) Associated Factor 4 (TAF4) is one of several TAFs that bind TBP and is involved in forming Transcription Factor IID (TFIID) complex. The TATA Binding Protein (TBP) Associated Factor 4 (TAF4) is one of several TAFs that bind TBP and are involved in forming the Transcription Factor IID (TFIID) complex. TFIID is one of seven General Transcription Factors (GTF) (TFIIA, TFIIB, TFIID, TFIIE, TFIIF, and TFIID) that are involved in accurate initiation of transcription by RNA polymerase II in eukaryote. TFIID plays an important role in the recognition of promoter DNA and assembly of the pre-initiation complex. TFIID complex is composed of the TBP and at least 13 TAFs. TAFs from various species were originally named by their predicted molecular weight or their electrophoretic mobility in polyacrylamide gels. A new, unified nomenclature for the pol II TAFs has been suggested to show the relationship between TAF orthologs and paralogs. Several hypotheses are
Probab=40.97 E-value=15 Score=41.86 Aligned_cols=44 Identities=30% Similarity=0.357 Sum_probs=31.5
Q ss_pred chHHHHHHHHHHHhcCCccccchhhhHHHHHHHHHHHHHHHHHHHHhcCCcceEEeeCCCCCccCccc
Q 000135 1353 NREAAAMAAAVRAVGGDSVLEDSFARERVSSIARRIRTAQLARRALQTGITGAICVLDDEPTTSGRHC 1420 (2087)
Q Consensus 1353 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1420 (2087)
.|--||=++|.-|+||+.-.- ....+...+|+|++||+.+..+.
T Consensus 166 ~r~r~AN~tA~~AiG~~kk~~------------------------~~i~~rD~l~~LE~e~~~~~s~l 209 (212)
T cd08045 166 MRHRAANATALAAIGGRKKKK------------------------RRITMRDVLFVLEREPRYSKSAL 209 (212)
T ss_pred HHHHHHHHHHHHHhCCCCccc------------------------ceeeHHHHHHHHHhCchhhhhhh
Confidence 344566677777899987765 33445677889999999876653
No 35
>PRK00068 hypothetical protein; Validated
Probab=40.95 E-value=29 Score=47.61 Aligned_cols=40 Identities=30% Similarity=0.542 Sum_probs=29.1
Q ss_pred hhhHHHHHHhhcccc---ee--------------ecCccccccceeeeehhHHHHHH
Q 000135 895 LVCIPALLSLCSGLL---KW--------------KDDDWKLSRGVYVFITIGLVLLL 934 (2087)
Q Consensus 895 l~~ipa~~~l~~gl~---kw--------------~dd~w~~s~~~y~f~~~gl~ll~ 934 (2087)
++.||++++|..|+. .| +|--....-|-|+|.-=-+-+|+
T Consensus 114 ~~~i~~~~gl~~g~~~~~~W~~~L~fln~~~Fg~~DP~Fg~DigFY~F~LPf~~~l~ 170 (970)
T PRK00068 114 LIGIPSFIGLLAGIFAQSYWYRIQLFLNGVDFGVKDPQFGKDLSFYAFKLPFYRSLL 170 (970)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHHhCCCCCCCCCCCCCCcceEEEEehHHHHHHH
Confidence 467788888888876 36 78888889999999754443333
No 36
>PTZ00266 NIMA-related protein kinase; Provisional
Probab=40.95 E-value=45 Score=46.21 Aligned_cols=16 Identities=25% Similarity=0.607 Sum_probs=6.8
Q ss_pred HHHHhhHHHhhhcccc
Q 000135 823 WLMASAIALVVTGVLP 838 (2087)
Q Consensus 823 w~~~s~i~lv~t~~~p 838 (2087)
|.++..+--++||-.|
T Consensus 227 WSLG~ILYELLTGk~P 242 (1021)
T PTZ00266 227 WALGCIIYELCSGKTP 242 (1021)
T ss_pred HHHHHHHHHHHHCCCC
Confidence 4444333334444444
No 37
>KOG1029 consensus Endocytic adaptor protein intersectin [Signal transduction mechanisms; Intracellular trafficking, secretion, and vesicular transport]
Probab=40.23 E-value=37 Score=45.35 Aligned_cols=43 Identities=23% Similarity=0.360 Sum_probs=20.0
Q ss_pred hhhhHHHHHHhhhhhhhhHHHHHHHHHhhhcccHHHHHHHHHH
Q 000135 1290 DRRQFEIIQESYIREKEMEEEILMQRREEEGRGKERRKALLEK 1332 (2087)
Q Consensus 1290 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1332 (2087)
|||+=|.-....-++-|.|.++-+||--|.+|-.||||.+.++
T Consensus 356 ekkererqEqErk~qlElekqLerQReiE~qrEEerkkeie~r 398 (1118)
T KOG1029|consen 356 EKKERERQEQERKAQLELEKQLERQREIERQREEERKKEIERR 398 (1118)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 4444443333344444555555555544444444555544443
No 38
>PLN02316 synthase/transferase
Probab=40.20 E-value=45 Score=46.24 Aligned_cols=45 Identities=29% Similarity=0.325 Sum_probs=21.6
Q ss_pred HHHHHHHHhhhccc-----HHHHHHHHHHHHhhHHhhhhhh----------cccCCCCCc
Q 000135 1309 EEILMQRREEEGRG-----KERRKALLEKEERKWKEIEASL----------ISSIPNAGN 1353 (2087)
Q Consensus 1309 ~~~~~~~~~~~~~~-----~~~~~~~~~~~~~~~~~~~~~~----------~~~~~~~~~ 1353 (2087)
++..+|||.||.|- +.+.||-.||..||-+|+=... .++.|.||+
T Consensus 270 ~~~ee~~r~~~~kaa~~a~~a~akae~~~~~~~~~~~~~~~~~~~~~~~~~~P~~~~aG~ 329 (1036)
T PLN02316 270 RQAEEQRRREEEKAAMEADRAQAKAEVEKRREKLQNLLKKASRSADNVWYIEPSEFKAGD 329 (1036)
T ss_pred HHHHHHHHHHHHhhhhhhhhhhhhHHHHHHHHHHHHHHhhhhhcccceEEecCCCcCCCC
Confidence 33345666665553 3444544444444444443222 356666664
No 39
>cd02620 Peptidase_C1A_CathepsinB Cathepsin B group; composed of cathepsin B and similar proteins, including tubulointerstitial nephritis antigen (TIN-Ag). Cathepsin B is a lysosomal papain-like cysteine peptidase which is expressed in all tissues and functions primarily as an exopeptidase through its carboxydipeptidyl activity. Together with other cathepsins, it is involved in the degradation of proteins, proenzyme activation, Ag processing, metabolism and apoptosis. Cathepsin B has been implicated in a number of human diseases such as cancer, rheumatoid arthritis, osteoporosis and Alzheimer's disease. The unique carboxydipeptidyl activity of cathepsin B is attributed to the presence of an occluding loop in its active site which favors the binding of the C-termini of substrate proteins. Some members of this group do not possess the occluding loop. TIN-Ag is an extracellular matrix basement protein which was originally identified as a target Ag involved in anti-tubular basement membrane
Probab=38.66 E-value=69 Score=36.70 Aligned_cols=27 Identities=26% Similarity=0.389 Sum_probs=23.6
Q ss_pred cCceeEEEEEEEECCEEEEEEecCCCC
Q 000135 1927 QGHAYSILQVREVDGHKLVQIRNPWAN 1953 (2087)
Q Consensus 1927 sGHAYSVLDVrEVdG~RLVRLRNPWG~ 1953 (2087)
.+||=.|++..+-+|.+...+||-||.
T Consensus 184 ~~HaV~iVGyg~~~g~~YWivrNSWG~ 210 (236)
T cd02620 184 GGHAVKIIGWGVENGVPYWLAANSWGT 210 (236)
T ss_pred CCeEEEEEEEeccCCeeEEEEEeCCCC
Confidence 479999999976678889999999994
No 40
>PF05297 Herpes_LMP1: Herpesvirus latent membrane protein 1 (LMP1); InterPro: IPR007961 This family consists of several latent membrane protein 1 or LMP1s mostly from Epstein-Barr virus (strain GD1) (HHV-4) (Human herpesvirus 4). LMP1 of HHV-4 is a 62-65 kDa plasma membrane protein possessing six membrane spanning regions, a short cytoplasmic N terminus and a long cytoplasmic carboxy tail of 200 amino acids. HHV-4 virus latent membrane protein 1 (LMP1) is essential for HHV-4 mediated transformation and has been associated with several cases of malignancies. HHV-4-like viruses in Macaca fascicularis (Cynomolgus monkeys) have been associated with high lymphoma rates in immunosuppressed monkeys [].; GO: 0019087 transformation of host cell by virus, 0016021 integral to membrane; PDB: 1CZY_E 1ZMS_B.
Probab=37.70 E-value=11 Score=45.42 Aligned_cols=52 Identities=21% Similarity=0.425 Sum_probs=0.0
Q ss_pred HHHHHHHHHHHHHHHHhhhhcccccceeeehhhHHHHHHHHHHHHHHHHHhhhc
Q 000135 949 IGVAFLLLLLLIVLAIGVIHHWASNNFYLTRTQMFFVCFLAFLLGLAAFLVGWF 1002 (2087)
Q Consensus 949 ~gvafll~~~~~v~~igvih~wasnnfyl~r~~~~~~~~~~~~~~~~~~~~~~~ 1002 (2087)
+|..|+.+.++++++|=.. .|-=.++=-|-.+ ++..++||+||+.-.++..+
T Consensus 107 ~Gi~~l~l~~lLaL~vW~Y-m~lLr~~GAs~Wt-iLaFcLAF~LaivlLIIAv~ 158 (381)
T PF05297_consen 107 VGIVILFLCCLLALGVWFY-MWLLRELGASFWT-ILAFCLAFLLAIVLLIIAVL 158 (381)
T ss_dssp ------------------------------------------------------
T ss_pred HHHHHHHHHHHHHHHHHHH-HHHHHHhhhHHHH-HHHHHHHHHHHHHHHHHHHH
Confidence 6777777777776665332 4433332223333 34445678887766665554
No 41
>PF09472 MtrF: Tetrahydromethanopterin S-methyltransferase, F subunit (MtrF); InterPro: IPR013347 Many archaea have evolved energy-yielding pathways marked by one-carbon biochemistry featuring novel cofactors and enzymes. This domain is mostly found in MtrF, where it covers the entire length of the protein. This polypeptide is one of eight subunits of the N5-methyltetrahydromethanopterin: coenzyme M methyltransferase complex found in methanogenic archaea. This is a membrane-associated enzyme complex that uses methyl-transfer reactions to drive a sodium-ion pump []. MtrF itself is involved in the transfer of the methyl group from N5-methyltetrahydromethanopterin to coenzyme M. Subsequently, methane is produced by two-electron reduction of the methyl moiety in methyl-coenzyme M by another enzyme, methyl-coenzyme M reductase. In some organisms this domain is found at the C-terminal region of what appears to be a fusion of the MtrA and MtrF proteins [, ]. The function of these proteins is unknown, though it is likely that they are involved in C1 metabolism.; GO: 0030269 tetrahydromethanopterin S-methyltransferase activity, 0015948 methanogenesis, 0016020 membrane
Probab=36.67 E-value=12 Score=36.69 Aligned_cols=47 Identities=21% Similarity=0.360 Sum_probs=39.9
Q ss_pred cccccCCcc-cCCCCccccccccchhHHHHHHhhHHHhhhcccchhhc
Q 000135 796 LEDLGYKGW-TGEPNSFASPYASSVYLGWLMASAIALVVTGVLPIVSW 842 (2087)
Q Consensus 796 ~~~~~~~~~-~~~~~~~~spy~~~~~~gw~~~s~i~lv~t~~~p~vsw 842 (2087)
.||++||.= -++.+...|--.++-..|.++...+|+|+.++.|+.-|
T Consensus 17 vedi~Yk~qLiaR~~kL~SGv~~~~~~GfaiG~~~AlvLv~ip~~l~~ 64 (64)
T PF09472_consen 17 VEDIRYKAQLIARDQKLESGVMATGIKGFAIGFLFALVLVGIPILLMF 64 (64)
T ss_pred HHHHHHHHHHhhhcchhHHHHhhhhhHHHHHHHHHHHHHHHHHHHHhC
Confidence 489999863 45667788888899999999999999999999888766
No 42
>cd02698 Peptidase_C1A_CathepsinX Cathepsin X; the only papain-like lysosomal cysteine peptidase exhibiting carboxymonopeptidase activity. It can also act as a carboxydipeptidase, like cathepsin B, but has been shown to preferentially cleave substrates through a monopeptidyl carboxypeptidase pathway. The propeptide region of cathepsin X, the shortest among papain-like peptidases, is covalently attached to the active site cysteine in the inactive form of the enzyme. Little is known about the biological function of cathepsin X. Some studies point to a role in early tumorigenesis. A more recent study indicates that cathepsin X expression is restricted to immune cells suggesting a role in phagocytosis and the regulation of the immune response.
Probab=36.57 E-value=83 Score=36.16 Aligned_cols=42 Identities=21% Similarity=0.464 Sum_probs=33.3
Q ss_pred cCceeEEEEEEEEC-CEEEEEEecCCCCCccccCCCCCCCcccchHHhhhhcCCCCCCCCeEEEehhhh
Q 000135 1927 QGHAYSILQVREVD-GHKLVQIRNPWANEVEWNGPWSDSSPEWTDRMKHKLKHVPQSKDGIFWMSWQDF 1994 (2087)
Q Consensus 1927 sGHAYSVLDVrEVd-G~RLVRLRNPWG~~~EWKGdWSD~S~eWTeeLKkkL~~~~~sDDGeFWMSfEDF 1994 (2087)
.+|+=.|++.-+.+ |.+.-.+||-||. .| .++|-|+|....+
T Consensus 178 ~~HaV~IVGyG~~~~g~~YWiikNSWG~--~W------------------------Ge~Gy~~i~rg~~ 220 (239)
T cd02698 178 INHIISVAGWGVDENGVEYWIVRNSWGE--PW------------------------GERGWFRIVTSSY 220 (239)
T ss_pred CCeEEEEEEEEecCCCCEEEEEEcCCCc--cc------------------------CcCceEEEEccCC
Confidence 48999999997665 7899999999994 44 3578899976653
No 43
>KOG1144 consensus Translation initiation factor 5B (eIF-5B) [Translation, ribosomal structure and biogenesis]
Probab=35.07 E-value=92 Score=42.15 Aligned_cols=17 Identities=24% Similarity=0.251 Sum_probs=9.6
Q ss_pred eeeeecccccccccccc
Q 000135 1521 ATCYLDGGFDGYQTGLA 1537 (2087)
Q Consensus 1521 ~~~~~~~~~~~~~~~~~ 1537 (2087)
++--+||-+|-+-.-+.
T Consensus 397 ~~~~~~~d~dd~ee~~~ 413 (1064)
T KOG1144|consen 397 VDLAIDGDDDDDEEELQ 413 (1064)
T ss_pred ccccccccccchhhhhc
Confidence 33446666776655444
No 44
>PF09323 DUF1980: Domain of unknown function (DUF1980); InterPro: IPR015402 Members of this occur in gene pairs with members of PF03773 from PFAM. The N-terminal region contains several predicted transmembrane helix regions while the few invariant residues (G, CxxD, and W) occur in the C-terminal region. Members of this family are found in a set of prokaryotic hypothetical proteins. Their exact function has not, as yet, been defined.
Probab=34.89 E-value=70 Score=35.69 Aligned_cols=65 Identities=28% Similarity=0.348 Sum_probs=37.9
Q ss_pred HHHHHHHHHhhhhcccccce--eee-hhhHHHHHHHHHHHHHHHH-HhhhcCCCCcc-----------ccchhHHHHHHH
Q 000135 956 LLLLIVLAIGVIHHWASNNF--YLT-RTQMFFVCFLAFLLGLAAF-LVGWFDDKPFV-----------GASVGYFTFLFL 1020 (2087)
Q Consensus 956 ~~~~~v~~igvih~wasnnf--yl~-r~~~~~~~~~~~~~~~~~~-~~~~~~~~~~~-----------~~~~~~~~~~~~ 1020 (2087)
+|+|+.+++-.+|.|.+.+. |+. |+.-+.+....+++.||.+ +..|+..+.-. .-..+|+.|++-
T Consensus 4 ~liL~~~~~l~~~l~~sG~i~~YI~P~~~~~~~~a~i~l~ilai~q~~~~~~~~~~~~~~h~h~~~~~~~~~~y~l~~iP 83 (182)
T PF09323_consen 4 FLILLGFGILLFYLILSGKILLYIHPRYIPLLYFAAILLLILAIVQLWRWFRPKRRKEDCHDHGHSKSKKLWSYFLFLIP 83 (182)
T ss_pred HHHHHHHHHHHHHHHHhCcHHHHhCccHHHHHHHHHHHHHHHHHHHHHHHHhcccccccccccccccccccHHHHHHHHH
Confidence 46677777888899998864 554 4444444444444444444 34556555443 345667776663
No 45
>PF05875 Ceramidase: Ceramidase; InterPro: IPR008901 This entry consists of several ceramidases. Ceramidases are enzymes involved in regulating cellular levels of ceramides, sphingoid bases, and their phosphates.; GO: 0016811 hydrolase activity, acting on carbon-nitrogen (but not peptide) bonds, in linear amides, 0006672 ceramide metabolic process, 0016021 integral to membrane
Probab=32.88 E-value=50 Score=38.37 Aligned_cols=143 Identities=21% Similarity=0.195 Sum_probs=73.8
Q ss_pred CCCCccccccccchhHHHHHHhhHHHhhhcccchhhceeecccccchhhHHHHHHHHHhhhcCceEEEEEeccCCCCCCh
Q 000135 806 GEPNSFASPYASSVYLGWLMASAIALVVTGVLPIVSWFSTYRFSLSSAICVGIFAAVLVAFCGASYLEVVKSREDQVPTK 885 (2087)
Q Consensus 806 ~~~~~~~spy~~~~~~gw~~~s~i~lv~t~~~p~vswf~tyrf~~~sav~~~~f~~vl~~fc~~sy~~v~~sr~~~~p~~ 885 (2087)
|++|+..|||-...+= + -|-++ ..++++.-|....|-.+.....+....+++|.+++.-|=- --++..|
T Consensus 14 CE~nY~~s~yiAEf~N--t-lSNl~---fi~~al~gl~~~~~~~~~~~~~l~~~~l~~VGiGS~~FHa-Tl~~~~q---- 82 (262)
T PF05875_consen 14 CEENYVVSPYIAEFWN--T-LSNLA---FIVAALYGLYLARRRGLERRFALLYLGLALVGIGSFLFHA-TLSYWTQ---- 82 (262)
T ss_pred chhccccCcccchHHH--H-HHHHH---HHHHHHHHHHHHhhccccchhHHHHHHHHHHHHhHHHHHh-ChhhhHH----
Confidence 6889999999765432 1 22222 3335566666666666666666666667777665554432 2222223
Q ss_pred hhHHHhhhhhhhHHHHHHhhcccceeecCccccccceeeeehhHHHHHHhhhhheeeee--chhHHHHHHHHHHHHHHHH
Q 000135 886 GDFLAALLPLVCIPALLSLCSGLLKWKDDDWKLSRGVYVFITIGLVLLLGAISAVIVVI--TPWTIGVAFLLLLLLIVLA 963 (2087)
Q Consensus 886 ~dfl~allpl~~ipa~~~l~~gl~kw~dd~w~~s~~~y~f~~~gl~ll~~aisa~~~~~--~pw~~gvafll~~~~~v~~ 963 (2087)
|.--|| -+...++-+|-|-++.. -+++.-..+++.|.... +++.+.... +|..-.++|..+.+++++-
T Consensus 83 ---l~DelP-----Ml~~~~~~~~~~~~~~~-~~~~~~~~~~~~L~~~~-~~~t~~~~~~~~p~~~~~~f~~~~~~~~~~ 152 (262)
T PF05875_consen 83 ---LLDELP-----MLWATLLFLYIVLTRRY-SSPRYRLALPLLLFIYA-VVVTVLYFVLDNPVFHQIAFASLVLLVILR 152 (262)
T ss_pred ---Hhhhhh-----HHHHHHHHHHHHhcccc-cCchhhHHHHHHHHHHH-HHHHHHHhhhccchhhhhhHHHHHHHHHHH
Confidence 222233 33333444444444433 12222223344443333 334434444 7888778887776666655
Q ss_pred Hhh-hhc
Q 000135 964 IGV-IHH 969 (2087)
Q Consensus 964 igv-ih~ 969 (2087)
... +++
T Consensus 153 ~~~~~~~ 159 (262)
T PF05875_consen 153 SIYLIRR 159 (262)
T ss_pred HHHHHHH
Confidence 554 444
No 46
>COG4870 Cysteine protease [Posttranslational modification, protein turnover, chaperones]
Probab=31.90 E-value=45 Score=41.58 Aligned_cols=49 Identities=31% Similarity=0.577 Sum_probs=35.5
Q ss_pred cCcccCceeEEEEEEEEC----------CEEEEEEecCCCCCccccCCCCCCCcccchHHhhhhcCCCCCCCCeEEEehh
Q 000135 1923 SGIVQGHAYSILQVREVD----------GHKLVQIRNPWANEVEWNGPWSDSSPEWTDRMKHKLKHVPQSKDGIFWMSWQ 1992 (2087)
Q Consensus 1923 ~GLVsGHAYSVLDVrEVd----------G~RLVRLRNPWG~~~EWKGdWSD~S~eWTeeLKkkL~~~~~sDDGeFWMSfE 1992 (2087)
.+...|||=.|++..+-. |.-=+++||-||. .| .++|-|||+++
T Consensus 260 s~~~~gHAv~iVGyDDs~~~n~~~~~~~g~GAfiikNSWGt--~w------------------------G~~GYfwisY~ 313 (372)
T COG4870 260 SGENWGHAVLIVGYDDSFDINNFKYGPPGDGAFIIKNSWGT--NW------------------------GENGYFWISYY 313 (372)
T ss_pred ccccccceEEEEeccccccccccccCCCCCceEEEECcccc--cc------------------------ccCceEEEEee
Confidence 345679999999886531 2236889999994 33 35799999998
Q ss_pred hhhhc
Q 000135 1993 DFQIH 1997 (2087)
Q Consensus 1993 DFLky 1997 (2087)
+-..-
T Consensus 314 ya~~g 318 (372)
T COG4870 314 YALNG 318 (372)
T ss_pred ecccc
Confidence 87654
No 47
>PF14023 DUF4239: Protein of unknown function (DUF4239)
Probab=31.88 E-value=1.2e+02 Score=33.93 Aligned_cols=31 Identities=32% Similarity=0.531 Sum_probs=25.0
Q ss_pred hhhHHHHHHHHHHHHHHHHHhhhcCCCCcccc
Q 000135 979 RTQMFFVCFLAFLLGLAAFLVGWFDDKPFVGA 1010 (2087)
Q Consensus 979 r~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1010 (2087)
+.+++.++++|++++++-|++= --|.||.|.
T Consensus 167 ~~~~~~~~l~a~~i~~~l~li~-~ld~Pf~G~ 197 (209)
T PF14023_consen 167 RAHLIAIALFAASIALALFLIL-DLDNPFSGP 197 (209)
T ss_pred hHHHHHHHHHHHHHHHHHHHHH-HhcCCCCCC
Confidence 5778888888888888888864 468999995
No 48
>PF09586 YfhO: Bacterial membrane protein YfhO; InterPro: IPR018580 The yfhO gene is transcribed in Difco sporulation medium and the transcription is affected by the YvrGHb two-component system []. Some members of this family have been annotated as putative ABC transporter permease proteins.
Probab=31.02 E-value=1.3e+02 Score=40.13 Aligned_cols=24 Identities=33% Similarity=0.375 Sum_probs=17.1
Q ss_pred cccccchhhHHHHHHHHHh-hhcCc
Q 000135 846 YRFSLSSAICVGIFAAVLV-AFCGA 869 (2087)
Q Consensus 846 yrf~~~sav~~~~f~~vl~-~fc~~ 869 (2087)
.||-.++.+.+|+-+++|+ ++++-
T Consensus 214 ~~~~~~~ilg~~lsa~~llP~~~~~ 238 (843)
T PF09586_consen 214 LRFIGSSILGVGLSAFLLLPTILSL 238 (843)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 5677777777788788777 66543
No 49
>PF09991 DUF2232: Predicted membrane protein (DUF2232); InterPro: IPR018710 This family of bacterial and eukaryotic proteins has no known fucntion; however this signature belongs to a Pfam Gx transporter clan.
Probab=30.95 E-value=51 Score=37.62 Aligned_cols=87 Identities=17% Similarity=0.305 Sum_probs=45.7
Q ss_pred ccccccceeeeehhHHHHHHhhhhheeeeechhHHHHHHHHHHHHHHHHHhhhhcccccceeeehhhHHHHHHHHHHHH-
Q 000135 915 DWKLSRGVYVFITIGLVLLLGAISAVIVVITPWTIGVAFLLLLLLIVLAIGVIHHWASNNFYLTRTQMFFVCFLAFLLG- 993 (2087)
Q Consensus 915 ~w~~s~~~y~f~~~gl~ll~~aisa~~~~~~pw~~gvafll~~~~~v~~igvih~wasnnfyl~r~~~~~~~~~~~~~~- 993 (2087)
.|++++..-.+..+++++.+-.....+-...-...-+..++..++++-+++++|+|..+. -++|.=-.+..++.+++.
T Consensus 199 ~~~lP~~~~~~~i~~~~~~l~~~~~~~~~~~~i~~Nl~~v~~~l~~~qGla~~~~~~~~~-~~~~~~~~l~~~~~i~~~~ 277 (290)
T PF09991_consen 199 EWRLPRWLIWLLIVALALSLVGGGFGGSWLQIIGLNLLIVLSFLFFIQGLAVIHFFLKRR-KMSKFLRVLLYILLILFPF 277 (290)
T ss_pred HHhCcHHHHHHHHHHHHHHHHhcccchHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHc-CCcHHHHHHHHHHHHHHHH
Confidence 488887654333333333321111100011223345666777788888999999998776 666654333333333332
Q ss_pred --HHHHHhhhc
Q 000135 994 --LAAFLVGWF 1002 (2087)
Q Consensus 994 --~~~~~~~~~ 1002 (2087)
..-.++|.+
T Consensus 278 ~~~~l~~lG~~ 288 (290)
T PF09991_consen 278 LIVILALLGLI 288 (290)
T ss_pred HHHHHHHHHhh
Confidence 334445544
No 50
>COG0815 Lnt Apolipoprotein N-acyltransferase [Cell envelope biogenesis, outer membrane]
Probab=30.69 E-value=1.2e+02 Score=39.43 Aligned_cols=77 Identities=18% Similarity=0.123 Sum_probs=44.2
Q ss_pred hHHHHHHHHHHHHHHHHHHhhhhccchhHHHHHHHHHHHHHhhcceEEEEEecC-------CCCC-----CCCCCcceeh
Q 000135 99 AIIMAGTALLLAFYSIMLWWRTQWQSSRAVAVLLLLAVALLCAYELSAVYVTAG-------SHAS-----DRYSPSGFFF 166 (2087)
Q Consensus 99 a~~m~g~~~~la~y~i~~w~~tqwqs~~a~a~ll~~a~~l~~~~~~~~~yvt~~-------~~~~-----~~~sps~~ff 166 (2087)
..++.+.++.+++|-.+..|-.+ +.+.+..+.. ++--++|..--.+=+| -+.. .++-|-+=-.
T Consensus 97 ~~~~~ll~~~lal~~~l~~~~~~-~~~~~~~~~~----~~w~~~E~lR~~~~tGFpW~~~Gy~q~~~~~l~q~a~i~Gv~ 171 (518)
T COG0815 97 PLLVLLLAAWLALFLLLVAVLTC-RLWFALLVVP----SAWVAAEWLRGWSLTGFPWLLLGYSQWSPSPLLQLASLGGVW 171 (518)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHH-HHhhhhHHHH----HHHHHHHHHHhccCcCCchhhhchhhccCccccceeeccCHH
Confidence 44567888888888888777665 7766665544 3333445333222222 2222 2333333445
Q ss_pred hhhHHHHHhhhhhh
Q 000135 167 GVSAIALAINMLFI 180 (2087)
Q Consensus 167 ~~sai~~~in~l~i 180 (2087)
++|.+.+++|+++.
T Consensus 172 ~lsflvv~~~~~~a 185 (518)
T COG0815 172 LLSFLVVAVNALLA 185 (518)
T ss_pred HHHHHHHHHHHHHH
Confidence 67888888888753
No 51
>KOG2341 consensus TATA box binding protein (TBP)-associated factor, RNA polymerase II [Transcription]
Probab=30.13 E-value=54 Score=42.82 Aligned_cols=27 Identities=19% Similarity=0.031 Sum_probs=17.2
Q ss_pred CCchhhHHHHHHHHhhhccceeEEEEE
Q 000135 1055 KNVSVAFLVLYGVALAIEGWGVVASLK 1081 (2087)
Q Consensus 1055 ~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1081 (2087)
|-+++.++-+++-..-.+|=+.=+.+.
T Consensus 189 ~t~~~~~~~~p~s~~~~~g~~~ppq~~ 215 (563)
T KOG2341|consen 189 KTLPALRLAVPPSNTFSEGSDPPPQLV 215 (563)
T ss_pred hcchHhhccCCCcccccCCCCCCcccc
Confidence 556777777777766666665544443
No 52
>PF12065 DUF3545: Protein of unknown function (DUF3545); InterPro: IPR021932 This family of proteins is functionally uncharacterised. This protein is found in bacteria. Proteins in this family are typically between 60 to 77 amino acids in length. This protein has two completely conserved residues (R and L) that may be functionally important.
Probab=29.14 E-value=24 Score=34.27 Aligned_cols=10 Identities=70% Similarity=1.358 Sum_probs=8.6
Q ss_pred HHhhHHhhhh
Q 000135 1333 EERKWKEIEA 1342 (2087)
Q Consensus 1333 ~~~~~~~~~~ 1342 (2087)
..|||+||||
T Consensus 23 ~KRKWREIEA 32 (59)
T PF12065_consen 23 KKRKWREIEA 32 (59)
T ss_pred cchhHHHHHH
Confidence 4589999998
No 53
>TIGR00570 cdk7 CDK-activating kinase assembly factor MAT1. All proteins in this family for which functions are known are cyclin dependent protein kinases that are components of TFIIH, a complex that is involved in nucleotide excision repair and transcription initiation. Also known as MAT1 (menage a trois 1). This family is based on the phylogenomic analysis of JA Eisen (1999, Ph.D. Thesis, Stanford University).
Probab=29.02 E-value=85 Score=38.56 Aligned_cols=104 Identities=30% Similarity=0.467 Sum_probs=53.4
Q ss_pred ccccCCCCCchhhHhhhhhhhhhhh----hhhcccceeeeeccchhhhHhhhccchhhhhhhhhhhhhhhhcccCCCcCC
Q 000135 1204 RFRHELSSDYDYRREMCTHARILAL----EEAIDTEWVYMWDKFGGYLLLLLGLTAKAERVQDEVRLRLFLDSIGFSDLS 1279 (2087)
Q Consensus 1204 ~~~~~~~~~~~~~~~~~~~~~~~~~----~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1279 (2087)
.|+.-.-.|...-|++-.--||+.. ||-.+|- ..|.-|| |+|.|=| ..| ...|.- .-.
T Consensus 57 ~fr~q~F~D~~vekEV~iRkrv~~i~Nk~e~dF~~l-----~~yNdYL----------E~vEdii-~nL-~~~~d~-~~t 118 (309)
T TIGR00570 57 NFRVQLFEDPTVEKEVDIRKRVLKIYNKREEDFPSL-----REYNDYL----------EEVEDIV-YNL-TNNIDL-ENT 118 (309)
T ss_pred hccccccccHHHHHHHHHHHHHHHHHccchhccCCH-----HHHHHHH----------HHHHHHH-HHh-hcCCcH-HHH
Confidence 3555566777777888887787765 3333321 2344555 2332211 000 001100 113
Q ss_pred hhhhhccCchhhhhHHHHHHhhhhhhhhHHHHHHHHHhhhcccHHHHHHH
Q 000135 1280 AKKIKKWMPEDRRQFEIIQESYIREKEMEEEILMQRREEEGRGKERRKAL 1329 (2087)
Q Consensus 1280 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1329 (2087)
..+|++|--|.+ +.|+++-.|+++ |++.++|+.++|.+-++.|+..
T Consensus 119 e~~l~~y~~~n~---~~I~~n~~~~~~-e~~~~~~~~~~E~~~~~~rr~~ 164 (309)
T TIGR00570 119 KKKIETYQKENK---DVIQKNKEKSTR-EQEELEEALEFEKEEEEQRRLL 164 (309)
T ss_pred HHHHHHHHHHhH---HHHHHHHHHHHh-HHHHHHHHHHHHHHHHHHHHHH
Confidence 456666655544 458888888776 4455555555555555444333
No 54
>PF04405 ScdA_N: Domain of Unknown function (DUF542) ; InterPro: IPR007500 This is a domain of unknown function found at the N terminus of genes involved in cell wall development and nitrous oxide protection. ScdA is required for normal cell growth and development; mutants have an increased level of peptidoglycan cross-linking and aberrant cellular morphology suggesting a role for ScdA in cell wall metabolism []. NorA1, NorA2, and YtfE are involved in the nitrous oxide response. NorA1 and NorA2, which are similar to YtfE, are co-transcribed with the membrane-bound nitrous oxide (NO) reductases. The genes appear to be involved in NO protection but their function is unknown [, ].
Probab=28.60 E-value=38 Score=32.12 Aligned_cols=33 Identities=36% Similarity=0.660 Sum_probs=29.3
Q ss_pred cChhhHHHHhhhc----ccCchHhHhhhhhcCCCcch
Q 000135 511 NDPRITSMLKKRA----REGDRELTSLLQDKGLDPNF 543 (2087)
Q Consensus 511 ~~p~~~~~lk~~~----~~g~~el~~llqdkgldpnf 543 (2087)
++|+-++.++|-+ -.|++-|..-.+.+|+||+-
T Consensus 11 ~~p~~a~vf~~~gIDfCCgG~~~L~eA~~~~~ld~~~ 47 (56)
T PF04405_consen 11 EDPRAARVFRKYGIDFCCGGNRSLEEACEEKGLDPEE 47 (56)
T ss_pred HChHHHHHHHHcCCcccCCCCchHHHHHHHcCCCHHH
Confidence 6899999999777 67999999999999999974
No 55
>cd06899 lectin_legume_LecRK_Arcelin_ConA legume lectins, lectin-like receptor kinases, arcelin, concanavalinA, and alpha-amylase inhibitor. This alignment model includes the legume lectins (also known as agglutinins), the arcelin (also known as phytohemagglutinin-L) family of lectin-like defense proteins, the LecRK family of lectin-like receptor kinases, concanavalinA (ConA), and an alpha-amylase inhibitor. Arcelin is a major seed glycoprotein discovered in kidney beans (Phaseolus vulgaris) that has insecticidal properties and protects the seeds from predation by larvae of various bruchids. Arcelin is devoid of monosaccharide binding properties and lacks a key metal-binding loop that is present in other members of this family. Phytohaemagglutinin (PHA) is a lectin found in plants, especially beans, that affects cell metabolism by inducing mitosis and by altering the permeability of the cell membrane to various proteins. PHA agglutinates most mammalian red blood cell types by bindin
Probab=28.36 E-value=2e+02 Score=33.37 Aligned_cols=37 Identities=14% Similarity=0.129 Sum_probs=30.3
Q ss_pred eeeeccccceecCCceEEEEEEEeccccceeeeeccc
Q 000135 1492 AKDWSISATSIADGRWHIVTMTIDADIGEATCYLDGG 1528 (2087)
Q Consensus 1492 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1528 (2087)
+..|......+.||++|.|.|.-|+.+..-+.||+..
T Consensus 150 ~~~~~~~~~~l~~g~~~~v~I~Y~~~~~~L~V~l~~~ 186 (236)
T cd06899 150 AGYWDDDGGKLKSGKPMQAWIDYDSSSKRLSVTLAYS 186 (236)
T ss_pred eeccccccccccCCCeEEEEEEEcCCCCEEEEEEEeC
Confidence 3556555445789999999999999999999999854
No 56
>PF05154 TM2: TM2 domain; InterPro: IPR007829 This domain is composed of a pair of transmembrane alpha helices connected by a short linker. The function of this domain is unknown, however it occurs in a wide range or protein contexts.
Probab=27.99 E-value=20 Score=32.91 Aligned_cols=33 Identities=36% Similarity=0.609 Sum_probs=22.4
Q ss_pred cchhhHHhhhccceeeeeeechhhhccchh---HHHHHHH
Q 000135 290 QSRVAALFVAGTSRVFLICFGVHYWYLGHC---ISYAVVA 326 (2087)
Q Consensus 290 ~~~~~~~~va~~~r~~li~fg~~~w~lghc---i~y~~~a 326 (2087)
||+.++.+.+- |+-.||+|.+|+||= +.|.++.
T Consensus 3 K~~~~a~lL~~----~lG~~G~hrfYlg~~~~g~~~l~~~ 38 (51)
T PF05154_consen 3 KSKWIAYLLSF----FLGWFGLHRFYLGKYGKGILYLLTF 38 (51)
T ss_pred cCHHHHHHHHH----HHhhccccceecCchHHHHHHHHHH
Confidence 56666666542 566899999999985 4444444
No 57
>PF11877 DUF3397: Protein of unknown function (DUF3397); InterPro: IPR024515 This family of bacterial proteins is currently functionally uncharacterised.
Probab=26.88 E-value=90 Score=32.90 Aligned_cols=96 Identities=18% Similarity=0.193 Sum_probs=55.7
Q ss_pred hhhHHHHHHhhcccceeecCccccccceeeeehhHHHHHHhhhhheeeeechhHHHHHHHHHHHHHHHHHhhhhcccccc
Q 000135 895 LVCIPALLSLCSGLLKWKDDDWKLSRGVYVFITIGLVLLLGAISAVIVVITPWTIGVAFLLLLLLIVLAIGVIHHWASNN 974 (2087)
Q Consensus 895 l~~ipa~~~l~~gl~kw~dd~w~~s~~~y~f~~~gl~ll~~aisa~~~~~~pw~~gvafll~~~~~v~~igvih~wasnn 974 (2087)
++.+|.+.-+...++++|..+|+.-+.+-+ ...++..++..+...+..-..+--.++++++++..+.+.|.....|
T Consensus 9 i~l~p~~~~iiv~~~~l~~~~~~~~~a~D~----~~~fli~~i~~ls~~~~~~s~lpy~~l~~~ll~i~l~~~~~~~~~~ 84 (116)
T PF11877_consen 9 IFLIPFLGFIIVYFFKLKRRKKAFHKAPDV----TTPFLIFSIHLLSNNIFGHSFLPYLLLVLLLLAIILAIYQARKKGE 84 (116)
T ss_pred HHHHHHHHHHHHHHHHhhhhhhhhhhhHHH----HHHHHHHHHHHHHHHHhchhHHHHHHHHHHHHHHHHHHHHHHHcCc
Confidence 344455555555558999988876554433 3445555666654444333344444455555555556688889999
Q ss_pred eeeehhh-----HHHHHHHHHHHHH
Q 000135 975 FYLTRTQ-----MFFVCFLAFLLGL 994 (2087)
Q Consensus 975 fyl~r~~-----~~~~~~~~~~~~~ 994 (2087)
|+..|.= +.|.|+..+=+++
T Consensus 85 i~~~k~~k~~WR~~Fll~~~~Yi~l 109 (116)
T PF11877_consen 85 ISYKKFFKKFWRLGFLLTFFLYIGL 109 (116)
T ss_pred chhhHHHHHHHHHHHHHHHHHHHHH
Confidence 9999863 4444444443333
No 58
>PF14402 7TM_transglut: 7 transmembrane helices usually fused to an inactive transglutaminase
Probab=26.30 E-value=88 Score=38.47 Aligned_cols=54 Identities=26% Similarity=0.503 Sum_probs=41.4
Q ss_pred eechhHHHHHHH-------HHHHHHHHHHhhhhcccccceeeehhhHHHHHHHHHHHHHHHHHhhh
Q 000135 943 VITPWTIGVAFL-------LLLLLIVLAIGVIHHWASNNFYLTRTQMFFVCFLAFLLGLAAFLVGW 1001 (2087)
Q Consensus 943 ~~~pw~~gvafl-------l~~~~~v~~igvih~wasnnfyl~r~~~~~~~~~~~~~~~~~~~~~~ 1001 (2087)
|.-|.-|.+||. ++++++++++|.+-+ +||+|..+++|-=+|-++....++++.
T Consensus 147 TFmPVLIAlAF~eT~L~~Gli~FllIV~~GL~iR-----~yLs~LnLLlV~RisaVli~VI~ii~~ 207 (313)
T PF14402_consen 147 TFMPVLIALAFRETQLLWGLILFLLIVAIGLLIR-----SYLSHLNLLLVPRISAVLIVVILIIAA 207 (313)
T ss_pred chHHHHHHHHHHHhhhHHHHHHHHHHHHHHHHHH-----HHHHhhhhHHHHHHHHHHHHHHHHHHH
Confidence 556777777775 678888999999766 699999999998777777666665554
No 59
>PF06439 DUF1080: Domain of Unknown Function (DUF1080); InterPro: IPR010496 This is a family of proteins of unknown function.; PDB: 3IMM_B 3NMB_A 3S5Q_A 3OSD_A 3HBK_A 3H3L_A 3U1X_A.
Probab=26.25 E-value=2.2e+02 Score=30.41 Aligned_cols=102 Identities=17% Similarity=0.248 Sum_probs=52.5
Q ss_pred ccCcccccccccccccceeEEEEEEEeecCCCceeee-cc-----cccchhhhheeecccccc----ccccceeEEEEEe
Q 000135 1415 TSGRHCGQIDASICQSQKVSFSIAVMIQPESGPVCLL-GT-----EFQKKVCWEILVAGSEQG----IEAGQVGLRLITK 1484 (2087)
Q Consensus 1415 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~-~~-----~~~~~~~~~~~~~~~~~~----~~~~~~~~~~~~~ 1484 (2087)
..+.+.|.+=... ......+++-++++| +|-..++ -. +.....|.|+-+.....+ -..|.+=-+
T Consensus 38 ~~~~~~~~l~~~~-~~~df~l~~d~k~~~-~~~sGi~~r~~~~~~~~~~~~gy~~~i~~~~~~~~~~~~~G~~~~~---- 111 (185)
T PF06439_consen 38 SSGSGGGYLYTDK-KFSDFELEVDFKITP-GGNSGIFFRAQSPGDGQDWNNGYEFQIDNSGGGTGLPNSTGSLYDE---- 111 (185)
T ss_dssp GGESSS--EEESS-EBSSEEEEEEEEE-T-T-EEEEEEEESSECCSSGGGTSEEEEEE-TTTCSTTTTSTTSBTTT----
T ss_pred cCCCCcceEEECC-ccccEEEEEEEEECC-CCCeEEEEEeccccCCCCcceEEEEEEECCCCccCCCCccceEEEe----
Confidence 3444555444443 556677888888754 4433332 22 245667888877766555 111111000
Q ss_pred cCCceeeeeeeccccceecCCceEEEEEEEeccccceeeeecccc
Q 000135 1485 GDRQTTVAKDWSISATSIADGRWHIVTMTIDADIGEATCYLDGGF 1529 (2087)
Q Consensus 1485 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1529 (2087)
-.++ .-.-....+..|+||.++|++..+. .++|+||..
T Consensus 112 ~~~~-----~~~~~~~~~~~~~W~~~~I~~~g~~--i~v~vnG~~ 149 (185)
T PF06439_consen 112 PPWQ-----LEPSVNVAIPPGEWNTVRIVVKGNR--ITVWVNGKP 149 (185)
T ss_dssp B-TC-----B-SSS--S--TTSEEEEEEEEETTE--EEEEETTEE
T ss_pred cccc-----ccccccccCCCCceEEEEEEEECCE--EEEEECCEE
Confidence 0000 0122344578899999999998776 889999964
No 60
>TIGR00917 2A060601 Niemann-Pick C type protein family. The model describes Niemann-Pick C type protein in eukaryotes. The defective protein has been associated with Niemann-Pick disease which is described in humans as autosomal recessive lipidosis. It is characterized by the lysosomal accumulation of unestrified cholesterol. It is an integral membrane protein, which indicates that this protein is most likely involved in cholesterol transport or acts as some component of cholesterol homeostasis.
Probab=26.01 E-value=35 Score=47.79 Aligned_cols=79 Identities=23% Similarity=0.296 Sum_probs=43.3
Q ss_pred HHHHHHHHHhh------hhcccccceee-----------ehhh----HH----HHHHHHHHHHHHHHHhhhcCCCCcccc
Q 000135 956 LLLLIVLAIGV------IHHWASNNFYL-----------TRTQ----MF----FVCFLAFLLGLAAFLVGWFDDKPFVGA 1010 (2087)
Q Consensus 956 ~~~~~v~~igv------ih~wasnnfyl-----------~r~~----~~----~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1010 (2087)
++-++|+|||| +|.|...+-.- +..| ++ --.+++-+.-.+||++|.+-+-|-+-
T Consensus 640 v~PFLvL~IGVD~ifilv~~~~r~~~~~~~~~~~~~~~~~~~~ri~~~l~~~G~sI~ltslt~~~aF~~g~~s~~Pavr- 718 (1204)
T TIGR00917 640 VIPFLVLAVGVDNIFILVQTYQRLERFYREVGVDNEQELTLEQQLGRALGEVGPSITLASLSESLAFFLGALSKMPAVR- 718 (1204)
T ss_pred HHHHHHHHHHhhHHHHHHHHHHHhhhccccccccccccCCHHHHHHHHHHHhhHHHHHHHHHHHHHHHHHhccCChHHH-
Confidence 45567889998 56675433210 2212 11 34667777888899999998776442
Q ss_pred chhHHHHHHHhhccceeeeccCCEE
Q 000135 1011 SVGYFTFLFLLAGRALTVLLSPPIV 1035 (2087)
Q Consensus 1011 ~~~~~~~~~~~~~~~~~~~~~~~~~ 1035 (2087)
..|.++-+.++.-=.+++.+-|+++
T Consensus 719 ~F~~~aa~av~~~fll~it~f~alL 743 (1204)
T TIGR00917 719 AFSLFAGLAVFIDFLLQITAFVALL 743 (1204)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 3344443333333333333334433
No 61
>PF02460 Patched: Patched family; InterPro: IPR003392 The transmembrane protein, patched, is a receptor for the morphogene Sonic Hedgehog. In Drosophila melanogaster, this protein associates with the smoothened protein to transduce hedgehog signals, leading to the activation of wingless, decapentaplegic and patched itself. It participates in cell interactions that establish pattern within the segment and imaginal disks during development. The mouse homologue may play a role in epidermal development. The human Niemann-Pick C1 protein, defects in which cause Niemann-Pick type II disease, is also a member of this family. This protein is involved in the intracellular trafficking of cholesterol, and may play a role in vesicular trafficking in glia, a process that may be crucial for maintaining the structural functional integrity of nerve terminals.; GO: 0008158 hedgehog receptor activity, 0016020 membrane
Probab=25.32 E-value=92 Score=41.61 Aligned_cols=53 Identities=28% Similarity=0.419 Sum_probs=35.9
Q ss_pred HHHHHHHHHHhh------hhcccccceeeehhhHH--------HHHHHHHHHHHHHHHhhhcCCCCc
Q 000135 955 LLLLLIVLAIGV------IHHWASNNFYLTRTQMF--------FVCFLAFLLGLAAFLVGWFDDKPF 1007 (2087)
Q Consensus 955 l~~~~~v~~igv------ih~wasnnfyl~r~~~~--------~~~~~~~~~~~~~~~~~~~~~~~~ 1007 (2087)
.+.-++++|||| +|.|-...-..+..+-+ --.++.-+--.+||++|.+-.-|=
T Consensus 282 ~v~PFLvlgIGvDd~Fi~~~~~~~~~~~~~~~er~~~~l~~~g~SitiTslT~~~aF~ig~~t~~pa 348 (798)
T PF02460_consen 282 LVIPFLVLGIGVDDMFIMIHAWRRTSPDLSVEERMAETLAEAGPSITITSLTNALAFAIGAITPIPA 348 (798)
T ss_pred HHHHHHHHHHHHhceEEeHHHHhhhchhccHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHhcCCcHH
Confidence 356677889999 89998776665543222 223444555567899999887773
No 62
>TIGR02916 PEP_his_kin putative PEP-CTERM system histidine kinase. Members of this protein family have a novel N-terminal domain, a single predicted membrane-spanning helix, and a predicted cystosolic histidine kinase domain. We designate this protein PrsK, and its companion DNA-binding response regulator protein (TIGR02915) PrsR. These predicted signal-transducing proteins appear to enable enhancer-dependent transcriptional activation. The prsK gene is often associated with exopolysaccharide biosynthesis genes.
Probab=24.95 E-value=39 Score=43.77 Aligned_cols=36 Identities=11% Similarity=-0.006 Sum_probs=20.3
Q ss_pred hhHHHhhhhhhhHHHHHHhhcccceeecCccccccce
Q 000135 886 GDFLAALLPLVCIPALLSLCSGLLKWKDDDWKLSRGV 922 (2087)
Q Consensus 886 ~dfl~allpl~~ipa~~~l~~gl~kw~dd~w~~s~~~ 922 (2087)
..++..+.|..-++.++.+ .+...+.++++.-++..
T Consensus 58 ~~~~~~l~~~~w~~~l~~~-~~~~~~~~~~~~~~~~~ 93 (679)
T TIGR02916 58 VLVLEVFRDAAWLAFLLTL-LRRPATSGKPFNQRPKL 93 (679)
T ss_pred HHHHHHHHHHHHHHHHHHH-hcccccccCcccchHHH
Confidence 3455555666655555543 34466677777665544
No 63
>PF13801 Metal_resist: Heavy-metal resistance; PDB: 3EPV_C 2Y3D_A 2Y3H_D 2Y3G_B 2Y3B_A 2Y39_A 3LAY_H.
Probab=24.34 E-value=2.6e+02 Score=27.51 Aligned_cols=20 Identities=15% Similarity=0.451 Sum_probs=13.2
Q ss_pred CchhhhhHHHHHHhhhhhhh
Q 000135 1287 MPEDRRQFEIIQESYIREKE 1306 (2087)
Q Consensus 1287 ~~~~~~~~~~~~~~~~~~~~ 1306 (2087)
+||++++++-+.+.|..+-+
T Consensus 43 t~eQ~~~l~~~~~~~~~~~~ 62 (125)
T PF13801_consen 43 TPEQQAKLRALMDEFRQEMR 62 (125)
T ss_dssp THHHHHHHHHHHHHHHHHHH
T ss_pred CHHHHHHHHHHHHHHHHHHH
Confidence 57777777777766665443
No 64
>PF11911 DUF3429: Protein of unknown function (DUF3429); InterPro: IPR021836 This family of proteins are functionally uncharacterised. This protein is found in bacteria and eukaryotes. Proteins in this family are typically between 147 to 245 amino acids in length.
Probab=23.93 E-value=87 Score=34.00 Aligned_cols=44 Identities=16% Similarity=0.286 Sum_probs=35.3
Q ss_pred hhhHHHHHHHHHhhhcCceEEEEEeccCCCCCChhhHHHhhhhh
Q 000135 852 SAICVGIFAAVLVAFCGASYLEVVKSREDQVPTKGDFLAALLPL 895 (2087)
Q Consensus 852 sav~~~~f~~vl~~fc~~sy~~v~~sr~~~~p~~~dfl~allpl 895 (2087)
........++||++|-||+|++.--++++..+....+..+.+|-
T Consensus 35 ~~~~~~~Y~AvILSFLgGv~WG~al~~~~~~~~~~~l~~sv~p~ 78 (142)
T PF11911_consen 35 ALYAFLAYGAVILSFLGGVHWGLALSQPSASPSWRRLIWSVVPS 78 (142)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHhcccccchHHHHHHHHHHH
Confidence 45667778999999999999999888886677777777766553
No 65
>cd01951 lectin_L-type legume lectins. The L-type (legume-type) lectins are a highly diverse family of carbohydrate binding proteins that generally display no enzymatic activity toward the sugars they bind. This family includes arcelin, concanavalinA, the lectin-like receptor kinases, the ERGIC-53/VIP36/EMP46 type1 transmembrane proteins, and an alpha-amylase inhibitor. L-type lectins have a dome-shaped beta-barrel carbohydrate recognition domain with a curved seven-stranded beta-sheet referred to as the "front face" and a flat six-stranded beta-sheet referred to as the "back face". This domain homodimerizes so that adjacent back sheets form a contiguous 12-stranded sheet and homotetramers occur by a back-to-back association of these homodimers. Though L-type lectins exhibit both sequence and structural similarity to one another, their carbohydrate binding specificities differ widely.
Probab=23.64 E-value=3.2e+02 Score=30.92 Aligned_cols=50 Identities=22% Similarity=0.286 Sum_probs=32.5
Q ss_pred CceEEEEEEEeccccceeeeecccccccccccccccccccccCCceEEeec
Q 000135 1505 GRWHIVTMTIDADIGEATCYLDGGFDGYQTGLALSAGNSIWEEGAEVWVGV 1555 (2087)
Q Consensus 1505 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1555 (2087)
|+||.|.|+.|+.++.-+.++|+.-.....-+..++.-.- ....+++||+
T Consensus 154 g~~~~v~I~Y~~~~~~L~v~l~~~~~~~~~~l~~~~~l~~-~~~~~~yvGF 203 (223)
T cd01951 154 GNEHTVRITYDPTTNTLTVYLDNGSTLTSLDITIPVDLIQ-LGPTKAYFGF 203 (223)
T ss_pred CCEEEEEEEEeCCCCEEEEEECCCCccccccEEEeeeecc-cCCCcEEEEE
Confidence 9999999999999999999999764312122222222221 2246777765
No 66
>KOG3011 consensus Ubiquitin-conjugating enzyme [Posttranslational modification, protein turnover, chaperones]
Probab=23.21 E-value=2.1e+02 Score=34.73 Aligned_cols=115 Identities=22% Similarity=0.302 Sum_probs=60.6
Q ss_pred hhhHHHHHHHHHhhhcCceEEEEEeccCCCCCChhhHHHhhhhhhhHHHHHHhhcccceeecCcc---------------
Q 000135 852 SAICVGIFAAVLVAFCGASYLEVVKSREDQVPTKGDFLAALLPLVCIPALLSLCSGLLKWKDDDW--------------- 916 (2087)
Q Consensus 852 sav~~~~f~~vl~~fc~~sy~~v~~sr~~~~p~~~dfl~allpl~~ipa~~~l~~gl~kw~dd~w--------------- 916 (2087)
.|.|.++|+.++...-|+-=.+ +.+..+|--+|=-..-=|++|+|.|--|.|
T Consensus 83 ~~~c~~lf~~~~~~ii~~~~s~-------------~~~~~~La~~aG~i~AD~~SGl~HWaaD~~Gsv~tP~vG~~f~rf 149 (293)
T KOG3011|consen 83 AAGCTTLFVSFAKSIIGGFGSH-------------LWLEPALAAYAGYITADLGSGVYHWAADNYGSVSTPWVGRQFERF 149 (293)
T ss_pred HhhhHHHHHHHHHHHHHhhhhh-------------hhHHHHHHHHHHHHHHhhhcceeEeeccccCccccchhHHHHHHH
Confidence 4568888888777655543211 223333333333334468999999966655
Q ss_pred --------ccccceeeeehhHHHHHHhhhhheeeeechhH-----HHHHHHHHHHHHHHHHhhhhcccccceeeehhhHH
Q 000135 917 --------KLSRGVYVFITIGLVLLLGAISAVIVVITPWT-----IGVAFLLLLLLIVLAIGVIHHWASNNFYLTRTQMF 983 (2087)
Q Consensus 917 --------~~s~~~y~f~~~gl~ll~~aisa~~~~~~pw~-----~gvafll~~~~~v~~igvih~wasnnfyl~r~~~~ 983 (2087)
.+.|.-++=. +-|+--|+-++ |..|=. .=-+|.+.+-+.|+----||.|+---|=|+|.-++
T Consensus 150 reHH~dP~tITr~~f~~~---~~ll~~a~~f~--v~~~d~~~q~~~~h~fV~~~~i~v~~tnQiHkWsHTy~gLP~wVv~ 224 (293)
T KOG3011|consen 150 QEHHKDPWTITRRQFANN---LHLLARAYTFI--VLPLDLAFQDPVFHGFVFLFAICVLFTNQIHKWSHTYSGLPPWVVL 224 (293)
T ss_pred HhccCCcceeeHHHHhhh---hHHHHHhheeE--ecCHHHHhhcccHHHHHHHHHHHHHHHHHHHHHHhhhccCchHHHH
Confidence 4444443333 22222233332 222211 22233333333344445599999988889986554
Q ss_pred H
Q 000135 984 F 984 (2087)
Q Consensus 984 ~ 984 (2087)
+
T Consensus 225 L 225 (293)
T KOG3011|consen 225 L 225 (293)
T ss_pred H
Confidence 3
No 67
>PRK15097 cytochrome d terminal oxidase subunit 1; Provisional
Probab=22.85 E-value=2.4e+02 Score=37.10 Aligned_cols=91 Identities=22% Similarity=0.358 Sum_probs=58.7
Q ss_pred HHHHHHHHHHHHHHHHHhhhhcccccceeeehhhHHHHHHHHHHHHHHHHHhhhcCCCCccccchhHHHHHHHhhcccee
Q 000135 948 TIGVAFLLLLLLIVLAIGVIHHWASNNFYLTRTQMFFVCFLAFLLGLAAFLVGWFDDKPFVGASVGYFTFLFLLAGRALT 1027 (2087)
Q Consensus 948 ~~gvafll~~~~~v~~igvih~wasnnfyl~r~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1027 (2087)
|+|..++++++.+ +|++-.|- +..| +.+-.|-++.++..|...|...||+-.| .||
T Consensus 393 MVg~G~l~~~l~~---~~l~l~~r-~~l~-~~rw~L~~~~~~~plp~iA~~~GWi~tE----------------vGR--- 448 (522)
T PRK15097 393 MVACGFLMLAIIA---LSFWSVIR-NRIG-EKKWLLRAALYGIPLPWIAVEAGWFVAE----------------YGR--- 448 (522)
T ss_pred HHHHHHHHHHHHH---HHHHHHHc-Cccc-cCcHHHHHHHHHHHHHHHHHHhhhhhee----------------cCC---
Confidence 4777766554433 34444443 3434 3355777888899999999999998665 466
Q ss_pred eeccCCEEEecCceeeEEEeecccccCCCchh--------hHHHHHHHHhhhccc
Q 000135 1028 VLLSPPIVVYSPRVLPVYVYDAHADCGKNVSV--------AFLVLYGVALAIEGW 1074 (2087)
Q Consensus 1028 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~--------~~~~~~~~~~~~~~~ 1074 (2087)
-|=+|| .+||++ |+.. ||+. .|.++|++.+..+.|
T Consensus 449 ----QPWiVy--g~l~T~--~avS----~~s~~~v~~sl~~f~~~Y~~L~~~~~~ 491 (522)
T PRK15097 449 ----QPWAIG--EVLPTA--VANS----SLTAGDLLFSMVLICGLYTLFLVAELF 491 (522)
T ss_pred ----CCeEEe--ceeeHh--HhcC----CCCHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 578888 466653 5543 3543 688888877666544
No 68
>PLN00122 serine/threonine protein phosphatase 2A; Provisional
Probab=22.84 E-value=93 Score=35.39 Aligned_cols=22 Identities=32% Similarity=0.605 Sum_probs=16.2
Q ss_pred HHHHHHHHHHHHhhHHhhhhhh
Q 000135 1323 KERRKALLEKEERKWKEIEASL 1344 (2087)
Q Consensus 1323 ~~~~~~~~~~~~~~~~~~~~~~ 1344 (2087)
++++++..+|.|.+|+.||..-
T Consensus 142 ~~~~~~~~~~r~~~W~~le~~A 163 (170)
T PLN00122 142 EAKAKEVEEKREATWKRLEEAA 163 (170)
T ss_pred HHHHHHHHHHHHHHHHHHHHHH
Confidence 3456666688889999998643
No 69
>PF15412 Nse4-Nse3_bdg: Binding domain of Nse4/EID3 to Nse3-MAGE
Probab=22.40 E-value=62 Score=30.52 Aligned_cols=28 Identities=36% Similarity=0.662 Sum_probs=23.8
Q ss_pred eeeecCCCCCHHHHHHHHhhhccCCCcc
Q 000135 182 RMVFNGNGLDVDEYVRRAYKFAYPDGIE 209 (2087)
Q Consensus 182 ~~~~~g~~~dv~eyvr~~y~~a~~d~~e 209 (2087)
++-+.|+++|+||||.+..+|.-.+..+
T Consensus 18 ~lk~~~~~fd~deFv~~l~~fm~~~~~~ 45 (56)
T PF15412_consen 18 NLKFGGSGFDVDEFVSKLKTFMGGNRFE 45 (56)
T ss_pred HhccCCCccCHHHHHHHHHHHhCcccCC
Confidence 4567799999999999999998876665
No 70
>KOG3583 consensus Uncharacterized conserved protein [Function unknown]
Probab=22.34 E-value=1.6e+02 Score=35.08 Aligned_cols=122 Identities=22% Similarity=0.325 Sum_probs=64.4
Q ss_pred ccceeeeeccchhhhHhhhccchhhhhhh-----h--hhhhhhhhccc-CCCcCChhhhhccCchhhhhHHHHHHhhhhh
Q 000135 1233 DTEWVYMWDKFGGYLLLLLGLTAKAERVQ-----D--EVRLRLFLDSI-GFSDLSAKKIKKWMPEDRRQFEIIQESYIRE 1304 (2087)
Q Consensus 1233 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~-----~--~~~~~~~~~~~-~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1304 (2087)
.|-|--|-|||.-.--.+-||+.--..-| . -|-+|+-.|-= -.-....+..--|.- -|--.|+|-
T Consensus 38 ~~~wp~~le~fs~las~ms~l~~~~~k~~~p~lr~~~~~~~~~~~e~detl~r~TeGRVpvfsH-------~lVPdyLRT 110 (279)
T KOG3583|consen 38 KCPWPLMLEKFSTLASFMSSLQSSVRKSGMPHLRSHVLVTQRLQYEPDETLQRATEGRVPVFSH-------ALVPDYLRT 110 (279)
T ss_pred cCccHHHHHHHHHHHHHHHHHHHHHHHccCCccccchhhhhhhhcCchHHHHHHhcCccccccc-------ccchHhhcc
Confidence 35599999999988777888875322111 0 11122211100 000000111111111 123468987
Q ss_pred h---hhHHHHHHHHHhhhcccHHH---H-----------HHHHHHHHhhHHhhhhhhcccCCCCCch-HHHHHHHHH
Q 000135 1305 K---EMEEEILMQRREEEGRGKER---R-----------KALLEKEERKWKEIEASLISSIPNAGNR-EAAAMAAAV 1363 (2087)
Q Consensus 1305 ~---~~~~~~~~~~~~~~~~~~~~---~-----------~~~~~~~~~~~~~~~~~~~~~~~~~~~~-~~~~~~~~~ 1363 (2087)
| |||+|+.|---|...++..- . -.-+.|++|.| +|++..--|-..-|+ |.|++.|||
T Consensus 111 kPdPe~E~~e~ql~~~aa~~saDaa~kQI~~yNK~is~ll~~lsk~~re~--tEs~~~~piqQT~n~~dT~~lVaaV 185 (279)
T KOG3583|consen 111 KPDPEMENEEGQLDGEAAAKSADAAVKQIAAYNKNISGLLNHLSKVDREH--TESAIEKPIQQTYNRDDTAKLVAAV 185 (279)
T ss_pred CCChhhHHHHhhhhhHHhhhhhHHHHHHHHHHHHHHHHHHHHHHHHHHHH--HHhhhcCccccccChhHHHHHHHHH
Confidence 6 89999887665555554321 1 12356788888 888776655555554 456666655
No 71
>PF02387 IncFII_repA: IncFII RepA protein family; InterPro: IPR003446 These proteins are plasmid encoded and essential for plasmid replication, they are also involved in copy control functions [].; GO: 0006276 plasmid maintenance
Probab=22.31 E-value=98 Score=37.54 Aligned_cols=87 Identities=24% Similarity=0.387 Sum_probs=52.3
Q ss_pred hHhhhccchhhhhhhhhhhhhhhhc---ccCCCcCChhhhhccCchhhhhHHHHHHhhhhhhhhHHHHHHHHHhhhcccH
Q 000135 1247 LLLLLGLTAKAERVQDEVRLRLFLD---SIGFSDLSAKKIKKWMPEDRRQFEIIQESYIREKEMEEEILMQRREEEGRGK 1323 (2087)
Q Consensus 1247 ~~~~~~~~~~~~~~~~~~~~~~~~~---~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1323 (2087)
+..++|.+.+.-+-+.+-||+..=+ ..|-..+|..++++..- ++..+..++.|+++..+|+
T Consensus 159 ff~l~gi~~~kl~~~~~~~l~~~~~~~~~~~~~~is~~e~~~r~~----------------~~~~~~~~~~r~~~~~~~~ 222 (281)
T PF02387_consen 159 FFMLLGISEDKLRREQRQRLQWENNGLSKQGEEPISLHEARRRAK----------------EQHRKRALDYRKERRAKGK 222 (281)
T ss_pred HHHHhCCCHHHHHHHHHHHHHHHHHhhhhcccCCCcHHHHHHHHH----------------HHHHHHHHHHHHHhHHHHH
Confidence 3567899888766666666665533 44667777777643222 2335567778888887788
Q ss_pred HHHHH--HHHH-HHhhHHhhhhhhcccCC
Q 000135 1324 ERRKA--LLEK-EERKWKEIEASLISSIP 1349 (2087)
Q Consensus 1324 ~~~~~--~~~~-~~~~~~~~~~~~~~~~~ 1349 (2087)
+|++| +.+. |....++|=.-|+.+.|
T Consensus 223 krk~A~rl~~L~e~~ar~~I~~~Lik~ys 251 (281)
T PF02387_consen 223 KRKRARRLAKLDEDEARQEILRQLIKEYS 251 (281)
T ss_pred HHHHHhhccccCHHHHHHHHHHHHHHHcC
Confidence 77654 2222 22334555555665555
No 72
>PRK10263 DNA translocase FtsK; Provisional
Probab=21.93 E-value=58 Score=46.09 Aligned_cols=30 Identities=20% Similarity=0.412 Sum_probs=23.2
Q ss_pred EEEeeeeehhhchhccceeeeccccccccCCccc
Q 000135 772 VLVICITVFTGSVLALGAIVSAKPLEDLGYKGWT 805 (2087)
Q Consensus 772 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 805 (2087)
.-.+.|++++.+++.+.+++|+.|.|- +|+
T Consensus 23 ~E~~gIlLlllAlfL~lALiSYsPsDP----SwS 52 (1355)
T PRK10263 23 LEALLILIVLFAVWLMAALLSFNPSDP----SWS 52 (1355)
T ss_pred HHHHHHHHHHHHHHHHHHHHhCCccCC----ccc
Confidence 335567778888888999999999774 665
No 73
>PF04123 DUF373: Domain of unknown function (DUF373); InterPro: IPR007254 This archaeal family of unknown function is predicted to be an integral membrane protein with six transmembrane regions.
Probab=21.87 E-value=49 Score=40.90 Aligned_cols=138 Identities=23% Similarity=0.429 Sum_probs=76.1
Q ss_pred eee-ehhHHHHHHhhhhheeeeechhHHHHHHHHHHHHH-HHHHhh---hhccccc---ceeeehhhHHHHHHHHHHHHH
Q 000135 923 YVF-ITIGLVLLLGAISAVIVVITPWTIGVAFLLLLLLI-VLAIGV---IHHWASN---NFYLTRTQMFFVCFLAFLLGL 994 (2087)
Q Consensus 923 y~f-~~~gl~ll~~aisa~~~~~~pw~~gvafll~~~~~-v~~igv---ih~wasn---nfyl~r~~~~~~~~~~~~~~~ 994 (2087)
++| +- |++||+-++.+++-. ..+++++..+++++.+ .=+.|. +.+|.++ .+|-.|.... .-..|.++.+
T Consensus 161 ~~lGvP-G~~lLiy~i~~l~~~-~~~a~~~i~~~iG~yll~kGfgld~~~~~~~~~~~~~l~~g~it~i-tyvva~~l~i 237 (344)
T PF04123_consen 161 TFLGVP-GLILLIYAILALLGY-PAYALGIILLLIGLYLLYKGFGLDDYLREWLERFRESLYEGRITFI-TYVVALLLII 237 (344)
T ss_pred eeecch-HHHHHHHHHHHHHcc-hHHHHHHHHHHHHHHHHHHhcCcHHHHHHHHHHhccccccceeehH-HHHHHHHHHH
Confidence 455 55 999999999986543 3445555555554444 335555 5566554 4666654333 3344444555
Q ss_pred HHHHhhhcC------CCC------ccccchhHHHH--HHHhhccceeeeccCCEEEecCceeeEEEeecccccCCCchhh
Q 000135 995 AAFLVGWFD------DKP------FVGASVGYFTF--LFLLAGRALTVLLSPPIVVYSPRVLPVYVYDAHADCGKNVSVA 1060 (2087)
Q Consensus 995 ~~~~~~~~~------~~~------~~~~~~~~~~~--~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1060 (2087)
.+...|... ..+ |+=.++.||++ +...+||.+.-.+.--...|+--..|.++ .+.
T Consensus 238 ig~i~g~~~~~~~~~~~~~~~~~~f~~~~v~~~~~a~l~~~~G~iid~~l~~~~~~~~~i~~~~~~-----------~a~ 306 (344)
T PF04123_consen 238 IGIIYGYLTLWSYYSISGLIVPGTFLYGSVPWLALAALIASLGKIIDEYLRRDFRLWRYINAPFFV-----------IAI 306 (344)
T ss_pred HHHHHHHHHHHhhccccchHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHccCcchHHHHHHHHHH-----------HHH
Confidence 555555441 111 44455666655 44557887776666555555544444432 455
Q ss_pred HHHHHHHHhhhccc
Q 000135 1061 FLVLYGVALAIEGW 1074 (2087)
Q Consensus 1061 ~~~~~~~~~~~~~~ 1074 (2087)
++++|++..-....
T Consensus 307 ~~v~~~~~~~~l~~ 320 (344)
T PF04123_consen 307 GLVLYGFSAYFLSI 320 (344)
T ss_pred HHHHHHHHHHHHhh
Confidence 56677766554443
No 74
>KOG4661 consensus Hsp27-ERE-TATA-binding protein/Scaffold attachment factor (SAF-B) [Transcription]
Probab=21.82 E-value=1.2e+02 Score=39.69 Aligned_cols=30 Identities=40% Similarity=0.446 Sum_probs=19.8
Q ss_pred HHHHhhhcccHHHHHHHHHHHHhhHHhhhh
Q 000135 1313 MQRREEEGRGKERRKALLEKEERKWKEIEA 1342 (2087)
Q Consensus 1313 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1342 (2087)
+||-+||.--.||||+..|+||+.+-++|.
T Consensus 626 r~RirE~rerEqR~~a~~ERee~eRl~~er 655 (940)
T KOG4661|consen 626 RQRIREEREREQRRKAAVEREELERLKAER 655 (940)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 344444445567888888888887766653
No 75
>PRK11588 hypothetical protein; Provisional
Probab=21.53 E-value=2.6e+02 Score=36.62 Aligned_cols=46 Identities=15% Similarity=0.249 Sum_probs=35.9
Q ss_pred HhhhhhhhHHHHHHhhcccceeecCccccccceeeeehhHHHHHHhhhhheeeeechhHHHHH
Q 000135 890 AALLPLVCIPALLSLCSGLLKWKDDDWKLSRGVYVFITIGLVLLLGAISAVIVVITPWTIGVA 952 (2087)
Q Consensus 890 ~allpl~~ipa~~~l~~gl~kw~dd~w~~s~~~y~f~~~gl~ll~~aisa~~~~~~pw~~gva 952 (2087)
.++.|+ ++|-+.+|| .=-.+|.+++++-..+.+...++||.++|+|
T Consensus 172 i~f~pi-~v~l~~alG----------------yD~ivg~ai~~lg~~iGf~~s~~NPftvgIA 217 (506)
T PRK11588 172 IAFAII-IAPLMVRLG----------------YDSITTVLVTYVATQIGFATSWMNPFSVAIA 217 (506)
T ss_pred HHHHHH-HHHHHHHhC----------------CcHHHHHHHHHHHhhhhhcccccCccHHHHH
Confidence 366664 567666665 2247899999999899999999999998887
No 76
>PTZ00358 hypothetical protein; Provisional
Probab=20.36 E-value=67 Score=39.89 Aligned_cols=93 Identities=24% Similarity=0.384 Sum_probs=59.9
Q ss_pred cccccccchhHHHHHHhhHHHhhhcccchhhceeecccccchhhHHHHHHHHHh-hhcCceEEEEEeccCCCCCChhh--
Q 000135 811 FASPYASSVYLGWLMASAIALVVTGVLPIVSWFSTYRFSLSSAICVGIFAAVLV-AFCGASYLEVVKSREDQVPTKGD-- 887 (2087)
Q Consensus 811 ~~spy~~~~~~gw~~~s~i~lv~t~~~p~vswf~tyrf~~~sav~~~~f~~vl~-~fc~~sy~~v~~sr~~~~p~~~d-- 887 (2087)
..++..-+++---+++-+|.++..=.--..-|=..| .|.|++++|++ ++-+..+..+ +++++.--.++-
T Consensus 247 ~~~~~~Y~lLalHlIalgiTlycleyK~~fyWPKD~-------~CfGVLavV~ll~ll~v~v~~i-~~~~~~~~~~~~~Y 318 (367)
T PTZ00358 247 LLSTQGYPFLALHLVALGITLYCLEYKKVFYWPKDY-------MCFGVLAAVLLLVLVVVVVIII-DGFSQEAKNVGVKY 318 (367)
T ss_pred ccCCcceehHHHHHHHHHHHHhhheecccccccchh-------eeeHHHHHHHHHHHHHHHeeec-cccCccccccceEE
Confidence 444555577777777777777755444444554444 89999999887 8888887665 444442111111
Q ss_pred HHHhhhhhhhHHHHH-Hhhccccee
Q 000135 888 FLAALLPLVCIPALL-SLCSGLLKW 911 (2087)
Q Consensus 888 fl~allpl~~ipa~~-~l~~gl~kw 911 (2087)
-|..-.=+..+|+++ +-=|||++|
T Consensus 319 ~LlgYs~ilmvpTLwYaYRcGLltw 343 (367)
T PTZ00358 319 TLLGYSMLLMIPTLWYAYRCGLFTW 343 (367)
T ss_pred EEHHHHHHHHhHHHHHHHhhhhhee
Confidence 233444467789865 788999999
No 77
>KOG2751 consensus Beclin-like protein [Signal transduction mechanisms]
Probab=20.24 E-value=2e+02 Score=37.01 Aligned_cols=63 Identities=19% Similarity=0.292 Sum_probs=43.7
Q ss_pred hhccCchhhhhHHHHHHhhhhhhhhHHHHHHHHHhhhcccHHHHHHHHHHHHhhHHhhhhhhcccCCCCCc
Q 000135 1283 IKKWMPEDRRQFEIIQESYIREKEMEEEILMQRREEEGRGKERRKALLEKEERKWKEIEASLISSIPNAGN 1353 (2087)
Q Consensus 1283 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1353 (2087)
.++=.-|++++|..+.|.-++|++...++.. -++|+.++.|||.+.|+|--.+..+-+-.-++
T Consensus 185 ~~~l~~eE~~L~q~lk~le~~~~~l~~~l~e--------~~~~~~~~~e~~~~~~~ey~~~~~q~~~~~de 247 (447)
T KOG2751|consen 185 LKNLKEEEERLLQQLEELEKEEAELDHQLKE--------LEFKAERLNEEEDQYWREYNNFQRQLIEHQDE 247 (447)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH--------HHHHHHHHHHHHHHHHHHHHHHHHhhhcccch
Confidence 4455568888888887776666655544432 23466789999999999998887766654443
No 78
>PRK09776 putative diguanylate cyclase; Provisional
Probab=20.14 E-value=3.3e+02 Score=36.87 Aligned_cols=191 Identities=14% Similarity=0.121 Sum_probs=0.0
Q ss_pred HHHHHHhhHHHhhhcccchhhceeecccccchhhHHHHHHHHHhhhcCceEEEEEeccCCCCCChhhHHHhhhhhhhHHH
Q 000135 821 LGWLMASAIALVVTGVLPIVSWFSTYRFSLSSAICVGIFAAVLVAFCGASYLEVVKSREDQVPTKGDFLAALLPLVCIPA 900 (2087)
Q Consensus 821 ~gw~~~s~i~lv~t~~~p~vswf~tyrf~~~sav~~~~f~~vl~~fc~~sy~~v~~sr~~~~p~~~dfl~allpl~~ipa 900 (2087)
.|+++++.+|.++..++--.++. ++.+.+++-..-++++..-..++ .++.....+..|.+..++=-..+++
T Consensus 47 ~~~~~~~~~~~l~~~~~~~~~~~------~~~~~~~~~~~~~~~~~~ll~~~---~~~~~~~~~~~~~l~~~~~~~~~~~ 117 (1092)
T PRK09776 47 PGILLSCSLGNIAANILLFSTSS------LNLTWTTINLVEAVVGAVLLRKL---LPWYNPLQNLADWLRLALGSAIVPP 117 (1092)
T ss_pred HHHHHHHHHHHHhHhhhcCCcHH------HHHHHHHHHHHHHHHHHHHHHHh---cCccChhhCHHHHHHHHHHHHHHHH
Q ss_pred HHHhhcccceeecCc--------cccccceeeeehhHHHHHHhhhhheeeeechhHHHHHHHHHHHHHHHHHhhhhcccc
Q 000135 901 LLSLCSGLLKWKDDD--------WKLSRGVYVFITIGLVLLLGAISAVIVVITPWTIGVAFLLLLLLIVLAIGVIHHWAS 972 (2087)
Q Consensus 901 ~~~l~~gl~kw~dd~--------w~~s~~~y~f~~~gl~ll~~aisa~~~~~~pw~~gvafll~~~~~v~~igvih~was 972 (2087)
+++ +...+-+-... |-++--+.+++..-++|++ .-.-.-....|....-+.++++++++++..++++...
T Consensus 118 l~~-~~~~~~~~~~~~~~~~~~~w~~~~~~g~l~~~p~~l~~-~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 195 (1092)
T PRK09776 118 LLG-GVLVVLLTPGDDPLRAFLIWVLSEAIGMLALVPLGLLF-KPHYLLRHRNPRLLFESLLTLAITLTLSWLALLYLPW 195 (1092)
T ss_pred HHH-HHHHHHHcCCCchhhHHHHHHHHHHHHHHHHhhHhhhc-chHHHhhhcccchHHHHHHHHHHHHHHHHHHHHhCCC
Q ss_pred cceee-----------ehhhHHHHHHHHHHHHHHHHHhhhcCCCCccccchhHHHHHHHhh
Q 000135 973 NNFYL-----------TRTQMFFVCFLAFLLGLAAFLVGWFDDKPFVGASVGYFTFLFLLA 1022 (2087)
Q Consensus 973 nnfyl-----------~r~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1022 (2087)
--.++ ++....++++++.+.....+..|.+............+..+.++.
T Consensus 196 ~~~~~~~~~~~~a~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 256 (1092)
T PRK09776 196 PFTFIIVLLMWSAVRLPRMEAFLIFLTTVMMVSLMMAADPSLLATPRTYLMSHMPWLPFLL 256 (1092)
T ss_pred cHHHHHHHHHHHHHhccchHHHHHHHHHHHHHHHHHhCCccccCCcchhhhhhhhHHHHHH
Done!