Query 016344
Match_columns 391
No_of_seqs 140 out of 195
Neff 6.2
Searched_HMMs 46136
Date Fri Mar 29 06:00:36 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/016344.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/016344hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF09402 MSC: Man1-Src1p-C-ter 100.0 5E-62 1.1E-66 484.0 4.8 275 71-349 18-334 (334)
2 PF12946 EGF_MSP1_1: MSP1 EGF 95.5 0.0054 1.2E-07 42.1 0.9 28 94-121 5-37 (37)
3 PF01683 EB: EB module; Inter 74.3 5.3 0.00011 28.8 3.8 26 92-120 27-52 (52)
4 PF13314 DUF4083: Domain of un 64.0 31 0.00067 26.1 6.0 16 263-278 42-57 (58)
5 PF07127 Nodulin_late: Late no 63.2 12 0.00025 27.6 3.7 26 73-113 26-52 (54)
6 PTZ00382 Variant-specific surf 61.8 9.8 0.00021 31.6 3.4 34 90-123 19-56 (96)
7 PF06667 PspB: Phage shock pro 61.1 34 0.00074 27.2 6.2 30 243-272 14-51 (75)
8 PF07645 EGF_CA: Calcium-bindi 54.1 7.1 0.00015 27.0 1.1 21 95-115 11-35 (42)
9 COG2976 Uncharacterized protei 52.8 48 0.0011 31.3 6.7 50 227-277 14-65 (207)
10 PF06387 Calcyon: D1 dopamine 52.3 21 0.00046 32.9 4.1 16 109-124 113-128 (186)
11 KOG0196 Tyrosine kinase, EPH ( 50.0 13 0.00028 41.9 2.9 40 74-115 276-318 (996)
12 KOG1214 Nidogen and related ba 46.2 12 0.00027 42.0 2.0 34 91-124 828-867 (1289)
13 PF12947 EGF_3: EGF domain; I 44.6 12 0.00026 25.4 1.1 26 95-120 7-36 (36)
14 PF04891 NifQ: NifQ; InterPro 41.9 28 0.00062 31.8 3.4 57 39-104 107-167 (167)
15 PF06864 PAP_PilO: Pilin acces 41.7 50 0.0011 34.2 5.6 14 290-303 220-233 (414)
16 PF02009 Rifin_STEVOR: Rifin/s 41.2 41 0.00088 33.7 4.6 16 248-263 274-289 (299)
17 TIGR02976 phageshock_pspB phag 40.6 99 0.0021 24.6 5.8 28 245-272 16-51 (75)
18 PF01826 TIL: Trypsin Inhibito 40.4 15 0.00032 26.7 1.1 26 96-124 27-53 (55)
19 PF01102 Glycophorin_A: Glycop 40.2 39 0.00084 29.4 3.8 19 237-255 69-87 (122)
20 smart00179 EGF_CA Calcium-bind 39.8 27 0.00059 22.6 2.3 25 95-120 10-38 (39)
21 PF08563 P53_TAD: P53 transact 38.8 13 0.00029 23.4 0.5 14 176-189 8-21 (25)
22 PF10576 EndIII_4Fe-2S: Iron-s 37.7 15 0.00031 21.1 0.5 14 89-102 4-17 (17)
23 PF07974 EGF_2: EGF-like domai 36.3 33 0.00072 22.6 2.2 20 95-114 7-28 (32)
24 PRK09458 pspB phage shock prot 34.5 1E+02 0.0022 24.6 5.0 29 244-272 15-51 (75)
25 cd00053 EGF Epidermal growth f 34.0 40 0.00086 20.9 2.3 25 95-120 7-35 (36)
26 PRK11677 hypothetical protein; 33.9 1.6E+02 0.0034 26.1 6.6 6 244-249 14-19 (134)
27 PF00558 Vpu: Vpu protein; In 33.4 46 0.00099 27.0 2.9 19 256-274 30-48 (81)
28 PRK07597 secE preprotein trans 32.1 85 0.0018 23.7 4.1 28 38-65 25-52 (64)
29 PF06679 DUF1180: Protein of u 31.9 1.8E+02 0.004 26.5 6.9 31 35-65 84-114 (163)
30 TIGR00964 secE_bact preprotein 31.7 89 0.0019 23.0 4.1 26 39-64 17-42 (55)
31 PF07271 Cytadhesin_P30: Cytad 30.5 1.6E+02 0.0034 29.2 6.6 17 260-276 104-120 (279)
32 PF06143 Baculo_11_kDa: Baculo 29.9 3.1E+02 0.0067 22.4 7.9 22 230-251 32-53 (84)
33 PHA03399 pif3 per os infectivi 28.9 66 0.0014 30.4 3.6 32 71-114 47-86 (200)
34 PF07543 PGA2: Protein traffic 27.9 1.7E+02 0.0037 26.0 5.9 12 293-304 62-73 (140)
35 KOG0474 Cl- channel CLC-7 and 27.9 86 0.0019 34.6 4.7 24 90-113 396-420 (762)
36 COG0690 SecE Preprotein transl 27.8 1.2E+02 0.0027 23.8 4.5 28 38-65 35-62 (73)
37 PF05568 ASFV_J13L: African sw 27.6 1.5E+02 0.0033 26.7 5.5 10 230-239 26-35 (189)
38 PHA02817 EEV Host range protei 26.7 64 0.0014 31.0 3.2 51 72-125 66-137 (225)
39 KOG4403 Cell surface glycoprot 26.5 3.3E+02 0.0071 28.9 8.3 13 156-168 116-128 (575)
40 PF06247 Plasmod_Pvs28: Plasmo 26.1 25 0.00054 32.9 0.3 32 95-126 51-91 (197)
41 PF07466 DUF1517: Protein of u 25.7 1.9E+02 0.0042 28.7 6.5 23 43-65 62-84 (289)
42 PF14316 DUF4381: Domain of un 25.5 1.7E+02 0.0037 25.6 5.6 15 269-283 70-84 (146)
43 PF03672 UPF0154: Uncharacteri 25.3 2.4E+02 0.0052 21.9 5.5 18 243-260 10-27 (64)
44 PF00584 SecE: SecE/Sec61-gamm 25.2 1.7E+02 0.0036 21.4 4.6 21 39-59 18-38 (57)
45 PF14991 MLANA: Protein melan- 24.7 18 0.00039 31.1 -0.8 22 241-262 31-54 (118)
46 PF09402 MSC: Man1-Src1p-C-ter 23.7 26 0.00057 34.8 0.0 70 262-331 98-174 (334)
47 PF07699 GCC2_GCC3: GCC2 and G 23.6 65 0.0014 22.8 2.0 28 73-102 10-37 (48)
48 PF10588 NADH-G_4Fe-4S_3: NADH 23.2 35 0.00076 23.7 0.5 16 88-103 11-26 (41)
49 PF08114 PMP1_2: ATPase proteo 23.1 1.4E+02 0.0031 21.1 3.5 17 246-262 20-36 (43)
50 smart00032 CCP Domain abundant 22.6 51 0.0011 22.7 1.3 20 106-125 27-49 (57)
51 PF10500 SR-25: Nuclear RNA-sp 22.5 60 0.0013 31.1 2.1 9 177-185 159-167 (225)
52 PF09064 Tme5_EGF_like: Thromb 22.4 68 0.0015 21.8 1.7 17 101-117 11-30 (34)
53 PF12729 4HB_MCP_1: Four helix 21.7 3.6E+02 0.0079 22.6 6.9 8 298-305 64-71 (181)
54 cd00033 CCP Complement control 21.6 50 0.0011 22.9 1.1 20 106-125 26-48 (57)
55 PF01102 Glycophorin_A: Glycop 21.3 1.1E+02 0.0023 26.7 3.3 22 237-258 73-94 (122)
56 PRK15428 putative propanediol 21.2 82 0.0018 28.8 2.7 31 263-301 4-34 (163)
57 PF11392 DUF2877: Protein of u 21.2 52 0.0011 27.8 1.3 11 35-45 5-15 (110)
No 1
>PF09402 MSC: Man1-Src1p-C-terminal domain; InterPro: IPR018996 This entry represents the Inner nuclear membrane proteins MAN1 (also known as LEM domain-containing protein 3) and LEM domain-containing protein 2 (or LEM protein 2). Emerin and MAN1 are LEM domain-containing integral membrane proteins of the vertebrate nuclear envelope []. MAN1 is an integral protein of the inner nuclear membrane which binds to chromatin associated proteins and plays a role in nuclear organisation. The C-terminal nulceoplasmic region forms a DNA binding winged helix and binds to Smad []. LEM protein 2 is an essential protein involved in chromosome segregation and cell division, probably via its interaction with lmn-1, the main component of nuclear lamina. Has some overlapping function with emr-1.; GO: 0005639 integral to nuclear inner membrane; PDB: 2CH0_A.
Probab=100.00 E-value=5e-62 Score=484.04 Aligned_cols=275 Identities=28% Similarity=0.421 Sum_probs=71.2
Q ss_pred CCCCCCCCCCCCCCC------------CCCCCCCCccCCCCceecCC-eeeeCCCceec-----------CCCcccChhh
Q 016344 71 STSKPFCDSNLLLDS------------PQSPTDSCEPCPSNGECHQG-KLECFHGYRKH-----------GKLCVEDGDI 126 (391)
Q Consensus 71 ~~~~~fCdS~~~~~~------------~~~~~p~C~PCP~hAiC~~g-~l~C~~gYvl~-----------~~~CV~D~~k 126 (391)
+...+|||++.+..+ ...++|+|+|||+||+|++| ++.|++||++. +++|++|+++
T Consensus 18 ~~~vgyC~~~~~~~~~~~~~~~~~~~~~~~~~P~C~pCP~~a~C~~~~~~~C~~~y~~~~~~l~~~g~~p~~~Ci~D~~k 97 (334)
T PF09402_consen 18 KIAVGYCGTESPSPSFADDDISVPDWLLENFKPSCEPCPEHAICYPGLKLECEPGYVLKPSPLSLFGLIPPPKCIPDTEK 97 (334)
T ss_dssp --------------------------------------------------------------------------------
T ss_pred cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccHHH
Confidence 468999999972111 14678999999999999999 99999999998 9999999999
Q ss_pred hHHHHHHHHHHHHHHHHHhhcccccC---CCCcccchhhHHHHhhhhhhhhccCCChHHHHHHHHHHHHHHHhhcccccc
Q 016344 127 NETAGRLSRWVENRLCRAYAQFLCDG---TGSIWVEENDIWNDLEGHELMKIFELDNPVYLYTKKRTMETVGRYLESRTN 203 (391)
Q Consensus 127 ~~~i~~l~~~i~~~Lr~r~A~~~CG~---~~s~~i~e~dL~~~~~~~~~~k~~~ls~~~fe~l~~~al~~l~~~l~~~~~ 203 (391)
++.+.+|++++.++||++||+++||. ..+.+|+++||++++.+ ++++++++++|+++|..|+..+.+.-+..+.
T Consensus 98 ~~~i~~l~~~~~~~Lr~~~a~~~Cg~~~~~~~~~ls~~el~~~~~~---~~~~~~~~~efe~l~~~a~~~L~~~~ei~~~ 174 (334)
T PF09402_consen 98 EEKIEELAKKILDELRERNAQYECGDSEDDESPGLSEEELKDILSS---KKSPWISDEEFEELWSAALQELKKNPEIIIR 174 (334)
T ss_dssp --------------------------------------------------------------------------------
T ss_pred HHHHHHHHHHHHHHHHHHHhhcccCCCCCCCCCCCcHHHHHHHHHh---ccCccccHHHHHHHHHHHHHHHHhCCcEEEe
Confidence 99999999999999999999999993 34678999999999998 7789999999999999999888744332222
Q ss_pred ---------cCCceeeccccccccCccCccchhHHH----HHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Q 016344 204 ---------SYGMKELKCPELLAEHYKPLSCRIHQW----VSTHALIIVPVCSLLVGCLLLLWKVHRRRYFAIRVEELYH 270 (391)
Q Consensus 204 ---------sn~~~~~k~~~s~S~~~i~l~Cr~r~~----i~~~~~~i~~~l~~~v~i~~l~~~~~r~~~e~~~v~~Lv~ 270 (391)
.+.........+++++++||+|++++. +.+++..++++++++++++++++++++++.++++|++||+
T Consensus 175 ~~~~~~~~~~~~~~~~~~~~s~s~~~lpl~C~~~~~i~~~~~~~~~~i~~~~~~~~~~~~~~~~~~~~~~~~~~v~~lv~ 254 (334)
T PF09402_consen 175 DDIINSHSSDDSNEKDKYFRSSSLPYLPLKCRLRRQIRQFISRYRLIILGVLILLLLIKYIRYRYRKRREEKARVEELVK 254 (334)
T ss_dssp -----------------------------------------------------------------STHHHHHTTTTTTHH
T ss_pred cccccccccccccCCcEEEEeeCCCccccEEEEehHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 001111122234689999999976555 4556666777766777777777778888899999999999
Q ss_pred HHHHHHHHHHhhhhccCCCCCCcccccccccccCCCCC--cCchhhHHHHHHHHhcCCCcceeeeEEcCceeeeeEEeec
Q 016344 271 QVCEILEENALMSKSVNGECEPWVVASRLRDHLLLPKE--RKDPVIWKKVEELVQEDSRVDQYPKLLKGESKVVWEWQVE 348 (391)
Q Consensus 271 ~vl~~L~~q~~~~~~~~~~~~pyl~~~qLRD~LL~~~~--r~r~~LW~kV~k~Ve~nSnVrt~~~ev~GE~~~vWEWig~ 348 (391)
+|+++|++|+..+ ..+..++|||+++||||+||.+.+ +++++||++|+++||+|||||++++|+|||+|+||||||+
T Consensus 255 ~ii~~L~~~~~~~-~~~~~~~p~v~~~qLRD~ll~~~~~~~~~~~lW~~v~~~ve~ns~Vr~~~~e~~Ge~~~vWeWig~ 333 (334)
T PF09402_consen 255 KIIDRLQDQARAS-DPNSSPEPYVSISQLRDDLLPPEHRLKRRNRLWKKVVKKVEENSNVRTEVREVHGEIMRVWEWIGP 333 (334)
T ss_dssp HHHHHHHHHHHHH-TTSS-S-S-B-HHHHHHTT--STTGGG-GHHHHHHHHHHHTT---SEEEEEEETTEEEEEEE----
T ss_pred HHHHHHHHHhhhh-ccCCCCCCCccHHHHHHHhCCcccCHHHHHHHHHHHHHHHHcCCCeeEEEEEECCeEEEEEEecCC
Confidence 9999999998843 344678999999999999998765 3379999999999999999999999999999999999997
Q ss_pred C
Q 016344 349 G 349 (391)
Q Consensus 349 ~ 349 (391)
+
T Consensus 334 ~ 334 (334)
T PF09402_consen 334 N 334 (334)
T ss_dssp -
T ss_pred C
Confidence 5
No 2
>PF12946 EGF_MSP1_1: MSP1 EGF domain 1; InterPro: IPR024730 This EGF-like domain is found at the C terminus of the malaria parasite MSP1 protein. MSP1 is the merozoite surface protein 1. This domain is part of the C-terminal fragment that is proteolytically processed from the the rest of the protein and is left attached to the surface of the invading parasite [].; PDB: 1N1I_C 2FLG_A 1CEJ_A 2NPR_A 1B9W_A 1OB1_F.
Probab=95.51 E-value=0.0054 Score=42.08 Aligned_cols=28 Identities=39% Similarity=0.954 Sum_probs=20.2
Q ss_pred ccCCCCceecCC----e-eeeCCCceecCCCcc
Q 016344 94 EPCPSNGECHQG----K-LECFHGYRKHGKLCV 121 (391)
Q Consensus 94 ~PCP~hAiC~~g----~-l~C~~gYvl~~~~CV 121 (391)
++||+||-|+.+ + -.|..||.+.+.+|+
T Consensus 5 ~~cP~NA~C~~~~dG~eecrCllgyk~~~~~C~ 37 (37)
T PF12946_consen 5 TKCPANAGCFRYDDGSEECRCLLGYKKVGGKCV 37 (37)
T ss_dssp S---TTEEEEEETTSEEEEEE-TTEEEETTEEE
T ss_pred ccCCCCcccEEcCCCCEEEEeeCCccccCCCcC
Confidence 589999999864 3 399999999999886
No 3
>PF01683 EB: EB module; InterPro: IPR006149 The EB domain has no known function. It is found in several Caenorhabditis sp. and Drosophila sp. proteins. The domain contains 8 conserved cysteines that probably form four disulphide bridges and is found associated with kunitz domains IPR002223 from INTERPRO
Probab=74.28 E-value=5.3 Score=28.79 Aligned_cols=26 Identities=31% Similarity=0.782 Sum_probs=23.0
Q ss_pred CCccCCCCceecCCeeeeCCCceecCCCc
Q 016344 92 SCEPCPSNGECHQGKLECFHGYRKHGKLC 120 (391)
Q Consensus 92 ~C~PCP~hAiC~~g~l~C~~gYvl~~~~C 120 (391)
+|+ .++.|.+|.=.|.+||+..+.+|
T Consensus 27 qC~---~~s~C~~g~C~C~~g~~~~~~~C 52 (52)
T PF01683_consen 27 QCI---GGSVCVNGRCQCPPGYVEVGGRC 52 (52)
T ss_pred CCC---CcCEEcCCEeECCCCCEecCCCC
Confidence 555 99999998889999999988877
No 4
>PF13314 DUF4083: Domain of unknown function (DUF4083)
Probab=64.01 E-value=31 Score=26.12 Aligned_cols=16 Identities=25% Similarity=0.416 Sum_probs=10.3
Q ss_pred HHHHHHHHHHHHHHHH
Q 016344 263 IRVEELYHQVCEILEE 278 (391)
Q Consensus 263 ~~v~~Lv~~vl~~L~~ 278 (391)
..+++=.+.+++.|++
T Consensus 42 ~~~eqKLDrIIeLLEK 57 (58)
T PF13314_consen 42 DSMEQKLDRIIELLEK 57 (58)
T ss_pred hHHHHHHHHHHHHHcc
Confidence 3566666677777754
No 5
>PF07127 Nodulin_late: Late nodulin protein; InterPro: IPR009810 This family consists of several plant specific late nodulin sequences which are homologous to the Pisum sativum (Garden pea) ENOD3 protein. ENOD3 is expressed in the late stages of root nodule formation and contains two pairs of cysteine residues toward the proteins C terminus which may be involved in metal-binding [].; GO: 0046872 metal ion binding, 0009878 nodule morphogenesis
Probab=63.16 E-value=12 Score=27.62 Aligned_cols=26 Identities=19% Similarity=0.586 Sum_probs=19.2
Q ss_pred CCCCCCCCCCCCCCCCCCCCCccCCCCceecCC-eeeeCCCc
Q 016344 73 SKPFCDSNLLLDSPQSPTDSCEPCPSNGECHQG-KLECFHGY 113 (391)
Q Consensus 73 ~~~fCdS~~~~~~~~~~~p~C~PCP~hAiC~~g-~l~C~~gY 113 (391)
....|.++. .||.+ |..+ ..+|..|+
T Consensus 26 ~~~~C~~d~-------------DCp~~--c~~~~~~kCi~~~ 52 (54)
T PF07127_consen 26 AIIPCKTDS-------------DCPKD--CPPPFIPKCINNI 52 (54)
T ss_pred CCcccCccc-------------cCCCC--CCCCcCcEeCcCC
Confidence 467898885 78888 8877 45887663
No 6
>PTZ00382 Variant-specific surface protein (VSP); Provisional
Probab=61.77 E-value=9.8 Score=31.56 Aligned_cols=34 Identities=24% Similarity=0.574 Sum_probs=23.6
Q ss_pred CCCCccCCC--CceecCCee--eeCCCceecCCCcccC
Q 016344 90 TDSCEPCPS--NGECHQGKL--ECFHGYRKHGKLCVED 123 (391)
Q Consensus 90 ~p~C~PCP~--hAiC~~g~l--~C~~gYvl~~~~CV~D 123 (391)
...|.+||. =+.|..... .|..||.+.++.|+.+
T Consensus 19 ~~~C~~C~~~~C~~C~~~~~C~~C~~GY~~~~~~Cv~~ 56 (96)
T PTZ00382 19 GSGCVLCSVGNCKSCVVDGVCGECNSGFSLDNGKCVSS 56 (96)
T ss_pred CCcCCcCCCCCCcCCCCCCccccCcCCcccCCCccccc
Confidence 346999985 234433322 8999999999989864
No 7
>PF06667 PspB: Phage shock protein B; InterPro: IPR009554 This family consists of several bacterial phage shock protein B (PspB) sequences. The phage shock protein (psp) operon is induced in response to heat, ethanol, osmotic shock and infection by filamentous bacteriophages []. Expression of the operon requires the alternative sigma factor sigma54 and the transcriptional activator PspF. In addition, PspA plays a negative regulatory role, and the integral-membrane proteins PspB and PspC play a positive one [].; GO: 0006355 regulation of transcription, DNA-dependent, 0009271 phage shock
Probab=61.11 E-value=34 Score=27.23 Aligned_cols=30 Identities=27% Similarity=0.343 Sum_probs=17.5
Q ss_pred HHHHHHHHH-HHHHHHHH-------HHHHHHHHHHHHH
Q 016344 243 SLLVGCLLL-LWKVHRRR-------YFAIRVEELYHQV 272 (391)
Q Consensus 243 ~~~v~i~~l-~~~~~r~~-------~e~~~v~~Lv~~v 272 (391)
+++|+..|+ .+|..+++ .+.++.++|++.+
T Consensus 14 ~ifVap~WL~lHY~sk~~~~~gLs~~d~~~L~~L~~~a 51 (75)
T PF06667_consen 14 MIFVAPIWLILHYRSKWKSSQGLSEEDEQRLQELYEQA 51 (75)
T ss_pred HHHHHHHHHHHHHHHhcccCCCCCHHHHHHHHHHHHHH
Confidence 445555554 56665554 4566677777755
No 8
>PF07645 EGF_CA: Calcium-binding EGF domain; InterPro: IPR001881 A sequence of about forty amino-acid residues found in epidermal growth factor (EGF) has been shown [, , , , , ] to be present in a large number of membrane-bound and extracellular, mostly animal, proteins. Many of these proteins require calcium for their biological function and a calcium-binding site has been found at the N terminus of some EGF-like domains []. Calcium-binding may be crucial for numerous protein-protein interactions. For human coagulation factor IX it has been shown [] that the calcium-ligands form a pentagonal bipyramid. The first, third and fourth conserved negatively charged or polar residues are side chain ligands. The latter is possibly hydroxylated (see aspartic acid and asparagine hydroxylation site) []. A conserved aromatic residue, as well as the second conserved negative residue, are thought to be involved in stabilising the calcium-binding site. As in non-calcium binding EGF-like domains, there are six conserved cysteines and the structure of both types is very similar as calcium-binding induces only strictly local structural changes []. +------------------+ +---------+ | | | | nxnnC-x(3,14)-C-x(3,7)-CxxbxxxxaxC-x(1,6)-C-x(8,13)-Cx | | +------------------+ 'n': negatively charged or polar residue [DEQN] 'b': possibly beta-hydroxylated residue [DN] 'a': aromatic amino acid 'C': cysteine, involved in disulphide bond 'x': any amino acid. ; GO: 0005509 calcium ion binding; PDB: 2VJ3_A 1TOZ_A 1LMJ_A 1UZQ_A 1UZK_A 1UZJ_B 1UZP_A 1EMO_A 1EMN_A 2RR0_A ....
Probab=54.06 E-value=7.1 Score=26.98 Aligned_cols=21 Identities=38% Similarity=0.911 Sum_probs=17.6
Q ss_pred cCCCCceecCC--ee--eeCCCcee
Q 016344 95 PCPSNGECHQG--KL--ECFHGYRK 115 (391)
Q Consensus 95 PCP~hAiC~~g--~l--~C~~gYvl 115 (391)
+|+.++.|.+- .. .|.+||..
T Consensus 11 ~C~~~~~C~N~~Gsy~C~C~~Gy~~ 35 (42)
T PF07645_consen 11 NCPENGTCVNTEGSYSCSCPPGYEL 35 (42)
T ss_dssp SSSTTSEEEEETTEEEEEESTTEEE
T ss_pred cCCCCCEEEcCCCCEEeeCCCCcEE
Confidence 68999999876 33 99999994
No 9
>COG2976 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=52.82 E-value=48 Score=31.34 Aligned_cols=50 Identities=14% Similarity=0.295 Sum_probs=24.5
Q ss_pred hHHHHHHHHHHHHHHHHHHHHHHHHHHHHH-HHHHHH-HHHHHHHHHHHHHHH
Q 016344 227 IHQWVSTHALIIVPVCSLLVGCLLLLWKVH-RRRYFA-IRVEELYHQVCEILE 277 (391)
Q Consensus 227 ~r~~i~~~~~~i~~~l~~~v~i~~l~~~~~-r~~~e~-~~v~~Lv~~vl~~L~ 277 (391)
+++|++.+...++..+++.+|. ++-|.++ .++.++ +.....|+.+++.++
T Consensus 14 ik~wwkeNGk~li~gviLg~~~-lfGW~ywq~~q~~q~~~AS~~Y~~~i~~~~ 65 (207)
T COG2976 14 IKDWWKENGKALIVGVILGLGG-LFGWRYWQSHQVEQAQEASAQYQNAIKAVQ 65 (207)
T ss_pred HHHHHHHCCchhHHHHHHHHHH-HHHHHHHHHHHHHHHHHHHHHHHHHHHHHh
Confidence 5667766654332222222222 2345554 344332 345667777777763
No 10
>PF06387 Calcyon: D1 dopamine receptor-interacting protein (calcyon); InterPro: IPR009431 This family consists of several D1 dopamine receptor-interacting (calcyon) proteins. D1/D5 dopamine receptors in the basal ganglia, hippocampus, and cerebral cortex modulate motor, reward, and cognitive behaviour. D1-like dopamine receptors likely modulate neocortical and hippocampal neuronal excitability and synaptic function via Ca2+ as well as cAMP-dependent signalling []. Defective calcyon proteins have been implicated in both attention-deficit/hyperactivity disorder (ADHD) [] and schizophrenia.; GO: 0050780 dopamine receptor binding, 0007212 dopamine receptor signaling pathway, 0016021 integral to membrane
Probab=52.28 E-value=21 Score=32.89 Aligned_cols=16 Identities=25% Similarity=0.455 Sum_probs=11.5
Q ss_pred eCCCceecCCCcccCh
Q 016344 109 CFHGYRKHGKLCVEDG 124 (391)
Q Consensus 109 C~~gYvl~~~~CV~D~ 124 (391)
|-+||++..+.|.|-+
T Consensus 113 CPdGFv~khk~C~P~~ 128 (186)
T PF06387_consen 113 CPDGFVLKHKRCTPLT 128 (186)
T ss_pred CCCcceeecccccchh
Confidence 4458888888888744
No 11
>KOG0196 consensus Tyrosine kinase, EPH (ephrin) receptor family [Signal transduction mechanisms]
Probab=50.02 E-value=13 Score=41.86 Aligned_cols=40 Identities=25% Similarity=0.610 Sum_probs=28.1
Q ss_pred CCCCCCCCCCCCCCCCCCCCccCCCCcee-cCC-ee-eeCCCcee
Q 016344 74 KPFCDSNLLLDSPQSPTDSCEPCPSNGEC-HQG-KL-ECFHGYRK 115 (391)
Q Consensus 74 ~~fCdS~~~~~~~~~~~p~C~PCP~hAiC-~~g-~l-~C~~gYvl 115 (391)
..-|..+. + +.......|.|||+|.+= ..| .. .|..||-.
T Consensus 276 C~aCp~G~-y-K~~~~~~~C~~CP~~S~s~~ega~~C~C~~gyyR 318 (996)
T KOG0196|consen 276 CQACPPGT-Y-KASQGDSLCLPCPPNSHSSSEGATSCTCENGYYR 318 (996)
T ss_pred ceeCCCCc-c-cCCCCCCCCCCCCCCCCCCCCCCCcccccCCccc
Confidence 33455553 1 223456789999999998 556 55 99999998
No 12
>KOG1214 consensus Nidogen and related basement membrane protein proteins [Cell wall/membrane/envelope biogenesis; Extracellular structures]
Probab=46.24 E-value=12 Score=42.00 Aligned_cols=34 Identities=35% Similarity=0.842 Sum_probs=29.4
Q ss_pred CCCcc--CCCCceecCC--ee--eeCCCceecCCCcccCh
Q 016344 91 DSCEP--CPSNGECHQG--KL--ECFHGYRKHGKLCVEDG 124 (391)
Q Consensus 91 p~C~P--CP~hAiC~~g--~l--~C~~gYvl~~~~CV~D~ 124 (391)
++|.| |=++|.||+. .+ +|.+||.--+-.||||+
T Consensus 828 DeC~psrChp~A~CyntpgsfsC~C~pGy~GDGf~CVP~~ 867 (1289)
T KOG1214|consen 828 DECSPSRCHPAATCYNTPGSFSCRCQPGYYGDGFQCVPDT 867 (1289)
T ss_pred cccCccccCCCceEecCCCcceeecccCccCCCceecCCC
Confidence 67766 9999999987 33 99999999999999993
No 13
>PF12947 EGF_3: EGF domain; InterPro: IPR024731 This entry represents an EGF domain found in the the C terminus of malarial parasite merozoite surface protein 1 [], as well as other proteins.; PDB: 2NPR_A 1N1I_C 1B9W_A 1YO8_A 2RHP_A.
Probab=44.61 E-value=12 Score=25.37 Aligned_cols=26 Identities=31% Similarity=0.797 Sum_probs=17.7
Q ss_pred cCCCCceecCC--ee--eeCCCceecCCCc
Q 016344 95 PCPSNGECHQG--KL--ECFHGYRKHGKLC 120 (391)
Q Consensus 95 PCP~hAiC~~g--~l--~C~~gYvl~~~~C 120 (391)
.|=+||.|.+- .+ .|.+||.--+..|
T Consensus 7 ~C~~nA~C~~~~~~~~C~C~~Gy~GdG~~C 36 (36)
T PF12947_consen 7 GCHPNATCTNTGGSYTCTCKPGYEGDGFFC 36 (36)
T ss_dssp GS-TTCEEEE-TTSEEEEE-CEEECCSTCE
T ss_pred CCCCCcEeecCCCCEEeECCCCCccCCcCC
Confidence 67889999876 44 9999998665554
No 14
>PF04891 NifQ: NifQ; InterPro: IPR006975 NifQ is involved in early stages of the biosynthesis of the iron-molybdenum cofactor (FeMo-co) [], which is an integral part of the active site of dinitrogenase []. The conserved C-terminal cysteine residues may be involved in metal binding [].; GO: 0030151 molybdenum ion binding, 0009399 nitrogen fixation
Probab=41.93 E-value=28 Score=31.84 Aligned_cols=57 Identities=19% Similarity=0.306 Sum_probs=30.7
Q ss_pred CCChhhHHHHHHHHHHHHHHHHHHHHHHh----hhcCCCCCCCCCCCCCCCCCCCCCCCccCCCCceecC
Q 016344 39 FPSKQDLLRLITVVAIASSVALTCNYLAN----FLNSTSKPFCDSNLLLDSPQSPTDSCEPCPSNGECHQ 104 (391)
Q Consensus 39 ~~~~~~~~~~~~~~~ia~~~~~~c~~l~~----~~~~~~~~fCdS~~~~~~~~~~~p~C~PCP~hAiC~~ 104 (391)
+++++|+.+|+.=-+=+.+.. |.=.+ |||+ .-|.... -+.=..|+|.-|.+++.||.
T Consensus 107 L~~R~eLs~Lm~r~Fp~Laa~---N~~~MrWKKFfYr---qlCe~eG---~~~C~aPsC~~C~D~~~CFG 167 (167)
T PF04891_consen 107 LRSRAELSALMRRHFPPLAAR---NTRNMRWKKFFYR---QLCEREG---LYLCRAPSCEECSDYAVCFG 167 (167)
T ss_pred CCCHHHHHHHHHHHhHHHHHh---ccCCCcHHHHHHH---HHHHHcC---CCcCCCCCCCCcCCHhhcCC
Confidence 466777665554443333332 22111 5662 1243332 01123589999999999984
No 15
>PF06864 PAP_PilO: Pilin accessory protein (PilO); InterPro: IPR009663 This family consists of several enterobacterial PilO proteins. The function of PilO is unknown although it has been suggested that it is a cytoplasmic protein in the absence of other Pil proteins, but PilO protein is translocated to the outer membrane in the presence of other Pil proteins. Alternatively, PilO protein may form a complex with other Pil protein(s). PilO has been predicted to function as a component of the pilin transport apparatus and thin-pilus basal body []. This family does not seem to be related to IPR007445 from INTERPRO.
Probab=41.73 E-value=50 Score=34.19 Aligned_cols=14 Identities=21% Similarity=0.581 Sum_probs=9.5
Q ss_pred CCCccccccccccc
Q 016344 290 CEPWVVASRLRDHL 303 (391)
Q Consensus 290 ~~pyl~~~qLRD~L 303 (391)
++||...+..-+.|
T Consensus 220 ~~PW~~~P~~~~fl 233 (414)
T PF06864_consen 220 PHPWAKQPSVQAFL 233 (414)
T ss_pred CCCcccCCCHHHHH
Confidence 56887777666654
No 16
>PF02009 Rifin_STEVOR: Rifin/stevor family; InterPro: IPR002858 Malaria is still a major cause of mortality in many areas of the world. Plasmodium falciparum causes the most severe human form of the disease and is responsible for most fatalities. Severe cases of malaria can occur when the parasite invades and then proliferates within red blood cell erythrocytes. The parasite produces many variant antigenic proteins, encoded by multigene families, which are present on the surface of the infected erythrocyte and play important roles in virulence. A crucial survival mechanism for the malaria parasite is its ability to evade the immune response by switching these variant surface antigens. The high virulence of P. falciparum relative to other malarial parasites is in large part due to the fact that in this organism many of these surface antigens mediate the binding of infected erythrocytes to the vascular endothelium (cytoadherence) and non-infected erythrocytes (rosetting). This can lead to the accumulation of infected cells in the vasculature of a variety of organs, blocking the blood flow and reducing the oxygen supply. Clinical symptoms of severe infection can include fever, progressive anaemia, multi-organ dysfunction and coma. For more information see []. Several multicopy gene families have been described in Plasmodium falciparum, including the stevor family of subtelomeric open reading frames and the rif interspersed repetitive elements. Both families contain three predicted transmembrane segments. It has been proposed that stevor and rif are members of a larger superfamily that code for variant surface antigens [].
Probab=41.18 E-value=41 Score=33.67 Aligned_cols=16 Identities=13% Similarity=0.358 Sum_probs=8.4
Q ss_pred HHHHHHHHHHHHHHHH
Q 016344 248 CLLLLWKVHRRRYFAI 263 (391)
Q Consensus 248 i~~l~~~~~r~~~e~~ 263 (391)
|.||+++|+|+++.+.
T Consensus 274 IIYLILRYRRKKKmkK 289 (299)
T PF02009_consen 274 IIYLILRYRRKKKMKK 289 (299)
T ss_pred HHHHHHHHHHHhhhhH
Confidence 4445566666555443
No 17
>TIGR02976 phageshock_pspB phage shock protein B. This model describes the PspB protein of the psp (phage shock protein) operon, as found in Escherichia coli and many related species. Expression of a phage protein called secretin protein IV, and a number of other stresses including ethanol, heat shock, and defects in protein secretion trigger sigma-54-dependent expression of the phage shock regulon. PspB is both a regulator and an effector protein of the phage shock response.
Probab=40.60 E-value=99 Score=24.59 Aligned_cols=28 Identities=29% Similarity=0.338 Sum_probs=14.6
Q ss_pred HHHHHHH-HHHHHHHH-------HHHHHHHHHHHHH
Q 016344 245 LVGCLLL-LWKVHRRR-------YFAIRVEELYHQV 272 (391)
Q Consensus 245 ~v~i~~l-~~~~~r~~-------~e~~~v~~Lv~~v 272 (391)
+++..++ .+|..+++ .+.++..+|++.+
T Consensus 16 fVap~wl~lHY~~k~~~~~~ls~~d~~~L~~L~~~a 51 (75)
T TIGR02976 16 FVAPLWLILHYRSKRKTAASLSTDDQALLQELYAKA 51 (75)
T ss_pred HHHHHHHHHHHHhhhccCCCCCHHHHHHHHHHHHHH
Confidence 3444444 55554443 3555666666654
No 18
>PF01826 TIL: Trypsin Inhibitor like cysteine rich domain; InterPro: IPR002919 This domain is found in proteinase inhibitors as well as in many extracellular proteins. The domain typically contains ten cysteine residues that form five disulphide bonds. The cysteine residues that form the disulphide bonds are 1-7, 2-6, 3-5, 4-10 and 8-9. This inhibitor domain belongs to MEROPS inhibitor family I8 (clan IA). Proteins containing this domain inhibit peptidases belonging to families S1 (IPR001254 from INTERPRO), S8 (IPR000209 from INTERPRO), and M4 (IPR001570 from INTERPRO) [] and are restricted to the chordata, nematoda, arthropoda and echinodermata. Examples of proteins containing this domain are: chymotrypsin/elastase inhibitor from Ascaris suum (pig roundworm) Acp62F protein from Drosophila melanogaster Bombina trypsin inhibitor from Bombina maxima (large-webbed bell toad) Bombyx subtilisin inhibitor from Bombyx mori (silk moth) von Willebrand factor ; PDB: 2P3F_N 1HX2_A 1CCV_A 1EAI_D 2H9E_C 1COU_A 1ATE_A 1ATB_A 1ATD_A 1ATA_A ....
Probab=40.35 E-value=15 Score=26.67 Aligned_cols=26 Identities=31% Similarity=0.752 Sum_probs=20.7
Q ss_pred CCCCceecCCeeeeCCCceecCC-CcccCh
Q 016344 96 CPSNGECHQGKLECFHGYRKHGK-LCVEDG 124 (391)
Q Consensus 96 CP~hAiC~~g~l~C~~gYvl~~~-~CV~D~ 124 (391)
|+ ..|.+| =.|.+||++... .||+-.
T Consensus 27 C~--~~C~~g-C~C~~G~v~~~~~~CV~~~ 53 (55)
T PF01826_consen 27 CS--EPCVEG-CFCPPGYVRNDNGRCVPPS 53 (55)
T ss_dssp CS--SS-ESE-EEETTTEEEETTSEEEEGG
T ss_pred cC--CCCCcc-CCCCCCeeEcCCCCEEcHH
Confidence 55 778888 789999999876 999864
No 19
>PF01102 Glycophorin_A: Glycophorin A; InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=40.15 E-value=39 Score=29.42 Aligned_cols=19 Identities=32% Similarity=0.427 Sum_probs=8.6
Q ss_pred HHHHHHHHHHHHHHHHHHH
Q 016344 237 IIVPVCSLLVGCLLLLWKV 255 (391)
Q Consensus 237 ~i~~~l~~~v~i~~l~~~~ 255 (391)
+|+++++.++|+.+++.|.
T Consensus 69 Ii~gv~aGvIg~Illi~y~ 87 (122)
T PF01102_consen 69 IIFGVMAGVIGIILLISYC 87 (122)
T ss_dssp HHHHHHHHHHHHHHHHHHH
T ss_pred hhHHHHHHHHHHHHHHHHH
Confidence 4445544444444444443
No 20
>smart00179 EGF_CA Calcium-binding EGF-like domain.
Probab=39.81 E-value=27 Score=22.56 Aligned_cols=25 Identities=40% Similarity=1.046 Sum_probs=18.4
Q ss_pred cCCCCceecCC--ee--eeCCCceecCCCc
Q 016344 95 PCPSNGECHQG--KL--ECFHGYRKHGKLC 120 (391)
Q Consensus 95 PCP~hAiC~~g--~l--~C~~gYvl~~~~C 120 (391)
||..+|.|.+. .. .|.+||. .+..|
T Consensus 10 ~C~~~~~C~~~~g~~~C~C~~g~~-~g~~C 38 (39)
T smart00179 10 PCQNGGTCVNTVGSYRCECPPGYT-DGRNC 38 (39)
T ss_pred CcCCCCEeECCCCCeEeECCCCCc-cCCcC
Confidence 69899999854 22 7889987 55555
No 21
>PF08563 P53_TAD: P53 transactivation motif; InterPro: IPR013872 The binding of this protein by regulatory proteins regulates p53 transcription activation. This entry is comprised of a single amphipathic alpha helix and contains a highly conserved motif [, ]. ; GO: 0005515 protein binding; PDB: 1YCQ_B 2Z5T_R 3DAB_B 3DAC_B 2Z5S_Q 2K8F_B 2L14_B 1YCR_B.
Probab=38.76 E-value=13 Score=23.38 Aligned_cols=14 Identities=7% Similarity=0.007 Sum_probs=9.9
Q ss_pred cCCChHHHHHHHHH
Q 016344 176 FELDNPVYLYTKKR 189 (391)
Q Consensus 176 ~~ls~~~fe~l~~~ 189 (391)
+-|+++.|++||+.
T Consensus 8 ~PLSQeTF~~LW~~ 21 (25)
T PF08563_consen 8 LPLSQETFSDLWNL 21 (25)
T ss_dssp ---STCCHHHHHHT
T ss_pred CCccHHHHHHHHHh
Confidence 45889999999974
No 22
>PF10576 EndIII_4Fe-2S: Iron-sulfur binding domain of endonuclease III; InterPro: IPR003651 Endonuclease III (4.2.99.18 from EC) is a DNA repair enzyme which removes a number of damaged pyrimidines from DNA via its glycosylase activity and also cleaves the phosphodiester backbone at apurinic / apyrimidinic sites via a beta-elimination mechanism [, ]. The structurally related DNA glycosylase MutY recognises and excises the mutational intermediate 8-oxoguanine-adenine mispair []. The 3-D structures of Escherichia coli endonuclease III [] and catalytic domain of MutY [] have been determined. The structures contain two all-alpha domains: a sequence-continuous, six-helix domain (residues 22-132) and a Greek-key, four-helix domain formed by one N-terminal and three C-terminal helices (residues 1-21 and 133-211) together with the [Fe4S4] cluster. The cluster is bound entirely within the C-terminal loop by four cysteine residues with a ligation pattern Cys-(Xaa)6-Cys-(Xaa)2-Cys-(Xaa)5-Cys which is distinct from all other known Fe4S4 proteins. This structural motif is referred to as a [Fe4S4] cluster loop (FCL) []. Two DNA-binding motifs have been proposed, one at either end of the interdomain groove: the helix-hairpin-helix (HhH) and FCL motifs. The primary role of the iron-sulphur cluster appears to involve positioning conserved basic residues for interaction with the DNA phosphate backbone by forming the loop of the FCL motif [, ]. The iron-sulphur cluster loop (FCL) is also found in DNA-(apurinic or apyrimidinic site) lyase, a subfamily of endonuclease III. The enzyme has both apurinic and apyrimidinic endonuclease activity and a DNA N-glycosylase activity. It cuts damaged DNA at cytosines, thymines and guanines, and acts on the damaged strand 5' of the damaged site. The enzyme binds a 4Fe-4S cluster which is not important for the catalytic activity, but is probably involved in the alignment of the enzyme along the DNA strand.; GO: 0004519 endonuclease activity, 0051539 4 iron, 4 sulfur cluster binding; PDB: 1VRL_A 1RRQ_A 3G0Q_A 3FSQ_A 1RRS_A 3FSP_A 2ABK_A 1KG7_A 1KG2_A 1MUN_A ....
Probab=37.67 E-value=15 Score=21.06 Aligned_cols=14 Identities=36% Similarity=0.923 Sum_probs=8.7
Q ss_pred CCCCCccCCCCcee
Q 016344 89 PTDSCEPCPSNGEC 102 (391)
Q Consensus 89 ~~p~C~PCP~hAiC 102 (391)
..|.|.-||-+..|
T Consensus 4 r~P~C~~Cpl~~~C 17 (17)
T PF10576_consen 4 RKPKCEECPLADYC 17 (17)
T ss_dssp SS--GGG-TTGGG-
T ss_pred CCCccccCCCcccC
Confidence 47899999999887
No 23
>PF07974 EGF_2: EGF-like domain; InterPro: IPR013111 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length. This entry contains EGF domains found in a variety of extracellular and membrane proteins
Probab=36.33 E-value=33 Score=22.62 Aligned_cols=20 Identities=35% Similarity=0.856 Sum_probs=16.9
Q ss_pred cCCCCceec--CCeeeeCCCce
Q 016344 95 PCPSNGECH--QGKLECFHGYR 114 (391)
Q Consensus 95 PCP~hAiC~--~g~l~C~~gYv 114 (391)
.|=.||+|. .|.=.|++||.
T Consensus 7 ~C~~~G~C~~~~g~C~C~~g~~ 28 (32)
T PF07974_consen 7 ICSGHGTCVSPCGRCVCDSGYT 28 (32)
T ss_pred ccCCCCEEeCCCCEEECCCCCc
Confidence 588999999 56779999984
No 24
>PRK09458 pspB phage shock protein B; Provisional
Probab=34.52 E-value=1e+02 Score=24.63 Aligned_cols=29 Identities=24% Similarity=0.281 Sum_probs=16.5
Q ss_pred HHHHHHHH-HHHHHHHH-------HHHHHHHHHHHHH
Q 016344 244 LLVGCLLL-LWKVHRRR-------YFAIRVEELYHQV 272 (391)
Q Consensus 244 ~~v~i~~l-~~~~~r~~-------~e~~~v~~Lv~~v 272 (391)
++|+-.|+ .+|..+++ .+.++.++|++.+
T Consensus 15 ifVaPiWL~LHY~sk~~~~~~Ls~~d~~~L~~L~~~A 51 (75)
T PRK09458 15 LFVAPIWLWLHYRSKRQGSQGLSQEEQQRLAQLTEKA 51 (75)
T ss_pred HHHHHHHHHHhhcccccCCCCCCHHHHHHHHHHHHHH
Confidence 34444444 56655443 4566677777765
No 25
>cd00053 EGF Epidermal growth factor domain, found in epidermal growth factor (EGF) presents in a large number of proteins, mostly animal; the list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied; the functional significance of EGF-like domains in what appear to be unrelated proteins is not yet clear; a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase); the domain includes six cysteine residues which have been shown to be involved in disulfide bonds; the main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet; Subdomains between the conserved cysteines vary in length; the region between the 5th and 6th cysteine contains two conserved glycines of which at least one is present in most EGF-like domains; a subset of these bind calcium.
Probab=33.95 E-value=40 Score=20.86 Aligned_cols=25 Identities=32% Similarity=0.890 Sum_probs=17.9
Q ss_pred cCCCCceecCC--ee--eeCCCceecCCCc
Q 016344 95 PCPSNGECHQG--KL--ECFHGYRKHGKLC 120 (391)
Q Consensus 95 PCP~hAiC~~g--~l--~C~~gYvl~~~~C 120 (391)
+|..||+|.+. .. .|..||... ..|
T Consensus 7 ~C~~~~~C~~~~~~~~C~C~~g~~g~-~~C 35 (36)
T cd00053 7 PCSNGGTCVNTPGSYRCVCPPGYTGD-RSC 35 (36)
T ss_pred CCCCCCEEecCCCCeEeECCCCCccc-CCc
Confidence 67788999985 23 899998654 344
No 26
>PRK11677 hypothetical protein; Provisional
Probab=33.88 E-value=1.6e+02 Score=26.09 Aligned_cols=6 Identities=17% Similarity=0.977 Sum_probs=2.3
Q ss_pred HHHHHH
Q 016344 244 LLVGCL 249 (391)
Q Consensus 244 ~~v~i~ 249 (391)
+++|.+
T Consensus 14 ~iiG~~ 19 (134)
T PRK11677 14 IIIGAV 19 (134)
T ss_pred HHHHHH
Confidence 333433
No 27
>PF00558 Vpu: Vpu protein; InterPro: IPR008187 The Human immunodeficiency virus 1 (HIV-1) Vpu protein acts in the degradation of CD4 in the endoplasmic reticulum and in the enhancement of virion release from the plasma membrane of infected cells [].; GO: 0019076 release of virus from host; PDB: 2JPX_A 1PI8_A 2GOH_A 2GOF_A 1PI7_A 1PJE_A 1VPU_A 2K7Y_A.
Probab=33.39 E-value=46 Score=26.95 Aligned_cols=19 Identities=16% Similarity=0.238 Sum_probs=5.5
Q ss_pred HHHHHHHHHHHHHHHHHHH
Q 016344 256 HRRRYFAIRVEELYHQVCE 274 (391)
Q Consensus 256 ~r~~~e~~~v~~Lv~~vl~ 274 (391)
|++.+.+++++.+++.+.+
T Consensus 30 Yrk~~rqrkId~li~RIre 48 (81)
T PF00558_consen 30 YRKIKRQRKIDRLIERIRE 48 (81)
T ss_dssp ---------CHHHHHHHHC
T ss_pred HHHHHHHHhHHHHHHHHHc
Confidence 5555556667766654433
No 28
>PRK07597 secE preprotein translocase subunit SecE; Reviewed
Probab=32.12 E-value=85 Score=23.72 Aligned_cols=28 Identities=25% Similarity=0.420 Sum_probs=19.9
Q ss_pred CCCChhhHHHHHHHHHHHHHHHHHHHHH
Q 016344 38 LFPSKQDLLRLITVVAIASSVALTCNYL 65 (391)
Q Consensus 38 ~~~~~~~~~~~~~~~~ia~~~~~~c~~l 65 (391)
-.|+++|..+...+.++++++..+..++
T Consensus 25 ~WPs~~e~~~~t~~Vi~~~~~~~~~i~~ 52 (64)
T PRK07597 25 TWPTRKELVRSTIVVLVFVAFFALFFYL 52 (64)
T ss_pred cCcCHHHHHhHHHHHHHHHHHHHHHHHH
Confidence 3699999998888777777665444443
No 29
>PF06679 DUF1180: Protein of unknown function (DUF1180); InterPro: IPR009565 This entry consists of several hypothetical eukaryotic proteins thought to be membrane proteins. Their function is unknown.
Probab=31.90 E-value=1.8e+02 Score=26.51 Aligned_cols=31 Identities=23% Similarity=0.231 Sum_probs=23.2
Q ss_pred CCCCCCChhhHHHHHHHHHHHHHHHHHHHHH
Q 016344 35 PQSLFPSKQDLLRLITVVAIASSVALTCNYL 65 (391)
Q Consensus 35 ~~~~~~~~~~~~~~~~~~~ia~~~~~~c~~l 65 (391)
|..+-+.+.-+.|.++||..+++.+.+|+++
T Consensus 84 ~s~~~~d~~~l~R~~~Vl~g~s~l~i~yfvi 114 (163)
T PF06679_consen 84 PSPSSPDSPMLKRALYVLVGLSALAILYFVI 114 (163)
T ss_pred cCCCcCCccchhhhHHHHHHHHHHHHHHHHH
Confidence 3345567777888999988888888777774
No 30
>TIGR00964 secE_bact preprotein translocase, SecE subunit, bacterial. This model represents exclusively the bacterial (and some organellar) SecE protein. SecE is part of the core heterotrimer, SecYEG, of the Sec preprotein translocase system. Other components are the ATPase SecA, a cytosolic chaperone SecB, and an accessory complex of SecDF and YajC.
Probab=31.66 E-value=89 Score=22.96 Aligned_cols=26 Identities=19% Similarity=0.327 Sum_probs=18.5
Q ss_pred CCChhhHHHHHHHHHHHHHHHHHHHH
Q 016344 39 FPSKQDLLRLITVVAIASSVALTCNY 64 (391)
Q Consensus 39 ~~~~~~~~~~~~~~~ia~~~~~~c~~ 64 (391)
.|+|+|..+...+.++++++..+..+
T Consensus 17 WPt~~e~~~~t~~Vi~~~~~~~~~~~ 42 (55)
T TIGR00964 17 WPSRKELITYTIVVIVFVIFFSLFLF 42 (55)
T ss_pred CcCHHHHHhHHHHHHHHHHHHHHHHH
Confidence 69999998887777776666444433
No 31
>PF07271 Cytadhesin_P30: Cytadhesin P30/P32; InterPro: IPR009896 This family consists of several Mycoplasma species specific Cytadhesin P32 and P30 proteins. P30 has been found to be membrane associated and localised on the tip organelle. It is thought that it is important in cytadherence and virulence [].; GO: 0007157 heterophilic cell-cell adhesion, 0009405 pathogenesis, 0016021 integral to membrane
Probab=30.52 E-value=1.6e+02 Score=29.16 Aligned_cols=17 Identities=24% Similarity=0.022 Sum_probs=10.8
Q ss_pred HHHHHHHHHHHHHHHHH
Q 016344 260 YFAIRVEELYHQVCEIL 276 (391)
Q Consensus 260 ~e~~~v~~Lv~~vl~~L 276 (391)
+|+++.++++++.-.+-
T Consensus 104 ee~e~~~q~~e~~~~i~ 120 (279)
T PF07271_consen 104 EEKEEHEQLAEQLGRIS 120 (279)
T ss_pred HHHHHHHHHHHHHHHHH
Confidence 56667788887654433
No 32
>PF06143 Baculo_11_kDa: Baculovirus 11 kDa family; InterPro: IPR009313 This is a family of uncharacterised Baculovirus proteins that are all about 11 kDa in size.
Probab=29.87 E-value=3.1e+02 Score=22.39 Aligned_cols=22 Identities=14% Similarity=0.354 Sum_probs=12.3
Q ss_pred HHHHHHHHHHHHHHHHHHHHHH
Q 016344 230 WVSTHALIIVPVCSLLVGCLLL 251 (391)
Q Consensus 230 ~i~~~~~~i~~~l~~~v~i~~l 251 (391)
.++.+++.+.+++++++.++++
T Consensus 32 firdFvLVic~~lVfVii~lFi 53 (84)
T PF06143_consen 32 FIRDFVLVICCFLVFVIIVLFI 53 (84)
T ss_pred HHHHHHHHHHHHHHHHHHHHHH
Confidence 4566666666665555444444
No 33
>PHA03399 pif3 per os infectivity factor 3; Provisional
Probab=28.93 E-value=66 Score=30.36 Aligned_cols=32 Identities=19% Similarity=0.510 Sum_probs=21.8
Q ss_pred CCCCCCCCCCCCCCCCCCCCCCCccCCCCceecCC--------eeeeCCCce
Q 016344 71 STSKPFCDSNLLLDSPQSPTDSCEPCPSNGECHQG--------KLECFHGYR 114 (391)
Q Consensus 71 ~~~~~fCdS~~~~~~~~~~~p~C~PCP~hAiC~~g--------~l~C~~gYv 114 (391)
++..--|+++. +||=.+..|.++ .+.|+.||=
T Consensus 47 r~~ivDC~~t~------------lPCVtD~QC~dnC~~~~~~~~~~C~~GFC 86 (200)
T PHA03399 47 RNGIVDCSLTR------------LPCVTDQQCRDNCAIGSAAGVMTCDGGFC 86 (200)
T ss_pred ccCcccCcCCc------------CCcccHHHHHHHHHhccccceEECCCCee
Confidence 44555566653 388899888754 458988863
No 34
>PF07543 PGA2: Protein trafficking PGA2; InterPro: IPR011431 A Saccharomyces cerevisiae (Baker's yeast) member of this family (PGA2, P53903 from SWISSPROT) is a single pass membrane protein which has been implicated in protein trafficking [, ].
Probab=27.93 E-value=1.7e+02 Score=26.02 Aligned_cols=12 Identities=17% Similarity=0.094 Sum_probs=6.6
Q ss_pred cccccccccccC
Q 016344 293 WVVASRLRDHLL 304 (391)
Q Consensus 293 yl~~~qLRD~LL 304 (391)
=|+.+.||+..-
T Consensus 62 k~s~n~lRg~~~ 73 (140)
T PF07543_consen 62 KISPNALRGGKA 73 (140)
T ss_pred cCCchhhccccc
Confidence 356666666433
No 35
>KOG0474 consensus Cl- channel CLC-7 and related proteins (CLC superfamily) [Inorganic ion transport and metabolism]
Probab=27.93 E-value=86 Score=34.62 Aligned_cols=24 Identities=29% Similarity=0.584 Sum_probs=13.4
Q ss_pred CCCCccCCCCceecCC-eeeeCCCc
Q 016344 90 TDSCEPCPSNGECHQG-KLECFHGY 113 (391)
Q Consensus 90 ~p~C~PCP~hAiC~~g-~l~C~~gY 113 (391)
-..|.|||....=..- .+-|.+|+
T Consensus 396 l~~C~P~~~~~~~~~~p~f~Cp~~~ 420 (762)
T KOG0474|consen 396 LADCQPCPPSITEGQCPTFFCPDGE 420 (762)
T ss_pred HhcCCCCCCCcccccCccccCCCCc
Confidence 3478888876532211 25676664
No 36
>COG0690 SecE Preprotein translocase subunit SecE [Intracellular trafficking and secretion]
Probab=27.83 E-value=1.2e+02 Score=23.75 Aligned_cols=28 Identities=18% Similarity=0.376 Sum_probs=18.4
Q ss_pred CCCChhhHHHHHHHHHHHHHHHHHHHHH
Q 016344 38 LFPSKQDLLRLITVVAIASSVALTCNYL 65 (391)
Q Consensus 38 ~~~~~~~~~~~~~~~~ia~~~~~~c~~l 65 (391)
-+|++.|..+...+.++..+++.+..++
T Consensus 35 ~WPsrke~~~~t~~Vl~~v~~~s~~~~~ 62 (73)
T COG0690 35 VWPTRKELIRSTLIVLVVVAFFSLFLYG 62 (73)
T ss_pred cCCCHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 3699999888877666655554443333
No 37
>PF05568 ASFV_J13L: African swine fever virus J13L protein; InterPro: IPR008385 This family consists of several African swine fever virus (ASFV) j13L proteins [, , ].
Probab=27.61 E-value=1.5e+02 Score=26.71 Aligned_cols=10 Identities=40% Similarity=0.630 Sum_probs=4.5
Q ss_pred HHHHHHHHHH
Q 016344 230 WVSTHALIIV 239 (391)
Q Consensus 230 ~i~~~~~~i~ 239 (391)
++..|+..|+
T Consensus 26 ffsthm~tIL 35 (189)
T PF05568_consen 26 FFSTHMYTIL 35 (189)
T ss_pred HHHHHHHHHH
Confidence 3455554443
No 38
>PHA02817 EEV Host range protein; Provisional
Probab=26.67 E-value=64 Score=30.97 Aligned_cols=51 Identities=18% Similarity=0.421 Sum_probs=30.1
Q ss_pred CCCCCCCCCCCCCCCCCCCCCCc--cCC----CCc----------eecCCe--eeeCCCceecC---CCcccChh
Q 016344 72 TSKPFCDSNLLLDSPQSPTDSCE--PCP----SNG----------ECHQGK--LECFHGYRKHG---KLCVEDGD 125 (391)
Q Consensus 72 ~~~~fCdS~~~~~~~~~~~p~C~--PCP----~hA----------iC~~g~--l~C~~gYvl~~---~~CV~D~~ 125 (391)
...-.|..+. .-....|.|+ .|| +|| ..+... +.|++||.+.| -.|..|+.
T Consensus 66 ~~~i~C~~dG---~Ws~~~P~C~~v~C~~P~i~NG~v~~~~~~~~y~yg~~Vty~C~~Gy~L~G~~~~tC~~~G~ 137 (225)
T PHA02817 66 EKNIICEKDG---KWNKEFPVCKIIRCRFPALQNGFVNGIPDSKKFYYESEVSFSCKPGFVLIGTKYSVCGINSS 137 (225)
T ss_pred CCeEEECCCC---cCCCCCCeeeeeECCCCCCcCceeEccccCCceEcCCEEEEEcCCCCEEcCCCceEECCCCe
Confidence 3456787543 1223468997 686 344 234443 49999999975 34555543
No 39
>KOG4403 consensus Cell surface glycoprotein STIM, contains SAM domain [General function prediction only]
Probab=26.49 E-value=3.3e+02 Score=28.91 Aligned_cols=13 Identities=15% Similarity=0.562 Sum_probs=8.2
Q ss_pred cccchhhHHHHhh
Q 016344 156 IWVEENDIWNDLE 168 (391)
Q Consensus 156 ~~i~e~dL~~~~~ 168 (391)
..|+.+|||+.-.
T Consensus 116 ~~ItVedLWeaW~ 128 (575)
T KOG4403|consen 116 KHITVEDLWEAWK 128 (575)
T ss_pred cceeHHHHHHHHH
Confidence 3577777777643
No 40
>PF06247 Plasmod_Pvs28: Plasmodium ookinete surface protein Pvs28; InterPro: IPR010423 This family consists of several ookinete surface protein (Pvs28) from several species of Plasmodium. Pvs25 and Pvs28 are expressed on the surface of ookinetes. These proteins are potential candidates for vaccine and induce antibodies that block the infectivity of Plasmodium vivax in immunised animals [].; GO: 0009986 cell surface, 0016020 membrane; PDB: 1Z3G_B 1Z1Y_B 1Z27_A.
Probab=26.13 E-value=25 Score=32.88 Aligned_cols=32 Identities=25% Similarity=0.669 Sum_probs=25.2
Q ss_pred cCCCCceecCC-e------e--eeCCCceecCCCcccChhh
Q 016344 95 PCPSNGECHQG-K------L--ECFHGYRKHGKLCVEDGDI 126 (391)
Q Consensus 95 PCP~hAiC~~g-~------l--~C~~gYvl~~~~CV~D~~k 126 (391)
||=+.|.|... . + .|.+||++..+.|+|+.=.
T Consensus 51 ~Cgdya~C~~~~~~~~~~~~~C~C~~gY~~~~~vCvp~~C~ 91 (197)
T PF06247_consen 51 PCGDYAKCINQANKGEERAYKCDCINGYILKQGVCVPNKCN 91 (197)
T ss_dssp EEETTEEEEE-SSTTSSTSEEEEE-TTEEESSSSEEEGGGS
T ss_pred cccchhhhhcCCCcccceeEEEecccCceeeCCeEchhhcC
Confidence 78999999865 2 2 8999999999999998743
No 41
>PF07466 DUF1517: Protein of unknown function (DUF1517); InterPro: IPR010903 This family consists of several hypothetical glycine rich plant and bacterial proteins of around 300 residues in length. The function of this family is unknown.
Probab=25.75 E-value=1.9e+02 Score=28.72 Aligned_cols=23 Identities=4% Similarity=0.163 Sum_probs=12.4
Q ss_pred hhHHHHHHHHHHHHHHHHHHHHH
Q 016344 43 QDLLRLITVVAIASSVALTCNYL 65 (391)
Q Consensus 43 ~~~~~~~~~~~ia~~~~~~c~~l 65 (391)
..|.-++.+|+++.+++++..++
T Consensus 62 gg~~gl~~iLIl~~Ia~~vv~~~ 84 (289)
T PF07466_consen 62 GGFGGLFDILILFGIAFFVVRFF 84 (289)
T ss_pred cccchHHHHHHHHHHHHHHHHHH
Confidence 33455666666555555555444
No 42
>PF14316 DUF4381: Domain of unknown function (DUF4381)
Probab=25.53 E-value=1.7e+02 Score=25.59 Aligned_cols=15 Identities=27% Similarity=0.275 Sum_probs=6.9
Q ss_pred HHHHHHHHHHHHhhh
Q 016344 269 YHQVCEILEENALMS 283 (391)
Q Consensus 269 v~~vl~~L~~q~~~~ 283 (391)
..++-.+|+..+..+
T Consensus 70 ~~~l~~LLKr~a~~~ 84 (146)
T PF14316_consen 70 LAALNELLKRVALQY 84 (146)
T ss_pred HHHHHHHHHHHHHHh
Confidence 334455555444433
No 43
>PF03672 UPF0154: Uncharacterised protein family (UPF0154); InterPro: IPR005359 The proteins in this entry are functionally uncharacterised.
Probab=25.32 E-value=2.4e+02 Score=21.86 Aligned_cols=18 Identities=6% Similarity=0.143 Sum_probs=9.1
Q ss_pred HHHHHHHHHHHHHHHHHH
Q 016344 243 SLLVGCLLLLWKVHRRRY 260 (391)
Q Consensus 243 ~~~v~i~~l~~~~~r~~~ 260 (391)
++++|+++.++++.+.-+
T Consensus 10 G~~~Gff~ar~~~~k~l~ 27 (64)
T PF03672_consen 10 GAVIGFFIARKYMEKQLK 27 (64)
T ss_pred HHHHHHHHHHHHHHHHHH
Confidence 444555555565554443
No 44
>PF00584 SecE: SecE/Sec61-gamma subunits of protein translocation complex; InterPro: IPR001901 Secretion across the inner membrane in some Gram-negative bacteria occurs via the preprotein translocase pathway. Proteins are produced in the cytoplasm as precursors, and require a chaperone subunit to direct them to the translocase component []. From there, the mature proteins are either targeted to the outer membrane, or remain as periplasmic proteins. The translocase protein subunits are encoded on the bacterial chromosome. The translocase itself comprises 7 proteins, including a chaperone protein (SecB), an ATPase (SecA), an integral membrane complex (SecCY, SecE and SecG), and two additional membrane proteins that promote the release of the mature peptide into the periplasm (SecD and SecF) []. The chaperone protein SecB [] is a highly acidic homotetrameric protein that exists as a "dimer of dimers" in the bacterial cytoplasm. SecB maintains preproteins in an unfolded state after translation, and targets these to the peripheral membrane protein ATPase SecA for secretion []. SecE, part of the main SecYEG translocase complex, is ~106 residues in length, and spans the inner membrane of the Gram-negative bacterial envelope. Together with SecY and SecG, SecE forms a multimeric channel through which preproteins are translocated, using both proton motive forces and ATP-driven secretion. The latter is mediated by SecA. In eukaryotes, the evolutionary related protein sec61-gamma plays a role in protein translocation through the endoplasmic reticulum; it is part of a trimeric complex that also consist of sec61-alpha and beta []. Both secE and sec61-gamma are small proteins of about 60 to 90 amino acids that contain a single transmembrane region at their C-terminal extremity (Escherichia coli secE is an exception, in that it possess an extra N-terminal segment of 60 residues that contains two additional transmembrane domains) [].; GO: 0006605 protein targeting, 0006886 intracellular protein transport, 0016020 membrane; PDB: 3J01_B 2WW9_B 2WWA_B 3DL8_C 2WWB_B 3DIN_G 2ZJS_E 2ZQP_E.
Probab=25.25 E-value=1.7e+02 Score=21.39 Aligned_cols=21 Identities=24% Similarity=0.505 Sum_probs=15.2
Q ss_pred CCChhhHHHHHHHHHHHHHHH
Q 016344 39 FPSKQDLLRLITVVAIASSVA 59 (391)
Q Consensus 39 ~~~~~~~~~~~~~~~ia~~~~ 59 (391)
.|+++|..+.-.+.++..++.
T Consensus 18 WP~~~e~~~~t~~Vl~~~~i~ 38 (57)
T PF00584_consen 18 WPSRKELLKSTIIVLVFVIIF 38 (57)
T ss_dssp CCCTHHHHHHHHHHHHHHHHH
T ss_pred CCCHHHHHHHHHHHHHHHHHH
Confidence 599999888776666655553
No 45
>PF14991 MLANA: Protein melan-A; PDB: 2GTZ_F 2GT9_F 3MRO_P 2GUO_C 3MRQ_P 2GTW_C 3L6F_C 3MRP_P.
Probab=24.70 E-value=18 Score=31.11 Aligned_cols=22 Identities=32% Similarity=0.679 Sum_probs=0.0
Q ss_pred HHHHHHHHHHH--HHHHHHHHHHH
Q 016344 241 VCSLLVGCLLL--LWKVHRRRYFA 262 (391)
Q Consensus 241 ~l~~~v~i~~l--~~~~~r~~~e~ 262 (391)
++++++|++++ .||++||.-++
T Consensus 31 iL~VILgiLLliGCWYckRRSGYk 54 (118)
T PF14991_consen 31 ILIVILGILLLIGCWYCKRRSGYK 54 (118)
T ss_dssp ------------------------
T ss_pred eHHHHHHHHHHHhheeeeecchhh
Confidence 33445555444 68887765443
No 46
>PF09402 MSC: Man1-Src1p-C-terminal domain; InterPro: IPR018996 This entry represents the Inner nuclear membrane proteins MAN1 (also known as LEM domain-containing protein 3) and LEM domain-containing protein 2 (or LEM protein 2). Emerin and MAN1 are LEM domain-containing integral membrane proteins of the vertebrate nuclear envelope []. MAN1 is an integral protein of the inner nuclear membrane which binds to chromatin associated proteins and plays a role in nuclear organisation. The C-terminal nulceoplasmic region forms a DNA binding winged helix and binds to Smad []. LEM protein 2 is an essential protein involved in chromosome segregation and cell division, probably via its interaction with lmn-1, the main component of nuclear lamina. Has some overlapping function with emr-1.; GO: 0005639 integral to nuclear inner membrane; PDB: 2CH0_A.
Probab=23.70 E-value=26 Score=34.83 Aligned_cols=70 Identities=16% Similarity=0.221 Sum_probs=0.0
Q ss_pred HHHHHHHHHHHHHHHHHHHhhhhcc--CCCCCCcccccccccccCCCCC-----cCchhhHHHHHHHHhcCCCccee
Q 016344 262 AIRVEELYHQVCEILEENALMSKSV--NGECEPWVVASRLRDHLLLPKE-----RKDPVIWKKVEELVQEDSRVDQY 331 (391)
Q Consensus 262 ~~~v~~Lv~~vl~~L~~q~~~~~~~--~~~~~pyl~~~qLRD~LL~~~~-----r~r~~LW~kV~k~Ve~nSnVrt~ 331 (391)
.+.+..|++.+.+.|++++..+.=+ .....++++...|+|.+..... ..-+.+|+.+...+.++..|...
T Consensus 98 ~~~i~~l~~~~~~~Lr~~~a~~~Cg~~~~~~~~~ls~~el~~~~~~~~~~~~~~~efe~l~~~a~~~L~~~~ei~~~ 174 (334)
T PF09402_consen 98 EEKIEELAKKILDELRERNAQYECGDSEDDESPGLSEEELKDILSSKKSPWISDEEFEELWSAALQELKKNPEIIIR 174 (334)
T ss_dssp -----------------------------------------------------------------------------
T ss_pred HHHHHHHHHHHHHHHHHHHhhcccCCCCCCCCCCCcHHHHHHHHHhccCccccHHHHHHHHHHHHHHHHhCCcEEEe
Confidence 4568888999999998776655323 2456899999999999995441 23388999999998887666554
No 47
>PF07699 GCC2_GCC3: GCC2 and GCC3; InterPro: IPR011641 Protein phosphorylation, which plays a key role in most cellular activities, is a reversible process mediated by protein kinases and phosphoprotein phosphatases. Protein kinases catalyse the transfer of the gamma phosphate from nucleotide triphosphates (often ATP) to one or more amino acid residues in a protein substrate side chain, resulting in a conformational change affecting protein function. Phosphoprotein phosphatases catalyse the reverse process. Protein kinases fall into three broad classes, characterised with respect to substrate specificity []: Serine/threonine-protein kinases Tyrosine-protein kinases Dual specific protein kinases (e.g. MEK - phosphorylates both Thr and Tyr on target proteins) Protein kinase function has been evolutionarily conserved from Escherichia coli to human []. Protein kinases play a role in a multitude of cellular processes, including division, proliferation, apoptosis, and differentiation []. Phosphorylation usually results in a functional change of the target protein by changing enzyme activity, cellular location, or association with other proteins. The catalytic subunits of protein kinases are highly conserved, and several structures have been solved [], leading to large screens to develop kinase-specific inhibitors for the treatments of a number of diseases []. Tyrosine-protein kinases can transfer a phosphate group from ATP to a tyrosine residue in a protein. These enzymes can be divided into two main groups []: Receptor tyrosine kinases (RTK), which are transmembrane proteins involved in signal transduction; they play key roles in growth, differentiation, metabolism, adhesion, motility, death and oncogenesis []. RTKs are composed of 3 domains: an extracellular domain (binds ligand), a transmembrane (TM) domain, and an intracellular catalytic domain (phosphorylates substrate). The TM domain plays an important role in the dimerisation process necessary for signal transduction []. Cytoplasmic / non-receptor tyrosine kinases, which act as regulatory proteins, playing key roles in cell differentiation, motility, proliferation, and survival. For example, the Src-family of protein-tyrosine kinases []. This entry represents various ephrin type A and B receptors, which have tyrosine kinase activity.
Probab=23.59 E-value=65 Score=22.76 Aligned_cols=28 Identities=21% Similarity=0.525 Sum_probs=15.6
Q ss_pred CCCCCCCCCCCCCCCCCCCCCccCCCCcee
Q 016344 73 SKPFCDSNLLLDSPQSPTDSCEPCPSNGEC 102 (391)
Q Consensus 73 ~~~fCdS~~~~~~~~~~~p~C~PCP~hAiC 102 (391)
.-.-|.-+.+ +.......|++||.+-.-
T Consensus 10 ~C~~Cp~GtY--q~~~g~~~C~~Cp~g~~T 37 (48)
T PF07699_consen 10 KCQPCPKGTY--QDEEGQTSCTPCPPGSTT 37 (48)
T ss_pred ccCCCCCCcc--CCccCCccCccCcCCCcc
Confidence 4445666652 222334478888887544
No 48
>PF10588 NADH-G_4Fe-4S_3: NADH-ubiquinone oxidoreductase-G iron-sulfur binding region; InterPro: IPR019574 NADH:ubiquinone oxidoreductase (complex I) (1.6.5.3 from EC) is a respiratory-chain enzyme that catalyses the transfer of two electrons from NADH to ubiquinone in a reaction that is associated with proton translocation across the membrane (NADH + ubiquinone = NAD+ + ubiquinol) []. Complex I is a major source of reactive oxygen species (ROS) that are predominantly formed by electron transfer from FMNH(2). Complex I is found in bacteria, cyanobacteria (as a NADH-plastoquinone oxidoreductase), archaea [], mitochondira, and in the hydrogenosome, a mitochondria-derived organelle. In general, the bacterial complex consists of 14 different subunits, while the mitochondrial complex contains homologues to these subunits in addition to approximately 31 additional proteins []. Mitochondrial complex I, which is located in the inner mitochondrial membrane, is the largest multimeric respiratory enzyme in the mitochondria, consisting of more than 40 subunits, one FMN co-factor and eight FeS clusters []. The assembly of mitochondrial complex I is an intricate process that requires the cooperation of the nuclear and mitochondrial genomes [, ]. Mitochondrial complex I can cycle between active and deactive forms that can be distinguished by the reactivity towards divalent cations and thiol-reactive agents. All redox prosthetic groups reside in the peripheral arm of the L-shaped structure. The NADH oxidation domain harbouring the FMN cofactor is connected via a chain of iron-sulphur clusters to the ubiquinone reduction site that is located in a large pocket formed by the PSST and 49kDa subunits of complex I []. This entry describes the G subunit (one of 14 subunits, A to N) of the NADH-quinone oxidoreductase complex I which generally couples NADH and ubiquinone oxidation/reduction in bacteria and mammalian mitochondria while translocating protons, but may act on NADPH and/or plastoquinone in cyanobacteria and plant chloroplasts. This family does not contain related subunits from formate dehydrogenase complexes. This entry represents the iron-sulphur binding domain of the G subunit.; GO: 0016491 oxidoreductase activity, 0055114 oxidation-reduction process; PDB: 3M9S_C 2FUG_L 3IAS_L 2YBB_3 3IAM_3 3I9V_3.
Probab=23.20 E-value=35 Score=23.74 Aligned_cols=16 Identities=31% Similarity=0.866 Sum_probs=8.5
Q ss_pred CCCCCCccCCCCceec
Q 016344 88 SPTDSCEPCPSNGECH 103 (391)
Q Consensus 88 ~~~p~C~PCP~hAiC~ 103 (391)
+++-.|.-|+.+|.|.
T Consensus 11 ~H~~dC~~C~~~G~Ce 26 (41)
T PF10588_consen 11 NHPLDCPTCDKNGNCE 26 (41)
T ss_dssp T----TTT-TTGGG-H
T ss_pred CCCCcCcCCCCCCCCH
Confidence 4456899999999994
No 49
>PF08114 PMP1_2: ATPase proteolipid family; InterPro: IPR012589 This family consists of small proteolipids associated with the plasma membrane H+ ATPase. Two proteolipids (PMP1 and PMP2) are associated with the ATPase and both genes are similarly expressed in the wild-type strain of yeast. No modification of the level of transcription of one PMP gene is detected in a strain deleted of the other. Though both proteolipids show similarity with other small proteolipids associated with other cation -transporting ATPases, their functions remain unclear [].
Probab=23.13 E-value=1.4e+02 Score=21.13 Aligned_cols=17 Identities=18% Similarity=0.120 Sum_probs=7.0
Q ss_pred HHHHHHHHHHHHHHHHH
Q 016344 246 VGCLLLLWKVHRRRYFA 262 (391)
Q Consensus 246 v~i~~l~~~~~r~~~e~ 262 (391)
+++..+.-.+||+.+.+
T Consensus 20 v~i~iva~~iYRKw~aR 36 (43)
T PF08114_consen 20 VGIGIVALFIYRKWQAR 36 (43)
T ss_pred HHHHHHHHHHHHHHHHH
Confidence 33443434444444333
No 50
>smart00032 CCP Domain abundant in complement control proteins; SUSHI repeat; short complement-like repeat (SCR). The complement control protein (CCP) modules (also known as short consensus repeats SCRs or SUSHI repeats) contain approximately 60 amino acid residues and have been identified in several proteins of the complement system. A missense mutation in seventh CCP domain causes deficiency of the b subunit of factor XIII.
Probab=22.62 E-value=51 Score=22.70 Aligned_cols=20 Identities=40% Similarity=0.815 Sum_probs=14.9
Q ss_pred eeeeCCCceecC---CCcccChh
Q 016344 106 KLECFHGYRKHG---KLCVEDGD 125 (391)
Q Consensus 106 ~l~C~~gYvl~~---~~CV~D~~ 125 (391)
++.|++||.+.+ -.|..|+.
T Consensus 27 ~~~C~~Gy~l~g~~~~~C~~~g~ 49 (57)
T smart00032 27 TYSCNPGYTLIGSSTITCLEDGT 49 (57)
T ss_pred EEEcCCCCEEcCCCeeEECCCCE
Confidence 359999999975 56776653
No 51
>PF10500 SR-25: Nuclear RNA-splicing-associated protein; InterPro: IPR019532 SR-25, otherwise known as ADP-ribosylation factor-like factor 6-interacting protein 4, is expressed in virtually all tissue types. At the N terminus there is a repeat of serine-arginine (SR repeat), and towards the middle of the protein there are clusters of both serines and of basic amino acids. The presence of many nuclear localisation signals strongly implies that this is a nuclear protein that may contribute to RNA splicing []. SR-25 is also implicated, along with heat-shock-protein-27, as a mediator in the Rac1 (GTPase ras-related C3 botulinum toxin substrate 1; also see IPR019093 from INTERPRO) signalling pathway [].
Probab=22.47 E-value=60 Score=31.12 Aligned_cols=9 Identities=11% Similarity=0.132 Sum_probs=5.6
Q ss_pred CCChHHHHH
Q 016344 177 ELDNPVYLY 185 (391)
Q Consensus 177 ~ls~~~fe~ 185 (391)
-|+.|||+-
T Consensus 159 PmTkEEyea 167 (225)
T PF10500_consen 159 PMTKEEYEA 167 (225)
T ss_pred CCCHHHHHH
Confidence 467776653
No 52
>PF09064 Tme5_EGF_like: Thrombomodulin like fifth domain, EGF-like; InterPro: IPR015149 This domain adopts a fold similar to other EGF domains, with a flat major and a twisted minor beta sheet. Disulphide pairing, however, is not of the usual 1-3, 2-4, 5-6 type; rather 1-2, 3-4, 5-6 pairing is found. Its extended major sheet (strands beta-2 and beta-3 and the connecting loop) projects into thrombin's active site groove. This domain is required for interaction of thrombomodulin with thrombin, and subsequent activation of protein-C []. ; GO: 0004888 transmembrane signaling receptor activity, 0016021 integral to membrane
Probab=22.40 E-value=68 Score=21.77 Aligned_cols=17 Identities=24% Similarity=0.559 Sum_probs=12.4
Q ss_pred eecCC---eeeeCCCceecC
Q 016344 101 ECHQG---KLECFHGYRKHG 117 (391)
Q Consensus 101 iC~~g---~l~C~~gYvl~~ 117 (391)
.|-++ .-.|-+||++..
T Consensus 11 ~CDpn~~~~C~CPeGyIlde 30 (34)
T PF09064_consen 11 DCDPNSPGQCFCPEGYILDE 30 (34)
T ss_pred ccCCCCCCceeCCCceEecC
Confidence 67665 338889999854
No 53
>PF12729 4HB_MCP_1: Four helix bundle sensory module for signal transduction; InterPro: IPR024478 This entry represents a four-helix bundle that operates as a ubiquitous sensory module in prokaryotic signal-transduction, which is known as four-helix bundles methyl-accepting chemotaxis protein (4HB_MCP) domain. The 4HB_MCP is always found between two predicted transmembrane helices indicating that it detects only extracellular signals. In many cases the domain is associated with a cytoplasmic HAMP domain suggesting that most proteins carrying the bundle might share the mechanism of transmembrane signalling which is well-characterised in E coli chemoreceptors [].
Probab=21.74 E-value=3.6e+02 Score=22.64 Aligned_cols=8 Identities=50% Similarity=0.517 Sum_probs=3.9
Q ss_pred ccccccCC
Q 016344 298 RLRDHLLL 305 (391)
Q Consensus 298 qLRD~LL~ 305 (391)
.+++.++.
T Consensus 64 ~~~~~~~~ 71 (181)
T PF12729_consen 64 ALRRYLLA 71 (181)
T ss_pred HHHHhhhc
Confidence 44455553
No 54
>cd00033 CCP Complement control protein (CCP) modules (aka short consensus repeats SCRs or SUSHI repeats) have been identified in several proteins of the complement system. SUSHI repeats (short complement-like repeat, SCR) are abundant in complement control proteins. The complement control protein (CCP) modules (also known as short consensus repeats SCRs or SUSHI repeats) contain approximately 60 amino acid residues and have been identified in several proteins of the complement system. Typically, 2 to 4 modules contribute to a binding site, implying that the orientation of the modules to each other is critical for function.
Probab=21.65 E-value=50 Score=22.94 Aligned_cols=20 Identities=35% Similarity=0.773 Sum_probs=14.5
Q ss_pred eeeeCCCceecC---CCcccChh
Q 016344 106 KLECFHGYRKHG---KLCVEDGD 125 (391)
Q Consensus 106 ~l~C~~gYvl~~---~~CV~D~~ 125 (391)
.+.|++||.+.+ -.|..|+.
T Consensus 26 ~~~C~~Gy~~~g~~~~~C~~~g~ 48 (57)
T cd00033 26 TYSCNEGYTLVGSSTITCTENGG 48 (57)
T ss_pred EEECCCCCeEeCCCeeEECCCCe
Confidence 459999999975 46766553
No 55
>PF01102 Glycophorin_A: Glycophorin A; InterPro: IPR001195 Proteins in this group are responsible for the molecular basis of the blood group antigens, surface markers on the outside of the red blood cell membrane. Most of these markers are proteins, but some are carbohydrates attached to lipids or proteins [Reid M.E., Lomas-Francis C. The Blood Group Antigen FactsBook Academic Press, London / San Diego, (1997)]. Glycophorin A (PAS-2) and glycophorin B (PAS-3) belong to the MNS blood group system and are associated with antigens that include M/N, S/s, U, He, Mi(a), M(c), Vw, Mur, M(g), Vr, M(e), Mt(a), St(a), Ri(a), Cl(a), Ny(a), Hut, Hil, M(v), Far, Mit, Dantu, Hop, Nob, En(a), ENKT, amongst others. Glycophorin A is the major sialoglycoprotein of the erythrocyte membrane []. Structurally, glycophorin A consists of an N-terminal extracellular domain, heavily glycosylated on serine and threonine residues, followed by a transmembrane region and a C-terminal cytoplasmic domain. Other glycophorins in this entry such as Glycophorin B and Glycophorin E represent minor sialoglycoproteins in the erythrocyte membrane.; GO: 0016021 integral to membrane; PDB: 2KPF_B 1AFO_B 2KPE_A.
Probab=21.29 E-value=1.1e+02 Score=26.71 Aligned_cols=22 Identities=5% Similarity=0.256 Sum_probs=13.9
Q ss_pred HHHHHHHHHHHHHHHHHHHHHH
Q 016344 237 IIVPVCSLLVGCLLLLWKVHRR 258 (391)
Q Consensus 237 ~i~~~l~~~v~i~~l~~~~~r~ 258 (391)
.+++++++++.++|++++.+++
T Consensus 73 v~aGvIg~Illi~y~irR~~Kk 94 (122)
T PF01102_consen 73 VMAGVIGIILLISYCIRRLRKK 94 (122)
T ss_dssp HHHHHHHHHHHHHHHHHHHS--
T ss_pred HHHHHHHHHHHHHHHHHHHhcc
Confidence 5567777777777777766554
No 56
>PRK15428 putative propanediol utilization protein PduM; Provisional
Probab=21.25 E-value=82 Score=28.79 Aligned_cols=31 Identities=19% Similarity=0.266 Sum_probs=24.0
Q ss_pred HHHHHHHHHHHHHHHHHHhhhhccCCCCCCccccccccc
Q 016344 263 IRVEELYHQVCEILEENALMSKSVNGECEPWVVASRLRD 301 (391)
Q Consensus 263 ~~v~~Lv~~vl~~L~~q~~~~~~~~~~~~pyl~~~qLRD 301 (391)
..++.||++|+.+|++++.... -++..|||+
T Consensus 4 ~~~~~iV~~Vv~RLk~Ra~~~~--------~ls~~ql~~ 34 (163)
T PRK15428 4 EMLQRIVEEVVARLQRRAQSTA--------TLSVAQLRD 34 (163)
T ss_pred HHHHHHHHHHHHHHHHHhhceE--------EEEHHHccC
Confidence 4578899999999998776543 377778877
No 57
>PF11392 DUF2877: Protein of unknown function (DUF2877); InterPro: IPR021530 This bacterial family of proteins are putative carboxylase proteins however this cannot be confirmed.
Probab=21.17 E-value=52 Score=27.77 Aligned_cols=11 Identities=36% Similarity=0.507 Sum_probs=9.0
Q ss_pred CCCCCCChhhH
Q 016344 35 PQSLFPSKQDL 45 (391)
Q Consensus 35 ~~~~~~~~~~~ 45 (391)
=|||+||.+||
T Consensus 5 G~GLTPSGDD~ 15 (110)
T PF11392_consen 5 GPGLTPSGDDF 15 (110)
T ss_pred CCCCCCchHHH
Confidence 37899999995
Done!