Query psy17223
Match_columns 323
No_of_seqs 271 out of 586
Neff 5.4
Searched_HMMs 46136
Date Fri Aug 16 16:30:26 2013
Command hhsearch -i /work/01045/syshi/Psyhhblits/psy17223.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/17223hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 KOG3865|consensus 100.0 4.6E-62 1E-66 460.8 15.8 202 4-312 91-313 (402)
2 KOG3780|consensus 99.7 9.7E-18 2.1E-22 165.3 13.2 125 16-150 101-230 (427)
3 PF00339 Arrestin_N: Arrestin 99.0 1.1E-09 2.4E-14 91.9 6.2 42 15-64 88-129 (149)
4 PF02752 Arrestin_C: Arrestin 98.6 9E-08 2E-12 78.7 5.3 45 105-149 2-46 (136)
5 PF13002 LDB19: Arrestin_N ter 98.0 4.1E-05 9E-10 69.7 10.5 121 13-148 43-175 (191)
6 PF08737 Rgp1: Rgp1; InterPro 97.3 0.0044 9.5E-08 62.6 13.2 47 104-150 300-346 (415)
7 KOG3865|consensus 97.1 0.00042 9.2E-09 67.4 3.6 32 120-151 207-238 (402)
8 PF02752 Arrestin_C: Arrestin 96.8 0.0026 5.7E-08 52.0 5.7 60 216-278 8-76 (136)
9 PF03643 Vps26: Vacuolar prote 96.4 0.043 9.2E-07 52.8 12.0 113 19-150 95-207 (275)
10 KOG3780|consensus 92.8 1 2.2E-05 44.7 10.8 41 109-149 6-48 (427)
11 KOG2717|consensus 89.0 2.2 4.9E-05 40.7 8.4 127 17-150 84-220 (313)
12 PF01345 DUF11: Domain of unkn 88.7 1.5 3.3E-05 33.3 5.9 44 104-147 22-65 (76)
13 PF01835 A2M_N: MG2 domain; I 88.4 1 2.2E-05 35.7 5.0 26 110-135 2-27 (99)
14 TIGR01451 B_ant_repeat conserv 85.8 1.5 3.2E-05 31.7 4.2 34 114-147 3-36 (53)
15 PF09478 CBM49: Carbohydrate b 85.6 2.1 4.5E-05 33.4 5.2 46 232-282 19-74 (80)
16 PF00339 Arrestin_N: Arrestin 85.0 1 2.2E-05 37.4 3.4 39 111-149 2-42 (149)
17 PF00927 Transglut_C: Transglu 83.9 1.9 4E-05 34.9 4.4 54 228-282 13-69 (107)
18 PF07070 Spo0M: SpoOM protein; 80.5 2 4.4E-05 40.1 3.9 38 15-64 80-118 (218)
19 KOG3118|consensus 80.3 0.82 1.8E-05 47.1 1.3 32 168-199 106-138 (517)
20 PF00207 A2M: Alpha-2-macroglo 77.3 7.4 0.00016 30.7 5.7 39 105-145 53-91 (92)
21 PF00927 Transglut_C: Transglu 72.8 9.2 0.0002 30.8 5.3 38 109-147 2-39 (107)
22 KOG3063|consensus 69.9 9.9 0.00022 36.4 5.5 60 17-87 95-158 (301)
23 PF06159 DUF974: Protein of un 69.0 12 0.00025 35.5 5.9 59 223-282 7-70 (249)
24 PF14796 AP3B1_C: Clathrin-ada 66.5 8 0.00017 34.0 3.9 54 227-286 82-138 (145)
25 PF07070 Spo0M: SpoOM protein; 66.3 15 0.00032 34.4 5.8 47 103-149 8-55 (218)
26 PF10633 NPCBM_assoc: NPCBM-as 64.5 5.9 0.00013 30.2 2.4 29 120-148 2-30 (78)
27 PF07705 CARDB: CARDB; InterP 61.0 26 0.00057 26.8 5.6 41 108-148 4-44 (101)
28 PF00963 Cohesin: Cohesin doma 57.7 28 0.00061 29.2 5.7 38 110-148 1-38 (141)
29 PLN02171 endoglucanase 54.1 17 0.00036 39.2 4.4 22 231-252 554-575 (629)
30 PF07703 A2M_N_2: Alpha-2-macr 53.8 16 0.00035 30.1 3.5 32 107-138 94-125 (136)
31 KOG2625|consensus 52.4 8.4 0.00018 36.7 1.7 27 226-252 11-37 (348)
32 PF05688 DUF824: Salmonella re 49.7 18 0.0004 26.0 2.7 29 120-148 10-38 (47)
33 PF11611 DUF4352: Domain of un 49.5 30 0.00066 27.8 4.4 53 229-282 35-94 (123)
34 PF04744 Monooxygenase_B: Mono 48.9 20 0.00043 36.2 3.8 37 107-143 246-283 (381)
35 COG2373 Large extracellular al 48.9 33 0.00071 40.9 6.0 131 107-252 393-535 (1621)
36 PF09478 CBM49: Carbohydrate b 48.2 50 0.0011 25.5 5.2 39 109-147 2-41 (80)
37 PF07703 A2M_N_2: Alpha-2-macr 46.5 24 0.00052 29.1 3.4 25 110-134 1-25 (136)
38 smart00809 Alpha_adaptinC2 Ada 43.2 71 0.0015 25.1 5.6 41 103-147 2-42 (104)
39 TIGR03079 CH4_NH3mon_ox_B meth 42.0 37 0.0008 34.4 4.4 36 105-140 263-299 (399)
40 PF02014 Reeler: Reeler domain 41.8 43 0.00092 28.1 4.3 35 112-148 23-57 (132)
41 PF05753 TRAP_beta: Translocon 40.5 57 0.0012 29.5 5.1 74 226-304 33-110 (181)
42 PF09624 DUF2393: Protein of u 40.3 68 0.0015 27.4 5.4 29 119-147 58-86 (149)
43 TIGR01451 B_ant_repeat conserv 40.1 56 0.0012 23.5 4.1 29 226-254 8-36 (53)
44 cd08544 Reeler Reeler, the N-t 38.1 56 0.0012 27.3 4.4 37 110-148 21-57 (135)
45 PF01345 DUF11: Domain of unkn 38.1 57 0.0012 24.4 4.1 30 224-253 35-64 (76)
46 PF06030 DUF916: Bacterial pro 36.9 37 0.0008 28.7 3.1 29 120-149 24-52 (121)
47 PF14796 AP3B1_C: Clathrin-ada 36.1 85 0.0018 27.7 5.3 46 103-148 64-110 (145)
48 PF13199 Glyco_hydro_66: Glyco 36.0 53 0.0011 35.0 4.7 25 113-137 1-25 (559)
49 PF09624 DUF2393: Protein of u 35.6 52 0.0011 28.2 3.9 34 216-252 51-84 (149)
50 cd08548 Type_I_cohesin_like Ty 34.8 76 0.0016 27.1 4.7 40 110-149 1-40 (135)
51 PF06159 DUF974: Protein of un 34.0 51 0.0011 31.2 3.9 28 119-146 10-37 (249)
52 COG4326 Spo0M Sporulation cont 34.0 57 0.0012 30.8 4.0 21 18-38 104-124 (270)
53 PF02883 Alpha_adaptinC2: Adap 34.0 1.4E+02 0.003 24.0 6.0 45 100-146 3-47 (115)
54 COG2373 Large extracellular al 32.0 91 0.002 37.4 6.2 41 105-145 495-535 (1621)
55 PF13595 DUF4138: Domain of un 31.8 50 0.0011 31.4 3.4 21 224-244 136-156 (246)
56 PF04425 Bul1_N: Bul1 N termin 31.5 28 0.0006 36.0 1.7 22 12-33 252-273 (438)
57 PF10633 NPCBM_assoc: NPCBM-as 31.2 48 0.001 25.1 2.7 24 229-252 4-27 (78)
58 cd00258 GM2-AP GM2 activator p 30.7 61 0.0013 29.2 3.6 35 17-63 109-145 (162)
59 cd00917 PG-PI_TP The phosphati 30.7 87 0.0019 26.0 4.4 32 16-63 78-109 (122)
60 PF11355 DUF3157: Protein of u 29.9 1.1E+02 0.0024 28.4 5.2 37 110-147 100-136 (199)
61 PF04314 DUF461: Protein of un 28.3 87 0.0019 25.6 3.9 40 100-139 70-109 (110)
62 smart00809 Alpha_adaptinC2 Ada 28.0 2.9E+02 0.0063 21.5 6.8 55 224-282 12-66 (104)
63 PF07919 Gryzun: Gryzun, putat 27.4 2.1E+02 0.0046 29.3 7.4 40 101-140 166-207 (554)
64 PF11355 DUF3157: Protein of u 26.8 1.3E+02 0.0029 27.9 5.1 39 216-258 101-139 (199)
65 PF07705 CARDB: CARDB; InterP 26.8 2E+02 0.0044 21.7 5.6 34 226-262 15-48 (101)
66 cd08546 cohesin_like Cohesin d 25.5 1.6E+02 0.0035 23.8 5.1 37 110-149 3-39 (135)
67 PF11797 DUF3324: Protein of u 24.6 1.5E+02 0.0032 25.4 4.8 66 231-299 43-110 (140)
68 PF02019 WIF: WIF domain; Int 24.0 2.9E+02 0.0063 23.9 6.5 43 104-146 85-128 (132)
69 PF10437 Lip_prot_lig_C: Bacte 23.3 1.8E+02 0.0039 22.4 4.7 27 223-252 9-35 (86)
70 PF04744 Monooxygenase_B: Mono 23.1 1.1E+02 0.0024 31.0 4.2 67 214-282 248-328 (381)
71 TIGR03780 Bac_Flav_CT_N Bacter 23.0 88 0.0019 30.6 3.4 20 224-243 175-194 (285)
72 PF12389 Peptidase_M73: Camely 23.0 1.3E+02 0.0029 27.9 4.5 32 117-148 59-90 (199)
73 PF04442 CtaG_Cox11: Cytochrom 21.8 97 0.0021 27.5 3.2 29 119-147 63-91 (152)
74 KOG2540|consensus 21.1 82 0.0018 30.0 2.7 71 116-194 155-231 (269)
75 PF04425 Bul1_N: Bul1 N termin 20.9 1.3E+02 0.0028 31.2 4.2 44 105-148 131-191 (438)
76 PF01050 MannoseP_isomer: Mann 20.7 4.2E+02 0.0091 23.2 7.0 72 77-148 63-139 (151)
77 PRK05089 cytochrome C oxidase 20.6 2E+02 0.0044 26.4 5.1 28 119-146 90-117 (188)
78 PF04205 FMN_bind: FMN-binding 20.4 1.7E+02 0.0037 21.9 4.0 22 229-252 3-24 (81)
No 1
>KOG3865|consensus
Probab=100.00 E-value=4.6e-62 Score=460.85 Aligned_cols=202 Identities=62% Similarity=0.899 Sum_probs=193.4
Q ss_pred ccccCcHHHHHHcccCCCeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccCccccccceeeEEEeeeeee
Q psy17223 4 VIYKVTFGQERLMKKLGPNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGETAEDKIHKRNSVRLAIRKIM 83 (323)
Q Consensus 4 ~~~~lt~~Q~~l~~~~G~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~~~dk~~k~~tv~l~Ir~l~ 83 (323)
...|||++|++|++|+|.|+|||.|++|+++|+|+.+|+++++.||+|+|.|+||||++.+.+++++|+++|+|+||+++
T Consensus 91 ~~~plT~lQErLlkKLG~nAyPF~f~~pp~~P~SVtLQp~p~D~gKpcGVdyevkaF~~~s~edk~hKr~sVrL~IRKvq 170 (402)
T KOG3865|consen 91 DSRPLTRLQERLLKKLGSNAYPFTFEFPPNLPCSVTLQPGPEDTGKPCGVDYEVKAFVADSEEDKIHKRNSVRLVIRKVQ 170 (402)
T ss_pred CCCcccHHHHHHHHHhCCCCCceEEeCCCCCCceEEeccCCccCCCcccceEEEEEEecCCcccccccccceeeeeeeee
Confidence 34679999999999999999999999999999999999999999999999999999999998999999999999999999
Q ss_pred cCCCCCCCCCeEEEEEEeEecCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCCccchhhhhhhccCC
Q psy17223 84 YAPSKQGEQPSVEVSKEFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGAEDDQDLKDELADSD 163 (323)
Q Consensus 84 ~~P~~~~~~~~~e~~k~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~ed~~d~~~~~~~~d 163 (323)
++|..+.++|+.++.+.|+|++|+++|+|+|||
T Consensus 171 yAP~~~GpqP~~~v~k~FlmS~~~lhLevsLDk----------------------------------------------- 203 (402)
T KOG3865|consen 171 YAPLEPGPQPSAEVSKQFLMSDGPLHLEVSLDK----------------------------------------------- 203 (402)
T ss_pred ecCCCCCCCchhHhhHhhccCCCceEEEEEecc-----------------------------------------------
Confidence 999998899999999999999999999999886
Q ss_pred CCCCccCCCccccccccccceeeecccccCCCCCCCCCCcchhhhhhccCcchhhccceeeEEeecceEEEEEEEecCCc
Q psy17223 164 IDGMEEDDLPNIKAWGKNKRMYYNTDYVDDDHGGIQSGTGFVWFEWLKKGSKEKSKKKYLFLYYHGESIAVNVHVANNSN 243 (323)
Q Consensus 164 i~~~~e~~lp~~~awg~~~~~~y~td~~~~~~~~~~~~~~~~~~e~~~~~~~~~s~~k~~~~yyhge~i~v~v~v~N~s~ 243 (323)
++|||||+|.|||+|+||||
T Consensus 204 ------------------------------------------------------------EiYyHGE~isvnV~V~NNsn 223 (402)
T KOG3865|consen 204 ------------------------------------------------------------EIYYHGEPISVNVHVTNNSN 223 (402)
T ss_pred ------------------------------------------------------------hheecCCceeEEEEEecCCc
Confidence 57999999999999999999
Q ss_pred ceEeeEEee-----eEEEeeCCeEeeeeccccCC--CccCCCCccc--------------ccceeecccccccccCCccc
Q psy17223 244 RTVKKIKVS-----DICLFSTAQYKCTVAETESD--CPIAPVSMFD--------------TEDLAMLRHGFKRMFGHAFS 302 (323)
Q Consensus 244 k~vkkikv~-----dv~l~s~~~y~~~Va~~e~~--~~i~p~~t~~--------------~~~~~~~~~~~~~~~~~a~~ 302 (323)
|||||||++ ||||||++||+|+||.+|+. |||+||+||+ |+||||||+|||+|+|||||
T Consensus 224 KtVKkIK~~V~Q~adi~Lfs~aqy~~~VA~~E~~eGc~v~Pgstl~Kvf~l~PllanN~dkrGlALDG~lKhEDtnLASS 303 (402)
T KOG3865|consen 224 KTVKKIKISVRQVADICLFSTAQYKKPVAMEETDEGCPVAPGSTLSKVFTLTPLLANNKDKRGLALDGKLKHEDTNLASS 303 (402)
T ss_pred ceeeeeEEEeEeeceEEEEecccccceeeeeecccCCccCCCCeeeeeEEechhhhcCcccccccccccccccccccchh
Confidence 999999998 99999999999999999974 5999999999 99999999999999999999
Q ss_pred ccccCccccc
Q psy17223 303 TSLAMPRDEL 312 (323)
Q Consensus 303 t~~~~~~~~~ 312 (323)
|||+|+.||-
T Consensus 304 Tii~~~~~re 313 (402)
T KOG3865|consen 304 TIIREGADRE 313 (402)
T ss_pred heecCCCCcc
Confidence 9999999874
No 2
>KOG3780|consensus
Probab=99.75 E-value=9.7e-18 Score=165.35 Aligned_cols=125 Identities=29% Similarity=0.380 Sum_probs=98.1
Q ss_pred cccCCCeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccCccccccceeeEEEeeee---eecCCCCCCCC
Q psy17223 16 MKKLGPNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGETAEDKIHKRNSVRLAIRK---IMYAPSKQGEQ 92 (323)
Q Consensus 16 ~~~~G~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~~~dk~~k~~tv~l~Ir~---l~~~P~~~~~~ 92 (323)
+.++|.|+|||.|+||.+||+||+ |.+| +|||.|+|.|+|+ |+..+.....+.|.. ++..|....+.
T Consensus 101 ~l~~G~~~~pF~~~LP~~~P~Sfe-----g~~G---~irY~vk~~idr~--~~~~~~~~~~~~V~~~~~ln~~p~~~~~~ 170 (427)
T KOG3780|consen 101 VLPPGNYEFPFSFTLPLNLPPSFE-----GKFG---HVRYFVKAEIDRP--WKLNKKNRKPFTVIETVDLNSSPSLLEPI 170 (427)
T ss_pred ecCCCceEEeEeccCCCCCCCcee-----eCCc---eEEEEEEEEEecC--CCCCccceeeEEEecccccccCccccCcc
Confidence 578999999999999999999998 5566 9999999999995 455555444444322 22234333222
Q ss_pred C--eEEEEEEeEecCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCCc
Q psy17223 93 P--SVEVSKEFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGAE 150 (323)
Q Consensus 93 ~--~~e~~k~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~e 150 (323)
. ..+...++||..|+|.+++.|++.+|+|||.|.+++.|+|.|++.++.+++++.|..
T Consensus 171 ~~~~~k~~~~~~~~~g~v~~~~~ip~~~~~~ge~i~~~~~i~n~ss~~~~~~~~~l~q~~ 230 (427)
T KOG3780|consen 171 ISKASKKLGCVCFSSGPVSLELTIPKTGYVPGETIPVTLEIENKSSRTIKKVKAKLIQKI 230 (427)
T ss_pred hhhhhheeeEEEecCCcEEEEEEcccccCcCCccEEEEEEEecCCCCcceeeEEEEEEEE
Confidence 1 123334478999999999999999999999999999999999999999999998854
No 3
>PF00339 Arrestin_N: Arrestin (or S-antigen), N-terminal domain; InterPro: IPR011021 G protein-coupled receptors are a large family of signalling molecules that respond to a wide variety of extracellular stimuli. The receptors relay the information encoded by the ligand through the activation of heterotrimeric G proteins and intracellular effector molecules. To ensure the appropriate regulation of the signalling cascade, it is vital to properly inactivate the receptor. This inactivation is achieved, in part, by the binding of a soluble protein, arrestin, which uncouples the receptor from the downstream G protein after the receptors are phosphorylated by G protein-coupled receptor kinases. In addition to the inactivation of G protein-coupled receptors, arrestins have also been implicated in the endocytosis of receptors and cross talk with other signalling pathways. Arrestin (retinal S-antigen) is a major protein of the retinal rod outer segments. It interacts with photo-activated phosphorylated rhodopsin, inhibiting or 'arresting' its ability to interact with transducin []. The protein binds calcium, and shows similarity in its C terminus to alpha-transducin and other purine nucleotide-binding proteins. In mammals, arrestin is associated with autoimmune uveitis. Arrestins comprise a family of closely-related proteins that includes beta-arrestin-1 and -2, which regulate the function of beta-adrenergic receptors by binding to their phosphorylated forms, impairing their capacity to activate G(S) proteins; Cone photoreceptors C-arrestin (arrestin-X) [], which could bind to phosphorylated red/green opsins; and Drosophila phosrestins I and II, which undergo light-induced phosphorylation, and probably play a role in photoreceptor transduction [, , ]. The crystal structure of bovine retinal arrestin comprises two domains of antiparallel beta-sheets connected through a hinge region and one short alpha-helix on the back of the amino-terminal fold []. The binding region for phosphorylated light-activated rhodopsin is located at the N-terminal domain, as indicated by the docking of the photoreceptor to the three-dimensional structure of arrestin. The N-terminal domain consists of an immunoglobulin-like beta-sandwich structure. This entry represents proteins with immunoglobulin-like domains that are similar to those found in arrestin.; PDB: 1SUJ_A 3UGX_A 1CF1_B 1AYR_A 3UGU_A 3P2D_B 1ZSH_A 2WTR_B 3GC3_A 1G4R_A ....
Probab=98.96 E-value=1.1e-09 Score=91.90 Aligned_cols=42 Identities=33% Similarity=0.572 Sum_probs=28.2
Q ss_pred HcccCCCeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccC
Q psy17223 15 LMKKLGPNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGET 64 (323)
Q Consensus 15 l~~~~G~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~ 64 (323)
...++|.|.|||+|+||.+||+||+.. .| +|+|.|+|.|+++
T Consensus 88 ~~l~~G~~~fpF~f~LP~~lP~S~~~~-----~g---~I~Y~l~a~l~~~ 129 (149)
T PF00339_consen 88 NILPPGEYEFPFEFQLPSNLPSSFEGS-----HG---SIRYKLKATLDRP 129 (149)
T ss_dssp ----C-TTEEEEEE---TTS--SEEEE------S---EEEEEEEEEESST
T ss_pred ecccCCCEEEEEEEECCCCCCceEecc-----Cc---CEEEEEEEEEECC
Confidence 344599999999999999999999843 45 9999999999886
No 4
>PF02752 Arrestin_C: Arrestin (or S-antigen), C-terminal domain; InterPro: IPR011022 G protein-coupled receptors are a large family of signalling molecules that respond to a wide variety of extracellular stimuli. The receptors relay the information encoded by the ligand through the activation of heterotrimeric G proteins and intracellular effector molecules. To ensure the appropriate regulation of the signalling cascade, it is vital to properly inactivate the receptor. This inactivation is achieved, in part, by the binding of a soluble protein, arrestin, which uncouples the receptor from the downstream G protein after the receptors are phosphorylated by G protein-coupled receptor kinases. In addition to the inactivation of G protein-coupled receptors, arrestins have also been implicated in the endocytosis of receptors and cross talk with other signalling pathways. Arrestin (retinal S-antigen) is a major protein of the retinal rod outer segments. It interacts with photo-activated phosphorylated rhodopsin, inhibiting or 'arresting' its ability to interact with transducin []. The protein binds calcium, and shows similarity in its C terminus to alpha-transducin and other purine nucleotide-binding proteins. In mammals, arrestin is associated with autoimmune uveitis. Arrestins comprise a family of closely-related proteins that includes beta-arrestin-1 and -2, which regulate the function of beta-adrenergic receptors by binding to their phosphorylated forms, impairing their capacity to activate G(S) proteins; Cone photoreceptors C-arrestin (arrestin-X) [], which could bind to phosphorylated red/green opsins; and Drosophila phosrestins I and II, which undergo light-induced phosphorylation, and probably play a role in photoreceptor transduction [, , ]. The crystal structure of bovine retinal arrestin comprises two domains of antiparallel beta-sheets connected through a hinge region and one short alpha-helix on the back of the amino-terminal fold []. The binding region for phosphorylated light-activated rhodopsin is located at the N-terminal domain, as indicated by the docking of the photoreceptor to the three-dimensional structure of arrestin. The C-terminal domain consists of an immunoglobulin-like beta-sandwich structure. This entry represents proteins with immunoglobulin-like domains that are similar to those found in arrestin.; PDB: 1SUJ_A 3UGX_A 1CF1_B 1AYR_A 3UGU_A 3P2D_B 1ZSH_A 2WTR_B 3GC3_A 1G4R_A ....
Probab=98.56 E-value=9e-08 Score=78.73 Aligned_cols=45 Identities=40% Similarity=0.573 Sum_probs=39.4
Q ss_pred CCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCC
Q psy17223 105 PNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGA 149 (323)
Q Consensus 105 sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~ 149 (323)
+|++++++++||++|.|||+|+|++.|+|.|++.|++|++.|.+.
T Consensus 2 ~g~i~~~~~i~~~~~~~Ge~i~v~v~i~n~s~~~i~~I~v~L~~~ 46 (136)
T PF02752_consen 2 SGKISLSISIPRTAYVPGETIPVNVEIDNQSKKKIKKIKVSLVER 46 (136)
T ss_dssp TEEEEEEEEES-SEEETT--EEEEEEEEE-SSSEEEEEEEEEEEE
T ss_pred CCEEEEEEEECCCEECCCCEEEEEEEEEECCCCEEEEEEEEEEEE
Confidence 699999999999999999999999999999999999999999874
No 5
>PF13002 LDB19: Arrestin_N terminal like; InterPro: IPR024391 This entry represents a predicted Ig-like beta sandwich domain found towards the N terminus of protein LDB19 []. It is also found in other sequences and is related to the arrestin N-terminal fold [].
Probab=98.03 E-value=4.1e-05 Score=69.67 Aligned_cols=121 Identities=22% Similarity=0.337 Sum_probs=77.7
Q ss_pred HHHcccCCCeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccC--c-ccccccee----eEEEee-eeeec
Q psy17223 13 ERLMKKLGPNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGET--A-EDKIHKRN----SVRLAI-RKIMY 84 (323)
Q Consensus 13 ~~l~~~~G~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~--~-~dk~~k~~----tv~l~I-r~l~~ 84 (323)
..+..+.|.|.|||++-||-+||+|..+- .+-. ..|.|+++|..... . .....+.. ...+.| |-|.
T Consensus 43 ~~t~l~~G~h~fPFS~LiPG~LPaS~~lg--s~~l---~~I~Yel~A~a~~~~~~~~~~~~~~~~~~~~~pl~V~Rsi~- 116 (191)
T PF13002_consen 43 HPTTLTKGSHAFPFSYLIPGHLPASMDLG--STPL---VSIKYELKAEATYKDPRRGSSSSKPRVLKLKRPLPVKRSIL- 116 (191)
T ss_pred CccccCCCcccCCeeEECCCCCccccccC--CCCc---EEEEEEEEEEEEEccCccccCCCcceeEEEeeeEEEEEecC-
Confidence 34557899999999999999999999742 1223 48999999987541 0 00011111 112222 2221
Q ss_pred CCCCCCCCCeEEEEEEeEecCCceEEEEEeCccceecCCeEEEEEEEeccCcce----eeEEEeecCC
Q psy17223 85 APSKQGEQPSVEVSKEFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRT----VKKIKVSDSG 148 (323)
Q Consensus 85 ~P~~~~~~~~~e~~k~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~----Vk~Ikv~L~q 148 (323)
| .+ .....-.|..=.|...|.||. +.+|--+.+|++.++|-++.+ ++++.=++.+
T Consensus 117 -~-----gp--d~~S~RvFPPT~l~a~a~lP~-VI~P~gtfpvel~LdGv~~~~~rWRlrKltWRiEE 175 (191)
T PF13002_consen 117 -P-----GP--DKNSLRVFPPTNLTASAVLPN-VIHPKGTFPVELRLDGVVSKDRRWRLRKLTWRIEE 175 (191)
T ss_pred -C-----CC--CcccEEecCCCCcEEEEEcCC-eeCCCCcccEEEEEecccCCCCEEEEEeeeEEEee
Confidence 1 11 112334678888999999995 666888899999999987664 6666666544
No 6
>PF08737 Rgp1: Rgp1; InterPro: IPR014848 Rgp1 forms heterodimer with Ric1 (IPR009771 from INTERPRO) which associates with Golgi membranes and functions as a guanyl-nucleotide exchange factor [].
Probab=97.27 E-value=0.0044 Score=62.63 Aligned_cols=47 Identities=21% Similarity=0.212 Sum_probs=41.0
Q ss_pred cCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCCc
Q psy17223 104 SPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGAE 150 (323)
Q Consensus 104 ~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~e 150 (323)
.+|..-..++|.|..|.-||+|...+++++.....+..+.+.|.-.|
T Consensus 300 ~n~~~va~~~LsK~~yrlGE~I~g~idf~~~~~~~c~~v~~~LEs~E 346 (415)
T PF08737_consen 300 RNGQRVARLSLSKPAYRLGEDIVGTIDFNDASTIPCYQVSASLESEE 346 (415)
T ss_pred ECCeEEEEEEecCCCcccCCeEEEEEEcCCCCcceeEEEEEEEEEEE
Confidence 37888889999999999999999999999998677888888886544
No 7
>KOG3865|consensus
Probab=97.09 E-value=0.00042 Score=67.45 Aligned_cols=32 Identities=69% Similarity=0.957 Sum_probs=29.6
Q ss_pred ecCCeEEEEEEEeccCcceeeEEEeecCCCcc
Q psy17223 120 YHGESIAVNVHVANNSNRTVKKIKVSDSGAED 151 (323)
Q Consensus 120 ~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~ed 151 (323)
+|||+|.|+|.|.|+|+|.|++||+.+.|.-|
T Consensus 207 yHGE~isvnV~V~NNsnKtVKkIK~~V~Q~ad 238 (402)
T KOG3865|consen 207 YHGEPISVNVHVTNNSNKTVKKIKISVRQVAD 238 (402)
T ss_pred ecCCceeEEEEEecCCcceeeeeEEEeEeece
Confidence 58999999999999999999999999998543
No 8
>PF02752 Arrestin_C: Arrestin (or S-antigen), C-terminal domain; InterPro: IPR011022 G protein-coupled receptors are a large family of signalling molecules that respond to a wide variety of extracellular stimuli. The receptors relay the information encoded by the ligand through the activation of heterotrimeric G proteins and intracellular effector molecules. To ensure the appropriate regulation of the signalling cascade, it is vital to properly inactivate the receptor. This inactivation is achieved, in part, by the binding of a soluble protein, arrestin, which uncouples the receptor from the downstream G protein after the receptors are phosphorylated by G protein-coupled receptor kinases. In addition to the inactivation of G protein-coupled receptors, arrestins have also been implicated in the endocytosis of receptors and cross talk with other signalling pathways. Arrestin (retinal S-antigen) is a major protein of the retinal rod outer segments. It interacts with photo-activated phosphorylated rhodopsin, inhibiting or 'arresting' its ability to interact with transducin []. The protein binds calcium, and shows similarity in its C terminus to alpha-transducin and other purine nucleotide-binding proteins. In mammals, arrestin is associated with autoimmune uveitis. Arrestins comprise a family of closely-related proteins that includes beta-arrestin-1 and -2, which regulate the function of beta-adrenergic receptors by binding to their phosphorylated forms, impairing their capacity to activate G(S) proteins; Cone photoreceptors C-arrestin (arrestin-X) [], which could bind to phosphorylated red/green opsins; and Drosophila phosrestins I and II, which undergo light-induced phosphorylation, and probably play a role in photoreceptor transduction [, , ]. The crystal structure of bovine retinal arrestin comprises two domains of antiparallel beta-sheets connected through a hinge region and one short alpha-helix on the back of the amino-terminal fold []. The binding region for phosphorylated light-activated rhodopsin is located at the N-terminal domain, as indicated by the docking of the photoreceptor to the three-dimensional structure of arrestin. The C-terminal domain consists of an immunoglobulin-like beta-sandwich structure. This entry represents proteins with immunoglobulin-like domains that are similar to those found in arrestin.; PDB: 1SUJ_A 3UGX_A 1CF1_B 1AYR_A 3UGU_A 3P2D_B 1ZSH_A 2WTR_B 3GC3_A 1G4R_A ....
Probab=96.82 E-value=0.0026 Score=52.02 Aligned_cols=60 Identities=37% Similarity=0.504 Sum_probs=40.3
Q ss_pred hhhccceeeEEeecceEEEEEEEecCCcceEeeEEee---eEEEeeC------CeEeeeeccccCCCccCCC
Q psy17223 216 EKSKKKYLFLYYHGESIAVNVHVANNSNRTVKKIKVS---DICLFST------AQYKCTVAETESDCPIAPV 278 (323)
Q Consensus 216 ~~s~~k~~~~yyhge~i~v~v~v~N~s~k~vkkikv~---dv~l~s~------~~y~~~Va~~e~~~~i~p~ 278 (323)
..++.| ..|..||.|.|++.|+|.|++.|++|++. .+..+.. .++.+.|+. ...+.+.++
T Consensus 8 ~~~i~~--~~~~~Ge~i~v~v~i~n~s~~~i~~I~v~L~~~~~~~~~~~~~~~~~~~~~v~~-~~~~~~~~~ 76 (136)
T PF02752_consen 8 SISIPR--TAYVPGETIPVNVEIDNQSKKKIKKIKVSLVERITYKAKGGKDESKSEKRVVAK-SKNCGVDPG 76 (136)
T ss_dssp EEEES---SEEETT--EEEEEEEEE-SSSEEEEEEEEEEEEEEE-SS----S-EEEEEEEEE-EECCEB-B-
T ss_pred EEEECC--CEECCCCEEEEEEEEEECCCCEEEEEEEEEEEEEEEEEeeccccceEEEEEEEE-EecCCccCC
Confidence 566778 88999999999999999999999999998 5555544 346677777 334444333
No 9
>PF03643 Vps26: Vacuolar protein sorting-associated protein 26 ; InterPro: IPR005377 The movement of lipid and protein components between intracellular organelles requires the regulated interactions of many molecules. Vacuolar protein sorting-associated protein (Vps)5 is a yeast protein that is a subunit of a large multimeric complex, termed the retromer complex, involved in retrograde transport of proteins from endosomes to the trans-Golgi network. Sorting nexin (SNX) 1 and SNX2 are its mammalian orthologs []. To carry out its biological functions, Vps5 forms the retromer complex with at least four other proteins: Vps17, Vps26, Vps29, and Vps35 []. This family of Vps26-proteins also contains Down syndrome critical region 3/A.; GO: 0007034 vacuolar transport, 0030904 retromer complex; PDB: 3LHA_A 3LH9_A 2R51_A 3LH8_B 2FAU_A.
Probab=96.45 E-value=0.043 Score=52.82 Aligned_cols=113 Identities=19% Similarity=0.234 Sum_probs=64.7
Q ss_pred CCCeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccCccccccceeeEEEeeeeeecCCCCCCCCCeEEEE
Q psy17223 19 LGPNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGETAEDKIHKRNSVRLAIRKIMYAPSKQGEQPSVEVS 98 (323)
Q Consensus 19 ~G~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~~~dk~~k~~tv~l~Ir~l~~~P~~~~~~~~~e~~ 98 (323)
.|.. |||.|.+-.. | ++ .|+|....|+|.|+|.+.|+. ..-..++.|.|..+...|... ..+
T Consensus 95 ~~~t-~pFeF~~~~k-~--yE-----TY~G~~v~i~Y~lrv~v~R~~---~~i~k~~ef~V~~~~~~p~~~-----~~i- 156 (275)
T PF03643_consen 95 EGKT-FPFEFPLVEK-P--YE-----TYHGVNVNIRYFLRVTVKRSY---KDISKEQEFWVQNFSITPESN-----QPI- 156 (275)
T ss_dssp S-EE-EEEEE-SB------S-------EE-SSEEEEEEEEEEE--SS---S-EEEEEEEEEE-EB-------------E-
T ss_pred CCcE-EeeEeCCCCC-C--Cc-----cEeeeEEEEEEEEEEEEEccC---CCcceEEEEEEEeccCCCCCC-----CCc-
Confidence 4445 9999987432 1 32 467888899999999998864 222345566776554444332 111
Q ss_pred EEeEecCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCCc
Q psy17223 99 KEFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGAE 150 (323)
Q Consensus 99 k~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~e 150 (323)
+.=.--.+-++++..++|..|.=.+.|.=.+.+.-.. ..|++|.++|.+-|
T Consensus 157 k~evgie~~lhief~~~k~~~~l~d~i~G~i~f~lv~-~kIk~~elqLiR~E 207 (275)
T PF03643_consen 157 KMEVGIEDCLHIEFEYDKSKYHLKDVITGKIYFLLVR-IKIKSMELQLIRVE 207 (275)
T ss_dssp EEEECETTTEEEEEEES-SEEETT-EEEEEEEEEEES-S-EEEEEEEEEEEE
T ss_pred ccccCCCccEEEEEEEcccceECCCCEEEEEEEEEEe-ecceEEEEEEEEEE
Confidence 1111135678999999999999999987776664333 67999999998844
No 10
>KOG3780|consensus
Probab=92.84 E-value=1 Score=44.68 Aligned_cols=41 Identities=27% Similarity=0.482 Sum_probs=35.6
Q ss_pred EEEEEeCcc--ceecCCeEEEEEEEeccCcceeeEEEeecCCC
Q psy17223 109 HLEASLDKE--LYYHGESIAVNVHVANNSNRTVKKIKVSDSGA 149 (323)
Q Consensus 109 ~L~a~LdK~--~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~ 149 (323)
.+.+.+|+. +|.+||+|.=++.+.+.....++.|++++.+.
T Consensus 6 ~~~i~~d~~~~iy~~G~~vsG~v~l~~~~~~~~~~i~l~~~G~ 48 (427)
T KOG3780|consen 6 SFEIVLDNPEAIYFPGEPVSGSVVLSTKEPIKVRAIKLQLKGR 48 (427)
T ss_pred eEEEEeCCCccccCCCCeEEEEEEEEeCCccceeEEEEEEEEe
Confidence 345666666 59999999999999999999999999999874
No 11
>KOG2717|consensus
Probab=89.03 E-value=2.2 Score=40.67 Aligned_cols=127 Identities=13% Similarity=0.134 Sum_probs=74.4
Q ss_pred ccCCCeeeeeEeeCCC-CCCCcEEeccCCCCCCCceeeEEEEEEEEccCccccccceeeEEEeeee----eecCCCCC--
Q psy17223 17 KKLGPNAFPFFFELPP-SCPASVTLQPAPGDTGKPCGVDYELKAFVGETAEDKIHKRNSVRLAIRK----IMYAPSKQ-- 89 (323)
Q Consensus 17 ~~~G~h~FPFsFqLP~-~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~~~dk~~k~~tv~l~Ir~----l~~~P~~~-- 89 (323)
.|+|.-+|||.|.|-. +=|--|. ..++|-...|.|.+++.+.|.--.|+. ..++.|.|.. +...|...
T Consensus 84 ~p~G~tEipFelpL~~kge~~~lY----ETyHGvfiNiqY~LtcdikR~~L~K~l-tkt~eFiv~s~pv~l~e~~p~iV~ 158 (313)
T KOG2717|consen 84 IPPGTTEIPFELPLREKGEGEKLY----ETYHGVFINIQYLLTCDIKRGYLHKPL-TKTMEFIVESGPVDLPERPPEIVI 158 (313)
T ss_pred CCCCceeeeeeeeeccCCCccEee----eeecceEEEEEEEEEEecccchhcCch-hhhheeeeccCCcccccCCCcceE
Confidence 4789999999988764 2232222 147888889999999999886333332 2345555531 11001000
Q ss_pred ---CCCCeEEEEEEeEecCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCCc
Q psy17223 90 ---GEQPSVEVSKEFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGAE 150 (323)
Q Consensus 90 ---~~~~~~e~~k~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~e 150 (323)
.|-..+...+. -..-|..-+.-.||+.-++--+++.=.+.|+ .|...|++|.++|.+-|
T Consensus 159 F~itpdtlq~~~ke-r~~~p~FlvtG~Ld~t~c~~t~PltGeltVe-~seaaI~Sie~qLvRVE 220 (313)
T KOG2717|consen 159 FYITPDTLQHPLKE-RIKTPGFLVTGKLDATQCSLTDPLTGELTVE-ASEAAITSIEIQLVRVE 220 (313)
T ss_pred EEEChHHhhccchh-hccCCceEEEeeecceeeEecCCccceEEEE-eeccceeEEEEEEEEEE
Confidence 00000000000 1233455566778888888777777666665 45688999999998744
No 12
>PF01345 DUF11: Domain of unknown function DUF11; InterPro: IPR001434 This group of sequences is represented by a conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis (Streptococcus faecalis), and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydia trachomatis outer membrane proteins. In C. trachomatis, three cysteine-rich proteins (also believed to be lipoproteins), MOMP, OMP6 and OMP3, make up the extracellular matrix of the outer membrane []. They are involved in the essential structural integrity of both the elementary body (EB) and recticulate body (RB) phase. They are thought to be involved in porin formation and, as these bacteria lack the peptidoglycan layer common to most Gram-negative microbes, such proteins are highly important in the pathogenicity of the organism.; GO: 0005727 extrachromosomal circular DNA
Probab=88.68 E-value=1.5 Score=33.26 Aligned_cols=44 Identities=14% Similarity=0.311 Sum_probs=39.4
Q ss_pred cCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223 104 SPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDS 147 (323)
Q Consensus 104 ~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~ 147 (323)
....+.+.-+.++....+||.|..++.+.|..+.....+.+.+.
T Consensus 22 ~~~~~~~~k~~~~~~~~~Gd~v~ytitvtN~G~~~a~nv~v~D~ 65 (76)
T PF01345_consen 22 AIPDLSITKTVNPSTANPGDTVTYTITVTNTGPAPATNVVVTDT 65 (76)
T ss_pred CCCCEEEEEecCCCcccCCCEEEEEEEEEECCCCeeEeEEEEEc
Confidence 45678888889999999999999999999999999999988874
No 13
>PF01835 A2M_N: MG2 domain; InterPro: IPR002890 The proteinase-binding alpha-macroglobulins (A2M) [] are large glycoproteins found in the plasma of vertebrates, in the hemolymph of some invertebrates and in reptilian and avian egg white. A2M-like proteins are able to inhibit all four classes of proteinases by a 'trapping' mechanism. They have a peptide stretch, called the 'bait region', which contains specific cleavage sites for different proteinases. When a proteinase cleaves the bait region, a conformational change is induced in the protein, thus trapping the proteinase. The entrapped enzyme remains active against low molecular weight substrates, whilst its activity toward larger substrates is greatly reduced, due to steric hindrance. Following cleavage in the bait region, a thiol ester bond, formed between the side chains of a cysteine and a glutamine, is cleaved and mediates the covalent binding of the A2M-like protein to the proteinase. This family includes the N-terminal region of the alpha-2-macroglobulin family. The inhibitor domains belong to MEROPS inhibitor family I39.; GO: 0004866 endopeptidase inhibitor activity; PDB: 2B39_B 3KLS_B 3PRX_C 3KM9_B 3PVM_C 3CU7_A 4E0S_A 4A5W_A 4ACQ_C 2P9R_B ....
Probab=88.38 E-value=1 Score=35.67 Aligned_cols=26 Identities=19% Similarity=0.340 Sum_probs=21.8
Q ss_pred EEEEeCccceecCCeEEEEEEEeccC
Q psy17223 110 LEASLDKELYYHGESIAVNVHVANNS 135 (323)
Q Consensus 110 L~a~LdK~~Y~PGE~I~V~v~IdN~S 135 (323)
+-+..||..|.|||+|.+.+-+.+..
T Consensus 2 ~~i~TDr~iYrPGetV~~~~~~~~~~ 27 (99)
T PF01835_consen 2 IFIQTDRPIYRPGETVHFRAIVRDLD 27 (99)
T ss_dssp EEEEESSSEE-TTSEEEEEEEEEEEC
T ss_pred EEEECCccCcCCCCEEEEEEEEeccc
Confidence 45789999999999999999987776
No 14
>TIGR01451 B_ant_repeat conserved repeat domain. This model represents the conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis, and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydial outer membrane proteins.
Probab=85.79 E-value=1.5 Score=31.75 Aligned_cols=34 Identities=29% Similarity=0.460 Sum_probs=30.5
Q ss_pred eCccceecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223 114 LDKELYYHGESIAVNVHVANNSNRTVKKIKVSDS 147 (323)
Q Consensus 114 LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~ 147 (323)
.++....|||.|..++.|.|.....+..|.+.+.
T Consensus 3 ~d~~~~~~Gd~v~Yti~v~N~g~~~a~~v~v~D~ 36 (53)
T TIGR01451 3 VDKTVATIGDTITYTITVTNNGNVPATNVVVTDI 36 (53)
T ss_pred cCccccCCCCEEEEEEEEEECCCCceEeEEEEEc
Confidence 5778889999999999999999999998888763
No 15
>PF09478 CBM49: Carbohydrate binding domain CBM49; InterPro: IPR019028 A carbohydrate-binding module (CBM) is defined as a contiguous amino acid sequence within a carbohydrate-active enzyme with a discreet fold having carbohydrate-binding activity. A few exceptions are CBMs in cellulosomal scaffolding proteins and rare instances of independent putative CBMs. The requirement of CBMs existing as modules within larger enzymes sets this class of carbohydrate-binding protein apart from other non-catalytic sugar binding proteins such as lectins and sugar transport proteins. CBMs were previously classified as cellulose-binding domains (CBDs) based on the initial discovery of several modules that bound cellulose [, ]. However, additional modules in carbohydrate-active enzymes are continually being found that bind carbohydrates other than cellulose yet otherwise meet the CBM criteria, hence the need to reclassify these polypeptides using more inclusive terminology. Previous classification of cellulose-binding domains were based on amino acid similarity. Groupings of CBDs were called "Types" and numbered with roman numerals (e.g. Type I or Type II CBDs). In keeping with the glycoside hydrolase classification, these groupings are now called families and numbered with Arabic numerals. Families 1 to 13 are the same as Types I to XIII. For a detailed review on the structure and binding modes of CBMs see []. This domain is found at the C-terminal of cellulases and in vitro binding studies have shown it to binds to crystalline cellulose []. ; GO: 0030246 carbohydrate binding, 0005576 extracellular region
Probab=85.57 E-value=2.1 Score=33.36 Aligned_cols=46 Identities=28% Similarity=0.432 Sum_probs=30.4
Q ss_pred EEEEEEEecCCcceEeeEEee-e--------EEEeeCCeEeeeeccccC-CCccCCCCccc
Q psy17223 232 IAVNVHVANNSNRTVKKIKVS-D--------ICLFSTAQYKCTVAETES-DCPIAPVSMFD 282 (323)
Q Consensus 232 i~v~v~v~N~s~k~vkkikv~-d--------v~l~s~~~y~~~Va~~e~-~~~i~p~~t~~ 282 (323)
...+|.|.|+++++|+.+++. | +..-+++.|. +-. ..+|.||++++
T Consensus 19 ~qy~v~I~N~~~~~I~~~~i~~~~l~~~iW~l~~~~~~~y~-----lPs~~~~i~pg~s~~ 74 (80)
T PF09478_consen 19 TQYDVTITNNGSKPIKSLKISIDNLYGSIWGLDKVSGNTYT-----LPSYQPTIKPGQSFT 74 (80)
T ss_pred EEEEEEEEECCCCeEEEEEEEECccchhheeEEeccCCEEE-----CCccccccCCCCEEE
Confidence 457889999999999999997 4 2222233332 111 23889998864
No 16
>PF00339 Arrestin_N: Arrestin (or S-antigen), N-terminal domain; InterPro: IPR011021 G protein-coupled receptors are a large family of signalling molecules that respond to a wide variety of extracellular stimuli. The receptors relay the information encoded by the ligand through the activation of heterotrimeric G proteins and intracellular effector molecules. To ensure the appropriate regulation of the signalling cascade, it is vital to properly inactivate the receptor. This inactivation is achieved, in part, by the binding of a soluble protein, arrestin, which uncouples the receptor from the downstream G protein after the receptors are phosphorylated by G protein-coupled receptor kinases. In addition to the inactivation of G protein-coupled receptors, arrestins have also been implicated in the endocytosis of receptors and cross talk with other signalling pathways. Arrestin (retinal S-antigen) is a major protein of the retinal rod outer segments. It interacts with photo-activated phosphorylated rhodopsin, inhibiting or 'arresting' its ability to interact with transducin []. The protein binds calcium, and shows similarity in its C terminus to alpha-transducin and other purine nucleotide-binding proteins. In mammals, arrestin is associated with autoimmune uveitis. Arrestins comprise a family of closely-related proteins that includes beta-arrestin-1 and -2, which regulate the function of beta-adrenergic receptors by binding to their phosphorylated forms, impairing their capacity to activate G(S) proteins; Cone photoreceptors C-arrestin (arrestin-X) [], which could bind to phosphorylated red/green opsins; and Drosophila phosrestins I and II, which undergo light-induced phosphorylation, and probably play a role in photoreceptor transduction [, , ]. The crystal structure of bovine retinal arrestin comprises two domains of antiparallel beta-sheets connected through a hinge region and one short alpha-helix on the back of the amino-terminal fold []. The binding region for phosphorylated light-activated rhodopsin is located at the N-terminal domain, as indicated by the docking of the photoreceptor to the three-dimensional structure of arrestin. The N-terminal domain consists of an immunoglobulin-like beta-sandwich structure. This entry represents proteins with immunoglobulin-like domains that are similar to those found in arrestin.; PDB: 1SUJ_A 3UGX_A 1CF1_B 1AYR_A 3UGU_A 3P2D_B 1ZSH_A 2WTR_B 3GC3_A 1G4R_A ....
Probab=85.01 E-value=1 Score=37.40 Aligned_cols=39 Identities=36% Similarity=0.486 Sum_probs=31.3
Q ss_pred EEEeC--ccceecCCeEEEEEEEeccCcceeeEEEeecCCC
Q psy17223 111 EASLD--KELYYHGESIAVNVHVANNSNRTVKKIKVSDSGA 149 (323)
Q Consensus 111 ~a~Ld--K~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~ 149 (323)
++.|| +..|.|||.|.=+|.+.......++.|++++.+.
T Consensus 2 ~I~ld~~~~~y~~Ge~I~G~V~l~~~~~~~i~~i~v~l~G~ 42 (149)
T PF00339_consen 2 EIELDNPKPVYFPGEVISGKVVLELSKPIKIKSIKVRLKGR 42 (149)
T ss_dssp EEEES-SEEEEESS--EEEEEEECTTT-TTTSEEEEEEEEE
T ss_pred EEEECCCCCEECCCCEEEEEEEEEECCccceeEEEEEEEEE
Confidence 34555 9999999999999999888888999999999874
No 17
>PF00927 Transglut_C: Transglutaminase family, C-terminal ig like domain; InterPro: IPR008958 Synonym(s): Protein-glutamine gamma-glutamyltransferase, Fibrinoligase, TGase Transglutaminases catalyse the post-translational modification of proteins at glutamine residues, with formation of isopeptide bonds. Members of the transglutaminase family usually have three domains: N-terminal (IPR001102 from INTERPRO), middle (IPR013808 from INTERPRO) and C-terminal. The middle domain is usually well conserved, but family members can display major differences in their N- and C-terminal domains, although their overall structure is conserved []. This entry represents the C-terminal domain found in transglutaminases, which consists of an immunoglobulin-like beta-sandwich consisting of seven strands in two sheets with a Greek key topology. The best known transglutaminase is blood coagulation factor XIII, a plasma tetrameric protein composed of two catalytic A subunits and two non-catalytic B subunits. Factor XIII is responsible for cross-linking fibrin chains, thus stabilising the fibrin clot. Protein-glutamine gamma-glutamyltransferases (2.3.2.13 from EC) are calcium-dependent enzymes that catalyse the cross-linking of proteins by promoting the formation of isopeptide bonds between the gamma-carboxyl group of a glutamine in one polypeptide chain and the epsilon-amino group of a lysine in a second polypeptide chain. TGases also catalyse the conjugation of polyamines to proteins [, ].; GO: 0003810 protein-glutamine gamma-glutamyltransferase activity, 0018149 peptide cross-linking; PDB: 2XZZ_A 1GGY_B 1FIE_B 1GGU_B 1GGT_B 1F13_A 1QRK_B 1EVU_A 1EX0_B 1L9N_B ....
Probab=83.86 E-value=1.9 Score=34.88 Aligned_cols=54 Identities=13% Similarity=0.224 Sum_probs=38.8
Q ss_pred ecceEEEEEEEecCCcceEeeEEee---eEEEeeCCeEeeeeccccCCCccCCCCccc
Q psy17223 228 HGESIAVNVHVANNSNRTVKKIKVS---DICLFSTAQYKCTVAETESDCPIAPVSMFD 282 (323)
Q Consensus 228 hge~i~v~v~v~N~s~k~vkkikv~---dv~l~s~~~y~~~Va~~e~~~~i~p~~t~~ 282 (323)
-|+++.|.|.+.|.++..++.|++. ..+.| +|-.+...-.....-.|.||.+..
T Consensus 13 vG~d~~v~v~~~N~~~~~l~~v~~~l~~~~v~y-tG~~~~~~~~~~~~~~l~p~~~~~ 69 (107)
T PF00927_consen 13 VGQDFTVSVSFTNPSSEPLRNVSLNLCAFTVEY-TGLTRDQFKKEKFEVTLKPGETKS 69 (107)
T ss_dssp TTSEEEEEEEEEE-SSS-EECEEEEEEEEEEEC-TTTEEEEEEEEEEEEEE-TTEEEE
T ss_pred CCCCEEEEEEEEeCCcCccccceeEEEEEEEEE-CCcccccEeEEEcceeeCCCCEEE
Confidence 4999999999999999999999986 55556 676554444444455799998765
No 18
>PF07070 Spo0M: SpoOM protein; InterPro: IPR009776 This family consists of several bacterial SpoOM proteins which are thought to control sporulation in Bacillus subtilis.Spo0M exerts certain negative effects on sporulation and its gene expression is controlled by sigmaH [].
Probab=80.49 E-value=2 Score=40.11 Aligned_cols=38 Identities=26% Similarity=0.423 Sum_probs=29.3
Q ss_pred HcccCC-CeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccC
Q psy17223 15 LMKKLG-PNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGET 64 (323)
Q Consensus 15 l~~~~G-~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~ 64 (323)
+..++| .+.+||+|+||.++|-|. | +.+|+|+..++-.
T Consensus 80 f~I~~ge~~~iPF~~~lP~etPiT~---------~---~~~v~l~T~LdI~ 118 (218)
T PF07070_consen 80 FTIEPGEEKEIPFSFPLPWETPITE---------G---GMRVWLRTGLDIA 118 (218)
T ss_pred EEECCCCEEEEeEEEECCCCCCccC---------C---CcEEEEEEEEEeC
Confidence 333444 589999999999999976 2 5778888888764
No 19
>KOG3118|consensus
Probab=80.34 E-value=0.82 Score=47.13 Aligned_cols=32 Identities=28% Similarity=0.579 Sum_probs=26.6
Q ss_pred ccCCCccccccccccceeeecc-cccCCCCCCC
Q psy17223 168 EEDDLPNIKAWGKNKRMYYNTD-YVDDDHGGIQ 199 (323)
Q Consensus 168 ~e~~lp~~~awg~~~~~~y~td-~~~~~~~~~~ 199 (323)
++.++.+..+||.++..||.+| |.++++++-+
T Consensus 106 ~e~e~ddn~~WG~~s~~yyg~dd~dddd~s~e~ 138 (517)
T KOG3118|consen 106 KEEEEDDNSTWGGRSGLYYGGDDVDDDDLSSED 138 (517)
T ss_pred cchhhhcccccccccccccCCccccchhhccch
Confidence 3457889999999999999997 8888887644
No 20
>PF00207 A2M: Alpha-2-macroglobulin family; InterPro: IPR001599 This entry contains serum complement C3 and C4 precursors and alpha-macrogrobulins. The alpha-macroglobulin (aM) family of proteins includes protease inhibitors [], typified by the human tetrameric a2-macroglobulin (a2M); they belong to the MEROPS proteinase inhibitor family I39, clan IL. These protease inhibitors share several defining properties, which include (i) the ability to inhibit proteases from all catalytic classes, (ii) the presence of a 'bait region' and a thiol ester, (iii) a similar protease inhibitory mechanism and (iv) the inactivation of the inhibitory capacity by reaction of the thiol ester with small primary amines. aM protease inhibitors inhibit by steric hindrance []. The mechanism involves protease cleavage of the bait region, a segment of the aM that is particularly susceptible to proteolytic cleavage, which initiates a conformational change such that the aM collapses about the protease. In the resulting aM-protease complex, the active site of the protease is sterically shielded, thus substantially decreasing access to protein substrates. Two additional events occur as a consequence of bait region cleavage, namely (i) the h-cysteinyl-g-glutamyl thiol ester becomes highly reactive and (ii) a major conformational change exposes a conserved COOH-terminal receptor binding domain [] (RBD). RBD exposure allows the aM protease complex to bind to clearance receptors and be removed from circulation []. Tetrameric, dimeric, and, more recently, monomeric aM protease inhibitors have been identified [, ].; GO: 0004866 endopeptidase inhibitor activity; PDB: 3KLS_B 3PRX_C 3KM9_B 3PVM_C 3CU7_A 4E0S_A 4A5W_A 2PN5_A 3FRP_G 3HRZ_B ....
Probab=77.32 E-value=7.4 Score=30.70 Aligned_cols=39 Identities=18% Similarity=0.356 Sum_probs=27.4
Q ss_pred CCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEee
Q psy17223 105 PNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVS 145 (323)
Q Consensus 105 sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~ 145 (323)
.-++.++..||+ ....||.+.|.+.|-|+..+.+. ++|+
T Consensus 53 ~~p~~i~~~lP~-~l~~GD~~~i~v~v~N~~~~~~~-v~V~ 91 (92)
T PF00207_consen 53 FKPFFIQLNLPR-SLRRGDQIQIPVTVFNYTDKDQE-VTVT 91 (92)
T ss_dssp B-SEEEEEE--S-EEETTSEEEEEEEEEE-SSS-EE-EEEE
T ss_pred EeeEEEEcCCCc-EEecCCEEEEEEEEEeCCCCCEE-EEEE
Confidence 448889999997 45799999999999999887764 4443
No 21
>PF00927 Transglut_C: Transglutaminase family, C-terminal ig like domain; InterPro: IPR008958 Synonym(s): Protein-glutamine gamma-glutamyltransferase, Fibrinoligase, TGase Transglutaminases catalyse the post-translational modification of proteins at glutamine residues, with formation of isopeptide bonds. Members of the transglutaminase family usually have three domains: N-terminal (IPR001102 from INTERPRO), middle (IPR013808 from INTERPRO) and C-terminal. The middle domain is usually well conserved, but family members can display major differences in their N- and C-terminal domains, although their overall structure is conserved []. This entry represents the C-terminal domain found in transglutaminases, which consists of an immunoglobulin-like beta-sandwich consisting of seven strands in two sheets with a Greek key topology. The best known transglutaminase is blood coagulation factor XIII, a plasma tetrameric protein composed of two catalytic A subunits and two non-catalytic B subunits. Factor XIII is responsible for cross-linking fibrin chains, thus stabilising the fibrin clot. Protein-glutamine gamma-glutamyltransferases (2.3.2.13 from EC) are calcium-dependent enzymes that catalyse the cross-linking of proteins by promoting the formation of isopeptide bonds between the gamma-carboxyl group of a glutamine in one polypeptide chain and the epsilon-amino group of a lysine in a second polypeptide chain. TGases also catalyse the conjugation of polyamines to proteins [, ].; GO: 0003810 protein-glutamine gamma-glutamyltransferase activity, 0018149 peptide cross-linking; PDB: 2XZZ_A 1GGY_B 1FIE_B 1GGU_B 1GGT_B 1F13_A 1QRK_B 1EVU_A 1EX0_B 1L9N_B ....
Probab=72.78 E-value=9.2 Score=30.80 Aligned_cols=38 Identities=16% Similarity=0.314 Sum_probs=30.8
Q ss_pred EEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223 109 HLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDS 147 (323)
Q Consensus 109 ~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~ 147 (323)
++++.++..+. -|+.+.+.+.+.|.++..++.|++.+.
T Consensus 2 ~~~i~~~~~~~-vG~d~~v~v~~~N~~~~~l~~v~~~l~ 39 (107)
T PF00927_consen 2 EIKIKLPGDPV-VGQDFTVSVSFTNPSSEPLRNVSLNLC 39 (107)
T ss_dssp EEEEEEESEEB-TTSEEEEEEEEEE-SSS-EECEEEEEE
T ss_pred eEEEEECCCcc-CCCCEEEEEEEEeCCcCccccceeEEE
Confidence 56677766665 899999999999999999999998884
No 22
>KOG3063|consensus
Probab=69.90 E-value=9.9 Score=36.43 Aligned_cols=60 Identities=23% Similarity=0.361 Sum_probs=37.1
Q ss_pred ccCCC----eeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEccCccccccceeeEEEeeeeeecCCC
Q psy17223 17 KKLGP----NAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGETAEDKIHKRNSVRLAIRKIMYAPS 87 (323)
Q Consensus 17 ~~~G~----h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r~~~dk~~k~~tv~l~Ir~l~~~P~ 87 (323)
-+||. ..|||-|.= +---|+ .+.|+...+||.++|++.|...|-.. ...+.|.-+...|.
T Consensus 95 a~pGel~~~~~fpFeF~~---vekpyE-----sY~G~NV~lrY~lkvTv~Rr~~di~k---e~d~~V~~~~~~P~ 158 (301)
T KOG3063|consen 95 ARPGELTQSQSFPFEFPH---VEKPYE-----SYIGKNVRLRYFLKVTVSRRLTDIVK---EKDLVVHNLSTYPE 158 (301)
T ss_pred cCCcceeecccCCccccc---cccchh-----hhcCcceEEEEEEEEEEEechhhhhh---hhheeeEecccCCC
Confidence 35665 568887752 222233 57898889999999999887543222 23455655544443
No 23
>PF06159 DUF974: Protein of unknown function (DUF974); InterPro: IPR010378 This is a family of uncharacterised eukaryotic proteins.
Probab=68.98 E-value=12 Score=35.45 Aligned_cols=59 Identities=22% Similarity=0.411 Sum_probs=38.2
Q ss_pred eeEEeecceEEEEEEEecCCcceEeeEEeeeEEEeeCCeE-eeeec-cccC---CCccCCCCccc
Q psy17223 223 LFLYYHGESIAVNVHVANNSNRTVKKIKVSDICLFSTAQY-KCTVA-ETES---DCPIAPVSMFD 282 (323)
Q Consensus 223 ~~~yyhge~i~v~v~v~N~s~k~vkkikv~dv~l~s~~~y-~~~Va-~~e~---~~~i~p~~t~~ 282 (323)
.+--|=||+...-++|+|++++.|+.+.|. |-|-...+- +-... ..+. ...+.||.++.
T Consensus 7 fG~iylGEtF~~~l~~~N~s~~~v~~v~ik-vemqT~s~~~r~~L~~~~~~~~~~~~L~p~~~l~ 70 (249)
T PF06159_consen 7 FGSIYLGETFSCYLSVNNDSNKPVRNVRIK-VEMQTPSQSLRLPLSDNENSDSPVASLAPGESLD 70 (249)
T ss_pred cCCEeecCCEEEEEEeecCCCCceEEeEEE-EEEeCCCCCccccCCCCccccccccccCCCCeEe
Confidence 455677999999999999999999999886 233322220 11111 1111 23688998887
No 24
>PF14796 AP3B1_C: Clathrin-adaptor complex-3 beta-1 subunit C-terminal
Probab=66.47 E-value=8 Score=34.04 Aligned_cols=54 Identities=15% Similarity=0.289 Sum_probs=36.3
Q ss_pred eecceEEEEEEEecCCcceEeeEEeeeEEEeeCCeEeeeeccccCC--CccCCCCccc-ccce
Q psy17223 227 YHGESIAVNVHVANNSNRTVKKIKVSDICLFSTAQYKCTVAETESD--CPIAPVSMFD-TEDL 286 (323)
Q Consensus 227 yhge~i~v~v~v~N~s~k~vkkikv~dv~l~s~~~y~~~Va~~e~~--~~i~p~~t~~-~~~~ 286 (323)
|+.--+.|.+.++|+|...+++|+|. +-+..+-+..-|+. +.+.||++.+ ..||
T Consensus 82 ~s~~mvsIql~ftN~s~~~i~~I~i~------~k~l~~g~~i~~F~~I~~L~pg~s~t~~lgI 138 (145)
T PF14796_consen 82 YSPSMVSIQLTFTNNSDEPIKNIHIG------EKKLPAGMRIHEFPEIESLEPGASVTVSLGI 138 (145)
T ss_pred CCCCcEEEEEEEEecCCCeecceEEC------CCCCCCCcEeeccCcccccCCCCeEEEEEEE
Confidence 45667899999999999999999997 22222222222332 2688888766 4444
No 25
>PF07070 Spo0M: SpoOM protein; InterPro: IPR009776 This family consists of several bacterial SpoOM proteins which are thought to control sporulation in Bacillus subtilis.Spo0M exerts certain negative effects on sporulation and its gene expression is controlled by sigmaH [].
Probab=66.27 E-value=15 Score=34.44 Aligned_cols=47 Identities=17% Similarity=0.288 Sum_probs=42.1
Q ss_pred ecCCceEEEEEeCccceecCCeEEEEEEEeccC-cceeeEEEeecCCC
Q psy17223 103 MSPNKLHLEASLDKELYYHGESIAVNVHVANNS-NRTVKKIKVSDSGA 149 (323)
Q Consensus 103 f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~S-sk~Vk~Ikv~L~q~ 149 (323)
+.-|-.++...|++..|.|||+|.=.|.|.-.+ ...|.+|.+.|.-.
T Consensus 8 ~GiG~akVDT~L~~~~~~pGe~v~G~V~i~GG~v~Q~I~~I~l~L~t~ 55 (218)
T PF07070_consen 8 IGIGGAKVDTVLEKPSVRPGETVRGEVHIKGGSVDQEIDRIYLELVTR 55 (218)
T ss_pred cCCCCceEEEEECCCCccCCCEEEEEEEEEeCCcceEEeEEEEEEEEE
Confidence 456889999999999999999999999999996 55899999999753
No 26
>PF10633 NPCBM_assoc: NPCBM-associated, NEW3 domain of alpha-galactosidase; InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=64.50 E-value=5.9 Score=30.21 Aligned_cols=29 Identities=24% Similarity=0.348 Sum_probs=21.1
Q ss_pred ecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223 120 YHGESIAVNVHVANNSNRTVKKIKVSDSG 148 (323)
Q Consensus 120 ~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q 148 (323)
.|||++.+++.|.|.....+..+++.+.-
T Consensus 2 ~~G~~~~~~~tv~N~g~~~~~~v~~~l~~ 30 (78)
T PF10633_consen 2 TPGETVTVTLTVTNTGTAPLTNVSLSLSL 30 (78)
T ss_dssp -TTEEEEEEEEEE--SSS-BSS-EEEEE-
T ss_pred CCCCEEEEEEEEEECCCCceeeEEEEEeC
Confidence 48999999999999998889888888853
No 27
>PF07705 CARDB: CARDB; InterPro: IPR011635 The APHP (acidic peptide-dependent hydrolases/peptidase) domain is found in a variety of different proteins.; PDB: 2KUT_A 2L0D_A 3IDU_A 2KL6_A.
Probab=60.99 E-value=26 Score=26.77 Aligned_cols=41 Identities=20% Similarity=0.263 Sum_probs=30.7
Q ss_pred eEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223 108 LHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSG 148 (323)
Q Consensus 108 I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q 148 (323)
+.+.+........+|+.+.|++.|.|.-......+.+.+..
T Consensus 4 L~v~~~~~~~~~~~g~~~~i~~~V~N~G~~~~~~~~v~~~~ 44 (101)
T PF07705_consen 4 LTVSITVSPSNVVPGEPVTITVTVKNNGTADAENVTVRLYL 44 (101)
T ss_dssp EEE-EEEC-SEEETTSEEEEEEEEEE-SSS-BEEEEEEEEE
T ss_pred EEEEEeeCCCcccCCCEEEEEEEEEECCCCCCCCEEEEEEE
Confidence 34455677778889999999999999988888888888764
No 28
>PF00963 Cohesin: Cohesin domain; InterPro: IPR002102 Cohesin domains interact with a complementary domain, termed the dockerin domain (see IPR002105 from INTERPRO). The cohesin-dockerin interaction is the crucial interaction for complex formation in the cellulosome []. The scaffoldin component of the cellulolytic bacterium Clostridium thermocellum is a non-hydrolytic protein which organises the hydrolytic enzymes in a large complex, called the cellulosome. Scaffoldin comprises a series of functional domains, amongst which is a single cellulose-binding domain and nine cohesin domains which are responsible for integrating the individual enzymatic subunits into the complex.; GO: 0030246 carbohydrate binding, 0000272 polysaccharide catabolic process; PDB: 2BM3_A 3P0D_I 3KCP_A 2B59_A 3L8Q_B 3FNK_C 3GHP_B 2CCL_A 1ANU_A 1OHZ_A ....
Probab=57.66 E-value=28 Score=29.24 Aligned_cols=38 Identities=29% Similarity=0.315 Sum_probs=31.6
Q ss_pred EEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223 110 LEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSG 148 (323)
Q Consensus 110 L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q 148 (323)
|++.+++.--.|||++.|.+.++|-++. |..+.+.+.=
T Consensus 1 v~l~~~~~~a~~G~tv~V~V~v~~~~~~-i~~~~~~l~y 38 (141)
T PF00963_consen 1 VTLSVDSVSAKPGETVTVPVNVSNVSNS-IAGMQFTLSY 38 (141)
T ss_dssp EEEEESECEE-TTSEEEEEEEEESCTTT-EEEEEEEEEE
T ss_pred CEEEeCCceECCCCEEEEEEEEEcCCCc-EEEEEEEEEe
Confidence 5677888888999999999999999766 8888888864
No 29
>PLN02171 endoglucanase
Probab=54.10 E-value=17 Score=39.17 Aligned_cols=22 Identities=23% Similarity=0.354 Sum_probs=20.0
Q ss_pred eEEEEEEEecCCcceEeeEEee
Q psy17223 231 SIAVNVHVANNSNRTVKKIKVS 252 (323)
Q Consensus 231 ~i~v~v~v~N~s~k~vkkikv~ 252 (323)
-..+.|.|+|+|+|+||.|+|.
T Consensus 554 y~qy~v~I~N~s~~~ik~i~i~ 575 (629)
T PLN02171 554 YYRYSTTVTNRSAKTLKELHLG 575 (629)
T ss_pred EEEEEEEEEECCCCceeeeeee
Confidence 4678889999999999999997
No 30
>PF07703 A2M_N_2: Alpha-2-macroglobulin family N-terminal region; InterPro: IPR011625 This is a domain of the alpha-2-macroglobulin family. The alpha-macroglobulin (aM) family of proteins includes protease inhibitors [], typified by the human tetrameric a2-macroglobulin (a2M); they belong to the MEROPS proteinase inhibitor family I39, clan IL. These protease inhibitors share several defining properties, which include (i) the ability to inhibit proteases from all catalytic classes, (ii) the presence of a 'bait region' and a thiol ester, (iii) a similar protease inhibitory mechanism and (iv) the inactivation of the inhibitory capacity by reaction of the thiol ester with small primary amines. aM protease inhibitors inhibit by steric hindrance []. The mechanism involves protease cleavage of the bait region, a segment of the aM that is particularly susceptible to proteolytic cleavage, which initiates a conformational change such that the aM collapses about the protease. In the resulting aM-protease complex, the active site of the protease is sterically shielded, thus substantially decreasing access to protein substrates. Two additional events occur as a consequence of bait region cleavage, namely (i) the h-cysteinyl-g-glutamyl thiol ester becomes highly reactive and (ii) a major conformational change exposes a conserved COOH-terminal receptor binding domain [] (RBD). RBD exposure allows the aM protease complex to bind to clearance receptors and be removed from circulation []. Tetrameric, dimeric, and, more recently, monomeric aM protease inhibitors have been identified [, ].; PDB: 2QKI_D 3L3O_D 3NMS_A 2ICF_A 2A73_A 2ICE_D 2HR0_A 2A74_A 2XWJ_G 3OHX_A ....
Probab=53.80 E-value=16 Score=30.07 Aligned_cols=32 Identities=22% Similarity=0.356 Sum_probs=25.7
Q ss_pred ceEEEEEeCccceecCCeEEEEEEEeccCcce
Q psy17223 107 KLHLEASLDKELYYHGESIAVNVHVANNSNRT 138 (323)
Q Consensus 107 ~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~ 138 (323)
..++++..++..|.|||++.+++.....|...
T Consensus 94 ~~~v~l~~~~~~~~Pg~~~~~~i~~~~~s~v~ 125 (136)
T PF07703_consen 94 ELKVELTASPDEYKPGEEVTLRIKAPPNSLVG 125 (136)
T ss_dssp SSSEEEEESSSSBTTTSEEEEEEEESTTEEEE
T ss_pred cceEEEEEecceeCCCCEEEEEEEeCCCCEEE
Confidence 56788888999999999999999985554433
No 31
>KOG2625|consensus
Probab=52.44 E-value=8.4 Score=36.74 Aligned_cols=27 Identities=33% Similarity=0.582 Sum_probs=24.3
Q ss_pred EeecceEEEEEEEecCCcceEeeEEee
Q psy17223 226 YYHGESIAVNVHVANNSNRTVKKIKVS 252 (323)
Q Consensus 226 yyhge~i~v~v~v~N~s~k~vkkikv~ 252 (323)
-|-||+...-|+|.|.|+||||.|-+-
T Consensus 11 iflgetfs~yinv~nds~k~v~~i~lk 37 (348)
T KOG2625|consen 11 IFLGETFSFYINVHNDSEKTVKDILLK 37 (348)
T ss_pred eeeccceEEEEEEecchhhhhhhheee
Confidence 356999999999999999999999884
No 32
>PF05688 DUF824: Salmonella repeat of unknown function (DUF824); InterPro: IPR008542 This family consists of a series of repeated sequences (of around 180 residues) which are found in Salmonella typhimurium, Salmonella typhi and Escherichia coli. These repeats are almost always found with this entry. The repeats are associated with RatA and RatB, the coding sequences of which are found in the pathogeneicity island of Salmonella. The sequences may be determinants of pathogenicity [, ].
Probab=49.71 E-value=18 Score=25.96 Aligned_cols=29 Identities=21% Similarity=0.280 Sum_probs=26.2
Q ss_pred ecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223 120 YHGESIAVNVHVANNSNRTVKKIKVSDSG 148 (323)
Q Consensus 120 ~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q 148 (323)
.-||+|+++|.+.|.....|-..-+.|.+
T Consensus 10 K~Ge~I~ltVt~kda~G~pv~n~~f~l~r 38 (47)
T PF05688_consen 10 KVGETIPLTVTVKDANGNPVPNAPFTLTR 38 (47)
T ss_pred ecCCeEEEEEEEECCCCCCcCCceEEEEe
Confidence 45999999999999999999999888876
No 33
>PF11611 DUF4352: Domain of unknown function (DUF4352); InterPro: IPR021652 This entry is represented by Bacteriophage A118, Gp32. The characteristics of the protein distribution suggest prophage matches in addition to the phage matches. This entry represents a group of putative lipoproteins of unknown function.; PDB: 3CFU_A.
Probab=49.45 E-value=30 Score=27.78 Aligned_cols=53 Identities=21% Similarity=0.359 Sum_probs=30.2
Q ss_pred cceEEEEEEEecCCcceEeeEEeeeEEEeeC--CeEeeeeccccC-----CCccCCCCccc
Q psy17223 229 GESIAVNVHVANNSNRTVKKIKVSDICLFST--AQYKCTVAETES-----DCPIAPVSMFD 282 (323)
Q Consensus 229 ge~i~v~v~v~N~s~k~vkkikv~dv~l~s~--~~y~~~Va~~e~-----~~~i~p~~t~~ 282 (323)
+..+.|+|.|.|++++.+- +-..++.|+.. .+|......... ...|.||.+.+
T Consensus 35 ~~fv~v~v~v~N~~~~~~~-~~~~~f~l~d~~g~~~~~~~~~~~~~~~~~~~~i~pG~~~~ 94 (123)
T PF11611_consen 35 NKFVVVDVTVKNNGDEPLD-FSPSDFKLYDSDGNKYDPDFSASSNDNDLFSETIKPGESVT 94 (123)
T ss_dssp SEEEEEEEEEEE-SSS-EE-EEGGGEEEE-TT--B--EEE-CCCTTTB--EEEE-TT-EEE
T ss_pred CEEEEEEEEEEECCCCcEE-ecccceEEEeCCCCEEcccccchhccccccccEECCCCEEE
Confidence 5679999999999888774 43348889833 345543333332 35899998866
No 34
>PF04744 Monooxygenase_B: Monooxygenase subunit B protein; InterPro: IPR006833 Ammonia monooxygenase and the particulate methane monooxygenase are both integral membrane proteins, occurring in ammonia oxidisers and methanotrophs respectively, which are thought to be evolutionarily related []. These enzymes have a relatively wide substrate specificity and can catalyse the oxidation of a range of substrates including ammonia, methane, halogenated hydrocarbons and aromatic molecules []. These enzymes are composed of 3 subunits - A (IPR003393 from INTERPRO), B (IPR006833 from INTERPRO) and C (IPR006980 from INTERPRO) - and contain various metal centres, including copper. Particulate methane monooxygenase from Methylococcus capsulatus str. Bath is an ABC homotrimer, which contains mononuclear and dinuclear copper metal centres, and a third metal centre containing a metal ion whose identity in vivo is not certain[]. The soluble regions of these enzymes derive primarily from the B subunit. This subunit forms two antiparallel beta-barrel-like structures and contains the mono- and di- nuclear copper metal centres [].; PDB: 3CHX_E 3RFR_A 3RGB_A 1YEW_A.
Probab=48.91 E-value=20 Score=36.21 Aligned_cols=37 Identities=16% Similarity=0.328 Sum_probs=29.7
Q ss_pred ceEEEEEeCccce-ecCCeEEEEEEEeccCcceeeEEE
Q psy17223 107 KLHLEASLDKELY-YHGESIAVNVHVANNSNRTVKKIK 143 (323)
Q Consensus 107 ~I~L~a~LdK~~Y-~PGE~I~V~v~IdN~Ssk~Vk~Ik 143 (323)
+-.+++.+.+-.| .||-++.++++|+|+++..|+==+
T Consensus 246 ~~~V~~~v~~A~Y~vpgR~l~~~l~VtN~g~~pv~Lge 283 (381)
T PF04744_consen 246 PNSVKVKVTDATYRVPGRTLTMTLTVTNNGDSPVRLGE 283 (381)
T ss_dssp -SSEEEEEEEEEEESSSSEEEEEEEEEEESSS-BEEEE
T ss_pred CCceEEEEeccEEecCCcEEEEEEEEEcCCCCceEeee
Confidence 3338888888888 899999999999999999887544
No 35
>COG2373 Large extracellular alpha-helical protein [General function prediction only]
Probab=48.87 E-value=33 Score=40.93 Aligned_cols=131 Identities=17% Similarity=0.189 Sum_probs=72.6
Q ss_pred ceEEEEEeCccceecCCeEEEEEEEeccCcc-eeeEEEeec--CCCccchhhhhhhccCCCCCCccC------CCccc-c
Q psy17223 107 KLHLEASLDKELYYHGESIAVNVHVANNSNR-TVKKIKVSD--SGAEDDQDLKDELADSDIDGMEED------DLPNI-K 176 (323)
Q Consensus 107 ~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk-~Vk~Ikv~L--~q~ed~~d~~~~~~~~di~~~~e~------~lp~~-~ 176 (323)
.+++-++-||..|.|||++.+.+-....-.+ .+..+-+++ .+ +| + ..++-..+..++++ .||.. +
T Consensus 393 ~~k~y~ftDRglYRpGE~v~~~~~~R~~~~~~a~~~~p~~l~v~~-Pd-G---~~~~~~~~~~~~~G~~~~~~~l~~na~ 467 (1621)
T COG2373 393 GLKVYLFTDRGLYRPGETVHVNALLRDFDGKTALDNQPLKLRVLD-PD-G---SVLRTLTITLDEEGLYELSFPLPENAL 467 (1621)
T ss_pred ceEEEEecCcccCCCCceeeeeeeehhhcccccccCCCeEEEEEC-CC-C---cEEEEEEEeccccCceEEeeeCCCCCC
Confidence 5788899999999999999999988776655 444443333 32 11 1 22333333332222 44433 1
Q ss_pred ccccccceeeecccccCCCCCCCCCCcchhhhhhccCcc--hhhccceeeEEeecceEEEEEEEecCCcceEeeEEee
Q psy17223 177 AWGKNKRMYYNTDYVDDDHGGIQSGTGFVWFEWLKKGSK--EKSKKKYLFLYYHGESIAVNVHVANNSNRTVKKIKVS 252 (323)
Q Consensus 177 awg~~~~~~y~td~~~~~~~~~~~~~~~~~~e~~~~~~~--~~s~~k~~~~yyhge~i~v~v~v~N~s~k~vkkikv~ 252 (323)
.=|=.-+++++.. . ...+..+--.+++.+ +. ..+++| ..|.||+++.++|...|-.-.=+.+-++.
T Consensus 468 tG~w~l~~~~~~~------~-~~~s~~f~V~df~p~-r~~i~l~~~k--~~~~~g~~v~~~v~~~yL~GaPa~g~~~~ 535 (1621)
T COG2373 468 TGGYTLELYTGGK------S-AVISMSFRVEDFIPD-RFKINLTLDK--TEWVPGKDVKIKVDLRYLYGAPAAGLTVQ 535 (1621)
T ss_pred cceEEEEEEeCCc------c-ceeeeeEEhhHhCCc-eEEEeccccc--ccccCCCcEEEEEEEEecCCCcccCceee
Confidence 1111122333111 0 222222222222322 33 456666 55999999999999988876666666654
No 36
>PF09478 CBM49: Carbohydrate binding domain CBM49; InterPro: IPR019028 A carbohydrate-binding module (CBM) is defined as a contiguous amino acid sequence within a carbohydrate-active enzyme with a discreet fold having carbohydrate-binding activity. A few exceptions are CBMs in cellulosomal scaffolding proteins and rare instances of independent putative CBMs. The requirement of CBMs existing as modules within larger enzymes sets this class of carbohydrate-binding protein apart from other non-catalytic sugar binding proteins such as lectins and sugar transport proteins. CBMs were previously classified as cellulose-binding domains (CBDs) based on the initial discovery of several modules that bound cellulose [, ]. However, additional modules in carbohydrate-active enzymes are continually being found that bind carbohydrates other than cellulose yet otherwise meet the CBM criteria, hence the need to reclassify these polypeptides using more inclusive terminology. Previous classification of cellulose-binding domains were based on amino acid similarity. Groupings of CBDs were called "Types" and numbered with roman numerals (e.g. Type I or Type II CBDs). In keeping with the glycoside hydrolase classification, these groupings are now called families and numbered with Arabic numerals. Families 1 to 13 are the same as Types I to XIII. For a detailed review on the structure and binding modes of CBMs see []. This domain is found at the C-terminal of cellulases and in vitro binding studies have shown it to binds to crystalline cellulose []. ; GO: 0030246 carbohydrate binding, 0005576 extracellular region
Probab=48.15 E-value=50 Score=25.53 Aligned_cols=39 Identities=21% Similarity=0.354 Sum_probs=30.2
Q ss_pred EEEEEeCccceecCC-eEEEEEEEeccCcceeeEEEeecC
Q psy17223 109 HLEASLDKELYYHGE-SIAVNVHVANNSNRTVKKIKVSDS 147 (323)
Q Consensus 109 ~L~a~LdK~~Y~PGE-~I~V~v~IdN~Ssk~Vk~Ikv~L~ 147 (323)
+++-.+......-|. -....+.|.|+++++|+.+.+...
T Consensus 2 ~i~q~~~~sW~~~g~~y~qy~v~I~N~~~~~I~~~~i~~~ 41 (80)
T PF09478_consen 2 TITQTLVNSWTENGQTYTQYDVTITNNGSKPIKSLKISID 41 (80)
T ss_pred EEEEEEEeEEEeCCEEEEEEEEEEEECCCCeEEEEEEEEC
Confidence 455555666666665 357899999999999999999886
No 37
>PF07703 A2M_N_2: Alpha-2-macroglobulin family N-terminal region; InterPro: IPR011625 This is a domain of the alpha-2-macroglobulin family. The alpha-macroglobulin (aM) family of proteins includes protease inhibitors [], typified by the human tetrameric a2-macroglobulin (a2M); they belong to the MEROPS proteinase inhibitor family I39, clan IL. These protease inhibitors share several defining properties, which include (i) the ability to inhibit proteases from all catalytic classes, (ii) the presence of a 'bait region' and a thiol ester, (iii) a similar protease inhibitory mechanism and (iv) the inactivation of the inhibitory capacity by reaction of the thiol ester with small primary amines. aM protease inhibitors inhibit by steric hindrance []. The mechanism involves protease cleavage of the bait region, a segment of the aM that is particularly susceptible to proteolytic cleavage, which initiates a conformational change such that the aM collapses about the protease. In the resulting aM-protease complex, the active site of the protease is sterically shielded, thus substantially decreasing access to protein substrates. Two additional events occur as a consequence of bait region cleavage, namely (i) the h-cysteinyl-g-glutamyl thiol ester becomes highly reactive and (ii) a major conformational change exposes a conserved COOH-terminal receptor binding domain [] (RBD). RBD exposure allows the aM protease complex to bind to clearance receptors and be removed from circulation []. Tetrameric, dimeric, and, more recently, monomeric aM protease inhibitors have been identified [, ].; PDB: 2QKI_D 3L3O_D 3NMS_A 2ICF_A 2A73_A 2ICE_D 2HR0_A 2A74_A 2XWJ_G 3OHX_A ....
Probab=46.46 E-value=24 Score=29.06 Aligned_cols=25 Identities=36% Similarity=0.444 Sum_probs=20.2
Q ss_pred EEEEeCccceecCCeEEEEEEEecc
Q psy17223 110 LEASLDKELYYHGESIAVNVHVANN 134 (323)
Q Consensus 110 L~a~LdK~~Y~PGE~I~V~v~IdN~ 134 (323)
|++.+||..|.|||++.+.+.-.-.
T Consensus 1 l~i~~~~~~~~~Ge~~~v~v~~~~~ 25 (136)
T PF07703_consen 1 LQISTDKDSYKPGETAKVTVQSPFP 25 (136)
T ss_dssp EEEEE-SSSB-TTSEEEEEEEEESC
T ss_pred CEEEcCCCCcCCCCEEEEEEEcCCC
Confidence 6789999999999999999887665
No 38
>smart00809 Alpha_adaptinC2 Adaptin C-terminal domain. Adaptins are components of the adaptor complexes which link clathrin to receptors in coated vesicles. Clathrin-associated protein complexes are believed to interact with the cytoplasmic tails of membrane proteins, leading to their selection and concentration. Gamma-adaptin is a subunit of the golgi adaptor. Alpha adaptin is a heterotetramer that regulates clathrin-bud formation. The carboxyl-terminal appendage of the alpha subunit regulates translocation of endocytic accessory proteins to the bud site. This Ig-fold domain is found in alpha, beta and gamma adaptins and consists of a beta-sandwich containing 7 strands in 2 beta-sheets in a greek-key topology PUBMED:10430869, PUBMED:12176391. The adaptor appendage contains an additional N-terminal strand.
Probab=43.15 E-value=71 Score=25.12 Aligned_cols=41 Identities=12% Similarity=0.259 Sum_probs=33.4
Q ss_pred ecCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223 103 MSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDS 147 (323)
Q Consensus 103 f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~ 147 (323)
+.+..+++.+.+.++ +..+.|.+.+.|.|..+|.++.+.+.
T Consensus 2 ~~~~~l~I~~~~~~~----~~~~~i~~~~~N~s~~~it~f~~~~a 42 (104)
T smart00809 2 YEKNGLQIGFKFERR----PGLIRITLTFTNKSPSPITNFSFQAA 42 (104)
T ss_pred ccCCCEEEEEEEEcC----CCeEEEEEEEEeCCCCeeeeEEEEEE
Confidence 445668888888775 45689999999999999999998875
No 39
>TIGR03079 CH4_NH3mon_ox_B methane monooxygenase/ammonia monooxygenase, subunit B. Both ammonia oxidizers such as Nitrosomonas europaea and methanotrophs (obligate methane oxidizers) such as Methylococcus capsulatus each can grow only on their own characteristic substrate. However, both groups have the ability to oxidize both substrates, and so the relevant enzymes must be named here according to their ability to oxidze both. The protein family represented here reflects subunit B of both the particulate methane monooxygenase of methylotrophs and the ammonia monooxygenase of nitrifying bacteria.
Probab=42.03 E-value=37 Score=34.41 Aligned_cols=36 Identities=17% Similarity=0.350 Sum_probs=29.0
Q ss_pred CCceEEEEEeCccce-ecCCeEEEEEEEeccCcceee
Q psy17223 105 PNKLHLEASLDKELY-YHGESIAVNVHVANNSNRTVK 140 (323)
Q Consensus 105 sG~I~L~a~LdK~~Y-~PGE~I~V~v~IdN~Ssk~Vk 140 (323)
-++-.+++.+.+..| +||-++.++++|+|+++..|+
T Consensus 263 ~~~~~V~~kv~~a~Y~VPGR~l~~~~~VTN~g~~~vr 299 (399)
T TIGR03079 263 VAPNPVSINVTKANYDVPGRALRVTMEITNNGDQVIS 299 (399)
T ss_pred CCCCceEEEEeccEEecCCcEEEEEEEEEcCCCCceE
Confidence 344456666667666 899999999999999999886
No 40
>PF02014 Reeler: Reeler domain Schematic picture including Reeler domain; InterPro: IPR002861 Extracellular matrix (ECM) proteins play an important role in early cortical development, specifically in the formation of neural connections and in controlling the cyto-architecture of the central nervous system. The product of the reeler gene in mouse is reelin,a large extracellular protein secreted by pioneer neurons that coordinates cell positioning during neurodevelopment []. F-spondin and mindin are a family of matrix-attached adhesion molecules that share structural similarities and overlapping domains of expression. Both F-spondin and mindin promote adhesion and outgrowth of hippocampal embryonic neurons and bind to a putative receptor(s) expressed on both hippocampal and sensory neurons []. This domain of unknown function is found at the N terminus of reelin and F-spondin.; PDB: 2ZOT_B 2ZOU_B 3COO_A.
Probab=41.79 E-value=43 Score=28.07 Aligned_cols=35 Identities=11% Similarity=0.271 Sum_probs=25.3
Q ss_pred EEeCccceecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223 112 ASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSG 148 (323)
Q Consensus 112 a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q 148 (323)
+.++...|.||+.+.|++ .+.++...+++-++...
T Consensus 23 i~~~~~~y~pg~~~~Vtl--~~~~~~~F~GFllqAr~ 57 (132)
T PF02014_consen 23 ISVSPSSYEPGQTYTVTL--SSSGSSSFRGFLLQARD 57 (132)
T ss_dssp EEET-SSB-TTBEEEEEE--EETTTEEBSEEEEEEEE
T ss_pred EEeCCCeEcCCCEEEEEE--ECCCCCceeEEEEEEEe
Confidence 444599999999999998 66667788887776643
No 41
>PF05753 TRAP_beta: Translocon-associated protein beta (TRAPB); InterPro: IPR008856 This family consists of several eukaryotic translocon-associated protein beta (TRAPB) or signal sequence receptor beta subunit (SSR-beta) proteins. The normal translocation of nascent polypeptides into the lumen of the endoplasmic reticulum (ER) is thought to be aided in part by a translocon-associated protein (TRAP) complex consisting of 4 protein subunits. The association of mature proteins with the ER and Golgi, or other intracellular locales, such as lysosomes, depends on the initial targeting of the nascent polypeptide to the ER membrane. A similar scenario must also exist for proteins destined for secretion [].; GO: 0005783 endoplasmic reticulum, 0016021 integral to membrane
Probab=40.51 E-value=57 Score=29.54 Aligned_cols=74 Identities=16% Similarity=0.196 Sum_probs=45.7
Q ss_pred Eee-cceEEEEEEEecCCcceEeeEEeeeEEEeeCCeEeeeecc-ccC-CCccCCCCccc-ccceeecccccccccCCcc
Q psy17223 226 YYH-GESIAVNVHVANNSNRTVKKIKVSDICLFSTAQYKCTVAE-TES-DCPIAPVSMFD-TEDLAMLRHGFKRMFGHAF 301 (323)
Q Consensus 226 yyh-ge~i~v~v~v~N~s~k~vkkikv~dv~l~s~~~y~~~Va~-~e~-~~~i~p~~t~~-~~~~~~~~~~~~~~~~~a~ 301 (323)
|.. |+.+.|++.|-|.=+.+.-++++.| -=|..+.|. .|.. ... =+.|+||++.+ .+-|.+ .+...||+.|
T Consensus 33 ~~v~g~~v~V~~~iyN~G~~~A~dV~l~D-~~fp~~~F~-lvsG~~s~~~~~i~pg~~vsh~~vv~p---~~~G~f~~~~ 107 (181)
T PF05753_consen 33 YLVEGEDVTVTYTIYNVGSSAAYDVKLTD-DSFPPEDFE-LVSGSLSASWERIPPGENVSHSYVVRP---KKSGYFNFTP 107 (181)
T ss_pred cccCCcEEEEEEEEEECCCCeEEEEEEEC-CCCCccccE-eccCceEEEEEEECCCCeEEEEEEEee---eeeEEEEccC
Confidence 444 9999999999999999999999985 111112221 1111 011 14888998876 444443 4455666666
Q ss_pred ccc
Q psy17223 302 STS 304 (323)
Q Consensus 302 ~t~ 304 (323)
.++
T Consensus 108 a~V 110 (181)
T PF05753_consen 108 AVV 110 (181)
T ss_pred EEE
Confidence 544
No 42
>PF09624 DUF2393: Protein of unknown function (DUF2393); InterPro: IPR013417 The function of this protein is unknown. It is always found as part of a two-gene operon with IPR013416 from INTERPRO, a protein that appears to span the membrane seven times. It has so far been found in the bacteria Anabaena sp. (strain PCC 7120), Agrobacterium tumefaciens, Rhizobium meliloti, and Gloeobacter violaceus.
Probab=40.33 E-value=68 Score=27.40 Aligned_cols=29 Identities=31% Similarity=0.292 Sum_probs=25.3
Q ss_pred eecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223 119 YYHGESIAVNVHVANNSNRTVKKIKVSDS 147 (323)
Q Consensus 119 Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~ 147 (323)
..-+|.+-|...|.|.++++++++++.+.
T Consensus 58 l~~~~~~~v~g~V~N~g~~~i~~c~i~~~ 86 (149)
T PF09624_consen 58 LQYSESFYVDGTVTNTGKFTIKKCKITVK 86 (149)
T ss_pred eeeccEEEEEEEEEECCCCEeeEEEEEEE
Confidence 44689999999999999999998888774
No 43
>TIGR01451 B_ant_repeat conserved repeat domain. This model represents the conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis, and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydial outer membrane proteins.
Probab=40.07 E-value=56 Score=23.46 Aligned_cols=29 Identities=31% Similarity=0.475 Sum_probs=25.0
Q ss_pred EeecceEEEEEEEecCCcceEeeEEeeeE
Q psy17223 226 YYHGESIAVNVHVANNSNRTVKKIKVSDI 254 (323)
Q Consensus 226 yyhge~i~v~v~v~N~s~k~vkkikv~dv 254 (323)
..=|+.|...|.|.|+.......++|.|.
T Consensus 8 ~~~Gd~v~Yti~v~N~g~~~a~~v~v~D~ 36 (53)
T TIGR01451 8 ATIGDTITYTITVTNNGNVPATNVVVTDI 36 (53)
T ss_pred cCCCCEEEEEEEEEECCCCceEeEEEEEc
Confidence 34599999999999999999999988743
No 44
>cd08544 Reeler Reeler, the N-terminal domain of reelin, F-spondin, and a variety of other proteins. This domain is found at the N-terminus of F-spondin, a protein attached to the extracellular matrix, which plays roles in neuronal development and vascular remodelling. The F-spondin reeler domain has been reported to bind heparin. The reeler domain is also found at the N-terminus of reelin, an extracellular glycoprotein involved in the development of the brain cortex, and in a variety of other eukaryotic proteins with different domain architectures, including the animal ferric-chelate reductase 1 or stromal cell-derived receptor 2, a member of the cytochrome B561 family, which reduces ferric iron before its transport from the endosome to the cytoplasm. Also included is the insect putative defense protein 1, which is expressed upon bacterial infection and appears to contain a single reeler domain.
Probab=38.12 E-value=56 Score=27.34 Aligned_cols=37 Identities=11% Similarity=0.222 Sum_probs=27.1
Q ss_pred EEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223 110 LEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSG 148 (323)
Q Consensus 110 L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q 148 (323)
..+.++...|.|||.+.|++.-.|. ..-+++-++...
T Consensus 21 y~i~~~~~~y~pG~~~~Vtl~~~~~--~~F~GF~lqAr~ 57 (135)
T cd08544 21 YSITISGNSYVPGETYTVTLSGSSP--SPFRGFLLQARD 57 (135)
T ss_pred EEEEeCCCEECCCCEEEEEEECCCC--CceeEEEEEEEc
Confidence 4556667799999999999988776 456666555543
No 45
>PF01345 DUF11: Domain of unknown function DUF11; InterPro: IPR001434 This group of sequences is represented by a conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis (Streptococcus faecalis), and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydia trachomatis outer membrane proteins. In C. trachomatis, three cysteine-rich proteins (also believed to be lipoproteins), MOMP, OMP6 and OMP3, make up the extracellular matrix of the outer membrane []. They are involved in the essential structural integrity of both the elementary body (EB) and recticulate body (RB) phase. They are thought to be involved in porin formation and, as these bacteria lack the peptidoglycan layer common to most Gram-negative microbes, such proteins are highly important in the pathogenicity of the organism.; GO: 0005727 extrachromosomal circular DNA
Probab=38.10 E-value=57 Score=24.45 Aligned_cols=30 Identities=17% Similarity=0.304 Sum_probs=26.5
Q ss_pred eEEeecceEEEEEEEecCCcceEeeEEeee
Q psy17223 224 FLYYHGESIAVNVHVANNSNRTVKKIKVSD 253 (323)
Q Consensus 224 ~~yyhge~i~v~v~v~N~s~k~vkkikv~d 253 (323)
....=||.+...|.|+|..+....+++|.|
T Consensus 35 ~~~~~Gd~v~ytitvtN~G~~~a~nv~v~D 64 (76)
T PF01345_consen 35 STANPGDTVTYTITVTNTGPAPATNVVVTD 64 (76)
T ss_pred CcccCCCEEEEEEEEEECCCCeeEeEEEEE
Confidence 445669999999999999999999999974
No 46
>PF06030 DUF916: Bacterial protein of unknown function (DUF916); InterPro: IPR010317 This family consists of putative cell surface proteins, from Firmicutes, of unknown function.
Probab=36.91 E-value=37 Score=28.67 Aligned_cols=29 Identities=31% Similarity=0.528 Sum_probs=24.3
Q ss_pred ecCCeEEEEEEEeccCcceeeEEEeecCCC
Q psy17223 120 YHGESIAVNVHVANNSNRTVKKIKVSDSGA 149 (323)
Q Consensus 120 ~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~ 149 (323)
-||+...+++.|.|.|.+.++ +.+.+..+
T Consensus 24 ~P~q~~~l~v~i~N~s~~~~t-v~v~~~~A 52 (121)
T PF06030_consen 24 KPGQKQTLEVRITNNSDKEIT-VKVSANTA 52 (121)
T ss_pred CCCCEEEEEEEEEeCCCCCEE-EEEEEeee
Confidence 489999999999999998876 77777654
No 47
>PF14796 AP3B1_C: Clathrin-adaptor complex-3 beta-1 subunit C-terminal
Probab=36.11 E-value=85 Score=27.66 Aligned_cols=46 Identities=15% Similarity=0.340 Sum_probs=38.7
Q ss_pred ecCCceEEEEEeCcccee-cCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223 103 MSPNKLHLEASLDKELYY-HGESIAVNVHVANNSNRTVKKIKVSDSG 148 (323)
Q Consensus 103 f~sG~I~L~a~LdK~~Y~-PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q 148 (323)
...+=+.++-++.|.-+. ..-.+.|++.+.|+|...|++|.+....
T Consensus 64 v~G~GL~v~Y~F~RqP~~~s~~mvsIql~ftN~s~~~i~~I~i~~k~ 110 (145)
T PF14796_consen 64 VNGKGLSVEYRFSRQPSLYSPSMVSIQLTFTNNSDEPIKNIHIGEKK 110 (145)
T ss_pred cCCCceeEEEEEccCCcCCCCCcEEEEEEEEecCCCeecceEECCCC
Confidence 356778999999998884 4568899999999999999999987653
No 48
>PF13199 Glyco_hydro_66: Glycosyl hydrolase family 66; PDB: 3VMO_A 3VMN_A 3VMP_A.
Probab=36.04 E-value=53 Score=34.97 Aligned_cols=25 Identities=24% Similarity=0.460 Sum_probs=16.2
Q ss_pred EeCccceecCCeEEEEEEEeccCcc
Q psy17223 113 SLDKELYYHGESIAVNVHVANNSNR 137 (323)
Q Consensus 113 ~LdK~~Y~PGE~I~V~v~IdN~Ssk 137 (323)
+.||..|.|||+|.+++...|....
T Consensus 1 ~tDKA~Y~PGe~V~l~~~~~~~~~~ 25 (559)
T PF13199_consen 1 TTDKARYRPGEKVTLTASLKNTTGS 25 (559)
T ss_dssp EES-SSB-TTS-EEEE-EEE--SSS
T ss_pred CCCcceeCCCCeEEEEEEeccCccc
Confidence 4689999999999999999998544
No 49
>PF09624 DUF2393: Protein of unknown function (DUF2393); InterPro: IPR013417 The function of this protein is unknown. It is always found as part of a two-gene operon with IPR013416 from INTERPRO, a protein that appears to span the membrane seven times. It has so far been found in the bacteria Anabaena sp. (strain PCC 7120), Agrobacterium tumefaciens, Rhizobium meliloti, and Gloeobacter violaceus.
Probab=35.58 E-value=52 Score=28.17 Aligned_cols=34 Identities=35% Similarity=0.428 Sum_probs=27.7
Q ss_pred hhhccceeeEEeecceEEEEEEEecCCcceEeeEEee
Q psy17223 216 EKSKKKYLFLYYHGESIAVNVHVANNSNRTVKKIKVS 252 (323)
Q Consensus 216 ~~s~~k~~~~yyhge~i~v~v~v~N~s~k~vkkikv~ 252 (323)
....++++. | +|.+-|...|+|.+++++++.+|.
T Consensus 51 ~~~~~~~l~-~--~~~~~v~g~V~N~g~~~i~~c~i~ 84 (149)
T PF09624_consen 51 TLTSQKRLQ-Y--SESFYVDGTVTNTGKFTIKKCKIT 84 (149)
T ss_pred EEeeeeeee-e--ccEEEEEEEEEECCCCEeeEEEEE
Confidence 344555533 3 899999999999999999999998
No 50
>cd08548 Type_I_cohesin_like Type I cohesin domain, interaction partner of dockerin. Bacterial cohesin domains bind to a complementary protein domain named dockerin, and this interaction is required for the formation of the cellulosome, a cellulose-degrading complex. The cellulosome consists of scaffoldin, a noncatalytic scaffolding polypeptide, that comprises repeating cohesion modules and a single carbohydrate-binding module (CBM). Specific calcium-dependent interactions between cohesins and dockerins appear to be essential for cellulosome assembly. This subfamily represents type I cohesins; their interactions with dockerin mediate assembly of a range of dockerin-borne enzymes to the complex.
Probab=34.79 E-value=76 Score=27.07 Aligned_cols=40 Identities=15% Similarity=0.160 Sum_probs=34.1
Q ss_pred EEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCC
Q psy17223 110 LEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGA 149 (323)
Q Consensus 110 L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~ 149 (323)
+++.+++---.||+++.|-|.++|-.+..|..+.+.+.-+
T Consensus 1 ~~v~ig~v~~~~G~tv~VpV~~~~v~~~~i~~~~f~l~yD 40 (135)
T cd08548 1 VEVKIGSVSGKPGDTVTVPVTLSNVPSKGIGACDFVLSYD 40 (135)
T ss_pred CeEEeccEEecCCCEEEEEEEEecCCccCEEEEEEEEEeC
Confidence 3566777777899999999999999999999999988753
No 51
>PF06159 DUF974: Protein of unknown function (DUF974); InterPro: IPR010378 This is a family of uncharacterised eukaryotic proteins.
Probab=34.01 E-value=51 Score=31.16 Aligned_cols=28 Identities=29% Similarity=0.597 Sum_probs=25.3
Q ss_pred eecCCeEEEEEEEeccCcceeeEEEeec
Q psy17223 119 YYHGESIAVNVHVANNSNRTVKKIKVSD 146 (323)
Q Consensus 119 Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L 146 (323)
.|-||+....++|.|.|+..|+.+.++.
T Consensus 10 iylGEtF~~~l~~~N~s~~~v~~v~ikv 37 (249)
T PF06159_consen 10 IYLGETFSCYLSVNNDSNKPVRNVRIKV 37 (249)
T ss_pred EeecCCEEEEEEeecCCCCceEEeEEEE
Confidence 4679999999999999999999887776
No 52
>COG4326 Spo0M Sporulation control protein [General function prediction only]
Probab=34.00 E-value=57 Score=30.85 Aligned_cols=21 Identities=29% Similarity=0.523 Sum_probs=17.9
Q ss_pred cCCCeeeeeEeeCCCCCCCcE
Q psy17223 18 KLGPNAFPFFFELPPSCPASV 38 (323)
Q Consensus 18 ~~G~h~FPFsFqLP~~LP~SF 38 (323)
|--.+.|||+|.||-+.|=+|
T Consensus 104 pgEe~~fpf~l~lP~~tPvT~ 124 (270)
T COG4326 104 PGEERNFPFELSLPWNTPVTI 124 (270)
T ss_pred CCceEeccEEEecCCCCceee
Confidence 334599999999999999887
No 53
>PF02883 Alpha_adaptinC2: Adaptin C-terminal domain; InterPro: IPR008152 Proteins synthesized on the ribosome and processed in the endoplasmic reticulum are transported from the Golgi apparatus to the trans-Golgi network (TGN), and from there via small carrier vesicles to their final destination compartment. These vesicles have specific coat proteins (such as clathrin or coatomer) that are important for cargo selection and direction of transport []. Clathrin coats contain both clathrin (acts as a scaffold) and adaptor complexes that link clathrin to receptors in coated vesicles. Clathrin-associated protein complexes are believed to interact with the cytoplasmic tails of membrane proteins, leading to their selection and concentration. The two major types of clathrin adaptor complexes are the heterotetrameric adaptor protein (AP) complexes, and the monomeric GGA (Golgi-localising, Gamma-adaptin ear domain homology, ARF-binding proteins) adaptors [, ]. AP (adaptor protein) complexes are found in coated vesicles and clathrin-coated pits. AP complexes connect cargo proteins and lipids to clathrin at vesicle budding sites, as well as binding accessory proteins that regulate coat assembly and disassembly (such as AP180, epsins and auxilin). There are different AP complexes in mammals. AP1 is responsible for the transport of lysosomal hydrolases between the TGN and endosomes []. AP2 associates with the plasma membrane and is responsible for endocytosis []. AP3 is responsible for protein trafficking to lysosomes and other related organelles []. AP4 is less well characterised. AP complexes are heterotetramers composed of two large subunits (adaptins), a medium subunit (mu) and a small subunit (sigma). For example, in AP1 these subunits are gamma-1-adaptin, beta-1-adaptin, mu-1 and sigma-1, while in AP2 they are alpha-adaptin, beta-2-adaptin, mu-2 and sigma-2. Each subunit has a specific function. Adaptins recognise and bind to clathrin through their hinge region (clathrin box), and recruit accessory proteins that modulate AP function through their C-terminal ear (appendage) domains. Mu recognises tyrosine-based sorting signals within the cytoplasmic domains of transmembrane cargo proteins []. One function of clathrin and AP2 complex-mediated endocytosis is to regulate the number of GABA(A) receptors available at the cell surface []. GGAs (Golgi-localising, Gamma-adaptin ear domain homology, ARF-binding proteins) are a family of monomeric clathrin adaptor proteins that are conserved from yeasts to humans. GGAs regulate clathrin-mediated the transport of proteins (such as mannose 6-phosphate receptors) from the TGN to endosomes and lysosomes through interactions with TGN-sorting receptors, sometimes in conjunction with AP-1 [, ]. GGAs bind cargo, membranes, clathrin and accessory factors. GGA1, GGA2 and GGA3 all contain a domain homologous to the ear domain of gamma-adaptin. GGAs are composed of a single polypeptide with four domains: an N-terminal VHS (Vps27p/Hrs/Stam) domain, a GAT (GGA and Tom1) domain, a hinge region, and a C-terminal GAE (gamma-adaptin ear) domain. The VHS domain is responsible for endocytosis and signal transduction, recognising transmembrane cargo through the ACLL sequence in the cytoplasmic domains of sorting receptors []. The GAT domain (also found in Tom1 proteins) interacts with ARF (ADP-ribosylation factor) to regulate membrane trafficking [], and with ubiquitin for receptor sorting []. The hinge region contains a clathrin box for recognition and binding to clathrin, similar to that found in AP adaptins. The GAE domain is similar to the AP gamma-adaptin ear domain, and is responsible for the recruitment of accessory proteins that regulate clathrin-mediated endocytosis []. This entry represents a beta-sandwich structural motif found in the appendage (ear) domain of alpha-, beta- and gamma-adaptin from AP clathrin adaptor complexes, and the GAE (gamma-adaptin ear) domain of GGA adaptor proteins. These domains have an immunoglobulin-like beta-sandwich fold containing 7 or 8 strands in 2 beta-sheets in a Greek key topology [, ]. Although these domains share a similar fold, there is little sequence identity between the alpha/beta-adaptins and gamma-adaptin/GAE. More information about these proteins can be found at Protein of the Month: Clathrin [].; GO: 0006886 intracellular protein transport, 0016192 vesicle-mediated transport, 0030131 clathrin adaptor complex; PDB: 3MNM_B 3ZY7_B 1GYU_A 1GYW_B 2A7B_A 1GYV_A 2E9G_A 1E42_B 2G30_A 2IV9_B ....
Probab=33.97 E-value=1.4e+02 Score=24.02 Aligned_cols=45 Identities=11% Similarity=0.179 Sum_probs=36.1
Q ss_pred EeEecCCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEeec
Q psy17223 100 EFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSD 146 (323)
Q Consensus 100 ~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L 146 (323)
..++.+..++|.+.+.+ -..+..+.|.+.+.|.+...|.++.+.+
T Consensus 3 ~~~ye~~~l~I~~~~~~--~~~~~~~~i~~~f~N~s~~~it~f~~q~ 47 (115)
T PF02883_consen 3 GVLYEDNGLQIGFKSEK--SPNPNQGRIKLTFGNKSSQPITNFSFQA 47 (115)
T ss_dssp EEEEEETTEEEEEEEEE--CCETTEEEEEEEEEE-SSS-BEEEEEEE
T ss_pred EEEEeCCCEEEEEEEEe--cCCCCEEEEEEEEEECCCCCcceEEEEE
Confidence 34677788888888887 4567889999999999999999999987
No 54
>COG2373 Large extracellular alpha-helical protein [General function prediction only]
Probab=32.00 E-value=91 Score=37.36 Aligned_cols=41 Identities=17% Similarity=0.396 Sum_probs=35.9
Q ss_pred CCceEEEEEeCccceecCCeEEEEEEEeccCcceeeEEEee
Q psy17223 105 PNKLHLEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVS 145 (323)
Q Consensus 105 sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~ 145 (323)
.-.+.+.+++++..+.+|+.+.+++...|....++...++.
T Consensus 495 p~r~~i~l~~~k~~~~~g~~v~~~v~~~yL~GaPa~g~~~~ 535 (1621)
T COG2373 495 PDRFKINLTLDKTEWVPGKDVKIKVDLRYLYGAPAAGLTVQ 535 (1621)
T ss_pred CceEEEecccccccccCCCcEEEEEEEEecCCCcccCceee
Confidence 44567888999999999999999999999999888777766
No 55
>PF13595 DUF4138: Domain of unknown function (DUF4138)
Probab=31.80 E-value=50 Score=31.38 Aligned_cols=21 Identities=29% Similarity=0.543 Sum_probs=19.9
Q ss_pred eEEeecceEEEEEEEecCCcc
Q psy17223 224 FLYYHGESIAVNVHVANNSNR 244 (323)
Q Consensus 224 ~~yyhge~i~v~v~v~N~s~k 244 (323)
+||+||+-+-+.+.|.|+|+-
T Consensus 136 ~Iy~~~d~lyf~~~i~N~S~i 156 (246)
T PF13595_consen 136 NIYVHGDYLYFHLSIKNKSNI 156 (246)
T ss_pred eEEEECCEEEEEEEEEcCCCC
Confidence 899999999999999999984
No 56
>PF04425 Bul1_N: Bul1 N terminus; InterPro: IPR007519 This domain is the N terminus of Saccharomyces cerevisiae (Baker's yeast) Bul1. Bul1 binds the ubiquitin ligase Rsp5, via an N-terminal PPSY motif (157-160 in P48524 from SWISSPROT) []. The complex containing Bul1 and Rsp5 is involved in intracellular trafficking of the general amino acid permease Gap1 [], degradation of Rog1 in cooperation with Bul2 and GSK-3 [], and mitochondrial inheritance []. Bul1 may contain HEAT repeats. The C terminus is IPR007520 from INTERPRO.
Probab=31.47 E-value=28 Score=35.96 Aligned_cols=22 Identities=18% Similarity=0.285 Sum_probs=18.1
Q ss_pred HHHHcccCCCeeeeeEeeCCCC
Q psy17223 12 QERLMKKLGPNAFPFFFELPPS 33 (323)
Q Consensus 12 Q~~l~~~~G~h~FPFsFqLP~~ 33 (323)
..+.++|--.|.++|.|+||..
T Consensus 252 ~~r~l~p~~~Yk~fF~FkiP~~ 273 (438)
T PF04425_consen 252 NKRILEPGVKYKKFFTFKIPEQ 273 (438)
T ss_pred CCceecCCCeEeceeEEeCCch
Confidence 4566777778999999999984
No 57
>PF10633 NPCBM_assoc: NPCBM-associated, NEW3 domain of alpha-galactosidase; InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=31.19 E-value=48 Score=25.09 Aligned_cols=24 Identities=25% Similarity=0.548 Sum_probs=18.6
Q ss_pred cceEEEEEEEecCCcceEeeEEee
Q psy17223 229 GESIAVNVHVANNSNRTVKKIKVS 252 (323)
Q Consensus 229 ge~i~v~v~v~N~s~k~vkkikv~ 252 (323)
|+++.+.+.|.|+....+..++++
T Consensus 4 G~~~~~~~tv~N~g~~~~~~v~~~ 27 (78)
T PF10633_consen 4 GETVTVTLTVTNTGTAPLTNVSLS 27 (78)
T ss_dssp TEEEEEEEEEE--SSS-BSS-EEE
T ss_pred CCEEEEEEEEEECCCCceeeEEEE
Confidence 899999999999999999988886
No 58
>cd00258 GM2-AP GM2 activator protein (GM2-AP) is a non-enzymatic lysosomal protein that acts as cofactor in the sequential degradation of gangliosides. GM2A is an essential cofactor for beta-hexosaminidase A (Hex A) in the enzymatic hydrolysis of GM2 ganglioside to GM3. Mutation of the gene results in the AB variant of Tay-Sachs disease. GM2-AP and similar proteins belong to the ML domain family.
Probab=30.74 E-value=61 Score=29.18 Aligned_cols=35 Identities=29% Similarity=0.473 Sum_probs=25.8
Q ss_pred ccCCCeeeee-EeeCCC-CCCCcEEeccCCCCCCCceeeEEEEEEEEcc
Q psy17223 17 KKLGPNAFPF-FFELPP-SCPASVTLQPAPGDTGKPCGVDYELKAFVGE 63 (323)
Q Consensus 17 ~~~G~h~FPF-sFqLP~-~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r 63 (323)
.++|.|..|= +|.||. +||+... .|+ |++++.+++
T Consensus 109 ~~~G~y~lp~s~f~lP~~~LPs~l~-------~G~-----Y~i~~~l~~ 145 (162)
T cd00258 109 FKEGVYSLPDSTFTLPNVDLPSWLT-------NGN-----YRITGILMA 145 (162)
T ss_pred CCCcceEccceeeecccccCCCccC-------CCc-----EEEEEEECC
Confidence 5688888854 558987 6887653 562 999999853
No 59
>cd00917 PG-PI_TP The phosphatidylinositol/phosphatidylglycerol transfer protein (PG/PI-TP) has been shown to bind phosphatidylglycerol and phosphatidylinositol, but the biological significance of this is still obscure. These proteins belong to the ML domain family.
Probab=30.69 E-value=87 Score=26.05 Aligned_cols=32 Identities=19% Similarity=0.225 Sum_probs=23.6
Q ss_pred cccCCCeeeeeEeeCCCCCCCcEEeccCCCCCCCceeeEEEEEEEEcc
Q psy17223 16 MKKLGPNAFPFFFELPPSCPASVTLQPAPGDTGKPCGVDYELKAFVGE 63 (323)
Q Consensus 16 ~~~~G~h~FPFsFqLP~~LP~SF~~~~~~g~~Gk~c~IrY~VKA~I~r 63 (323)
=.++|.+.+..+..||...|+ +.|.|++.+..
T Consensus 78 Pi~~G~~~~~~~~~ip~~~P~----------------g~y~v~~~l~d 109 (122)
T cd00917 78 PIEPGDKFLTKLVDLPGEIPP----------------GKYTVSARAYT 109 (122)
T ss_pred CcCCCcEEEEEEeeCCCCCCC----------------ceEEEEEEEEC
Confidence 345788888888888876676 24888887743
No 60
>PF11355 DUF3157: Protein of unknown function (DUF3157); InterPro: IPR021501 This family of proteins with unknown function appears to be restricted to Gammaproteobacteria.
Probab=29.87 E-value=1.1e+02 Score=28.43 Aligned_cols=37 Identities=19% Similarity=0.389 Sum_probs=30.6
Q ss_pred EEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223 110 LEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDS 147 (323)
Q Consensus 110 L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~ 147 (323)
+++.|...-|--| .+-+...|.|+|+..|..|++.+.
T Consensus 100 VdV~l~~~~y~~~-~L~l~~~ltnqSsqsVv~Vel~v~ 136 (199)
T PF11355_consen 100 VDVSLGASQYEDG-QLGLPFSLTNQSSQSVVLVELEVT 136 (199)
T ss_pred eeEEEeccceeCC-eEEEEEEEecCCCceEEEEEEEEE
Confidence 6777777777777 899999999999999988877663
No 61
>PF04314 DUF461: Protein of unknown function (DUF461); InterPro: IPR007410 This entry represents a domain found in of proteins of unknown function, including DR1885 from Deinococcus radiodurans and CC3502 from Caulobacter crescentus (Caulobacter vibrioides), which share a potential metal binding motif H(M)X10MX21HXM. DR1885 was found to bind copper(I) through a histidine and three Mets in a cupredoxin-like fold []. The surface location of the copper-binding site as well as the type of coordination are well poised for metal transfer chemistry, suggesting that DR1885 might transfer copper, taking the role of Cox17 in bacteria (Cox17 being an accessory protein required for correct assembly of eukaryotic cyochrome c oxidase). ; PDB: 2K6W_A 2K6Z_A 2K6Y_A 2K70_A 1X9L_A 2JQA_A.
Probab=28.27 E-value=87 Score=25.65 Aligned_cols=40 Identities=18% Similarity=0.241 Sum_probs=30.9
Q ss_pred EeEecCCceEEEEEeCccceecCCeEEEEEEEeccCccee
Q psy17223 100 EFMMSPNKLHLEASLDKELYYHGESIAVNVHVANNSNRTV 139 (323)
Q Consensus 100 ~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~V 139 (323)
..-|..|-.+|-+.=.+....+||.+++++..+|.....|
T Consensus 70 ~v~l~pgg~HlmL~g~~~~l~~G~~v~ltL~f~~gg~v~v 109 (110)
T PF04314_consen 70 TVELKPGGYHLMLMGLKRPLKPGDTVPLTLTFEDGGKVTV 109 (110)
T ss_dssp EEEE-CCCCEEEEECESS-B-TTEEEEEEEEETTTEEEEE
T ss_pred eEEecCCCEEEEEeCCcccCCCCCEEEEEEEECCCCEEEe
Confidence 3457888888888777888999999999999999887665
No 62
>smart00809 Alpha_adaptinC2 Adaptin C-terminal domain. Adaptins are components of the adaptor complexes which link clathrin to receptors in coated vesicles. Clathrin-associated protein complexes are believed to interact with the cytoplasmic tails of membrane proteins, leading to their selection and concentration. Gamma-adaptin is a subunit of the golgi adaptor. Alpha adaptin is a heterotetramer that regulates clathrin-bud formation. The carboxyl-terminal appendage of the alpha subunit regulates translocation of endocytic accessory proteins to the bud site. This Ig-fold domain is found in alpha, beta and gamma adaptins and consists of a beta-sandwich containing 7 strands in 2 beta-sheets in a greek-key topology PUBMED:10430869, PUBMED:12176391. The adaptor appendage contains an additional N-terminal strand.
Probab=27.99 E-value=2.9e+02 Score=21.53 Aligned_cols=55 Identities=9% Similarity=0.059 Sum_probs=37.4
Q ss_pred eEEeecceEEEEEEEecCCcceEeeEEeeeEEEeeCCeEeeeeccccCCCccCCCCccc
Q psy17223 224 FLYYHGESIAVNVHVANNSNRTVKKIKVSDICLFSTAQYKCTVAETESDCPIAPVSMFD 282 (323)
Q Consensus 224 ~~yyhge~i~v~v~v~N~s~k~vkkikv~dv~l~s~~~y~~~Va~~e~~~~i~p~~t~~ 282 (323)
.+=+++..+.+.+...|++...+.++.+. +..-+|-+.-..--++..|.||+..+
T Consensus 12 ~~~~~~~~~~i~~~~~N~s~~~it~f~~~----~avpk~~~l~l~~~s~~~l~p~~~i~ 66 (104)
T smart00809 12 KFERRPGLIRITLTFTNKSPSPITNFSFQ----AAVPKSLKLQLQPPSSPTLPPGGQIT 66 (104)
T ss_pred EEEcCCCeEEEEEEEEeCCCCeeeeEEEE----EEcccceEEEEcCCCCCccCCCCCEE
Confidence 44456778899999999999999888875 22344444444434566899987633
No 63
>PF07919 Gryzun: Gryzun, putative trafficking through Golgi; InterPro: IPR012880 The proteins featured in this family are all hypothetical eukaryotic proteins of unknown function. The region in question is approximately 150 residues long.
Probab=27.38 E-value=2.1e+02 Score=29.31 Aligned_cols=40 Identities=18% Similarity=0.298 Sum_probs=32.0
Q ss_pred eEecCCceEEEEEe--CccceecCCeEEEEEEEeccCcceee
Q psy17223 101 FMMSPNKLHLEASL--DKELYYHGESIAVNVHVANNSNRTVK 140 (323)
Q Consensus 101 f~f~sG~I~L~a~L--dK~~Y~PGE~I~V~v~IdN~Ssk~Vk 140 (323)
+.+..-+-+|++.+ .+.-|+-||.+.|.+.|.|.......
T Consensus 166 i~I~p~pp~v~I~~~~~~~~~l~gE~~~i~i~I~n~e~~~~~ 207 (554)
T PF07919_consen 166 IRILPRPPKVSIKLPNHKPPALTGEFYPIPITISNNEDEEAS 207 (554)
T ss_pred EEEECCCCCeEEEeCCCCCCeEcCCEEEEEEEEEcCCCccce
Confidence 34556677777777 78889999999999999999977544
No 64
>PF11355 DUF3157: Protein of unknown function (DUF3157); InterPro: IPR021501 This family of proteins with unknown function appears to be restricted to Gammaproteobacteria.
Probab=26.79 E-value=1.3e+02 Score=27.93 Aligned_cols=39 Identities=18% Similarity=0.370 Sum_probs=30.3
Q ss_pred hhhccceeeEEeecceEEEEEEEecCCcceEeeEEeeeEEEee
Q psy17223 216 EKSKKKYLFLYYHGESIAVNVHVANNSNRTVKKIKVSDICLFS 258 (323)
Q Consensus 216 ~~s~~k~~~~yyhge~i~v~v~v~N~s~k~vkkikv~dv~l~s 258 (323)
..+++. -+|.|....+...++|+|++.|.-|.+ +|.||.
T Consensus 101 dV~l~~---~~y~~~~L~l~~~ltnqSsqsVv~Vel-~v~l~d 139 (199)
T PF11355_consen 101 DVSLGA---SQYEDGQLGLPFSLTNQSSQSVVLVEL-EVTLFD 139 (199)
T ss_pred eEEEec---cceeCCeEEEEEEEecCCCceEEEEEE-EEEEEc
Confidence 444543 456666999999999999999988876 588883
No 65
>PF07705 CARDB: CARDB; InterPro: IPR011635 The APHP (acidic peptide-dependent hydrolases/peptidase) domain is found in a variety of different proteins.; PDB: 2KUT_A 2L0D_A 3IDU_A 2KL6_A.
Probab=26.76 E-value=2e+02 Score=21.66 Aligned_cols=34 Identities=24% Similarity=0.427 Sum_probs=24.1
Q ss_pred EeecceEEEEEEEecCCcceEeeEEeeeEEEeeCCeE
Q psy17223 226 YYHGESIAVNVHVANNSNRTVKKIKVSDICLFSTAQY 262 (323)
Q Consensus 226 yyhge~i~v~v~v~N~s~k~vkkikv~dv~l~s~~~y 262 (323)
.+=|+++.|.+.|.|.-......++|. +|.++.-
T Consensus 15 ~~~g~~~~i~~~V~N~G~~~~~~~~v~---~~~~~~~ 48 (101)
T PF07705_consen 15 VVPGEPVTITVTVKNNGTADAENVTVR---LYLDGNS 48 (101)
T ss_dssp EETTSEEEEEEEEEE-SSS-BEEEEEE---EEETTEE
T ss_pred ccCCCEEEEEEEEEECCCCCCCCEEEE---EEECCce
Confidence 345999999999999988877776665 5555554
No 66
>cd08546 cohesin_like Cohesin domain, interaction parter of dockerin. Bacterial cohesin domains bind to a complementary protein domain named dockerin, and this interaction is required for the formation of the cellulosome, a cellulose-degrading complex. The cellulosome consists of scaffoldin, a noncatalytic scaffolding polypeptide, that comprises repeating cohesion modules and a single carbohydrate-binding module (CBM). Specific calcium-dependent interactions between cohesins and dockerins appear to be essential for cellulosome assembly. Cohesin modules are phylogenetically distributed into three groups: type I cohesin-dockerin interactions mediate assembly of a range of dockerin-borne enzymes to the complex, while type-II interactions mediate attachment of the cellulosome complex to the bacterial cell wall. Recently discovered type-III cohesins, such as found in the anchoring scaffoldin ScaE, appears to contribute to increased stability of the elaborate cellulosome complex. While the p
Probab=25.50 E-value=1.6e+02 Score=23.78 Aligned_cols=37 Identities=22% Similarity=0.215 Sum_probs=29.0
Q ss_pred EEEEeCccceecCCeEEEEEEEeccCcceeeEEEeecCCC
Q psy17223 110 LEASLDKELYYHGESIAVNVHVANNSNRTVKKIKVSDSGA 149 (323)
Q Consensus 110 L~a~LdK~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~ 149 (323)
+.+..+... .+||++.|.+.++|.+ .+..+.+.+.=+
T Consensus 3 ~~~~~~~~~-~~G~~~~v~v~~~~~~--~~~~~~~~l~yD 39 (135)
T cd08546 3 VSLGAPSTV-KVGETVTVTVKVNNVP--NVAAADFTLSYD 39 (135)
T ss_pred EEEeccccc-cCCCEEEEEEEEecCC--CeEEEEEEEEEC
Confidence 445555555 8999999999999998 888888877654
No 67
>PF11797 DUF3324: Protein of unknown function C-terminal (DUF3324); InterPro: IPR021759 This family consists of several hypothetical bacterial proteins of unknown function.
Probab=24.64 E-value=1.5e+02 Score=25.38 Aligned_cols=66 Identities=14% Similarity=0.143 Sum_probs=43.8
Q ss_pred eEEEEEEEecCCcceEeeEEeeeEEEeeCCeEeeeecccc-CCCccCCCCcccccceeec-ccccccccCC
Q psy17223 231 SIAVNVHVANNSNRTVKKIKVSDICLFSTAQYKCTVAETE-SDCPIAPVSMFDTEDLAML-RHGFKRMFGH 299 (323)
Q Consensus 231 ~i~v~v~v~N~s~k~vkkikv~dv~l~s~~~y~~~Va~~e-~~~~i~p~~t~~~~~~~~~-~~~~~~~~~~ 299 (323)
.-.|.+.|.|.....++++++. ..++..++ .+++...+ .+-.++|+|.|. ..|.+. ..|+..+|.+
T Consensus 43 ~~~i~~~l~N~~~~~l~~~~v~-a~V~~~~~-~k~~~~~~~~~~~mAPNS~f~-~~i~~~~~~lk~G~Y~l 110 (140)
T PF11797_consen 43 RNVIQANLQNPQPAILKKLTVD-AKVTKKGS-KKVLYTFKKENMQMAPNSNFN-FPIPLGGKKLKPGKYTL 110 (140)
T ss_pred eeEEEEEEECCCchhhcCcEEE-EEEEECCC-CeEEEEeeccCCEECCCCeEE-eEecCCCcCccCCEEEE
Confidence 4457778999999999999885 45554444 34555544 466999999986 344443 3555555543
No 68
>PF02019 WIF: WIF domain; InterPro: IPR003306 Wnt proteins constitute a large family of secreted molecules that are involved in intercellular signalling during development. The name derives from the first 2 members of the family to be discovered: int-1 (mouse) and wingless (Drosophila) []. It is now recognised that Wnt signalling controls many cell fate decisions in a variety of different organisms, including mammals []. Wnt signalling has been implicated in tumourigenesis, early mesodermal patterning of the embryo, morphogenesis of the brain and kidneys, regulation of mammary gland proliferation and Alzheimer's disease [, ]. Wnt-mediated signalling is believed to proceed initially through binding to cell surface receptors of the frizzled family; the signal is subsequently transduced through several cytoplasmic components to B-catenin, which enters the nucleus and activates the transcription of several genes important in development []. Several non-canonical Wnt signalling pathways have also been elucidated that act independently of B-catenin. Canonical and noncanonical Wnt signaling branches are highly interconnected, and cross-regulate each other []. Members of the Wnt gene family are defined by their sequence similarity to mouse Wnt-1 and Wingless in Drosophila. They encode proteins of ~350-400 residues in length, with orthologues identified in several, mostly vertebrate, species. Very little is known about the structure of Wnts as they are notoriously insoluble, but they share the following features characteristics of secretory proteins: a signal peptide, several potential N-glycosylation sites and 22 conserved cysteines [] that are probably involved in disulphide bonds. The Wnt proteins seem to adhere to the plasma membrane of the secreting cells and are therefore likely to signal over only few cell diameters. Fifteen major Wnt gene families have been identified in vertebrates, with multiple subtypes within some classes. This entry represents the WIF domain, and is found in the RYK tyrosine kinase receptors and WIF the Wnt-inhibitory-factor. The domain is extracellular and contains two conserved cysteines that may form a disulphide bridge. This domain is Wnt binding in WIF, and it has been suggested that RYK may also bind to Wnt [].; GO: 0004713 protein tyrosine kinase activity; PDB: 2YGP_A 2YGO_A 2YGN_A 2D3J_A 2YGQ_A.
Probab=23.99 E-value=2.9e+02 Score=23.89 Aligned_cols=43 Identities=9% Similarity=0.159 Sum_probs=25.5
Q ss_pred cCCceEEEEEeCccceecCC-eEEEEEEEeccCcceeeEEEeec
Q psy17223 104 SPNKLHLEASLDKELYYHGE-SIAVNVHVANNSNRTVKKIKVSD 146 (323)
Q Consensus 104 ~sG~I~L~a~LdK~~Y~PGE-~I~V~v~IdN~Ssk~Vk~Ikv~L 146 (323)
...+-..++.|+-.|...|| ++.|+++|.+.+++..+.+.++.
T Consensus 85 P~~~~~F~V~L~CtG~~~g~a~v~i~lni~~~~~~n~T~L~~kr 128 (132)
T PF02019_consen 85 PHSPQVFSVELPCTGKRSGEATVTIQLNITLPSSKNGTPLRFKR 128 (132)
T ss_dssp -SS-EEEEEE--B-SSS-EEEEEEEEEEEEETT--S-EEEE--T
T ss_pred cCCCEEEEEEEEecCccceEEEEEEEEEEEeCCCCCceEEEEee
Confidence 34466778889999999998 78999999999987666555543
No 69
>PF10437 Lip_prot_lig_C: Bacterial lipoate protein ligase C-terminus; InterPro: IPR019491 This is the C-terminal domain of a bacterial lipoate protein ligase. There is no conservation between this C terminus and that of vertebrate lipoate protein ligase C-termini, but both are associated with IPR004143 from INTERPRO, further upstream. This C-terminal domain is more stable than IPR004143 from INTERPRO and the hypothesis is that the C-terminal domain has a role in recognising the lipoyl domain and/or transferring the lipoyl group onto it from the lipoyl-AMP intermediate. C-terminal fragments of length 172 to 193 amino acid residues are observed in the eubacterial enzymes whereas in their archaeal counterparts the C-terminal segment is significantly smaller, ranging in size from 87 to 107 amino acid residues. ; PDB: 1X2G_A 3A7R_A 3A7A_A 1X2H_C 1VQZ_A 3R07_C.
Probab=23.31 E-value=1.8e+02 Score=22.41 Aligned_cols=27 Identities=22% Similarity=0.288 Sum_probs=21.9
Q ss_pred eeEEeecceEEEEEEEecCCcceEeeEEee
Q psy17223 223 LFLYYHGESIAVNVHVANNSNRTVKKIKVS 252 (323)
Q Consensus 223 ~~~yyhge~i~v~v~v~N~s~k~vkkikv~ 252 (323)
+..++.|-.|.|+++|.|. .|+.|++.
T Consensus 9 ~~~rf~~G~v~v~~~V~~G---~I~~i~i~ 35 (86)
T PF10437_consen 9 KERRFPWGTVEVHLNVKNG---IIKDIKIY 35 (86)
T ss_dssp EEEEETTEEEEEEEEEETT---EEEEEEEE
T ss_pred eeeEcCCceEEEEEEEECC---EEEEEEEE
Confidence 3678999999999999888 66666665
No 70
>PF04744 Monooxygenase_B: Monooxygenase subunit B protein; InterPro: IPR006833 Ammonia monooxygenase and the particulate methane monooxygenase are both integral membrane proteins, occurring in ammonia oxidisers and methanotrophs respectively, which are thought to be evolutionarily related []. These enzymes have a relatively wide substrate specificity and can catalyse the oxidation of a range of substrates including ammonia, methane, halogenated hydrocarbons and aromatic molecules []. These enzymes are composed of 3 subunits - A (IPR003393 from INTERPRO), B (IPR006833 from INTERPRO) and C (IPR006980 from INTERPRO) - and contain various metal centres, including copper. Particulate methane monooxygenase from Methylococcus capsulatus str. Bath is an ABC homotrimer, which contains mononuclear and dinuclear copper metal centres, and a third metal centre containing a metal ion whose identity in vivo is not certain[]. The soluble regions of these enzymes derive primarily from the B subunit. This subunit forms two antiparallel beta-barrel-like structures and contains the mono- and di- nuclear copper metal centres [].; PDB: 3CHX_E 3RFR_A 3RGB_A 1YEW_A.
Probab=23.11 E-value=1.1e+02 Score=31.04 Aligned_cols=67 Identities=21% Similarity=0.342 Sum_probs=34.7
Q ss_pred cchhhccceeeEEe-ecceEEEEEEEecCCcceEeeEEee--eEEEeeCC------eEee-eecc----ccCCCccCCCC
Q psy17223 214 SKEKSKKKYLFLYY-HGESIAVNVHVANNSNRTVKKIKVS--DICLFSTA------QYKC-TVAE----TESDCPIAPVS 279 (323)
Q Consensus 214 ~~~~s~~k~~~~yy-hge~i~v~v~v~N~s~k~vkkikv~--dv~l~s~~------~y~~-~Va~----~e~~~~i~p~~ 279 (323)
.++.-..+ |.|. -|..+.+++.|+||+++-|+==... +|..-.-+ .|-. .+|. +....||+||.
T Consensus 248 ~V~~~v~~--A~Y~vpgR~l~~~l~VtN~g~~pv~LgeF~tA~vrFln~~v~~~~~~~P~~l~A~~gL~vs~~~pI~PGE 325 (381)
T PF04744_consen 248 SVKVKVTD--ATYRVPGRTLTMTLTVTNNGDSPVRLGEFNTANVRFLNPDVPTDDPDYPDELLAERGLSVSDNSPIAPGE 325 (381)
T ss_dssp SEEEEEEE--EEEESSSSEEEEEEEEEEESSS-BEEEEEESSS-EEE-TTT-SS-S---TTTEETT-EEES--S-B-TT-
T ss_pred ceEEEEec--cEEecCCcEEEEEEEEEcCCCCceEeeeEEeccEEEeCcccccCCCCCchhhhccCcceeCCCCCcCCCc
Confidence 34444555 6665 6899999999999999987644443 33332111 1111 1332 23345999998
Q ss_pred ccc
Q psy17223 280 MFD 282 (323)
Q Consensus 280 t~~ 282 (323)
|-.
T Consensus 326 Trt 328 (381)
T PF04744_consen 326 TRT 328 (381)
T ss_dssp EEE
T ss_pred eEE
Confidence 744
No 71
>TIGR03780 Bac_Flav_CT_N Bacteroides conjugative transposon TraN protein. Members of this family are the TraN protein encoded by transfer region genes of conjugative transposons of Bacteroides. The family is related to conjugative transfer proteins VirB9 and TrbG of Agrobacterium Ti plasmids.
Probab=23.04 E-value=88 Score=30.57 Aligned_cols=20 Identities=20% Similarity=0.399 Sum_probs=19.4
Q ss_pred eEEeecceEEEEEEEecCCc
Q psy17223 224 FLYYHGESIAVNVHVANNSN 243 (323)
Q Consensus 224 ~~yyhge~i~v~v~v~N~s~ 243 (323)
+||.||+-+-+++.+.|+||
T Consensus 175 ~Iy~~~d~lyf~~~l~N~Sn 194 (285)
T TIGR03780 175 GIYTHNDLLYFHTSLENKTN 194 (285)
T ss_pred eEEEECCEEEEEEEEEcCCC
Confidence 89999999999999999987
No 72
>PF12389 Peptidase_M73: Camelysin metallo-endopeptidase; InterPro: IPR022121 Camelysin is a novel surface metallopeptidase from Bacillus cereus []. Camelysin prefers cleavage sites in front of aliphatic and hydrophilic amino acid residues (-OH, -SO3H, amido group), and requires zinc for activity [, ].
Probab=23.03 E-value=1.3e+02 Score=27.85 Aligned_cols=32 Identities=9% Similarity=0.188 Sum_probs=27.3
Q ss_pred cceecCCeEEEEEEEeccCcceeeEEEeecCC
Q psy17223 117 ELYYHGESIAVNVHVANNSNRTVKKIKVSDSG 148 (323)
Q Consensus 117 ~~Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q 148 (323)
.-..||+++.-.+.|.|.-+-+|+.|.+...-
T Consensus 59 ~nlkPGD~v~k~f~l~N~Gtldi~~v~l~~~y 90 (199)
T PF12389_consen 59 SNLKPGDTVEKEFTLKNSGTLDIKDVLLKTDY 90 (199)
T ss_pred ccCCCCCeEEEEEEEEeCCeeeeeeEEEEEEE
Confidence 35789999999999999999999888777643
No 73
>PF04442 CtaG_Cox11: Cytochrome c oxidase assembly protein CtaG/Cox11; InterPro: IPR007533 Cytochrome c oxidase assembly protein is essential for the assembly of functional cytochrome oxidase protein. In eukaryotes it is an integral protein of the mitochondrial inner membrane. Cox11 is essential for the insertion of Cu(I) ions to form the CuB site. This is essential for the stability of other structures in subunit I, for example haems a and a3, and the magnesium/manganese centre. Cox11 is probably only required in sub-stoichiometric amounts relative to the structural units []. The C-terminal region of the protein is known to form a dimer. Each monomer coordinates one Cu(I) ion via three conserved cysteine residues (111, 208 and 210) in Saccharomyces cerevisiae (P19516 from SWISSPROT). Met 224 is also thought to play a role in copper transfer or stabilising the copper site [].; GO: 0005507 copper ion binding; PDB: 1SO9_A 1SP0_A.
Probab=21.83 E-value=97 Score=27.53 Aligned_cols=29 Identities=17% Similarity=0.234 Sum_probs=18.5
Q ss_pred eecCCeEEEEEEEeccCcceeeEEEeecC
Q psy17223 119 YYHGESIAVNVHVANNSNRTVKKIKVSDS 147 (323)
Q Consensus 119 Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~ 147 (323)
..|||...+...+.|.|++.|..+-+-=.
T Consensus 63 V~pGe~~~~~y~a~N~s~~~i~g~A~~nV 91 (152)
T PF04442_consen 63 VHPGETALVFYEATNPSDKPITGQAIPNV 91 (152)
T ss_dssp EETT--EEEEEEEEE-SSS-EE---EEEE
T ss_pred eCCCCEEEEEEEEECCCCCcEEEEEeeeE
Confidence 36999999999999999999987765443
No 74
>KOG2540|consensus
Probab=21.06 E-value=82 Score=29.96 Aligned_cols=71 Identities=15% Similarity=0.252 Sum_probs=42.8
Q ss_pred ccce-ecCCeEEEEEEEeccCcceeeEEEeecCCCccchh----hhhhhccCCCCCC-ccCCCccccccccccceeeecc
Q psy17223 116 KELY-YHGESIAVNVHVANNSNRTVKKIKVSDSGAEDDQD----LKDELADSDIDGM-EEDDLPNIKAWGKNKRMYYNTD 189 (323)
Q Consensus 116 K~~Y-~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L~q~ed~~d----~~~~~~~~di~~~-~e~~lp~~~awg~~~~~~y~td 189 (323)
+.+| .|||+...-....|.|.++|.+|.---+--.+..- .+=.+.+.-.-.. |+-||| .=-|.|.|
T Consensus 155 rEiyV~PGEtALaFYta~N~sdkpIiGvstYni~P~~Aa~YFnKiqCFCFEEQ~L~pgE~vDmP--------VFFyIDPe 226 (269)
T KOG2540|consen 155 REIYVLPGETALAFYTAENPSDKPIIGVSTYNITPGQAAVYFNKIQCFCFEEQKLNPGEQVDMP--------VFFYIDPE 226 (269)
T ss_pred eEEEEcCCcceeeeEeccCCCCCCceeeEeeccCccHhhhheeceeEEeehhhccCCCcccCcc--------eEEEeCcc
Confidence 4556 69999999999999999999888754332111110 1112222222222 445788 55677777
Q ss_pred cccCC
Q psy17223 190 YVDDD 194 (323)
Q Consensus 190 ~~~~~ 194 (323)
|+++.
T Consensus 227 fa~DP 231 (269)
T KOG2540|consen 227 FATDP 231 (269)
T ss_pred cccCc
Confidence 76443
No 75
>PF04425 Bul1_N: Bul1 N terminus; InterPro: IPR007519 This domain is the N terminus of Saccharomyces cerevisiae (Baker's yeast) Bul1. Bul1 binds the ubiquitin ligase Rsp5, via an N-terminal PPSY motif (157-160 in P48524 from SWISSPROT) []. The complex containing Bul1 and Rsp5 is involved in intracellular trafficking of the general amino acid permease Gap1 [], degradation of Rog1 in cooperation with Bul2 and GSK-3 [], and mitochondrial inheritance []. Bul1 may contain HEAT repeats. The C terminus is IPR007520 from INTERPRO.
Probab=20.87 E-value=1.3e+02 Score=31.23 Aligned_cols=44 Identities=27% Similarity=0.381 Sum_probs=32.7
Q ss_pred CCceEEEEEeCc---------------cceecCCeEEEEEEEeccCcceee--EEEeecCC
Q psy17223 105 PNKLHLEASLDK---------------ELYYHGESIAVNVHVANNSNRTVK--KIKVSDSG 148 (323)
Q Consensus 105 sG~I~L~a~LdK---------------~~Y~PGE~I~V~v~IdN~Ssk~Vk--~Ikv~L~q 148 (323)
.-+|.+++.+-| .-|.+|+.|.=-|.|+|.|++.|. =+.|.|.+
T Consensus 131 s~~l~I~I~~Tk~v~~~g~p~~id~~l~Ey~qGD~I~GyvtI~N~S~~pIpFdMFyV~lEG 191 (438)
T PF04425_consen 131 SSPLEIEIYVTKDVGKPGKPPEIDPSLKEYTQGDIIHGYVTIENTSSKPIPFDMFYVSLEG 191 (438)
T ss_pred CCceEEEEEEeccCCCCCCCcccCcccccccCCCEEEEEEEEEECCCCCcccceEEEEEEE
Confidence 456777777766 468889999999999999999875 34444443
No 76
>PF01050 MannoseP_isomer: Mannose-6-phosphate isomerase; InterPro: IPR001538 Mannose-6-phosphate isomerase or phosphomannose isomerase (5.3.1.8 from EC) (PMI) is the enzyme that catalyses the interconversion of mannose-6-phosphate and fructose-6-phosphate. In eukaryotes PMI is involved in the synthesis of GDP-mannose, a constituent of N- and O-linked glycans and GPI anchors and in prokaryotes it participates in a variety of pathways, including capsular polysaccharide biosynthesis and D-mannose metabolism. PMI's belong to the cupin superfamily whose functions range from isomerase and epimerase activities involved in the modification of cell wall carbohydrates in bacteria and plants, to non-enzymatic storage proteins in plant seeds, and transcription factors linked to congenital baldness in mammals []. Three classes of PMI have been defined []. The type II phosphomannose isomerases are bifunctional enzymes 5.3.1.8 from EC. This entry covers the isomerase region of the protein []. The guanosine diphospho-D-mannose pyrophosphorylase region is described in another InterPro entry (see IPR005836 from INTERPRO).; GO: 0016779 nucleotidyltransferase activity, 0005976 polysaccharide metabolic process
Probab=20.66 E-value=4.2e+02 Score=23.19 Aligned_cols=72 Identities=14% Similarity=0.221 Sum_probs=45.3
Q ss_pred EeeeeeecCCCCC-CCCCeEEEEEEeEecCCceEEEEEeCccceecCCeEEEEE----EEeccCcceeeEEEeecCC
Q psy17223 77 LAIRKIMYAPSKQ-GEQPSVEVSKEFMMSPNKLHLEASLDKELYYHGESIAVNV----HVANNSNRTVKKIKVSDSG 148 (323)
Q Consensus 77 l~Ir~l~~~P~~~-~~~~~~e~~k~f~f~sG~I~L~a~LdK~~Y~PGE~I~V~v----~IdN~Ssk~Vk~Ikv~L~q 148 (323)
+.++.|...|-.. ..+....-...+.+-+|.-.+.+-=....+.+||.|.|-. .|.|.++.++.=|.|+.-.
T Consensus 63 ~~vkri~V~pG~~lSlq~H~~R~E~W~Vv~G~a~v~~~~~~~~~~~g~sv~Ip~g~~H~i~n~g~~~L~~IEVq~G~ 139 (151)
T PF01050_consen 63 YKVKRITVNPGKRLSLQYHHHRSEHWTVVSGTAEVTLDDEEFTLKEGDSVYIPRGAKHRIENPGKTPLEIIEVQTGE 139 (151)
T ss_pred EEEEEEEEcCCCccceeeecccccEEEEEeCeEEEEECCEEEEEcCCCEEEECCCCEEEEECCCCcCcEEEEEecCC
Confidence 3445555555432 2232223334566777877777655555668899886643 5889988888888888854
No 77
>PRK05089 cytochrome C oxidase assembly protein; Provisional
Probab=20.63 E-value=2e+02 Score=26.44 Aligned_cols=28 Identities=21% Similarity=0.138 Sum_probs=24.2
Q ss_pred eecCCeEEEEEEEeccCcceeeEEEeec
Q psy17223 119 YYHGESIAVNVHVANNSNRTVKKIKVSD 146 (323)
Q Consensus 119 Y~PGE~I~V~v~IdN~Ssk~Vk~Ikv~L 146 (323)
..|||+..+...+.|.|.+.|...-+-=
T Consensus 90 V~pGE~~~~~y~a~N~sd~~i~g~A~~n 117 (188)
T PRK05089 90 VHPGELNLVFYEAENLSDRPIVGQAIPS 117 (188)
T ss_pred EcCCCeEEEEEEEECCCCCcEEEEEecc
Confidence 3599999999999999999998776644
No 78
>PF04205 FMN_bind: FMN-binding domain; InterPro: IPR007329 This conserved region includes the FMN-binding site of the NqrC protein [] as well as the NosR and NirI regulatory proteins.; GO: 0010181 FMN binding, 0016020 membrane; PDB: 3LWX_A 2KZX_A 3DCZ_A 3O6U_D.
Probab=20.36 E-value=1.7e+02 Score=21.89 Aligned_cols=22 Identities=23% Similarity=0.511 Sum_probs=17.0
Q ss_pred cceEEEEEEEecCCcceEeeEEee
Q psy17223 229 GESIAVNVHVANNSNRTVKKIKVS 252 (323)
Q Consensus 229 ge~i~v~v~v~N~s~k~vkkikv~ 252 (323)
|.+|.|.|.|+++ ..|..|++.
T Consensus 3 ~g~i~v~v~i~~d--g~I~~v~~~ 24 (81)
T PF04205_consen 3 GGPITVTVTIDKD--GKITDVKIL 24 (81)
T ss_dssp EEEEEEEEEEETT--TEEEEEEEE
T ss_pred CceEEEEEEEeCC--CEEEEEEEe
Confidence 4489999999886 567777776
Done!