Query 029361
Match_columns 194
No_of_seqs 95 out of 109
Neff 4.6
Searched_HMMs 46136
Date Fri Mar 29 11:41:21 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/029361.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/029361hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF05753 TRAP_beta: Translocon 100.0 3.9E-55 8.5E-60 364.5 17.4 176 10-192 2-181 (181)
2 KOG3317 Translocon-associated 100.0 1.4E-53 2.9E-58 353.2 17.5 169 22-194 18-188 (188)
3 PF07705 CARDB: CARDB; InterP 97.8 0.00022 4.8E-09 51.2 9.5 78 29-116 2-81 (101)
4 PF10633 NPCBM_assoc: NPCBM-as 97.5 0.00058 1.3E-08 48.8 7.6 68 44-116 3-73 (78)
5 PF01345 DUF11: Domain of unkn 97.4 0.00075 1.6E-08 47.8 7.0 56 26-84 21-76 (76)
6 TIGR01451 B_ant_repeat conserv 97.2 0.0013 2.9E-08 44.7 6.0 50 35-87 1-50 (53)
7 PF13473 Cupredoxin_1: Cupredo 96.4 0.012 2.6E-07 44.0 6.4 48 50-115 44-91 (104)
8 PF13584 BatD: Oxygen toleranc 95.5 0.093 2E-06 48.8 9.2 77 37-123 19-98 (484)
9 TIGR02588 conserved hypothetic 95.4 0.41 8.8E-06 38.5 11.3 86 9-105 14-104 (122)
10 COG1361 S-layer domain [Cell e 93.7 0.87 1.9E-05 42.7 11.1 85 39-126 160-248 (500)
11 COG1721 Uncharacterized conser 93.2 0.66 1.4E-05 42.9 9.2 92 26-126 48-139 (416)
12 PF07919 Gryzun: Gryzun, putat 93.1 0.93 2E-05 42.5 10.3 85 27-120 469-553 (554)
13 COG1470 Predicted membrane pro 92.5 0.59 1.3E-05 45.2 8.1 92 42-139 393-490 (513)
14 PF03896 TRAP_alpha: Transloco 92.4 1.2 2.5E-05 40.3 9.5 93 31-128 86-184 (285)
15 PF00927 Transglut_C: Transglu 91.9 0.25 5.4E-06 36.9 4.0 70 39-115 8-85 (107)
16 PF12584 TRAPPC10: Trafficking 89.3 4.2 9.2E-05 32.5 9.2 77 41-123 26-114 (147)
17 PF14874 PapD-like: Flagellar- 89.1 7 0.00015 28.4 10.0 73 37-116 11-84 (102)
18 PF12690 BsuPI: Intracellular 88.4 2.6 5.6E-05 31.0 6.8 66 49-115 2-81 (82)
19 TIGR03079 CH4_NH3mon_ox_B meth 85.7 4 8.6E-05 38.6 7.9 73 29-105 265-353 (399)
20 PF13584 BatD: Oxygen toleranc 85.6 13 0.00028 34.7 11.3 114 28-144 125-256 (484)
21 COG1361 S-layer domain [Cell e 84.6 3.1 6.8E-05 39.0 6.8 80 37-119 38-123 (500)
22 PF04744 Monooxygenase_B: Mono 83.6 12 0.00025 35.5 9.9 72 32-105 249-334 (381)
23 PF03345 DDOST_48kD: Oligosacc 80.9 20 0.00044 34.1 10.6 79 95-183 329-413 (423)
24 PF09478 CBM49: Carbohydrate b 80.4 8.6 0.00019 27.7 6.4 65 40-104 9-79 (80)
25 PF09624 DUF2393: Protein of u 79.5 28 0.0006 27.5 9.6 74 40-114 56-142 (149)
26 PRK10378 inactive ferrous ion 74.2 19 0.00042 33.8 8.4 50 51-116 53-103 (375)
27 PF05506 DUF756: Domain of unk 73.5 27 0.00059 25.3 7.5 52 50-115 21-75 (89)
28 cd04036 C2_cPLA2 C2 domain pre 70.8 11 0.00024 28.1 5.0 68 41-116 47-115 (119)
29 COG1572 Uncharacterized conser 69.7 16 0.00034 36.5 7.0 68 41-118 418-487 (606)
30 KOG4386 Uncharacterized conser 68.9 10 0.00022 37.9 5.5 73 42-120 704-776 (809)
31 KOG2291 Oligosaccharyltransfer 67.5 16 0.00035 36.2 6.5 71 1-74 1-74 (602)
32 PF02102 Peptidase_M35: Deuter 65.2 2.1 4.5E-05 39.9 0.0 59 47-105 38-115 (359)
33 PF11611 DUF4352: Domain of un 65.0 46 0.001 24.5 7.3 67 43-109 32-105 (123)
34 PF14796 AP3B1_C: Clathrin-ada 64.6 17 0.00036 29.9 5.2 49 47-102 85-136 (145)
35 PF06159 DUF974: Protein of un 63.2 51 0.0011 28.8 8.3 79 44-124 12-94 (249)
36 PF07610 DUF1573: Protein of u 62.5 28 0.00061 22.5 5.0 41 52-102 1-43 (45)
37 PF13473 Cupredoxin_1: Cupredo 61.8 11 0.00023 27.9 3.4 42 62-106 21-63 (104)
38 TIGR02656 cyanin_plasto plasto 61.3 19 0.00041 26.7 4.6 53 56-114 30-82 (99)
39 PRK02710 plastocyanin; Provisi 60.4 76 0.0016 24.4 8.7 20 91-114 83-102 (119)
40 PF13860 FlgD_ig: FlgD Ig-like 58.1 29 0.00063 24.8 5.0 39 50-97 26-64 (81)
41 KOG1691 emp24/gp25L/p24 family 57.6 40 0.00086 29.5 6.5 32 39-71 36-69 (210)
42 PF07919 Gryzun: Gryzun, putat 56.3 1.2E+02 0.0027 28.4 10.1 85 37-123 181-283 (554)
43 PF08626 TRAPPC9-Trs120: Trans 55.2 89 0.0019 33.2 9.8 98 25-124 775-898 (1185)
44 PF07760 DUF1616: Protein of u 55.0 1.3E+02 0.0028 26.7 9.5 66 40-107 185-255 (287)
45 PRK15188 fimbrial chaperone pr 54.4 1.1E+02 0.0023 26.7 8.7 59 41-104 35-96 (228)
46 PF11797 DUF3324: Protein of u 53.3 40 0.00086 26.8 5.5 53 37-95 50-106 (140)
47 TIGR03096 nitroso_cyanin nitro 53.0 1.2E+02 0.0027 24.6 9.2 23 90-114 94-116 (135)
48 PF14263 DUF4354: Domain of un 52.4 39 0.00084 27.3 5.2 93 12-108 10-110 (124)
49 TIGR02745 ccoG_rdxA_fixG cytoc 49.5 2.4E+02 0.0053 26.9 11.6 55 45-106 344-399 (434)
50 PRK14740 kdbF potassium-transp 48.1 2.1 4.5E-05 26.5 -2.0 18 163-180 4-21 (29)
51 PF04495 GRASP55_65: GRASP55/6 48.1 76 0.0016 25.5 6.4 49 54-108 2-54 (138)
52 PF00127 Copper-bind: Copper b 47.2 42 0.00092 24.6 4.5 45 55-103 29-75 (99)
53 COG1470 Predicted membrane pro 46.6 1.3E+02 0.0028 29.7 8.6 72 47-120 284-360 (513)
54 PRK15208 long polar fimbrial c 46.2 1.4E+02 0.003 25.7 8.1 52 47-103 35-89 (228)
55 PRK15290 lfpB fimbrial chapero 44.0 1.6E+02 0.0034 25.9 8.2 55 5-65 15-69 (243)
56 cd08547 Type_II_cohesin Type I 42.6 1.5E+02 0.0031 22.4 8.7 39 42-84 12-50 (132)
57 cd08379 C2D_MCTP_PRT_plant C2 42.5 89 0.0019 24.4 5.9 54 50-108 53-113 (126)
58 PF00635 Motile_Sperm: MSP (Ma 42.0 1.2E+02 0.0027 21.8 6.3 51 47-106 18-69 (109)
59 PF00630 Filamin: Filamin/ABP2 41.4 1.3E+02 0.0028 21.4 8.2 67 40-115 15-87 (101)
60 PF00207 A2M: Alpha-2-macroglo 41.2 48 0.001 24.0 3.9 37 30-68 49-90 (92)
61 PRK15098 beta-D-glucoside gluc 39.4 95 0.0021 31.5 6.9 84 47-136 667-755 (765)
62 PF14310 Fn3-like: Fibronectin 39.1 28 0.00061 24.2 2.3 24 92-115 29-52 (71)
63 cd04049 C2_putative_Elicitor-r 38.9 1E+02 0.0022 22.9 5.5 76 41-123 46-122 (124)
64 PF10731 Anophelin: Thrombin i 37.6 19 0.00042 25.9 1.3 17 13-29 7-23 (65)
65 PF03314 DUF273: Protein of un 35.4 22 0.00047 31.4 1.5 48 47-96 168-216 (222)
66 cd08678 C2_C21orf25-like C2 do 34.2 2E+02 0.0043 21.5 7.5 59 41-107 43-102 (126)
67 PF06280 DUF1034: Fn3-like dom 33.2 2E+02 0.0044 21.3 8.5 81 47-127 8-104 (112)
68 PTZ00234 variable surface prot 32.8 21 0.00045 34.2 1.0 35 159-193 364-400 (433)
69 PF08626 TRAPPC9-Trs120: Trans 32.8 1.3E+02 0.0029 31.9 7.0 77 39-123 644-722 (1185)
70 PF15012 DUF4519: Domain of un 29.7 41 0.0009 23.7 1.9 19 165-183 38-56 (56)
71 PRK06655 flgD flagellar basal 29.3 2.5E+02 0.0053 24.4 7.0 30 81-110 149-182 (225)
72 PF03896 TRAP_alpha: Transloco 29.2 4.1E+02 0.0088 24.1 8.6 28 57-87 75-102 (285)
73 PF00345 PapD_N: Pili and flag 28.8 2.5E+02 0.0054 21.0 9.0 72 47-123 14-95 (122)
74 PRK09918 putative fimbrial cha 28.6 3.8E+02 0.0082 23.0 8.7 19 44-62 35-53 (230)
75 PF08441 Integrin_alpha2: Inte 28.5 1.2E+02 0.0026 27.9 5.3 45 30-76 169-218 (457)
76 PF12112 DUF3579: Protein of u 27.5 27 0.00058 26.9 0.7 24 136-159 4-30 (92)
77 PRK13792 lysozyme inhibitor; P 26.9 1.5E+02 0.0033 23.8 4.9 15 45-61 53-67 (127)
78 TIGR02781 VirB9 P-type conjuga 26.7 2.7E+02 0.0059 24.1 6.9 22 86-111 69-90 (243)
79 PF12034 DUF3520: Domain of un 25.4 3.1E+02 0.0068 23.4 6.8 62 90-152 47-128 (183)
80 PRK10737 FKBP-type peptidyl-pr 25.0 94 0.002 26.6 3.6 59 47-113 7-72 (196)
81 PF10989 DUF2808: Protein of u 24.7 76 0.0017 25.3 2.9 27 91-117 98-126 (146)
82 smart00557 IG_FLMN Filamin-typ 24.7 2.7E+02 0.0059 20.0 7.0 60 42-115 14-73 (93)
83 PF05984 Cytomega_UL20A: Cytom 24.3 1.9E+02 0.0041 22.4 4.7 21 10-30 7-28 (100)
84 PRK11385 putativi pili assembl 23.8 4.9E+02 0.011 22.7 10.9 25 40-64 33-57 (236)
85 COG2373 Large extracellular al 23.6 2.4E+02 0.0051 31.7 7.0 84 35-123 1493-1605(1621)
86 PF12099 DUF3575: Protein of u 22.8 3.6E+02 0.0079 22.6 6.7 83 27-114 22-106 (189)
87 PF13157 DUF3992: Protein of u 22.8 2.4E+02 0.0052 21.5 5.1 46 47-92 24-71 (92)
88 PRK15211 fimbrial chaperone pr 22.8 5.1E+02 0.011 22.5 8.9 57 43-105 32-92 (229)
89 PLN02171 endoglucanase 22.7 4.1E+02 0.0089 26.8 8.0 62 44-105 550-615 (629)
90 cd04458 CSP_CDS Cold-Shock Pro 22.2 1.8E+02 0.0038 19.5 3.9 39 27-67 22-64 (65)
91 KOG3865 Arrestin [Signal trans 21.8 1.1E+02 0.0024 29.0 3.6 18 91-108 261-278 (402)
92 TIGR03102 halo_cynanin halocya 21.4 3E+02 0.0064 21.5 5.5 39 53-103 52-91 (115)
93 PF10528 PA14_2: GLEYA domain; 21.3 1.7E+02 0.0038 22.6 4.2 32 39-71 63-94 (113)
94 PF06030 DUF916: Bacterial pro 21.1 4E+02 0.0088 20.7 8.0 63 43-108 24-106 (121)
95 cd08546 cohesin_like Cohesin d 21.0 3.5E+02 0.0076 19.9 8.7 37 43-83 12-48 (135)
96 cd08373 C2A_Ferlin C2 domain f 20.5 3.7E+02 0.0079 20.0 7.6 81 41-127 38-121 (127)
97 PHA02668 GM-CSF/IL-2 inhibitio 20.1 2.4E+02 0.0052 25.5 5.2 96 11-123 5-104 (265)
98 PF00963 Cohesin: Cohesin doma 20.1 2.4E+02 0.0051 21.7 4.7 41 39-83 7-48 (141)
99 PF08441 Integrin_alpha2: Inte 20.1 1.4E+02 0.0031 27.4 4.1 31 44-76 341-371 (457)
No 1
>PF05753 TRAP_beta: Translocon-associated protein beta (TRAPB); InterPro: IPR008856 This family consists of several eukaryotic translocon-associated protein beta (TRAPB) or signal sequence receptor beta subunit (SSR-beta) proteins. The normal translocation of nascent polypeptides into the lumen of the endoplasmic reticulum (ER) is thought to be aided in part by a translocon-associated protein (TRAP) complex consisting of 4 protein subunits. The association of mature proteins with the ER and Golgi, or other intracellular locales, such as lysosomes, depends on the initial targeting of the nascent polypeptide to the ER membrane. A similar scenario must also exist for proteins destined for secretion [].; GO: 0005783 endoplasmic reticulum, 0016021 integral to membrane
Probab=100.00 E-value=3.9e-55 Score=364.55 Aligned_cols=176 Identities=36% Similarity=0.512 Sum_probs=159.0
Q ss_pred HHHHHHHHHHhhhcccCCCceEEEEeecccccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCceeeEE
Q 029361 10 ISVLIALFLISSSFASSDVPFIVAHKKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQSW 89 (194)
Q Consensus 10 ~~~lla~~~v~~~~~~~~~a~LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~~~ 89 (194)
+++++++++++.+++.+++|+|+++|+++++++++| +|++|+|+|||+|+++|+||+|+||+||+|+|++++|+++++|
T Consensus 2 ~~~~~~~l~~~~~~~~~~~a~llv~K~il~~~~v~g-~~v~V~~~iyN~G~~~A~dV~l~D~~fp~~~F~lvsG~~s~~~ 80 (181)
T PF05753_consen 2 ALFLLALLALASVAQEDSPARLLVSKQILNKYLVEG-EDVTVTYTIYNVGSSAAYDVKLTDDSFPPEDFELVSGSLSASW 80 (181)
T ss_pred hhhhHHHHHHHHhccCCCCcEEEEEEeeccccccCC-cEEEEEEEEEECCCCeEEEEEEECCCCCccccEeccCceEEEE
Confidence 455666666666788899999999999999999999 9999999999999999999999999999999999999999999
Q ss_pred EEecCCCceEEEEEEEecceeeEeeecEEEEEEcCCC-cceeeEeecCCCcceeeecCchhhhhHHHHHHHhhhchhhhh
Q 029361 90 ERLDAGGILSHSFELDAKVKGMFHGSPALITFRIPTK-AALQEAYSTPMLPLDVLAEKPTENKLELAKRLLAKYGSQISV 168 (194)
Q Consensus 90 erI~pg~nvsH~vvv~Pk~~G~fn~t~A~VtY~~se~-~~~q~a~Ss~pg~~~I~~~~~ydrkfewa~~l~~~y~~~~~v 168 (194)
||||||+|++|+|+|+|++.|+||+++|+|+|+.+++ .++|+++||+||+++|+++|+|||+| ..|+.+|..+
T Consensus 81 ~~i~pg~~vsh~~vv~p~~~G~f~~~~a~VtY~~~~~~~~~~~a~Ss~~~~~~I~~~~~~~k~f------~~~~~~w~~f 154 (181)
T PF05753_consen 81 ERIPPGENVSHSYVVRPKKSGYFNFTPAVVTYRDSEGAKELQVAYSSPPGEGDILAERDYDKKF------SSHVMDWGAF 154 (181)
T ss_pred EEECCCCeEEEEEEEeeeeeEEEEccCEEEEEECCCCCceeEEEEecCCCcceEEeccccchhh------hhhHHHHHhH
Confidence 9999999999999999999999999999999999999 77999999999999999999999999 4456777665
Q ss_pred HH---HheeeeEEEeCcCccccccccc
Q 029361 169 IS---IIVLFVYLITSPSKSAAKGSKK 192 (194)
Q Consensus 169 ~s---~~~~~v~~~~~~~~s~~~~~~~ 192 (194)
.+ .++++.|++..+|||+....||
T Consensus 155 ~~~~~~~~~~p~ll~~~sKsky~~~k~ 181 (181)
T PF05753_consen 155 AIMTLPVLLIPYLLWYSSKSKYEKSKK 181 (181)
T ss_pred HHHHHHHHHHHHHhhhhhhhhccccCC
Confidence 44 4558999999999999433354
No 2
>KOG3317 consensus Translocon-associated complex TRAP, beta subunit [Intracellular trafficking, secretion, and vesicular transport]
Probab=100.00 E-value=1.4e-53 Score=353.18 Aligned_cols=169 Identities=38% Similarity=0.643 Sum_probs=152.7
Q ss_pred hcccCCCceEEEEeecccccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEE
Q 029361 22 SFASSDVPFIVAHKKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHS 101 (194)
Q Consensus 22 ~~~~~~~a~LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~ 101 (194)
++++...++||.+|+.+|+|.|++ +|++++|+|||+|+++|+||+|+|+|||++.||||+|+++++|||||||+|++|+
T Consensus 18 a~~at~~a~ll~kk~~lnry~v~~-rd~~leY~IyNvGsspAldVtLsD~Sfpt~~FeIvkG~~~~swerIpags~vsHs 96 (188)
T KOG3317|consen 18 ASFATSEAMLLAKKATLNRYAVEA-RDVSLEYDIYNVGSSPALDVTLSDNSFPTKTFEIVKGNLSVSWERIPAGSNVSHS 96 (188)
T ss_pred hhhcccceEEEeeccchhhccccc-eeeEEEEeeEEcCCCcceeEEecCCCCCccceeeeccccccceeecCCCCceEEE
Confidence 355566699999999999999999 9999999999999999999999999999999999999999999999999999999
Q ss_pred EEEEecceeeEeeecEEEEEEcCCCcceeeEeecCCCcceeeecCchhhhhHHHHHHHhhhchhhhhHHHheeeeEEEeC
Q 029361 102 FELDAKVKGMFHGSPALITFRIPTKAALQEAYSTPMLPLDVLAEKPTENKLELAKRLLAKYGSQISVISIIVLFVYLITS 181 (194)
Q Consensus 102 vvv~Pk~~G~fn~t~A~VtY~~se~~~~q~a~Ss~pg~~~I~~~~~ydrkfewa~~l~~~y~~~~~v~s~~~~~v~~~~~ 181 (194)
+||||++.|.||+++|+|||+.+|+..+|++++|+||+|+|+++|||||+| ..|+...||..++++..+++++||+++
T Consensus 97 ivl~prv~g~f~~t~atVty~~~e~g~~~~~~ts~~~~gyila~re~~rr~--~~~~l~flgfgviv~p~t~ip~lL~~~ 174 (188)
T KOG3317|consen 97 IVLRPRVKGVFNGTPATVTYRIPEKGALQEAYTSPPGPGYILAQREPDRRF--DPRLLAFLGFGVIVIPMTVIPILLVAT 174 (188)
T ss_pred EEEeecccceeccCceEEEEEcCCCCceeEEeecCCCCcceeeecCccccc--ChhHHHHHhhhhhhhhhhheeeeEEEe
Confidence 999999999999999999999999988899999999999999999999999 225566666666777778899999998
Q ss_pred cCccc-c-cccccCC
Q 029361 182 PSKSA-A-KGSKKKR 194 (194)
Q Consensus 182 ~~~s~-~-~~~~~~~ 194 (194)
| |++ . +.+||||
T Consensus 175 s-Krrysn~~kkkk~ 188 (188)
T KOG3317|consen 175 S-KRRYSNASKKKKR 188 (188)
T ss_pred c-ccccccccccccC
Confidence 7 777 4 5555554
No 3
>PF07705 CARDB: CARDB; InterPro: IPR011635 The APHP (acidic peptide-dependent hydrolases/peptidase) domain is found in a variety of different proteins.; PDB: 2KUT_A 2L0D_A 3IDU_A 2KL6_A.
Probab=97.84 E-value=0.00022 Score=51.17 Aligned_cols=78 Identities=19% Similarity=0.339 Sum_probs=54.3
Q ss_pred ceEEEEeecccccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCce--eeEEEEecCCCceEEEEEEEe
Q 029361 29 PFIVAHKKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNI--SQSWERLDAGGILSHSFELDA 106 (194)
Q Consensus 29 a~LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~--s~~~erI~pg~nvsH~vvv~P 106 (194)
|-|.+.-......+..| ++++|+++|.|.|..+|.++.+. | ..+|.. +.....|+||+..+..+.+.+
T Consensus 2 pDL~v~~~~~~~~~~~g-~~~~i~~~V~N~G~~~~~~~~v~--------~-~~~~~~~~~~~i~~L~~g~~~~v~~~~~~ 71 (101)
T PF07705_consen 2 PDLTVSITVSPSNVVPG-EPVTITVTVKNNGTADAENVTVR--------L-YLDGNSVSTVTIPSLAPGESETVTFTWTP 71 (101)
T ss_dssp --EEE-EEEC-SEEETT-SEEEEEEEEEE-SSS-BEEEEEE--------E-EETTEEEEEEEESEB-TTEEEEEEEEEE-
T ss_pred CCEEEEEeeCCCcccCC-CEEEEEEEEEECCCCCCCCEEEE--------E-EECCceeccEEECCcCCCcEEEEEEEEEe
Confidence 34555555566777888 99999999999999998888776 3 233433 445679999999999999999
Q ss_pred cceeeEeeec
Q 029361 107 KVKGMFHGSP 116 (194)
Q Consensus 107 k~~G~fn~t~ 116 (194)
...|.|.+..
T Consensus 72 ~~~G~~~i~~ 81 (101)
T PF07705_consen 72 PSPGSYTIRV 81 (101)
T ss_dssp SS-CEEEEEE
T ss_pred CCCCeEEEEE
Confidence 9999988653
No 4
>PF10633 NPCBM_assoc: NPCBM-associated, NEW3 domain of alpha-galactosidase; InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=97.53 E-value=0.00058 Score=48.82 Aligned_cols=68 Identities=22% Similarity=0.378 Sum_probs=44.2
Q ss_pred ccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEecc---eeeEeeec
Q 029361 44 SGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHSFELDAKV---KGMFHGSP 116 (194)
Q Consensus 44 ~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~---~G~fn~t~ 116 (194)
.| +.++++.++-|.|+.++.++++.=+. |+.|+ +. .-..+...|+||++++.++.|+|-. .|.|+++.
T Consensus 3 ~G-~~~~~~~tv~N~g~~~~~~v~~~l~~--P~GW~-~~-~~~~~~~~l~pG~s~~~~~~V~vp~~a~~G~y~v~~ 73 (78)
T PF10633_consen 3 PG-ETVTVTLTVTNTGTAPLTNVSLSLSL--PEGWT-VS-ASPASVPSLPPGESVTVTFTVTVPADAAPGTYTVTV 73 (78)
T ss_dssp TT-EEEEEEEEEE--SSS-BSS-EEEEE----TTSE-----EEEEE--B-TTSEEEEEEEEEE-TT--SEEEEEEE
T ss_pred CC-CEEEEEEEEEECCCCceeeEEEEEeC--CCCcc-cc-CCccccccCCCCCEEEEEEEEECCCCCCCceEEEEE
Confidence 46 99999999999999999999987432 56666 22 2234555999999999999999643 58888763
No 5
>PF01345 DUF11: Domain of unknown function DUF11; InterPro: IPR001434 This group of sequences is represented by a conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis (Streptococcus faecalis), and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydia trachomatis outer membrane proteins. In C. trachomatis, three cysteine-rich proteins (also believed to be lipoproteins), MOMP, OMP6 and OMP3, make up the extracellular matrix of the outer membrane []. They are involved in the essential structural integrity of both the elementary body (EB) and recticulate body (RB) phase. They are thought to be involved in porin formation and, as these bacteria lack the peptidoglycan layer common to most Gram-negative microbes, such proteins are highly important in the pathogenicity of the organism.; GO: 0005727 extrachromosomal circular DNA
Probab=97.42 E-value=0.00075 Score=47.83 Aligned_cols=56 Identities=21% Similarity=0.399 Sum_probs=48.4
Q ss_pred CCCceEEEEeecccccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCc
Q 029361 26 SDVPFIVAHKKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGN 84 (194)
Q Consensus 26 ~~~a~LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~ 84 (194)
...+.+.+.|......+..| +.+++++++-|.|+.+|.+|.|.|. + +..+++++|+
T Consensus 21 ~~~~~~~~~k~~~~~~~~~G-d~v~ytitvtN~G~~~a~nv~v~D~-l-p~g~~~v~~S 76 (76)
T PF01345_consen 21 VAIPDLSITKTVNPSTANPG-DTVTYTITVTNTGPAPATNVVVTDT-L-PAGLTFVSGS 76 (76)
T ss_pred cCCCCEEEEEecCCCcccCC-CEEEEEEEEEECCCCeeEeEEEEEc-C-CCCCEEeCCC
Confidence 34466999999999999999 9999999999999999999999996 5 5567777774
No 6
>TIGR01451 B_ant_repeat conserved repeat domain. This model represents the conserved region of about 53 amino acids shared between regions, usually repeated, of proteins from a small number of phylogenetically distant prokaryotes. Examples include a 132-residue region found repeated in three of the five longest proteins of Bacillus anthracis, a 131-residue repeat in a cell wall-anchored protein of Enterococcus faecalis, and a 120-residue repeat in Methanobacterium thermoautotrophicum. A similar region is found in some Chlamydial outer membrane proteins.
Probab=97.20 E-value=0.0013 Score=44.71 Aligned_cols=50 Identities=24% Similarity=0.359 Sum_probs=40.6
Q ss_pred eecccccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCceee
Q 029361 35 KKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQ 87 (194)
Q Consensus 35 K~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~ 87 (194)
|......+..| ..++.++++-|.|..+|.+|.|+|. .| +.++.++|+++.
T Consensus 1 Kt~d~~~~~~G-d~v~Yti~v~N~g~~~a~~v~v~D~-lP-~g~~~v~~S~~~ 50 (53)
T TIGR01451 1 KTVDKTVATIG-DTITYTITVTNNGNVPATNVVVTDI-LP-SGTTFVSNSVTV 50 (53)
T ss_pred CccCccccCCC-CEEEEEEEEEECCCCceEeEEEEEc-CC-CCCEEEeCcEEE
Confidence 44555667788 9999999999999999999999985 44 557788887653
No 7
>PF13473 Cupredoxin_1: Cupredoxin-like domain; PDB: 1IBZ_D 1IC0_E 1IBY_D.
Probab=96.42 E-value=0.012 Score=44.00 Aligned_cols=48 Identities=15% Similarity=0.236 Sum_probs=29.4
Q ss_pred EEEEEEEecCCcceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEecceeeEeee
Q 029361 50 SVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHSFELDAKVKGMFHGS 115 (194)
Q Consensus 50 tV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~~G~fn~t 115 (194)
.|+.++-|.|+.+ .++.+.| +.. ...|+||++.+.+| +|.+.|.|.|.
T Consensus 44 ~v~l~~~N~~~~~-h~~~i~~--------------~~~-~~~l~~g~~~~~~f--~~~~~G~y~~~ 91 (104)
T PF13473_consen 44 PVTLTFTNNDSRP-HEFVIPD--------------LGI-SKVLPPGETATVTF--TPLKPGEYEFY 91 (104)
T ss_dssp EEEEEEEE-SSS--EEEEEGG--------------GTE-EEEE-TT-EEEEEE--EE-S-EEEEEB
T ss_pred eEEEEEEECCCCc-EEEEECC--------------Cce-EEEECCCCEEEEEE--cCCCCEEEEEE
Confidence 3445677998876 6666655 222 27899999986665 79999999875
No 8
>PF13584 BatD: Oxygen tolerance
Probab=95.49 E-value=0.093 Score=48.79 Aligned_cols=77 Identities=22% Similarity=0.366 Sum_probs=53.7
Q ss_pred cccccccccceeEEEEEEEEecCCcceeeeEEecCCCCC-CCeeeecCceeeEEEEecCC--CceEEEEEEEecceeeEe
Q 029361 37 ASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQ-DKFDVISGNISQSWERLDAG--GILSHSFELDAKVKGMFH 113 (194)
Q Consensus 37 i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~-e~Felv~G~~s~~~erI~pg--~nvsH~vvv~Pk~~G~fn 113 (194)
+..+.+..| +.+++++++.+-|+ +..+|+ ++|++.+.+.+.+..-+.-. ...+..|++.|++.|.|.
T Consensus 19 vd~~~v~~g-e~~~l~i~~~~~~~---------~~~~p~l~~f~v~~~~~s~~~~~inG~~~~~~~~~~~l~p~~~G~~~ 88 (484)
T PF13584_consen 19 VDRNEVGLG-ETFQLTITINGDGD---------DPDLPELDGFEVLGPSQSSSTSIINGKVSSSTTYTYTLQPKKTGTFT 88 (484)
T ss_pred ECCcEEcCC-CEEEEEEEEecCcc---------cCCCCCCCCeEEcceEEEEEEEEecCceEEEEEEEEEEEecccceEE
Confidence 455677788 99999999976332 233444 88998444455555444322 236778899999999999
Q ss_pred eecEEEEEEc
Q 029361 114 GSPALITFRI 123 (194)
Q Consensus 114 ~t~A~VtY~~ 123 (194)
+.++.|++..
T Consensus 89 IP~~~v~v~G 98 (484)
T PF13584_consen 89 IPPFTVEVDG 98 (484)
T ss_pred EceEEEEECC
Confidence 9999997643
No 9
>TIGR02588 conserved hypothetical protein TIGR02588. The function of this protein is unknown. It is always found as part of a two-gene operon with TIGR02587, a protein that appears to span the membrane seven times. It is found in Nostoc sp. PCC 7120, Agrobacterium tumefaciens, Sinorhizobium meliloti, and Gloeobacter violaceus, so far, all of which are bacterial.
Probab=95.41 E-value=0.41 Score=38.51 Aligned_cols=86 Identities=16% Similarity=0.254 Sum_probs=55.3
Q ss_pred HHHHHHHHHHHhhhcccCCCceEEEEeecccccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecC-----
Q 029361 9 LISVLIALFLISSSFASSDVPFIVAHKKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISG----- 83 (194)
Q Consensus 9 ~~~~lla~~~v~~~~~~~~~a~LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G----- 83 (194)
++.+++++++--..++.+..|.|.+...= ..+-+. ...-|.++|-|-|+.+|-.|++.-. +-.|
T Consensus 14 ill~viglv~y~~l~~~~~pp~l~v~~~~-~~r~~~--gqyyVpF~V~N~gg~TAasV~V~ge--------L~~~~~v~E 82 (122)
T TIGR02588 14 ILAAMFGLVAYDWLRYSNKAAVLEVAPAE-VERMQT--GQYYVPFAIHNLGGTTAAAVNIRGE--------LRQAGAVVE 82 (122)
T ss_pred HHHHHHHHHHHHhhccCCCCCeEEEeehh-eeEEeC--CEEEEEEEEEeCCCcEEEEEEEEEE--------EccCCceeE
Confidence 44444444444446788888988777622 233333 4799999999999999999998753 2222
Q ss_pred ceeeEEEEecCCCceEEEEEEE
Q 029361 84 NISQSWERLDAGGILSHSFELD 105 (194)
Q Consensus 84 ~~s~~~erI~pg~nvsH~vvv~ 105 (194)
+-..++|=||-|+..+-.++-+
T Consensus 83 ~~e~tiDfl~g~e~~~G~~IF~ 104 (122)
T TIGR02588 83 NAEVTIDYLASGSKENGTLIFR 104 (122)
T ss_pred EeeEEEEEcCCCCeEeEEEEEc
Confidence 3466777777776655444443
No 10
>COG1361 S-layer domain [Cell envelope biogenesis, outer membrane]
Probab=93.71 E-value=0.87 Score=42.71 Aligned_cols=85 Identities=22% Similarity=0.289 Sum_probs=67.9
Q ss_pred cccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCce-eeEEEEecCCCceEEEEEEEec---ceeeEee
Q 029361 39 LKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNI-SQSWERLDAGGILSHSFELDAK---VKGMFHG 114 (194)
Q Consensus 39 ~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~-s~~~erI~pg~nvsH~vvv~Pk---~~G~fn~ 114 (194)
...+..| +..+++++|.|.|+.+|.++.|...+ |...+.-+.+.. ...++-|.||+++.-++.+... ..|.|..
T Consensus 160 ~~~i~~G-~~~~l~~~I~N~G~~~~~~v~l~~~~-~~~~~~~i~~~~~~~~i~~l~p~es~~v~f~v~~~~~a~~g~y~i 237 (500)
T COG1361 160 PEAIIPG-ETNTLTLTIKNPGEGPAKNVSLSLES-PTSYLGPIYSANDTPYIGALGPGESVNVTFSVYAGSNAEPGTYTI 237 (500)
T ss_pred ccccCCC-CccEEEEEEEeCCcccccceEEEEeC-CcceeccccccccceeeeeeCCCceEEEEEEEEeecCCCCccEEE
Confidence 4455677 77799999999999999999999864 555566666666 6889999999999999999987 5777776
Q ss_pred ecEEEEEEcCCC
Q 029361 115 SPALITFRIPTK 126 (194)
Q Consensus 115 t~A~VtY~~se~ 126 (194)
. ..++|+..+.
T Consensus 238 ~-i~i~~~~~~~ 248 (500)
T COG1361 238 N-LEITYKDEEG 248 (500)
T ss_pred E-EEEEEecCCc
Confidence 4 6788888443
No 11
>COG1721 Uncharacterized conserved protein (some members contain a von Willebrand factor type A (vWA) domain) [General function prediction only]
Probab=93.15 E-value=0.66 Score=42.88 Aligned_cols=92 Identities=20% Similarity=0.298 Sum_probs=65.2
Q ss_pred CCCceEEEEeecccccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEEEEEE
Q 029361 26 SDVPFIVAHKKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHSFELD 105 (194)
Q Consensus 26 ~~~a~LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~ 105 (194)
...+.+-+.+.+....+.+| +++++++.+-| .+.-.+.+.|+ +|++.+ .+.|.-. ..-.+.+|+. ..+.+.
T Consensus 48 ~~~~~~~v~r~~~~~~~~~g-~~~~v~~~v~~---r~~~~~~~~~~-~~~~~~-~~~~~~~-~~~~~~~~~~--~~~~~~ 118 (416)
T COG1721 48 RSLPGARVERSLEKRRLFAG-EEVEVTLRVRN---RGRPRLLLVDD-IPPSFL-GVEGTEE-VSLRLGPGER--VAYKVT 118 (416)
T ss_pred hcccceEeeccccccccccC-ccceeEEEEEe---cCccceEeeec-cCCccc-ccccCcc-eeeccCCCce--EEEEEe
Confidence 44456777887765558888 99999999999 33444556653 666644 4444322 2234555555 999999
Q ss_pred ecceeeEeeecEEEEEEcCCC
Q 029361 106 AKVKGMFHGSPALITFRIPTK 126 (194)
Q Consensus 106 Pk~~G~fn~t~A~VtY~~se~ 126 (194)
|.+-|.|.+.+..+.....-+
T Consensus 119 ~~~rG~~~~~~v~~~~~~~~g 139 (416)
T COG1721 119 PLRRGEYRLPPVRVRAEDPFG 139 (416)
T ss_pred cccCCcccccceEEEccCccc
Confidence 999999999999999887654
No 12
>PF07919 Gryzun: Gryzun, putative trafficking through Golgi; InterPro: IPR012880 The proteins featured in this family are all hypothetical eukaryotic proteins of unknown function. The region in question is approximately 150 residues long.
Probab=93.12 E-value=0.93 Score=42.46 Aligned_cols=85 Identities=20% Similarity=0.206 Sum_probs=65.2
Q ss_pred CCceEEEEeecccccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEe
Q 029361 27 DVPFIVAHKKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHSFELDA 106 (194)
Q Consensus 27 ~~a~LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~P 106 (194)
..++++++- ..+...| ..++++|+|.|- +.-..+.++.=+ ++|+| +.+|.-+.++- |.|++.-+-.|.+.|
T Consensus 469 ~~~~v~~~~---p~~~~~~-~~~~l~~~I~N~-T~~~~~~~~~me--~s~~F-~fsG~k~~~~~-llP~s~~~~~y~l~p 539 (554)
T PF07919_consen 469 SPLRVLASV---PPSAIVG-EPFTLSYTIENP-TNHFQTFELSME--PSDDF-MFSGPKQTTFS-LLPFSRHTVRYNLLP 539 (554)
T ss_pred CCcEEEEec---CCccccC-cEEEEEEEEECC-CCccEEEEEEEc--cCCCE-EEECCCcCceE-ECCCCcEEEEEEEEE
Confidence 344555554 6777888 999999999994 445555555432 45669 99999888887 999999999999999
Q ss_pred cceeeEeeecEEEE
Q 029361 107 KVKGMFHGSPALIT 120 (194)
Q Consensus 107 k~~G~fn~t~A~Vt 120 (194)
...|...+..=.|.
T Consensus 540 l~~G~~~lP~l~v~ 553 (554)
T PF07919_consen 540 LVAGWWILPRLKVR 553 (554)
T ss_pred ccCCcEECCcEEEe
Confidence 99999987765553
No 13
>COG1470 Predicted membrane protein [Function unknown]
Probab=92.54 E-value=0.59 Score=45.24 Aligned_cols=92 Identities=18% Similarity=0.291 Sum_probs=62.6
Q ss_pred ccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeee-ecCceeeEEEEecCCCceEEEEEEE-ec--ceeeEeeecE
Q 029361 42 LKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDV-ISGNISQSWERLDAGGILSHSFELD-AK--VKGMFHGSPA 117 (194)
Q Consensus 42 ~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Fel-v~G~~s~~~erI~pg~nvsH~vvv~-Pk--~~G~fn~t~A 117 (194)
+-.| ++.++...|.|.|+.|=.||.|+=+ =|++ |++ |+++ +++.|+||++.+-.++++ |. ..|-|..+-.
T Consensus 393 ~taG-ee~~i~i~I~NsGna~LtdIkl~v~-~Pqg-Wei~Vd~~---~I~sL~pge~~tV~ltI~vP~~a~aGdY~i~i~ 466 (513)
T COG1470 393 ITAG-EEKTIRISIENSGNAPLTDIKLTVN-GPQG-WEIEVDES---TIPSLEPGESKTVSLTITVPEDAGAGDYRITIT 466 (513)
T ss_pred ecCC-ccceEEEEEEecCCCccceeeEEec-CCcc-ceEEECcc---cccccCCCCcceEEEEEEcCCCCCCCcEEEEEE
Confidence 4567 9999999999999999999999865 3443 654 3332 899999999999999999 43 4566655444
Q ss_pred EEEEEcCCCcce--eeEeecCCCc
Q 029361 118 LITFRIPTKAAL--QEAYSTPMLP 139 (194)
Q Consensus 118 ~VtY~~se~~~~--q~a~Ss~pg~ 139 (194)
..+=..+.++.+ .++-||.-+-
T Consensus 467 ~ksDq~s~e~tlrV~V~~sS~st~ 490 (513)
T COG1470 467 AKSDQASSEDTLRVVVGQSSTSTY 490 (513)
T ss_pred EeeccccccceEEEEEeccccchh
Confidence 433333333322 2444555443
No 14
>PF03896 TRAP_alpha: Translocon-associated protein (TRAP), alpha subunit; InterPro: IPR005595 The alpha-subunit of the TRAP complex (TRAP alpha) is a single-spanning membrane protein of the endoplasmic reticulum (ER) which is found in proximity of nascent polypeptide chains translocating across the membrane [].; GO: 0005783 endoplasmic reticulum
Probab=92.43 E-value=1.2 Score=40.29 Aligned_cols=93 Identities=14% Similarity=0.249 Sum_probs=68.2
Q ss_pred EEEEeecccccccccceeEEEEEEEEecCCcceeeeEEecCCCC-CCCeeeecCce-eeEE-EEecCCCceEEEEEEEe-
Q 029361 31 IVAHKKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWP-QDKFDVISGNI-SQSW-ERLDAGGILSHSFELDA- 106 (194)
Q Consensus 31 LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp-~e~Felv~G~~-s~~~-erI~pg~nvsH~vvv~P- 106 (194)
++.-|. ...++.| +...+-+.+-|-|+ ..+.|...+-+|. +++|...==++ ...+ -.|+||+..|-.|...|
T Consensus 86 ~~F~~~--~~~l~aG-~~~~~LvgftN~g~-~~~~V~~i~aSl~~p~d~~~~iqNfTa~~y~~~V~pg~~aT~~YsF~~~ 161 (285)
T PF03896_consen 86 ILFPKP--TKKLPAG-EPVKFLVGFTNKGS-EPFTVESIEASLRYPQDYSYYIQNFTAVRYNREVPPGEEATFPYSFTPS 161 (285)
T ss_pred EEeccc--cccccCC-CeEEEEEEEEeCCC-CCEEEEEEeeeecCccccceEEEeecccccCcccCCCCeEEEEEEEecc
Confidence 444454 4667777 99999999999999 5899999998884 56665543333 2222 36899999999999997
Q ss_pred --cceeeEeeecEEEEEEcCCCcc
Q 029361 107 --KVKGMFHGSPALITFRIPTKAA 128 (194)
Q Consensus 107 --k~~G~fn~t~A~VtY~~se~~~ 128 (194)
-..+.|.+.-. +.|+..++..
T Consensus 162 ~~l~pr~f~L~i~-l~y~d~~g~~ 184 (285)
T PF03896_consen 162 EELAPRPFGLVIN-LIYEDSDGNQ 184 (285)
T ss_pred hhcCCcceEEEEE-EEEEeCCCCE
Confidence 45677887774 5598887754
No 15
>PF00927 Transglut_C: Transglutaminase family, C-terminal ig like domain; InterPro: IPR008958 Synonym(s): Protein-glutamine gamma-glutamyltransferase, Fibrinoligase, TGase Transglutaminases catalyse the post-translational modification of proteins at glutamine residues, with formation of isopeptide bonds. Members of the transglutaminase family usually have three domains: N-terminal (IPR001102 from INTERPRO), middle (IPR013808 from INTERPRO) and C-terminal. The middle domain is usually well conserved, but family members can display major differences in their N- and C-terminal domains, although their overall structure is conserved []. This entry represents the C-terminal domain found in transglutaminases, which consists of an immunoglobulin-like beta-sandwich consisting of seven strands in two sheets with a Greek key topology. The best known transglutaminase is blood coagulation factor XIII, a plasma tetrameric protein composed of two catalytic A subunits and two non-catalytic B subunits. Factor XIII is responsible for cross-linking fibrin chains, thus stabilising the fibrin clot. Protein-glutamine gamma-glutamyltransferases (2.3.2.13 from EC) are calcium-dependent enzymes that catalyse the cross-linking of proteins by promoting the formation of isopeptide bonds between the gamma-carboxyl group of a glutamine in one polypeptide chain and the epsilon-amino group of a lysine in a second polypeptide chain. TGases also catalyse the conjugation of polyamines to proteins [, ].; GO: 0003810 protein-glutamine gamma-glutamyltransferase activity, 0018149 peptide cross-linking; PDB: 2XZZ_A 1GGY_B 1FIE_B 1GGU_B 1GGT_B 1F13_A 1QRK_B 1EVU_A 1EX0_B 1L9N_B ....
Probab=91.92 E-value=0.25 Score=36.92 Aligned_cols=70 Identities=17% Similarity=0.122 Sum_probs=49.4
Q ss_pred cccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeec--Cc------eeeEEEEecCCCceEEEEEEEeccee
Q 029361 39 LKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVIS--GN------ISQSWERLDAGGILSHSFELDAKVKG 110 (194)
Q Consensus 39 ~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~--G~------~s~~~erI~pg~nvsH~vvv~Pk~~G 110 (194)
.+.++.| +|+++..++.|-.+.+-.+|++.=-.+ .+. |. .....-.|+||+..++.+.+.|..+|
T Consensus 8 ~~~~~vG-~d~~v~v~~~N~~~~~l~~v~~~l~~~------~v~ytG~~~~~~~~~~~~~~l~p~~~~~~~~~i~p~~yG 80 (107)
T PF00927_consen 8 PGDPVVG-QDFTVSVSFTNPSSEPLRNVSLNLCAF------TVEYTGLTRDQFKKEKFEVTLKPGETKSVEVTITPSQYG 80 (107)
T ss_dssp ESEEBTT-SEEEEEEEEEE-SSS-EECEEEEEEEE------EEECTTTEEEEEEEEEEEEEE-TTEEEEEEEEE-HHSHE
T ss_pred CCCccCC-CCEEEEEEEEeCCcCccccceeEEEEE------EEEECCcccccEeEEEcceeeCCCCEEEEEEEEEceeEe
Confidence 5677899 999999999999999878877653100 122 22 24456679999999999999999999
Q ss_pred eEeee
Q 029361 111 MFHGS 115 (194)
Q Consensus 111 ~fn~t 115 (194)
.-..-
T Consensus 81 ~~~~l 85 (107)
T PF00927_consen 81 PKQLL 85 (107)
T ss_dssp EECCE
T ss_pred cchhc
Confidence 84443
No 16
>PF12584 TRAPPC10: Trafficking protein particle complex subunit 10, TRAPPC10; InterPro: IPR022233 The trafficking protein particle complex TRAPP is a multi-protein complex needed in the early stages of the secretory pathway. To date, two kinds of TRAPP complexes have been studied, TRAPPI and TRAPP II. These complexes differ in subunit composition []. TRAPP I binds vesicles derived from the endoplasmic reticulum bringing them closer to the acceptor membrane. This entry represents a domain which forms part of the TRAPP complex for mediating vesicle docking and fusion in the Golgi apparatus. The fungal version is referred to as Trs130, and an alternative vertebrate alias is TMEM1 [, ].
Probab=89.35 E-value=4.2 Score=32.51 Aligned_cols=77 Identities=17% Similarity=0.203 Sum_probs=58.0
Q ss_pred cccccceeEEEEEEEEecC------------CcceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEecc
Q 029361 41 RLKSGAERISVSIDIHNQG------------TSTAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHSFELDAKV 108 (194)
Q Consensus 41 ~~v~g~~ditV~ytIYNvG------------~s~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~ 108 (194)
-..+| +-+.++..|-|.. ....+-.++.+| ++.| +|+|--...+.- ..|+..+-.++|.|.+
T Consensus 26 ~~~vG-qpi~~~l~I~~~~~W~~~~~~~~~~~~~~~~yei~a~---~~~W-lV~Grrrg~f~~-~~~~~~~~~l~LIPL~ 99 (147)
T PF12584_consen 26 PCRVG-QPIPAELRIKNSRKWSSEDQEESSNEDTEFMYEIVAD---SDNW-LVSGRRRGVFSL-SDGSEHEIPLTLIPLR 99 (147)
T ss_pred ceEeC-CeEEEEEEEEEcccCCccccccccCCCccEEEEEecC---CCcE-EEeccCcceEEe-cCCCeEEEEEEEEecc
Confidence 34688 9999999999972 122333444332 3445 899988777766 8888889999999999
Q ss_pred eeeEeeecEEEEEEc
Q 029361 109 KGMFHGSPALITFRI 123 (194)
Q Consensus 109 ~G~fn~t~A~VtY~~ 123 (194)
.|+..+...+|.=..
T Consensus 100 ~G~L~lP~V~i~~~~ 114 (147)
T PF12584_consen 100 AGYLPLPKVEIRPYD 114 (147)
T ss_pred cceecCCEEEEEecc
Confidence 999999999887665
No 17
>PF14874 PapD-like: Flagellar-associated PapD-like
Probab=89.08 E-value=7 Score=28.44 Aligned_cols=73 Identities=16% Similarity=0.178 Sum_probs=50.6
Q ss_pred cccccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEEEEEE-ecceeeEeee
Q 029361 37 ASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHSFELD-AKVKGMFHGS 115 (194)
Q Consensus 37 i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~-Pk~~G~fn~t 115 (194)
+.=.....| +..+...+|-|.|..++ ...+..+.-..+.|.+--. =..|+||+.++-.+++. +...|.|...
T Consensus 11 ldFG~v~~g-~~~~~~v~l~N~s~~p~-~f~v~~~~~~~~~~~v~~~-----~g~l~PG~~~~~~V~~~~~~~~g~~~~~ 83 (102)
T PF14874_consen 11 LDFGNVFVG-QTYSRTVTLTNTSSIPA-RFRVRQPESLSSFFSVEPP-----SGFLAPGESVELEVTFSPTKPLGDYEGS 83 (102)
T ss_pred EEeeEEccC-CEEEEEEEEEECCCCCE-EEEEEeCCcCCCCEEEECC-----CCEECCCCEEEEEEEEEeCCCCceEEEE
Confidence 333445577 99999999999999875 3333333323455655321 14599999999999999 6777988755
Q ss_pred c
Q 029361 116 P 116 (194)
Q Consensus 116 ~ 116 (194)
-
T Consensus 84 l 84 (102)
T PF14874_consen 84 L 84 (102)
T ss_pred E
Confidence 4
No 18
>PF12690 BsuPI: Intracellular proteinase inhibitor; InterPro: IPR020481 BsuPI is a intracellular proteinase inhibitor that directly regulates the major intracellular proteinase (ISP-1) activity in vivo. It inhibits ISP-1 in the early stages of sporulation and then may be inactivated by a membrane-bound proteinase [].; PDB: 3ISY_A.
Probab=88.41 E-value=2.6 Score=31.04 Aligned_cols=66 Identities=18% Similarity=0.288 Sum_probs=36.2
Q ss_pred EEEEEEEEecCCcc---------eeeeEEecCCCCCCCeeeecCce---eeEEEEecCCCceEEEEEEEecc--eeeEee
Q 029361 49 ISVSIDIHNQGTST---------AYDVSLTDDSWPQDKFDVISGNI---SQSWERLDAGGILSHSFELDAKV--KGMFHG 114 (194)
Q Consensus 49 itV~ytIYNvG~s~---------A~dV~L~D~sfp~e~Felv~G~~---s~~~erI~pg~nvsH~vvv~Pk~--~G~fn~ 114 (194)
+.++++|-|.|+.+ -+|+.|.|. =..+.|.--.|.+ -..=..|+||+..++..++.... .|.|..
T Consensus 2 v~~~l~v~N~s~~~v~l~f~sgq~~D~~v~d~-~g~~vwrwS~~~~FtQal~~~~l~pGe~~~~~~~~~~~~~~~G~Y~~ 80 (82)
T PF12690_consen 2 VEFTLTVTNNSDEPVTLQFPSGQRYDFVVKDK-EGKEVWRWSDGKMFTQALQEETLEPGESLTYEETWDLKDLSPGEYTL 80 (82)
T ss_dssp EEEEEEEEE-SSS-EEEEESSS--EEEEEE-T-T--EEEETTTT-------EEEEE-TT-EEEEEEEESS----SEEEEE
T ss_pred EEEEEEEEeCCCCeEEEEeCCCCEEEEEEECC-CCCEEEEecCCchhhheeeEEEECCCCEEEEEEEECCCCCCCceEEE
Confidence 56788888887743 456666652 1222222223332 23557899999999999998777 798876
Q ss_pred e
Q 029361 115 S 115 (194)
Q Consensus 115 t 115 (194)
.
T Consensus 81 ~ 81 (82)
T PF12690_consen 81 E 81 (82)
T ss_dssp E
T ss_pred e
Confidence 4
No 19
>TIGR03079 CH4_NH3mon_ox_B methane monooxygenase/ammonia monooxygenase, subunit B. Both ammonia oxidizers such as Nitrosomonas europaea and methanotrophs (obligate methane oxidizers) such as Methylococcus capsulatus each can grow only on their own characteristic substrate. However, both groups have the ability to oxidize both substrates, and so the relevant enzymes must be named here according to their ability to oxidze both. The protein family represented here reflects subunit B of both the particulate methane monooxygenase of methylotrophs and the ammonia monooxygenase of nitrifying bacteria.
Probab=85.68 E-value=4 Score=38.60 Aligned_cols=73 Identities=16% Similarity=0.324 Sum_probs=51.5
Q ss_pred ceEEEEeecccccccccceeEEEEEEEEecCCccee-------eeEEe--------cCCCCCCCeeeecCceeeE-EEEe
Q 029361 29 PFIVAHKKASLKRLKSGAERISVSIDIHNQGTSTAY-------DVSLT--------DDSWPQDKFDVISGNISQS-WERL 92 (194)
Q Consensus 29 a~LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~-------dV~L~--------D~sfp~e~Felv~G~~s~~-~erI 92 (194)
+.-+.-|-..-+|-|.| +.+.++++|-|.|+.+-+ +|.+. ++.||+|--. .| ++.+ =+-|
T Consensus 265 ~~~V~~kv~~a~Y~VPG-R~l~~~~~VTN~g~~~vrlgEF~TA~vRFlN~~~v~~~~~~yP~~lla--~G-L~v~d~~pI 340 (399)
T TIGR03079 265 PNPVSINVTKANYDVPG-RALRVTMEITNNGDQVISIGEFTTAGIRFMNANGVRVLDPDYPRELLA--EG-LEVDDQSAI 340 (399)
T ss_pred CCceEEEEeccEEecCC-cEEEEEEEEEcCCCCceEEEeEeecceEeeCcccccccCCCChHHHhh--cc-ceeCCCCCc
Confidence 34667787888999999 999999999999998754 33333 3455555322 23 3433 3359
Q ss_pred cCCCceEEEEEEE
Q 029361 93 DAGGILSHSFELD 105 (194)
Q Consensus 93 ~pg~nvsH~vvv~ 105 (194)
.||++.+-++...
T Consensus 341 ~PGETr~v~v~aq 353 (399)
T TIGR03079 341 APGETVEVKMEAK 353 (399)
T ss_pred CCCcceEEEEEEe
Confidence 9999998887765
No 20
>PF13584 BatD: Oxygen tolerance
Probab=85.64 E-value=13 Score=34.68 Aligned_cols=114 Identities=14% Similarity=0.161 Sum_probs=72.6
Q ss_pred CceEEEEeecccccccccceeEEEEEEEEecCCcceee-eEEecCCCCCCCeeeecCceeeEEEEe-cCCC---ceE-EE
Q 029361 28 VPFIVAHKKASLKRLKSGAERISVSIDIHNQGTSTAYD-VSLTDDSWPQDKFDVISGNISQSWERL-DAGG---ILS-HS 101 (194)
Q Consensus 28 ~a~LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~d-V~L~D~sfp~e~Felv~G~~s~~~erI-~pg~---nvs-H~ 101 (194)
...+.+.=.+..+.+-+| +.+.++|.+|=...-...+ ..+..+.++ +|.+..=.-..++.+- --|. .+. +.
T Consensus 125 ~~~~~l~~~v~~~~~Yvg-e~v~lt~~ly~~~~~~~~~~~~~~~p~~~--~~~~~~~~~~~~~~~~~i~G~~y~~~~~~~ 201 (484)
T PF13584_consen 125 DDDVFLEAEVSKKSVYVG-EPVILTLRLYTRNNFRQLGIEELPPPDFE--GFWVEQLGDDRQYEEERINGRRYRVIELRR 201 (484)
T ss_pred cccEEEEEEeCCCceecC-CcEEEEEEEEEecCchhccccccCCCCCC--CcEEEECCCCCceeEEEECCEEEEEEEEEE
Confidence 344677777778889999 9999999999877765333 233333333 3432222223344432 2222 233 56
Q ss_pred EEEEecceeeEeeecEEEEEEcCCC------------cceeeEeecCCCcceeee
Q 029361 102 FELDAKVKGMFHGSPALITFRIPTK------------AALQEAYSTPMLPLDVLA 144 (194)
Q Consensus 102 vvv~Pk~~G~fn~t~A~VtY~~se~------------~~~q~a~Ss~pg~~~I~~ 144 (194)
+.|.|.+.|.+...++.++...... ...+.-+++++....|.+
T Consensus 202 ~~l~P~ksG~l~I~~~~~~~~~~~~~~~~~~fg~~~~~~~~~~~~s~~~~i~V~p 256 (484)
T PF13584_consen 202 YALFPQKSGTLTIPPATFEVTVSDPSGRRDFFGGNFGRSRPVSISSEPLTITVKP 256 (484)
T ss_pred EEEEeCCceeEEecCEEEEEEEecccCccCccccccccceeEEecCCCeEEEecc
Confidence 8999999999999999998876532 123466777777776654
No 21
>COG1361 S-layer domain [Cell envelope biogenesis, outer membrane]
Probab=84.63 E-value=3.1 Score=39.02 Aligned_cols=80 Identities=23% Similarity=0.253 Sum_probs=54.4
Q ss_pred cccccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCce-eeEEEEecC--CCceEEEEEEE---eccee
Q 029361 37 ASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNI-SQSWERLDA--GGILSHSFELD---AKVKG 110 (194)
Q Consensus 37 i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~-s~~~erI~p--g~nvsH~vvv~---Pk~~G 110 (194)
.....+-.| .+..+..+++|.|...+.|+.+....-.+ |....+.. ......+.. |+-.++.+.+. ..+.|
T Consensus 38 ~~p~~~~~~-~~~~l~v~~~n~~~~~~~~v~v~i~~~~~--~~~~~~~~~~~~~~~~~~l~~~~~~~~~~~~V~~~a~~g 114 (500)
T COG1361 38 YSPNVARPG-EDVDLTVTIENVGELLAEDVKVEITPEYP--FSLVSGETLLLSIGTLNFLGGEPATVKFKLTVDENAKSG 114 (500)
T ss_pred ccCcccCcc-cceEEEEEeccccccccccEEEEEEeccc--ceeeEEEeecCCCceeeecCCCcceEEEEEEEcCCCCCC
Confidence 334445555 99999999999999988888877632222 88888875 333444444 55555555443 67788
Q ss_pred eEeeecEEE
Q 029361 111 MFHGSPALI 119 (194)
Q Consensus 111 ~fn~t~A~V 119 (194)
.|++.-..-
T Consensus 115 ~y~i~v~~~ 123 (500)
T COG1361 115 DYEIDVYVS 123 (500)
T ss_pred cEEEeEEEE
Confidence 888887774
No 22
>PF04744 Monooxygenase_B: Monooxygenase subunit B protein; InterPro: IPR006833 Ammonia monooxygenase and the particulate methane monooxygenase are both integral membrane proteins, occurring in ammonia oxidisers and methanotrophs respectively, which are thought to be evolutionarily related []. These enzymes have a relatively wide substrate specificity and can catalyse the oxidation of a range of substrates including ammonia, methane, halogenated hydrocarbons and aromatic molecules []. These enzymes are composed of 3 subunits - A (IPR003393 from INTERPRO), B (IPR006833 from INTERPRO) and C (IPR006980 from INTERPRO) - and contain various metal centres, including copper. Particulate methane monooxygenase from Methylococcus capsulatus str. Bath is an ABC homotrimer, which contains mononuclear and dinuclear copper metal centres, and a third metal centre containing a metal ion whose identity in vivo is not certain[]. The soluble regions of these enzymes derive primarily from the B subunit. This subunit forms two antiparallel beta-barrel-like structures and contains the mono- and di- nuclear copper metal centres [].; PDB: 3CHX_E 3RFR_A 3RGB_A 1YEW_A.
Probab=83.60 E-value=12 Score=35.46 Aligned_cols=72 Identities=18% Similarity=0.285 Sum_probs=46.4
Q ss_pred EEEeecccccccccceeEEEEEEEEecCCccee-------eeEEecCCCCCC------CeeeecCceeeEEE-EecCCCc
Q 029361 32 VAHKKASLKRLKSGAERISVSIDIHNQGTSTAY-------DVSLTDDSWPQD------KFDVISGNISQSWE-RLDAGGI 97 (194)
Q Consensus 32 lvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~-------dV~L~D~sfp~e------~Felv~G~~s~~~e-rI~pg~n 97 (194)
+.-|-..-+|-|.| +.++++.+|-|.|+++.. +|.+.|+..+.+ +. +-.+-++++=+ -|+||++
T Consensus 249 V~~~v~~A~Y~vpg-R~l~~~l~VtN~g~~pv~LgeF~tA~vrFln~~v~~~~~~~P~~l-~A~~gL~vs~~~pI~PGET 326 (381)
T PF04744_consen 249 VKVKVTDATYRVPG-RTLTMTLTVTNNGDSPVRLGEFNTANVRFLNPDVPTDDPDYPDEL-LAERGLSVSDNSPIAPGET 326 (381)
T ss_dssp EEEEEEEEEEESSS-SEEEEEEEEEEESSS-BEEEEEESSS-EEE-TTT-SS-S---TTT-EETT-EEES--S-B-TT-E
T ss_pred eEEEEeccEEecCC-cEEEEEEEEEcCCCCceEeeeEEeccEEEeCcccccCCCCCchhh-hccCcceeCCCCCcCCCce
Confidence 56666778899999 999999999999999875 477777555422 22 33323555544 7999999
Q ss_pred eEEEEEEE
Q 029361 98 LSHSFELD 105 (194)
Q Consensus 98 vsH~vvv~ 105 (194)
.+-++.+.
T Consensus 327 rtl~V~a~ 334 (381)
T PF04744_consen 327 RTLTVEAQ 334 (381)
T ss_dssp EEEEEEEE
T ss_pred EEEEEEee
Confidence 99988875
No 23
>PF03345 DDOST_48kD: Oligosaccharyltransferase 48 kDa subunit beta; InterPro: IPR005013 During N-linked glycosylation of proteins, oligosaccharide chains are assembled on the carrier molecule dolichyl pyrophosphate in the following order: 2 molecules of N-acetylglucosamine (GlcNAc), 9 molecules of mannose, and 3 molecules of glucose. These 14-residue oligosaccharide cores are then transferred to asparagine residues on nascent polypeptide chains in the endoplasmic reticulum (ER). As proteins progress through the Golgi apparatus, the oligosaccharide cores are modified by trimming and extension to generate a diverse array of glycosylated proteins [, ]. The oligosaccharyl transferase complex (OST complex) 2.4.1.119 from EC transfers 14-sugar branched oligosaccharides from dolichyl pyrophosphate to asparagine residues []. The complex contains nine protein subunits: Ost1p, Ost2p, Ost3p, Ost4p, Ost5p, Ost6p, Stt3p, Swp1p, and Wbp1p, all of which are integral membrane proteins of the ER. The OST complex interacts with the Sec61p pore complex [] involved in protein import into the ER. This entry represents subunits OST3 and OST6. OST3 is homologous to OST6 [], and several lines of evidence indicate that they are alternative members of the OST complex. Disruption of both OST3 and OST6 causes severe underglycosylation of soluble and membrane-bound glycoproteins and a defect in the assembly of the complex. Hence, the function of these genes seems to be essential for recruiting a fully active complex necessary for efficient N-glycosylation []. This entry also includes the magnesium transporter protein 1, also known as OST3 homologue B, which might be involved in N-glycosylation through its association with the oligosaccharyl transferase (OST) complex. Wbp1p is the beta subunit of the OST complex, one of the original six subunits purified []. Wbp1 is essential [, ], but conditional mutants have decreased transferase activity [, ]. Wbp1p is homologous to mammalian OST48 [].; GO: 0004579 dolichyl-diphosphooligosaccharide-protein glycotransferase activity, 0018279 protein N-linked glycosylation via asparagine, 0005789 endoplasmic reticulum membrane
Probab=80.85 E-value=20 Score=34.14 Aligned_cols=79 Identities=22% Similarity=0.326 Sum_probs=42.6
Q ss_pred CCceEEEEEEE-ecceeeEeeecEEEEEEcCCCcceeeEeecCCCcceeeecCchhhhhHHHHHHHhhhchhhhhHHHhe
Q 029361 95 GGILSHSFELD-AKVKGMFHGSPALITFRIPTKAALQEAYSTPMLPLDVLAEKPTENKLELAKRLLAKYGSQISVISIIV 173 (194)
Q Consensus 95 g~nvsH~vvv~-Pk~~G~fn~t~A~VtY~~se~~~~q~a~Ss~pg~~~I~~~~~ydrkfewa~~l~~~y~~~~~v~s~~~ 173 (194)
+.+-.++...+ |-..|.|+| .|.|+-.-=.-+.....-+.++ ++-.+|+|.+ +|...|--..+++|+++
T Consensus 329 ~~~~~Y~~~FklPD~hGVF~F---~vdY~R~G~t~l~~~~~v~VRp---l~Hdey~Rs~----fI~~A~PYyas~~s~m~ 398 (423)
T PF03345_consen 329 DDNGTYSTTFKLPDVHGVFTF---KVDYKRPGYTFLEEKTQVSVRP---LAHDEYPRSW----FITNAYPYYASAFSMMI 398 (423)
T ss_pred CCCCEEEEEEECCCccceEEE---EEEEecCceeeEEEEEEEeccC---CccccCcccc----ccccccHHHHHHHHHHH
Confidence 34444555555 999999999 5888853211111122222222 2346788855 55555555555555444
Q ss_pred -----eeeEEEeCcC
Q 029361 174 -----LFVYLITSPS 183 (194)
Q Consensus 174 -----~~v~~~~~~~ 183 (194)
.++||.-.|.
T Consensus 399 gf~lF~~~fL~~~~~ 413 (423)
T PF03345_consen 399 GFFLFVFVFLYHKPV 413 (423)
T ss_pred HHHhheeeEEEecCc
Confidence 3445555554
No 24
>PF09478 CBM49: Carbohydrate binding domain CBM49; InterPro: IPR019028 A carbohydrate-binding module (CBM) is defined as a contiguous amino acid sequence within a carbohydrate-active enzyme with a discreet fold having carbohydrate-binding activity. A few exceptions are CBMs in cellulosomal scaffolding proteins and rare instances of independent putative CBMs. The requirement of CBMs existing as modules within larger enzymes sets this class of carbohydrate-binding protein apart from other non-catalytic sugar binding proteins such as lectins and sugar transport proteins. CBMs were previously classified as cellulose-binding domains (CBDs) based on the initial discovery of several modules that bound cellulose [, ]. However, additional modules in carbohydrate-active enzymes are continually being found that bind carbohydrates other than cellulose yet otherwise meet the CBM criteria, hence the need to reclassify these polypeptides using more inclusive terminology. Previous classification of cellulose-binding domains were based on amino acid similarity. Groupings of CBDs were called "Types" and numbered with roman numerals (e.g. Type I or Type II CBDs). In keeping with the glycoside hydrolase classification, these groupings are now called families and numbered with Arabic numerals. Families 1 to 13 are the same as Types I to XIII. For a detailed review on the structure and binding modes of CBMs see []. This domain is found at the C-terminal of cellulases and in vitro binding studies have shown it to binds to crystalline cellulose []. ; GO: 0030246 carbohydrate binding, 0005576 extracellular region
Probab=80.39 E-value=8.6 Score=27.71 Aligned_cols=65 Identities=9% Similarity=0.215 Sum_probs=45.7
Q ss_pred cccccccee-EEEEEEEEecCCcceeeeEEecCCCCCCCeeeec---Cceee-EEE-EecCCCceEEEEEE
Q 029361 40 KRLKSGAER-ISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVIS---GNISQ-SWE-RLDAGGILSHSFEL 104 (194)
Q Consensus 40 ~~~v~g~~d-itV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~---G~~s~-~~e-rI~pg~nvsH~vvv 104 (194)
+.-.+|+.. .-+..+|.|.|..+-.++.|.=+.+..+-+++.. |.... +|- .|+||++.+--|+.
T Consensus 9 ~sW~~~g~~y~qy~v~I~N~~~~~I~~~~i~~~~l~~~iW~l~~~~~~~y~lPs~~~~i~pg~s~~FGYI~ 79 (80)
T PF09478_consen 9 NSWTENGQTYTQYDVTITNNGSKPIKSLKISIDNLYGSIWGLDKVSGNTYTLPSYQPTIKPGQSFTFGYIS 79 (80)
T ss_pred eEEEeCCEEEEEEEEEEEECCCCeEEEEEEEECccchhheeEEeccCCEEECCccccccCCCCEEEEEEEe
Confidence 333444333 3467889999999999999988777777777766 22333 564 89999988766653
No 25
>PF09624 DUF2393: Protein of unknown function (DUF2393); InterPro: IPR013417 The function of this protein is unknown. It is always found as part of a two-gene operon with IPR013416 from INTERPRO, a protein that appears to span the membrane seven times. It has so far been found in the bacteria Anabaena sp. (strain PCC 7120), Agrobacterium tumefaciens, Rhizobium meliloti, and Gloeobacter violaceus.
Probab=79.50 E-value=28 Score=27.52 Aligned_cols=74 Identities=20% Similarity=0.131 Sum_probs=43.0
Q ss_pred ccccccceeEEEEEEEEecCCcceeeeEEecC----CCCCCC-----eeeecCc--eeeEEEE-ecCCCceEEEEEEE-e
Q 029361 40 KRLKSGAERISVSIDIHNQGTSTAYDVSLTDD----SWPQDK-----FDVISGN--ISQSWER-LDAGGILSHSFELD-A 106 (194)
Q Consensus 40 ~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~----sfp~e~-----Felv~G~--~s~~~er-I~pg~nvsH~vvv~-P 106 (194)
+.+-.+ +.+.|..+|-|.|+-++.++.++=+ +...+. ++-..+- .+..++. |+||+...-++.+. |
T Consensus 56 ~~l~~~-~~~~v~g~V~N~g~~~i~~c~i~~~l~~~~~~~~n~~~~~~~~~~~f~~~~~~i~~~L~~~e~~~f~~~~~~~ 134 (149)
T PF09624_consen 56 KRLQYS-ESFYVDGTVTNTGKFTIKKCKITVKLYNDKQVSGNKFKEIFYQQIPFVKKSIPIADNLKPGESKEFRFIFPYP 134 (149)
T ss_pred eeeeec-cEEEEEEEEEECCCCEeeEEEEEEEEEeCCCccCchhhhhhccccchhccceeHHhhcCcccceeEEEEecCC
Confidence 333345 8999999999999999999887643 111111 1111110 0122222 88888888877766 3
Q ss_pred cceeeEee
Q 029361 107 KVKGMFHG 114 (194)
Q Consensus 107 k~~G~fn~ 114 (194)
...|.+++
T Consensus 135 p~~~~~~~ 142 (149)
T PF09624_consen 135 PYFGNYNI 142 (149)
T ss_pred ccCCCceE
Confidence 33444443
No 26
>PRK10378 inactive ferrous ion transporter periplasmic protein EfeO; Provisional
Probab=74.24 E-value=19 Score=33.81 Aligned_cols=50 Identities=12% Similarity=0.236 Sum_probs=32.0
Q ss_pred EEEEEEecCCcceeeeEEecCCCCCCCeeeecCce-eeEEEEecCCCceEEEEEEEecceeeEeeec
Q 029361 51 VSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNI-SQSWERLDAGGILSHSFELDAKVKGMFHGSP 116 (194)
Q Consensus 51 V~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~-s~~~erI~pg~nvsH~vvv~Pk~~G~fn~t~ 116 (194)
+.+.|.|.|..+ .+ |+++.|.+ -...|.|.||.+-+.+ .+.+.|.|.|.=
T Consensus 53 ~~f~V~N~~~~~-~E------------fe~~~~~~vv~e~EnIaPG~s~~l~---~~L~pGtY~~~C 103 (375)
T PRK10378 53 TQFIIQNHSQKA-LE------------WEILKGVMVVEERENIAPGFSQKMT---ANLQPGEYDMTC 103 (375)
T ss_pred EEEEEEeCCCCc-ce------------EEeeccccccccccccCCCCceEEE---EecCCceEEeec
Confidence 567778877654 33 44444331 2246899999887744 344678888865
No 27
>PF05506 DUF756: Domain of unknown function (DUF756); InterPro: IPR008475 This domain is found, normally as a tandem repeat, at the C terminus of bacterial phospholipase C proteins.; GO: 0004629 phospholipase C activity, 0016042 lipid catabolic process
Probab=73.54 E-value=27 Score=25.30 Aligned_cols=52 Identities=19% Similarity=0.391 Sum_probs=34.8
Q ss_pred EEEEEEEecCCcceeeeEEecCCCC---CCCeeeecCceeeEEEEecCCCceEEEEEEEecceeeEeee
Q 029361 50 SVSIDIHNQGTSTAYDVSLTDDSWP---QDKFDVISGNISQSWERLDAGGILSHSFELDAKVKGMFHGS 115 (194)
Q Consensus 50 tV~ytIYNvG~s~A~dV~L~D~sfp---~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~~G~fn~t 115 (194)
.+..+|-|.|. .+..+++.|+... +..+ .|+||+++++.+-+ ....|-|-|+
T Consensus 21 ~l~l~l~N~g~-~~~~~~v~~~~y~~~~~~~~------------~v~ag~~~~~~w~l-~~s~gwYDl~ 75 (89)
T PF05506_consen 21 NLRLTLSNPGS-AAVTFTVYDNAYGGGGPWTY------------TVAAGQTVSLTWPL-AASGGWYDLT 75 (89)
T ss_pred EEEEEEEeCCC-CcEEEEEEeCCcCCCCCEEE------------EECCCCEEEEEEee-cCCCCcEEEE
Confidence 78889999987 4558999986553 2222 46677777777766 4555666554
No 28
>cd04036 C2_cPLA2 C2 domain present in cytosolic PhosphoLipase A2 (cPLA2). A single copy of the C2 domain is present in cPLA2 which releases arachidonic acid from membranes initiating the biosynthesis of potent inflammatory mediators such as prostaglandins, leukotrienes, and platelet-activating factor. C2 domains fold into an 8-standed beta-sandwich that can adopt 2 structural arrangements: Type I and Type II, distinguished by a circular permutation involving their N- and C-terminal beta strands. Many C2 domains are Ca2+-dependent membrane-targeting modules that bind a wide variety of substances including bind phospholipids, inositol polyphosphates, and intracellular proteins. Most C2 domain proteins are either signal transduction enzymes that contain a single C2 domain, such as protein kinase C, or membrane trafficking proteins which contain at least two C2 domains, such as synaptotagmin 1. However, there are a few exceptions to this including RIM isoforms and some splice variants o
Probab=70.77 E-value=11 Score=28.11 Aligned_cols=68 Identities=18% Similarity=0.216 Sum_probs=49.2
Q ss_pred cccccceeEEEEEEEEecCCcceeeeEEec-CCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEecceeeEeeec
Q 029361 41 RLKSGAERISVSIDIHNQGTSTAYDVSLTD-DSWPQDKFDVISGNISQSWERLDAGGILSHSFELDAKVKGMFHGSP 116 (194)
Q Consensus 41 ~~v~g~~ditV~ytIYNvG~s~A~dV~L~D-~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~~G~fn~t~ 116 (194)
.++-+ +.+++ .+.+- ......|++-| +.+ .++| =|.....++.|.+|......+.+.++..|..++.-
T Consensus 47 nP~Wn-e~f~f--~i~~~-~~~~l~v~v~d~d~~-~~~~---iG~~~~~l~~l~~g~~~~~~~~L~~~~~g~l~~~~ 115 (119)
T cd04036 47 NPVWN-ETFEF--RIQSQ-VKNVLELTVMDEDYV-MDDH---LGTVLFDVSKLKLGEKVRVTFSLNPQGKEELEVEF 115 (119)
T ss_pred CCccc-eEEEE--EeCcc-cCCEEEEEEEECCCC-CCcc---cEEEEEEHHHCCCCCcEEEEEECCCCCCceEEEEE
Confidence 34555 45444 44443 33568899988 444 4443 47888899999999999999999999899887753
No 29
>COG1572 Uncharacterized conserved protein [Function unknown]
Probab=69.65 E-value=16 Score=36.53 Aligned_cols=68 Identities=16% Similarity=0.204 Sum_probs=55.7
Q ss_pred cccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCc--eeeEEEEecCCCceEEEEEEEecceeeEeeecEE
Q 029361 41 RLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGN--ISQSWERLDAGGILSHSFELDAKVKGMFHGSPAL 118 (194)
Q Consensus 41 ~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~--~s~~~erI~pg~nvsH~vvv~Pk~~G~fn~t~A~ 118 (194)
...++ +.+++.++|-|.|.+.|.+++..+ ++.|. ++.+..-++||+..+-.|.-.+...|.++++.-.
T Consensus 418 ~~~~~-k~~~i~l~i~N~G~~~a~~~~v~l---------~lnG~~~~~~~i~~l~~~~s~e~~v~~~~~s~G~~~Ls~~~ 487 (606)
T COG1572 418 QESVN-KALTITLNIKNLGEAYASGFQVDL---------VLNGTIVTVDSIPGLESGESREVVVNEVSTSGGSHTLSVVI 487 (606)
T ss_pred eEeec-ceEEEEEEEEeccccccCCceEEE---------EEcCceeeeEecccCCCCCceEEEEEEEecCCCceEEEEEe
Confidence 34456 999999999999999999988876 67776 3667777889998888888778999999887543
No 30
>KOG4386 consensus Uncharacterized conserved protein [Function unknown]
Probab=68.90 E-value=10 Score=37.92 Aligned_cols=73 Identities=19% Similarity=0.251 Sum_probs=59.8
Q ss_pred ccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEecceeeEeeecEEEE
Q 029361 42 LKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHSFELDAKVKGMFHGSPALIT 120 (194)
Q Consensus 42 ~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~~G~fn~t~A~Vt 120 (194)
..+- +.+.|.|.+-|--+ -+.||++.=+ |+|.| .-+|.-..+. ||-||+.-.-.|.+-|.--|+.||.+--+.
T Consensus 704 grVR-eslpvkyhLqnktd-lvqdveisve--psDaF-MFSGlkqirl-riLPGteqemlynfypLmAGyqqlPslnin 776 (809)
T KOG4386|consen 704 GRVR-ESLPVKYHLQNKTD-LVQDVEISVE--PSDAF-MFSGLKQIRL-RILPGTEQEMLYNFYPLMAGYQQLPSLNIN 776 (809)
T ss_pred ceec-ccccEEEEeccccc-eeeeEEeecc--cchhh-eecccceEEE-EEcCCCceEEEEEEehhhchhhhCCccccc
Confidence 3445 68999999999766 7889988643 78889 8888877765 788999999999999999999999876554
No 31
>KOG2291 consensus Oligosaccharyltransferase, alpha subunit (ribophorin I) [Posttranslational modification, protein turnover, chaperones]
Probab=67.53 E-value=16 Score=36.21 Aligned_cols=71 Identities=23% Similarity=0.208 Sum_probs=41.1
Q ss_pred CCCchhhHHHHHHHHHHHHhhhcccCCCceEE---EEeecccccccccceeEEEEEEEEecCCcceeeeEEecCCCC
Q 029361 1 MASPISKSLISVLIALFLISSSFASSDVPFIV---AHKKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWP 74 (194)
Q Consensus 1 ~~~~~~~~~~~~lla~~~v~~~~~~~~~a~Ll---vsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp 74 (194)
|++.....+..+++.++++++.++....+-+. +-+.|....-..- .+.++.|-|+|+.||...-+.=...+
T Consensus 1 M~~~~~~~~~~l~l~l~aia~~~a~~a~~~w~n~nv~RTIDlsS~ivK---~tt~l~i~N~g~ePatey~~a~~~~~ 74 (602)
T KOG2291|consen 1 MAQVSASWALVLVLLLFAIASGAASSAEQDWVNVNVERTIDLSSQIVK---VTTELSIENIGSEPATEYLLAFEKEL 74 (602)
T ss_pred CcchhhHHHHHHHHHHHHHhhccccCCccccccccceEEEehhhhhhh---heeEEEEEecCCCchheEEEeccCcc
Confidence 77655544445555556666554443333233 2233433322222 57889999999999999888754443
No 32
>PF02102 Peptidase_M35: Deuterolysin metalloprotease (M35) family; InterPro: IPR001384 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Metalloproteases are the most diverse of the four main types of protease, with more than 50 families identified to date. In these enzymes, a divalent cation, usually zinc, activates the water molecule. The metal ion is held in place by amino acid ligands, usually three in number. The known metal ligands are His, Glu, Asp or Lys and at least one other residue is required for catalysis, which may play an electrophillic role. Of the known metalloproteases, around half contain an HEXXH motif, which has been shown in crystallographic studies to form part of the metal-binding site []. The HEXXH motif is relatively common, but can be more stringently defined for metalloproteases as 'abXHEbbHbc', where 'a' is most often valine or threonine and forms part of the S1' subsite in thermolysin and neprilysin, 'b' is an uncharged residue, and 'c' a hydrophobic residue. Proline is never found in this site, possibly because it would break the helical structure adopted by this motif in metalloproteases []. This group of metallopeptidases belong to MEROPS peptidase family M35 (deuterolysin family, clan MA(M)). The protein fold of the peptidase domain for members of this family resembles that of thermolysin, the type example for clan MA. Deuterolysin is a microbial zinc-containing metalloprotease that shows some similarity to thermolysin []. The protein is expressed with a possible 19-residue signal sequence, a 155-residue propeptide, and an active peptide of 177 residues []. The latter contains an HEXXH motif towards the C terminus, but the other zinc ligands are as yet undetermined [, ].; GO: 0004222 metalloendopeptidase activity, 0006508 proteolysis; PDB: 1EB6_A.
Probab=65.19 E-value=2.1 Score=39.93 Aligned_cols=59 Identities=19% Similarity=0.286 Sum_probs=0.0
Q ss_pred eeEEEEEEEEecCCcceeeeEE--ecCCCCCCCeeeecCc-----------------eeeEEEEecCCCceEEEEEEE
Q 029361 47 ERISVSIDIHNQGTSTAYDVSL--TDDSWPQDKFDVISGN-----------------ISQSWERLDAGGILSHSFELD 105 (194)
Q Consensus 47 ~ditV~ytIYNvG~s~A~dV~L--~D~sfp~e~Felv~G~-----------------~s~~~erI~pg~nvsH~vvv~ 105 (194)
.+..|+-+|.|.|+.+-.=++. ..|+.|-+.|+|-++. ..--|..|+||++++|+|-+-
T Consensus 38 ~nt~VkA~VTNtG~e~l~llK~ntilD~~Pv~kv~V~~~g~~V~F~Gi~~~~~~~~L~~d~F~~L~pG~sve~~fDiA 115 (359)
T PF02102_consen 38 GNTRVKATVTNTGSEDLKLLKYNTILDSAPVKKVSVYKDGKEVPFTGIRLRYDTSGLTEDAFQTLAPGESVEVEFDIA 115 (359)
T ss_dssp ------------------------------------------------------------------------------
T ss_pred CCcEEEEEEEeCCCcceEEEeeceecCCCceeEEEEEcCCcccccccEEEEEecCCCCHHHceecCCCCeEEEEEcch
Confidence 6778999999999987332222 2346788888876653 344678999999999998765
No 33
>PF11611 DUF4352: Domain of unknown function (DUF4352); InterPro: IPR021652 This entry is represented by Bacteriophage A118, Gp32. The characteristics of the protein distribution suggest prophage matches in addition to the phage matches. This entry represents a group of putative lipoproteins of unknown function.; PDB: 3CFU_A.
Probab=64.97 E-value=46 Score=24.49 Aligned_cols=67 Identities=21% Similarity=0.301 Sum_probs=33.5
Q ss_pred cccceeEEEEEEEEecCCcce----eeeEEecC-CCCCC-CeeeecCceeeEEEEecCCCceEEEEEEE-ecce
Q 029361 43 KSGAERISVSIDIHNQGTSTA----YDVSLTDD-SWPQD-KFDVISGNISQSWERLDAGGILSHSFELD-AKVK 109 (194)
Q Consensus 43 v~g~~ditV~ytIYNvG~s~A----~dV~L~D~-sfp~e-~Felv~G~~s~~~erI~pg~nvsH~vvv~-Pk~~ 109 (194)
.+|++=+.|.++|-|.|+.+- .+.+|.|+ +-.-+ .+....-........|+||+.++=.++-. |+..
T Consensus 32 ~~g~~fv~v~v~v~N~~~~~~~~~~~~f~l~d~~g~~~~~~~~~~~~~~~~~~~~i~pG~~~~g~l~F~vp~~~ 105 (123)
T PF11611_consen 32 KEGNKFVVVDVTVKNNGDEPLDFSPSDFKLYDSDGNKYDPDFSASSNDNDLFSETIKPGESVTGKLVFEVPKDD 105 (123)
T ss_dssp ---SEEEEEEEEEEE-SSS-EEEEGGGEEEE-TT--B--EEE-CCCTTTB--EEEE-TT-EEEEEEEEEESTT-
T ss_pred CCCCEEEEEEEEEEECCCCcEEecccceEEEeCCCCEEcccccchhccccccccEECCCCEEEEEEEEEECCCC
Confidence 466688999999999999753 35666542 11111 11111101114678999999998887776 5444
No 34
>PF14796 AP3B1_C: Clathrin-adaptor complex-3 beta-1 subunit C-terminal
Probab=64.61 E-value=17 Score=29.93 Aligned_cols=49 Identities=14% Similarity=0.270 Sum_probs=37.3
Q ss_pred eeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCce---eeEEEEecCCCceEEEE
Q 029361 47 ERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNI---SQSWERLDAGGILSHSF 102 (194)
Q Consensus 47 ~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~---s~~~erI~pg~nvsH~v 102 (194)
.-+.|+.++-|.++.+-.+|.+.+... ..|.- -..+++|+||++++-.+
T Consensus 85 ~mvsIql~ftN~s~~~i~~I~i~~k~l-------~~g~~i~~F~~I~~L~pg~s~t~~l 136 (145)
T PF14796_consen 85 SMVSIQLTFTNNSDEPIKNIHIGEKKL-------PAGMRIHEFPEIESLEPGASVTVSL 136 (145)
T ss_pred CcEEEEEEEEecCCCeecceEECCCCC-------CCCcEeeccCcccccCCCCeEEEEE
Confidence 456677789999999999999999643 33432 24788999999988544
No 35
>PF06159 DUF974: Protein of unknown function (DUF974); InterPro: IPR010378 This is a family of uncharacterised eukaryotic proteins.
Probab=63.16 E-value=51 Score=28.83 Aligned_cols=79 Identities=16% Similarity=0.241 Sum_probs=61.2
Q ss_pred ccceeEEEEEEEEecCCcceeeeEEecCC-CCCC--CeeeecCcee-eEEEEecCCCceEEEEEEEecceeeEeeecEEE
Q 029361 44 SGAERISVSIDIHNQGTSTAYDVSLTDDS-WPQD--KFDVISGNIS-QSWERLDAGGILSHSFELDAKVKGMFHGSPALI 119 (194)
Q Consensus 44 ~g~~ditV~ytIYNvG~s~A~dV~L~D~s-fp~e--~Felv~G~~s-~~~erI~pg~nvsH~vvv~Pk~~G~fn~t~A~V 119 (194)
.| +.+...+.+-|--+.+..+|.|.=+= =|+. .+.+...... .....|+||+++...+.-.=+..|.|... -.|
T Consensus 12 lG-EtF~~~l~~~N~s~~~v~~v~ikvemqT~s~~~r~~L~~~~~~~~~~~~L~p~~~l~~iv~~~lkE~G~h~L~-c~V 89 (249)
T PF06159_consen 12 LG-ETFSCYLSVNNDSNKPVRNVRIKVEMQTPSQSLRLPLSDNENSDSPVASLAPGESLDFIVSHELKELGNHTLV-CTV 89 (249)
T ss_pred ec-CCEEEEEEeecCCCCceEEeEEEEEEeCCCCCccccCCCCccccccccccCCCCeEeEEEEEEeeecCceEEE-EEE
Confidence 57 88999999999888899998776532 2344 5655554432 35778999999999999999999999994 568
Q ss_pred EEEcC
Q 029361 120 TFRIP 124 (194)
Q Consensus 120 tY~~s 124 (194)
+|...
T Consensus 90 sY~~~ 94 (249)
T PF06159_consen 90 SYTDP 94 (249)
T ss_pred EEecC
Confidence 88877
No 36
>PF07610 DUF1573: Protein of unknown function (DUF1573); InterPro: IPR011467 These hypothetical proteins from bacteria, such as Rhodopirellula baltica, Bacteroides thetaiotaomicron and Porphyromonas gingivalis, share a region of conserved sequence towards their N termini.
Probab=62.53 E-value=28 Score=22.51 Aligned_cols=41 Identities=17% Similarity=0.272 Sum_probs=25.0
Q ss_pred EEEEEecCCcceeeeEEecCCCCCCCeeeecCceeeEE--EEecCCCceEEEE
Q 029361 52 SIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQSW--ERLDAGGILSHSF 102 (194)
Q Consensus 52 ~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~~~--erI~pg~nvsH~v 102 (194)
+|+|.|.|+.+- .|.|- +---|-+.++| +.|+||+...-.+
T Consensus 1 ~F~~~N~g~~~L---~I~~v-------~tsCgCt~~~~~~~~i~PGes~~i~v 43 (45)
T PF07610_consen 1 TFEFTNTGDSPL---VITDV-------QTSCGCTTAEYSKKPIAPGESGKIKV 43 (45)
T ss_pred CEEEEECCCCcE---EEEEe-------eEccCCEEeeCCcceECCCCEEEEEE
Confidence 478999999873 44441 12234444444 5689998866544
No 37
>PF13473 Cupredoxin_1: Cupredoxin-like domain; PDB: 1IBZ_D 1IC0_E 1IBY_D.
Probab=61.85 E-value=11 Score=27.93 Aligned_cols=42 Identities=17% Similarity=0.364 Sum_probs=24.6
Q ss_pred ceeeeEEecCCCCCCCeeeecCc-eeeEEEEecCCCceEEEEEEEe
Q 029361 62 TAYDVSLTDDSWPQDKFDVISGN-ISQSWERLDAGGILSHSFELDA 106 (194)
Q Consensus 62 ~A~dV~L~D~sfp~e~Felv~G~-~s~~~erI~pg~nvsH~vvv~P 106 (194)
....|++.|.+|.|+..++-.|+ ....|.....+. |.+++.-
T Consensus 21 ~~v~I~~~~~~f~P~~i~v~~G~~v~l~~~N~~~~~---h~~~i~~ 63 (104)
T PF13473_consen 21 QTVTITVTDFGFSPSTITVKAGQPVTLTFTNNDSRP---HEFVIPD 63 (104)
T ss_dssp ---------EEEES-EEEEETTCEEEEEEEE-SSS----EEEEEGG
T ss_pred ccccccccCCeEecCEEEEcCCCeEEEEEEECCCCc---EEEEECC
Confidence 44677778889999999999999 588888775554 8887664
No 38
>TIGR02656 cyanin_plasto plastocyanin. Members of this family are plastocyanin, a blue copper protein related to pseudoazurin, halocyanin, amicyanin, etc. This protein, located in the thylakoid luman, performs electron transport to photosystem I in Cyanobacteria and chloroplasts.
Probab=61.26 E-value=19 Score=26.66 Aligned_cols=53 Identities=15% Similarity=0.167 Sum_probs=33.1
Q ss_pred EecCCcceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEecceeeEee
Q 029361 56 HNQGTSTAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHSFELDAKVKGMFHG 114 (194)
Q Consensus 56 YNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~~G~fn~ 114 (194)
-|.|. ...++.+.+..+|.+..+. .+.+..+--.+.||++.+++|.- .|.|.|
T Consensus 30 ~N~~~-~~H~~~~~~~~~~~~~~~~-~~~~~~~~~~~~pG~t~~~tF~~----~G~y~y 82 (99)
T TIGR02656 30 VNNKG-GPHNVVFDEDAVPAGVKEL-AKSLSHKDLLNSPGESYEVTFST----PGTYTF 82 (99)
T ss_pred EECCC-CCceEEECCCCCccchhhh-cccccccccccCCCCEEEEEeCC----CEEEEE
Confidence 38764 6799999887777665432 12222222457899999887663 465544
No 39
>PRK02710 plastocyanin; Provisional
Probab=60.37 E-value=76 Score=24.36 Aligned_cols=20 Identities=15% Similarity=0.295 Sum_probs=13.1
Q ss_pred EecCCCceEEEEEEEecceeeEee
Q 029361 91 RLDAGGILSHSFELDAKVKGMFHG 114 (194)
Q Consensus 91 rI~pg~nvsH~vvv~Pk~~G~fn~ 114 (194)
.+.||+..+++|.- .|.|.|
T Consensus 83 ~~~pg~t~~~tF~~----~G~y~y 102 (119)
T PRK02710 83 AFAPGESWEETFSE----AGTYTY 102 (119)
T ss_pred ccCCCCEEEEEecC----CEEEEE
Confidence 46788888776663 465544
No 40
>PF13860 FlgD_ig: FlgD Ig-like domain; PDB: 3C12_A 3OSV_A.
Probab=58.09 E-value=29 Score=24.81 Aligned_cols=39 Identities=23% Similarity=0.394 Sum_probs=21.5
Q ss_pred EEEEEEEecCCcceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCc
Q 029361 50 SVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGI 97 (194)
Q Consensus 50 tV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~n 97 (194)
.++.+|||.-......+++.. .-.|.-+..|+-.....+
T Consensus 26 ~v~v~I~d~~G~~V~t~~~~~---------~~~G~~~~~WdG~d~~G~ 64 (81)
T PF13860_consen 26 NVTVTIYDSNGQVVRTISLGS---------QSAGEHSFTWDGKDDDGN 64 (81)
T ss_dssp EEEEEEEETTS-EEEEEEEEE---------CSSEEEEEEE-SB-TTS-
T ss_pred EEEEEEEcCCCCEEEEEEcCC---------cCCceEEEEECCCCCCcC
Confidence 456677777655656666543 344677888885555443
No 41
>KOG1691 consensus emp24/gp25L/p24 family of membrane trafficking proteins [Intracellular trafficking, secretion, and vesicular transport]
Probab=57.58 E-value=40 Score=29.54 Aligned_cols=32 Identities=13% Similarity=0.141 Sum_probs=22.0
Q ss_pred cccccccceeEEEEEEEEecCCc--ceeeeEEecC
Q 029361 39 LKRLKSGAERISVSIDIHNQGTS--TAYDVSLTDD 71 (194)
Q Consensus 39 ~~~~v~g~~ditV~ytIYNvG~s--~A~dV~L~D~ 71 (194)
.+++-++ .=++-+|.+.|.-+. +..+|.++|+
T Consensus 36 ~EeI~~n-~lv~g~y~i~~~~~~~~~~~~~~Vts~ 69 (210)
T KOG1691|consen 36 SEEIHEN-VLVVGDYEIINPNGDHSHKLSVKVTSP 69 (210)
T ss_pred hhhhccC-eEEEEEEEEecCCCCccceEEEEEEcC
Confidence 4444444 444448999987666 6899999984
No 42
>PF07919 Gryzun: Gryzun, putative trafficking through Golgi; InterPro: IPR012880 The proteins featured in this family are all hypothetical eukaryotic proteins of unknown function. The region in question is approximately 150 residues long.
Probab=56.33 E-value=1.2e+02 Score=28.37 Aligned_cols=85 Identities=15% Similarity=0.205 Sum_probs=58.8
Q ss_pred cccccccccceeEEEEEEEEecCCcceeee---EEe---cCCC-CCCCeeee----c----C---ceeeEEEEecCCCce
Q 029361 37 ASLKRLKSGAERISVSIDIHNQGTSTAYDV---SLT---DDSW-PQDKFDVI----S----G---NISQSWERLDAGGIL 98 (194)
Q Consensus 37 i~~~~~v~g~~ditV~ytIYNvG~s~A~dV---~L~---D~sf-p~e~Felv----~----G---~~s~~~erI~pg~nv 98 (194)
-...-+..| +.+.+.++|.|..+..+..+ .+. +..+ .++.=++. . + ........|++|+..
T Consensus 181 ~~~~~~l~g-E~~~i~i~I~n~e~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~lg~l~~~~s~ 259 (554)
T PF07919_consen 181 NHKPPALTG-EFYPIPITISNNEDEEASGVLEVRLLHPSQLGVSSEETEDLSQVNWDSDKDDEPLFLGIPLGELAPGSSI 259 (554)
T ss_pred CCCCCeEcC-CEEEEEEEEEcCCCccceeEEEEEEecccccccccccCccceecccccccccchhccCcccccCCCCCcE
Confidence 345666788 99999999999998766533 233 0111 11111121 0 1 134567789999999
Q ss_pred EEEEEEEecceeeEeeecEEEEEEc
Q 029361 99 SHSFELDAKVKGMFHGSPALITFRI 123 (194)
Q Consensus 99 sH~vvv~Pk~~G~fn~t~A~VtY~~ 123 (194)
++.+.++....|.+.+. -.++|..
T Consensus 260 ~~~l~i~~~~~~~~~L~-i~~~Y~l 283 (554)
T PF07919_consen 260 TVTLYIRTSRPGEYELS-ISVSYHL 283 (554)
T ss_pred EEEEEEEeCCceeEEEE-EEEEEEE
Confidence 99999999999999998 6678876
No 43
>PF08626 TRAPPC9-Trs120: Transport protein Trs120 or TRAPPC9, TRAPP II complex subunit; InterPro: IPR013935 The trafficking protein particle complex TRAPP is a multi-protein complex needed in the early stages of the secretory pathway. To date, two kinds of TRAPP complexes have been studied, TRAPPI and TRAPP II. These complexes differ in subunit composition []. TRAPP I binds vesicles derived from the endoplasmic reticulum bringing them closer to the acceptor membrane. Trs120 is a subunit specific to the TRAPP II complex [] along with Trs65p and Trs130p(TRAPPC10). It is suggested that Trs120p is required for the stability of the Trs130p subunit, suggesting that these two proteins might interact in some way []. It is likely that there is a complex function for TRAPP II in multiple pathways [].
Probab=55.16 E-value=89 Score=33.21 Aligned_cols=98 Identities=15% Similarity=0.229 Sum_probs=61.1
Q ss_pred cCCCceEEEEeeccc---ccccccceeEEEEEEEEecCCcceeeeEEe--cCCC-------------CCCCeeeecCce-
Q 029361 25 SSDVPFIVAHKKASL---KRLKSGAERISVSIDIHNQGTSTAYDVSLT--DDSW-------------PQDKFDVISGNI- 85 (194)
Q Consensus 25 ~~~~a~LlvsK~i~~---~~~v~g~~ditV~ytIYNvG~s~A~dV~L~--D~sf-------------p~e~Felv~G~~- 85 (194)
..+.|.|-+...-+. -.+.+| +.-+++++|.|.|+.|.-.+.++ |..- |.|-+|+---..
T Consensus 775 Ip~qP~L~v~~~sl~~~~~mlleG-E~~~~~ItL~N~S~~pvd~l~~sf~DS~~~~~~~~l~~k~l~~~e~yelE~~l~~ 853 (1185)
T PF08626_consen 775 IPPQPLLEVKSSSLTQGALMLLEG-EKQTFTITLRNTSSVPVDFLSFSFQDSTIEPLQKALSNKDLSPDELYELEWQLFK 853 (1185)
T ss_pred ECCCCeEEEEeccCCCcceEEECC-cEEEEEEEEEECCccccceEEEEEEeccHHHHhhhhhcccCChhhhhhhhhhhhc
Confidence 456677777775222 245788 99999999999998888777766 4111 122222211111
Q ss_pred --eeEE---EEecCCCceEEEEEEEecceeeEeeecEE--EEEEcC
Q 029361 86 --SQSW---ERLDAGGILSHSFELDAKVKGMFHGSPAL--ITFRIP 124 (194)
Q Consensus 86 --s~~~---erI~pg~nvsH~vvv~Pk~~G~fn~t~A~--VtY~~s 124 (194)
..+| +.|+||+.++-.+.+.-+ .|.+.++.+. +.|...
T Consensus 854 ~~~~~i~~~~~I~Pg~~~~~~~~~~~~-~~~~~~~~~~i~l~y~~~ 898 (1185)
T PF08626_consen 854 LPAFRILNKPPIPPGESATFTVEVDGK-PGPIQLTYADIQLEYGYS 898 (1185)
T ss_pred CcceeecccCccCCCCEEEEEEEecCc-ccccceeeeeEEEEeccc
Confidence 1233 389999999999997644 4555555554 477643
No 44
>PF07760 DUF1616: Protein of unknown function (DUF1616); InterPro: IPR011674 This is a group of sequences from hypothetical archaeal proteins. The region in question is approximately 330 amino acid residues long.
Probab=54.99 E-value=1.3e+02 Score=26.65 Aligned_cols=66 Identities=18% Similarity=0.379 Sum_probs=46.4
Q ss_pred ccccccceeEEEEEEEEecCCc-ceeeeEE--ecCCCCCCCeeeecCce--eeEEEEecCCCceEEEEEEEec
Q 029361 40 KRLKSGAERISVSIDIHNQGTS-TAYDVSL--TDDSWPQDKFDVISGNI--SQSWERLDAGGILSHSFELDAK 107 (194)
Q Consensus 40 ~~~v~g~~ditV~ytIYNvG~s-~A~dV~L--~D~sfp~e~Felv~G~~--s~~~erI~pg~nvsH~vvv~Pk 107 (194)
+.+..| ++.++...|+|.... ..|.|++ .+..|.++...+..... .... .|+.|++.+..+.+.|.
T Consensus 185 t~l~~g-e~~~v~vgI~NhE~~~~~Ytv~v~l~~~~~~~~~~~~~~~~~l~~~~~-~L~~n~t~~~~~~~~~~ 255 (287)
T PF07760_consen 185 TNLTSG-EPGTVIVGIENHEGRPENYTVVVVLQNVTWNPNNYNVMESTVLDRPIV-TLADNETWEQPYKFTPF 255 (287)
T ss_pred eeEEcC-CcEEEEEEEEcCCCCcEEEEEEEEEeccccccccccccchhcccceEE-EeCCCCeEEEEEEEEEe
Confidence 344577 999999999998754 5555554 55666655555544443 3333 89999999999999983
No 45
>PRK15188 fimbrial chaperone protein BcfB; Provisional
Probab=54.40 E-value=1.1e+02 Score=26.71 Aligned_cols=59 Identities=12% Similarity=0.047 Sum_probs=30.4
Q ss_pred cccccceeEEEEEEEEecCCcceeeeEE-ec--CCCCCCCeeeecCceeeEEEEecCCCceEEEEEE
Q 029361 41 RLKSGAERISVSIDIHNQGTSTAYDVSL-TD--DSWPQDKFDVISGNISQSWERLDAGGILSHSFEL 104 (194)
Q Consensus 41 ~~v~g~~ditV~ytIYNvG~s~A~dV~L-~D--~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv 104 (194)
+++-.+.+-.++++|.|.|+.+.+=|+- +| ++=....|-+ +--+-||+||+.-+-.+..
T Consensus 35 RvIy~~~~~~~sv~i~N~~~~~p~LvQsWv~~~~~~~~~pFiv-----tPPlfrl~~~~~~~lRI~~ 96 (228)
T PRK15188 35 RVIYPQGSKQTSLPIINSSASNVFLIQSWVANADGSRSTDFII-----TPPLFVIQPKKENILRIMY 96 (228)
T ss_pred EEEEcCCCceEEEEEEeCCCCccEEEEEEEecCCCCccCCEEE-----cCCeEEECCCCceEEEEEE
Confidence 3333336678899999999765443432 22 1111123421 2235566666665555443
No 46
>PF11797 DUF3324: Protein of unknown function C-terminal (DUF3324); InterPro: IPR021759 This family consists of several hypothetical bacterial proteins of unknown function.
Probab=53.30 E-value=40 Score=26.77 Aligned_cols=53 Identities=13% Similarity=0.307 Sum_probs=37.4
Q ss_pred cccccccccceeEEEEEEEEecCCc-ceeeeEEecCCC-CCCCeeeecCceeeEE--EEecCC
Q 029361 37 ASLKRLKSGAERISVSIDIHNQGTS-TAYDVSLTDDSW-PQDKFDVISGNISQSW--ERLDAG 95 (194)
Q Consensus 37 i~~~~~v~g~~ditV~ytIYNvG~s-~A~dV~L~D~sf-p~e~Felv~G~~s~~~--erI~pg 95 (194)
+.|..+..= .++++++.||+.|+. .-+.-+..+-.+ |...|++ ...| ++|+||
T Consensus 50 l~N~~~~~l-~~~~v~a~V~~~~~~k~~~~~~~~~~~mAPNS~f~~-----~i~~~~~~lk~G 106 (140)
T PF11797_consen 50 LQNPQPAIL-KKLTVDAKVTKKGSKKVLYTFKKENMQMAPNSNFNF-----PIPLGGKKLKPG 106 (140)
T ss_pred EECCCchhh-cCcEEEEEEEECCCCeEEEEeeccCCEECCCCeEEe-----EecCCCcCccCC
Confidence 667777777 789999999999975 666666666556 3445644 4455 477777
No 47
>TIGR03096 nitroso_cyanin nitrosocyanin. Nitrosocyanin, as described from the obligate chemolithoautotroph Nitrosomonas europaea, is a red copper protein of unknown function with sequence similarity to a number of blue copper redox proteins.
Probab=52.95 E-value=1.2e+02 Score=24.60 Aligned_cols=23 Identities=26% Similarity=0.356 Sum_probs=10.8
Q ss_pred EEecCCCceEEEEEEEecceeeEee
Q 029361 90 ERLDAGGILSHSFELDAKVKGMFHG 114 (194)
Q Consensus 90 erI~pg~nvsH~vvv~Pk~~G~fn~ 114 (194)
..|+||+..+ +...|.+.|.|.|
T Consensus 94 ~~I~pGet~T--itF~adKpG~Y~y 116 (135)
T TIGR03096 94 EVIKAGETKT--ISFKADKAGAFTI 116 (135)
T ss_pred eEECCCCeEE--EEEECCCCEEEEE
Confidence 3455554433 3344555555543
No 48
>PF14263 DUF4354: Domain of unknown function (DUF4354); PDB: 3NRF_B 3SB3_A.
Probab=52.39 E-value=39 Score=27.27 Aligned_cols=93 Identities=14% Similarity=0.173 Sum_probs=40.5
Q ss_pred HHHHHHHHhhhcccCCCceEEEEeeccccccccccee---EEEEEEEEecCCcceeeeEEec---CCCCCCCeeeecCce
Q 029361 12 VLIALFLISSSFASSDVPFIVAHKKASLKRLKSGAER---ISVSIDIHNQGTSTAYDVSLTD---DSWPQDKFDVISGNI 85 (194)
Q Consensus 12 ~lla~~~v~~~~~~~~~a~LlvsK~i~~~~~v~g~~d---itV~ytIYNvG~s~A~dV~L~D---~sfp~e~Felv~G~~ 85 (194)
++|+.+|....+...+.-.+.+.++-...--+.| +. -++++.+.|.++.+. +|.. .-|.++.-++.-...
T Consensus 10 ~~l~~~~~~a~a~~~d~i~V~At~~~~Gs~sv~~-k~~ytktF~V~vaN~s~~~i---dLsk~Cf~a~~~~gk~f~ldTV 85 (124)
T PF14263_consen 10 VALASFSFSANASAPDNIAVYATEKSQGSVSVGG-KSFYTKTFDVTVANLSDKDI---DLSKMCFKAYSPDGKEFKLDTV 85 (124)
T ss_dssp ---------------SSEEEEEEEEEEEEEEETT-EEEEEEEEEEEEEE-SSS-E---E-TT-EEEEEETTS-EEEEEEE
T ss_pred HHHHHHHHhhhhccCCCeEEEEEecCCccEeecC-ccceEEEEEEEEecCCCCcc---ccccchhhhccccCCEEEeccc
Confidence 4444444433344445455777776644443334 43 578889999999764 4443 223444333333222
Q ss_pred eeEE--EEecCCCceEEEEEEEecc
Q 029361 86 SQSW--ERLDAGGILSHSFELDAKV 108 (194)
Q Consensus 86 s~~~--erI~pg~nvsH~vvv~Pk~ 108 (194)
..+. ..|.||+++.=.++--...
T Consensus 86 d~~L~~g~lK~g~s~kG~avFaS~d 110 (124)
T PF14263_consen 86 DEELTSGTLKPGESVKGIAVFASDD 110 (124)
T ss_dssp -GGGG-SEE-TT-EEEEEEEEEESS
T ss_pred chhhhhccccCCCceeEEEEEeeCC
Confidence 2222 4688999988777665443
No 49
>TIGR02745 ccoG_rdxA_fixG cytochrome c oxidase accessory protein FixG. Member of this ferredoxin-like protein family are found exclusively in species with an operon encoding the cbb3 type of cytochrome c oxidase (cco-cbb3), and near the cco-cbb3 operon in about half the cases. The cco-cbb3 is found in a variety of proteobacteria and almost nowhere else, and is associated with oxygen use under microaerobic conditions. Some (but not all) of these proteobacteria are also nitrogen-fixing, hence the gene symbol fixG. FixG was shown essential for functional cco-cbb3 expression in Bradyrhizobium japonicum.
Probab=49.51 E-value=2.4e+02 Score=26.93 Aligned_cols=55 Identities=20% Similarity=0.257 Sum_probs=36.0
Q ss_pred cceeEEEEEEEEecCCc-ceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEe
Q 029361 45 GAERISVSIDIHNQGTS-TAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHSFELDA 106 (194)
Q Consensus 45 g~~ditV~ytIYNvG~s-~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~P 106 (194)
|.-+...++.|.|-.+. -.+++++.+ -|....+ +.-+ . =.++||+..+..+.|+.
T Consensus 344 g~i~N~Y~~~i~Nk~~~~~~~~l~v~g--~~~~~~~---~~~~-~-i~v~~g~~~~~~v~v~~ 399 (434)
T TIGR02745 344 GVVENTYTLKILNKTEQPHEYYLSVLG--LPGIKIE---GPGA-P-IHVKAGEKVKLPVFLRT 399 (434)
T ss_pred CcEEEEEEEEEEECCCCCEEEEEEEec--CCCcEEE---cCCc-e-EEECCCCEEEEEEEEEe
Confidence 44577889999998876 355555554 3443222 1111 2 27999999999999985
No 50
>PRK14740 kdbF potassium-transporting ATPase subunit F; Provisional
Probab=48.06 E-value=2.1 Score=26.47 Aligned_cols=18 Identities=33% Similarity=0.401 Sum_probs=15.7
Q ss_pred chhhhhHHHheeeeEEEe
Q 029361 163 GSQISVISIIVLFVYLIT 180 (194)
Q Consensus 163 ~~~~~v~s~~~~~v~~~~ 180 (194)
.+|++--+.+++|+||++
T Consensus 4 ~~wls~a~a~~Lf~YLv~ 21 (29)
T PRK14740 4 LDWLSLALATGLFVYLLV 21 (29)
T ss_pred HHHHHHHHHHHHHHHHHH
Confidence 578999999999999975
No 51
>PF04495 GRASP55_65: GRASP55/65 PDZ-like domain ; InterPro: IPR007583 GRASP55 (Golgi reassembly stacking protein of 55 kDa) and GRASP65 (a 65 kDa) protein are highly homologous. GRASP55 is a component of the Golgi stacking machinery. GRASP65, an N-ethylmaleimide-sensitive membrane protein required for the stacking of Golgi cisternae in a cell-free system [].; PDB: 3RLE_A 4EDJ_A.
Probab=48.06 E-value=76 Score=25.54 Aligned_cols=49 Identities=22% Similarity=0.333 Sum_probs=37.0
Q ss_pred EEEecCCcceeeeEEec-CCCCCCCeeeecCc--eeeEEEEec-CCCceEEEEEEEecc
Q 029361 54 DIHNQGTSTAYDVSLTD-DSWPQDKFDVISGN--ISQSWERLD-AGGILSHSFELDAKV 108 (194)
Q Consensus 54 tIYNvG~s~A~dV~L~D-~sfp~e~Felv~G~--~s~~~erI~-pg~nvsH~vvv~Pk~ 108 (194)
++||.=...-.||.++. +.|... |. .+++|+... ..+..-|.+.|.|-.
T Consensus 2 ~v~~~k~~~~R~v~i~ps~~w~~~------g~LG~sv~~~~~~~~~~~~~~Vl~V~p~S 54 (138)
T PF04495_consen 2 NVYNAKGQTTREVSIVPSKKWGGQ------GLLGISVRFESFEGAEEEGWHVLRVAPNS 54 (138)
T ss_dssp EEEETTTSSEEEEEE---SSSSSS------SSS-EEEEEEE-TTGCCCEEEEEEE-TTS
T ss_pred ceEECCCCeEEEEEEccCcccCCC------CCCcEEEEEecccccccceEEEeEecCCC
Confidence 68999999999999987 556554 55 499999999 788888988888654
No 52
>PF00127 Copper-bind: Copper binding proteins, plastocyanin/azurin family; InterPro: IPR000923 Blue (type 1) copper proteins are small proteins which bind a single copper atom and which are characterised by an intense electronic absorption band near 600 nm [, ]. The most well known members of this class of proteins are the plant chloroplastic plastocyanins, which exchange electrons with cytochrome c6, and the distantly related bacterial azurins, which exchange electrons with cytochrome c551. This family of proteins also includes amicyanin from bacteria such as Methylobacterium extorquens or Paracoccus versutus (Thiobacillus versutus) that can grow on methylamine; auracyanins A and B from Chloroflexus aurantiacus []; blue copper protein from Alcaligenes faecalis; cupredoxin (CPC) from Cucumis sativus (Cucumber) peelings []; cusacyanin (basic blue protein; plantacyanin, CBP) from cucumber; halocyanin from Natronomonas pharaonis (Natronobacterium pharaonis) [], a membrane associated copper-binding protein; pseudoazurin from Pseudomonas; rusticyanin from Thiobacillus ferrooxidans []; stellacyanin from Rhus vernicifera (Japanese lacquer tree); umecyanin from the roots of Armoracia rusticana (Horseradish); and allergen Ra3 from ragweed. This pollen protein is evolutionary related to the above proteins, but seems to have lost the ability to bind copper. Although there is an appreciable amount of divergence in the sequences of all these proteins, the copper ligand sites are conserved.; GO: 0005507 copper ion binding, 0009055 electron carrier activity; PDB: 1UAT_A 1CUO_A 1PLC_A 4PCY_A 3PCY_A 1PND_A 1PNC_A 1JXG_A 6PCY_A 1TKW_A ....
Probab=47.20 E-value=42 Score=24.65 Aligned_cols=45 Identities=29% Similarity=0.316 Sum_probs=27.6
Q ss_pred EEecCCcceeeeEEecCCCCC--CCeeeecCceeeEEEEecCCCceEEEEE
Q 029361 55 IHNQGTSTAYDVSLTDDSWPQ--DKFDVISGNISQSWERLDAGGILSHSFE 103 (194)
Q Consensus 55 IYNvG~s~A~dV~L~D~sfp~--e~Felv~G~~s~~~erI~pg~nvsH~vv 103 (194)
.-|. +...-++.+.+++++. +.+..-.+. .=..+.||++.+++|.
T Consensus 29 ~~n~-~~~~Hnv~~~~~~~~~~~~~~~~~~~~---~~~~~~~G~~~~~tF~ 75 (99)
T PF00127_consen 29 FVNN-DSMPHNVVFVADGMPAGADSDYVPPGD---SSPLLAPGETYSVTFT 75 (99)
T ss_dssp EEEE-SSSSBEEEEETTSSHTTGGHCHHSTTC---EEEEBSTTEEEEEEEE
T ss_pred EEEC-CCCCceEEEecccccccccccccCccc---cceecCCCCEEEEEeC
Confidence 4455 4456888888876644 222222222 4456889999888877
No 53
>COG1470 Predicted membrane protein [Function unknown]
Probab=46.58 E-value=1.3e+02 Score=29.68 Aligned_cols=72 Identities=17% Similarity=0.343 Sum_probs=55.6
Q ss_pred eeEEEEEEEEecCCc-ceeeeEEecCCCCCC-CeeeecCceeeEEEEecCCCceEEEEEEEecc---eeeEeeecEEEE
Q 029361 47 ERISVSIDIHNQGTS-TAYDVSLTDDSWPQD-KFDVISGNISQSWERLDAGGILSHSFELDAKV---KGMFHGSPALIT 120 (194)
Q Consensus 47 ~ditV~ytIYNvG~s-~A~dV~L~D~sfp~e-~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~---~G~fn~t~A~Vt 120 (194)
....+..+|=|-|.. .-|+.++. ++|+. ..++..|....+==.|.||+.-.-++.|+|.. .|.||++-+..+
T Consensus 284 ~t~sf~V~IeN~g~~~d~y~Le~~--g~pe~w~~~Fteg~~~vt~vkL~~gE~kdvtleV~ps~na~pG~Ynv~I~A~s 360 (513)
T COG1470 284 TTASFTVSIENRGKQDDEYALELS--GLPEGWTAEFTEGELRVTSVKLKPGEEKDVTLEVYPSLNATPGTYNVTITASS 360 (513)
T ss_pred CceEEEEEEccCCCCCceeEEEec--cCCCCcceEEeeCceEEEEEEecCCCceEEEEEEecCCCCCCCceeEEEEEec
Confidence 455778889999987 45666665 34443 12356999999999999999999999999864 699999887766
No 54
>PRK15208 long polar fimbrial chaperone LpfB; Provisional
Probab=46.17 E-value=1.4e+02 Score=25.75 Aligned_cols=52 Identities=17% Similarity=0.257 Sum_probs=27.5
Q ss_pred eeEEEEEEEEecCCcceeee-EEecCCCC--CCCeeeecCceeeEEEEecCCCceEEEEE
Q 029361 47 ERISVSIDIHNQGTSTAYDV-SLTDDSWP--QDKFDVISGNISQSWERLDAGGILSHSFE 103 (194)
Q Consensus 47 ~ditV~ytIYNvG~s~A~dV-~L~D~sfp--~e~Felv~G~~s~~~erI~pg~nvsH~vv 103 (194)
.+-.++++|.|.|+...+=| .-+|++=. ...| ++ +--+-||+||+.-+-.++
T Consensus 35 ~~~~~si~i~N~~~~~~~LvQsWv~~~~~~~~~pf-iv----tPPl~rl~p~~~q~lRIi 89 (228)
T PRK15208 35 SKKEASLTVNNKSKTEEFLIQSWIDDANGNKKTPF-II----TPPLFKLDPTKNNVLRIV 89 (228)
T ss_pred CCceEEEEEEeCCCCCcEEEEEEEECCCCCccCCE-EE----CCCeEEECCCCccEEEEE
Confidence 56678899999997644444 33332111 1124 11 223556666666555544
No 55
>PRK15290 lfpB fimbrial chaperone protein; Provisional
Probab=44.02 E-value=1.6e+02 Score=25.92 Aligned_cols=55 Identities=15% Similarity=0.133 Sum_probs=30.4
Q ss_pred hhhHHHHHHHHHHHHhhhcccCCCceEEEEeecccccccccceeEEEEEEEEecCCcceee
Q 029361 5 ISKSLISVLIALFLISSSFASSDVPFIVAHKKASLKRLKSGAERISVSIDIHNQGTSTAYD 65 (194)
Q Consensus 5 ~~~~~~~~lla~~~v~~~~~~~~~a~LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~d 65 (194)
.+|.++.+++.++++++.+++. |-|.++ ..+++-.+.+-.++++|.|.++...+=
T Consensus 15 ~~~~~~~~~~~~~~~~~~~~a~--Agv~l~----~TRvIy~~~~~~~sl~v~N~~~~~p~L 69 (243)
T PRK15290 15 VSCKLFTAIILSVFLGQPALTY--AGVVIG----GTRVVYLSNNPDKSISVFSKEEKIPYL 69 (243)
T ss_pred HHHhHHHHHHHHHHHhchhhhe--EeEEEC----ceEEEEeCCCceEEEEEEeCCCCCcEE
Confidence 4556666655555554443332 223333 334444446777889999999754333
No 56
>cd08547 Type_II_cohesin Type II cohesin domain, interaction partner of dockerin. Bacterial cohesin domains bind to a complementary protein domain named dockerin, and this interaction is required for the formation of the cellulosome, a cellulose-degrading complex. The cellulosome consists of scaffoldin, a noncatalytic scaffolding polypeptide, that comprises repeating cohesion modules and a single carbohydrate-binding module (CBM). Specific calcium-dependent interactions between cohesins and dockerins appear to be essential for cellulosome assembly. This subfamily represents type II cohesins; their interactions with dockerin mediate attachment of the cellulosome complex to the bacterial cell wall.
Probab=42.59 E-value=1.5e+02 Score=22.37 Aligned_cols=39 Identities=21% Similarity=0.483 Sum_probs=32.8
Q ss_pred ccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCc
Q 029361 42 LKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGN 84 (194)
Q Consensus 42 ~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~ 84 (194)
...| +.++|...+-|...-.+++++|. |+++.+++++..
T Consensus 12 v~~G-~~~~v~v~~~~~~~~~~~~~~l~---YD~~~l~~~~~~ 50 (132)
T cd08547 12 VKVG-ETFTVTVKVNNATNLAGYQFTLS---YDPSVLEFVSVT 50 (132)
T ss_pred cCCC-CEEEEEEEEeccCceEEEEEEEE---ECcceEEEEecc
Confidence 5678 99999999999997788888886 778888888754
No 57
>cd08379 C2D_MCTP_PRT_plant C2 domain fourth repeat found in Multiple C2 domain and Transmembrane region Proteins (MCTP); plant subset. MCTPs are involved in Ca2+ signaling at the membrane. Plant-MCTPs are composed of a variable N-terminal sequence, four C2 domains, two transmembrane regions (TMRs), and a short C-terminal sequence. It is one of four protein classes that are anchored to membranes via a transmembrane region; the others being synaptotagmins, extended synaptotagmins, and ferlins. MCTPs are the only membrane-bound C2 domain proteins that contain two functional TMRs. MCTPs are unique in that they bind Ca2+ but not phospholipids. C2 domains fold into an 8-standed beta-sandwich that can adopt 2 structural arrangements: Type I and Type II, distinguished by a circular permutation involving their N- and C-terminal beta strands. Many C2 domains are Ca2+-dependent membrane-targeting modules that bind a wide variety of substances including bind phospholipids, inositol polyphosphate
Probab=42.53 E-value=89 Score=24.42 Aligned_cols=54 Identities=19% Similarity=0.289 Sum_probs=37.8
Q ss_pred EEEEEEEecCCcceeeeEEecC-CC------CCCCeeeecCceeeEEEEecCCCceEEEEEEEecc
Q 029361 50 SVSIDIHNQGTSTAYDVSLTDD-SW------PQDKFDVISGNISQSWERLDAGGILSHSFELDAKV 108 (194)
Q Consensus 50 tV~ytIYNvG~s~A~dV~L~D~-sf------p~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~ 108 (194)
++.+.+.+ .....+|++-|. .+ ..++| + |........+.+|....|.|.|++..
T Consensus 53 ~f~f~v~~--~~~~l~v~V~d~d~~~~~~~~~~dd~--l-G~~~i~l~~l~~~~~~~~~~~L~~~~ 113 (126)
T cd08379 53 QYTWPVYD--PCTVLTVGVFDNSQSHWKEAVQPDVL--I-GKVRIRLSTLEDDRVYAHSYPLLSLN 113 (126)
T ss_pred EEEEEecC--CCCEEEEEEEECCCccccccCCCCce--E-EEEEEEHHHccCCCEEeeEEEeEeCC
Confidence 33444443 345899999883 33 24544 3 78888888999999999999999654
No 58
>PF00635 Motile_Sperm: MSP (Major sperm protein) domain; InterPro: IPR000535 Major sperm proteins (MSP) are central components in molecular interactions underlying sperm motility in Caenorhabditis elegans, whose sperm employ an amoebae-like crawling motion using a MSP-containing lamellipod, rather than the flagellar-based swimming motion associated with other sperm. These proteins oligomerise to form an extensive filament system that extends from sperm villipoda, along the leading edge of the pseudopod. About 30 MSP isoforms may exist in C. elegans. MSPs form a fibrous network, whereby MSP dimers form helical subfilaments that coil around one another to produce filaments, which in turn form supercoils to produce bundles. The crystal structure of MSP from C. elegans reveals an immunoglobulin (Ig)-like seven-stranded beta sandwich fold []. ; GO: 0005198 structural molecule activity; PDB: 1MSP_A 3MSP_B 2BVU_B 2MSP_C 1Z9O_F 1Z9L_A 3IKK_A 1WIC_A 2CRI_A 2RR3_A ....
Probab=42.03 E-value=1.2e+02 Score=21.75 Aligned_cols=51 Identities=12% Similarity=0.373 Sum_probs=35.9
Q ss_pred eeEEEEEEEEecCCc-ceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEe
Q 029361 47 ERISVSIDIHNQGTS-TAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHSFELDA 106 (194)
Q Consensus 47 ~ditV~ytIYNvG~s-~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~P 106 (194)
+......+|.|.++. -|+.|+-+.+ +.|. ..-...-|.||+.++-.++..|
T Consensus 18 ~~~~~~l~l~N~s~~~i~fKiktt~~----~~y~-----v~P~~G~i~p~~~~~i~I~~~~ 69 (109)
T PF00635_consen 18 KQQSCELTLTNPSDKPIAFKIKTTNP----NRYR-----VKPSYGIIEPGESVEITITFQP 69 (109)
T ss_dssp S-EEEEEEEEE-SSSEEEEEEEES-T----TTEE-----EESSEEEE-TTEEEEEEEEE-S
T ss_pred ceEEEEEEEECCCCCcEEEEEEcCCC----ceEE-----ecCCCEEECCCCEEEEEEEEEe
Confidence 678899999999998 6788887754 4553 2345688999999999998887
No 59
>PF00630 Filamin: Filamin/ABP280 repeat; InterPro: IPR017868 The many different actin cross-linking proteins share a common architecture, consisting of a globular actin-binding domain and an extended rod. Whereas their actin-binding domains consist of two calponin homology domains (see IPR001715 from INTERPRO), their rods fall into three families. The rod domain of the family including the Dictyostelium discoideum (Slime mould) gelation factor (ABP120) and human filamin (ABP280) is constructed from tandem repeats of a 100-residue motif that is glycine and proline rich []. The gelation factor's rod contains 6 copies of the repeat, whereas filamin has a rod constructed from 24 repeats. The resolution of the 3D structure of rod repeats from the gelation factor has shown that they consist of a beta-sandwich, formed by two beta-sheets arranged in an immunoglobulin-like fold [, ]. Because conserved residues that form the core of the repeats are preserved in filamin, the repeat structure should be common to the members of the gelation factor/filamin family. The head to tail homodimerisation is crucial to the function of the ABP120 and ABP280 proteins. This interaction involves a small portion at the distal end of the rod domains. For the gelation factor it has been shown that the carboxy-terminal repeat 6 dimerises through a double edge-to-edge extension of the beta-sheet and that repeat 5 contributes to dimerisation to some extent [, , ].; PDB: 2DI9_A 2EEC_A 2DIC_A 2EEA_A 2DMC_A 2EE9_A 2D7O_A 2D7N_A 2K7P_A 2NQC_A ....
Probab=41.37 E-value=1.3e+02 Score=21.37 Aligned_cols=67 Identities=15% Similarity=0.238 Sum_probs=42.8
Q ss_pred ccccccceeEEEEEEEEecCCcc------eeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEecceeeEe
Q 029361 40 KRLKSGAERISVSIDIHNQGTST------AYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHSFELDAKVKGMFH 113 (194)
Q Consensus 40 ~~~v~g~~ditV~ytIYNvG~s~------A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~~G~fn 113 (194)
+....| +..++.++..+.+..+ ...|++.+++=..+. ....++ +....+=++.+.-+|+..|.|+
T Consensus 15 ~~~~~g-~~~~F~V~~~d~~g~~~~~~~~~~~v~i~~p~~~~~~-------~~~~~~-v~~~~~G~y~v~y~p~~~G~y~ 85 (101)
T PF00630_consen 15 EPAVVG-EPATFTVDTRDAGGNPVSSGGDEFQVTITSPDGKEEP-------VPVPVE-VIDNGDGTYTVSYTPTEPGKYK 85 (101)
T ss_dssp TEEETT-SEEEEEEEETTTTSSBEESTSSEEEEEEESSSSESS---------EEEEE-EEEESSSEEEEEEEESSSEEEE
T ss_pred CCeECC-CcEEEEEEEccCCCCccccCCceeEEEEeCCCCCccc-------cccceE-EEECCCCEEEEEEEeCccEeEE
Confidence 445788 9999999999996553 356777664111100 033343 2233444888888899999988
Q ss_pred ee
Q 029361 114 GS 115 (194)
Q Consensus 114 ~t 115 (194)
..
T Consensus 86 i~ 87 (101)
T PF00630_consen 86 IS 87 (101)
T ss_dssp EE
T ss_pred EE
Confidence 75
No 60
>PF00207 A2M: Alpha-2-macroglobulin family; InterPro: IPR001599 This entry contains serum complement C3 and C4 precursors and alpha-macrogrobulins. The alpha-macroglobulin (aM) family of proteins includes protease inhibitors [], typified by the human tetrameric a2-macroglobulin (a2M); they belong to the MEROPS proteinase inhibitor family I39, clan IL. These protease inhibitors share several defining properties, which include (i) the ability to inhibit proteases from all catalytic classes, (ii) the presence of a 'bait region' and a thiol ester, (iii) a similar protease inhibitory mechanism and (iv) the inactivation of the inhibitory capacity by reaction of the thiol ester with small primary amines. aM protease inhibitors inhibit by steric hindrance []. The mechanism involves protease cleavage of the bait region, a segment of the aM that is particularly susceptible to proteolytic cleavage, which initiates a conformational change such that the aM collapses about the protease. In the resulting aM-protease complex, the active site of the protease is sterically shielded, thus substantially decreasing access to protein substrates. Two additional events occur as a consequence of bait region cleavage, namely (i) the h-cysteinyl-g-glutamyl thiol ester becomes highly reactive and (ii) a major conformational change exposes a conserved COOH-terminal receptor binding domain [] (RBD). RBD exposure allows the aM protease complex to bind to clearance receptors and be removed from circulation []. Tetrameric, dimeric, and, more recently, monomeric aM protease inhibitors have been identified [, ].; GO: 0004866 endopeptidase inhibitor activity; PDB: 3KLS_B 3PRX_C 3KM9_B 3PVM_C 3CU7_A 4E0S_A 4A5W_A 2PN5_A 3FRP_G 3HRZ_B ....
Probab=41.19 E-value=48 Score=24.03 Aligned_cols=37 Identities=19% Similarity=0.345 Sum_probs=23.2
Q ss_pred eEEEEeec-----ccccccccceeEEEEEEEEecCCcceeeeEE
Q 029361 30 FIVAHKKA-----SLKRLKSGAERISVSIDIHNQGTSTAYDVSL 68 (194)
Q Consensus 30 ~LlvsK~i-----~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L 68 (194)
.+.+.|.+ +...+..| +.+.+..+|+|.++. ..+|++
T Consensus 49 ~~~v~~p~~i~~~lP~~l~~G-D~~~i~v~v~N~~~~-~~~v~V 90 (92)
T PF00207_consen 49 EITVFKPFFIQLNLPRSLRRG-DQIQIPVTVFNYTDK-DQEVTV 90 (92)
T ss_dssp EEEEB-SEEEEEE--SEEETT-SEEEEEEEEEE-SSS--EEEEE
T ss_pred EEEEEeeEEEEcCCCcEEecC-CEEEEEEEEEeCCCC-CEEEEE
Confidence 45555543 35677888 999999999999874 344443
No 61
>PRK15098 beta-D-glucoside glucohydrolase; Provisional
Probab=39.38 E-value=95 Score=31.50 Aligned_cols=84 Identities=12% Similarity=0.145 Sum_probs=58.3
Q ss_pred eeEEEEEEEEecCCcceeeeEEecCCCCCCCee----eecCceeeEEEEecCCCceEEEEEEEecceeeEeeecEEEEEE
Q 029361 47 ERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFD----VISGNISQSWERLDAGGILSHSFELDAKVKGMFHGSPALITFR 122 (194)
Q Consensus 47 ~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Fe----lv~G~~s~~~erI~pg~nvsH~vvv~Pk~~G~fn~t~A~VtY~ 122 (194)
..++|+.++-|.|+-+..+|.-.=-+.|...-+ -..|= .+. .|+||++.+-+|.|.++..++|+-.. .|.
T Consensus 667 ~~i~v~v~V~NtG~~~G~EVvQlYv~~~~~~~~~P~k~L~gF--~Kv-~L~pGes~~V~~~l~~~~L~~~d~~~---~~~ 740 (765)
T PRK15098 667 GKVTASVTVTNTGKREGATVVQLYLQDVTASMSRPVKELKGF--EKI-MLKPGETQTVSFPIDIEALKFWNQQM---KYV 740 (765)
T ss_pred CeEEEEEEEEECCCCCccEEEEEeccCCCCCCCCHHHhccCc--eeE-eECCCCeEEEEEeecHHHhceECCCC---cEE
Confidence 679999999999999888876543333322110 01111 233 49999999999999999999998753 566
Q ss_pred cCCC-cceeeEeecC
Q 029361 123 IPTK-AALQEAYSTP 136 (194)
Q Consensus 123 ~se~-~~~q~a~Ss~ 136 (194)
.+.+ =.+.+|-||.
T Consensus 741 ~e~G~y~v~vG~ss~ 755 (765)
T PRK15098 741 AEPGKFNVFIGLDSA 755 (765)
T ss_pred EeCceEEEEEECCCC
Confidence 6666 3366777764
No 62
>PF14310 Fn3-like: Fibronectin type III-like domain; PDB: 3ABZ_D 3AC0_D 2X40_A 2X41_A 2X42_A.
Probab=39.08 E-value=28 Score=24.20 Aligned_cols=24 Identities=21% Similarity=0.256 Sum_probs=20.6
Q ss_pred ecCCCceEEEEEEEecceeeEeee
Q 029361 92 LDAGGILSHSFELDAKVKGMFHGS 115 (194)
Q Consensus 92 I~pg~nvsH~vvv~Pk~~G~fn~t 115 (194)
|+||++.+.++.|.|...+.++-.
T Consensus 29 l~pGes~~v~~~l~~~~l~~~d~~ 52 (71)
T PF14310_consen 29 LAPGESKTVSFTLPPEDLAYWDED 52 (71)
T ss_dssp E-TT-EEEEEEEEEHHHHEEEETT
T ss_pred ECCCCEEEEEEEECHHHEeeEcCC
Confidence 999999999999999999998876
No 63
>cd04049 C2_putative_Elicitor-responsive_gene C2 domain present in the putative elicitor-responsive gene. In plants elicitor-responsive proteins are triggered in response to specific elicitor molecules such as glycolproteins, peptides, carbohydrates and lipids. A host of defensive responses are also triggered resulting in localized cell death. Antimicrobial secondary metabolites, such as phytoalexins, or defense-related proteins, including pathogenesis-related (PR) proteins are also produced. There is a single C2 domain present here. C2 domains fold into an 8-standed beta-sandwich that can adopt 2 structural arrangements: Type I and Type II, distinguished by a circular permutation involving their N- and C-terminal beta strands. Many C2 domains are Ca2+-dependent membrane-targeting modules that bind a wide variety of substances including bind phospholipids, inositol polyphosphates, and intracellular proteins. Most C2 domain proteins are either signal transduction enzymes that contai
Probab=38.94 E-value=1e+02 Score=22.94 Aligned_cols=76 Identities=17% Similarity=0.225 Sum_probs=48.2
Q ss_pred cccccceeEEEEEEEEecCCcceeeeEEec-CCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEecceeeEeeecEEE
Q 029361 41 RLKSGAERISVSIDIHNQGTSTAYDVSLTD-DSWPQDKFDVISGNISQSWERLDAGGILSHSFELDAKVKGMFHGSPALI 119 (194)
Q Consensus 41 ~~v~g~~ditV~ytIYNvG~s~A~dV~L~D-~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~~G~fn~t~A~V 119 (194)
.++=+ +.+.+.+.--+........|++-| +.+..++| =|.....++.+..++....-+.+.|. -|++..+..
T Consensus 46 nP~Wn-e~f~f~v~~~~~~~~~~l~v~V~d~~~~~~d~~---iG~~~i~l~~l~~~~~~~~~~~l~p~---~~~~~~~~~ 118 (124)
T cd04049 46 NPEWN-EKFKFTVEYPGWGGDTKLILRIMDKDNFSDDDF---IGEATIHLKGLFEEGVEPGTAELVPA---KYNVVLEDD 118 (124)
T ss_pred CCccc-ceEEEEecCcccCCCCEEEEEEEECccCCCCCe---EEEEEEEhHHhhhCCCCcCceEeecc---ceEEEEece
Confidence 44445 555444332221134677788777 55666654 27788888888888888999999986 345555555
Q ss_pred EEEc
Q 029361 120 TFRI 123 (194)
Q Consensus 120 tY~~ 123 (194)
+|+-
T Consensus 119 ~~~~ 122 (124)
T cd04049 119 TYKG 122 (124)
T ss_pred EEEe
Confidence 7763
No 64
>PF10731 Anophelin: Thrombin inhibitor from mosquito; InterPro: IPR018932 Members of this family are all inhibitors of thrombin, the peptidase that is at the end of the blood coagulation cascade and which creates the clot by cleaving fibrinogen. The interaction between thrombin and fibrinogen involves two different areas of contact - via the thrombin active site and via a second substrate-binding site known as an exosite. The inhibitor acts by blocking the exosite, rather than by interacting with the active site. The inhibitors are from mosquitoes that feed on human blood and which, by inhibiting thrombin, prevent the blood from clotting and keep it flowing.
Probab=37.65 E-value=19 Score=25.94 Aligned_cols=17 Identities=24% Similarity=0.356 Sum_probs=9.9
Q ss_pred HHHHHHHhhhcccCCCc
Q 029361 13 LIALFLISSSFASSDVP 29 (194)
Q Consensus 13 lla~~~v~~~~~~~~~a 29 (194)
+++++|+.+++-.+.+|
T Consensus 7 vialLC~aLva~vQ~AP 23 (65)
T PF10731_consen 7 VIALLCVALVAIVQSAP 23 (65)
T ss_pred HHHHHHHHHHHHHhcCc
Confidence 66777776665344433
No 65
>PF03314 DUF273: Protein of unknown function, DUF273; InterPro: IPR004988 This is a family of proteins of unknown function.
Probab=35.38 E-value=22 Score=31.38 Aligned_cols=48 Identities=23% Similarity=0.385 Sum_probs=41.1
Q ss_pred eeEEEEEEEEecCCcceeeeEEecCCCCCC-CeeeecCceeeEEEEecCCC
Q 029361 47 ERISVSIDIHNQGTSTAYDVSLTDDSWPQD-KFDVISGNISQSWERLDAGG 96 (194)
Q Consensus 47 ~ditV~ytIYNvG~s~A~dV~L~D~sfp~e-~Felv~G~~s~~~erI~pg~ 96 (194)
.+.- +..|+.-|.+=|.|.=|+|..|+++ +| +.+|--......+|.|.
T Consensus 168 t~F~-kvrIl~KGtgWaRD~WLT~s~Ws~~~DF-MlHGwK~~~l~~~p~~~ 216 (222)
T PF03314_consen 168 TDFP-KVRILKKGTGWARDGWLTSSVWSPERDF-MLHGWKTKQLKPTPNGT 216 (222)
T ss_pred cccc-ceEEeeccccceecccccccccCCccch-hhhhhhhhccccCCCCc
Confidence 4444 7899999999999999999999999 99 88998777777777764
No 66
>cd08678 C2_C21orf25-like C2 domain found in the Human chromosome 21 open reading frame 25 (C21orf25) protein. The members in this cd are named after the Human C21orf25 which contains a single C2 domain. Several other members contain a C1 domain downstream of the C2 domain. No other information on this protein is currently known. The C2 domain was first identified in PKC. C2 domains fold into an 8-standed beta-sandwich that can adopt 2 structural arrangements: Type I and Type II, distinguished by a circular permutation involving their N- and C-terminal beta strands. Many C2 domains are Ca2+-dependent membrane-targeting modules that bind a wide variety of substances including bind phospholipids, inositol polyphosphates, and intracellular proteins. Most C2 domain proteins are either signal transduction enzymes that contain a single C2 domain, such as protein kinase C, or membrane trafficking proteins which contain at least two C2 domains, such as synaptotagmin 1. However, there are a
Probab=34.22 E-value=2e+02 Score=21.53 Aligned_cols=59 Identities=15% Similarity=0.155 Sum_probs=42.4
Q ss_pred cccccceeEEEEEEEEecCCcceeeeEEec-CCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEec
Q 029361 41 RLKSGAERISVSIDIHNQGTSTAYDVSLTD-DSWPQDKFDVISGNISQSWERLDAGGILSHSFELDAK 107 (194)
Q Consensus 41 ~~v~g~~ditV~ytIYNvG~s~A~dV~L~D-~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk 107 (194)
.++-+ +.+++.+. .......|++-| +.+..++| + |........|..+.+..+.+.+.|+
T Consensus 43 nP~Wn-e~f~f~~~----~~~~~l~~~v~d~~~~~~~~~-l--G~~~i~l~~l~~~~~~~~~~~L~~~ 102 (126)
T cd08678 43 NPFWD-EHFLFELS----PNSKELLFEVYDNGKKSDSKF-L--GLAIVPFDELRKNPSGRQIFPLQGR 102 (126)
T ss_pred CCccC-ceEEEEeC----CCCCEEEEEEEECCCCCCCce-E--EEEEEeHHHhccCCceeEEEEecCC
Confidence 55555 66655441 234568888888 55555655 3 8889999999999999999999876
No 67
>PF06280 DUF1034: Fn3-like domain (DUF1034); InterPro: IPR010435 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This domain of unknown function is present in bacterial and plant peptidases belonging to MEROPS peptidase family S8 (subfamily S8A subtilisin, clan SB). It is C-terminal to and adjacent to the S8 peptidase domain and can be found in conjunction with the PA (Protease associated) domain (IPR003137 from INTERPRO) and additionally in Gram-positive bacteria with the surface protein anchor domain (IPR001899 from INTERPRO).; GO: 0004252 serine-type endopeptidase activity, 0005618 cell wall, 0016020 membrane; PDB: 3EIF_A 1XF1_B.
Probab=33.19 E-value=2e+02 Score=21.34 Aligned_cols=81 Identities=17% Similarity=0.147 Sum_probs=40.6
Q ss_pred eeEEEEEEEEecCCccee-eeEEe----cCCCCCCC-e----eee---cCceeeEEEEecCCCceEEEEEEEe-cceee-
Q 029361 47 ERISVSIDIHNQGTSTAY-DVSLT----DDSWPQDK-F----DVI---SGNISQSWERLDAGGILSHSFELDA-KVKGM- 111 (194)
Q Consensus 47 ~ditV~ytIYNvG~s~A~-dV~L~----D~sfp~e~-F----elv---~G~~s~~~erI~pg~nvsH~vvv~P-k~~G~- 111 (194)
+..+++++|.|.|+.+.. .++-. |..+..+. + ... ....+...=.|+||++.+-.+.+.| ...-.
T Consensus 8 ~~~~~~itl~N~~~~~~ty~~~~~~~~t~~~~~~~~~~~~~~~~~~~~~~~~~~~~vTV~ag~s~~v~vti~~p~~~~~~ 87 (112)
T PF06280_consen 8 NKFSFTITLHNYGDKPVTYTLSHVPVLTDKTDTEEGYSILVPPVPSISTVSFSPDTVTVPAGQSKTVTVTITPPSGLDAS 87 (112)
T ss_dssp SEEEEEEEEEE-SSS-EEEEEEEE-EEEEEE--ETTEEEEEEEE----EEE---EEEEE-TTEEEEEEEEEE--GGGHHT
T ss_pred CceEEEEEEEECCCCCEEEEEeeEEEEeeEeeccCCcccccccccceeeEEeCCCeEEECCCCEEEEEEEEEehhcCCcc
Confidence 458899999999998543 33332 32221111 1 111 1223445557899999999999997 41111
Q ss_pred -EeeecEEEEEEcCCCc
Q 029361 112 -FHGSPALITFRIPTKA 127 (194)
Q Consensus 112 -fn~t~A~VtY~~se~~ 127 (194)
-.|-.--|....++++
T Consensus 88 ~~~~~eG~I~~~~~~~~ 104 (112)
T PF06280_consen 88 NGPFYEGFITFKSSDGE 104 (112)
T ss_dssp T-EEEEEEEEEESSTTS
T ss_pred cCCEEEEEEEEEcCCCC
Confidence 2333355666666554
No 68
>PTZ00234 variable surface protein Vir12; Provisional
Probab=32.82 E-value=21 Score=34.18 Aligned_cols=35 Identities=34% Similarity=0.579 Sum_probs=21.1
Q ss_pred HhhhchhhhhHHHhe-eeeEEEeCcCccc-ccccccC
Q 029361 159 LAKYGSQISVISIIV-LFVYLITSPSKSA-AKGSKKK 193 (194)
Q Consensus 159 ~~~y~~~~~v~s~~~-~~v~~~~~~~~s~-~~~~~~~ 193 (194)
..+.+-=++++.+|. +|-|-+.||=||+ .|+.+||
T Consensus 364 ~rniim~~ailGtifFlfyyn~ss~lks~~~krkrkk 400 (433)
T PTZ00234 364 FRHSIVGASIIGVLVFLFFFFKSTPIRSQTNKGEKKK 400 (433)
T ss_pred HHHHHHHHHHHHHHHHhhhhhcccchhccccchhhcc
Confidence 334433444444443 7888899999999 4444443
No 69
>PF08626 TRAPPC9-Trs120: Transport protein Trs120 or TRAPPC9, TRAPP II complex subunit; InterPro: IPR013935 The trafficking protein particle complex TRAPP is a multi-protein complex needed in the early stages of the secretory pathway. To date, two kinds of TRAPP complexes have been studied, TRAPPI and TRAPP II. These complexes differ in subunit composition []. TRAPP I binds vesicles derived from the endoplasmic reticulum bringing them closer to the acceptor membrane. Trs120 is a subunit specific to the TRAPP II complex [] along with Trs65p and Trs130p(TRAPPC10). It is suggested that Trs120p is required for the stability of the Trs130p subunit, suggesting that these two proteins might interact in some way []. It is likely that there is a complex function for TRAPP II in multiple pathways [].
Probab=32.77 E-value=1.3e+02 Score=31.91 Aligned_cols=77 Identities=13% Similarity=0.124 Sum_probs=55.5
Q ss_pred cccccccceeEEEEEEEEecCCcceeeeEEecCCCCCC--CeeeecCceeeEEEEecCCCceEEEEEEEecceeeEeeec
Q 029361 39 LKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQD--KFDVISGNISQSWERLDAGGILSHSFELDAKVKGMFHGSP 116 (194)
Q Consensus 39 ~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e--~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~~G~fn~t~ 116 (194)
..-+|+| +..+|..++.|==. .||+|.|-....+ .|+ ....+.--++|.+..+..+..+|+..|....+.
T Consensus 644 ~~~~V~g-E~~~v~VtLqNPf~---fel~I~~I~L~~egv~fe----s~~~s~~l~~p~s~~~v~L~g~P~~~G~L~I~G 715 (1185)
T PF08626_consen 644 EPLWVVG-EPAEVKVTLQNPFK---FELEISSISLSTEGVPFE----SYPVSIVLLPPNSTQTVRLSGTPLETGTLKITG 715 (1185)
T ss_pred CccEEcC-CeEEEEEEEECCcc---ceEEEEEEEEEEcCCccc----cceeeeEecCCCcceEEEEEEEECccceEEEEE
Confidence 4577888 99999999999754 5677776444332 231 112233224999999999999999999999999
Q ss_pred EEEEEEc
Q 029361 117 ALITFRI 123 (194)
Q Consensus 117 A~VtY~~ 123 (194)
..|+...
T Consensus 716 ~~i~v~g 722 (1185)
T PF08626_consen 716 CIIKVFG 722 (1185)
T ss_pred EEEEEcc
Confidence 9887653
No 70
>PF15012 DUF4519: Domain of unknown function (DUF4519)
Probab=29.68 E-value=41 Score=23.73 Aligned_cols=19 Identities=47% Similarity=0.719 Sum_probs=14.7
Q ss_pred hhhhHHHheeeeEEEeCcC
Q 029361 165 QISVISIIVLFVYLITSPS 183 (194)
Q Consensus 165 ~~~v~s~~~~~v~~~~~~~ 183 (194)
..+++.++++|||+-..|.
T Consensus 38 l~~~~~~Ivv~vy~kTRP~ 56 (56)
T PF15012_consen 38 LAAVFLFIVVFVYLKTRPR 56 (56)
T ss_pred HHHHHHHHhheeEEeccCC
Confidence 3567778889999988773
No 71
>PRK06655 flgD flagellar basal body rod modification protein; Reviewed
Probab=29.28 E-value=2.5e+02 Score=24.41 Aligned_cols=30 Identities=27% Similarity=0.485 Sum_probs=18.4
Q ss_pred ecCceeeEEEEecCCCce----EEEEEEEeccee
Q 029361 81 ISGNISQSWERLDAGGIL----SHSFELDAKVKG 110 (194)
Q Consensus 81 v~G~~s~~~erI~pg~nv----sH~vvv~Pk~~G 110 (194)
-.|.....||.++...+. .++|.|..+..|
T Consensus 149 ~aG~~~f~WDG~d~~G~~lp~G~Yt~~V~A~~~g 182 (225)
T PRK06655 149 SAGVVSFTWDGTDTDGNALPDGNYTIKASASVGG 182 (225)
T ss_pred CCCceeEEECCCCCCCCcCCCeeEEEEEEEEeCC
Confidence 367778999997776552 344444444333
No 72
>PF03896 TRAP_alpha: Translocon-associated protein (TRAP), alpha subunit; InterPro: IPR005595 The alpha-subunit of the TRAP complex (TRAP alpha) is a single-spanning membrane protein of the endoplasmic reticulum (ER) which is found in proximity of nascent polypeptide chains translocating across the membrane [].; GO: 0005783 endoplasmic reticulum
Probab=29.15 E-value=4.1e+02 Score=24.11 Aligned_cols=28 Identities=14% Similarity=0.249 Sum_probs=16.2
Q ss_pred ecCCcceeeeEEecCCCCCCCeeeecCceee
Q 029361 57 NQGTSTAYDVSLTDDSWPQDKFDVISGNISQ 87 (194)
Q Consensus 57 NvG~s~A~dV~L~D~sfp~e~Felv~G~~s~ 87 (194)
+.+.+|.-|+.+. ||....+++.|+...
T Consensus 75 ~~~~sP~adt~~~---F~~~~~~l~aG~~~~ 102 (285)
T PF03896_consen 75 ELKPSPDADTTIL---FPKPTKKLPAGEPVK 102 (285)
T ss_pred cccccCCceEEEE---eccccccccCCCeEE
Confidence 4555555555554 544466777777533
No 73
>PF00345 PapD_N: Pili and flagellar-assembly chaperone, PapD N-terminal domain; InterPro: IPR016147 Most Gram-negative bacteria possess a supramolecular structure - the pili - on their surface, which mediates attachment to specific receptors. Many interactive subunits are required to assemble pili, but their assembly only takes place after translocation across the cytoplasmic membrane. Periplasmic chaperones assist pili assembly by binding to the subunits, thereby preventing premature aggregation [, ]. Pili chaperones are structurally, and possibly evolutionarily, related to the immunoglobulin superfamily [, ]: they contain two globular domains, with a topology identical to an immunoglobulin fold. This entry represents the N-terminal domain of pili assembly chaperone, and has a beta-sandwich fold consisting of seven strands in two sheets with a Greek key topology.; GO: 0007047 cellular cell wall organization, 0030288 outer membrane-bounded periplasmic space; PDB: 2CO6_B 2CO7_B 1L4I_B 3GFU_A 3F65_F 3F6L_A 3F6I_A 3GEW_B 3DSN_D 2OS7_B ....
Probab=28.85 E-value=2.5e+02 Score=20.98 Aligned_cols=72 Identities=22% Similarity=0.188 Sum_probs=46.2
Q ss_pred eeEEEEEEEEecCCcc-eeeeEEecC-C---C-CCCCeeeecCceeeEEEEecCCCceEEEEEEEec----ceeeEeeec
Q 029361 47 ERISVSIDIHNQGTST-AYDVSLTDD-S---W-PQDKFDVISGNISQSWERLDAGGILSHSFELDAK----VKGMFHGSP 116 (194)
Q Consensus 47 ~ditV~ytIYNvG~s~-A~dV~L~D~-s---f-p~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk----~~G~fn~t~ 116 (194)
.+=+..++|+|.|+.+ .+.+.+.|. . - +.+.| ..+-..-+|+||+.-+-.+...+. +...|.+.-
T Consensus 14 ~~~~~~i~v~N~~~~~~~vq~~v~~~~~~~~~~~~~~~-----~vsPp~~~L~pg~~q~vRv~~~~~~~~~~E~~yrl~~ 88 (122)
T PF00345_consen 14 SQRSASITVTNNSDQPYLVQVWVYDQDDEDEDEPTDPF-----IVSPPIFRLEPGESQTVRVYRGSKLPIDRESLYRLSF 88 (122)
T ss_dssp TSSEEEEEEEESSSSEEEEEEEEEETTSTTSSSSSSSE-----EEESSEEEEETTEEEEEEEEECSGS-SSS-EEEEEEE
T ss_pred CCCEEEEEEEcCCCCcEEEEEEEEcCCCcccccccccE-----EEeCCceEeCCCCcEEEEEEecCCCCCCceEEEEEEE
Confidence 3447799999999984 567777761 1 1 11134 134567799999999999944333 335666666
Q ss_pred EEEEEEc
Q 029361 117 ALITFRI 123 (194)
Q Consensus 117 A~VtY~~ 123 (194)
.+|-...
T Consensus 89 ~~iP~~~ 95 (122)
T PF00345_consen 89 REIPPSE 95 (122)
T ss_dssp EEEESCC
T ss_pred EEEeccc
Confidence 6666655
No 74
>PRK09918 putative fimbrial chaperone protein; Provisional
Probab=28.56 E-value=3.8e+02 Score=23.00 Aligned_cols=19 Identities=16% Similarity=0.160 Sum_probs=14.4
Q ss_pred ccceeEEEEEEEEecCCcc
Q 029361 44 SGAERISVSIDIHNQGTST 62 (194)
Q Consensus 44 ~g~~ditV~ytIYNvG~s~ 62 (194)
-.+.+-...++|.|.|+.+
T Consensus 35 ~~~~~~~~si~v~N~~~~p 53 (230)
T PRK09918 35 VEESDGEGSINVKNTDSNP 53 (230)
T ss_pred EECCCCeEEEEEEcCCCCc
Confidence 3336677888999999875
No 75
>PF08441 Integrin_alpha2: Integrin alpha; InterPro: IPR013649 This domain is found in integrin alpha and integrin alpha precursors to the C terminus of a number of IPR013517 from INTERPRO repeats and to the N terminus of the IPR013513 from INTERPRO cytoplasmic region. ; PDB: 1M1X_A 1U8C_A 1L5G_A 3IJE_A 1JV2_A 2VDN_A 2VC2_A 3NIF_A 3NIG_C 2VDM_A ....
Probab=28.46 E-value=1.2e+02 Score=27.87 Aligned_cols=45 Identities=22% Similarity=0.401 Sum_probs=28.7
Q ss_pred eEEEEeeccccc----cccc-ceeEEEEEEEEecCCcceeeeEEecCCCCCC
Q 029361 30 FIVAHKKASLKR----LKSG-AERISVSIDIHNQGTSTAYDVSLTDDSWPQD 76 (194)
Q Consensus 30 ~LlvsK~i~~~~----~v~g-~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e 76 (194)
.|-++=++.... ++.| .++++++++|-|.|+ +||+-+|.= .+|++
T Consensus 169 dL~l~~~~~~~~~~~~l~lg~~~~l~l~v~v~N~GE-~AY~a~l~v-~~P~~ 218 (457)
T PF08441_consen 169 DLQLSASFSNSESSDVLVLGSDNTLNLNVTVTNKGE-DAYEAKLTV-TYPSG 218 (457)
T ss_dssp -EEEEEEETS-CS---EECSS-EEEEEEEEEEESSS--BSSEEEEE-EEETT
T ss_pred CeEEEEEecCccceeEEEECCCCEEEEEEEEEECCC-CCCceeEEE-ECCCC
Confidence 344444444444 5554 389999999999997 899888873 35554
No 76
>PF12112 DUF3579: Protein of unknown function (DUF3579); InterPro: IPR021969 This family of proteins is functionally uncharacterised. This protein is found in bacteria. Proteins in this family are typically between 98 to 126 amino acids in length. This protein has a conserved FRP sequence motif. ; PDB: 2L9D_A.
Probab=27.48 E-value=27 Score=26.88 Aligned_cols=24 Identities=13% Similarity=0.123 Sum_probs=15.0
Q ss_pred CCCcceeeecCchhhhh---HHHHHHH
Q 029361 136 PMLPLDVLAEKPTENKL---ELAKRLL 159 (194)
Q Consensus 136 ~pg~~~I~~~~~ydrkf---ewa~~l~ 159 (194)
.+.+.-|.+--.-.|+| |||+||.
T Consensus 4 ~~~e~~I~GiT~~Gk~FRPSDWaERL~ 30 (92)
T PF12112_consen 4 NPKEIVIQGITSDGKTFRPSDWAERLC 30 (92)
T ss_dssp ---EEEEEEEETTS-B-S-TTHHHHHH
T ss_pred CccEEEEEeEcCCCCCcCCccHHHHHH
Confidence 34555666666777889 9999998
No 77
>PRK13792 lysozyme inhibitor; Provisional
Probab=26.89 E-value=1.5e+02 Score=23.84 Aligned_cols=15 Identities=20% Similarity=0.406 Sum_probs=7.8
Q ss_pred cceeEEEEEEEEecCCc
Q 029361 45 GAERISVSIDIHNQGTS 61 (194)
Q Consensus 45 g~~ditV~ytIYNvG~s 61 (194)
++++++|+|. |.++.
T Consensus 53 ~~~~~tV~y~--n~~~~ 67 (127)
T PRK13792 53 NGRKFTVQYL--NKGDN 67 (127)
T ss_pred CCCEEEEEEe--CCCCC
Confidence 3355555543 76653
No 78
>TIGR02781 VirB9 P-type conjugative transfer protein VirB9. The VirB9 protein is found in the vir locus of Agrobacterium Ti plasmids where it is involved in a type IV secretion system. VirB9 is a homolog of the F-type conjugative transfer system TraK protein (which is believed to be an outer membrane pore-forming secretin, TIGR02756) as well as the Ti system TrbG protein.
Probab=26.67 E-value=2.7e+02 Score=24.13 Aligned_cols=22 Identities=23% Similarity=0.431 Sum_probs=15.2
Q ss_pred eeEEEEecCCCceEEEEEEEecceee
Q 029361 86 SQSWERLDAGGILSHSFELDAKVKGM 111 (194)
Q Consensus 86 s~~~erI~pg~nvsH~vvv~Pk~~G~ 111 (194)
+..|+-.+.| +.+.|+|+..|.
T Consensus 69 t~~W~v~~~~----n~i~IKP~~~~~ 90 (243)
T TIGR02781 69 SKAWEVTPNG----NKLFIKPTEKDW 90 (243)
T ss_pred CcceEEEcCC----CEEEEEECCCCC
Confidence 5678887774 347777876664
No 79
>PF12034 DUF3520: Domain of unknown function (DUF3520); InterPro: IPR021908 This presumed domain is functionally uncharacterised. This domain is found in bacteria. This domain is about 180 amino acids in length. This domain is found associated with PF00092 from PFAM.
Probab=25.39 E-value=3.1e+02 Score=23.41 Aligned_cols=62 Identities=10% Similarity=0.132 Sum_probs=39.4
Q ss_pred EEecCCCceEEEEEEEecce-------------------eeEeeecEEEEEEcCCCcceeeEeecC-CCcceeeecCchh
Q 029361 90 ERLDAGGILSHSFELDAKVK-------------------GMFHGSPALITFRIPTKAALQEAYSTP-MLPLDVLAEKPTE 149 (194)
Q Consensus 90 erI~pg~nvsH~vvv~Pk~~-------------------G~fn~t~A~VtY~~se~~~~q~a~Ss~-pg~~~I~~~~~yd 149 (194)
..|-+|-+||--|.|+|... +.=.+.-..|-|+.+++.+.+ -.+-+ ........+...|
T Consensus 47 GEIGAGHsVTALYEi~p~g~~~~~~~~lkY~~~~~~~~~~~~el~tvklRYK~P~~~~s~-l~~~~v~~~~~~~~~~s~d 125 (183)
T PF12034_consen 47 GEIGAGHSVTALYEIVPAGSKGEVVDDLKYQDNEAAPASNSGELATVKLRYKDPDGDKSR-LIEQPVADASSSFAQASDD 125 (183)
T ss_pred cccCCCCEEEEEEEEEECCCCccccccccccccccCCCCCCCceEEEEEEeeCCCCCccE-EEEEeecccccccccCCcc
Confidence 45778888888888888843 345566778889988874321 11111 3344555666666
Q ss_pred hhh
Q 029361 150 NKL 152 (194)
Q Consensus 150 rkf 152 (194)
-+|
T Consensus 126 ~rf 128 (183)
T PF12034_consen 126 FRF 128 (183)
T ss_pred hhH
Confidence 677
No 80
>PRK10737 FKBP-type peptidyl-prolyl cis-trans isomerase; Provisional
Probab=25.02 E-value=94 Score=26.65 Aligned_cols=59 Identities=17% Similarity=0.327 Sum_probs=39.8
Q ss_pred eeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCc------eeeEEEEecCCCceEEEEEEEe-cceeeEe
Q 029361 47 ERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGN------ISQSWERLDAGGILSHSFELDA-KVKGMFH 113 (194)
Q Consensus 47 ~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~------~s~~~erI~pg~nvsH~vvv~P-k~~G~fn 113 (194)
.-+++.|++.+... ++.|.++..+-++++-|. +.....-..+|+..+ |+|.| ..+|.|+
T Consensus 7 ~vV~l~Y~l~~~dG------~v~dst~~~~Pl~~~~G~g~lipglE~aL~G~~~Gd~~~--v~l~peeAyGe~d 72 (196)
T PRK10737 7 LVVSLAYQVRTEDG------VLVDESPVSAPLDYLHGHGSLISGLETALEGHEVGDKFD--VAVGANDAYGQYD 72 (196)
T ss_pred CEEEEEEEEEeCCC------CEEEecCCCCCeEEEeCCCcchHHHHHHHcCCCCCCEEE--EEEChHHhcCCCC
Confidence 78999999999532 356777777777777775 345667788888776 55554 3344443
No 81
>PF10989 DUF2808: Protein of unknown function (DUF2808); InterPro: IPR021256 This family of proteins with unknown function appears to be restricted to Cyanobacteria.
Probab=24.75 E-value=76 Score=25.32 Aligned_cols=27 Identities=7% Similarity=0.204 Sum_probs=21.9
Q ss_pred EecCCCceEEEE-EEE-ecceeeEeeecE
Q 029361 91 RLDAGGILSHSF-ELD-AKVKGMFHGSPA 117 (194)
Q Consensus 91 rI~pg~nvsH~v-vv~-Pk~~G~fn~t~A 117 (194)
-|+||++++-.+ -++ |...|.|.|...
T Consensus 98 PV~pG~tv~V~l~~v~NP~~~G~Y~f~v~ 126 (146)
T PF10989_consen 98 PVPPGTTVTVVLSPVRNPRSGGTYQFNVT 126 (146)
T ss_pred CCCCCCEEEEEEEeeeCCCCCCeEEEEEE
Confidence 389999988887 554 889999999754
No 82
>smart00557 IG_FLMN Filamin-type immunoglobulin domains. These form a rod-like structure in the actin-binding cytoskeleton protein, filamin. The C-terminal repeats of filamin bind beta1-integrin (CD29).
Probab=24.66 E-value=2.7e+02 Score=20.03 Aligned_cols=60 Identities=17% Similarity=0.189 Sum_probs=41.2
Q ss_pred ccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEecceeeEeee
Q 029361 42 LKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQSWERLDAGGILSHSFELDAKVKGMFHGS 115 (194)
Q Consensus 42 ~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~~G~fn~t 115 (194)
..+| +..++.++-.+.|. ...+|.|++++-. .... ++....+=++.+.-+|+..|.|...
T Consensus 14 ~~vg-~~~~f~v~~~d~G~-~~~~v~i~~p~g~---------~~~~---~v~d~~dGty~v~y~P~~~G~~~i~ 73 (93)
T smart00557 14 GVVG-EPAEFTIDTRGAGG-GELEVEVTGPSGK---------KVPV---EVKDNGDGTYTVSYTPTEPGDYTVT 73 (93)
T ss_pred eecC-CCEEEEEEcCCCCC-CcEEEEEECCCCC---------eeEe---EEEeCCCCEEEEEEEeCCCEeEEEE
Confidence 4677 78888888888875 7888999886321 1122 2334455578888889999988654
No 83
>PF05984 Cytomega_UL20A: Cytomegalovirus UL20A protein; InterPro: IPR009245 This family consists of several Cytomegalovirus UL20A proteins. UL20A is thought to be a glycoprotein [].
Probab=24.30 E-value=1.9e+02 Score=22.37 Aligned_cols=21 Identities=19% Similarity=0.142 Sum_probs=9.6
Q ss_pred HHHHHH-HHHHhhhcccCCCce
Q 029361 10 ISVLIA-LFLISSSFASSDVPF 30 (194)
Q Consensus 10 ~~~lla-~~~v~~~~~~~~~a~ 30 (194)
+.-||| .+||++++-++..-|
T Consensus 7 iLslLAVtLtVALAAPsQKsKR 28 (100)
T PF05984_consen 7 ILSLLAVTLTVALAAPSQKSKR 28 (100)
T ss_pred HHHHHHHHHHHHhhcccccccc
Confidence 333444 445544444454443
No 84
>PRK11385 putativi pili assembly chaperone; Provisional
Probab=23.80 E-value=4.9e+02 Score=22.68 Aligned_cols=25 Identities=28% Similarity=0.264 Sum_probs=17.3
Q ss_pred ccccccceeEEEEEEEEecCCccee
Q 029361 40 KRLKSGAERISVSIDIHNQGTSTAY 64 (194)
Q Consensus 40 ~~~v~g~~ditV~ytIYNvG~s~A~ 64 (194)
.+++-.+.+-.++++|.|.|+.+..
T Consensus 33 TRvIy~~~~~~~sv~l~N~~~~p~L 57 (236)
T PRK11385 33 TRFIFPADRESISILLTNTSQESWL 57 (236)
T ss_pred eEEEEcCCCceEEEEEEeCCCCcEE
Confidence 3444444667788899999998743
No 85
>COG2373 Large extracellular alpha-helical protein [General function prediction only]
Probab=23.56 E-value=2.4e+02 Score=31.68 Aligned_cols=84 Identities=23% Similarity=0.335 Sum_probs=58.8
Q ss_pred eecccccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCceee----------------------EEEEe
Q 029361 35 KKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNISQ----------------------SWERL 92 (194)
Q Consensus 35 K~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s~----------------------~~erI 92 (194)
|-+....+-.| ..+.|..++-+..+.+ |+-|+|. =|..||+..=+... +-||+
T Consensus 1493 k~v~~~~l~~g-~~~~v~l~v~~~~~~~--~~~v~Dl--LPaG~Ev~~~~~~~~~~~n~~~~~~~~~~~~~~~e~r~DR~ 1567 (1621)
T COG2373 1493 KPVDPVELRSG-DLYLVVLTVTAQNDVP--DLLVEDL--LPAGFEVENTTLGIGSAPNEALLSWLESADAEHGEIRDDRF 1567 (1621)
T ss_pred eECccccccCC-CEEEEEEEEEecCCcc--ceEEEec--CCCceEEeccccccccccccchhhHHHHHHhhhhhhccceE
Confidence 34545566677 8888889998888877 8888873 34456665433311 11222
Q ss_pred -----cCCCceEEEEEEEecceeeEeeecEEEE--EEc
Q 029361 93 -----DAGGILSHSFELDAKVKGMFHGSPALIT--FRI 123 (194)
Q Consensus 93 -----~pg~nvsH~vvv~Pk~~G~fn~t~A~Vt--Y~~ 123 (194)
.-+...+..++||....|.|...+|.|. |++
T Consensus 1568 va~~~~~~~~~~l~Y~vRAvtpGtf~lPpa~ve~MY~p 1605 (1621)
T COG2373 1568 VAALDDEGEPVTLAYLVRAVTPGTFQLPPARVEDMYRP 1605 (1621)
T ss_pred EEEeccCCCceEEEEEEEEecCceecCChhHhhhhcCh
Confidence 2457799999999999999999999874 554
No 86
>PF12099 DUF3575: Protein of unknown function (DUF3575); InterPro: IPR021958 This family of proteins are functionally uncharacterised. This family is only found in bacteria. Proteins in this family are typically between 187 to 236 amino acids in length.
Probab=22.78 E-value=3.6e+02 Score=22.56 Aligned_cols=83 Identities=13% Similarity=0.002 Sum_probs=45.3
Q ss_pred CCceEEEEeecccccccccceeEEEEEEEEecCCcceeeeEEecCCC--CCCCeeeecCceeeEEEEecCCCceEEEEEE
Q 029361 27 DVPFIVAHKKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSW--PQDKFDVISGNISQSWERLDAGGILSHSFEL 104 (194)
Q Consensus 27 ~~a~LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sf--p~e~Felv~G~~s~~~erI~pg~nvsH~vvv 104 (194)
..+.-++=|.-...-+ .+.-++.+++.+ |--.+-..++....-.+ ....+.+...++..++=- ++.....|+=
T Consensus 22 ~~~q~~avKtN~l~~~-~~tpNlg~E~~l-~~~~Sl~l~~~yn~w~~~~~~~~~~~~~vqpE~Ryw~---~~~~~G~f~G 96 (189)
T PF12099_consen 22 ARAQKVAVKTNLLYWA-TGTPNLGVEFAL-GNRWSLDLSGSYNPWKFKSDNKKMKHWAVQPEYRYWF---CEPFNGHFIG 96 (189)
T ss_pred ccceEEEEEeHHhHHH-HhCCceEEEEEE-CCCEEEEEEEEECCccccCCCceEEEEEecceeEEEe---cccccceEEE
Confidence 3344555565545554 444889999997 22333444444443223 234577777777765444 4444544544
Q ss_pred EecceeeEee
Q 029361 105 DAKVKGMFHG 114 (194)
Q Consensus 105 ~Pk~~G~fn~ 114 (194)
.=-..|.||+
T Consensus 97 ~~~~~~~yn~ 106 (189)
T PF12099_consen 97 AHAGYGQYNI 106 (189)
T ss_pred EEEeEEEEEc
Confidence 4445566666
No 87
>PF13157 DUF3992: Protein of unknown function (DUF3992)
Probab=22.77 E-value=2.4e+02 Score=21.52 Aligned_cols=46 Identities=24% Similarity=0.268 Sum_probs=30.7
Q ss_pred eeEEEEEEEEecCCcc-eeeeEEecCCCCCCCeeeecCce-eeEEEEe
Q 029361 47 ERISVSIDIHNQGTST-AYDVSLTDDSWPQDKFDVISGNI-SQSWERL 92 (194)
Q Consensus 47 ~ditV~ytIYNvG~s~-A~dV~L~D~sfp~e~Felv~G~~-s~~~erI 92 (194)
.++.-+..+||-+.+. +..|.+..++=.-+.|.+..|+. |.+..++
T Consensus 24 ~~i~gTi~V~n~~~~~~~itV~i~~~g~~v~tftV~pG~S~S~T~~~~ 71 (92)
T PF13157_consen 24 QSISGTIYVYNDTGSGNPITVTILQNGTAVNTFTVQPGNSRSFTVRDF 71 (92)
T ss_pred EEEEEEEEEEECCCCCCCEEEEEEECCcEEeEEEECCCceEEEEeccc
Confidence 5666777888777776 99999876655556676666654 5444443
No 88
>PRK15211 fimbrial chaperone protein PefD; Provisional
Probab=22.75 E-value=5.1e+02 Score=22.47 Aligned_cols=57 Identities=14% Similarity=0.059 Sum_probs=32.0
Q ss_pred cccceeEEEEEEEEecCCcceeeeEEecCCCCCC----CeeeecCceeeEEEEecCCCceEEEEEEE
Q 029361 43 KSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQD----KFDVISGNISQSWERLDAGGILSHSFELD 105 (194)
Q Consensus 43 v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e----~Felv~G~~s~~~erI~pg~nvsH~vvv~ 105 (194)
+-.+.+-.++++|.|.|+.+..==.-.|+ ++++ .| ++ +--+-||+||+.-+-.++-.
T Consensus 32 Iy~~~~~~~si~i~N~~~~p~LvQswv~~-~~~~~~~~pF-iv----tPPlfrl~p~~~q~lRI~~~ 92 (229)
T PRK15211 32 IYDEGRKNISFEVTNQADQTYGGQVWIDN-TTQGSSTVYM-VP----APPFFKVRPKEKQIIRIMKT 92 (229)
T ss_pred EEcCCCceEEEEEEeCCCCcEEEEEEEec-CCCCCccCCE-EE----cCCeEEECCCCceEEEEEEC
Confidence 33335667888899999986432223332 3322 24 11 23466777777766666544
No 89
>PLN02171 endoglucanase
Probab=22.65 E-value=4.1e+02 Score=26.77 Aligned_cols=62 Identities=19% Similarity=0.268 Sum_probs=45.9
Q ss_pred ccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeec---CceeeEE-EEecCCCceEEEEEEE
Q 029361 44 SGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVIS---GNISQSW-ERLDAGGILSHSFELD 105 (194)
Q Consensus 44 ~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~---G~~s~~~-erI~pg~nvsH~vvv~ 105 (194)
.|..-..++.+|+|.+..++.++.|.-..+..+-++|.. |-+=-+| ..|++|++.+-.++.+
T Consensus 550 ~g~~y~qy~v~I~N~s~~~ik~i~i~~~~~~~~iW~v~~~~ngytlPs~~~sL~aG~s~tFgyI~~ 615 (629)
T PLN02171 550 KGRTYYRYSTTVTNRSAKTLKELHLGISKLYGPLWGLTKAGYGYVLPSWMPSLPAGKSLEFVYVHS 615 (629)
T ss_pred CCceEEEEEEEEEECCCCceeeeeeeeccccccchheeecCCcccCchhhcccCCCCeeEEEeecC
Confidence 344556788899999999999999986667777777764 1111234 4899999999888855
No 90
>cd04458 CSP_CDS Cold-Shock Protein (CSP) contains an S1-like cold-shock domain (CSD) that is found in eukaryotes, prokaryotes, and archaea. CSP's include the major cold-shock proteins CspA and CspB in bacteria and the eukaryotic gene regulatory factor Y-box protein. CSP expression is up-regulated by an abrupt drop in growth temperature. CSP's are also expressed under normal condition at lower level. The function of cold-shock proteins is not fully understood. They preferentially bind poly-pyrimidine region of single-stranded RNA and DNA. CSP's are thought to bind mRNA and regulate ribosomal translation, mRNA degradation, and the rate of transcription termination. The human Y-box protein, which contains a CSD, regulates transcription and translation of genes that contain the Y-box sequence in their promoters. This specific ssDNA-binding properties of CSD are required for the binding of Y-box protein to the promoter's Y-box sequence, thereby regulating transcription.
Probab=22.24 E-value=1.8e+02 Score=19.53 Aligned_cols=39 Identities=21% Similarity=0.318 Sum_probs=29.6
Q ss_pred CCceEEEEeeccccc----ccccceeEEEEEEEEecCCcceeeeE
Q 029361 27 DVPFIVAHKKASLKR----LKSGAERISVSIDIHNQGTSTAYDVS 67 (194)
Q Consensus 27 ~~a~LlvsK~i~~~~----~v~g~~ditV~ytIYNvG~s~A~dV~ 67 (194)
.+..+++|++...+. +.+| ..+++++.-.+-| --|.+|+
T Consensus 22 ~g~diffh~~~~~~~~~~~~~~G-~~V~f~~~~~~~g-~~A~~V~ 64 (65)
T cd04458 22 GGEDVFVHISALEGDGFRSLEEG-DRVEFELEEGDKG-PQAVNVR 64 (65)
T ss_pred CCcCEEEEhhHhhccCCCcCCCC-CEEEEEEEECCCC-CeEEEeE
Confidence 367899999887764 8888 8888888887544 4677665
No 91
>KOG3865 consensus Arrestin [Signal transduction mechanisms]
Probab=21.79 E-value=1.1e+02 Score=28.96 Aligned_cols=18 Identities=28% Similarity=0.446 Sum_probs=16.0
Q ss_pred EecCCCceEEEEEEEecc
Q 029361 91 RLDAGGILSHSFELDAKV 108 (194)
Q Consensus 91 rI~pg~nvsH~vvv~Pk~ 108 (194)
.++||++.+.++.|.|.-
T Consensus 261 ~v~Pgstl~Kvf~l~Pll 278 (402)
T KOG3865|consen 261 PVAPGSTLSKVFTLTPLL 278 (402)
T ss_pred ccCCCCeeeeeEEechhh
Confidence 588999999999999864
No 92
>TIGR03102 halo_cynanin halocyanin domain. Halocyanins are blue (type I) copper redox proteins found in halophilic archaea such as Natronobacterium pharaonis. This model represents a domain duplicated in some halocyanins, while appearing once in others. This domain includes the characteristic copper ligand residues. This family does not include plastocyanins, and does not include certain divergent paralogs of halocyanin.
Probab=21.43 E-value=3e+02 Score=21.49 Aligned_cols=39 Identities=15% Similarity=0.228 Sum_probs=20.7
Q ss_pred EEEEecCCcceeeeEEec-CCCCCCCeeeecCceeeEEEEecCCCceEEEEE
Q 029361 53 IDIHNQGTSTAYDVSLTD-DSWPQDKFDVISGNISQSWERLDAGGILSHSFE 103 (194)
Q Consensus 53 ytIYNvG~s~A~dV~L~D-~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vv 103 (194)
++.-|.++..+.+|+..+ ..|... . ..+++|++.+|+|.
T Consensus 52 Vtw~~~~d~~~HnV~s~~~~~f~s~-------~-----~~~~~G~t~s~Tf~ 91 (115)
T TIGR03102 52 VVWEWTGEGGGHNVVSDGDGDLDES-------E-----RVSEEGTTYEHTFE 91 (115)
T ss_pred EEEEECCCCCCEEEEECCCCCcccc-------c-----cccCCCCEEEEEec
Confidence 334566666677776543 223211 0 02457777777774
No 93
>PF10528 PA14_2: GLEYA domain; InterPro: IPR018871 This presumed domain is found in fungal adhesins and is related to the PA14 domain. ; PDB: 4A3X_A.
Probab=21.31 E-value=1.7e+02 Score=22.62 Aligned_cols=32 Identities=28% Similarity=0.391 Sum_probs=25.6
Q ss_pred cccccccceeEEEEEEEEecCCcceeeeEEecC
Q 029361 39 LKRLKSGAERISVSIDIHNQGTSTAYDVSLTDD 71 (194)
Q Consensus 39 ~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~ 71 (194)
.-++..| .=+.|++-..|.|.....+.+++|+
T Consensus 63 tv~L~aG-~yyPiRi~~~N~~g~~~~~~~i~~P 94 (113)
T PF10528_consen 63 TVYLTAG-TYYPIRIVYANGGGPGSFDFSITDP 94 (113)
T ss_dssp EEEE-TT--BEEEEEEEEE-SS-EEEEEEEEET
T ss_pred EEEEECC-cEEEEEEEEEcCCCceEEEEEEECC
Confidence 6677888 9999999999999999999999995
No 94
>PF06030 DUF916: Bacterial protein of unknown function (DUF916); InterPro: IPR010317 This family consists of putative cell surface proteins, from Firmicutes, of unknown function.
Probab=21.15 E-value=4e+02 Score=20.73 Aligned_cols=63 Identities=16% Similarity=0.187 Sum_probs=35.7
Q ss_pred cccceeEEEEEEEEecCCcc-eeeeEEecCCCCCCCeeeecCce------------------eeEEEEecCCCceEEEEE
Q 029361 43 KSGAERISVSIDIHNQGTST-AYDVSLTDDSWPQDKFDVISGNI------------------SQSWERLDAGGILSHSFE 103 (194)
Q Consensus 43 v~g~~ditV~ytIYNvG~s~-A~dV~L~D~sfp~e~Felv~G~~------------------s~~~erI~pg~nvsH~vv 103 (194)
..| +..++++.|.|.++.+ -++|++.+- ...+.-.|.=+.. +.. =.|+||++.+-++.
T Consensus 24 ~P~-q~~~l~v~i~N~s~~~~tv~v~~~~A-~Tn~nG~I~Y~~~~~~~d~sl~~~~~~~v~~~~~-Vtl~~~~sk~V~~~ 100 (121)
T PF06030_consen 24 KPG-QKQTLEVRITNNSDKEITVKVSANTA-TTNDNGVIDYSQNNPKKDKSLKYPFSDLVKIPKE-VTLPPNESKTVTFT 100 (121)
T ss_pred CCC-CEEEEEEEEEeCCCCCEEEEEEEeee-EecCCEEEEECCCCcccCcccCcchHHhccCCcE-EEECCCCEEEEEEE
Confidence 345 7888888888887763 345555431 2222221111111 112 46788888888888
Q ss_pred EE-ecc
Q 029361 104 LD-AKV 108 (194)
Q Consensus 104 v~-Pk~ 108 (194)
|. |..
T Consensus 101 i~~P~~ 106 (121)
T PF06030_consen 101 IKMPKK 106 (121)
T ss_pred EEcCCC
Confidence 87 655
No 95
>cd08546 cohesin_like Cohesin domain, interaction parter of dockerin. Bacterial cohesin domains bind to a complementary protein domain named dockerin, and this interaction is required for the formation of the cellulosome, a cellulose-degrading complex. The cellulosome consists of scaffoldin, a noncatalytic scaffolding polypeptide, that comprises repeating cohesion modules and a single carbohydrate-binding module (CBM). Specific calcium-dependent interactions between cohesins and dockerins appear to be essential for cellulosome assembly. Cohesin modules are phylogenetically distributed into three groups: type I cohesin-dockerin interactions mediate assembly of a range of dockerin-borne enzymes to the complex, while type-II interactions mediate attachment of the cellulosome complex to the bacterial cell wall. Recently discovered type-III cohesins, such as found in the anchoring scaffoldin ScaE, appears to contribute to increased stability of the elaborate cellulosome complex. While the p
Probab=20.95 E-value=3.5e+02 Score=19.95 Aligned_cols=37 Identities=24% Similarity=0.462 Sum_probs=30.8
Q ss_pred cccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecC
Q 029361 43 KSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISG 83 (194)
Q Consensus 43 v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G 83 (194)
..| +.++|...+-|...-.+.++.|. |+++.+++++-
T Consensus 12 ~~G-~~~~v~v~~~~~~~~~~~~~~l~---yD~~~l~~~~~ 48 (135)
T cd08546 12 KVG-ETVTVTVKVNNVPNVAAADFTLS---YDPSVLEFVSV 48 (135)
T ss_pred cCC-CEEEEEEEEecCCCeEEEEEEEE---ECcccEEEEec
Confidence 577 99999999999997778888876 77888888874
No 96
>cd08373 C2A_Ferlin C2 domain first repeat in Ferlin. Ferlins are involved in vesicle fusion events. Ferlins and other proteins, such as Synaptotagmins, are implicated in facilitating the fusion process when cell membranes fuse together. There are six known human Ferlins: Dysferlin (Fer1L1), Otoferlin (Fer1L2), Myoferlin (Fer1L3), Fer1L4, Fer1L5, and Fer1L6. Defects in these genes can lead to a wide range of diseases including muscular dystrophy (dysferlin), deafness (otoferlin), and infertility (fer-1, fertilization factor-1). Structurally they have 6 tandem C2 domains, designated as (C2A-C2F) and a single C-terminal transmembrane domain, though there is a new study that disputes this and claims that there are actually 7 tandem C2 domains with another C2 domain inserted between C2D and C2E. In a subset of them (Dysferlin, Myoferlin, and Fer1) there is an additional conserved domain called DysF. C2 domains fold into an 8-standed beta-sandwich that can adopt 2 structural arrangemen
Probab=20.53 E-value=3.7e+02 Score=19.98 Aligned_cols=81 Identities=12% Similarity=0.165 Sum_probs=50.5
Q ss_pred cccccceeEEEEEEEEec-CCcceeeeEEec-CCCCCCCeeeecCceeeEEEEecCCCceEEEEEEEecc-eeeEeeecE
Q 029361 41 RLKSGAERISVSIDIHNQ-GTSTAYDVSLTD-DSWPQDKFDVISGNISQSWERLDAGGILSHSFELDAKV-KGMFHGSPA 117 (194)
Q Consensus 41 ~~v~g~~ditV~ytIYNv-G~s~A~dV~L~D-~sfp~e~Felv~G~~s~~~erI~pg~nvsH~vvv~Pk~-~G~fn~t~A 117 (194)
.++-+ +.+.+. +-+. -......+++-| +.+..++| =|......+.|..+......+-|.+.+ .+.-..-.-
T Consensus 38 nP~Wn-e~f~f~--~~~~~~~~~~l~~~v~d~~~~~~d~~---iG~~~~~l~~l~~~~~~~~~~~L~~~~~~~~~~~l~l 111 (127)
T cd08373 38 NPVWN-ETFEWP--LAGSPDPDESLEIVVKDYEKVGRNRL---IGSATVSLQDLVSEGLLEVTEPLLDSNGRPTGATISL 111 (127)
T ss_pred CCccc-ceEEEE--eCCCcCCCCEEEEEEEECCCCCCCce---EEEEEEEhhHcccCCceEEEEeCcCCCCCcccEEEEE
Confidence 34444 444444 3332 345678888888 44544443 378888888999999988888887443 222234445
Q ss_pred EEEEEcCCCc
Q 029361 118 LITFRIPTKA 127 (194)
Q Consensus 118 ~VtY~~se~~ 127 (194)
++.|.+.++.
T Consensus 112 ~~~~~~~~~~ 121 (127)
T cd08373 112 EVSYQPPDGA 121 (127)
T ss_pred EEEEeCCCCc
Confidence 7778777664
No 97
>PHA02668 GM-CSF/IL-2 inhibition factor; Provisional
Probab=20.11 E-value=2.4e+02 Score=25.46 Aligned_cols=96 Identities=20% Similarity=0.245 Sum_probs=62.2
Q ss_pred HHHHHHHHHhhhcc----cCCCceEEEEeecccccccccceeEEEEEEEEecCCcceeeeEEecCCCCCCCeeeecCcee
Q 029361 11 SVLIALFLISSSFA----SSDVPFIVAHKKASLKRLKSGAERISVSIDIHNQGTSTAYDVSLTDDSWPQDKFDVISGNIS 86 (194)
Q Consensus 11 ~~lla~~~v~~~~~----~~~~a~LlvsK~i~~~~~v~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e~Felv~G~~s 86 (194)
..+||+.++|-... ..++-|-.+|++-.=.++-.- |+|-+..+.|---+.-|+|++=.-..|..++.
T Consensus 5 r~~la~~~~cg~~~s~~~~~~~~FC~aH~~evyarfrl~-MRI~~~~~~~~ps~mCmLdIE~~~~d~D~~a~-------- 75 (265)
T PHA02668 5 RVFLAVSALCGSVHSYRWIGERDFCRAHAQDVFTRLQVW-MRIDRNVTTADNASACALAIETPPSNFDADVY-------- 75 (265)
T ss_pred HHHHHHHHHhcCcccccccccchhHHhhhHHHHhheeeE-EEecccccccCCccceeEEccCCcccchhhhe--------
Confidence 34666555543211 234455666665555566667 78878888888888888888765544544221
Q ss_pred eEEEEecCCCceEEEEEEEecceeeEeeecEEEEEEc
Q 029361 87 QSWERLDAGGILSHSFELDAKVKGMFHGSPALITFRI 123 (194)
Q Consensus 87 ~~~erI~pg~nvsH~vvv~Pk~~G~fn~t~A~VtY~~ 123 (194)
-=+.|=|||-+.+-. |.|++.--+++|..
T Consensus 76 ----~~AaGINVSVali~~----gvip~syigv~fnp 104 (265)
T PHA02668 76 ----VAAAGINVSVSAINC----GFFDMRQVEVTYDT 104 (265)
T ss_pred ----eeccceEEEEEEeec----ceeeeeeEeeeecC
Confidence 124567777666555 88999999999987
No 98
>PF00963 Cohesin: Cohesin domain; InterPro: IPR002102 Cohesin domains interact with a complementary domain, termed the dockerin domain (see IPR002105 from INTERPRO). The cohesin-dockerin interaction is the crucial interaction for complex formation in the cellulosome []. The scaffoldin component of the cellulolytic bacterium Clostridium thermocellum is a non-hydrolytic protein which organises the hydrolytic enzymes in a large complex, called the cellulosome. Scaffoldin comprises a series of functional domains, amongst which is a single cellulose-binding domain and nine cohesin domains which are responsible for integrating the individual enzymatic subunits into the complex.; GO: 0030246 carbohydrate binding, 0000272 polysaccharide catabolic process; PDB: 2BM3_A 3P0D_I 3KCP_A 2B59_A 3L8Q_B 3FNK_C 3GHP_B 2CCL_A 1ANU_A 1OHZ_A ....
Probab=20.10 E-value=2.4e+02 Score=21.67 Aligned_cols=41 Identities=17% Similarity=0.360 Sum_probs=31.7
Q ss_pred cccccccceeEEEEEEEEecCC-cceeeeEEecCCCCCCCeeeecC
Q 029361 39 LKRLKSGAERISVSIDIHNQGT-STAYDVSLTDDSWPQDKFDVISG 83 (194)
Q Consensus 39 ~~~~v~g~~ditV~ytIYNvG~-s~A~dV~L~D~sfp~e~Felv~G 83 (194)
......| +.++|.+.+-|..+ -.+.+.+|. |+++.+++++.
T Consensus 7 ~~~a~~G-~tv~V~V~v~~~~~~i~~~~~~l~---yDp~~Le~~~v 48 (141)
T PF00963_consen 7 SVSAKPG-ETVTVPVNVSNVSNSIAGMQFTLS---YDPSVLEFVSV 48 (141)
T ss_dssp ECEE-TT-SEEEEEEEEESCTTTEEEEEEEEE---E-TTTEEEEEC
T ss_pred CceECCC-CEEEEEEEEEcCCCcEEEEEEEEE---eCCceEEEEee
Confidence 3445678 99999999999988 677777775 88898888875
No 99
>PF08441 Integrin_alpha2: Integrin alpha; InterPro: IPR013649 This domain is found in integrin alpha and integrin alpha precursors to the C terminus of a number of IPR013517 from INTERPRO repeats and to the N terminus of the IPR013513 from INTERPRO cytoplasmic region. ; PDB: 1M1X_A 1U8C_A 1L5G_A 3IJE_A 1JV2_A 2VDN_A 2VC2_A 3NIF_A 3NIG_C 2VDM_A ....
Probab=20.06 E-value=1.4e+02 Score=27.39 Aligned_cols=31 Identities=29% Similarity=0.534 Sum_probs=21.1
Q ss_pred ccceeEEEEEEEEecCCcceeeeEEecCCCCCC
Q 029361 44 SGAERISVSIDIHNQGTSTAYDVSLTDDSWPQD 76 (194)
Q Consensus 44 ~g~~ditV~ytIYNvG~s~A~dV~L~D~sfp~e 76 (194)
.| .++...|.|.|.|.++.-+++|.= .||..
T Consensus 341 ig-~~v~h~y~V~N~Gps~i~~~~l~i-~~P~~ 371 (457)
T PF08441_consen 341 IG-PEVTHTYEVRNNGPSTIPSASLNI-MWPYQ 371 (457)
T ss_dssp H---EEEEEEEEEE-SSS-EEEEEEEE-EEECE
T ss_pred CC-CcEEEEEEeeecCCCccccEEEEE-eeChh
Confidence 45 789999999999999877777763 46543
Done!