Query 048192
Match_columns 422
No_of_seqs 155 out of 1419
Neff 7.8
Searched_HMMs 46136
Date Fri Mar 29 07:15:07 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/048192.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/048192hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF01453 B_lectin: D-mannose b 100.0 5.5E-30 1.2E-34 215.3 5.4 101 43-143 2-114 (114)
2 cd00028 B_lectin Bulb-type man 99.9 7.3E-25 1.6E-29 185.0 14.0 103 2-117 12-116 (116)
3 smart00108 B_lectin Bulb-type 99.9 6.3E-24 1.4E-28 178.8 13.5 102 2-116 12-114 (114)
4 PF00954 S_locus_glycop: S-loc 99.8 4.1E-19 8.9E-24 148.5 8.8 76 202-283 31-110 (110)
5 PF08276 PAN_2: PAN-like domai 99.4 2.1E-13 4.5E-18 103.1 5.2 62 303-368 3-66 (66)
6 cd01098 PAN_AP_plant Plant PAN 99.4 1.2E-12 2.6E-17 103.4 8.1 81 296-387 2-84 (84)
7 cd00129 PAN_APPLE PAN/APPLE-li 99.1 9.2E-11 2E-15 91.8 5.5 67 305-386 9-80 (80)
8 smart00108 B_lectin Bulb-type 98.7 8.6E-08 1.9E-12 80.4 9.1 85 60-172 23-110 (114)
9 cd00028 B_lectin Bulb-type man 98.6 2E-07 4.3E-12 78.5 9.0 85 60-172 23-111 (116)
10 smart00473 PAN_AP divergent su 98.1 5.5E-06 1.2E-10 63.7 6.2 71 305-385 4-77 (78)
11 PF01453 B_lectin: D-mannose b 98.0 8.5E-05 1.8E-09 62.3 11.4 75 42-118 37-114 (114)
12 cd01100 APPLE_Factor_XI_like S 97.6 6E-05 1.3E-09 57.9 3.8 51 309-365 8-58 (73)
13 smart00605 CW CW domain. 91.8 0.95 2E-05 36.3 7.6 56 326-389 20-77 (94)
14 PF08277 PAN_3: PAN-like domai 90.9 0.8 1.7E-05 34.4 5.9 41 326-372 18-60 (71)
15 smart00223 APPLE APPLE domain. 89.7 0.4 8.7E-06 37.3 3.4 51 311-364 7-57 (79)
16 PF14295 PAN_4: PAN domain; PD 89.6 0.23 5E-06 34.5 1.8 37 326-362 14-51 (51)
17 PF00024 PAN_1: PAN domain Thi 89.1 0.21 4.6E-06 37.9 1.5 51 307-363 4-55 (79)
18 cd01099 PAN_AP_HGF Subfamily o 83.6 1.5 3.3E-05 34.0 3.7 34 326-363 23-58 (80)
19 PF01683 EB: EB module; Inter 83.5 1.4 3.1E-05 31.0 3.2 33 250-283 17-49 (52)
20 PF07645 EGF_CA: Calcium-bindi 68.6 3.3 7.2E-05 27.8 1.5 28 253-281 3-35 (42)
21 PF13360 PQQ_2: PQQ-like domai 65.4 92 0.002 28.3 11.3 72 42-113 12-102 (238)
22 cd00053 EGF Epidermal growth f 64.2 5.3 0.00012 24.6 1.8 25 255-280 2-30 (36)
23 cd05845 Ig2_L1-CAM_like Second 62.1 14 0.00029 29.8 4.2 35 39-73 30-64 (95)
24 PF13360 PQQ_2: PQQ-like domai 60.1 23 0.0005 32.4 6.1 48 65-112 1-61 (238)
25 KOG4649 PQQ (pyrrolo-quinoline 58.8 19 0.00042 34.6 5.2 47 40-86 166-218 (354)
26 PF07354 Sp38: Zona-pellucida- 58.5 14 0.00031 35.3 4.2 35 39-73 9-43 (271)
27 PF07974 EGF_2: EGF-like domai 57.6 9.1 0.0002 24.3 2.0 19 259-277 6-26 (32)
28 TIGR03300 assembly_YfgL outer 55.4 47 0.001 33.2 7.8 70 43-112 40-130 (377)
29 PF04478 Mid2: Mid2 like cell 54.7 2.7 5.9E-05 36.7 -1.1 20 401-420 45-64 (154)
30 PRK11138 outer membrane biogen 54.2 37 0.0008 34.3 6.9 56 58-113 120-186 (394)
31 TIGR03066 Gem_osc_para_1 Gemma 51.9 42 0.00091 27.9 5.5 52 57-109 34-104 (111)
32 smart00179 EGF_CA Calcium-bind 50.2 13 0.00027 23.7 1.9 27 253-280 3-33 (39)
33 PF01436 NHL: NHL repeat; Int 49.9 29 0.00064 20.9 3.3 21 60-80 6-26 (28)
34 PRK11138 outer membrane biogen 48.7 53 0.0011 33.2 7.0 19 95-113 342-361 (394)
35 smart00765 MANEC The MANEC dom 45.8 25 0.00055 28.2 3.3 38 326-363 36-73 (93)
36 PF12661 hEGF: Human growth fa 45.2 5.8 0.00013 19.9 -0.3 9 271-280 1-9 (13)
37 cd00054 EGF_CA Calcium-binding 44.4 18 0.00038 22.5 1.9 27 253-280 3-33 (38)
38 cd05852 Ig5_Contactin-1 Fifth 42.5 90 0.002 23.2 5.8 34 40-74 13-46 (73)
39 TIGR03300 assembly_YfgL outer 36.7 1.9E+02 0.0041 28.7 8.9 71 42-112 84-170 (377)
40 PHA00149 DNA encapsidation pro 35.7 1.2E+02 0.0026 29.7 6.7 55 2-73 235-292 (331)
41 PF10681 Rot1: Chaperone for p 35.6 1.4E+02 0.0029 27.7 6.6 84 44-129 48-159 (212)
42 PF12662 cEGF: Complement Clr- 34.7 16 0.00035 21.6 0.4 10 271-281 3-12 (24)
43 PF12690 BsuPI: Intracellular 34.1 46 0.001 25.9 3.0 15 69-83 28-42 (82)
44 cd00216 PQQ_DH Dehydrogenases 32.0 1.4E+02 0.0029 31.4 7.1 72 41-112 37-135 (488)
45 PF05935 Arylsulfotrans: Aryls 31.3 1.3E+02 0.0029 31.5 6.8 62 42-103 136-206 (477)
46 PF06006 DUF905: Bacterial pro 30.7 48 0.0011 25.0 2.4 18 100-117 35-52 (70)
47 PF13570 PQQ_3: PQQ-like domai 29.9 55 0.0012 21.2 2.5 11 76-86 1-11 (40)
48 PF09064 Tme5_EGF_like: Thromb 27.9 42 0.00091 21.6 1.5 11 270-281 18-28 (34)
49 PF05935 Arylsulfotrans: Aryls 27.8 61 0.0013 34.0 3.6 53 66-119 127-187 (477)
50 PLN00033 photosystem II stabil 27.7 1.9E+02 0.0042 29.6 7.1 51 62-112 254-306 (398)
51 COG1520 FOG: WD40-like repeat 25.0 2.9E+02 0.0063 27.5 7.9 73 42-114 130-226 (370)
52 TIGR02513 type_III_yscB type I 24.4 1.3E+02 0.0029 25.8 4.2 45 59-103 22-93 (139)
53 PF12947 EGF_3: EGF domain; I 24.3 17 0.00036 23.7 -0.8 21 259-280 6-30 (36)
54 COG3236 Uncharacterized protei 23.2 45 0.00097 29.0 1.3 20 93-112 117-137 (162)
55 PF02237 BPL_C: Biotin protein 22.4 63 0.0014 22.2 1.7 15 93-107 21-35 (48)
56 smart00564 PQQ beta-propeller 21.8 1E+02 0.0022 18.6 2.5 19 64-82 13-32 (33)
57 PF00008 EGF: EGF-like domain 21.7 33 0.00072 21.5 0.2 20 260-280 5-29 (32)
58 smart00181 EGF Epidermal growt 20.9 72 0.0016 19.6 1.7 21 259-281 6-30 (35)
59 PF14870 PSII_BNR: Photosynthe 20.0 2.6E+02 0.0056 27.5 6.1 25 87-112 187-211 (302)
No 1
>PF01453 B_lectin: D-mannose binding lectin; InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]: Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity. Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=99.96 E-value=5.5e-30 Score=215.25 Aligned_cols=101 Identities=53% Similarity=0.808 Sum_probs=75.6
Q ss_pred CcEEEEcCCCCCCCC---CcEEEEecCccEEEEcCCCCEEEee-CCCCCc--eeEEEEecCCCeeEEcCCCcEEEeccCC
Q 048192 43 PQVVWSANRNNLVRI---NATLELTSDGNLVLQDADGAIAWST-NTSGKS--VVGLNLTDMGNLVLFDKNNAAVWQSFDH 116 (422)
Q Consensus 43 ~~vVW~ANr~~pv~~---~~~L~l~~~G~LvL~~~~~~~vWst-~~~~~~--~~~~~LldsGNLVL~~~~~~~lWQSFd~ 116 (422)
++|||+|||+.|+.. ..+|.|+.||+|+|.+..++.+|++ ++.+.+ ...|.|+|+|||||+|..+.+|||||||
T Consensus 2 ~tvvW~an~~~p~~~~s~~~~L~l~~dGnLvl~~~~~~~iWss~~t~~~~~~~~~~~L~~~GNlvl~d~~~~~lW~Sf~~ 81 (114)
T PF01453_consen 2 RTVVWVANRNSPLTSSSGNYTLILQSDGNLVLYDSNGSVIWSSNNTSGRGNSGCYLVLQDDGNLVLYDSSGNVLWQSFDY 81 (114)
T ss_dssp --------TTEEEEECETTEEEEEETTSEEEEEETTTEEEEE--S-TTSS-SSEEEEEETTSEEEEEETTSEEEEESTTS
T ss_pred cccccccccccccccccccccceECCCCeEEEEcCCCCEEEEecccCCccccCeEEEEeCCCCEEEEeecceEEEeecCC
Confidence 789999999999943 3899999999999999988899999 666554 7889999999999999999999999999
Q ss_pred CCCccCCCceecCC------CeeeeecCCCCCC
Q 048192 117 PTDSLVPGQKLLEG------KKLTASVSTTNWT 143 (422)
Q Consensus 117 PTDTlLpgq~l~~~------~~L~S~~s~~d~s 143 (422)
||||+||+|+|+.+ ..|+||++.+|||
T Consensus 82 ptdt~L~~q~l~~~~~~~~~~~~~sw~s~~dps 114 (114)
T PF01453_consen 82 PTDTLLPGQKLGDGNVTGKNDSLTSWSSNTDPS 114 (114)
T ss_dssp SS-EEEEEET--TSEEEEESTSSEEEESS----
T ss_pred CccEEEeccCcccCCCccccceEEeECCCCCCC
Confidence 99999999999873 3499999999986
No 2
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=99.92 E-value=7.3e-25 Score=185.04 Aligned_cols=103 Identities=41% Similarity=0.633 Sum_probs=91.1
Q ss_pred CCCCeEEEeeecCCCCCc-eEEEEEeeccccccccccccCCCCcEEEEcCCCCCCCCCcEEEEecCccEEEEcCCCCEEE
Q 048192 2 TFGPTYACGFFCNGTCDS-YLFAVFIVHAYDASLIEYQHTEFPQVVWSANRNNLVRINATLELTSDGNLVLQDADGAIAW 80 (422)
Q Consensus 2 ~~~~~F~~GF~~~~~~~~-~~l~Iw~~~~~~~~~~~~~~~~~~~vVW~ANr~~pv~~~~~L~l~~~G~LvL~~~~~~~vW 80 (422)
|.++.|++|||.+.. .. ++.+|||.. .+ .++||.||++.|....++|.|++||+|+|.|.++.++|
T Consensus 12 s~~~~f~~G~~~~~~-q~~dgnlv~~~~-----------~~-~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW 78 (116)
T cd00028 12 SSGSLFELGFFKLIM-QSRDYNLILYKG-----------SS-RTVVWVANRDNPSGSSCTLTLQSDGNLVIYDGSGTVVW 78 (116)
T ss_pred eCCCcEEEecccCCC-CCCeEEEEEEeC-----------CC-CeEEEECCCCCCCCCCEEEEEecCCCeEEEcCCCcEEE
Confidence 678999999999865 44 899999864 23 68999999999966668999999999999999999999
Q ss_pred eeCCCC-CceeEEEEecCCCeeEEcCCCcEEEeccCCC
Q 048192 81 STNTSG-KSVVGLNLTDMGNLVLFDKNNAAVWQSFDHP 117 (422)
Q Consensus 81 st~~~~-~~~~~~~LldsGNLVL~~~~~~~lWQSFd~P 117 (422)
++++.+ .....|+|+|+|||||++.++.+||||||||
T Consensus 79 ~S~~~~~~~~~~~~L~ddGnlvl~~~~~~~~W~Sf~~P 116 (116)
T cd00028 79 SSNTTRVNGNYVLVLLDDGNLVLYDSDGNFLWQSFDYP 116 (116)
T ss_pred EecccCCCCceEEEEeCCCCEEEECCCCCEEEcCCCCC
Confidence 999876 5677889999999999999999999999999
No 3
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=99.91 E-value=6.3e-24 Score=178.77 Aligned_cols=102 Identities=40% Similarity=0.619 Sum_probs=90.7
Q ss_pred CCCCeEEEeeecCCCCCceEEEEEeeccccccccccccCCCCcEEEEcCCCCCCCCCcEEEEecCccEEEEcCCCCEEEe
Q 048192 2 TFGPTYACGFFCNGTCDSYLFAVFIVHAYDASLIEYQHTEFPQVVWSANRNNLVRINATLELTSDGNLVLQDADGAIAWS 81 (422)
Q Consensus 2 ~~~~~F~~GF~~~~~~~~~~l~Iw~~~~~~~~~~~~~~~~~~~vVW~ANr~~pv~~~~~L~l~~~G~LvL~~~~~~~vWs 81 (422)
|.++.|++|||.+.. ..++.+|||.. .+ .++||+|||+.|+..+++|.|++||+|+|.+.++.++|+
T Consensus 12 s~~~~f~~G~~~~~~-q~dgnlV~~~~-----------~~-~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~ 78 (114)
T smart00108 12 SGNSLFELGFFTLIM-QNDYNLILYKS-----------SS-RTVVWVANRDNPVSDSCTLTLQSDGNLVLYDGDGRVVWS 78 (114)
T ss_pred cCCCcEeeeccccCC-CCCEEEEEEEC-----------CC-CcEEEECCCCCCCCCCEEEEEeCCCCEEEEeCCCCEEEE
Confidence 678999999998865 56889999864 23 689999999999887789999999999999998999999
Q ss_pred eCCC-CCceeEEEEecCCCeeEEcCCCcEEEeccCC
Q 048192 82 TNTS-GKSVVGLNLTDMGNLVLFDKNNAAVWQSFDH 116 (422)
Q Consensus 82 t~~~-~~~~~~~~LldsGNLVL~~~~~~~lWQSFd~ 116 (422)
+++. +.+...|+|+|+|||||++..+.+|||||||
T Consensus 79 S~t~~~~~~~~~~L~ddGnlvl~~~~~~~~W~Sf~~ 114 (114)
T smart00108 79 SNTTGANGNYVLVLLDDGNLVIYDSDGNFLWQSFDY 114 (114)
T ss_pred ecccCCCCceEEEEeCCCCEEEECCCCCEEeCCCCC
Confidence 9986 5567789999999999999999999999997
No 4
>PF00954 S_locus_glycop: S-locus glycoprotein family; InterPro: IPR000858 In Brassicaceae, self-incompatible plants have a self/non-self recognition system, which involves the inability of flowering plants to achieve self-fertilisation. This is sporophytically controlled by multiple alleles at a single locus (S). There are a total of 50 different S alleles in Brassica oleracea. S-locus glycoproteins, as well as S-receptor kinases, are in linkage with the S-alleles []. Most of the proteins within this family contain apple-like domain (IPR003609 from INTERPRO), which is predicted to possess protein- and/or carbohydrate-binding functions.; GO: 0048544 recognition of pollen
Probab=99.78 E-value=4.1e-19 Score=148.46 Aligned_cols=76 Identities=26% Similarity=0.508 Sum_probs=64.8
Q ss_pred CCCCceEEEcCCCCCCCeEEEEEccCCcEEEEEEe-CCCCeeEEeecccccCCCCCCCcCCCCCceeCC---CccCCCCC
Q 048192 202 PREPDGAVPVPPASSSPGQYMRLWPDGHLRVYEWQ-ASIGWTQVADLLEGYHGECGYPMVCGKYGICSQ---GQCSCPAT 277 (422)
Q Consensus 202 ~~~~~~~~s~~~~~~~~~~rl~Ld~dG~lr~y~w~-~~~~W~~~~~~~~~p~d~C~v~g~CG~~giC~~---~~C~C~~g 277 (422)
..+.++.|...+... ++|++||++|++++|.|. ..++|.++ |.+|.|+||+|+.||+||+|+. +.|+||+|
T Consensus 31 ~~e~~~t~~~~~~s~--~~r~~ld~~G~l~~~~w~~~~~~W~~~---~~~p~d~Cd~y~~CG~~g~C~~~~~~~C~Cl~G 105 (110)
T PF00954_consen 31 NEEVYYTYSLSNSSV--LSRLVLDSDGQLQRYIWNESTQSWSVF---WSAPKDQCDVYGFCGPNGICNSNNSPKCSCLPG 105 (110)
T ss_pred CCeEEEEEecCCCce--EEEEEEeeeeEEEEEEEecCCCcEEEE---EEecccCCCCccccCCccEeCCCCCCceECCCC
Confidence 445566666655554 899999999999999999 88999996 6789999999999999999983 78999998
Q ss_pred CceecC
Q 048192 278 YFKLLN 283 (422)
Q Consensus 278 ~F~~~~ 283 (422)
|+|++
T Consensus 106 -F~P~n 110 (110)
T PF00954_consen 106 -FEPKN 110 (110)
T ss_pred -cCCCc
Confidence 99974
No 5
>PF08276 PAN_2: PAN-like domain; InterPro: IPR013227 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs
Probab=99.42 E-value=2.1e-13 Score=103.14 Aligned_cols=62 Identities=27% Similarity=0.620 Sum_probs=48.5
Q ss_pred CCcceEEeCCcccCCCCCCCCCc-CCCCHHHHHHHhhccCCccceecccccCCCCCcee-ecceeece
Q 048192 303 QDHSFVELNDVAYFAFSSPSSDL-TNTDPETCKQACLKNCSCKAALFLYGLNLSPGDCY-LPSELFSM 368 (422)
Q Consensus 303 ~~~~f~~l~~~~~~~~~~~~~~~-~~~s~~~C~~~CL~nCsC~a~~y~~~~~~~~g~C~-~~~~l~~~ 368 (422)
+.++|++|+++++|+. ....+ .+.++++|++.||+||||+| |+|.+..+++.|+ |.++|+|+
T Consensus 3 ~~d~F~~l~~~~~p~~--~~~~~~~~~s~~~C~~~Cl~nCsC~A--yay~~~~~~~~C~lW~~~L~d~ 66 (66)
T PF08276_consen 3 SGDGFLKLPNMKLPDF--DNAIVDSSVSLEECEKACLSNCSCTA--YAYSNLSGGGGCLLWYGDLVDL 66 (66)
T ss_pred CCCEEEEECCeeCCCC--cceeeecCCCHHHHHhhcCCCCCEee--EEeeccCCCCEEEEEcCEeecC
Confidence 3578999999999876 33332 56899999999999999999 6665432356799 77899875
No 6
>cd01098 PAN_AP_plant Plant PAN/APPLE-like domain; present in plant S-receptor protein kinases and secreted glycoproteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions. S-receptor protein kinases and S-locus glycoproteins are involved in sporophytic self-incompatibility response in Brassica, one of probably many molecular mechanisms, by which hermaphrodite flowering plants avoid self-fertilization.
Probab=99.39 E-value=1.2e-12 Score=103.42 Aligned_cols=81 Identities=31% Similarity=0.562 Sum_probs=57.4
Q ss_pred cCCCCCCC-CcceEEeCCcccCCCCCCCCCcCCCCHHHHHHHhhccCCccceecccccCCCCCceeec-ceeeceeeccc
Q 048192 296 PLSCEASQ-DHSFVELNDVAYFAFSSPSSDLTNTDPETCKQACLKNCSCKAALFLYGLNLSPGDCYLP-SELFSMMNNEK 373 (422)
Q Consensus 296 ~l~C~~~~-~~~f~~l~~~~~~~~~~~~~~~~~~s~~~C~~~CL~nCsC~a~~y~~~~~~~~g~C~~~-~~l~~~~~~~~ 373 (422)
++.|.... .+.|++++++++++. .+.. ...++++|++.||+||+|.|++|.. +++.|+++ ..+.+......
T Consensus 2 ~~~C~~~~~~~~f~~~~~~~~~~~--~~~~-~~~s~~~C~~~Cl~nCsC~a~~~~~----~~~~C~~~~~~~~~~~~~~~ 74 (84)
T cd01098 2 PLNCGGDGSTDGFLKLPDVKLPDN--ASAI-TAISLEECREACLSNCSCTAYAYNN----GSGGCLLWNGLLNNLRSLSS 74 (84)
T ss_pred CcccCCCCCCCEEEEeCCeeCCCc--hhhh-ccCCHHHHHHHHhcCCCcceeeecC----CCCeEEEEeceecceEeecC
Confidence 45675322 468999999999875 3333 6789999999999999999955543 25689954 56666554321
Q ss_pred cCCCcCceEEEEEc
Q 048192 374 ERTHYNSTAYIKVQ 387 (422)
Q Consensus 374 ~~~~~~~~~yikv~ 387 (422)
. +.++||||+
T Consensus 75 ~----~~~~yiKv~ 84 (84)
T cd01098 75 G----GGTLYLRLA 84 (84)
T ss_pred C----CcEEEEEeC
Confidence 1 568999985
No 7
>cd00129 PAN_APPLE PAN/APPLE-like domain; present in N-terminal (N) domains of plasminogen/ hepatocyte growth factor proteins, plasma prekallikrein/coagulation factor XI and microneme antigen proteins, plant receptor-like protein kinases, and various nematode and leech anti-platelet proteins. Common structural features include two disulfide bonds that link the alpha-helix to the central region of the protein. PAN domains have significant functional versatility, fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=99.12 E-value=9.2e-11 Score=91.83 Aligned_cols=67 Identities=16% Similarity=0.280 Sum_probs=52.2
Q ss_pred cceEEeCCcccCCCCCCCCCcCCCCHHHHHHHhhc---cCCccceecccccCCCCCcee-eccee-eceeeccccCCCcC
Q 048192 305 HSFVELNDVAYFAFSSPSSDLTNTDPETCKQACLK---NCSCKAALFLYGLNLSPGDCY-LPSEL-FSMMNNEKERTHYN 379 (422)
Q Consensus 305 ~~f~~l~~~~~~~~~~~~~~~~~~s~~~C~~~CL~---nCsC~a~~y~~~~~~~~g~C~-~~~~l-~~~~~~~~~~~~~~ 379 (422)
..|+++.+++.|+. ...++++|++.|++ ||||.| |+|.+. .+.|+ |.+++ +++++...+ +
T Consensus 9 g~fl~~~~~klpd~-------~~~s~~eC~~~Cl~~~~nCsC~A--ya~~~~--~~gC~~W~~~l~~d~~~~~~~----g 73 (80)
T cd00129 9 GTTLIKIALKIKTT-------KANTADECANRCEKNGLPFSCKA--FVFAKA--RKQCLWFPFNSMSGVRKEFSH----G 73 (80)
T ss_pred CeEEEeecccCCcc-------cccCHHHHHHHHhcCCCCCCcee--eeccCC--CCCeEEecCcchhhHHhccCC----C
Confidence 46888888988764 33689999999999 999999 777543 23598 77889 888765332 7
Q ss_pred ceEEEEE
Q 048192 380 STAYIKV 386 (422)
Q Consensus 380 ~~~yikv 386 (422)
.++|+|.
T Consensus 74 ~~Ly~r~ 80 (80)
T cd00129 74 FDLYENK 80 (80)
T ss_pred ceeEeEC
Confidence 8999983
No 8
>smart00108 B_lectin Bulb-type mannose-specific lectin.
Probab=98.70 E-value=8.6e-08 Score=80.44 Aligned_cols=85 Identities=27% Similarity=0.519 Sum_probs=62.8
Q ss_pred EEEEecCccEEEEcCC-CCEEEeeCCCCC--ceeEEEEecCCCeeEEcCCCcEEEeccCCCCCccCCCceecCCCeeeee
Q 048192 60 TLELTSDGNLVLQDAD-GAIAWSTNTSGK--SVVGLNLTDMGNLVLFDKNNAAVWQSFDHPTDSLVPGQKLLEGKKLTAS 136 (422)
Q Consensus 60 ~L~l~~~G~LvL~~~~-~~~vWst~~~~~--~~~~~~LldsGNLVL~~~~~~~lWQSFd~PTDTlLpgq~l~~~~~L~S~ 136 (422)
++.+..||+||+.+.. +.++|++++... ....+.|+++|||||++.++.++|+|= |- +
T Consensus 23 ~~~~q~dgnlV~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S~-----t~-~------------- 83 (114)
T smart00108 23 TLIMQNDYNLILYKSSSRTVVWVANRDNPVSDSCTLTLQSDGNLVLYDGDGRVVWSSN-----TT-G------------- 83 (114)
T ss_pred ccCCCCCEEEEEEECCCCcEEEECCCCCCCCCCEEEEEeCCCCEEEEeCCCCEEEEec-----cc-C-------------
Confidence 4556689999999754 579999998533 236789999999999999899999971 11 1
Q ss_pred cCCCCCCCCCceEEEecCCCceEEEeccCCcceEEE
Q 048192 137 VSTTNWTDGGLFSLSVSNKGLFAFIESNNTSIRYYE 172 (422)
Q Consensus 137 ~s~~d~s~~G~y~l~~~~~g~~~~~~~~~~~~~Yw~ 172 (422)
.. |.|.+.|+++|.+.+... ..++ .|.
T Consensus 84 ------~~-~~~~~~L~ddGnlvl~~~-~~~~-~W~ 110 (114)
T smart00108 84 ------AN-GNYVLVLLDDGNLVIYDS-DGNF-LWQ 110 (114)
T ss_pred ------CC-CceEEEEeCCCCEEEECC-CCCE-EeC
Confidence 12 678999999998766533 2345 775
No 9
>cd00028 B_lectin Bulb-type mannose-specific lectin. The domain contains a three-fold internal repeat (beta-prism architecture). The consensus sequence motif QXDXNXVXY is involved in alpha-D-mannose recognition. Lectins are carbohydrate-binding proteins which specifically recognize diverse carbohydrates and mediate a wide variety of biological processes, such as cell-cell and host-pathogen interactions, serum glycoprotein turnover, and innate immune responses.
Probab=98.61 E-value=2e-07 Score=78.48 Aligned_cols=85 Identities=27% Similarity=0.514 Sum_probs=63.2
Q ss_pred EEEEec-CccEEEEcCC-CCEEEeeCCCC--CceeEEEEecCCCeeEEcCCCcEEEeccCCCCCccCCCceecCCCeeee
Q 048192 60 TLELTS-DGNLVLQDAD-GAIAWSTNTSG--KSVVGLNLTDMGNLVLFDKNNAAVWQSFDHPTDSLVPGQKLLEGKKLTA 135 (422)
Q Consensus 60 ~L~l~~-~G~LvL~~~~-~~~vWst~~~~--~~~~~~~LldsGNLVL~~~~~~~lWQSFd~PTDTlLpgq~l~~~~~L~S 135 (422)
++.... +|+||+.+.. +.++|++++.. .....+.|+++|||||++.++.++|+|=-..
T Consensus 23 ~~~~q~~dgnlv~~~~~~~~~vW~snt~~~~~~~~~l~l~~dGnLvl~~~~g~~vW~S~~~~------------------ 84 (116)
T cd00028 23 KLIMQSRDYNLILYKGSSRTVVWVANRDNPSGSSCTLTLQSDGNLVIYDGSGTVVWSSNTTR------------------ 84 (116)
T ss_pred cCCCCCCeEEEEEEeCCCCeEEEECCCCCCCCCCEEEEEecCCCeEEEcCCCcEEEEecccC------------------
Confidence 344565 9999999754 57899999854 3457789999999999999999999954210
Q ss_pred ecCCCCCCCCCceEEEecCCCceEEEeccCCcceEEE
Q 048192 136 SVSTTNWTDGGLFSLSVSNKGLFAFIESNNTSIRYYE 172 (422)
Q Consensus 136 ~~s~~d~s~~G~y~l~~~~~g~~~~~~~~~~~~~Yw~ 172 (422)
.. +.+.+.|+++|.+.+.... ..+ .|.
T Consensus 85 -------~~-~~~~~~L~ddGnlvl~~~~-~~~-~W~ 111 (116)
T cd00028 85 -------VN-GNYVLVLLDDGNLVLYDSD-GNF-LWQ 111 (116)
T ss_pred -------CC-CceEEEEeCCCCEEEECCC-CCE-EEc
Confidence 13 6789999999987765433 345 776
No 10
>smart00473 PAN_AP divergent subfamily of APPLE domains. Apple-like domains present in Plasminogen, C. elegans hypothetical ORFs and the extracellular portion of plant receptor-like protein kinases. Predicted to possess protein- and/or carbohydrate-binding functions.
Probab=98.14 E-value=5.5e-06 Score=63.71 Aligned_cols=71 Identities=25% Similarity=0.388 Sum_probs=47.3
Q ss_pred cceEEeCCcccCCCCCCCCCcCCCCHHHHHHHhhc-cCCccceecccccCCCCCceeecc--eeeceeeccccCCCcCce
Q 048192 305 HSFVELNDVAYFAFSSPSSDLTNTDPETCKQACLK-NCSCKAALFLYGLNLSPGDCYLPS--ELFSMMNNEKERTHYNST 381 (422)
Q Consensus 305 ~~f~~l~~~~~~~~~~~~~~~~~~s~~~C~~~CL~-nCsC~a~~y~~~~~~~~g~C~~~~--~l~~~~~~~~~~~~~~~~ 381 (422)
..|++++++.++.. ........++++|++.|++ +|+|.|+.|.+ +++.|+++. .+.+.... ...+.+
T Consensus 4 ~~f~~~~~~~l~~~--~~~~~~~~s~~~C~~~C~~~~~~C~s~~y~~----~~~~C~l~~~~~~~~~~~~----~~~~~~ 73 (78)
T smart00473 4 DCFVRLPNTKLPGF--SRIVISVASLEECASKCLNSNCSCRSFTYNN----GTKGCLLWSESSLGDARLF----PSGGVD 73 (78)
T ss_pred ceeEEecCccCCCC--cceeEcCCCHHHHHHHhCCCCCceEEEEEcC----CCCEEEEeeCCccccceec----ccCCce
Confidence 46889999988753 2222356799999999999 99999965543 256799544 44444422 122456
Q ss_pred EEEE
Q 048192 382 AYIK 385 (422)
Q Consensus 382 ~yik 385 (422)
+|.|
T Consensus 74 ~y~~ 77 (78)
T smart00473 74 LYEK 77 (78)
T ss_pred eEEe
Confidence 7766
No 11
>PF01453 B_lectin: D-mannose binding lectin; InterPro: IPR001480 A bulb lectin super-family (Amaryllidaceae, Orchidaceae and Aliaceae) contains a ~115-residue-long domain whose overall three dimensional fold is very similar to that of [, ]: Dictyostelium discoideum comitin, an actin binding protein Curculigo latifolia curculin, a sweet tasting and taste-modifying protein This domain generally binds mannose, but in at least one protein, curculin, it is apparently devoid of mannose-binding activity. Each bulb-type lectin domain consists of three sequential beta-sheet subdomains (I, II, III) that are inter-related by pseudo three-fold symmetry. The three subdomains are flat four-stranded, antiparrallel beta-sheets. Together they form a 12-stranded beta-barrel in which the barrel axis coincides with the pseudo 3-fold axis.; GO: 0005529 sugar binding; PDB: 3M7H_A 3M7J_B 3MEZ_D 1DLP_A 1BWU_D 1KJ1_A 1B2P_A 1XD6_A 2DPF_C 2D04_B ....
Probab=98.02 E-value=8.5e-05 Score=62.34 Aligned_cols=75 Identities=28% Similarity=0.418 Sum_probs=51.8
Q ss_pred CCcEEEEc-CCCCCCCCCcEEEEecCccEEEEcCCCCEEEeeCCCCCceeEEEEec--CCCeeEEcCCCcEEEeccCCCC
Q 048192 42 FPQVVWSA-NRNNLVRINATLELTSDGNLVLQDADGAIAWSTNTSGKSVVGLNLTD--MGNLVLFDKNNAAVWQSFDHPT 118 (422)
Q Consensus 42 ~~~vVW~A-Nr~~pv~~~~~L~l~~~G~LvL~~~~~~~vWst~~~~~~~~~~~Lld--sGNLVL~~~~~~~lWQSFd~PT 118 (422)
..++||.. +........+.+.|.++|||||.|..+.++|++.. ..+.+.+.+++ .||++ ......+.|.|=+.|.
T Consensus 37 ~~~~iWss~~t~~~~~~~~~~~L~~~GNlvl~d~~~~~lW~Sf~-~ptdt~L~~q~l~~~~~~-~~~~~~~sw~s~~dps 114 (114)
T PF01453_consen 37 NGSVIWSSNNTSGRGNSGCYLVLQDDGNLVLYDSSGNVLWQSFD-YPTDTLLPGQKLGDGNVT-GKNDSLTSWSSNTDPS 114 (114)
T ss_dssp TTEEEEE--S-TTSS-SSEEEEEETTSEEEEEETTSEEEEESTT-SSS-EEEEEET--TSEEE-EESTSSEEEESS----
T ss_pred CCCEEEEecccCCccccCeEEEEeCCCCEEEEeecceEEEeecC-CCccEEEeccCcccCCCc-cccceEEeECCCCCCC
Confidence 35679999 43433334588999999999999988999999943 33455566777 88998 6656679999877663
No 12
>cd01100 APPLE_Factor_XI_like Subfamily of PAN/APPLE-like domains; present in plasma prekallikrein/coagulation factor XI, microneme antigen proteins, and a few prokaryotic proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=97.61 E-value=6e-05 Score=57.89 Aligned_cols=51 Identities=22% Similarity=0.378 Sum_probs=36.3
Q ss_pred EeCCcccCCCCCCCCCcCCCCHHHHHHHhhccCCccceecccccCCCCCceeeccee
Q 048192 309 ELNDVAYFAFSSPSSDLTNTDPETCKQACLKNCSCKAALFLYGLNLSPGDCYLPSEL 365 (422)
Q Consensus 309 ~l~~~~~~~~~~~~~~~~~~s~~~C~~~CL~nCsC~a~~y~~~~~~~~g~C~~~~~l 365 (422)
.++++++++. +.......+.++|++.|+.+|+|.|++|.. +.+.|+++...
T Consensus 8 ~~~~~~~~g~--d~~~~~~~s~~~Cq~~C~~~~~C~afT~~~----~~~~C~lk~~~ 58 (73)
T cd01100 8 QGSNVDFRGG--DLSTVFASSAEQCQAACTADPGCLAFTYNT----KSKKCFLKSSE 58 (73)
T ss_pred ccCCCccccC--CcceeecCCHHHHHHHcCCCCCceEEEEEC----CCCeEEcccCC
Confidence 3457777764 333334668999999999999999966643 35789986544
No 13
>smart00605 CW CW domain.
Probab=91.79 E-value=0.95 Score=36.26 Aligned_cols=56 Identities=18% Similarity=0.309 Sum_probs=40.0
Q ss_pred CCCCHHHHHHHhhccCCccceecccccCCCCCceeec--ceeeceeeccccCCCcCceEEEEEcCC
Q 048192 326 TNTDPETCKQACLKNCSCKAALFLYGLNLSPGDCYLP--SELFSMMNNEKERTHYNSTAYIKVQNF 389 (422)
Q Consensus 326 ~~~s~~~C~~~CL~nCsC~a~~y~~~~~~~~g~C~~~--~~l~~~~~~~~~~~~~~~~~yikv~~s 389 (422)
...+.++|.+.|..+..|..|.+.. ...|.|+ ++++.+++...+. +..+=||+..+
T Consensus 20 ~~~sw~~Ci~~C~~~~~Cvlay~~~-----~~~C~~f~~~~~~~v~~~~~~~---~~~VAfK~~~~ 77 (94)
T smart00605 20 ATLSWDECIQKCYEDSNCVLAYGNS-----SETCYLFSYGTVLTVKKLSSSS---GKKVAFKVSTD 77 (94)
T ss_pred cCCCHHHHHHHHhCCCceEEEecCC-----CCceEEEEcCCeEEEEEccCCC---CcEEEEEEeCC
Confidence 4668899999999999999865431 2569874 5666777664432 66788887644
No 14
>PF08277 PAN_3: PAN-like domain; InterPro: IPR006583 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs The PAN-3 or CW is a domain associated with a number of Caenorhabditis elegans hypothetical proteins.
Probab=90.86 E-value=0.8 Score=34.36 Aligned_cols=41 Identities=32% Similarity=0.650 Sum_probs=31.0
Q ss_pred CCCCHHHHHHHhhccCCccceecccccCCCCCceeec--ceeeceeecc
Q 048192 326 TNTDPETCKQACLKNCSCKAALFLYGLNLSPGDCYLP--SELFSMMNNE 372 (422)
Q Consensus 326 ~~~s~~~C~~~CL~nCsC~a~~y~~~~~~~~g~C~~~--~~l~~~~~~~ 372 (422)
...+.++|-+.|+.+=.|.+|.+. ++.|.++ +++..+.+..
T Consensus 18 ~~~sw~~Cv~~C~~~~~C~la~~~------~~~C~~y~~~~i~~v~~~~ 60 (71)
T PF08277_consen 18 TNTSWDDCVQKCYNDENCVLAYFD------SGKCYLYNYGSISTVQKTD 60 (71)
T ss_pred cCCCHHHHhHHhCCCCEEEEEEeC------CCCEEEEEcCCEEEEEEee
Confidence 567889999999999999998875 2469974 5555555543
No 15
>smart00223 APPLE APPLE domain. Four-fold repeat in plasma kallikrein and coagulation factor XI. Factor XI apple 3 mediates binding to platelets. Factor XI apple 1 binds high-molecular-mass kininogen. Apple 4 in factor XI mediates dimer formation and binds to factor XIIa. Mutations in apple 4 cause factor XI deficiency, an inherited bleeding disorder.
Probab=89.66 E-value=0.4 Score=37.34 Aligned_cols=51 Identities=14% Similarity=0.295 Sum_probs=35.0
Q ss_pred CCcccCCCCCCCCCcCCCCHHHHHHHhhccCCccceecccccCCCCCceeecce
Q 048192 311 NDVAYFAFSSPSSDLTNTDPETCKQACLKNCSCKAALFLYGLNLSPGDCYLPSE 364 (422)
Q Consensus 311 ~~~~~~~~~~~~~~~~~~s~~~C~~~CL~nCsC~a~~y~~~~~~~~g~C~~~~~ 364 (422)
++++|++. +...+...+.++|++.|..+=.|.+++|..... ....|+++..
T Consensus 7 ~~~df~G~--Dl~~~~~~~~~~Cq~~Ct~~~~C~~FTf~~~~~-~~~~C~LK~s 57 (79)
T smart00223 7 KNVDFRGS--DINTVYVPSAQVCQKRCTSHPRCLFFTFSTNEP-PEEKCLLKDS 57 (79)
T ss_pred cCccccCc--eeeeeecCCHHHHHHhhcCCCCccEEEeeCCCC-CCCEeEeCcC
Confidence 56777765 344446778999999999999999966643321 0117998654
No 16
>PF14295 PAN_4: PAN domain; PDB: 2YIL_E 2YIP_C 2YIO_A.
Probab=89.57 E-value=0.23 Score=34.49 Aligned_cols=37 Identities=38% Similarity=0.678 Sum_probs=16.6
Q ss_pred CCCCHHHHHHHhhccCCccceeccccc-CCCCCceeec
Q 048192 326 TNTDPETCKQACLKNCSCKAALFLYGL-NLSPGDCYLP 362 (422)
Q Consensus 326 ~~~s~~~C~~~CL~nCsC~a~~y~~~~-~~~~g~C~~~ 362 (422)
...++++|.+.|..+=.|.++.|.... ..+.+.|+++
T Consensus 14 ~~~s~~~C~~~C~~~~~C~~~~~~~~~~~~~~~~C~LK 51 (51)
T PF14295_consen 14 TASSPEECQAACAADPGCQAFTFNPPGCPSSSGRCYLK 51 (51)
T ss_dssp ----HHHHHHHHHTSTT--EEEEETTEE----------
T ss_pred cCCCHHHHHHHccCCCCCCEEEEECCCcccccccccCC
Confidence 566899999999999999995554311 1134678763
No 17
>PF00024 PAN_1: PAN domain This Prosite entry concerns apple domains, a subset of PAN domains; InterPro: IPR003014 PAN domains have significant functional versatility fulfilling diverse biological functions by mediating protein-protein or protein-carbohydrate interactions []. These domains contain a hair-pin loop like structure, similar to knottins, but the pattern of disulphide bonds differs It has been shown that, the N-terminal N domains of members of the plasminogen/hepatocyte growth factor family, the apple domains of the plasma prekallikrein/coagulation factor XI family, and domains of various nematode proteins belong to the same module superfamily, the PAN module []. PAN contains a conserved core of three disulphide bridges. In some members of the family there is an additional fourth disulphide bridge that links the N and C termini of the domain.; PDB: 1GP9_C 2QJ2_B 1GMO_H 1NK1_B 3MKP_B 1BHT_B 3HN4_A 1GMN_A 3HMS_A 3HMT_B ....
Probab=89.07 E-value=0.21 Score=37.89 Aligned_cols=51 Identities=25% Similarity=0.459 Sum_probs=34.3
Q ss_pred eEEeCCcccCCCCCCCCCcCCCCHHHHHHHhhccCC-ccceecccccCCCCCceeecc
Q 048192 307 FVELNDVAYFAFSSPSSDLTNTDPETCKQACLKNCS-CKAALFLYGLNLSPGDCYLPS 363 (422)
Q Consensus 307 f~~l~~~~~~~~~~~~~~~~~~s~~~C~~~CL~nCs-C~a~~y~~~~~~~~g~C~~~~ 363 (422)
|.++++..+... ........++++|.+.|+.+=. |.++.|.. .++.|+++.
T Consensus 4 f~~~~~~~l~~~--~~~~~~v~s~~~C~~~C~~~~~~C~s~~y~~----~~~~C~L~~ 55 (79)
T PF00024_consen 4 FERIPGYRLSGH--SIKEINVPSLEECAQLCLNEPRRCKSFNYDP----SSKTCYLSS 55 (79)
T ss_dssp EEEEEEEEEESC--EEEEEEESSHHHHHHHHHHSTT-ESEEEEET----TTTEEEEEC
T ss_pred eEEECCEEEeCC--cceEEcCCCHHHHHhhcCcCcccCCeEEEEC----CCCEEEEcC
Confidence 556666555442 1222244589999999999999 99955543 257899753
No 18
>cd01099 PAN_AP_HGF Subfamily of PAN/APPLE-like domains; present in N-terminal (N) domains of plasminogen/hepatocyte growth factor proteins, and various proteins found in Bilateria, such as leech anti-platelet proteins. PAN/APPLE domains fulfill diverse biological functions by mediating protein-protein or protein-carbohydrate interactions.
Probab=83.57 E-value=1.5 Score=33.99 Aligned_cols=34 Identities=26% Similarity=0.666 Sum_probs=27.4
Q ss_pred CCCCHHHHHHHhhc--cCCccceecccccCCCCCceeecc
Q 048192 326 TNTDPETCKQACLK--NCSCKAALFLYGLNLSPGDCYLPS 363 (422)
Q Consensus 326 ~~~s~~~C~~~CL~--nCsC~a~~y~~~~~~~~g~C~~~~ 363 (422)
...++++|.++|++ +=.|.++.|.+. ++.|.+..
T Consensus 23 ~~~s~~~C~~~C~~~~~f~CrSf~y~~~----~~~C~L~~ 58 (80)
T cd01099 23 TVASLEECLRKCLEETEFTCRSFNYNYK----SKECILSD 58 (80)
T ss_pred ecCCHHHHHHHhCCCCCceEeEEEEEcC----CCEEEEeC
Confidence 45799999999999 999999666553 57899853
No 19
>PF01683 EB: EB module; InterPro: IPR006149 The EB domain has no known function. It is found in several Caenorhabditis sp. and Drosophila sp. proteins. The domain contains 8 conserved cysteines that probably form four disulphide bridges and is found associated with kunitz domains IPR002223 from INTERPRO
Probab=83.55 E-value=1.4 Score=31.00 Aligned_cols=33 Identities=21% Similarity=0.566 Sum_probs=27.7
Q ss_pred ccCCCCCCCcCCCCCceeCCCccCCCCCCceecC
Q 048192 250 GYHGECGYPMVCGKYGICSQGQCSCPATYFKLLN 283 (422)
Q Consensus 250 ~p~d~C~v~g~CG~~giC~~~~C~C~~g~F~~~~ 283 (422)
.+.+.|....-|-.++.|....|.|++| |.+.+
T Consensus 17 ~~g~~C~~~~qC~~~s~C~~g~C~C~~g-~~~~~ 49 (52)
T PF01683_consen 17 QPGESCESDEQCIGGSVCVNGRCQCPPG-YVEVG 49 (52)
T ss_pred CCCCCCCCcCCCCCcCEEcCCEeECCCC-CEecC
Confidence 3557899999999999998899999998 76643
No 20
>PF07645 EGF_CA: Calcium-binding EGF domain; InterPro: IPR001881 A sequence of about forty amino-acid residues found in epidermal growth factor (EGF) has been shown [, , , , , ] to be present in a large number of membrane-bound and extracellular, mostly animal, proteins. Many of these proteins require calcium for their biological function and a calcium-binding site has been found at the N terminus of some EGF-like domains []. Calcium-binding may be crucial for numerous protein-protein interactions. For human coagulation factor IX it has been shown [] that the calcium-ligands form a pentagonal bipyramid. The first, third and fourth conserved negatively charged or polar residues are side chain ligands. The latter is possibly hydroxylated (see aspartic acid and asparagine hydroxylation site) []. A conserved aromatic residue, as well as the second conserved negative residue, are thought to be involved in stabilising the calcium-binding site. As in non-calcium binding EGF-like domains, there are six conserved cysteines and the structure of both types is very similar as calcium-binding induces only strictly local structural changes []. +------------------+ +---------+ | | | | nxnnC-x(3,14)-C-x(3,7)-CxxbxxxxaxC-x(1,6)-C-x(8,13)-Cx | | +------------------+ 'n': negatively charged or polar residue [DEQN] 'b': possibly beta-hydroxylated residue [DN] 'a': aromatic amino acid 'C': cysteine, involved in disulphide bond 'x': any amino acid. ; GO: 0005509 calcium ion binding; PDB: 2VJ3_A 1TOZ_A 1LMJ_A 1UZQ_A 1UZK_A 1UZJ_B 1UZP_A 1EMO_A 1EMN_A 2RR0_A ....
Probab=68.58 E-value=3.3 Score=27.80 Aligned_cols=28 Identities=36% Similarity=0.872 Sum_probs=22.5
Q ss_pred CCCCCC-cCCCCCceeCC----CccCCCCCCcee
Q 048192 253 GECGYP-MVCGKYGICSQ----GQCSCPATYFKL 281 (422)
Q Consensus 253 d~C~v~-g~CG~~giC~~----~~C~C~~g~F~~ 281 (422)
|+|... ..|..++.|.+ -.|.|++| |+.
T Consensus 3 dEC~~~~~~C~~~~~C~N~~Gsy~C~C~~G-y~~ 35 (42)
T PF07645_consen 3 DECAEGPHNCPENGTCVNTEGSYSCSCPPG-YEL 35 (42)
T ss_dssp STTTTTSSSSSTTSEEEEETTEEEEEESTT-EEE
T ss_pred cccCCCCCcCCCCCEEEcCCCCEEeeCCCC-cEE
Confidence 788875 48999999974 36999998 874
No 21
>PF13360 PQQ_2: PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=65.38 E-value=92 Score=28.30 Aligned_cols=72 Identities=25% Similarity=0.427 Sum_probs=44.2
Q ss_pred CCcEEEEcCC----CCC----CCCCcEEE-EecCccEEEEcC-CCCEEEeeCCCCC---c-e---eEEEEe-cCCCeeEE
Q 048192 42 FPQVVWSANR----NNL----VRINATLE-LTSDGNLVLQDA-DGAIAWSTNTSGK---S-V---VGLNLT-DMGNLVLF 103 (422)
Q Consensus 42 ~~~vVW~ANr----~~p----v~~~~~L~-l~~~G~LvL~~~-~~~~vWst~~~~~---~-~---~~~~Ll-dsGNLVL~ 103 (422)
....+|..+- ..+ +.++..+- .+.+|.|+..|. .|.++|+...... . . ..+.+. .+|-|...
T Consensus 12 tG~~~W~~~~~~~~~~~~~~~~~~~~~v~~~~~~~~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~v~v~~~~~~l~~~ 91 (238)
T PF13360_consen 12 TGKELWSYDLGPGIGGPVATAVPDGGRVYVASGDGNLYALDAKTGKVLWRFDLPGPISGAPVVDGGRVYVGTSDGSLYAL 91 (238)
T ss_dssp TTEEEEEEECSSSCSSEEETEEEETTEEEEEETTSEEEEEETTTSEEEEEEECSSCGGSGEEEETTEEEEEETTSEEEEE
T ss_pred CCCEEEEEECCCCCCCccceEEEeCCEEEEEcCCCEEEEEECCCCCEEEEeeccccccceeeecccccccccceeeeEec
Confidence 4568888753 222 22333333 358899999996 7999999876432 1 0 012222 34456666
Q ss_pred c-CCCcEEEec
Q 048192 104 D-KNNAAVWQS 113 (422)
Q Consensus 104 ~-~~~~~lWQS 113 (422)
| .+|.++|+.
T Consensus 92 d~~tG~~~W~~ 102 (238)
T PF13360_consen 92 DAKTGKVLWSI 102 (238)
T ss_dssp ETTTSCEEEEE
T ss_pred ccCCcceeeee
Confidence 7 678899995
No 22
>cd00053 EGF Epidermal growth factor domain, found in epidermal growth factor (EGF) presents in a large number of proteins, mostly animal; the list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied; the functional significance of EGF-like domains in what appear to be unrelated proteins is not yet clear; a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase); the domain includes six cysteine residues which have been shown to be involved in disulfide bonds; the main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet; Subdomains between the conserved cysteines vary in length; the region between the 5th and 6th cysteine contains two conserved glycines of which at least one is present in most EGF-like domains; a subset of these bind calcium.
Probab=64.24 E-value=5.3 Score=24.62 Aligned_cols=25 Identities=28% Similarity=0.848 Sum_probs=18.4
Q ss_pred CCCCcCCCCCceeCC----CccCCCCCCce
Q 048192 255 CGYPMVCGKYGICSQ----GQCSCPATYFK 280 (422)
Q Consensus 255 C~v~g~CG~~giC~~----~~C~C~~g~F~ 280 (422)
|.....|...+.|.. ..|.|++| |.
T Consensus 2 C~~~~~C~~~~~C~~~~~~~~C~C~~g-~~ 30 (36)
T cd00053 2 CAASNPCSNGGTCVNTPGSYRCVCPPG-YT 30 (36)
T ss_pred CCCCCCCCCCCEEecCCCCeEeECCCC-Cc
Confidence 443567888899973 57999997 54
No 23
>cd05845 Ig2_L1-CAM_like Second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM) and similar proteins. Ig2_L1-CAM_like: domain similar to the second immunoglobulin (Ig)-like domain of the L1 cell adhesion molecule (CAM). L1 belongs to the L1 subfamily of cell adhesion molecules (CAMs) and is comprised of an extracellular region having six Ig-like domains, five fibronectin type III domains, a transmembrane region and an intracellular domain. L1 is primarily expressed in the nervous system and is involved in its development and function. L1 is associated with an X-linked recessive disorder, X-linked hydrocephalus, MASA syndrome, or spastic paraplegia type 1, that involves abnormalities of axonal growth.
Probab=62.10 E-value=14 Score=29.85 Aligned_cols=35 Identities=11% Similarity=0.264 Sum_probs=24.7
Q ss_pred cCCCCcEEEEcCCCCCCCCCcEEEEecCccEEEEc
Q 048192 39 HTEFPQVVWSANRNNLVRINATLELTSDGNLVLQD 73 (422)
Q Consensus 39 ~~~~~~vVW~ANr~~pv~~~~~L~l~~~G~LvL~~ 73 (422)
..+..++.|+-+....+..+..+.++.+|+|.+.+
T Consensus 30 g~P~P~i~W~~~~~~~i~~~~Ri~~~~~GnL~fs~ 64 (95)
T cd05845 30 SAVPLRIYWMNSDLLHITQDERVSMGQNGNLYFAN 64 (95)
T ss_pred CCCCCEEEEECCCCccccccccEEECCCceEEEEE
Confidence 45778899995544456555677777788888754
No 24
>PF13360 PQQ_2: PQQ-like domain; PDB: 3HXJ_B 1YIQ_A 1KV9_A 3Q54_A 2YH3_A 3PRW_A 3P1L_A 3Q7M_A 3Q7O_A 3Q7N_A ....
Probab=60.13 E-value=23 Score=32.42 Aligned_cols=48 Identities=29% Similarity=0.565 Sum_probs=32.0
Q ss_pred cCccEEEEcC-CCCEEEeeCCC---CCce--e-----EEEE-ecCCCeeEEcC-CCcEEEe
Q 048192 65 SDGNLVLQDA-DGAIAWSTNTS---GKSV--V-----GLNL-TDMGNLVLFDK-NNAAVWQ 112 (422)
Q Consensus 65 ~~G~LvL~~~-~~~~vWst~~~---~~~~--~-----~~~L-ldsGNLVL~~~-~~~~lWQ 112 (422)
++|.|...|. .|.++|+.... ...+ + .+.. ..+|+|+..|. +|+++|+
T Consensus 1 ~~g~l~~~d~~tG~~~W~~~~~~~~~~~~~~~~~~~~~v~~~~~~~~l~~~d~~tG~~~W~ 61 (238)
T PF13360_consen 1 DDGTLSALDPRTGKELWSYDLGPGIGGPVATAVPDGGRVYVASGDGNLYALDAKTGKVLWR 61 (238)
T ss_dssp -TSEEEEEETTTTEEEEEEECSSSCSSEEETEEEETTEEEEEETTSEEEEEETTTSEEEEE
T ss_pred CCCEEEEEECCCCCEEEEEECCCCCCCccceEEEeCCEEEEEcCCCEEEEEECCCCCEEEE
Confidence 3688988897 78999998652 1111 1 1122 47777788885 7889997
No 25
>KOG4649 consensus PQQ (pyrrolo-quinoline quinone) repeat protein [Secondary metabolites biosynthesis, transport and catabolism]
Probab=58.84 E-value=19 Score=34.56 Aligned_cols=47 Identities=26% Similarity=0.412 Sum_probs=36.1
Q ss_pred CCCCcEEEEcCCCCCCCCC-----cEEEE-ecCccEEEEcCCCCEEEeeCCCC
Q 048192 40 TEFPQVVWSANRNNLVRIN-----ATLEL-TSDGNLVLQDADGAIAWSTNTSG 86 (422)
Q Consensus 40 ~~~~~vVW~ANr~~pv~~~-----~~L~l-~~~G~LvL~~~~~~~vWst~~~~ 86 (422)
..+.+..|-|.|..|+-.+ ..+.+ +-||+|.-+|+.|+.||...+.+
T Consensus 166 ~~~~~~~w~~~~~~PiF~splcv~~sv~i~~VdG~l~~f~~sG~qvwr~~t~G 218 (354)
T KOG4649|consen 166 PYSSTEFWAATRFGPIFASPLCVGSSVIITTVDGVLTSFDESGRQVWRPATKG 218 (354)
T ss_pred CCCcceehhhhcCCccccCceeccceEEEEEeccEEEEEcCCCcEEEeecCCC
Confidence 3456889999999998654 23444 46899999999999999876654
No 26
>PF07354 Sp38: Zona-pellucida-binding protein (Sp38); InterPro: IPR010857 This family contains a number of zona-pellucida-binding proteins that seem to be restricted to mammals. These are sperm proteins that bind to the 90 kDa family of zona pellucida glycoproteins in a calcium-dependent manner []. These represent some of the specific molecules that mediate the first steps of gamete interaction, allowing fertilisation to occur [].; GO: 0007339 binding of sperm to zona pellucida, 0005576 extracellular region
Probab=58.51 E-value=14 Score=35.32 Aligned_cols=35 Identities=17% Similarity=0.398 Sum_probs=31.2
Q ss_pred cCCCCcEEEEcCCCCCCCCCcEEEEecCccEEEEc
Q 048192 39 HTEFPQVVWSANRNNLVRINATLELTSDGNLVLQD 73 (422)
Q Consensus 39 ~~~~~~vVW~ANr~~pv~~~~~L~l~~~G~LvL~~ 73 (422)
++.+++..|+--.++++++++.+.||+.|.|++.|
T Consensus 9 E~iDP~y~W~GP~g~~l~gn~~~nIT~TG~L~~~~ 43 (271)
T PF07354_consen 9 ELIDPTYLWTGPNGKPLSGNSYVNITETGKLMFKN 43 (271)
T ss_pred ccCCCceEEECCCCcccCCCCeEEEccCceEEeec
Confidence 46678899999999999999999999999999876
No 27
>PF07974 EGF_2: EGF-like domain; InterPro: IPR013111 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length. This entry contains EGF domains found in a variety of extracellular and membrane proteins
Probab=57.62 E-value=9.1 Score=24.27 Aligned_cols=19 Identities=32% Similarity=0.929 Sum_probs=16.4
Q ss_pred cCCCCCceeC--CCccCCCCC
Q 048192 259 MVCGKYGICS--QGQCSCPAT 277 (422)
Q Consensus 259 g~CG~~giC~--~~~C~C~~g 277 (422)
..|...|.|. ..+|.|.+|
T Consensus 6 ~~C~~~G~C~~~~g~C~C~~g 26 (32)
T PF07974_consen 6 NICSGHGTCVSPCGRCVCDSG 26 (32)
T ss_pred CccCCCCEEeCCCCEEECCCC
Confidence 4799999998 379999998
No 28
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=55.39 E-value=47 Score=33.18 Aligned_cols=70 Identities=23% Similarity=0.429 Sum_probs=43.0
Q ss_pred CcEEEEcCCCCCC----------CCCcEEEE-ecCccEEEEc-CCCCEEEeeCCCCC---cee-----EEEEecCCCeeE
Q 048192 43 PQVVWSANRNNLV----------RINATLEL-TSDGNLVLQD-ADGAIAWSTNTSGK---SVV-----GLNLTDMGNLVL 102 (422)
Q Consensus 43 ~~vVW~ANr~~pv----------~~~~~L~l-~~~G~LvL~~-~~~~~vWst~~~~~---~~~-----~~~LldsGNLVL 102 (422)
..++|..+-..++ -....+-+ +.+|.|.-+| .+|.++|+.+.... +++ ...-..+|+|+-
T Consensus 40 ~~~~W~~~~~~~~~~~~~~~~p~v~~~~v~v~~~~g~v~a~d~~tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~a 119 (377)
T TIGR03300 40 VDQVWSASVGDGVGHYYLRLQPAVAGGKVYAADADGTVVALDAETGKRLWRVDLDERLSGGVGADGGLVFVGTEKGEVIA 119 (377)
T ss_pred ceeeeEEEcCCCcCccccccceEEECCEEEEECCCCeEEEEEccCCcEeeeecCCCCcccceEEcCCEEEEEcCCCEEEE
Confidence 4578887754433 22344444 4568888888 57899998775432 111 111234677777
Q ss_pred EcC-CCcEEEe
Q 048192 103 FDK-NNAAVWQ 112 (422)
Q Consensus 103 ~~~-~~~~lWQ 112 (422)
+|. +|+++|+
T Consensus 120 ld~~tG~~~W~ 130 (377)
T TIGR03300 120 LDAEDGKELWR 130 (377)
T ss_pred EECCCCcEeee
Confidence 775 6889997
No 29
>PF04478 Mid2: Mid2 like cell wall stress sensor; InterPro: IPR007567 This family represents a region near the C terminus of Mid2, which contains a transmembrane region. The remainder of the protein sequence is serine-rich and of low complexity, and is therefore impossible to align accurately. Mid2 is thought to act as a mechanosensor of cell wall stress. The C-terminal cytoplasmic region of Mid2 is known to interact with Rom2, a guanine nucleotide exchange factor (GEF) for Rho1, which is part of the cell wall integrity signalling pathway [].
Probab=54.71 E-value=2.7 Score=36.74 Aligned_cols=20 Identities=15% Similarity=0.389 Sum_probs=13.6
Q ss_pred cccceEEEEehhhhheeeee
Q 048192 401 TSHRKRIMGFILGSFFGLLV 420 (422)
Q Consensus 401 ~~~~~~i~~~~v~~~~~~~~ 420 (422)
++.|++|||++||+.+++||
T Consensus 45 ~knknIVIGvVVGVGg~ill 64 (154)
T PF04478_consen 45 SKNKNIVIGVVVGVGGPILL 64 (154)
T ss_pred cCCccEEEEEEecccHHHHH
Confidence 34557889999986555544
No 30
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=54.22 E-value=37 Score=34.32 Aligned_cols=56 Identities=29% Similarity=0.544 Sum_probs=35.5
Q ss_pred CcEEEE-ecCccEEEEcC-CCCEEEeeCCCCCc----ee----EEEEecCCCeeEEcC-CCcEEEec
Q 048192 58 NATLEL-TSDGNLVLQDA-DGAIAWSTNTSGKS----VV----GLNLTDMGNLVLFDK-NNAAVWQS 113 (422)
Q Consensus 58 ~~~L~l-~~~G~LvL~~~-~~~~vWst~~~~~~----~~----~~~LldsGNLVL~~~-~~~~lWQS 113 (422)
...+-+ +.+|.|+-+|. .|.++|+....+.. +. ...-..+|.|+-+|. +|+++|+-
T Consensus 120 ~~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~ssP~v~~~~v~v~~~~g~l~ald~~tG~~~W~~ 186 (394)
T PRK11138 120 GGKVYIGSEKGQVYALNAEDGEVAWQTKVAGEALSRPVVSDGLVLVHTSNGMLQALNESDGAVKWTV 186 (394)
T ss_pred CCEEEEEcCCCEEEEEECCCCCCcccccCCCceecCCEEECCEEEEECCCCEEEEEEccCCCEeeee
Confidence 344444 46788988885 68999998764321 11 112234567777775 68899984
No 31
>TIGR03066 Gem_osc_para_1 Gemmata obscuriglobus paralogous family TIGR03066. This model represents an uncharacterized paralogous family in Gemmata obscuriglobus UQM 2246, a member of the Planctomycetes. This family shows sequence similarity to TIGR03067, which is also found in Gemmata obscuriglobus as well as in a few other species.
Probab=51.91 E-value=42 Score=27.90 Aligned_cols=52 Identities=19% Similarity=0.278 Sum_probs=32.0
Q ss_pred CCcEEEEecCccEEEEcCCCCE------EEeeC---------CCCC----ceeEEEEecCCCeeEEcCCCcE
Q 048192 57 INATLELTSDGNLVLQDADGAI------AWSTN---------TSGK----SVVGLNLTDMGNLVLFDKNNAA 109 (422)
Q Consensus 57 ~~~~L~l~~~G~LvL~~~~~~~------vWst~---------~~~~----~~~~~~LldsGNLVL~~~~~~~ 109 (422)
+...|+|..+|.|+|..+++.- -|+-. ..+. .++. .=++.|-|||.|++|.+
T Consensus 34 ~~~~leF~~dGKL~v~~gnng~~~~~~Gty~L~G~kLtL~~~p~g~t~k~~Vtv-~~l~~~~Lvl~d~dg~~ 104 (111)
T TIGR03066 34 DDVVIEFAKDGKLVVTIGEKGKEVKADGTYKLDGNKLTLTLKAGGKEKKETLTV-KKLTDDELVGKDPDGKK 104 (111)
T ss_pred CceEEEEcCCCeEEEecCCCCcEeccCceEEEECCEEEEEEcCCCccccceEEE-EEecCCeEEEEcCCCCE
Confidence 4578999999999987654331 13321 1111 1222 23688999999998863
No 32
>smart00179 EGF_CA Calcium-binding EGF-like domain.
Probab=50.16 E-value=13 Score=23.67 Aligned_cols=27 Identities=30% Similarity=0.807 Sum_probs=20.2
Q ss_pred CCCCCCcCCCCCceeCC----CccCCCCCCce
Q 048192 253 GECGYPMVCGKYGICSQ----GQCSCPATYFK 280 (422)
Q Consensus 253 d~C~v~g~CG~~giC~~----~~C~C~~g~F~ 280 (422)
++|.....|...+.|.. -.|.|++| |.
T Consensus 3 ~~C~~~~~C~~~~~C~~~~g~~~C~C~~g-~~ 33 (39)
T smart00179 3 DECASGNPCQNGGTCVNTVGSYRCECPPG-YT 33 (39)
T ss_pred ccCcCCCCcCCCCEeECCCCCeEeECCCC-Cc
Confidence 66765567888889963 36999997 64
No 33
>PF01436 NHL: NHL repeat; InterPro: IPR001258 The NHL repeat, named after NCL-1, HT2A and Lin-41, is found largely in a large number of eukaryotic and prokaryotic proteins. For example, the repeat is found in a variety of enzymes of the copper type II, ascorbate-dependent monooxygenase family which catalyse the C terminus alpha-amidation of biological peptides []. In many it occurs in tandem arrays, for example in the ringfinger beta-box, coiled-coil (RBCC) eukaryotic growth regulators []. The 'Brain Tumor' protein (Brat) is one such growth regulator that contains a 6-bladed NHL-repeat beta-propeller [, ]. The NHL repeats are also found in serine/threonine protein kinase (STPK) in diverse range of pathogenic bacteria. These STPK are transmembrane receptors with a intracellular N-terminal kinase domain and extracellular C-terminal sensor domain. In the STPK, PknD, from Mycobacterium tuberculosis, the sensor domain forms a rigid, six-bladed b-propeller composed of NHL repeats with a flexible tether to the transmembrane domain.; GO: 0005515 protein binding; PDB: 3FVZ_A 3FW0_A 1RWL_A 1RWI_A 1Q7F_A.
Probab=49.93 E-value=29 Score=20.92 Aligned_cols=21 Identities=24% Similarity=0.376 Sum_probs=14.2
Q ss_pred EEEEecCccEEEEcCCCCEEE
Q 048192 60 TLELTSDGNLVLQDADGAIAW 80 (422)
Q Consensus 60 ~L~l~~~G~LvL~~~~~~~vW 80 (422)
-+.++.+|+|++.|..+.-||
T Consensus 6 gvav~~~g~i~VaD~~n~rV~ 26 (28)
T PF01436_consen 6 GVAVDSDGNIYVADSGNHRVQ 26 (28)
T ss_dssp EEEEETTSEEEEEECCCTEEE
T ss_pred EEEEeCCCCEEEEECCCCEEE
Confidence 466777888888876554444
No 34
>PRK11138 outer membrane biogenesis protein BamB; Provisional
Probab=48.74 E-value=53 Score=33.19 Aligned_cols=19 Identities=21% Similarity=0.235 Sum_probs=10.6
Q ss_pred ecCCCeeEEcC-CCcEEEec
Q 048192 95 TDMGNLVLFDK-NNAAVWQS 113 (422)
Q Consensus 95 ldsGNLVL~~~-~~~~lWQS 113 (422)
.++|.|...|. +|+++|+-
T Consensus 342 ~~~G~l~~ld~~tG~~~~~~ 361 (394)
T PRK11138 342 DSEGYLHWINREDGRFVAQQ 361 (394)
T ss_pred eCCCEEEEEECCCCCEEEEE
Confidence 44566665553 45566654
No 35
>smart00765 MANEC The MANEC domain was formerly called MANSC. This domain, comprising 8 conserved cysteines, is found in the N terminus of higher multicellular animal membrane and extracellular proteins. It is postulated that this domain may play a role in the formation of protein complexes involving various protease activators and inhibitors. It is possible that some of the cysteine residues in the MANSC domain form structurally important disulfide bridges. All of the MANSC-containing proteins contain predicted transmembrane regions and signal peptides. It has been proposed that the MANSC domain in HAI-1 might function through binding with hepatocyte growth factor activator and matriptase.
Probab=45.81 E-value=25 Score=28.24 Aligned_cols=38 Identities=29% Similarity=0.550 Sum_probs=28.9
Q ss_pred CCCCHHHHHHHhhccCCccceecccccCCCCCceeecc
Q 048192 326 TNTDPETCKQACLKNCSCKAALFLYGLNLSPGDCYLPS 363 (422)
Q Consensus 326 ~~~s~~~C~~~CL~nCsC~a~~y~~~~~~~~g~C~~~~ 363 (422)
...+.++|..+|=+.=+|..|+|.....++.+.|++..
T Consensus 36 ~~~s~edC~~aCC~~~~CnlAv~e~~~~~~~~~CyLf~ 73 (93)
T smart00765 36 AVNTWEDCVRACCSTPNCNLAVFELRREDAEGNCYLFN 73 (93)
T ss_pred ccCCHHHHHHHHcCCCCCcEEEEeccCCCCCCceEEEE
Confidence 34578999999999999999998653333467899743
No 36
>PF12661 hEGF: Human growth factor-like EGF; PDB: 2YGQ_A 2E26_A 3A7Q_A 2YGP_A 2YGO_A 1HRE_A 1HAE_A 1HAF_A 1HRF_A.
Probab=45.24 E-value=5.8 Score=19.87 Aligned_cols=9 Identities=33% Similarity=1.246 Sum_probs=5.9
Q ss_pred ccCCCCCCce
Q 048192 271 QCSCPATYFK 280 (422)
Q Consensus 271 ~C~C~~g~F~ 280 (422)
.|.|++| |.
T Consensus 1 ~C~C~~G-~~ 9 (13)
T PF12661_consen 1 TCQCPPG-WT 9 (13)
T ss_dssp EEEE-TT-EE
T ss_pred CccCcCC-Cc
Confidence 4899997 64
No 37
>cd00054 EGF_CA Calcium-binding EGF-like domain, present in a large number of membrane-bound and extracellular (mostly animal) proteins. Many of these proteins require calcium for their biological function and calcium-binding sites have been found to be located at the N-terminus of particular EGF-like domains; calcium-binding may be crucial for numerous protein-protein interactions. Six conserved core cysteines form three disulfide bridges as in non calcium-binding EGF domains, whose structures are very similar. EGF_CA can be found in tandem repeat arrangements.
Probab=44.42 E-value=18 Score=22.54 Aligned_cols=27 Identities=33% Similarity=0.832 Sum_probs=19.5
Q ss_pred CCCCCCcCCCCCceeCC----CccCCCCCCce
Q 048192 253 GECGYPMVCGKYGICSQ----GQCSCPATYFK 280 (422)
Q Consensus 253 d~C~v~g~CG~~giC~~----~~C~C~~g~F~ 280 (422)
++|.....|...+.|.. -.|.|++| |.
T Consensus 3 ~~C~~~~~C~~~~~C~~~~~~~~C~C~~g-~~ 33 (38)
T cd00054 3 DECASGNPCQNGGTCVNTVGSYRCSCPPG-YT 33 (38)
T ss_pred ccCCCCCCcCCCCEeECCCCCeEeECCCC-Cc
Confidence 56765457888889963 36999997 54
No 38
>cd05852 Ig5_Contactin-1 Fifth Ig domain of contactin-1. Ig5_Contactin-1: fifth Ig domain of the neural cell adhesion molecule contactin-1. Contactins are comprised of six Ig domains followed by four fibronectin type III (FnIII) domains anchored to the membrane by glycosylphosphatidylinositol. Contactin-1 is differentially expressed in tumor tissues and may through a RhoA mechanism, facilitate invasion and metastasis of human lung adenocarcinoma.
Probab=42.52 E-value=90 Score=23.24 Aligned_cols=34 Identities=24% Similarity=0.467 Sum_probs=23.8
Q ss_pred CCCCcEEEEcCCCCCCCCCcEEEEecCccEEEEcC
Q 048192 40 TEFPQVVWSANRNNLVRINATLELTSDGNLVLQDA 74 (422)
Q Consensus 40 ~~~~~vVW~ANr~~pv~~~~~L~l~~~G~LvL~~~ 74 (422)
.|.+++.|.=+.. ++..+..+.+..+|.|+|.+.
T Consensus 13 ~P~p~v~W~k~~~-~l~~~~r~~~~~~g~L~I~~v 46 (73)
T cd05852 13 APKPKFSWSKGTE-LLVNNSRISIWDDGSLEILNI 46 (73)
T ss_pred eCCCEEEEEeCCE-ecccCCCEEEcCCCEEEECcC
Confidence 3567899986643 555555677777899988753
No 39
>TIGR03300 assembly_YfgL outer membrane assembly lipoprotein YfgL. Members of this protein family are YfgL, a lipoprotein component of a complex that acts protein insertion into the bacterial outer membrane. Other members of this complex are NlpB, YfiO, and YaeT. This protein contains multiple copies of a repeat that, in other contexts, are associated with binding of the coenzyme PQQ.
Probab=36.72 E-value=1.9e+02 Score=28.70 Aligned_cols=71 Identities=20% Similarity=0.394 Sum_probs=44.3
Q ss_pred CCcEEEEcCCCC-----CCCCCcEEEE-ecCccEEEEcC-CCCEEEeeCCCCCce--------eEEEEecCCCeeEEcC-
Q 048192 42 FPQVVWSANRNN-----LVRINATLEL-TSDGNLVLQDA-DGAIAWSTNTSGKSV--------VGLNLTDMGNLVLFDK- 105 (422)
Q Consensus 42 ~~~vVW~ANr~~-----pv~~~~~L~l-~~~G~LvL~~~-~~~~vWst~~~~~~~--------~~~~LldsGNLVL~~~- 105 (422)
.-.++|.-+-.. |+-++..+-+ +.+|.|+.+|. .|.++|+....+... ....-..+|.|+..|.
T Consensus 84 tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~ald~~tG~~~W~~~~~~~~~~~p~v~~~~v~v~~~~g~l~a~d~~ 163 (377)
T TIGR03300 84 TGKRLWRVDLDERLSGGVGADGGLVFVGTEKGEVIALDAEDGKELWRAKLSSEVLSPPLVANGLVVVRTNDGRLTALDAA 163 (377)
T ss_pred CCcEeeeecCCCCcccceEEcCCEEEEEcCCCEEEEEECCCCcEeeeeccCceeecCCEEECCEEEEECCCCeEEEEEcC
Confidence 456789755443 3333444544 46899998886 689999877543211 1112235677887775
Q ss_pred CCcEEEe
Q 048192 106 NNAAVWQ 112 (422)
Q Consensus 106 ~~~~lWQ 112 (422)
+|+++|+
T Consensus 164 tG~~~W~ 170 (377)
T TIGR03300 164 TGERLWT 170 (377)
T ss_pred CCceeeE
Confidence 6789998
No 40
>PHA00149 DNA encapsidation protein
Probab=35.67 E-value=1.2e+02 Score=29.71 Aligned_cols=55 Identities=15% Similarity=0.286 Sum_probs=37.6
Q ss_pred CCCCeEEEeeecCCCCCceEEEEEeeccccccccccccCCCCcEEEEcCCCCCCCCC-cEEEEe--cCccEEEEc
Q 048192 2 TFGPTYACGFFCNGTCDSYLFAVFIVHAYDASLIEYQHTEFPQVVWSANRNNLVRIN-ATLELT--SDGNLVLQD 73 (422)
Q Consensus 2 ~~~~~F~~GF~~~~~~~~~~l~Iw~~~~~~~~~~~~~~~~~~~vVW~ANr~~pv~~~-~~L~l~--~~G~LvL~~ 73 (422)
+.++.|.++++-+++ +++||... .+-.||.|.+-.|=+.. -.|+.. ++|...|.+
T Consensus 235 ~~~~k~~ysi~~~g~----~~~vwvd~-------------~~~~~y~~~~~dp~~~~v~alt~~dl~e~~vll~~ 292 (331)
T PHA00149 235 SKNSKFVFSIRYNGN----YYTVWVDL-------------TQMLVYIATAHDPSTKRVYALTVDDLEEGMVLLIN 292 (331)
T ss_pred ccCceEEEEEEECCe----EEEEEEEc-------------cceEEEEecccCCCCCceEEEEccccccCcEEehh
Confidence 467899999998774 78999743 56789999998885554 234432 345555433
No 41
>PF10681 Rot1: Chaperone for protein-folding within the ER, fungal; InterPro: IPR019623 This conserved fungal family is an essential molecular chaperone in the endoplasmic reticulum. Molecular chaperones transiently interact with unfolded proteins to inhibit their self-aggregation and to support their folding and/or assembly. Rot1 is a general chaperone with some substrate specificity, its substrates being the structurally unrelated Kre5 Kre6 Big1 Atg22, which are type I, type II, and polytopic membrane proteins. The dependencies of each for Rot1 do not share similarities. However, their folding does require BiP, and one of these proteins was simultaneously associated with both Rot1 and BiP. In addition, Rot1 may cooperate with BiP/Kar2 in the folding of Kre6 [].
Probab=35.58 E-value=1.4e+02 Score=27.72 Aligned_cols=84 Identities=26% Similarity=0.331 Sum_probs=50.2
Q ss_pred cEEEEcCCCCCC--C----CC-cEEEEecCccEEEE--cCCCCEEEeeCCCCCcee------------EEEEecC--C--
Q 048192 44 QVVWSANRNNLV--R----IN-ATLELTSDGNLVLQ--DADGAIAWSTNTSGKSVV------------GLNLTDM--G-- 98 (422)
Q Consensus 44 ~vVW~ANr~~pv--~----~~-~~L~l~~~G~LvL~--~~~~~~vWst~~~~~~~~------------~~~Llds--G-- 98 (422)
....++|..+|- + .- ++-+|.++|.|+|. ..||+...|..++..... -....|. |
T Consensus 48 ~Yr~~~Np~~p~C~~a~l~wQHGtY~l~~nGsl~L~P~~~DGrQl~sdPC~~~~s~y~rYnq~e~f~~~~v~~D~y~~~~ 127 (212)
T PF10681_consen 48 QYRVTSNPTNPSCPTAVLIWQHGTYELNSNGSLTLTPFAVDGRQLVSDPCADDSSTYTRYNQTELFKSFDVYVDPYHGRY 127 (212)
T ss_pred EEEEccCCCCCCCCceEEEEecceEEECCCCcEEEeecCCCCceeccCCCCCCcccEEEEcceEEEEEEEEEEeCCCCee
Confidence 457788877762 1 12 66778789999986 467888777766432111 0122332 2
Q ss_pred CeeEEcCCCc---EEEeccCCCCCccCCCceecC
Q 048192 99 NLVLFDKNNA---AVWQSFDHPTDSLVPGQKLLE 129 (422)
Q Consensus 99 NLVL~~~~~~---~lWQSFd~PTDTlLpgq~l~~ 129 (422)
.|.|.+-+|+ ++|-=..-|. |||.+.|..
T Consensus 128 ~L~L~~fDGsp~~pmyL~y~pP~--MLPT~tLnp 159 (212)
T PF10681_consen 128 RLQLYQFDGSPMQPMYLAYRPPM--MLPTQTLNP 159 (212)
T ss_pred EEEEEccCCCcCCcchhccCCcc--cCcCcccCc
Confidence 3555555553 5677666663 777777754
No 42
>PF12662 cEGF: Complement Clr-like EGF-like
Probab=34.70 E-value=16 Score=21.64 Aligned_cols=10 Identities=50% Similarity=1.418 Sum_probs=8.2
Q ss_pred ccCCCCCCcee
Q 048192 271 QCSCPATYFKL 281 (422)
Q Consensus 271 ~C~C~~g~F~~ 281 (422)
.|+|++| |+.
T Consensus 3 ~C~C~~G-y~l 12 (24)
T PF12662_consen 3 TCSCPPG-YQL 12 (24)
T ss_pred EeeCCCC-CcC
Confidence 5999998 765
No 43
>PF12690 BsuPI: Intracellular proteinase inhibitor; InterPro: IPR020481 BsuPI is a intracellular proteinase inhibitor that directly regulates the major intracellular proteinase (ISP-1) activity in vivo. It inhibits ISP-1 in the early stages of sporulation and then may be inactivated by a membrane-bound proteinase [].; PDB: 3ISY_A.
Probab=34.10 E-value=46 Score=25.94 Aligned_cols=15 Identities=27% Similarity=0.683 Sum_probs=8.2
Q ss_pred EEEEcCCCCEEEeeC
Q 048192 69 LVLQDADGAIAWSTN 83 (422)
Q Consensus 69 LvL~~~~~~~vWst~ 83 (422)
|+|.|.+|..||.-.
T Consensus 28 ~~v~d~~g~~vwrwS 42 (82)
T PF12690_consen 28 FVVKDKEGKEVWRWS 42 (82)
T ss_dssp EEEE-TT--EEEETT
T ss_pred EEEECCCCCEEEEec
Confidence 677777777777554
No 44
>cd00216 PQQ_DH Dehydrogenases with pyrrolo-quinoline quinone (PQQ) as cofactor, like ethanol, methanol, and membrane bound glucose dehydrogenases. The alignment model contains an 8-bladed beta-propeller.
Probab=31.95 E-value=1.4e+02 Score=31.40 Aligned_cols=72 Identities=24% Similarity=0.386 Sum_probs=44.2
Q ss_pred CCCcEEEEcCCC-------CCCCCCcEEEE-ecCccEEEEcC-CCCEEEeeCCCCC-----------cee-----EEE-E
Q 048192 41 EFPQVVWSANRN-------NLVRINATLEL-TSDGNLVLQDA-DGAIAWSTNTSGK-----------SVV-----GLN-L 94 (422)
Q Consensus 41 ~~~~vVW~ANr~-------~pv~~~~~L~l-~~~G~LvL~~~-~~~~vWst~~~~~-----------~~~-----~~~-L 94 (422)
....++|..+-. .|+-.+.++-+ +.+|.|+-+|. .|.++|+...... +++ .+. -
T Consensus 37 ~~~~~~W~~~~~~~~~~~~sPvv~~g~vy~~~~~g~l~AlD~~tG~~~W~~~~~~~~~~~~~~~~~~g~~~~~~~~V~v~ 116 (488)
T cd00216 37 KKLKVAWTFSTGDERGQEGTPLVVDGDMYFTTSHSALFALDAATGKVLWRYDPKLPADRGCCDVVNRGVAYWDPRKVFFG 116 (488)
T ss_pred hcceeeEEEECCCCCCcccCCEEECCEEEEeCCCCcEEEEECCCChhhceeCCCCCccccccccccCCcEEccCCeEEEe
Confidence 445689987654 35444455544 45799988885 6889998765321 000 011 1
Q ss_pred ecCCCeeEEcC-CCcEEEe
Q 048192 95 TDMGNLVLFDK-NNAAVWQ 112 (422)
Q Consensus 95 ldsGNLVL~~~-~~~~lWQ 112 (422)
..+|.++-+|. +|+++|+
T Consensus 117 ~~~g~v~AlD~~TG~~~W~ 135 (488)
T cd00216 117 TFDGRLVALDAETGKQVWK 135 (488)
T ss_pred cCCCeEEEEECCCCCEeee
Confidence 24677777775 5789999
No 45
>PF05935 Arylsulfotrans: Arylsulfotransferase (ASST); InterPro: IPR010262 This family consists of several bacterial arylsulphotransferase proteins. Arylsulphotransferase (ASST) transfers a sulphate group from phenolic sulphate esters to a phenolic acceptor substrate [].; PDB: 3ETT_B 3ELQ_A 3ETS_A.
Probab=31.25 E-value=1.3e+02 Score=31.50 Aligned_cols=62 Identities=21% Similarity=0.269 Sum_probs=30.4
Q ss_pred CCcEEEEcCCCCCCC------CCcEEEEecCccEEEEcCCCCEEEeeCCCCCc---eeEEEEecCCCeeEE
Q 048192 42 FPQVVWSANRNNLVR------INATLELTSDGNLVLQDADGAIAWSTNTSGKS---VVGLNLTDMGNLVLF 103 (422)
Q Consensus 42 ~~~vVW~ANr~~pv~------~~~~L~l~~~G~LvL~~~~~~~vWst~~~~~~---~~~~~LldsGNLVL~ 103 (422)
.-.|+|.-....... .++.|.+.....|...|-.|.++|.-...+.. .-.+..+++||+.++
T Consensus 136 ~G~Vrw~~~~~~~~~~~~~~l~nG~ll~~~~~~~~e~D~~G~v~~~~~l~~~~~~~HHD~~~l~nGn~L~l 206 (477)
T PF05935_consen 136 NGDVRWYLPLDSGSDNSFKQLPNGNLLIGSGNRLYEIDLLGKVIWEYDLPGGYYDFHHDIDELPNGNLLIL 206 (477)
T ss_dssp TS-EEEEE-GGGT--SSEEE-TTS-EEEEEBTEEEEE-TT--EEEEEE--TTEE-B-S-EEE-TTS-EEEE
T ss_pred CccEEEEEccCccccceeeEcCCCCEEEecCCceEEEcCCCCEEEeeecCCcccccccccEECCCCCEEEE
Confidence 456888876654322 22333333345556667789999987654432 345678899999986
No 46
>PF06006 DUF905: Bacterial protein of unknown function (DUF905); InterPro: IPR009253 This family consists of several short hypothetical proteobacterial proteins of unknown function.; PDB: 2HJJ_A.
Probab=30.71 E-value=48 Score=24.96 Aligned_cols=18 Identities=28% Similarity=0.678 Sum_probs=10.1
Q ss_pred eeEEcCCCcEEEeccCCC
Q 048192 100 LVLFDKNNAAVWQSFDHP 117 (422)
Q Consensus 100 LVL~~~~~~~lWQSFd~P 117 (422)
||+||.+|..+|..|.+-
T Consensus 35 lvvRd~~g~mvWRaWNFE 52 (70)
T PF06006_consen 35 LVVRDTEGQMVWRAWNFE 52 (70)
T ss_dssp EEEE-SS--EEEEEESSS
T ss_pred EEEEcCCCcEEEEeeccC
Confidence 577777777777766653
No 47
>PF13570 PQQ_3: PQQ-like domain; PDB: 3HXJ_B 3Q54_A.
Probab=29.87 E-value=55 Score=21.22 Aligned_cols=11 Identities=45% Similarity=1.039 Sum_probs=4.9
Q ss_pred CCEEEeeCCCC
Q 048192 76 GAIAWSTNTSG 86 (422)
Q Consensus 76 ~~~vWst~~~~ 86 (422)
|.++|+..+.+
T Consensus 1 G~~~W~~~~~~ 11 (40)
T PF13570_consen 1 GKVLWSYDTGG 11 (40)
T ss_dssp S-EEEEEE-SS
T ss_pred CceeEEEECCC
Confidence 34566665543
No 48
>PF09064 Tme5_EGF_like: Thrombomodulin like fifth domain, EGF-like; InterPro: IPR015149 This domain adopts a fold similar to other EGF domains, with a flat major and a twisted minor beta sheet. Disulphide pairing, however, is not of the usual 1-3, 2-4, 5-6 type; rather 1-2, 3-4, 5-6 pairing is found. Its extended major sheet (strands beta-2 and beta-3 and the connecting loop) projects into thrombin's active site groove. This domain is required for interaction of thrombomodulin with thrombin, and subsequent activation of protein-C []. ; GO: 0004888 transmembrane signaling receptor activity, 0016021 integral to membrane
Probab=27.89 E-value=42 Score=21.65 Aligned_cols=11 Identities=55% Similarity=1.301 Sum_probs=8.5
Q ss_pred CccCCCCCCcee
Q 048192 270 GQCSCPATYFKL 281 (422)
Q Consensus 270 ~~C~C~~g~F~~ 281 (422)
.+|.||.| |-.
T Consensus 18 ~~C~CPeG-yIl 28 (34)
T PF09064_consen 18 GQCFCPEG-YIL 28 (34)
T ss_pred CceeCCCc-eEe
Confidence 58999998 643
No 49
>PF05935 Arylsulfotrans: Arylsulfotransferase (ASST); InterPro: IPR010262 This family consists of several bacterial arylsulphotransferase proteins. Arylsulphotransferase (ASST) transfers a sulphate group from phenolic sulphate esters to a phenolic acceptor substrate [].; PDB: 3ETT_B 3ELQ_A 3ETS_A.
Probab=27.76 E-value=61 Score=34.00 Aligned_cols=53 Identities=23% Similarity=0.408 Sum_probs=31.1
Q ss_pred CccEEEEcCCCCEEEeeCCCCCceeEEEEecCCCeeEE--------cCCCcEEEeccCCCCC
Q 048192 66 DGNLVLQDADGAIAWSTNTSGKSVVGLNLTDMGNLVLF--------DKNNAAVWQSFDHPTD 119 (422)
Q Consensus 66 ~G~LvL~~~~~~~vWst~~~~~~~~~~~LldsGNLVL~--------~~~~~~lWQSFd~PTD 119 (422)
.+..++.|.+|.++|-..........+..+++|+|... |-.|+++|+ ++.|..
T Consensus 127 ~~~~~~iD~~G~Vrw~~~~~~~~~~~~~~l~nG~ll~~~~~~~~e~D~~G~v~~~-~~l~~~ 187 (477)
T PF05935_consen 127 SSYTYLIDNNGDVRWYLPLDSGSDNSFKQLPNGNLLIGSGNRLYEIDLLGKVIWE-YDLPGG 187 (477)
T ss_dssp EEEEEEEETTS-EEEEE-GGGT--SSEEE-TTS-EEEEEBTEEEEE-TT--EEEE-EE--TT
T ss_pred CceEEEECCCccEEEEEccCccccceeeEcCCCCEEEecCCceEEEcCCCCEEEe-eecCCc
Confidence 47788999999999987653332222678899999865 345789999 776663
No 50
>PLN00033 photosystem II stability/assembly factor; Provisional
Probab=27.66 E-value=1.9e+02 Score=29.58 Aligned_cols=51 Identities=18% Similarity=0.328 Sum_probs=29.5
Q ss_pred EEecCccEEEEcCCCCEEEeeCCC--CCceeEEEEecCCCeeEEcCCCcEEEe
Q 048192 62 ELTSDGNLVLQDADGAIAWSTNTS--GKSVVGLNLTDMGNLVLFDKNNAAVWQ 112 (422)
Q Consensus 62 ~l~~~G~LvL~~~~~~~vWst~~~--~~~~~~~~LldsGNLVL~~~~~~~lWQ 112 (422)
.+...|++++.+.+|..-|..-.. ......+...++|.|||....|.++|.
T Consensus 254 ~vg~~G~~~~s~d~G~~~W~~~~~~~~~~l~~v~~~~dg~l~l~g~~G~l~~S 306 (398)
T PLN00033 254 AVSSRGNFYLTWEPGQPYWQPHNRASARRIQNMGWRADGGLWLLTRGGGLYVS 306 (398)
T ss_pred EEECCccEEEecCCCCcceEEecCCCccceeeeeEcCCCCEEEEeCCceEEEe
Confidence 334445555444455555654322 223455667789999998877766553
No 51
>COG1520 FOG: WD40-like repeat [Function unknown]
Probab=24.97 E-value=2.9e+02 Score=27.51 Aligned_cols=73 Identities=26% Similarity=0.426 Sum_probs=44.3
Q ss_pred CCcEEEEcCCCC-------CCCCCcEEEEe-cCccEEEEcCC-CCEEEeeCCCC---C----cee----EEEE-ec--CC
Q 048192 42 FPQVVWSANRNN-------LVRINATLELT-SDGNLVLQDAD-GAIAWSTNTSG---K----SVV----GLNL-TD--MG 98 (422)
Q Consensus 42 ~~~vVW~ANr~~-------pv~~~~~L~l~-~~G~LvL~~~~-~~~vWst~~~~---~----~~~----~~~L-ld--sG 98 (422)
.-+.+|..+... |+.....+-+. .+|.|+-+|++ |..+|...... . ... .+.+ .+ +|
T Consensus 130 ~G~~~W~~~~~~~~~~~~~~v~~~~~v~~~s~~g~~~al~~~tG~~~W~~~~~~~~~~~~~~~~~~~~~~vy~~~~~~~~ 209 (370)
T COG1520 130 TGTLVWSRNVGGSPYYASPPVVGDGTVYVGTDDGHLYALNADTGTLKWTYETPAPLSLSIYGSPAIASGTVYVGSDGYDG 209 (370)
T ss_pred CCcEEEEEecCCCeEEecCcEEcCcEEEEecCCCeEEEEEccCCcEEEEEecCCccccccccCceeecceEEEecCCCcc
Confidence 456889887766 23334555555 67999988876 89999854421 1 000 1111 22 45
Q ss_pred CeeEEcC-CCcEEEecc
Q 048192 99 NLVLFDK-NNAAVWQSF 114 (422)
Q Consensus 99 NLVL~~~-~~~~lWQSF 114 (422)
+|+=.|. +|..+|+-+
T Consensus 210 ~~~a~~~~~G~~~w~~~ 226 (370)
T COG1520 210 ILYALNAEDGTLKWSQK 226 (370)
T ss_pred eEEEEEccCCcEeeeee
Confidence 6776666 678899854
No 52
>TIGR02513 type_III_yscB type III secretion system chaperone, YscB family. Members of this family include YscB of Yersinia and functionally equivalent (but differently named) proteins from type III secretion systems of other pathogens that affect animal cells. YscB acts, along with SycN (TIGR02503), as a chaperone for YopN, a key part of a complex that regulates type III secretion so it responds to contact with the eukaryotic target cell.
Probab=24.37 E-value=1.3e+02 Score=25.79 Aligned_cols=45 Identities=24% Similarity=0.296 Sum_probs=29.0
Q ss_pred cEEEEecCccEEE-EcCCCCEEEeeCCCCC--------------------------ceeEEEEecCCCeeEE
Q 048192 59 ATLELTSDGNLVL-QDADGAIAWSTNTSGK--------------------------SVVGLNLTDMGNLVLF 103 (422)
Q Consensus 59 ~~L~l~~~G~LvL-~~~~~~~vWst~~~~~--------------------------~~~~~~LldsGNLVL~ 103 (422)
++-.|.-||-++. .-..+..+|+|..... ....++|.|+|||+|.
T Consensus 22 G~Yhl~iD~~~l~l~q~~sellletpL~~~~~~~~d~q~~~lLk~lmQq~l~w~R~~p~aLvld~~~qLiLe 93 (139)
T TIGR02513 22 GVYHLTIDQHLVMLAQHGSELVLETPLDARMLRPGDNQNVTLLRSLMQQVLAWARRYPQALVLDADGQLILE 93 (139)
T ss_pred CceEEEEcCcEEEeeccCceEEEeccccchhhCccccccHHHHHHHHHHHHHHHhcCCceEEEcCccchhHH
Confidence 4444444666554 4445568999986320 1236789999999986
No 53
>PF12947 EGF_3: EGF domain; InterPro: IPR024731 This entry represents an EGF domain found in the the C terminus of malarial parasite merozoite surface protein 1 [], as well as other proteins.; PDB: 2NPR_A 1N1I_C 1B9W_A 1YO8_A 2RHP_A.
Probab=24.30 E-value=17 Score=23.71 Aligned_cols=21 Identities=19% Similarity=0.640 Sum_probs=14.6
Q ss_pred cCCCCCceeCC----CccCCCCCCce
Q 048192 259 MVCGKYGICSQ----GQCSCPATYFK 280 (422)
Q Consensus 259 g~CG~~giC~~----~~C~C~~g~F~ 280 (422)
+-|.++..|.. -.|.|.+| |+
T Consensus 6 ~~C~~nA~C~~~~~~~~C~C~~G-y~ 30 (36)
T PF12947_consen 6 GGCHPNATCTNTGGSYTCTCKPG-YE 30 (36)
T ss_dssp GGS-TTCEEEE-TTSEEEEE-CE-EE
T ss_pred CCCCCCcEeecCCCCEEeECCCC-Cc
Confidence 56889999973 46999997 64
No 54
>COG3236 Uncharacterized protein conserved in bacteria [Function unknown]
Probab=23.21 E-value=45 Score=28.97 Aligned_cols=20 Identities=35% Similarity=0.634 Sum_probs=15.9
Q ss_pred EEecCCCeeEEc-CCCcEEEe
Q 048192 93 NLTDMGNLVLFD-KNNAAVWQ 112 (422)
Q Consensus 93 ~LldsGNLVL~~-~~~~~lWQ 112 (422)
.||+||+.||.. +.+..+|-
T Consensus 117 ~LL~Tgd~vLVE~s~~D~~WG 137 (162)
T COG3236 117 LLLATGDAVLVEASPNDAIWG 137 (162)
T ss_pred HHHhcCCeeEEecCCCcceee
Confidence 589999999995 45667884
No 55
>PF02237 BPL_C: Biotin protein ligase C terminal domain; InterPro: IPR003142 This C-terminal domain has an SH3-like barrel fold, the function of which is unknown. It is found associated with prokaryotic bifunctional transcriptional repressors [] and eukaryotic enzymes involved in biotin utilization [, ]. In Escherichia coli the biotin operon repressor (BirA) is a bifunctional protein. BirA acts both as the acetyl-coA carboxylase biotin holoenzyme synthetase (6.3.4.15 from EC) and as the biotin operon repressor. DNA sequence analysis of mutations indicates that the helix-turn-helix DNA binding region is located at the N terminus while mutations affecting enzyme function, although mapping over a large region, are found mainly in the central part of the protein's primary sequence [].; GO: 0006464 protein modification process; PDB: 3RUX_A 2CGH_A 3L1A_B 3L2Z_A 1HXD_A 1BIB_A 2EWN_B 1BIA_A 2EJ9_A 3FJP_A ....
Probab=22.42 E-value=63 Score=22.23 Aligned_cols=15 Identities=20% Similarity=0.432 Sum_probs=8.0
Q ss_pred EEecCCCeeEEcCCC
Q 048192 93 NLTDMGNLVLFDKNN 107 (422)
Q Consensus 93 ~LldsGNLVL~~~~~ 107 (422)
-+.|+|.|+|+.+++
T Consensus 21 gId~~G~L~v~~~~g 35 (48)
T PF02237_consen 21 GIDDDGALLVRTEDG 35 (48)
T ss_dssp EEETTSEEEEEETTE
T ss_pred EECCCCEEEEEECCC
Confidence 445555555555444
No 56
>smart00564 PQQ beta-propeller repeat. Beta-propeller repeat occurring in enzymes with pyrrolo-quinoline quinone (PQQ) as cofactor, in Ire1p-like Ser/Thr kinases, and in prokaryotic dehydrogenases.
Probab=21.84 E-value=1e+02 Score=18.61 Aligned_cols=19 Identities=42% Similarity=0.793 Sum_probs=11.3
Q ss_pred ecCccEEEEcC-CCCEEEee
Q 048192 64 TSDGNLVLQDA-DGAIAWST 82 (422)
Q Consensus 64 ~~~G~LvL~~~-~~~~vWst 82 (422)
+.+|.|+-.|. +|.++|..
T Consensus 13 ~~~g~l~a~d~~~G~~~W~~ 32 (33)
T smart00564 13 STDGTLYALDAKTGEILWTY 32 (33)
T ss_pred cCCCEEEEEEcccCcEEEEc
Confidence 34566666664 56667753
No 57
>PF00008 EGF: EGF-like domain This is a sub-family of the Pfam entry This is a sub-family of the Pfam entry; InterPro: IPR006209 A sequence of about thirty to forty amino-acid residues long found in the sequence of epidermal growth factor (EGF) has been shown [, , , , ] to be present, in a more or less conserved form, in a large number of other, mostly animal proteins. The list of proteins currently known to contain one or more copies of an EGF-like pattern is large and varied. The functional significance of EGF domains in what appear to be unrelated proteins is not yet clear. However, a common feature is that these repeats are found in the extracellular domain of membrane-bound proteins or in proteins known to be secreted (exception: prostaglandin G/H synthase). The EGF domain includes six cysteine residues which have been shown (in EGF) to be involved in disulphide bonds. The main structure is a two-stranded beta-sheet followed by a loop to a C-terminal short two-stranded sheet. Subdomains between the conserved cysteines vary in length.; GO: 0005515 protein binding; PDB: 1WHE_A 1CCF_A 1APO_A 1WHF_A 2VJ3_A 1TOZ_A 4D90_B 3CFW_A 1EDM_B 1IXA_A ....
Probab=21.68 E-value=33 Score=21.46 Aligned_cols=20 Identities=30% Similarity=0.888 Sum_probs=15.1
Q ss_pred CCCCCceeCC-----CccCCCCCCce
Q 048192 260 VCGKYGICSQ-----GQCSCPATYFK 280 (422)
Q Consensus 260 ~CG~~giC~~-----~~C~C~~g~F~ 280 (422)
.|...|.|.. -.|.|++| |.
T Consensus 5 ~C~n~g~C~~~~~~~y~C~C~~G-~~ 29 (32)
T PF00008_consen 5 PCQNGGTCIDLPGGGYTCECPPG-YT 29 (32)
T ss_dssp SSTTTEEEEEESTSEEEEEEBTT-EE
T ss_pred cCCCCeEEEeCCCCCEEeECCCC-Cc
Confidence 6777888862 47999997 64
No 58
>smart00181 EGF Epidermal growth factor-like domain.
Probab=20.92 E-value=72 Score=19.64 Aligned_cols=21 Identities=29% Similarity=0.667 Sum_probs=14.3
Q ss_pred cCCCCCceeCC----CccCCCCCCcee
Q 048192 259 MVCGKYGICSQ----GQCSCPATYFKL 281 (422)
Q Consensus 259 g~CG~~giC~~----~~C~C~~g~F~~ 281 (422)
..|... .|.. ..|.|++| |+-
T Consensus 6 ~~C~~~-~C~~~~~~~~C~C~~g-~~g 30 (35)
T smart00181 6 GPCSNG-TCINTPGSYTCSCPPG-YTG 30 (35)
T ss_pred CCCCCC-EEECCCCCeEeECCCC-Ccc
Confidence 456666 7752 57999997 643
No 59
>PF14870 PSII_BNR: Photosynthesis system II assembly factor YCF48; PDB: 2XBG_A.
Probab=20.04 E-value=2.6e+02 Score=27.54 Aligned_cols=25 Identities=12% Similarity=0.303 Sum_probs=12.2
Q ss_pred CceeEEEEecCCCeeEEcCCCcEEEe
Q 048192 87 KSVVGLNLTDMGNLVLFDKNNAAVWQ 112 (422)
Q Consensus 87 ~~~~~~~LldsGNLVL~~~~~~~lWQ 112 (422)
+....|....+|+|.+.. .|..|..
T Consensus 187 ~riq~~gf~~~~~lw~~~-~Gg~~~~ 211 (302)
T PF14870_consen 187 RRIQSMGFSPDGNLWMLA-RGGQIQF 211 (302)
T ss_dssp S-EEEEEE-TTS-EEEEE-TTTEEEE
T ss_pred ceehhceecCCCCEEEEe-CCcEEEE
Confidence 345556666777777765 3334444
Done!