Query 016558
Match_columns 387
No_of_seqs 35 out of 37
Neff 2.8
Searched_HMMs 46136
Date Fri Mar 29 07:56:00 2013
Command hhsearch -i /work/01045/syshi/csienesis_hhblits_a3m/016558.a3m -d /work/01045/syshi/HHdatabase/Cdd.hhm -o /work/01045/syshi/hhsearch_cdd/016558hhsearch_cdd -cpu 12 -v 0
No Hit Prob E-value P-value Score SS Cols Query HMM Template HMM
1 PF12273 RCR: Chitin synthesis 90.0 0.047 1E-06 46.7 -1.7 23 291-313 5-27 (130)
2 PF05454 DAG1: Dystroglycan (D 89.4 0.11 2.3E-06 51.4 0.0 23 292-314 154-176 (290)
3 PHA03283 envelope glycoprotein 89.0 0.77 1.7E-05 49.1 5.8 70 288-360 405-482 (542)
4 PF07213 DAP10: DAP10 membrane 84.2 0.38 8.3E-06 40.1 0.5 28 292-319 44-72 (79)
5 PF12259 DUF3609: Protein of u 78.2 1.2 2.5E-05 45.2 1.6 27 288-314 303-329 (361)
6 PF13908 Shisa: Wnt and FGF in 75.6 1.6 3.4E-05 39.1 1.5 22 286-307 81-102 (179)
7 PF11614 FixG_C: IG-like fold 73.3 16 0.00035 30.2 6.9 47 197-243 33-83 (118)
8 PF02480 Herpes_gE: Alphaherpe 69.0 1.6 3.4E-05 45.2 0.0 21 294-314 364-384 (439)
9 PF06280 DUF1034: Fn3-like dom 68.1 20 0.00044 29.4 6.4 68 197-264 10-110 (112)
10 PF11359 gpUL132: Glycoprotein 67.0 2.3 5E-05 41.4 0.7 21 286-306 58-78 (235)
11 PF10633 NPCBM_assoc: NPCBM-as 65.7 29 0.00063 26.8 6.4 51 196-246 6-62 (78)
12 PF05506 DUF756: Domain of unk 62.9 73 0.0016 25.4 8.4 47 196-242 19-65 (89)
13 PF14283 DUF4366: Domain of un 61.4 8.1 0.00018 36.9 3.2 25 290-314 164-188 (218)
14 PF01299 Lamp: Lysosome-associ 60.3 4.5 9.7E-05 39.1 1.3 36 281-322 270-306 (306)
15 PF00974 Rhabdo_glycop: Rhabdo 60.2 2.9 6.3E-05 43.9 0.0 39 286-324 456-496 (501)
16 COG1470 Predicted membrane pro 58.8 24 0.00052 38.0 6.3 51 197-247 399-455 (513)
17 PF14874 PapD-like: Flagellar- 57.8 79 0.0017 25.0 7.8 48 196-243 21-72 (102)
18 PF15102 TMEM154: TMEM154 prot 57.2 10 0.00022 34.8 3.0 8 318-325 99-106 (146)
19 PF03896 TRAP_alpha: Transloco 53.8 1.1E+02 0.0023 30.6 9.6 20 197-216 101-120 (285)
20 KOG4818 Lysosomal-associated m 53.0 8.8 0.00019 39.6 2.0 37 280-322 325-362 (362)
21 TIGR00806 rfc RFC reduced fola 50.9 14 0.00031 39.5 3.3 38 286-323 425-464 (511)
22 TIGR02866 CoxB cytochrome c ox 50.6 9.3 0.0002 35.0 1.6 40 289-328 19-65 (201)
23 PF11770 GAPT: GRB2-binding ad 46.7 14 0.0003 34.5 2.1 20 286-305 13-33 (158)
24 PF06365 CD34_antigen: CD34/Po 44.7 9.9 0.00021 36.3 0.9 23 279-301 99-121 (202)
25 PF11669 WBP-1: WW domain-bind 44.3 3.5 7.7E-05 34.9 -1.9 10 300-309 35-44 (102)
26 PF02480 Herpes_gE: Alphaherpe 42.3 8.4 0.00018 40.1 0.0 41 286-327 353-394 (439)
27 PF07010 Endomucin: Endomucin; 41.0 6.1 0.00013 39.0 -1.1 38 273-311 180-217 (259)
28 PF09972 DUF2207: Predicted me 40.8 90 0.0019 30.7 6.7 17 207-223 130-146 (511)
29 PF07610 DUF1573: Protein of u 39.1 73 0.0016 23.0 4.4 42 201-242 2-45 (45)
30 PHA03282 envelope glycoprotein 38.8 42 0.00091 36.4 4.4 15 299-313 424-438 (540)
31 PHA03281 envelope glycoprotein 38.7 14 0.0003 40.4 1.0 45 288-332 562-609 (642)
32 PF07790 DUF1628: Protein of u 37.3 12 0.00027 29.3 0.3 43 283-325 3-45 (80)
33 PF05083 LST1: LST-1 protein; 37.3 8.1 0.00018 32.1 -0.8 41 290-332 4-53 (74)
34 PF07705 CARDB: CARDB; InterP 36.0 1.9E+02 0.0042 21.9 7.0 50 195-245 19-72 (101)
35 PF12768 Rax2: Cortical protei 35.8 15 0.00032 36.2 0.6 43 286-328 234-281 (281)
36 PF14610 DUF4448: Protein of u 35.8 18 0.00038 32.8 1.0 12 232-243 96-107 (189)
37 PF00635 Motile_Sperm: MSP (Ma 35.5 1.3E+02 0.0028 23.8 5.8 50 195-244 18-69 (109)
38 TIGR01433 CyoA cytochrome o ub 35.2 43 0.00094 31.9 3.5 37 288-324 37-78 (226)
39 KOG4222 Axon guidance receptor 34.2 1.6E+02 0.0034 35.3 8.2 115 269-385 855-980 (1281)
40 PF13908 Shisa: Wnt and FGF in 32.5 35 0.00076 30.5 2.4 41 281-322 80-120 (179)
41 PF04478 Mid2: Mid2 like cell 32.3 34 0.00074 31.8 2.3 16 297-312 65-80 (154)
42 PF05545 FixQ: Cbb3-type cytoc 32.1 7.4 0.00016 28.5 -1.6 28 279-306 3-30 (49)
43 PF12297 EVC2_like: Ellis van 31.5 10 0.00022 40.0 -1.4 26 286-311 65-90 (429)
44 PF06030 DUF916: Bacterial pro 31.4 83 0.0018 27.3 4.4 43 189-231 21-63 (121)
45 PF10989 DUF2808: Protein of u 30.0 1.2E+02 0.0027 26.6 5.3 24 200-223 31-54 (146)
46 PF15102 TMEM154: TMEM154 prot 29.7 50 0.0011 30.5 2.8 28 302-330 76-103 (146)
47 KOG4764 Uncharacterized conser 28.8 22 0.00047 29.4 0.4 9 340-348 40-48 (70)
48 PF15065 NCU-G1: Lysosomal tra 26.7 34 0.00074 35.0 1.4 26 286-311 322-348 (350)
49 TIGR02745 ccoG_rdxA_fixG cytoc 26.0 2.4E+02 0.0052 29.6 7.3 49 196-244 347-399 (434)
50 PTZ00364 dipeptidyl-peptidase 25.9 1.4E+02 0.003 32.4 5.8 16 248-263 437-452 (548)
51 PF07204 Orthoreo_P10: Orthore 25.8 13 0.00028 32.3 -1.5 69 250-322 12-81 (98)
52 PF14316 DUF4381: Domain of un 25.5 8.8 0.00019 33.5 -2.6 32 287-320 20-51 (146)
53 PF13980 UPF0370: Uncharacteri 25.0 14 0.00031 29.8 -1.2 54 289-351 6-59 (63)
54 PF06682 DUF1183: Protein of u 24.9 1.3E+02 0.0027 30.8 5.0 21 339-359 198-218 (318)
55 smart00557 IG_FLMN Filamin-typ 23.1 3.9E+02 0.0084 21.3 6.6 44 198-244 21-64 (93)
56 PHA03286 envelope glycoprotein 22.4 29 0.00062 37.3 -0.0 39 293-331 400-441 (492)
57 PHA03291 envelope glycoprotein 22.0 31 0.00066 36.1 0.1 16 296-311 299-314 (401)
58 TIGR01732 tiny_TM_bacill conse 21.0 72 0.0016 22.1 1.6 14 290-303 12-25 (26)
59 PF07889 DUF1664: Protein of u 20.7 51 0.0011 29.5 1.2 27 288-314 5-31 (126)
60 PF12273 RCR: Chitin synthesis 20.7 30 0.00066 29.7 -0.2 25 288-313 6-30 (130)
61 TIGR02537 arch_flag_Nterm arch 20.6 65 0.0014 21.9 1.4 20 283-302 4-23 (26)
62 PF11980 DUF3481: Domain of un 20.6 41 0.0009 28.8 0.6 46 273-319 9-56 (87)
63 PF15234 LAT: Linker for activ 20.0 34 0.00074 33.3 -0.1 22 286-307 10-31 (230)
No 1
>PF12273 RCR: Chitin synthesis regulation, resistance to Congo red; InterPro: IPR020999 RCR proteins are ER membrane proteins that regulate chitin deposition in fungal cell walls. Although chitin, a linear polymer of beta-1,4-linked N-acetylglucosamine, constitutes only 2% of the cell wall it plays a vital role in the overall protection of the cell wall against stress, noxious chemicals and osmotic pressure changes. Congo red is a cell wall-disrupting benzidine-type dye extensively used in many cell wall mutant studies that specifically targets chitin in yeast cells and inhibits growth. RCR proteins render the yeasts resistant to Congo red by diminishing the content of chitin in the cell wall []. RCR proteins are probably regulating chitin synthase III interact directly with ubiquitin ligase Rsp5, and the VPEY motif is necessary for this, via interaction with the WW domains of Rsp5 [].
Probab=90.04 E-value=0.047 Score=46.69 Aligned_cols=23 Identities=26% Similarity=0.309 Sum_probs=9.2
Q ss_pred HHHHHHHhhhceeEEEeeccccc
Q 016558 291 FLILSVLIFGVTWACCKCRKRRW 313 (387)
Q Consensus 291 fLv~tvVliGgvwaCCkfRkrr~ 313 (387)
|+||+++||..+.+||+++|||+
T Consensus 5 ~~iii~~i~l~~~~~~~~~rRR~ 27 (130)
T PF12273_consen 5 FAIIIVAILLFLFLFYCHNRRRR 27 (130)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHh
Confidence 33433333333334444444443
No 2
>PF05454 DAG1: Dystroglycan (Dystrophin-associated glycoprotein 1); InterPro: IPR008465 Dystroglycan is one of the dystrophin-associated glycoproteins, which is encoded by a 5.5 kb transcript in Homo sapiens. The protein product is cleaved into two non-covalently associated subunits, [alpha] (N-terminal) and [beta] (C-terminal). In skeletal muscle the dystroglycan complex works as a transmembrane linkage between the extracellular matrix and the cytoskeleton [alpha]-dystroglycan is extracellular and binds to merosin ([alpha]-2 laminin) in the basement membrane, while [beta]-dystroglycan is a transmembrane protein and binds to dystrophin, which is a large rod-like cytoskeletal protein, absent in Duchenne muscular dystrophy patients. Dystrophin binds to intracellular actin cables. In this way, the dystroglycan complex, which links the extracellular matrix to the intracellular actin cables, is thought to provide structural integrity in muscle tissues. The dystroglycan complex is also known to serve as an agrin receptor in muscle, where it may regulate agrin-induced acetylcholine receptor clustering at the neuromuscular junction. There is also evidence which suggests the function of dystroglycan as a part of the signal transduction pathway because it is shown that Grb2, a mediator of the Ras-related signal pathway, can interact with the cytoplasmic domain of dystroglycan. In general, aberrant expression of dystrophin-associated protein complex underlies the pathogenesis of Duchenne muscular dystrophy, Becker muscular dystrophy and severe childhood autosomal recessive muscular dystrophy. Interestingly, no genetic disease has been described for either [alpha]- or [beta]-dystroglycan. Dystroglycan is widely distributed in non-muscle tissues as well as in muscle tissues. During epithelial morphogenesis of kidney, the dystroglycan complex is shown to act as a receptor for the basement membrane. Dystroglycan expression in Mus musculus brain and neural retina has also been reported. However, the physiological role of dystroglycan in non-muscle tissues has remained unclear [].; PDB: 1EG4_P.
Probab=89.45 E-value=0.11 Score=51.42 Aligned_cols=23 Identities=26% Similarity=0.474 Sum_probs=0.0
Q ss_pred HHHHHHhhhceeEEEeecccccC
Q 016558 292 LILSVLIFGVTWACCKCRKRRWN 314 (387)
Q Consensus 292 Lv~tvVliGgvwaCCkfRkrr~~ 314 (387)
+|+++|||+++.|||.+||||.-
T Consensus 154 VI~~iLLIA~iIa~icyrrkR~G 176 (290)
T PF05454_consen 154 VIAAILLIAGIIACICYRRKRKG 176 (290)
T ss_dssp -----------------------
T ss_pred HHHHHHHHHHHHHHHhhhhhhcc
Confidence 45556667777788888877653
No 3
>PHA03283 envelope glycoprotein E; Provisional
Probab=88.97 E-value=0.77 Score=49.08 Aligned_cols=70 Identities=24% Similarity=0.405 Sum_probs=38.4
Q ss_pred hhHHHHHHHHhhhceeEEEeecccccCCCCCceeeec------CCCCccCCccc-c-cCCCcCCCCCCCCCcccCCCCCC
Q 016558 288 GAYFLILSVLIFGVTWACCKCRKRRWNDGVPYQELEM------GLPESVSAMNV-E-TAEGWDEGWDDDWDENNAVKSPG 359 (387)
Q Consensus 288 GAYfLv~tvVliGgvwaCCkfRkrr~~~GvpYQELEM------~LP~S~ga~ev-E-taDGWDdgWDDDWDDEEApKSPs 359 (387)
++--++.++|+..++|+|+.||++++. +|.=|-= .||.-..-..+ | -+.-=||..|+|=|||-++.+|.
T Consensus 405 ~~~~~~~~~~~~l~vw~c~~~r~~~~~---~y~ilnpf~~vytslptn~~~~~~f~~~~~~~ddsf~~~~de~~~~~~~~ 481 (542)
T PHA03283 405 AIICTCAALLVALVVWGCILYRRSNRK---PYEVLNPFETVYTSVPSNDPEVLVFERLASDSDDSFDSSSDEELEPPPPP 481 (542)
T ss_pred HHHHHHHHHHHHHhhhheeeehhhcCC---cccccCCCccceeccCCCCCcccceeecccCccccccccccccccCCCCC
Confidence 333344556677789999998777665 4443332 24433332111 1 12223567777766666665555
Q ss_pred C
Q 016558 360 A 360 (387)
Q Consensus 360 ~ 360 (387)
.
T Consensus 482 ~ 482 (542)
T PHA03283 482 G 482 (542)
T ss_pred C
Confidence 3
No 4
>PF07213 DAP10: DAP10 membrane protein; InterPro: IPR009861 This family consists of several mammalian DAP10 membrane proteins. In activated mouse natural killer (NK) cells, the NKG2D receptor associates with two intracellular adaptors, DAP10 and DAP12, which trigger phosphatidyl inositol 3 kinase (PI3K) and Syk family protein tyrosine kinases, respectively. It has been suggested that the DAP10-PI3K pathway is sufficient to initiate NKG2D-mediated killing of target cells [].
Probab=84.19 E-value=0.38 Score=40.05 Aligned_cols=28 Identities=32% Similarity=0.564 Sum_probs=23.3
Q ss_pred HHHHHHhhhceeEEEeecccccC-CCCCc
Q 016558 292 LILSVLIFGVTWACCKCRKRRWN-DGVPY 319 (387)
Q Consensus 292 Lv~tvVliGgvwaCCkfRkrr~~-~GvpY 319 (387)
+++|+||+++++.|-++|||++| ++--|
T Consensus 44 ~vlTLLIv~~vy~car~r~r~~~~~~kvY 72 (79)
T PF07213_consen 44 AVLTLLIVLVVYYCARPRRRPTQEDDKVY 72 (79)
T ss_pred HHHHHHHHHHHHhhcccccCCcccCCEEE
Confidence 57999999999999999999888 54333
No 5
>PF12259 DUF3609: Protein of unknown function (DUF3609); InterPro: IPR022048 This domain family is found in eukaryotes and viruses, and is typically between 348 and 360 amino acids in length.
Probab=78.21 E-value=1.2 Score=45.20 Aligned_cols=27 Identities=15% Similarity=0.412 Sum_probs=21.3
Q ss_pred hhHHHHHHHHhhhceeEEEeecccccC
Q 016558 288 GAYFLILSVLIFGVTWACCKCRKRRWN 314 (387)
Q Consensus 288 GAYfLv~tvVliGgvwaCCkfRkrr~~ 314 (387)
-++.+++++|+++++|.|++||||+.+
T Consensus 303 v~~~~vli~vl~~~~~~~~~~~~~~~~ 329 (361)
T PF12259_consen 303 VCGAIVLIIVLISLAWLYRTFRRRQLR 329 (361)
T ss_pred hhHHHHHHHHHHHHHhheeehHHHHhh
Confidence 344566667889999999999998765
No 6
>PF13908 Shisa: Wnt and FGF inhibitory regulator
Probab=75.64 E-value=1.6 Score=39.05 Aligned_cols=22 Identities=23% Similarity=0.641 Sum_probs=10.5
Q ss_pred chhhHHHHHHHHhhhceeEEEe
Q 016558 286 INGAYFLILSVLIFGVTWACCK 307 (387)
Q Consensus 286 I~GAYfLv~tvVliGgvwaCCk 307 (387)
|.|+.++|++||++-+++.||+
T Consensus 81 ivgvi~~Vi~Iv~~Iv~~~Cc~ 102 (179)
T PF13908_consen 81 IVGVICGVIAIVVLIVCFCCCC 102 (179)
T ss_pred eeehhhHHHHHHHhHhhheecc
Confidence 3344444444444445555544
No 7
>PF11614 FixG_C: IG-like fold at C-terminal of FixG, putative oxidoreductase; PDB: 2R39_A.
Probab=73.32 E-value=16 Score=30.23 Aligned_cols=47 Identities=15% Similarity=0.315 Sum_probs=31.9
Q ss_pred cceEEEEEcCCCceEEEEEEcC--ccccC--CCceeeecccceeEEEEEEe
Q 016558 197 GELTILVQNEGEKTLIVTITIP--TAVEN--PLKQLKISKHQTQKINISLS 243 (387)
Q Consensus 197 ~~lsLLVQNkG~~~L~V~ItaP--d~V~~--~~~~L~L~K~qskKV~IS~s 243 (387)
-.|.|-+.|+.+.+..+.|++. ..+.+ ....|+|..++..++.|.+.
T Consensus 33 N~Y~lkl~Nkt~~~~~~~i~~~g~~~~~l~~~~~~i~v~~g~~~~~~v~v~ 83 (118)
T PF11614_consen 33 NQYTLKLTNKTNQPRTYTISVEGLPGAELQGPENTITVPPGETREVPVFVT 83 (118)
T ss_dssp EEEEEEEEE-SSS-EEEEEEEES-SS-EE-ES--EEEE-TT-EEEEEEEEE
T ss_pred EEEEEEEEECCCCCEEEEEEEecCCCeEEECCCcceEECCCCEEEEEEEEE
Confidence 3688999999999888888655 44444 44788898899988888887
No 8
>PF02480 Herpes_gE: Alphaherpesvirus glycoprotein E; InterPro: IPR003404 Glycoprotein E (gE) of Alphaherpesvirus forms a complex with glycoprotein I (gI), functioning as an immunoglobulin G (IgG) Fc binding protein. gE is involved in virus spread but is not essential for propagation [].; GO: 0016020 membrane; PDB: 2GJ7_F 2GIY_B.
Probab=68.96 E-value=1.6 Score=45.24 Aligned_cols=21 Identities=38% Similarity=0.944 Sum_probs=0.0
Q ss_pred HHHHhhhceeEEEeecccccC
Q 016558 294 LSVLIFGVTWACCKCRKRRWN 314 (387)
Q Consensus 294 ~tvVliGgvwaCCkfRkrr~~ 314 (387)
+++||+.++|+|+++||||++
T Consensus 364 livVv~viv~vc~~~rrrR~~ 384 (439)
T PF02480_consen 364 LIVVVGVIVWVCLRCRRRRRQ 384 (439)
T ss_dssp ---------------------
T ss_pred HHHHHHHHhheeeeehhcccc
Confidence 334444555555555555554
No 9
>PF06280 DUF1034: Fn3-like domain (DUF1034); InterPro: IPR010435 In the MEROPS database peptidases and peptidase homologues are grouped into clans and families. Clans are groups of families for which there is evidence of common ancestry based on a common structural fold: Each clan is identified with two letters, the first representing the catalytic type of the families included in the clan (with the letter 'P' being used for a clan containing families of more than one of the catalytic types serine, threonine and cysteine). Some families cannot yet be assigned to clans, and when a formal assignment is required, such a family is described as belonging to clan A-, C-, M-, N-, S-, T- or U-, according to the catalytic type. Some clans are divided into subclans because there is evidence of a very ancient divergence within the clan, for example MA(E), the gluzincins, and MA(M), the metzincins. Peptidase families are grouped by their catalytic type, the first character representing the catalytic type: A, aspartic; C, cysteine; G, glutamic acid; M, metallo; N, asparagine; S, serine; T, threonine; and U, unknown. The serine, threonine and cysteine peptidases utilise the amino acid as a nucleophile and form an acyl intermediate - these peptidases can also readily act as transferases. In the case of aspartic, glutamic and metallopeptidases, the nucleophile is an activated water molecule. In the case of the asparagine endopeptidases, the nucleophile is asparagine and all are self-processing endopeptidases. In many instances the structural protein fold that characterises the clan or family may have lost its catalytic activity, yet retain its function in protein recognition and binding. Proteolytic enzymes that exploit serine in their catalytic activity are ubiquitous, being found in viruses, bacteria and eukaryotes []. They include a wide range of peptidase activity, including exopeptidase, endopeptidase, oligopeptidase and omega-peptidase activity. Over 20 families (denoted S1 - S66) of serine protease have been identified, these being grouped into clans on the basis of structural similarity and other functional evidence []. Structures are known for members of the clans and the structures indicate that some appear to be totally unrelated, suggesting different evolutionary origins for the serine peptidases []. Not withstanding their different evolutionary origins, there are similarities in the reaction mechanisms of several peptidases. Chymotrypsin, subtilisin and carboxypeptidase C have a catalytic triad of serine, aspartate and histidine in common: serine acts as a nucleophile, aspartate as an electrophile, and histidine as a base []. The geometric orientations of the catalytic residues are similar between families, despite different protein folds []. The linear arrangements of the catalytic residues commonly reflect clan relationships. For example the catalytic triad in the chymotrypsin clan (PA) is ordered HDS, but is ordered DHS in the subtilisin clan (SB) and SDH in the carboxypeptidase clan (SC) [, ]. This domain of unknown function is present in bacterial and plant peptidases belonging to MEROPS peptidase family S8 (subfamily S8A subtilisin, clan SB). It is C-terminal to and adjacent to the S8 peptidase domain and can be found in conjunction with the PA (Protease associated) domain (IPR003137 from INTERPRO) and additionally in Gram-positive bacteria with the surface protein anchor domain (IPR001899 from INTERPRO).; GO: 0004252 serine-type endopeptidase activity, 0005618 cell wall, 0016020 membrane; PDB: 3EIF_A 1XF1_B.
Probab=68.06 E-value=20 Score=29.42 Aligned_cols=68 Identities=16% Similarity=0.291 Sum_probs=39.3
Q ss_pred cceEEEEEcCCCceEEEEEEcC----c-------------------cccCCCceeeecccceeEEEEEEecCC-------
Q 016558 197 GELTILVQNEGEKTLIVTITIP----T-------------------AVENPLKQLKISKHQTQKINISLSARK------- 246 (387)
Q Consensus 197 ~~lsLLVQNkG~~~L~V~ItaP----d-------------------~V~~~~~~L~L~K~qskKV~IS~s~~~------- 246 (387)
..+.|.++|.|..++..+|..- + .+......|.|.-++++.|+|++..+.
T Consensus 10 ~~~~itl~N~~~~~~ty~~~~~~~~t~~~~~~~~~~~~~~~~~~~~~~~~~~~~vTV~ag~s~~v~vti~~p~~~~~~~~ 89 (112)
T PF06280_consen 10 FSFTITLHNYGDKPVTYTLSHVPVLTDKTDTEEGYSILVPPVPSISTVSFSPDTVTVPAGQSKTVTVTITPPSGLDASNG 89 (112)
T ss_dssp EEEEEEEEE-SSS-EEEEEEEE-EEEEEE--ETTEEEEEEEE----EEE---EEEEE-TTEEEEEEEEEE--GGGHHTT-
T ss_pred eEEEEEEEECCCCCEEEEEeeEEEEeeEeeccCCcccccccccceeeEEeCCCeEEECCCCEEEEEEEEEehhcCCcccC
Confidence 5667777777777666555221 0 234455677888888888888888522
Q ss_pred ---CceEEEEeccCceEEecC
Q 016558 247 ---NSKLVLNAGNGECVLHMG 264 (387)
Q Consensus 247 ---s~~IvL~aGkG~C~Lhi~ 264 (387)
++-|.|+...+.+.|+|+
T Consensus 90 ~~~eG~I~~~~~~~~~~lsIP 110 (112)
T PF06280_consen 90 PFYEGFITFKSSDGEPDLSIP 110 (112)
T ss_dssp EEEEEEEEEESSTTSEEEEEE
T ss_pred CEEEEEEEEEcCCCCEEEEee
Confidence 134777777776666653
No 10
>PF11359 gpUL132: Glycoprotein UL132; InterPro: IPR021023 Glycoprotein UL132 is a low-abundance structural component of Human herpesvirus 5 []. The function of this protein is not fully understood.
Probab=67.00 E-value=2.3 Score=41.39 Aligned_cols=21 Identities=24% Similarity=0.528 Sum_probs=16.4
Q ss_pred chhhHHHHHHHHhhhceeEEE
Q 016558 286 INGAYFLILSVLIFGVTWACC 306 (387)
Q Consensus 286 I~GAYfLv~tvVliGgvwaCC 306 (387)
+.|..+|-|.+|++++...-|
T Consensus 58 VTg~sllsli~VtvaalYsSC 78 (235)
T PF11359_consen 58 VTGFSLLSLIVVTVAALYSSC 78 (235)
T ss_pred ehhHHHHHHHHHHHHHHHHHH
Confidence 558888888888888877655
No 11
>PF10633 NPCBM_assoc: NPCBM-associated, NEW3 domain of alpha-galactosidase; InterPro: IPR018905 This domain has been named NEW3, but its function is not known. It is found on proteins which are bacterial galactosidases [].; PDB: 1EUT_A 2BZD_A 1WCQ_C 2BER_A 1W8O_A 1EUU_A 1W8N_A.
Probab=65.67 E-value=29 Score=26.82 Aligned_cols=51 Identities=10% Similarity=0.289 Sum_probs=31.2
Q ss_pred CcceEEEEEcCCCc---eEEEEEEcCcccc--CCCcee-eecccceeEEEEEEecCC
Q 016558 196 SGELTILVQNEGEK---TLIVTITIPTAVE--NPLKQL-KISKHQTQKINISLSARK 246 (387)
Q Consensus 196 S~~lsLLVQNkG~~---~L~V~ItaPd~V~--~~~~~L-~L~K~qskKV~IS~s~~~ 246 (387)
...+.|-|.|.|.. .+.|.+..|+.+. ..+..+ .|.-+++..+.+.++.+.
T Consensus 6 ~~~~~~tv~N~g~~~~~~v~~~l~~P~GW~~~~~~~~~~~l~pG~s~~~~~~V~vp~ 62 (78)
T PF10633_consen 6 TVTVTLTVTNTGTAPLTNVSLSLSLPEGWTVSASPASVPSLPPGESVTVTFTVTVPA 62 (78)
T ss_dssp EEEEEEEEE--SSS-BSS-EEEEE--TTSE---EEEEE--B-TTSEEEEEEEEEE-T
T ss_pred EEEEEEEEEECCCCceeeEEEEEeCCCCccccCCccccccCCCCCEEEEEEEEECCC
Confidence 45688999999975 4788889998887 333333 567888888888888443
No 12
>PF05506 DUF756: Domain of unknown function (DUF756); InterPro: IPR008475 This domain is found, normally as a tandem repeat, at the C terminus of bacterial phospholipase C proteins.; GO: 0004629 phospholipase C activity, 0016042 lipid catabolic process
Probab=62.93 E-value=73 Score=25.43 Aligned_cols=47 Identities=17% Similarity=0.189 Sum_probs=33.6
Q ss_pred CcceEEEEEcCCCceEEEEEEcCccccCCCceeeecccceeEEEEEE
Q 016558 196 SGELTILVQNEGEKTLIVTITIPTAVENPLKQLKISKHQTQKINISL 242 (387)
Q Consensus 196 S~~lsLLVQNkG~~~L~V~ItaPd~V~~~~~~L~L~K~qskKV~IS~ 242 (387)
...+.|.+.|.|...+.|+|..-.+-...+..+.|.-+++..+.+..
T Consensus 19 ~g~l~l~l~N~g~~~~~~~v~~~~y~~~~~~~~~v~ag~~~~~~w~l 65 (89)
T PF05506_consen 19 TGNLRLTLSNPGSAAVTFTVYDNAYGGGGPWTYTVAAGQTVSLTWPL 65 (89)
T ss_pred CCEEEEEEEeCCCCcEEEEEEeCCcCCCCCEEEEECCCCEEEEEEee
Confidence 34899999999999999999875454344556666666665555543
No 13
>PF14283 DUF4366: Domain of unknown function (DUF4366)
Probab=61.45 E-value=8.1 Score=36.91 Aligned_cols=25 Identities=28% Similarity=0.308 Sum_probs=16.0
Q ss_pred HHHHHHHHhhhceeEEEeecccccC
Q 016558 290 YFLILSVLIFGVTWACCKCRKRRWN 314 (387)
Q Consensus 290 YfLv~tvVliGgvwaCCkfRkrr~~ 314 (387)
.+|++++|+.||++++.||+|.+++
T Consensus 164 l~lllv~l~gGGa~yYfK~~K~K~~ 188 (218)
T PF14283_consen 164 LLLLLVALIGGGAYYYFKFYKPKQE 188 (218)
T ss_pred HHHHHHHHhhcceEEEEEEeccccc
Confidence 3344455556667777778887766
No 14
>PF01299 Lamp: Lysosome-associated membrane glycoprotein (Lamp); InterPro: IPR002000 Lysosome-associated membrane glycoproteins (lamp) [] are integral membrane proteins, specific to lysosomes, and whose exact biological function is not yet clear. Structurally, the lamp proteins consist of two internally homologous lysosome-luminal domains separated by a proline-rich hinge region; at the C-terminal extremity there is a transmembrane region (TM) followed by a very short cytoplasmic tail (C). In each of the duplicated domains, there are two conserved disulphide bonds. This structure is schematically represented in the figure below. +-----+ +-----+ +-----+ +-----+ | | | | | | | | xCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxxxCxxxxxCxxxxxxxxxxxxCxxxxxCxxxxxxxx +--------------------------++Hinge++--------------------------++TM++C+ In mammals, there are two closely related types of lamp: lamp-1 and lamp-2, which form major components of the lysosome membrane. In chicken lamp-1 is known as LEP100. Also included in this entry is the macrophage protein CD68 (or macrosialin) [] is a heavily glycosylated integral membrane protein whose structure consists of a mucin-like domain followed by a proline-rich hinge; a single lamp-like domain; a transmembrane region and a short cytoplasmic tail. Similar to CD68, mammalian lamp-3, which is expressed in lymphoid organs, dendritic cells and in lung, contains all the C-terminal regions but lacks the N-terminal lamp-like region []. In a lamp-family protein from nematodes [] only the part C-terminal to the hinge is conserved. ; GO: 0016020 membrane
Probab=60.33 E-value=4.5 Score=39.15 Aligned_cols=36 Identities=31% Similarity=0.359 Sum_probs=18.1
Q ss_pred eecccc-hhhHHHHHHHHhhhceeEEEeecccccCCCCCceee
Q 016558 281 KILTPI-NGAYFLILSVLIFGVTWACCKCRKRRWNDGVPYQEL 322 (387)
Q Consensus 281 ~iltPI-~GAYfLv~tvVliGgvwaCCkfRkrr~~~GvpYQEL 322 (387)
.++-|| .|+-+.+++||+ .-|||..|||++. -||.+
T Consensus 270 ~~~vPIaVG~~La~lvliv---LiaYli~Rrr~~~---gYq~~ 306 (306)
T PF01299_consen 270 SDLVPIAVGAALAGLVLIV---LIAYLIGRRRSRA---GYQSI 306 (306)
T ss_pred cchHHHHHHHHHHHHHHHH---HHhheeEeccccc---ccccC
Confidence 456776 566543332222 2245545554444 58864
No 15
>PF00974 Rhabdo_glycop: Rhabdovirus spike glycoprotein; InterPro: IPR001903 Different families of ssRNA negative-strand viruses contain glycoproteins responsible for forming spikes on the surface of the virion. The glycoprotein spike is made up of a trimer of glycoproteins. These proteins are frequently abbreviated to G protein. Channel formed by glycoprotein spike is thought to function in a similar manner to Influenza virus M2 protein channel, thus allowing a signal to pass across the viral membrane to signal for viral uncoating [, ].; GO: 0019031 viral envelope; PDB: 2CMZ_C 2J6J_A 3EGD_D.
Probab=60.24 E-value=2.9 Score=43.88 Aligned_cols=39 Identities=28% Similarity=0.477 Sum_probs=0.0
Q ss_pred chhhHHHHHHHHhhhceeEEEeeccccc-C-CCCCceeeec
Q 016558 286 INGAYFLILSVLIFGVTWACCKCRKRRW-N-DGVPYQELEM 324 (387)
Q Consensus 286 I~GAYfLv~tvVliGgvwaCCkfRkrr~-~-~GvpYQELEM 324 (387)
..+++.+++.+|||.++..||+|||+++ + .-.-|...+|
T Consensus 456 ~~~~~~vi~~illi~l~~cc~~~~r~~~~~~~~~i~~~~~~ 496 (501)
T PF00974_consen 456 SIIAIAVILLILLILLIRCCCRCRRRRRPKRKRGIYESKVS 496 (501)
T ss_dssp -----------------------------------------
T ss_pred HHHHHHHHHHHHHHHHHHHhhhhccccccccCCcccccccc
Confidence 3355555555666655545555664433 2 3355666666
No 16
>COG1470 Predicted membrane protein [Function unknown]
Probab=58.77 E-value=24 Score=38.01 Aligned_cols=51 Identities=14% Similarity=0.350 Sum_probs=33.5
Q ss_pred cceEEEEEcCCCc---eEEEEEEcCccccCCCceeee---cccceeEEEEEEecCCC
Q 016558 197 GELTILVQNEGEK---TLIVTITIPTAVENPLKQLKI---SKHQTQKINISLSARKN 247 (387)
Q Consensus 197 ~~lsLLVQNkG~~---~L~V~ItaPd~V~~~~~~L~L---~K~qskKV~IS~s~~~s 247 (387)
...-+-|-|.|.- .++++|..|..++..-.+-++ .-+..+.|+++++.+..
T Consensus 399 ~~i~i~I~NsGna~LtdIkl~v~~PqgWei~Vd~~~I~sL~pge~~tV~ltI~vP~~ 455 (513)
T COG1470 399 KTIRISIENSGNAPLTDIKLTVNGPQGWEIEVDESTIPSLEPGESKTVSLTITVPED 455 (513)
T ss_pred ceEEEEEEecCCCccceeeEEecCCccceEEECcccccccCCCCcceEEEEEEcCCC
Confidence 4667778888865 456788888666554444333 45667788888775443
No 17
>PF14874 PapD-like: Flagellar-associated PapD-like
Probab=57.77 E-value=79 Score=25.03 Aligned_cols=48 Identities=8% Similarity=0.078 Sum_probs=36.8
Q ss_pred CcceEEEEEcCCCceEEEEEEcCc----cccCCCceeeecccceeEEEEEEe
Q 016558 196 SGELTILVQNEGEKTLIVTITIPT----AVENPLKQLKISKHQTQKINISLS 243 (387)
Q Consensus 196 S~~lsLLVQNkG~~~L~V~ItaPd----~V~~~~~~L~L~K~qskKV~IS~s 243 (387)
.+...|.+.|.|..++.+.|..|. .+...+..=.|.-+.+..|+|.+.
T Consensus 21 ~~~~~v~l~N~s~~p~~f~v~~~~~~~~~~~v~~~~g~l~PG~~~~~~V~~~ 72 (102)
T PF14874_consen 21 TYSRTVTLTNTSSIPARFRVRQPESLSSFFSVEPPSGFLAPGESVELEVTFS 72 (102)
T ss_pred EEEEEEEEEECCCCCEEEEEEeCCcCCCCEEEECCCCEECCCCEEEEEEEEE
Confidence 456899999999999999997774 344444444667788888888888
No 18
>PF15102 TMEM154: TMEM154 protein family
Probab=57.15 E-value=10 Score=34.79 Aligned_cols=8 Identities=38% Similarity=0.663 Sum_probs=5.0
Q ss_pred CceeeecC
Q 016558 318 PYQELEMG 325 (387)
Q Consensus 318 pYQELEM~ 325 (387)
.||..|++
T Consensus 99 ~~qt~e~~ 106 (146)
T PF15102_consen 99 ALQTYELG 106 (146)
T ss_pred cccccccC
Confidence 66666663
No 19
>PF03896 TRAP_alpha: Translocon-associated protein (TRAP), alpha subunit; InterPro: IPR005595 The alpha-subunit of the TRAP complex (TRAP alpha) is a single-spanning membrane protein of the endoplasmic reticulum (ER) which is found in proximity of nascent polypeptide chains translocating across the membrane [].; GO: 0005783 endoplasmic reticulum
Probab=53.79 E-value=1.1e+02 Score=30.65 Aligned_cols=20 Identities=15% Similarity=0.215 Sum_probs=15.9
Q ss_pred cceEEEEEcCCCceEEEEEE
Q 016558 197 GELTILVQNEGEKTLIVTIT 216 (387)
Q Consensus 197 ~~lsLLVQNkG~~~L~V~It 216 (387)
....|=+.|+|..++.|...
T Consensus 101 ~~~LvgftN~g~~~~~V~~i 120 (285)
T PF03896_consen 101 VKFLVGFTNKGSEPFTVESI 120 (285)
T ss_pred EEEEEEEEeCCCCCEEEEEE
Confidence 46677789999999988763
No 20
>KOG4818 consensus Lysosomal-associated membrane protein [General function prediction only]
Probab=52.99 E-value=8.8 Score=39.65 Aligned_cols=37 Identities=24% Similarity=0.321 Sum_probs=22.0
Q ss_pred eeeccc-chhhHHHHHHHHhhhceeEEEeecccccCCCCCceee
Q 016558 280 DKILTP-INGAYFLILSVLIFGVTWACCKCRKRRWNDGVPYQEL 322 (387)
Q Consensus 280 ~~iltP-I~GAYfLv~tvVliGgvwaCCkfRkrr~~~GvpYQEL 322 (387)
..++.| |.|+-+..+.++|+.+ .||. ||||++ -||.|
T Consensus 325 ~siv~PivVg~~l~gl~~~vlia--ylIg-rr~~~~---gYq~i 362 (362)
T KOG4818|consen 325 LNIVLPIAVGAILAGLVLVVLIA--YLIG-RRRSHS---GYQTI 362 (362)
T ss_pred cceecchHHHHHHHHHHHHHHHH--hhee-heeccc---ccccC
Confidence 457788 6677766655555544 3554 555554 38764
No 21
>TIGR00806 rfc RFC reduced folate carrier. Proteins of the RFC family are so-far restricted to animals. RFC proteins possess 12 putative transmembrane a-helical spanners (TMSs) and evidence for a 12 TMS topology has been published for the human RFC. The RFC transporters appear to transport reduced folate by an energy-dependent, pH-dependent, Na+-independent mechanism. Folate:H+ symport, folate:OH- antiport and folate:anion antiport mechanisms have been proposed, but the energetic mechanism is not well defined.
Probab=50.93 E-value=14 Score=39.51 Aligned_cols=38 Identities=34% Similarity=0.502 Sum_probs=24.5
Q ss_pred chhhHHHHHHHHh-hhceeEEEeecc-cccCCCCCceeee
Q 016558 286 INGAYFLILSVLI-FGVTWACCKCRK-RRWNDGVPYQELE 323 (387)
Q Consensus 286 I~GAYfLv~tvVl-iGgvwaCCkfRk-rr~~~GvpYQELE 323 (387)
+||.||++++++. +++++.|+++-+ .|++.-.+=|++.
T Consensus 425 vY~~yf~~~~~i~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 464 (511)
T TIGR00806 425 IYSVYFLVLSIICFFGAGLDGLRYCKRGTHQPLAPAQELR 464 (511)
T ss_pred ehhhHHHHHHHHHHHHHHHHHhhhhcccccCCCCcccccc
Confidence 6789999877655 444677777443 3444455666665
No 22
>TIGR02866 CoxB cytochrome c oxidase, subunit II. Cytochrome c oxidase is the terminal electron acceptor of mitochondria (and one of several possible acceptors in prokaryotes) in the electron transport chain of aerobic respiration. The enzyme couples the oxidation of reduced cytochrome c with the reduction of molecular oxygen to water. This process results in the pumping of four protons across the membrane which are used in the proton gradient powered synthesis of ATP. The oxidase contains two heme a cofactors and three copper atoms as well as other bound ions.
Probab=50.57 E-value=9.3 Score=35.00 Aligned_cols=40 Identities=15% Similarity=0.122 Sum_probs=23.6
Q ss_pred hHHHHHHHHhhhceeEEEeecccccCCCCCc----eeeec---CCCC
Q 016558 289 AYFLILSVLIFGVTWACCKCRKRRWNDGVPY----QELEM---GLPE 328 (387)
Q Consensus 289 AYfLv~tvVliGgvwaCCkfRkrr~~~GvpY----QELEM---~LP~ 328 (387)
+-++|+++|....+|++++||+++++.-.+| +.||+ .+|.
T Consensus 19 i~~iI~v~V~~~l~~~~~k~r~~~~~~~~~~~~~~~~lEi~wtiiP~ 65 (201)
T TIGR02866 19 VATTISLLVAALLAYVVWKFRRKGDEEKPSKIHGNRALEYTWTVIPL 65 (201)
T ss_pred HHHHHHHHHHHHHHHhhhhhhcccccCCCccccCCceEEEEeehHhH
Confidence 3445556666677788888887533212233 56887 3664
No 23
>PF11770 GAPT: GRB2-binding adapter (GAPT); InterPro: IPR021082 This entry represents a family of transmembrane proteins which bind the growth factor receptor-bound protein 2 (GRB2) in B cells []. In contrast to other transmembrane adaptor proteins, GAPT, which this entry represents, is not phosphorylated upon BCR ligation. It associates with GRB2 constitutively through its proline-rich region [].
Probab=46.71 E-value=14 Score=34.49 Aligned_cols=20 Identities=30% Similarity=0.579 Sum_probs=9.7
Q ss_pred chhhHHHHHHHHh-hhceeEE
Q 016558 286 INGAYFLILSVLI-FGVTWAC 305 (387)
Q Consensus 286 I~GAYfLv~tvVl-iGgvwaC 305 (387)
..|++|||+.||+ ||.+|.|
T Consensus 13 ~igi~Ll~lLl~cgiGcvwhw 33 (158)
T PF11770_consen 13 SIGISLLLLLLLCGIGCVWHW 33 (158)
T ss_pred HHHHHHHHHHHHHhcceEEEe
Confidence 3466666533333 4445543
No 24
>PF06365 CD34_antigen: CD34/Podocalyxin family; InterPro: IPR013836 This family consists of several mammalian CD34 antigen proteins. The CD34 antigen is a human leukocyte membrane protein expressed specifically by lymphohematopoietic progenitor cells. CD34 is a phosphoprotein. Activation of protein kinase C (PKC) has been found to enhance CD34 phosphorylation [, ]. This family contains several eukaryotic podocalyxin proteins. Podocalyxin is a major membrane protein of the glomerular epithelium and is thought to be involved in maintenance of the architecture of the foot processes and filtration slits characteristic of this unique epithelium by virtue of its high negative charge. Podocalyxin functions as an anti-adhesin that maintains an open filtration pathway between neighbouring foot processes in the glomerular epithelium by charge repulsion [].
Probab=44.74 E-value=9.9 Score=36.31 Aligned_cols=23 Identities=22% Similarity=0.501 Sum_probs=14.1
Q ss_pred ceeecccchhhHHHHHHHHhhhc
Q 016558 279 YDKILTPINGAYFLILSVLIFGV 301 (387)
Q Consensus 279 Y~~iltPI~GAYfLv~tvVliGg 301 (387)
|..++.-+..+.||+++++++++
T Consensus 99 ~~~lI~lv~~g~~lLla~~~~~~ 121 (202)
T PF06365_consen 99 YPTLIALVTSGSFLLLAILLGAG 121 (202)
T ss_pred ceEEEehHHhhHHHHHHHHHHHH
Confidence 55666666666666666555554
No 25
>PF11669 WBP-1: WW domain-binding protein 1; InterPro: IPR021684 This family of proteins represents WBP-1, a ligand of the WW domain of Yes-associated protein. This protein has a proline-rich domain. WBP-1 does not bind to the SH3 domain [].
Probab=44.27 E-value=3.5 Score=34.94 Aligned_cols=10 Identities=30% Similarity=0.458 Sum_probs=3.9
Q ss_pred hceeEEEeec
Q 016558 300 GVTWACCKCR 309 (387)
Q Consensus 300 GgvwaCCkfR 309 (387)
+..++|..+|
T Consensus 35 ~c~c~~~~~r 44 (102)
T PF11669_consen 35 SCCCACRHRR 44 (102)
T ss_pred HHHHHHHHHH
Confidence 3333443333
No 26
>PF02480 Herpes_gE: Alphaherpesvirus glycoprotein E; InterPro: IPR003404 Glycoprotein E (gE) of Alphaherpesvirus forms a complex with glycoprotein I (gI), functioning as an immunoglobulin G (IgG) Fc binding protein. gE is involved in virus spread but is not essential for propagation [].; GO: 0016020 membrane; PDB: 2GJ7_F 2GIY_B.
Probab=42.27 E-value=8.4 Score=40.06 Aligned_cols=41 Identities=22% Similarity=0.062 Sum_probs=0.0
Q ss_pred chhhHHHHHHHHhhhceeEEEeecccccCCCCCce-eeecCCC
Q 016558 286 INGAYFLILSVLIFGVTWACCKCRKRRWNDGVPYQ-ELEMGLP 327 (387)
Q Consensus 286 I~GAYfLv~tvVliGgvwaCCkfRkrr~~~GvpYQ-ELEM~LP 327 (387)
+..++++.++++|+.++-+||.+.++|++ --+|+ .+++.-|
T Consensus 353 ~~l~vVlgvavlivVv~viv~vc~~~rrr-R~~~~~~~~~~~~ 394 (439)
T PF02480_consen 353 ALLGVVLGVAVLIVVVGVIVWVCLRCRRR-RRQRDKILNPFSP 394 (439)
T ss_dssp -------------------------------------------
T ss_pred chHHHHHHHHHHHHHHHHHhheeeeehhc-ccccccccCcCCC
Confidence 33444444555555555566667777777 67777 6666433
No 27
>PF07010 Endomucin: Endomucin; InterPro: IPR010740 This family consists of several mammalian endomucin proteins. Endomucin is an early endothelial-specific antigen that is also expressed on putative hematopoietic progenitor cells.
Probab=41.03 E-value=6.1 Score=39.01 Aligned_cols=38 Identities=21% Similarity=0.395 Sum_probs=22.1
Q ss_pred ccccccceeecccchhhHHHHHHHHhhhceeEEEeeccc
Q 016558 273 FIYLPSYDKILTPINGAYFLILSVLIFGVTWACCKCRKR 311 (387)
Q Consensus 273 f~~~pSY~~iltPI~GAYfLv~tvVliGgvwaCCkfRkr 311 (387)
+.-.|+|+.++-|+..|. +|+++++|-.+..+.+|||+
T Consensus 180 ~stspS~S~vilpvvIal-iVitl~vf~LvgLyr~C~k~ 217 (259)
T PF07010_consen 180 SSTSPSYSSVILPVVIAL-IVITLSVFTLVGLYRMCWKT 217 (259)
T ss_pred ccCCccccchhHHHHHHH-HHHHHHHHHHHHHHHHhhcC
Confidence 455789999999987655 44444444333333333344
No 28
>PF09972 DUF2207: Predicted membrane protein (DUF2207); InterPro: IPR018702 This domain has no known function.
Probab=40.75 E-value=90 Score=30.70 Aligned_cols=17 Identities=41% Similarity=0.622 Sum_probs=11.6
Q ss_pred CCceEEEEEEcCccccC
Q 016558 207 GEKTLIVTITIPTAVEN 223 (387)
Q Consensus 207 G~~~L~V~ItaPd~V~~ 223 (387)
.-+.++|+|..|..+..
T Consensus 130 ~i~~v~v~i~~P~~~~~ 146 (511)
T PF09972_consen 130 PIENVTVTITLPKPVDN 146 (511)
T ss_pred ccceEEEEEECCCCCcc
Confidence 34578889999955433
No 29
>PF07610 DUF1573: Protein of unknown function (DUF1573); InterPro: IPR011467 These hypothetical proteins from bacteria, such as Rhodopirellula baltica, Bacteroides thetaiotaomicron and Porphyromonas gingivalis, share a region of conserved sequence towards their N termini.
Probab=39.13 E-value=73 Score=22.96 Aligned_cols=42 Identities=17% Similarity=0.338 Sum_probs=28.9
Q ss_pred EEEEcCCCceEEEE-EEcC-ccccCCCceeeecccceeEEEEEE
Q 016558 201 ILVQNEGEKTLIVT-ITIP-TAVENPLKQLKISKHQTQKINISL 242 (387)
Q Consensus 201 LLVQNkG~~~L~V~-ItaP-d~V~~~~~~L~L~K~qskKV~IS~ 242 (387)
+-+.|.|+.+|.+. |.++ .=+.+....-.|.-+++.+|+|+|
T Consensus 2 F~~~N~g~~~L~I~~v~tsCgCt~~~~~~~~i~PGes~~i~v~y 45 (45)
T PF07610_consen 2 FEFTNTGDSPLVITDVQTSCGCTTAEYSKKPIAPGESGKIKVTY 45 (45)
T ss_pred EEEEECCCCcEEEEEeeEccCCEEeeCCcceECCCCEEEEEEEC
Confidence 56899999999885 4444 334444455556688888888764
No 30
>PHA03282 envelope glycoprotein E; Provisional
Probab=38.77 E-value=42 Score=36.35 Aligned_cols=15 Identities=40% Similarity=0.822 Sum_probs=11.9
Q ss_pred hhceeEEEeeccccc
Q 016558 299 FGVTWACCKCRKRRW 313 (387)
Q Consensus 299 iGgvwaCCkfRkrr~ 313 (387)
-..+|+|..+||+|.
T Consensus 424 glsvw~C~~c~r~ra 438 (540)
T PHA03282 424 GLSVWACVTCRRARA 438 (540)
T ss_pred Hhhheeeeeehhhhh
Confidence 346899999998865
No 31
>PHA03281 envelope glycoprotein E; Provisional
Probab=38.70 E-value=14 Score=40.42 Aligned_cols=45 Identities=22% Similarity=0.264 Sum_probs=30.1
Q ss_pred hhHHHHHHHHhhhceeEEEeecccccC-CCCCceeee--cCCCCccCC
Q 016558 288 GAYFLILSVLIFGVTWACCKCRKRRWN-DGVPYQELE--MGLPESVSA 332 (387)
Q Consensus 288 GAYfLv~tvVliGgvwaCCkfRkrr~~-~GvpYQELE--M~LP~S~ga 332 (387)
|...+++++|+++++|.-.+||+|+++ +.-+||+-- |+||+-.-.
T Consensus 562 ~~a~~~ll~l~~~~~c~~~~~~~~~~~~~~~~~~~s~~Y~~lP~~d~e 609 (642)
T PHA03281 562 GFAALALLCLAIALICTAKKFGHKAYRSDKAAYGQSMYYAGLPVDDFE 609 (642)
T ss_pred hhHHHHHHHHHHHHHHHHHHhhhheeeccccccccccccccCCCcccc
Confidence 344455666666776666788888665 777888753 589986544
No 32
>PF07790 DUF1628: Protein of unknown function (DUF1628); InterPro: IPR012859 The sequences making up this family are derived from hypothetical proteins of unknown function expressed by various archaeal species. The region in question is approximately 160 residues long.
Probab=37.26 E-value=12 Score=29.33 Aligned_cols=43 Identities=14% Similarity=0.248 Sum_probs=30.6
Q ss_pred cccchhhHHHHHHHHhhhceeEEEeecccccCCCCCceeeecC
Q 016558 283 LTPINGAYFLILSVLIFGVTWACCKCRKRRWNDGVPYQELEMG 325 (387)
Q Consensus 283 ltPI~GAYfLv~tvVliGgvwaCCkfRkrr~~~GvpYQELEM~ 325 (387)
++|+.|+-+|++..|+++++-+...|---......|+-.+++.
T Consensus 3 vS~viGviLliaitVilaavv~~~~~~~~~~~~~~P~~~~~~~ 45 (80)
T PF07790_consen 3 VSPVIGVILLIAITVILAAVVGAFVFGLDSSPESPPQASISVD 45 (80)
T ss_pred ccHHHHHHHHHHHHHHHHHHHHHHHhcccCCCCCCCEEEEEEE
Confidence 5799999999988888888877776665222245666666554
No 33
>PF05083 LST1: LST-1 protein; InterPro: IPR007775 B144/LST1 is a gene encoded in the human major histocompatibility complex that produces multiple forms of alternatively spliced mRNA and encodes peptides fewer than 100 amino acids in length. B144/LST1 is strongly expressed in dendritic cells. Transfection of B144/LST1 into a variety of cells induces morphologic changes including the production of long, thin filopodia []. A possible role in modulating immune responses. Induces morphological changes including production of filopodia and microspikes when overexpressed in a variety of cell types and may be involved in dendritic cell maturation. Isoform 1 and isoform 2 have an inhibitory effect on lymphocyte proliferation [, ]. ; GO: 0000902 cell morphogenesis, 0006955 immune response, 0016020 membrane
Probab=37.25 E-value=8.1 Score=32.08 Aligned_cols=41 Identities=27% Similarity=0.316 Sum_probs=22.4
Q ss_pred HHHHHHHHhhhceeEEEeecccccC-----CCCCceeeecC----CCCccCC
Q 016558 290 YFLILSVLIFGVTWACCKCRKRRWN-----DGVPYQELEMG----LPESVSA 332 (387)
Q Consensus 290 YfLv~tvVliGgvwaCCkfRkrr~~-----~GvpYQELEM~----LP~S~ga 332 (387)
.+|++++|++ +|.|..-||.++- -+.--|||-|+ ||++...
T Consensus 4 llll~vvll~--~clC~lsrRvkrLErs~~~~~~eQE~hyasLqrLPv~~se 53 (74)
T PF05083_consen 4 LLLLAVVLLS--ACLCRLSRRVKRLERSWEQLSSEQELHYASLQRLPVPSSE 53 (74)
T ss_pred hhhHHHHHHH--HHHHHHHhhhhhcccchhccccccchHHHHHHhCCCCCCC
Confidence 3344333333 3666665555421 22234888884 8888763
No 34
>PF07705 CARDB: CARDB; InterPro: IPR011635 The APHP (acidic peptide-dependent hydrolases/peptidase) domain is found in a variety of different proteins.; PDB: 2KUT_A 2L0D_A 3IDU_A 2KL6_A.
Probab=36.01 E-value=1.9e+02 Score=21.91 Aligned_cols=50 Identities=10% Similarity=0.240 Sum_probs=32.1
Q ss_pred CCcceEEEEEcCCCc---eEEEEEEcCccccCCCcee-eecccceeEEEEEEecC
Q 016558 195 GSGELTILVQNEGEK---TLIVTITIPTAVENPLKQL-KISKHQTQKINISLSAR 245 (387)
Q Consensus 195 ~S~~lsLLVQNkG~~---~L~V~ItaPd~V~~~~~~L-~L~K~qskKV~IS~s~~ 245 (387)
....+.+.|+|.|.. .+.|.+...... .....| .|..+++..|.+.+...
T Consensus 19 ~~~~i~~~V~N~G~~~~~~~~v~~~~~~~~-~~~~~i~~L~~g~~~~v~~~~~~~ 72 (101)
T PF07705_consen 19 EPVTITVTVKNNGTADAENVTVRLYLDGNS-VSTVTIPSLAPGESETVTFTWTPP 72 (101)
T ss_dssp SEEEEEEEEEE-SSS-BEEEEEEEEETTEE-EEEEEESEB-TTEEEEEEEEEE-S
T ss_pred CEEEEEEEEEECCCCCCCCEEEEEEECCce-eccEEECCcCCCcEEEEEEEEEeC
Confidence 456788999999986 466766555332 233344 66788888888888853
No 35
>PF12768 Rax2: Cortical protein marker for cell polarity
Probab=35.85 E-value=15 Score=36.16 Aligned_cols=43 Identities=19% Similarity=0.289 Sum_probs=26.5
Q ss_pred chhhHHHHHHHHhhhceeEEEeecccccC---CCCCceeeec--CCCC
Q 016558 286 INGAYFLILSVLIFGVTWACCKCRKRRWN---DGVPYQELEM--GLPE 328 (387)
Q Consensus 286 I~GAYfLv~tvVliGgvwaCCkfRkrr~~---~GvpYQELEM--~LP~ 328 (387)
+--|-=++|.++|+|++++++++||.... -..+|.|-|| .+|.
T Consensus 234 lAiALG~v~ll~l~Gii~~~~~r~~~~~~~~p~~~~~d~~~~~~~vpP 281 (281)
T PF12768_consen 234 LAIALGTVFLLVLIGIILAYIRRRRQGYVPAPTSPRIDEDEMMQRVPP 281 (281)
T ss_pred hHHHHHHHHHHHHHHHHHHHHHhhhccCcCCCcccccCcccccccCCC
Confidence 33444556777888888877644433222 1247999999 4663
No 36
>PF14610 DUF4448: Protein of unknown function (DUF4448)
Probab=35.83 E-value=18 Score=32.82 Aligned_cols=12 Identities=25% Similarity=0.390 Sum_probs=6.2
Q ss_pred ccceeEEEEEEe
Q 016558 232 KHQTQKINISLS 243 (387)
Q Consensus 232 K~qskKV~IS~s 243 (387)
.....+++|++.
T Consensus 96 ~~~~~~~~itl~ 107 (189)
T PF14610_consen 96 GEKYERNNITLQ 107 (189)
T ss_pred CCccceEEEEEE
Confidence 443434666665
No 37
>PF00635 Motile_Sperm: MSP (Major sperm protein) domain; InterPro: IPR000535 Major sperm proteins (MSP) are central components in molecular interactions underlying sperm motility in Caenorhabditis elegans, whose sperm employ an amoebae-like crawling motion using a MSP-containing lamellipod, rather than the flagellar-based swimming motion associated with other sperm. These proteins oligomerise to form an extensive filament system that extends from sperm villipoda, along the leading edge of the pseudopod. About 30 MSP isoforms may exist in C. elegans. MSPs form a fibrous network, whereby MSP dimers form helical subfilaments that coil around one another to produce filaments, which in turn form supercoils to produce bundles. The crystal structure of MSP from C. elegans reveals an immunoglobulin (Ig)-like seven-stranded beta sandwich fold []. ; GO: 0005198 structural molecule activity; PDB: 1MSP_A 3MSP_B 2BVU_B 2MSP_C 1Z9O_F 1Z9L_A 3IKK_A 1WIC_A 2CRI_A 2RR3_A ....
Probab=35.45 E-value=1.3e+02 Score=23.81 Aligned_cols=50 Identities=12% Similarity=0.129 Sum_probs=35.8
Q ss_pred CCcceEEEEEcCCCceEEEEEEcC--ccccCCCceeeecccceeEEEEEEec
Q 016558 195 GSGELTILVQNEGEKTLIVTITIP--TAVENPLKQLKISKHQTQKINISLSA 244 (387)
Q Consensus 195 ~S~~lsLLVQNkG~~~L~V~ItaP--d~V~~~~~~L~L~K~qskKV~IS~s~ 244 (387)
......|.+.|.+..++-.+|.+. ....+.+..=.|.-+++..|.|++..
T Consensus 18 ~~~~~~l~l~N~s~~~i~fKiktt~~~~y~v~P~~G~i~p~~~~~i~I~~~~ 69 (109)
T PF00635_consen 18 KQQSCELTLTNPSDKPIAFKIKTTNPNRYRVKPSYGIIEPGESVEITITFQP 69 (109)
T ss_dssp S-EEEEEEEEE-SSSEEEEEEEES-TTTEEEESSEEEE-TTEEEEEEEEE-S
T ss_pred ceEEEEEEEECCCCCcEEEEEEcCCCceEEecCCCEEECCCCEEEEEEEEEe
Confidence 345678899999999998888544 55567777667788999999998774
No 38
>TIGR01433 CyoA cytochrome o ubiquinol oxidase subunit II. This enzyme catalyzes the oxidation of ubiquinol with the concomitant reduction of molecular oxygen to water. This acts as the terminal electron acceptor in the respiratory chain. Subunit II is responsible for binding and oxidation of the ubiquinone substrate. This sequence is closely related to QoxA, which oxidizes quinol in gram positive bacteria but which is in complex with subunits which utilize cytochromes a in the reduction of molecular oxygen. Slightly more distantly related is subunit II of cytochrome c oxidase which uses cyt. c as the oxidant.
Probab=35.22 E-value=43 Score=31.87 Aligned_cols=37 Identities=16% Similarity=0.247 Sum_probs=22.0
Q ss_pred hhHHHHHHHHhhhceeEEEeecccccCCC--C---Cceeeec
Q 016558 288 GAYFLILSVLIFGVTWACCKCRKRRWNDG--V---PYQELEM 324 (387)
Q Consensus 288 GAYfLv~tvVliGgvwaCCkfRkrr~~~G--v---pYQELEM 324 (387)
++.++|+++|.+..+|...+||+++.... . .-+.||+
T Consensus 37 ~~~~ii~v~v~~~~~~~~~r~r~~~~~~~~~p~~~~~~~lE~ 78 (226)
T TIGR01433 37 GLMLLVVIPVILMTLFFAWKYRATNKDADYSPNWHHSTKIEI 78 (226)
T ss_pred HHHHHHHHHHHHHHheeeEEEeccCCcCCCCCcccCCceeeh
Confidence 34444555555556888888988765421 1 2245885
No 39
>KOG4222 consensus Axon guidance receptor Dscam [Signal transduction mechanisms]
Probab=34.16 E-value=1.6e+02 Score=35.26 Aligned_cols=115 Identities=16% Similarity=0.052 Sum_probs=53.3
Q ss_pred ccccccccccceeecccchhhHHHHHHHHhhhcee-EEEeecccccC-------CCCCceeeecCCCCccCCc-ccccCC
Q 016558 269 EEKIFIYLPSYDKILTPINGAYFLILSVLIFGVTW-ACCKCRKRRWN-------DGVPYQELEMGLPESVSAM-NVETAE 339 (387)
Q Consensus 269 d~~~f~~~pSY~~iltPI~GAYfLv~tvVliGgvw-aCCkfRkrr~~-------~GvpYQELEM~LP~S~ga~-evEtaD 339 (387)
+.+.-...++|+.+--|-..|-.-++.+||+++.- +||.|||+++. ..++-|.|=|.++++.+.. --.-..
T Consensus 855 ~~ns~~~~~s~~v~~qp~f~a~v~~a~~ii~~v~s~~~~y~~rk~~~~~~~~t~~~s~~d~~f~s~n~~~~~~~~~~~~~ 934 (1281)
T KOG4222|consen 855 DRNSETEQISVDVVNQPAFIAGVHRACLIIVMVFSIIWLYWRRKEPLSGKDLTAGLSRLDNLFTSLNVNQGKGYLPCYSP 934 (1281)
T ss_pred ccchhhhhheeeeecCcchheeeeeeeeeeeeeeeeeeeeecccccccccccccccccCCcceecccccccccccccccc
Confidence 33334444566666655333332333344444443 78888888665 3445667777777443321 111122
Q ss_pred CcCCCCCCCCCcccCC-C-CCCCCCccccCcCcccCCCCCCCCCccCC
Q 016558 340 GWDEGWDDDWDENNAV-K-SPGASRIGSISANGLTSRSPNRDGWEHDW 385 (387)
Q Consensus 340 GWDdgWDDDWDDEEAp-K-SPs~~~t~SlSSnGLaSRrssKDGWk~dW 385 (387)
+|-.. +-|=+++.|- + -|--+.+..+++ =..-|=..-+||.-+|
T Consensus 935 ~W~~~-~~~~~~~~ag~~l~~~vP~s~~~~n-~~~~~~~~s~~~n~~s 980 (1281)
T KOG4222|consen 935 GWRTA-RLDHQNERAGQGLLPPVPNSQDNHN-DISERGLGSIGWNTDS 980 (1281)
T ss_pred ccccc-ccccccccccCcccCCCCCcccccc-cccccccccccccccc
Confidence 22111 1222233331 2 122334445555 2222336667888777
No 40
>PF13908 Shisa: Wnt and FGF inhibitory regulator
Probab=32.54 E-value=35 Score=30.52 Aligned_cols=41 Identities=20% Similarity=0.292 Sum_probs=24.3
Q ss_pred eecccchhhHHHHHHHHhhhceeEEEeecccccCCCCCceee
Q 016558 281 KILTPINGAYFLILSVLIFGVTWACCKCRKRRWNDGVPYQEL 322 (387)
Q Consensus 281 ~iltPI~GAYfLv~tvVliGgvwaCCkfRkrr~~~GvpYQEL 322 (387)
.++.-|.|+.|+ +++|++..-+-||+-++.|++.....+.+
T Consensus 80 iivgvi~~Vi~I-v~~Iv~~~Cc~c~~~K~~~~~~~~~~~~~ 120 (179)
T PF13908_consen 80 IIVGVICGVIAI-VVLIVCFCCCCCCLYKKCRSQRPNRSRAL 120 (179)
T ss_pred eeeehhhHHHHH-HHhHhhheeccccccccccCccccccccc
Confidence 445666666655 55555567677888886555433444443
No 41
>PF04478 Mid2: Mid2 like cell wall stress sensor; InterPro: IPR007567 This family represents a region near the C terminus of Mid2, which contains a transmembrane region. The remainder of the protein sequence is serine-rich and of low complexity, and is therefore impossible to align accurately. Mid2 is thought to act as a mechanosensor of cell wall stress. The C-terminal cytoplasmic region of Mid2 is known to interact with Rom2, a guanine nucleotide exchange factor (GEF) for Rho1, which is part of the cell wall integrity signalling pathway [].
Probab=32.30 E-value=34 Score=31.78 Aligned_cols=16 Identities=25% Similarity=0.490 Sum_probs=8.7
Q ss_pred HhhhceeEEEeecccc
Q 016558 297 LIFGVTWACCKCRKRR 312 (387)
Q Consensus 297 VliGgvwaCCkfRkrr 312 (387)
+|++.+|.||.-|||.
T Consensus 65 ~il~lvf~~c~r~kkt 80 (154)
T PF04478_consen 65 GILALVFIFCIRRKKT 80 (154)
T ss_pred HHHHhheeEEEecccC
Confidence 4455566666554443
No 42
>PF05545 FixQ: Cbb3-type cytochrome oxidase component FixQ; InterPro: IPR008621 This family consists of several Cbb3-type cytochrome oxidase components (FixQ/CcoQ). FixQ is found in nitrogen fixing bacteria. Since nitrogen fixation is an energy-consuming process, effective symbioses depend on operation of a respiratory chain with a high affinity for O2, closely coupled to ATP production. This requirement is fulfilled by a special three-subunit terminal oxidase (cytochrome terminal oxidase cbb3), which was first identified in Bradyrhizobium japonicum as the product of the fixNOQP operon [].
Probab=32.08 E-value=7.4 Score=28.54 Aligned_cols=28 Identities=11% Similarity=0.184 Sum_probs=14.2
Q ss_pred ceeecccchhhHHHHHHHHhhhceeEEE
Q 016558 279 YDKILTPINGAYFLILSVLIFGVTWACC 306 (387)
Q Consensus 279 Y~~iltPI~GAYfLv~tvVliGgvwaCC 306 (387)
|..+..=+.+..++++.++.+|.+|-.+
T Consensus 3 ~~~~~~~~~~~~~v~~~~~F~gi~~w~~ 30 (49)
T PF05545_consen 3 YETLQGFARSIGTVLFFVFFIGIVIWAY 30 (49)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 3334444445566666666666444333
No 43
>PF12297 EVC2_like: Ellis van Creveld protein 2 like protein; InterPro: IPR022076 This family of proteins is found in eukaryotes. Proteins in this family are typically between 571 and 1310 amino acids in length. There are two conserved sequence motifs: LPA and ELH. EVC2 is implicated in Ellis van Creveld chondrodysplastic dwarfism in humans. Mutations in this protein can give rise to this congenital condition. LIMBIN is a protein which shares around 80% sequence homology with EVC2 and it is implicated in a similar condition in bovine chondrodysplastic dwarfism.
Probab=31.54 E-value=10 Score=40.02 Aligned_cols=26 Identities=23% Similarity=0.473 Sum_probs=21.7
Q ss_pred chhhHHHHHHHHhhhceeEEEeeccc
Q 016558 286 INGAYFLILSVLIFGVTWACCKCRKR 311 (387)
Q Consensus 286 I~GAYfLv~tvVliGgvwaCCkfRkr 311 (387)
+++|-|+|+.+|-|..+|+||.|-.+
T Consensus 65 lhaagFfvaflvslVL~~l~~f~l~r 90 (429)
T PF12297_consen 65 LHAAGFFVAFLVSLVLTWLCFFLLAR 90 (429)
T ss_pred hHHHHHHHHHHHHHHHHHHHHHHHHH
Confidence 66788888889999999999987554
No 44
>PF06030 DUF916: Bacterial protein of unknown function (DUF916); InterPro: IPR010317 This family consists of putative cell surface proteins, from Firmicutes, of unknown function.
Probab=31.36 E-value=83 Score=27.26 Aligned_cols=43 Identities=19% Similarity=0.230 Sum_probs=32.4
Q ss_pred eeccCCCCcceEEEEEcCCCceEEEEEEcCccccCCCceeeec
Q 016558 189 IQNFDTGSGELTILVQNEGEKTLIVTITIPTAVENPLKQLKIS 231 (387)
Q Consensus 189 L~v~gn~S~~lsLLVQNkG~~~L~V~ItaPd~V~~~~~~L~L~ 231 (387)
|++.-.....+.|.|+|....+++|.|.+-+..+..-..|...
T Consensus 21 L~~~P~q~~~l~v~i~N~s~~~~tv~v~~~~A~Tn~nG~I~Y~ 63 (121)
T PF06030_consen 21 LKVKPGQKQTLEVRITNNSDKEITVKVSANTATTNDNGVIDYS 63 (121)
T ss_pred EEeCCCCEEEEEEEEEeCCCCCEEEEEEEeeeEecCCEEEEEC
Confidence 4555556678999999999999999997776666666666553
No 45
>PF10989 DUF2808: Protein of unknown function (DUF2808); InterPro: IPR021256 This family of proteins with unknown function appears to be restricted to Cyanobacteria.
Probab=30.01 E-value=1.2e+02 Score=26.59 Aligned_cols=24 Identities=29% Similarity=0.450 Sum_probs=19.1
Q ss_pred EEEEEcCCCceEEEEEEcCccccC
Q 016558 200 TILVQNEGEKTLIVTITIPTAVEN 223 (387)
Q Consensus 200 sLLVQNkG~~~L~V~ItaPd~V~~ 223 (387)
.++-++.|+.-.+|+|+.|++++.
T Consensus 31 ~~~p~~~~~~L~~l~I~~p~~~~~ 54 (146)
T PF10989_consen 31 IIVPQDAGEALQKLTISQPDGFDG 54 (146)
T ss_pred EEccccCCCcceeEEEEccccccc
Confidence 344568899999999999988755
No 46
>PF15102 TMEM154: TMEM154 protein family
Probab=29.68 E-value=50 Score=30.48 Aligned_cols=28 Identities=7% Similarity=0.035 Sum_probs=14.7
Q ss_pred eeEEEeecccccCCCCCceeeecCCCCcc
Q 016558 302 TWACCKCRKRRWNDGVPYQELEMGLPESV 330 (387)
Q Consensus 302 vwaCCkfRkrr~~~GvpYQELEM~LP~S~ 330 (387)
++.-+.+||||.. .-|||+..=+.+-+.
T Consensus 76 V~lv~~~kRkr~K-~~~ss~gsq~~~qt~ 103 (146)
T PF15102_consen 76 VCLVIYYKRKRTK-QEPSSQGSQSALQTY 103 (146)
T ss_pred HHheeEEeecccC-CCCcccccccccccc
Confidence 3344444555554 467777666544433
No 47
>KOG4764 consensus Uncharacterized conserved protein [Function unknown]
Probab=28.82 E-value=22 Score=29.39 Aligned_cols=9 Identities=56% Similarity=1.664 Sum_probs=4.9
Q ss_pred CcCCCCCCC
Q 016558 340 GWDEGWDDD 348 (387)
Q Consensus 340 GWDdgWDDD 348 (387)
-|.++||||
T Consensus 40 vWEdnWDDd 48 (70)
T KOG4764|consen 40 VWEDNWDDD 48 (70)
T ss_pred hhhhcCCcc
Confidence 566666443
No 48
>PF15065 NCU-G1: Lysosomal transcription factor, NCU-G1
Probab=26.68 E-value=34 Score=35.02 Aligned_cols=26 Identities=23% Similarity=0.428 Sum_probs=13.8
Q ss_pred chhhHHHHHHH-HhhhceeEEEeeccc
Q 016558 286 INGAYFLILSV-LIFGVTWACCKCRKR 311 (387)
Q Consensus 286 I~GAYfLv~tv-VliGgvwaCCkfRkr 311 (387)
|..+-|.++.+ ||+|++..|++-+|+
T Consensus 322 i~~vgLG~P~l~li~Ggl~v~~~r~r~ 348 (350)
T PF15065_consen 322 IMAVGLGVPLLLLILGGLYVCLRRRRK 348 (350)
T ss_pred HHHHHhhHHHHHHHHhhheEEEecccc
Confidence 34455556655 556666666543333
No 49
>TIGR02745 ccoG_rdxA_fixG cytochrome c oxidase accessory protein FixG. Member of this ferredoxin-like protein family are found exclusively in species with an operon encoding the cbb3 type of cytochrome c oxidase (cco-cbb3), and near the cco-cbb3 operon in about half the cases. The cco-cbb3 is found in a variety of proteobacteria and almost nowhere else, and is associated with oxygen use under microaerobic conditions. Some (but not all) of these proteobacteria are also nitrogen-fixing, hence the gene symbol fixG. FixG was shown essential for functional cco-cbb3 expression in Bradyrhizobium japonicum.
Probab=25.95 E-value=2.4e+02 Score=29.64 Aligned_cols=49 Identities=10% Similarity=0.160 Sum_probs=34.3
Q ss_pred CcceEEEEEcCCCceEEEEEEcC--ccccCCC--ceeeecccceeEEEEEEec
Q 016558 196 SGELTILVQNEGEKTLIVTITIP--TAVENPL--KQLKISKHQTQKINISLSA 244 (387)
Q Consensus 196 S~~lsLLVQNkG~~~L~V~ItaP--d~V~~~~--~~L~L~K~qskKV~IS~s~ 244 (387)
.-.|.|.++|+.+.+..+.|+.. +.+.... .+++|..++..++.|.+..
T Consensus 347 ~N~Y~~~i~Nk~~~~~~~~l~v~g~~~~~~~~~~~~i~v~~g~~~~~~v~v~~ 399 (434)
T TIGR02745 347 ENTYTLKILNKTEQPHEYYLSVLGLPGIKIEGPGAPIHVKAGEKVKLPVFLRT 399 (434)
T ss_pred EEEEEEEEEECCCCCEEEEEEEecCCCcEEEcCCceEEECCCCEEEEEEEEEe
Confidence 34689999999999777777654 2222222 3788888888877777763
No 50
>PTZ00364 dipeptidyl-peptidase I precursor; Provisional
Probab=25.86 E-value=1.4e+02 Score=32.44 Aligned_cols=16 Identities=6% Similarity=0.052 Sum_probs=11.6
Q ss_pred ceEEEEeccCceEEec
Q 016558 248 SKLVLNAGNGECVLHM 263 (387)
Q Consensus 248 ~~IvL~aGkG~C~Lhi 263 (387)
+-+-|..|...|-|..
T Consensus 437 GYfRI~RG~N~CGIes 452 (548)
T PTZ00364 437 GTRKIARGVNAYNIES 452 (548)
T ss_pred CeEEEEcCCCcccccc
Confidence 3466777778898875
No 51
>PF07204 Orthoreo_P10: Orthoreovirus membrane fusion protein p10; InterPro: IPR009854 This family consists of several Orthoreovirus membrane fusion protein p10 sequences. p10 is thought to be a multifunctional protein that plays a key role in virus-host interaction [].
Probab=25.83 E-value=13 Score=32.34 Aligned_cols=69 Identities=19% Similarity=0.307 Sum_probs=29.1
Q ss_pred EEEEeccCceEEecCCCCcccccccccccceeecccchhhHHHHHHHHhhhceeEEEeecccccC-CCCCceee
Q 016558 250 LVLNAGNGECVLHMGRPASEEKIFIYLPSYDKILTPINGAYFLILSVLIFGVTWACCKCRKRRWN-DGVPYQEL 322 (387)
Q Consensus 250 IvL~aGkG~C~Lhi~~~vsd~~~f~~~pSY~~iltPI~GAYfLv~tvVliGgvwaCCkfRkrr~~-~GvpYQEL 322 (387)
++---|+-.|.-.-.++-.+=++-..|-+|--++.+- |+++||+++ |+.+ .||+.|++..+ -.+-+.||
T Consensus 12 ~~svfg~vhcqa~~nsaGgdL~atS~~~ayWpyLA~G-GG~iLilIi--i~Lv-~CC~~K~K~~~~r~~~~reL 81 (98)
T PF07204_consen 12 ATSVFGNVHCQASQNSAGGDLQATSSFVAYWPYLAAG-GGLILILII--IALV-CCCRAKHKTSAARNTFHREL 81 (98)
T ss_pred HHHhccchheeccccCCCCCeEEeehHHhhhHHhhcc-chhhhHHHH--HHHH-HHhhhhhhhHhhhhHHHHHH
Confidence 3333455556544332222312223333455555554 433333222 4444 46655555444 23344444
No 52
>PF14316 DUF4381: Domain of unknown function (DUF4381)
Probab=25.45 E-value=8.8 Score=33.46 Aligned_cols=32 Identities=16% Similarity=0.330 Sum_probs=17.4
Q ss_pred hhhHHHHHHHHhhhceeEEEeecccccCCCCCce
Q 016558 287 NGAYFLILSVLIFGVTWACCKCRKRRWNDGVPYQ 320 (387)
Q Consensus 287 ~GAYfLv~tvVliGgvwaCCkfRkrr~~~GvpYQ 320 (387)
.-.+-+++++||++.++..+.++|++++ .+|.
T Consensus 20 a~GWwll~~lll~~~~~~~~~~~r~~~~--~~yr 51 (146)
T PF14316_consen 20 APGWWLLLALLLLLLILLLWRLWRRWRR--NRYR 51 (146)
T ss_pred cHHHHHHHHHHHHHHHHHHHHHHHHHHc--cHHH
Confidence 3344455555555556666665555554 3554
No 53
>PF13980 UPF0370: Uncharacterised protein family (UPF0370)
Probab=24.99 E-value=14 Score=29.82 Aligned_cols=54 Identities=24% Similarity=0.554 Sum_probs=32.4
Q ss_pred hHHHHHHHHhhhceeEEEeecccccCCCCCceeeecCCCCccCCcccccCCCcCCCCCCCCCc
Q 016558 289 AYFLILSVLIFGVTWACCKCRKRRWNDGVPYQELEMGLPESVSAMNVETAEGWDEGWDDDWDE 351 (387)
Q Consensus 289 AYfLv~tvVliGgvwaCCkfRkrr~~~GvpYQELEM~LP~S~ga~evEtaDGWDdgWDDDWDD 351 (387)
-|.-|+.++|+|.+|--++=-+|- +--+|-.=-=+||.- -+-++.||+ +|||-.
T Consensus 6 dYWWiiLl~lvG~i~n~iK~L~Rv--D~K~fL~nKP~lPPH-----RDnN~~WDd--eDDwPk 59 (63)
T PF13980_consen 6 DYWWIILLILVGMIINGIKELRRV--DHKKFLDNKPELPPH-----RDNNAKWDD--EDDWPK 59 (63)
T ss_pred HHHHHHHHHHHHHHHHHHHHHHhc--CHHHHhcCCCCCCCC-----Ccccccccc--cccccc
Confidence 477788888999988887633331 112332222246642 245677888 788854
No 54
>PF06682 DUF1183: Protein of unknown function (DUF1183); InterPro: IPR009567 This family consists of several eukaryotic proteins of around 360 residues in length. The function of this family is unknown.
Probab=24.88 E-value=1.3e+02 Score=30.75 Aligned_cols=21 Identities=19% Similarity=0.264 Sum_probs=14.2
Q ss_pred CCcCCCCCCCCCcccCCCCCC
Q 016558 339 EGWDEGWDDDWDENNAVKSPG 359 (387)
Q Consensus 339 DGWDdgWDDDWDDEEApKSPs 359 (387)
.||-=+|+.+|+.-..|-+|-
T Consensus 198 ggggGGgg~~~~~~~~PPPPy 218 (318)
T PF06682_consen 198 GGGGGGGGGGWGGYPDPPPPY 218 (318)
T ss_pred cccccCCCCCCCCCCCCCCCC
Confidence 455556777777766776666
No 55
>smart00557 IG_FLMN Filamin-type immunoglobulin domains. These form a rod-like structure in the actin-binding cytoskeleton protein, filamin. The C-terminal repeats of filamin bind beta1-integrin (CD29).
Probab=23.06 E-value=3.9e+02 Score=21.32 Aligned_cols=44 Identities=20% Similarity=0.291 Sum_probs=30.9
Q ss_pred ceEEEEEcCCCceEEEEEEcCccccCCCceeeecccceeEEEEEEec
Q 016558 198 ELTILVQNEGEKTLIVTITIPTAVENPLKQLKISKHQTQKINISLSA 244 (387)
Q Consensus 198 ~lsLLVQNkG~~~L~V~ItaPd~V~~~~~~L~L~K~qskKV~IS~s~ 244 (387)
.+.|...+.|...|.|.|+-|+. ...++++.....-...|+|+-
T Consensus 21 ~f~v~~~d~G~~~~~v~i~~p~g---~~~~~~v~d~~dGty~v~y~P 64 (93)
T smart00557 21 EFTIDTRGAGGGELEVEVTGPSG---KKVPVEVKDNGDGTYTVSYTP 64 (93)
T ss_pred EEEEEcCCCCCCcEEEEEECCCC---CeeEeEEEeCCCCEEEEEEEe
Confidence 55555666688999999999965 224566666666677777773
No 56
>PHA03286 envelope glycoprotein E; Provisional
Probab=22.36 E-value=29 Score=37.25 Aligned_cols=39 Identities=21% Similarity=0.239 Sum_probs=22.5
Q ss_pred HHHHHhhhceeEEEeecccccC-CCCCceee--ecCCCCccC
Q 016558 293 ILSVLIFGVTWACCKCRKRRWN-DGVPYQEL--EMGLPESVS 331 (387)
Q Consensus 293 v~tvVliGgvwaCCkfRkrr~~-~GvpYQEL--EM~LP~S~g 331 (387)
++++|++++.|+-|.|||||++ -.-.+|+- =|.||--.-
T Consensus 400 ~~~~~~~~~~~~~~~~~r~~~~r~~~~~~~~~ky~~lp~n~~ 441 (492)
T PHA03286 400 AILVVLLFALCIAGLYRRRRRHRTNGYFQAYPKYMSLPSNDE 441 (492)
T ss_pred HHHHHHHHHHHhHhHhhhhhhhhcccccccCcccccCCCccc
Confidence 3566677777777888877665 11122221 277885443
No 57
>PHA03291 envelope glycoprotein I; Provisional
Probab=22.00 E-value=31 Score=36.13 Aligned_cols=16 Identities=31% Similarity=0.598 Sum_probs=11.0
Q ss_pred HHhhhceeEEEeeccc
Q 016558 296 VLIFGVTWACCKCRKR 311 (387)
Q Consensus 296 vVliGgvwaCCkfRkr 311 (387)
+.+|-|.|+||..|+.
T Consensus 299 ~cV~lGSC~Ccl~R~~ 314 (401)
T PHA03291 299 ACVFLGSCACCLHRRC 314 (401)
T ss_pred HHhhhhhhhhhhhhhh
Confidence 3445678999986544
No 58
>TIGR01732 tiny_TM_bacill conserved hypothetical tiny transmembrane protein. This model represents a family of hypothetical proteins, half of which are 40 residues or less in length. Members are found only in spore-forming species. A Gly-rich variable region is followed by a strongly conserved, highly hydrophobic region, predicted to form a transmembrane helix, ending with an invariant Gly. The consensus for this stretch is FALLVVFILLIIV.
Probab=20.97 E-value=72 Score=22.10 Aligned_cols=14 Identities=21% Similarity=0.555 Sum_probs=9.1
Q ss_pred HHHHHHHHhhhcee
Q 016558 290 YFLILSVLIFGVTW 303 (387)
Q Consensus 290 YfLv~tvVliGgvw 303 (387)
..||+.++|+|++|
T Consensus 12 vVLFILLIIiga~~ 25 (26)
T TIGR01732 12 VVLFILLVIVGAAF 25 (26)
T ss_pred HHHHHHHHHhheee
Confidence 34566667777766
No 59
>PF07889 DUF1664: Protein of unknown function (DUF1664); InterPro: IPR012458 The members of this family are hypothetical plant proteins of unknown function. The region featured in this family is approximately 100 amino acids long.
Probab=20.70 E-value=51 Score=29.45 Aligned_cols=27 Identities=7% Similarity=0.069 Sum_probs=19.6
Q ss_pred hhHHHHHHHHhhhceeEEEeecccccC
Q 016558 288 GAYFLILSVLIFGVTWACCKCRKRRWN 314 (387)
Q Consensus 288 GAYfLv~tvVliGgvwaCCkfRkrr~~ 314 (387)
++|+++++++|.++.+.|++|+..+..
T Consensus 5 ~~~~i~paa~~gavGY~Y~wwKGws~s 31 (126)
T PF07889_consen 5 WSSLIVPAAAIGAVGYGYMWWKGWSFS 31 (126)
T ss_pred ccchhhHHHHHHHHHheeeeecCCchh
Confidence 467778888888887777777766543
No 60
>PF12273 RCR: Chitin synthesis regulation, resistance to Congo red; InterPro: IPR020999 RCR proteins are ER membrane proteins that regulate chitin deposition in fungal cell walls. Although chitin, a linear polymer of beta-1,4-linked N-acetylglucosamine, constitutes only 2% of the cell wall it plays a vital role in the overall protection of the cell wall against stress, noxious chemicals and osmotic pressure changes. Congo red is a cell wall-disrupting benzidine-type dye extensively used in many cell wall mutant studies that specifically targets chitin in yeast cells and inhibits growth. RCR proteins render the yeasts resistant to Congo red by diminishing the content of chitin in the cell wall []. RCR proteins are probably regulating chitin synthase III interact directly with ubiquitin ligase Rsp5, and the VPEY motif is necessary for this, via interaction with the WW domains of Rsp5 [].
Probab=20.68 E-value=30 Score=29.66 Aligned_cols=25 Identities=16% Similarity=0.252 Sum_probs=12.2
Q ss_pred hhHHHHHHHHhhhceeEEEeeccccc
Q 016558 288 GAYFLILSVLIFGVTWACCKCRKRRW 313 (387)
Q Consensus 288 GAYfLv~tvVliGgvwaCCkfRkrr~ 313 (387)
++++++|+|+||+..|. -+-|+||.
T Consensus 6 ~iii~~i~l~~~~~~~~-~rRR~r~G 30 (130)
T PF12273_consen 6 AIIIVAILLFLFLFYCH-NRRRRRRG 30 (130)
T ss_pred HHHHHHHHHHHHHHHHH-HHHHhhcC
Confidence 34444444444444444 55555553
No 61
>TIGR02537 arch_flag_Nterm archaeal flagellin N-terminal-like domain. This model describes a hydrophobic N-terminal sequence of archaeal flagellins and other archaeal proteins. The sequence is directly analogous to bacterial sequences recognized by TIGR02532, which has cleavage motif resembling G^FxxxE followed by strongly hydrophobic sequence. Such sequences are the recognized for cleavage and methylation, and include pilins and other pilus components and competence and type II secretion secretion proteins. In the present family, the E is not conversed and sequence differs enough that there is no overlap between this family and TIGR02532.
Probab=20.60 E-value=65 Score=21.86 Aligned_cols=20 Identities=25% Similarity=0.579 Sum_probs=14.2
Q ss_pred cccchhhHHHHHHHHhhhce
Q 016558 283 LTPINGAYFLILSVLIFGVT 302 (387)
Q Consensus 283 ltPI~GAYfLv~tvVliGgv 302 (387)
++||.|+-+|+..+++++++
T Consensus 4 is~I~~~iiliai~ivla~~ 23 (26)
T TIGR02537 4 ISPIIGTIILIAITIVLAAA 23 (26)
T ss_pred chhHHHHHHHHHHHHHHHHh
Confidence 57888888777666666554
No 62
>PF11980 DUF3481: Domain of unknown function (DUF3481); InterPro: IPR022579 This domain of unknown function is located in the C terminus of the eukaryotic neuropilin receptor family of proteins. It is found in association with PF00754 from PFAM, PF00431 from PFAM and PF00629 from PFAM. There are two completely conserved residues (Y and E) that may be functionally important.
Probab=20.59 E-value=41 Score=28.80 Aligned_cols=46 Identities=26% Similarity=0.469 Sum_probs=25.9
Q ss_pred ccccccceeecccchhhHHHHHHHHhhhceeEEEeecc--cccCCCCCc
Q 016558 273 FIYLPSYDKILTPINGAYFLILSVLIFGVTWACCKCRK--RRWNDGVPY 319 (387)
Q Consensus 273 f~~~pSY~~iltPI~GAYfLv~tvVliGgvwaCCkfRk--rr~~~GvpY 319 (387)
+|.+|.|--+++---||- |++..|++|++-.|++|+- ++++--+.|
T Consensus 9 vqplp~~~yyiiA~gga~-llL~~v~l~vvL~C~r~~~a~kk~~~s~~y 56 (87)
T PF11980_consen 9 VQPLPPYWYYIIAMGGAL-LLLVAVCLGVVLYCHRFHWAAKKRSHSVLY 56 (87)
T ss_pred cCCCCceeeHHHhhccHH-HHHHHHHHHHHHhhhhhccccccCccceee
Confidence 355777666665554444 5556666677766666553 444423444
No 63
>PF15234 LAT: Linker for activation of T-cells
Probab=20.00 E-value=34 Score=33.25 Aligned_cols=22 Identities=18% Similarity=0.363 Sum_probs=17.3
Q ss_pred chhhHHHHHHHHhhhceeEEEe
Q 016558 286 INGAYFLILSVLIFGVTWACCK 307 (387)
Q Consensus 286 I~GAYfLv~tvVliGgvwaCCk 307 (387)
+.|..+|-|.+||+++.|.||+
T Consensus 10 ~LgLLlLplla~LlmALCvrCR 31 (230)
T PF15234_consen 10 VLGLLLLPLLAVLLMALCVRCR 31 (230)
T ss_pred HHHHHHHHHHHHHHHHHHHHHh
Confidence 4577777788888888888884
Done!