Sequence Search

Find iBGCs by protein sequence similarity

The Sequence search takes a protein sequence and finds iBGCs that encode a similar protein. It uses phmmer (a sensitive profile-based aligner from the HMMER suite) to align your query against the proteins in the catalogue’s clusters. Use it when you have a specific protein in hand — a characterised enzyme, a hit from your own genome, or a sequence from the literature.

Unlike the domain and find-similar searches, which compare clusters by shared domains, this search compares raw protein sequence — so it can find matches that domain membership alone would miss.

Supplying a query

Open the sequence chip and paste a protein sequence as FASTA or raw amino acids. A counter shows the length; the maximum is 5,000 amino acids.

Three cut-off sliders control what counts as a hit. A protein must pass all three to count:

Cut-off What it controls
Min bitscore Overall alignment strength (0 = permissive, 500 = strict).
Min % identity The fraction of aligned positions that are identical.
Min query coverage How much of your query the alignment spans.

Loosen the cut-offs to cast a wide net for distant homologues; tighten them to keep only close matches.

Running and reading results

The search runs on Run Query and runs as a background job — the discovery platform shows progress and fills in the results when it finishes. It combines with any active filters and domain conditions, so you can restrict the search to, say, marine Actinomycetota.

Results rank by match strength, and the roster adapts:

  • The similarity column becomes Bitscore (the score of the iBGC’s best-matching protein).
  • A Best hit column shows the protein ID of that matching gene.

On the Variables map you gain Bitscore, Identity %, and Query coverage % axes — useful for separating strong, full-length matches from short, weak ones.

Tips

  • Copy a query straight from the protein panel: click a gene in any iBGC, then use its Copy button.
  • Start with moderate cut-offs (e.g. a low-to-mid bitscore, low identity, high coverage) to find divergent homologues, then raise identity to focus.
  • Use the Best-hit column to jump into the matching gene and confirm the alignment in the protein panel.