BotShelf Vampire BOTSHELF VAMPIRE Register

Biotech: open data, real tools, reproducible steps

Screen known actives with RDKit, check AlphaFold confidence, profile UniProt sequences, digest PubMed, then pull PubChem, MyGene, PDBe and STRING neighbors — public APIs only, research/education, no pathogen protocols.

Research and education only. Calculated properties and public annotations are not evidence of efficacy or safety, not clinical advice, and not pathogen or dual-use guidance. BSV recipes here use public compound/gene/structure APIs — no enhancement or synthesis protocols.

Interactive tool, runs in your browser (English / Japanese)
Research reproducibility worksheet →
Identify missing evidence and export an analysis record.

Build recipes

Step-by-step workflows written by BSV. Each one combines real open tools into something you can make today.

Drug-likeness screen of known actives with ChEMBL and RDKit

Tested by BSV

Pull potent compounds for one target from ChEMBL, compute standard descriptors and QED with RDKit, and rank them with a rule-of-five flag.

Input
ChEMBL target id (default CHEMBL203, EGFR), minimum pChEMBL value (default 7), number of compounds.
Output
screen_<target>.csv: MW, cLogP, H-bond donors and acceptors, TPSA, rotatable bonds, QED, rule-of-five violations and pass flag, SMILES.
Prerequisites
Python 3.10 or later and pip install rdkit. The public ChEMBL API can be slow; the script retries.
Steps
  1. Save the code below as chembl_screen.py.
  2. Run python chembl_screen.py CHEMBL203 7 60.
  3. Open the CSV, sort by QED and look at the structures of the top rows (paste a SMILES into any structure viewer).
  4. Change the target id to one you study; find ids by searching the ChEMBL website.
  5. Keep the CSV with the date: ChEMBL releases change the data.
Expected result
Most potent, published actives pass the rule of five; the useful part is seeing which ones do not and why. Rule of five and QED are rough filters, not predictions.
Next step
Check the target's predicted structure confidence before any docking idea (the next recipe uses the same target, EGFR, UniProt P00533). Open
Code (BSV original, MIT licence)

Download .py · chembl_screen.py

#!/usr/bin/env python3
"""BSV recipe: drug-likeness screen of active compounds for one target (ChEMBL API + RDKit).

Input : ChEMBL target id (default CHEMBL203 = EGFR), minimum pChEMBL value (default 7), max compounds.
Output: screen_<target>.csv with MW, cLogP, HBD, HBA, TPSA, rotatable bonds, QED and a Lipinski rule-of-five flag.
Original BSV code, MIT. Data: ChEMBL (EMBL-EBI, CC BY-SA 3.0). Library: RDKit (BSD-3-Clause).
This is a filtering exercise, not a prediction of efficacy or safety.
"""
import csv, json, sys, time, urllib.error, urllib.parse, urllib.request
from rdkit import Chem
from rdkit.Chem import Descriptors, Crippen, Lipinski, QED, rdMolDescriptors

def get_json(url: str, tries: int = 4) -> dict:
    for i in range(tries):   # the public ChEMBL API is sometimes slow or returns 5xx; back off and retry
        try:
            with urllib.request.urlopen(url, timeout=180) as r:
                return json.load(r)
        except (urllib.error.HTTPError, urllib.error.URLError, TimeoutError) as e:
            if i == tries - 1 or (isinstance(e, urllib.error.HTTPError) and e.code < 500):
                raise
            time.sleep(5 * (i + 1))

def fetch(target: str, min_p: float, limit: int) -> dict:
    q = urllib.parse.urlencode({"target_chembl_id": target, "pchembl_value__gte": min_p, "limit": 100,
                                "only": "molecule_chembl_id,canonical_smiles,pchembl_value,standard_type"})
    url = f"https://www.ebi.ac.uk/chembl/api/data/activity.json?{q}"
    best = {}
    while url and len(best) < limit:
        page = get_json(url)
        for a in page["activities"]:
            mid, smi = a["molecule_chembl_id"], a["canonical_smiles"]
            if smi and (mid not in best or float(a["pchembl_value"]) > best[mid][1]):
                best[mid] = (smi, float(a["pchembl_value"]), a["standard_type"])
        nxt = page["page_meta"]["next"]
        url = f"https://www.ebi.ac.uk{nxt}" if nxt else None
    return dict(list(best.items())[:limit])

def profile(smi: str) -> dict | None:
    m = Chem.MolFromSmiles(smi)
    if m is None:
        return None
    p = {"mw": Descriptors.MolWt(m), "clogp": Crippen.MolLogP(m), "hbd": Lipinski.NumHDonors(m),
         "hba": Lipinski.NumHAcceptors(m), "tpsa": rdMolDescriptors.CalcTPSA(m),
         "rot_bonds": Lipinski.NumRotatableBonds(m), "qed": QED.qed(m)}
    violations = sum([p["mw"] > 500, p["clogp"] > 5, p["hbd"] > 5, p["hba"] > 10])
    p["ro5_violations"] = violations
    p["ro5_pass"] = violations <= 1   # Lipinski: poor absorption more likely with 2+ violations
    return p

def main(target="CHEMBL203", min_p="7", limit="100"):
    mols = fetch(target, float(min_p), int(limit))
    rows = []
    for mid, (smi, pval, stype) in mols.items():
        p = profile(smi)
        if p:
            rows.append({"molecule": mid, "pchembl": pval, "type": stype, **{k: round(v, 3) if isinstance(v, float) else v for k, v in p.items()}, "smiles": smi})
    rows.sort(key=lambda r: (-r["ro5_pass"], -r["qed"]))
    with open(f"screen_{target}.csv", "w", newline="") as f:
        w = csv.DictWriter(f, fieldnames=list(rows[0].keys())); w.writeheader(); w.writerows(rows)
    passed = sum(r["ro5_pass"] for r in rows)
    print(f"{target}: {len(rows)} compounds profiled, {passed} pass rule-of-five (<=1 violation) -> screen_{target}.csv")
    print("top by QED:", rows[0]["molecule"], rows[0]["qed"])

if __name__ == "__main__":
    main(*sys.argv[1:])
Output of BSV's own test run

Run on 2026-10-06 23:38 JST · Python 3.13.5, rdkit 2026.3.6

CHEMBL203: 60 compounds profiled, 59 pass rule-of-five (<=1 violation) -> screen_CHEMBL203.csv
top by QED: CHEMBL52829 0.795

How far can you trust an AlphaFold model? Confidence per residue

Tested by BSV

Download a predicted structure from AlphaFold DB, read the per-residue pLDDT and summarise it in the database's own four confidence bands, with low-confidence stretches listed.

Input
UniProt accession (default P00533, human EGFR).
Output
afdb_<acc>.pdb, afdb_<acc>_confidence.csv (residue, amino acid, pLDDT, band) and a printed band summary.
Prerequisites
Python 3.10 or later; standard library only.
Steps
  1. Save the code below as afdb_confidence.py.
  2. Run python afdb_confidence.py P00533.
  3. Read the band summary: very high (pLDDT ≥ 90), confident (70–90), low (50–70), very low (< 50).
  4. Open the PDB in a viewer (for example Mol* on rcsb.org) coloured by B-factor; the low stretches are usually flexible or disordered regions.
  5. Before using a region for docking or design, check whether an experimental structure exists in the PDB.
Expected result
A per-residue table and a band summary. For a large multi-domain protein, expect well-predicted domains next to long low-confidence tails — a map of where the model is useful.
Next step
Look up experimental structures of the same protein on RCSB PDB and compare. Open
Code (BSV original, MIT licence)

Download .py · afdb_confidence.py

#!/usr/bin/env python3
"""BSV recipe: how much of a predicted protein structure can you trust? (AlphaFold DB API, pLDDT bands)

Input : UniProt accession (default P00533 = human EGFR).
Output: afdb_<acc>.pdb (downloaded model) and afdb_<acc>_confidence.csv with per-residue pLDDT,
        plus a printed summary of the four AlphaFold DB confidence bands and the low-confidence stretches.
Original BSV code, MIT. Data: AlphaFold Protein Structure Database (DeepMind / EMBL-EBI, CC BY 4.0).
"""
import csv, json, sys, urllib.request

BANDS = [(90, "very high"), (70, "confident"), (50, "low"), (0, "very low")]   # AlphaFold DB's published bands

def band(v: float) -> str:
    return next(name for lo, name in BANDS if v >= lo)

def main(acc="P00533"):
    with urllib.request.urlopen(f"https://alphafold.ebi.ac.uk/api/prediction/{acc}", timeout=60) as r:
        meta = json.load(r)[0]
    pdb_url = meta["pdbUrl"]
    pdb = urllib.request.urlopen(pdb_url, timeout=120).read().decode()
    open(f"afdb_{acc}.pdb", "w").write(pdb)
    res = {}
    for line in pdb.splitlines():
        if line.startswith("ATOM") and line[12:16].strip() == "CA":   # pLDDT is stored in the B-factor column
            res[int(line[22:26])] = (line[17:20], float(line[60:66]))
    with open(f"afdb_{acc}_confidence.csv", "w", newline="") as f:
        w = csv.writer(f); w.writerow(["residue", "aa", "plddt", "band"])
        for i, (aa, v) in sorted(res.items()):
            w.writerow([i, aa, v, band(v)])
    n = len(res)
    print(f"{acc} {meta.get('uniprotDescription', '')} model {meta.get('entryId')} ({n} residues), source {pdb_url}")
    for lo, name in BANDS:
        k = sum(1 for _, v in res.values() if band(v) == name)
        print(f"  {name:9s} {k:5d} residues ({100 * k / n:.1f}%)")
    low, start = [], None
    for i in sorted(res):
        if res[i][1] < 50 and start is None:
            start = i
        if (res[i][1] >= 50 or i == max(res)) and start is not None:
            end = i - 1 if res[i][1] >= 50 else i
            if end - start + 1 >= 10:
                low.append((start, end))
            start = None
    print("  stretches of >=10 residues below pLDDT 50:", ", ".join(f"{a}-{b}" for a, b in low) or "none")

if __name__ == "__main__":
    main(*sys.argv[1:])
Output of BSV's own test run

Run on 2026-10-06 23:36 JST · Python 3.13.5

P00533 Epidermal growth factor receptor model AF-P00533-F1 (1210 residues), source https://alphafold.ebi.ac.uk/files/AF-P00533-F1-model_v6.pdb
  very high   573 residues (47.4%)
  confident   282 residues (23.3%)
  low          79 residues (6.5%)
  very low    276 residues (22.8%)
  stretches of >=10 residues below pLDDT 50: 16-25, 639-649, 682-702, 990-1006, 1019-1210

One-page protein profile from UniProt

Tested by BSV

Fetch a sequence from UniProt and compute length, molecular weight, theoretical pI, hydropathy, instability index and aromaticity with Biopython.

Input
UniProt accession (default P69905, human haemoglobin subunit alpha).
Output
profile_<acc>.md with the six values and the source.
Prerequisites
Python 3.10 or later and pip install biopython.
Steps
  1. Save the code below as uniprot_profile.py.
  2. Run python uniprot_profile.py P69905.
  3. Compare with the same protein on the Expasy ProtParam web tool; the methods are the same, so values should agree closely.
  4. Loop over a list of accessions to build a comparison table.
Expected result
A short Markdown profile. These are sequence-based calculations, not measurements; the instability index in particular is only a rough guide.
Next step
Feed the accession into the AlphaFold recipe above to add structure confidence to the profile. Open
Code (BSV original, MIT licence)

Download .py · uniprot_profile.py

#!/usr/bin/env python3
"""BSV recipe: one-page physico-chemical profile of a protein (UniProt REST + Biopython ProtParam).

Input : UniProt accession (default P69905 = human haemoglobin subunit alpha).
Output: profile_<acc>.md with length, molecular weight, theoretical pI, GRAVY, instability index and aromaticity.
Original BSV code, MIT. Data: UniProt (CC BY 4.0). Library: Biopython.
Values are sequence-based calculations, not lab measurements.
"""
import sys, urllib.request
from Bio.SeqUtils.ProtParam import ProteinAnalysis

def main(acc="P69905"):
    fasta = urllib.request.urlopen(f"https://rest.uniprot.org/uniprotkb/{acc}.fasta", timeout=60).read().decode()
    header, *seq = fasta.strip().splitlines()
    s = "".join(seq)
    pa = ProteinAnalysis(s)
    ii = pa.instability_index()
    lines = [f"# {header[1:]}", "", f"- Length: {len(s)} aa", f"- Molecular weight: {pa.molecular_weight():,.1f} Da",
             f"- Theoretical pI: {pa.isoelectric_point():.2f}", f"- GRAVY (hydropathy): {pa.gravy():.3f}",
             f"- Instability index: {ii:.1f} ({'predicted unstable (>40)' if ii > 40 else 'predicted stable (<=40)'})",
             f"- Aromaticity: {pa.aromaticity():.3f}", "",
             "Computed from sequence with Biopython ProtParam (Expasy ProtParam methods). Source: rest.uniprot.org."]
    open(f"profile_{acc}.md", "w").write("\n".join(lines) + "\n")
    print("\n".join(lines[:9]))

if __name__ == "__main__":
    main(*sys.argv[1:])
Output of BSV's own test run

Run on 2026-10-06 23:37 JST · Python 3.13.5, biopython 1.88

- Length: 142 aa
- Molecular weight: 15,257.4 Da
- Theoretical pI: 8.72
- GRAVY (hydropathy): 0.048
- Instability index: 7.0 (predicted stable (<=40))
- Aromaticity: 0.077

Weekly PubMed digest for your topic, ready for an AI summary

Tested by BSV

Query PubMed through NCBI E-utilities for the last seven days and write a linked Markdown list you can skim or hand to an AI agent.

Input
A PubMed query, days back (default 7), maximum papers (default 20). Set NCBI_EMAIL as NCBI asks; an NCBI API key is optional.
Output
digest_<date>.md: title, journal, publication date and PubMed link for each paper.
Prerequisites
Python 3.10 or later; standard library only. Keep within NCBI's request-rate guidance.
Steps
  1. Save the code below as pubmed_digest.py.
  2. Run NCBI_EMAIL=you@example.com python pubmed_digest.py "AlphaFold protein structure" 7 20.
  3. Skim the list and open the abstracts that matter.
  4. Optional AI step: give the Markdown to a local model and ask for grouping by topic, with the instruction not to add facts.
  5. Schedule it weekly and keep the files: they become your reading log.
Expected result
A dated list of recent papers with working links. Search terms matter: refine the query until the list matches what you actually follow.
Next step
Pull PubChem properties for a common compound name next. Open

The optional AI-summary step was not run by BSV.

Code (BSV original, MIT licence)

Download .py · pubmed_digest.py

#!/usr/bin/env python3
"""BSV recipe: weekly literature digest from PubMed (NCBI E-utilities).

Input : a PubMed query (default: AlphaFold protein structure), days back (default 7), max papers (default 20).
        Set NCBI_EMAIL (and optionally NCBI_API_KEY) as NCBI asks.
Output: digest_<date>.md (title, journal, date, PMID link) ready to hand to a summarising AI agent.
Original BSV code, MIT. Data: NCBI PubMed via E-utilities.
"""
import datetime as dt, json, os, sys, urllib.parse, urllib.request

BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/"

def get(path: str, **params) -> dict:
    params.update({"retmode": "json", "tool": "bsv-recipe", "email": os.environ.get("NCBI_EMAIL", "")})
    if os.environ.get("NCBI_API_KEY"):
        params["api_key"] = os.environ["NCBI_API_KEY"]
    with urllib.request.urlopen(BASE + path + "?" + urllib.parse.urlencode(params), timeout=60) as r:
        return json.load(r)

def main(query="AlphaFold protein structure", days="7", n="20"):
    ids = get("esearch.fcgi", db="pubmed", term=query, reldate=days, datetype="pdat", retmax=n, sort="pub_date")["esearchresult"]["idlist"]
    today = dt.date.today().isoformat()
    out = [f"# PubMed digest: {query}", "", f"Last {days} days, {len(ids)} papers (newest first). Generated {today}.", ""]
    if ids:
        summ = get("esummary.fcgi", db="pubmed", id=",".join(ids))["result"]
        for pmid in ids:
            s = summ[pmid]
            out.append(f"- **{s['title']}** — {s.get('fulljournalname', s.get('source', ''))}, {s.get('pubdate', '')}. "
                       f"https://pubmed.ncbi.nlm.nih.gov/{pmid}/")
    out += ["", "Read the abstracts yourself before relying on any AI summary of this list."]
    open(f"digest_{today}.md", "w").write("\n".join(out) + "\n")
    print(f"{len(ids)} papers -> digest_{today}.md")

if __name__ == "__main__":
    main(*sys.argv[1:])
Output of BSV's own test run

Run on 2026-10-06 23:37 JST · Python 3.13.5

19 papers -> digest_2026-10-06.md

PubChem properties for a compound name

Tested by BSV

Fetch CID, formula, molecular weight, IUPAC name and canonical SMILES into CSV + markdown. Research/education only — not clinical advice.

Input
Compound name (default aspirin).
Output
pubchem_<name>.csv and pubchem_<name>.md.
Prerequisites
Python 3 stdlib + network access to pubchem.ncbi.nlm.nih.gov.
Steps
  1. Save the script.
  2. Run: python pubchem_compound_props.py aspirin
  3. Cite PubChem if you republish; do not use the numbers as care guidance.
Expected result
One or a few CID rows for a common name. Properties are PubChem's.
Next step
Look up a human gene symbol on MyGene.info next. Open
Code (BSV original, MIT licence)

Download .py · pubchem_compound_props.py

#!/usr/bin/env python3
"""BSV recipe: pull basic PubChem compound properties for a common name.

Research / education cheminformatics only. Not efficacy, safety, or dosing advice.
Input : compound name (default aspirin).
Output: pubchem_<name>.csv and pubchem_<name>.md (CID, formula, MW, IUPAC).
Original BSV code, MIT. Data: NCBI PubChem PUG REST (public).
"""
import csv, json, re, sys, urllib.parse, urllib.request

UA = {"User-Agent": "BSV-fields-recipe/1.0 (research education; +https://botshelfvampire.com)"}

def main(name="aspirin"):
    name = (name or "aspirin").strip()
    q = urllib.parse.quote(name)
    url = (f"https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/name/{q}/"
           "property/MolecularWeight,MolecularFormula,IUPACName,CanonicalSMILES/JSON")
    req = urllib.request.Request(url, headers=UA)
    with urllib.request.urlopen(req, timeout=45) as r:
        data = json.loads(r.read().decode("utf-8"))
    props = (data.get("PropertyTable") or {}).get("Properties") or []
    if not props:
        raise SystemExit(f"no PubChem properties for {name!r}")
    safe = re.sub(r"[^A-Za-z0-9._-]+", "_", name)[:40]
    rows = []
    for p in props:
        rows.append({
            "query": name,
            "cid": p.get("CID"),
            "molecular_formula": p.get("MolecularFormula"),
            "molecular_weight": p.get("MolecularWeight"),
            "iupac_name": p.get("IUPACName"),
            "canonical_smiles": p.get("CanonicalSMILES"),
        })
    with open(f"pubchem_{safe}.csv", "w", newline="", encoding="utf-8") as f:
        w = csv.DictWriter(f, fieldnames=list(rows[0].keys()))
        w.writeheader(); w.writerows(rows)
    md = [f"# PubChem properties — {name}", "",
          "**Research/education only. Not clinical advice.**", "",
          f"Hits: {len(rows)}", ""]
    for r in rows:
        md += [f"## CID {r['cid']}", f"- formula: {r['molecular_formula']}",
               f"- MW: {r['molecular_weight']}", f"- IUPAC: {r['iupac_name']}", ""]
    md += ["Source: PubChem PUG REST — https://pubchem.ncbi.nlm.nih.gov/", ""]
    open(f"pubchem_{safe}.md", "w", encoding="utf-8").write("\n".join(md) + "\n")
    print(f"wrote {len(rows)} PubChem row(s) -> pubchem_{safe}.csv + .md")

if __name__ == "__main__":
    main(*sys.argv[1:2])
Output of BSV's own test run

Run on 2026-10-07 02:22 JST · Python 3.13.5

wrote 1 PubChem row(s) -> pubchem_aspirin.csv + .md

MyGene.info lookup for a gene symbol

Tested by BSV

Write mygene_<symbol>.md/.json with name, Entrez id and a truncated public summary. Not clinical advice.

Input
Gene symbol (default TP53) and species (default human).
Output
mygene_<symbol>.md and mygene_<symbol>.json.
Prerequisites
Python 3 stdlib + network access to mygene.info.
Steps
  1. Save the script.
  2. Run: python mygene_symbol_lookup.py TP53 human
  3. Read the full record on NCBI Gene before citing biology claims.
Expected result
At least one hit for common human symbols. Summary text is truncated on purpose.
Next step
Fetch a PDBe entry summary for a public PDB id. Open
Code (BSV original, MIT licence)

Download .py · mygene_symbol_lookup.py

#!/usr/bin/env python3
"""BSV recipe: look up a human gene symbol via MyGene.info.

Research / education only. Summaries are public annotations — not diagnosis or care advice.
Input : gene symbol (default TP53), species (default human).
Output: mygene_<symbol>.json and mygene_<symbol>.md
Original BSV code, MIT. Data: MyGene.info (open gene annotation API).
"""
import json, re, sys, urllib.parse, urllib.request

UA = {"User-Agent": "BSV-fields-recipe/1.0 (research education; +https://botshelfvampire.com)"}

def main(symbol="TP53", species="human"):
    symbol = (symbol or "TP53").strip()
    species = (species or "human").strip()
    q = urllib.parse.urlencode({
        "q": f"symbol:{symbol}", "species": species,
        "fields": "symbol,name,entrezgene,summary,type_of_gene,taxid",
    })
    url = f"https://mygene.info/v3/query?{q}"
    req = urllib.request.Request(url, headers=UA)
    with urllib.request.urlopen(req, timeout=45) as r:
        data = json.loads(r.read().decode("utf-8"))
    hits = data.get("hits") or []
    if not hits:
        raise SystemExit(f"no MyGene hits for {symbol!r} ({species})")
    safe = re.sub(r"[^A-Za-z0-9._-]+", "_", symbol)[:40]
    json.dump(data, open(f"mygene_{safe}.json", "w", encoding="utf-8"), indent=1)
    h0 = hits[0]
    summary = (h0.get("summary") or "")[:800]
    md = [f"# MyGene.info — {h0.get('symbol') or symbol}", "",
          "**Research/education only. Not clinical advice.**", "",
          f"- name: {h0.get('name')}",
          f"- entrezgene: {h0.get('entrezgene')}",
          f"- type: {h0.get('type_of_gene')}",
          f"- taxid: {h0.get('taxid')}",
          f"- hits returned: {len(hits)}", "",
          "## Summary (truncated from API)", summary or "(none)", "",
          f"Source: {url}", ""]
    open(f"mygene_{safe}.md", "w", encoding="utf-8").write("\n".join(md) + "\n")
    print(f"wrote mygene_{safe}.md + .json ({len(hits)} hit(s)) for {symbol!r}")

if __name__ == "__main__":
    main(*sys.argv[1:3])
Output of BSV's own test run

Run on 2026-10-07 02:22 JST · Python 3.13.5

wrote mygene_TP53.md + .json (1 hit(s)) for 'TP53'

PDBe entry summary for one PDB id

Tested by BSV

Download title, method, resolution and dates for one archive entry. Structural literacy only — not a design claim.

Input
PDB id (default 1cbs).
Output
pdbe_<id>.md and pdbe_<id>.json.
Prerequisites
Python 3 stdlib + network access to www.ebi.ac.uk/pdbe.
Steps
  1. Save the script.
  2. Run: python pdbe_entry_summary.py 1cbs
  3. Open the RCSB or PDBe page for the authoritative viewer.
Expected result
A short markdown file plus the raw PDBe JSON.
Next step
List a few STRING interaction neighbors for TP53 (computational scores only). Open
Code (BSV original, MIT licence)

Download .py · pdbe_entry_summary.py

#!/usr/bin/env python3
"""BSV recipe: fetch a PDBe entry summary for one PDB id.

Research / education structural biology only. Experimental metadata as published — not a drug design claim.
Input : PDB id (default 1cbs = cellular retinoic acid-binding protein).
Output: pdbe_<id>.json and pdbe_<id>.md
Original BSV code, MIT. Data: PDBe / EMBL-EBI (PDB archive).
"""
import json, re, sys, urllib.request

UA = {"User-Agent": "BSV-fields-recipe/1.0 (research education; +https://botshelfvampire.com)"}

def main(pdb_id="1cbs"):
    pdb_id = (pdb_id or "1cbs").strip().lower()
    url = f"https://www.ebi.ac.uk/pdbe/api/pdb/entry/summary/{pdb_id}"
    req = urllib.request.Request(url, headers=UA)
    with urllib.request.urlopen(req, timeout=45) as r:
        data = json.loads(r.read().decode("utf-8"))
    entries = data.get(pdb_id) or []
    if not entries:
        raise SystemExit(f"no PDBe summary for {pdb_id!r}")
    e0 = entries[0]
    safe = re.sub(r"[^A-Za-z0-9._-]+", "_", pdb_id)[:40]
    json.dump(data, open(f"pdbe_{safe}.json", "w", encoding="utf-8"), indent=1)
    title = e0.get("title") or ""
    md = [f"# PDBe summary — {pdb_id}", "",
          "**Research/education only.**", "",
          f"- title: {title}",
          f"- experimental method: {e0.get('experimental_method')}",
          f"- resolution_summary: {e0.get('resolution')}",
          f"- entry authors: {', '.join(e0.get('entry_authors') or [])[:200]}",
          f"- deposition date: {e0.get('deposition_date')}",
          f"- release date: {e0.get('release_date')}", "",
          f"Source: {url}",
          f"RCSB mirror: https://www.rcsb.org/structure/{pdb_id.upper()}", ""]
    open(f"pdbe_{safe}.md", "w", encoding="utf-8").write("\n".join(md) + "\n")
    print(f"wrote pdbe_{safe}.md + .json for {pdb_id!r}")

if __name__ == "__main__":
    main(*sys.argv[1:2])
Output of BSV's own test run

Run on 2026-10-07 02:22 JST · Python 3.13.5

wrote pdbe_1cbs.md + .json for '1cbs'

STRING neighbors for one protein symbol

Tested by BSV

Write string_neighbors.csv of scored edges from STRING. Computational evidence only — not wet-lab proof. No pathogen or dual-use protocols.

Input
Identifier (default TP53), NCBI taxid (default 9606), limit 1–20 (default 8).
Output
string_neighbors.csv.
Prerequisites
Python 3 stdlib + network access to string-db.org.
Steps
  1. Save the script.
  2. Run: python string_neighbors_brief.py TP53 9606 8
  3. Cite STRING (CC BY 4.0) if you republish the edge list.
Expected result
A small table of protein–protein edges with combined scores.
Next step
Look up a gene symbol on Ensembl next. Open
Code (BSV original, MIT licence)

Download .py · string_neighbors_brief.py

#!/usr/bin/env python3
"""BSV recipe: list a few STRING DB interaction neighbors for one protein symbol.

Research / education network literacy only. Scores are computational evidence aggregates —
not wet-lab proof and not pathogen or dual-use guidance. Default is human TP53.

Input : protein identifier (default TP53), species NCBI taxid (default 9606), limit (default 8).
Output: string_neighbors.csv
Original BSV code, MIT. Data: STRING Consortium (CC BY 4.0).
"""
import csv, json, sys, urllib.parse, urllib.request

UA = {"User-Agent": "BSV-fields-recipe/1.0 (research education; +https://botshelfvampire.com)"}

def main(ident="TP53", species="9606", limit="8"):
    ident = (ident or "TP53").strip()
    species = str(species or "9606").strip()
    limit = max(1, min(20, int(limit)))
    q = urllib.parse.urlencode({"identifiers": ident, "species": species, "limit": limit})
    url = f"https://string-db.org/api/json/network?{q}"
    req = urllib.request.Request(url, headers=UA)
    with urllib.request.urlopen(req, timeout=60) as r:
        data = json.loads(r.read().decode("utf-8"))
    if not isinstance(data, list):
        raise SystemExit(f"unexpected STRING response for {ident!r}")
    rows = []
    for e in data:
        rows.append({
            "preferredName_A": e.get("preferredName_A"),
            "preferredName_B": e.get("preferredName_B"),
            "score": e.get("score"),
            "nscore": e.get("nscore"),
            "fscore": e.get("fscore"),
            "pscore": e.get("pscore"),
            "ascore": e.get("ascore"),
            "escore": e.get("escore"),
            "dscore": e.get("dscore"),
            "tscore": e.get("tscore"),
        })
    with open("string_neighbors.csv", "w", newline="", encoding="utf-8") as f:
        fields = list(rows[0].keys()) if rows else [
            "preferredName_A", "preferredName_B", "score", "nscore", "fscore",
            "pscore", "ascore", "escore", "dscore", "tscore"]
        w = csv.DictWriter(f, fieldnames=fields)
        w.writeheader(); w.writerows(rows)
    print(f"wrote {len(rows)} STRING edges for {ident!r} (taxid {species}) -> string_neighbors.csv")

if __name__ == "__main__":
    main(*sys.argv[1:4])
Output of BSV's own test run

Run on 2026-10-07 02:22 JST · Python 3.13.5

wrote 17 STRING edges for 'TP53' (taxid 9606) -> string_neighbors.csv

Ensembl gene symbol lookup

Tested by BSV

REST lookup for one human gene symbol into ensembl_gene.md/.json. Research/education only — not clinical genetics.

Input
symbol (default BRCA2).
Output
ensembl_gene.md and ensembl_gene.json.
Prerequisites
Python 3 stdlib + network access to rest.ensembl.org.
Steps
  1. Save the script.
  2. Run: python ensembl_gene_lookup.py BRCA2
  3. Read id / biotype.
Expected result
Markdown + JSON with an Ensembl id when the symbol resolves.
Next step
Search EBI OLS for an ontology term next. Open
Code (BSV original, MIT licence)

Download .py · ensembl_gene_lookup.py

#!/usr/bin/env python3
"""BSV recipe: Ensembl REST lookup for a gene symbol (human).

Research / education only — not clinical genetics.
Input : symbol (default BRCA2).
Output: ensembl_gene.md + ensembl_gene.json.
Original BSV code, MIT. Data: rest.ensembl.org.
"""
import json, sys, urllib.request
UA = {"User-Agent": "BSV-fields-recipe/1.0 (research-education; +https://botshelfvampire.com)", "Content-Type": "application/json"}

def main(symbol="BRCA2"):
    symbol = (symbol or "BRCA2").strip()
    url = f"https://rest.ensembl.org/lookup/symbol/homo_sapiens/{symbol}?content-type=application/json"
    req = urllib.request.Request(url, headers=UA)
    with urllib.request.urlopen(req, timeout=45) as r:
        d = json.loads(r.read().decode())
    with open("ensembl_gene.json", "w", encoding="utf-8") as f:
        json.dump(d, f, indent=2)
    lines = [f"# Ensembl lookup: {symbol}", "",
             f"- id: {d.get('id')}", f"- display_name: {d.get('display_name')}",
             f"- biotype: {d.get('biotype')}", f"- species: {d.get('species')}",
             f"- description: {d.get('description')}", "",
             "_Research/education only. Not clinical advice._", ""]
    open("ensembl_gene.md", "w", encoding="utf-8").write("\n".join(lines))
    print(f"Ensembl {symbol} id={d.get('id')} -> ensembl_gene.md / .json")

if __name__ == "__main__":
    main(*sys.argv[1:2])
Output of BSV's own test run

Run on 2026-10-07 03:20 JST · Python 3.13.5

Ensembl BRCA2 id=ENSG00000139618 -> ensembl_gene.md / .json

Reactome pathway brief for a gene

Tested by BSV

Search Reactome ContentService for pathway hits. Research/education only — not medical advice. No pathogen protocols.

Input
gene symbol (default TP53).
Output
reactome_pathways.csv.
Prerequisites
Python 3 stdlib + network access to reactome.org.
Steps
  1. Save the script.
  2. Run: python reactome_pathway_brief.py TP53
  3. Skim stId / name.
Expected result
Up to a dozen pathway-ish hits.
Next step
NCBI taxonomy brief for a model organism next. Open
Code (BSV original, MIT licence)

Download .py · reactome_pathway_brief.py

#!/usr/bin/env python3
"""BSV recipe: Reactome content service — pathways for a gene symbol.

Research / education only — not a treatment pathway.
Input : gene symbol (default TP53).
Output: reactome_pathways.csv.
Original BSV code, MIT. Data: reactome.org ContentService.
"""
import csv, json, sys, urllib.request
UA = {"User-Agent": "BSV-fields-recipe/1.0 (research-education; +https://botshelfvampire.com)"}

def main(symbol="TP53"):
    symbol = (symbol or "TP53").strip()
    url = f"https://reactome.org/ContentService/data/pathways/top/homo%20sapiens/{symbol}"
    # Actually endpoint: /data/pathways/top/{species}/{gene} may vary — use query by name
    url = f"https://reactome.org/ContentService/search/query?query={symbol}&species=Homo%20sapiens&types=Pathway&cluster=true"
    req = urllib.request.Request(url, headers=UA)
    with urllib.request.urlopen(req, timeout=45) as r:
        d = json.loads(r.read().decode())
    rows = []
    for group in d.get("results") or []:
        for entry in group.get("entries") or []:
            rows.append({"stId": entry.get("stId"), "name": entry.get("name"),
                         "exactType": entry.get("exactType")})
            if len(rows) >= 12:
                break
        if len(rows) >= 12:
            break
    if not rows:
        raise SystemExit("no Reactome pathway hits")
    with open("reactome_pathways.csv", "w", newline="", encoding="utf-8") as f:
        w = csv.DictWriter(f, fieldnames=["stId", "name", "exactType"]); w.writeheader(); w.writerows(rows)
    print(f"wrote {len(rows)} Reactome hits for {symbol!r} -> reactome_pathways.csv")

if __name__ == "__main__":
    main(*sys.argv[1:2])
Output of BSV's own test run

Run on 2026-10-07 03:20 JST · Python 3.13.5

wrote 10 Reactome hits for 'TP53' -> reactome_pathways.csv

NCBI taxonomy brief (model organism)

Tested by BSV

Pull a taxonomy summary for a taxon name into md+json. Research/education only — model-organism literacy; no pathogen or dual-use protocols.

Input
taxon (default Drosophila melanogaster).
Output
ncbi_taxonomy.md and ncbi_taxonomy.json.
Prerequisites
Python 3 stdlib + network access to api.ncbi.nlm.nih.gov.
Steps
  1. Save the script.
  2. Run: python ncbi_taxonomy_brief.py "Drosophila melanogaster"
  3. Read tax_id / rank if present.
Expected result
Markdown + JSON from NCBI Datasets taxonomy.
Next step
Return to the ChEMBL/RDKit screen, or profile another UniProt accession. Open
Code (BSV original, MIT licence)

Download .py · ncbi_taxonomy_brief.py

#!/usr/bin/env python3
"""BSV recipe: NCBI Datasets taxonomy summary for a taxon name.

Research / education only — model organisms / taxonomy literacy; no pathogen protocols.
Input : taxon (default Drosophila melanogaster).
Output: ncbi_taxonomy.md + ncbi_taxonomy.json.
Original BSV code, MIT. Data: api.ncbi.nlm.nih.gov/datasets/v2.
"""
import json, sys, urllib.parse, urllib.request
UA = {"User-Agent": "BSV-fields-recipe/1.0 (research-education; +https://botshelfvampire.com)"}

def main(taxon="Drosophila melanogaster"):
    taxon = (taxon or "Drosophila melanogaster").strip()
    # taxonomy suggest/name endpoint
    q = urllib.parse.quote(taxon)
    url = f"https://api.ncbi.nlm.nih.gov/datasets/v2/taxonomy/taxon/{q}"
    req = urllib.request.Request(url, headers=UA)
    try:
        with urllib.request.urlopen(req, timeout=45) as r:
            d = json.loads(r.read().decode())
    except Exception:
        # fallback taxonomy report by name search
        url2 = f"https://api.ncbi.nlm.nih.gov/datasets/v2/taxonomy/taxon_suggest/{q}"
        req2 = urllib.request.Request(url2, headers=UA)
        with urllib.request.urlopen(req2, timeout=45) as r:
            d = json.loads(r.read().decode())
    with open("ncbi_taxonomy.json", "w", encoding="utf-8") as f:
        json.dump(d, f, indent=2)
    # normalize a few fields if present
    tax = d
    if isinstance(d, dict) and "taxonomy_nodes" in d:
        node = (d.get("taxonomy_nodes") or [{}])[0]
        tax = node.get("taxonomy") or node
    elif isinstance(d, dict) and "sci_name" in str(d):
        pass
    lines = ["# NCBI taxonomy brief", "", f"- query: {taxon}", f"- raw_keys: {', '.join(list(d)[:12]) if isinstance(d, dict) else type(d).__name__}", "",
             "_Research/education only. Not a biosafety manual._", ""]
    # try common fields
    for k in ("tax_id", "organism_name", "common_name", "rank", "species", "genus"):
        if isinstance(tax, dict) and tax.get(k) is not None:
            lines.insert(-3, f"- {k}: {tax.get(k)}")
    open("ncbi_taxonomy.md", "w", encoding="utf-8").write("\n".join(lines))
    print(f"NCBI taxonomy brief for {taxon!r} -> ncbi_taxonomy.md / .json")

if __name__ == "__main__":
    main(*sys.argv[1:2])
Output of BSV's own test run

Run on 2026-10-07 03:20 JST · Python 3.13.5

NCBI taxonomy brief for 'Drosophila melanogaster' -> ncbi_taxonomy.md / .json

Compare the tools

Difficulty is BSV's own rating for a first project. Check each licence on the official page before you ship anything.

ToolJobLicenceDifficultyLocal / cloud
RDKitCheminformatics: descriptors, fingerprints, substructure searchBSD-3-ClauseIntermediateLocal
ChEMBL web servicesBioactivity data for drug-like moleculesCC BY-SA 3.0 (data)BeginnerCloud / web API
PubChem PUG RESTCompound properties by name or CIDPublic (NCBI)BeginnerCloud / web API
MyGene.infoGene annotation query APIPublicBeginnerCloud / web API
PDBe APIPDB entry summaries from EMBL-EBIPublic (PDBe)BeginnerCloud / web API
STRING APIProtein–protein association networks (computational scores)CC BY 4.0BeginnerCloud / web API
AlphaFold DBPredicted protein structures with confidence scoresCC BY 4.0 (data)BeginnerCloud / web API
UniProt REST APIProtein sequences and functional annotationCC BY 4.0 (data)BeginnerCloud / web API
BiopythonSequences, structures and bioinformatics file formatsBiopython License / BSD-3-ClauseBeginnerLocal
NCBI E-utilitiesProgrammatic access to PubMed and other NCBI databasesPublic service (NCBI usage guidelines)BeginnerCloud / web API
RCSB PDBExperimentally determined 3D structuresCC0 1.0 (PDB data)BeginnerCloud / web API

Starter stack: verified sources

Official pages only. BSV opened each link and recorded the HTTP status and date shown. We link out and summarise; we do not copy or rehost their code.

Related on BSV