Drug-likeness screen of known actives with ChEMBL and RDKit
Tested by BSVPull potent compounds for one target from ChEMBL, compute standard descriptors and QED with RDKit, and rank them with a rule-of-five flag.
- Input
- ChEMBL target id (default CHEMBL203, EGFR), minimum pChEMBL value (default 7), number of compounds.ChEMBLの標的ID(既定はCHEMBL203=EGFR)、pChEMBL値の下限(既定7)、化合物の数。
- Output
- screen_<target>.csv: MW, cLogP, H-bond donors and acceptors, TPSA, rotatable bonds, QED, rule-of-five violations and pass flag, SMILES.screen_<標的>.csv:分子量、cLogP、水素結合ドナー・アクセプター数、TPSA、回転可能結合数、QED、Rule of Fiveの違反数と判定、SMILES。
- Prerequisites
- Python 3.10 or later and
pip install rdkit. The public ChEMBL API can be slow; the script retries.Python 3.10以上とpip install rdkit。公開されているChEMBLのAPIは遅いことがあるため、スクリプトは自動で再試行します。 - Steps
-
- Save the code below as chembl_screen.py.
- Run
python chembl_screen.py CHEMBL203 7 60. - Open the CSV, sort by QED and look at the structures of the top rows (paste a SMILES into any structure viewer).
- Change the target id to one you study; find ids by searching the ChEMBL website.
- Keep the CSV with the date: ChEMBL releases change the data.
- 下のコードを chembl_screen.py として保存します。
python chembl_screen.py CHEMBL203 7 60を実行します。- CSVを開いてQEDで並べ替え、上位の構造を見てみます(SMILESを構造ビューアに貼り付けます)。
- 標的IDを、ご自身が扱っているものに変えます。IDはChEMBLのサイトで検索すると見つかります。
- CSVは日付つきで保存してください。ChEMBLは版が変わるとデータも変わります。
- Expected result
- Most potent, published actives pass the rule of five; the useful part is seeing which ones do not and why. Rule of five and QED are rough filters, not predictions.公表されている強い活性化合物の多くはRule of Fiveを満たします。役に立つのは、満たさないものがどれで、なぜかを見ることです。Rule of FiveもQEDも大まかなふるいで、予測ではありません。
- Next step
- Check the target's predicted structure confidence before any docking idea (the next recipe uses the same target, EGFR, UniProt P00533).ドッキングなどを考える前に、標的の予測構造の信頼度を確かめてください(次のレシピでは同じEGFR、UniProt P00533を使います)。 Open
Code (BSV original, MIT licence)
Download .py · chembl_screen.py
#!/usr/bin/env python3
"""BSV recipe: drug-likeness screen of active compounds for one target (ChEMBL API + RDKit).
Input : ChEMBL target id (default CHEMBL203 = EGFR), minimum pChEMBL value (default 7), max compounds.
Output: screen_<target>.csv with MW, cLogP, HBD, HBA, TPSA, rotatable bonds, QED and a Lipinski rule-of-five flag.
Original BSV code, MIT. Data: ChEMBL (EMBL-EBI, CC BY-SA 3.0). Library: RDKit (BSD-3-Clause).
This is a filtering exercise, not a prediction of efficacy or safety.
"""
import csv, json, sys, time, urllib.error, urllib.parse, urllib.request
from rdkit import Chem
from rdkit.Chem import Descriptors, Crippen, Lipinski, QED, rdMolDescriptors
def get_json(url: str, tries: int = 4) -> dict:
for i in range(tries): # the public ChEMBL API is sometimes slow or returns 5xx; back off and retry
try:
with urllib.request.urlopen(url, timeout=180) as r:
return json.load(r)
except (urllib.error.HTTPError, urllib.error.URLError, TimeoutError) as e:
if i == tries - 1 or (isinstance(e, urllib.error.HTTPError) and e.code < 500):
raise
time.sleep(5 * (i + 1))
def fetch(target: str, min_p: float, limit: int) -> dict:
q = urllib.parse.urlencode({"target_chembl_id": target, "pchembl_value__gte": min_p, "limit": 100,
"only": "molecule_chembl_id,canonical_smiles,pchembl_value,standard_type"})
url = f"https://www.ebi.ac.uk/chembl/api/data/activity.json?{q}"
best = {}
while url and len(best) < limit:
page = get_json(url)
for a in page["activities"]:
mid, smi = a["molecule_chembl_id"], a["canonical_smiles"]
if smi and (mid not in best or float(a["pchembl_value"]) > best[mid][1]):
best[mid] = (smi, float(a["pchembl_value"]), a["standard_type"])
nxt = page["page_meta"]["next"]
url = f"https://www.ebi.ac.uk{nxt}" if nxt else None
return dict(list(best.items())[:limit])
def profile(smi: str) -> dict | None:
m = Chem.MolFromSmiles(smi)
if m is None:
return None
p = {"mw": Descriptors.MolWt(m), "clogp": Crippen.MolLogP(m), "hbd": Lipinski.NumHDonors(m),
"hba": Lipinski.NumHAcceptors(m), "tpsa": rdMolDescriptors.CalcTPSA(m),
"rot_bonds": Lipinski.NumRotatableBonds(m), "qed": QED.qed(m)}
violations = sum([p["mw"] > 500, p["clogp"] > 5, p["hbd"] > 5, p["hba"] > 10])
p["ro5_violations"] = violations
p["ro5_pass"] = violations <= 1 # Lipinski: poor absorption more likely with 2+ violations
return p
def main(target="CHEMBL203", min_p="7", limit="100"):
mols = fetch(target, float(min_p), int(limit))
rows = []
for mid, (smi, pval, stype) in mols.items():
p = profile(smi)
if p:
rows.append({"molecule": mid, "pchembl": pval, "type": stype, **{k: round(v, 3) if isinstance(v, float) else v for k, v in p.items()}, "smiles": smi})
rows.sort(key=lambda r: (-r["ro5_pass"], -r["qed"]))
with open(f"screen_{target}.csv", "w", newline="") as f:
w = csv.DictWriter(f, fieldnames=list(rows[0].keys())); w.writeheader(); w.writerows(rows)
passed = sum(r["ro5_pass"] for r in rows)
print(f"{target}: {len(rows)} compounds profiled, {passed} pass rule-of-five (<=1 violation) -> screen_{target}.csv")
print("top by QED:", rows[0]["molecule"], rows[0]["qed"])
if __name__ == "__main__":
main(*sys.argv[1:])
Output of BSV's own test run
Run on 2026-10-06 23:38 JST · Python 3.13.5, rdkit 2026.3.6
CHEMBL203: 60 compounds profiled, 59 pass rule-of-five (<=1 violation) -> screen_CHEMBL203.csv
top by QED: CHEMBL52829 0.795