# Repeatable molecular workflow

Use [molecule_pipeline.py](molecule_pipeline.py) for a new dataset; see
[PIPELINE.md](PIPELINE.md). It automatically enumerates unspecified
stereochemistry and collects local, hosted and private Colab predictions into a
shared versioned schema. The lower-level commands below also preserve the
original study workflows; their individual defaults do not replace the general
pipeline's automatic enumeration policy.

Use the project `.venv` or `uv run --frozen`. Preparation is entirely local and uses CPU chemistry tools. Neural inference requires MPS, or Metal through llama.cpp for LlaSMol and ChemDFM, and should run outside the Codex sandbox on this Mac.

## Prepare any dataset

```sh
uv run --frozen prepare_molecules.py --input my_molecules.csv --output runs/my_batch/prepared
uv run --frozen verify_prepared.py runs/my_batch/prepared
```

CSV and TSV accept `id,smiles` columns; select other names with `--id-column` and `--smiles-column`. Also supported: SMI/TXT lines (`SMILES identifier`), JSON lists of records, MOL and SDF. Missing IDs receive `mol_1`, `mol_2`, etc.; duplicate IDs are rejected. Invalid SMILES and undefined atoms such as `R`/`*` are saved in `invalid_records.json` and cause a nonzero exit. Duplicate structures with different IDs are allowed.

Each prepared dataset contains:

- `structures.json` with original and canonical isomeric SMILES, Kekulé SMILES, InChIKey, charge, formula and unspecified stereo elements.
- `inference_inputs.csv`, suitable for the local adapters.
- `equivalent_smiles.csv`: canonical plus up to 16 unique alternate atom-order spellings. Every spelling must round-trip to the exact same canonical molecular graph and stereochemistry. Small molecules may have fewer alternatives.
- Per-molecule SVG and 224/600 px PNG drawings, 2D MOL files and atom/bond JSON graphs.
- A Morgan radius-2, 2048-bit chiral fingerprint matrix with the exact ID order. This convenience fingerprint does not replace any model's trained featurizer.
- A browsable `index.html` and a `manifest.json` containing settings, source/script hashes, RDKit version and output hashes.

Charge, salts/fragments, isotopes, atom maps and specified stereochemistry are retained. Canonicalization does not neutralize, choose tautomers or fill in missing stereochemistry. Model-specific preprocessing belongs in each adapter. The same graph drawn differently supplies a different representation, not new experimental information.

## Optional 3D and stereoisomer candidates

```sh
uv run --frozen prepare_molecules.py --input my_molecules.csv --output runs/my_3d/prepared --conformers 10
uv run --frozen prepare_molecules.py --input my_molecules.csv --output runs/my_stereo/prepared --conformers 10 --stereo-policy enumerate --max-isomers 64
```

The default `preserve` policy skips 3D generation when stereochemistry is undefined. `enumerate` explicitly creates candidate stereoisomers without altering specified centers; their SDF records are marked as hypotheses. Enumeration exceeding the configured limit is reported, not silently truncated. ETKDGv3 conformers use a recorded seed and RMS pruning; optimization uses MMFF94s, with an explicitly recorded UFF fallback when available. Failures and unconverged optimizations remain visible. Generated geometry and force-field energy are not observed conformations, populations or binding affinities.

`--stereo-policy enumerate` also writes `stereoisomer_candidates.csv`, even without 3D generation. This file is accepted directly by the preparation and inference CLIs as a separate hypothesis dataset. It never replaces the original unresolved input.

Changing `--random-smiles`, `--seed`, `--image-sizes`, `--conformers` or `--max-isomers` creates an inspectable new run. Output directories must be empty to prevent stale files from being mistaken for current results. `verify_prepared.py` checks hashes, graph/stereo roundtrips, fingerprint ordering and every generated SDF record.

## Run the installed local model suite

The focused predictor accepts additional food molecules and retains the trained activity definitions:

```sh
.venv/bin/python predict_function.py --input my_molecules.csv --output runs/my_function --device mps
# Optional formal completions of unspecified stereo, explicitly labelled:
.venv/bin/python predict_function.py --input my_molecules.csv --output runs/my_function_stereo --device mps --enumerate-stereo
```

The CSV requires unique `id` and `smiles` columns. The default runs the three fitted constituents and their validation-selected combination, preserving all scores and nearest measured-reference records. It uses the training pipeline's MolVS fragment-parent and canonical isomeric graph rules; this model-specific standardization is recorded separately from the representation-preparation workflow. More than 64 formal stereo completions is rejected, not silently truncated. The panel covers 17 measured activity questions across eight proteins; it is not a universal effect predictor. Use the saved chemical-coverage diagnostics and concentration thresholds when interpreting it.

The broader suite remains available for deliberate comparison:

```sh
uv run --frozen run_dataset.py --input my_molecules.csv --output runs/my_inference
# Continue a run after repairing a failed stage:
uv run --frozen run_dataset.py --input my_molecules.csv --output runs/my_inference --resume
# List or run selected stages:
uv run --frozen run_dataset.py --list-stages
uv run --frozen run_dataset.py --input my_molecules.csv --output runs/my_herg --stages bayesherg_and_mt_tox,cardiotox
```

Use a fresh output directory inside this project for `run_dataset.py`; it creates its own `prepared` directory. Inputs may be anywhere readable. This command requires the already downloaded assets for the selected stages. The default runs the whole suite; `--stages` selects a subset, and `--resume` can add more stages later. It validates all inputs, runs the adapters sequentially on the GPU, records per-stage commands/logs, and exports every completed prediction. It makes no hosted submissions or identity lookups. A failed stage stops the pipeline and keeps its diagnostics. Resume checks input hashes, column selections, preparation integrity, hashes of recursively imported local code, and saved output hashes before reusing successful stages. Changed code or missing/changed outputs trigger a rerun, and each attempt receives a separate log. Older stages without recorded hashes are rerun once.

For completed native outputs, rebuild reports without inference:

```sh
uv run --frozen export_dataset.py --results runs/my_inference/results --structures runs/my_inference/prepared/structures.json --output runs/my_inference/report
```

Reports contain per-model native and normalized predictions, model notes, settings, source hashes, all-input summaries and large long/wide tables. Scores remain on their native endpoint scales. Notes about the original known controls are identified as reference checks, not accuracy estimates on a new batch.

To combine completed batches without repeating inference:

```sh
uv run --frozen combine_datasets.py --dataset runs/batch_a/results runs/batch_a/prepared/structures.json --dataset runs/batch_b/results runs/batch_b/prepared/structures.json --output runs/combined
```

Input IDs must be unique across batches, and both batches must have identical completed model/endpoint panels. The combined report retains each constituent run's settings and input hash.

The suite includes PIDGIN v5's four concentration panels and TargetNet's seven original fingerprint panels. To run only those:

```sh
uv run --frozen run_dataset.py --list-stages
uv run --frozen run_dataset.py --input my_molecules.csv --output runs/my_target_panels --stages pidgin,targetnet
```

Their adapters reproduce the trained fingerprint definitions; do not substitute the preparation workflow's convenience fingerprint. Numerical MPS/native-reference parity checks are retained with each run. TargetNet's hosted ECFP4 control values match the local implementation to the server's rounding, so the two are not independent evidence.

## Private Colab and Google Drive

The private project is [food_molecules_prediction](https://drive.google.com/drive/folders/1mON-2yhPX7_RqjJIjMD6CRwpq8aon_X3). No sharing permissions were changed. All completed study structures are authorized in this private project. Public prediction services still receive only the known controls.

- [00: Prepare molecules](https://colab.research.google.com/drive/130ovlHJ7IK0cKl78L6aRxCwgWa-NnXJx) accepts new input files and creates drawings, fingerprints, alternate SMILES and optional 3D candidates. Its code cells were executed successfully in the managed Colab runtime.
- [01: TxGemma 27B](https://colab.research.google.com/drive/1KBij9-lQMoNu5xndy8m79pjqsWSW5Leb) uses the pinned author release, full bfloat16 and an 80 GB A100. Select no examples or ten retrieved examples, the saved study or a custom CSV, and a distinct run name. The ten-example mode covers hERG plus 12 Tox21 tasks, not every task in the 671-endpoint no-example run.
- [02: Endpoint training preparation](https://colab.research.google.com/drive/1-0DbrR2S2eXYB62fBt47Ohg7aBvtMbw7) prepares a locked split from real measured labels. The notebook prepares data for a user-defined training task.
- [03: Fitted function models and training](https://colab.research.google.com/drive/1op7k1OFeF1RIS9KDQBmu2f9ciHbJOuE7) runs our trained combination on a custom CSV, generates repeatable representations, and saves a private result archive. Both its default inference path and `RUN_TRAINING=True` path were executed top to bottom. Training uses the bundled measured labels and locked split. The portable bundle contains weights, data, scripts and their hashes, without credentials.

The same deterministic TxGemma input construction is available without a notebook:

```sh
uv run --frozen prepare_txgemma.py --input my_molecules.csv --output runs/my_txgemma/requests --no-include-benchmarks
uv run --frozen prepare_txgemma_fewshot.py --requests runs/my_txgemma/requests/requests.jsonl --output runs/my_txgemma/ten_examples
```

The latter script accepts explicit `--tox21-pool` and `--herg-pool` paths. It excludes all query parents from the example pools, removes conflicting labels per parent/task, and saves every selected reference, similarity, label and source hash. These are actual public measurements supplied as context to this specific model. This is distinct from merely generating another drawing or text spelling.

`txgemma_colab.py` records the author prompt, actual generated text, answer-letter score, formatting recovery if needed, checkpoint revision, GPU and progress. It refuses silent truncation and resumes by request hash. A regression answer on the author's normalized 000–1000 scale is not converted to physical units without the original transform. Downloaded complete runs can be evaluated with `evaluate_txgemma.py RESULTS_FOLDER --name RUN_NAME` against the saved matched control and public panels.

`build_colabs.py` regenerates the notebooks and an explicit upload manifest. `cloud_project.py --manifest output/jupyter-notebook/drive_upload_manifest.json` updates those files in the existing private Drive folder. Credentials remain in the existing private authorization store; never place them in a notebook, artifact bundle or manifest.

For notebook 03, `package_function_project.py` builds the deterministic bundle and `build_function_notebook.py` inserts its checksum into the notebook. `cloud_project.py --manifest results/function_followup/portable/upload.json` updates the explicit private files. The executed notebook, per-cell status, complete log and numerical outputs are preserved under `results/function_followup/notebook_validation_v3/` and `results/function_followup/notebook_training_validation/`. `analyze_function_training_repeat.py` compares the repeated fit without replacing the primary models.

`prepare_target_context.py` obtains public human protein sequences for a selected target panel and writes explicit protein/ligand input files. These choose a biological question; they do not discover the correct target. `boltz_colab.py` provides a separate Python 3.12 environment for a bounded Boltz-2 control check. Its MSA request contains only a public protein sequence, with no ligand structure. Structure and affinity inference run on the private GPU. The saved manifest distinguishes full canonical protein sequences from experimentally chosen constructs.

## Generated 3D activity comparison

`extract_unimol_features.py` accepts unique `structure_id,smiles` rows containing the same standardized canonical inputs as the activity panel. It records seeded ETKDGv3/MMFF geometry, specified-stereo checks, any excluded structures and native 512-dimensional Uni-Mol features. All atoms are retained; structures exceeding the supported size or lacking valid geometry are explicitly excluded. Its model weights are pinned in `sources/function_followup/unimol/weights_manifest.json`.

```sh
.venv/bin/python extract_unimol_features.py --input my_canonical_structures.csv --output runs/my_unimol_features --device mps
.venv/bin/python predict_unimol_function.py --features runs/my_unimol_features --output runs/my_unimol_scores --device mps
```

The feature extractor preserves the input spelling's atom order, so inputs should first pass the activity panel's canonicalization. Missing stereo can receive one generated assignment; the SDF and metadata retain that assignment without changing the input identity. For an explicit sensitivity study, supply the enumerated candidates and repeat extraction with `--seed 42`, `43` and `44`, saving each run separately. Multiple `--features` arguments combine those runs into complete per-conformation scores and ranges.

`prepare_unimol_classification.py` retains the existing measured labels and scaffold splits for structures with valid geometry. `package_unimol_experiment.py` and `function_colab_unimol_pipeline.py` run matched fingerprint, graph and 3D activity heads in the private project. `prepare_unimol_comparators.py`, `analyze_function_models.py` and `analyze_unimol_comparison.py` preserve the identical held-out comparison. The complete study results are under `results/function_followup/unimol_analysis/`; this 3D configuration has lower predictive accuracy than the graph panel.

## Rebuild the reader-facing study

```sh
uv run --frozen run_all.py --reports-only
uv run --frozen build_explorer.py
```

Open `results/study/index.html`. The five views separate effects, evidence and disagreement, model choice, terminology, and complete outputs. `results/study/all_methods_comparison.csv` includes local, hosted and private-Colab results; missing hosted unknown results are marked as not submitted. The local-only long/wide tables remain in `results/study/all_molecules/`.

Short explanations and top predictions expand inside the overview; multiple cards can remain open. Complete result tables also expand inside the results list. Separate method links support browser Back/Forward, middle-click and direct URLs that retain the selected molecule. Switching views preserves open cards and searches; the molecule selector keeps the current scroll position.

After rebuilding, `uv run --frozen verify_study.py` checks saved data, source hashes and cloud-run completeness and refreshes the source snapshot. `node verify_explorer.mjs` and `node verify_navigation.mjs` run headless Chrome checks using the installed Playwright package. Their screenshots and verification receipts are saved under `output/playwright/`.

Saved PharmMapper results can be reparsed without contacting the service or creating another job:

```sh
uv run --frozen parse_pharmmapper.py results/hosted/pharmmapper
```

The parser retains the original ordering and separately exports all three score rankings, including score ties. Hosted service output is a comparator, not an additional measurement.

## Extract source drawings from Word

```sh
uv run --frozen extract_document_images.py --input 'my_structures.docx' --output runs/my_document_images
```

This retains the original embedded images and extracts native ChemDraw CDX streams when present. EMF previews use `emf2svg-conv` and `rsvg-convert` when installed (`libemf2svg` and `librsvg` via Homebrew). Extraction preserves artwork; it does not infer a unique molecule from an undefined R group or add missing stereochemistry. Inspect the generated gallery and keep any image-to-structure transcription review with the dataset.

## Search the local reference library

```sh
uv run --frozen fetch_chembl_database.py --structures-only
uv run --frozen chembl_neighbors.py build --workers 6
uv run --frozen chembl_neighbors.py search --input runs/my_batch/prepared/structures.json --output runs/my_batch/neighbors --neighbors 30
uv run --frozen chembl_neighbors.py evidence --output runs/my_batch/neighbors
```

The search uses the complete downloaded ChEMBL 37 structure catalogue and a recorded Morgan fingerprint protocol. It preserves exact-match, exact-parent-excluded, and scaffold-excluded results. Query structures stay local. The optional `evidence` stage requests existing measurements by the already-public ChEMBL IDs of the saved reference neighbors; it does not send query SMILES or run a hosted predictor. This is a transparent similarity baseline, not a replica of SwissTargetPrediction.

## Recreate the environment

`uv.lock` pins the Python packages. For the optional LlaSMol/ChemDFM Metal backend on macOS:

```sh
CMAKE_ARGS='-DGGML_METAL=on' CMAKE_BUILD_PARALLEL_LEVEL=4 uv sync --frozen --extra metal-llm
```

The verified llama.cpp source archive and its PyPI hash are also retained in `sources/llama_cpp/`. Model assets and pinned author source files are recorded in their manifests under `models/` and `sources/`. The fetch scripts retrieve those public releases; preparation alone does not need model downloads. Checkpoint licenses are retained per model.

OPERA's descriptor stage also needs Java 8. On another machine, set `OPERA_JAVA` to the absolute path of the Java executable; the adapter records the executable and original descriptor-jar hashes. Java descriptor generation is CPU work; the ported regression calculations use MPS.

The optional CardioTox native-reference audit uses a separate environment so historical TensorFlow dependencies do not change the inference environment:

```sh
uv venv .venv-tensorflow --python 3.11
uv pip install --python .venv-tensorflow/bin/python -r requirements-tensorflow-audit.txt
.venv-tensorflow/bin/python verify_cardiotox_tensorflow.py
```

This checks every converted numerical array against the original TensorFlow checkpoints and compares native Keras inference with the saved MPS predictions on identical saved features. It does not rerun chemistry preprocessing or replace the MPS model workflow. `--results` accepts one or more directories containing `cardiotox_features.npz` and `cardiotox.csv`.


## Standalone report

The report has a separate static distribution. It contains all interactive prediction values, scientific reference tables, model explanations and structure images. CSV exports are generated from the stored values, so the same large prediction tables need not be included twice.

```sh
.venv/bin/python build_explorer.py
.venv/bin/python build_report_package.py
.venv/bin/python verify_report_package.py
node verify_report_package.mjs
```

The directory is `dist/food_molecules_report/`; its ZIP is `dist/food_molecules_report.zip`. It works when opened directly or served beneath a directory prefix. Upload the directory to a static host such as Cloudflare Pages without a build command. Package creation and validation do not publish anything.

`control_report.py` assembles the source-linked known-control dossiers during `build_explorer.py`. It joins saved experimental records, broad-effect checks, native target ranks and per-assay Tox21 predictions without rerunning inference. Its source hashes and complete data are retained in `results/study/known_controls.json`. The browser checks cover separate molecule and whole-study navigation, the cross-molecule comparison, control dossiers and inline scientific details.

`control_literature.py` contains the curated study interpretations and distinguishes primary observations from measurements transcribed through reviews. `method_review.py` combines those examples with task-matched benchmarks, method availability and training priorities. Its dedicated global `#reliability` view includes every retained method; `results/study/method_reliability.csv` and `method_review.json` export the same analysis. `node verify_reliability.mjs` checks that view, source-linked control dossiers, filtering, native links and responsive layout.

`package_pipeline_source.py` creates a deterministic source ZIP and checksum manifest from the project scripts, report assets, schema, documentation and pinned environment. Model assets and experimental results remain in their separately versioned archives.

`package_method_review.py` packages the PRnet, BioTransformer, AmIActive and admetSAR evaluations, cached literature and candidate assay data, and unified-run receipts. It excludes model weights and session credentials, preserves input lineage, and verifies every archived file against its checksum. The research archive is separate from the static report.

For a larger package containing the additional native research downloads, use `build_report_package.py --include-research-downloads --output dist/food_molecules_report_with_archives`. The compact package retains the same displayed model values.

`REPORT_GUIDELINES.md` governs scientific wording and navigation. `results/storage/cleanup_receipt.json` records the approved storage cleanup; original outputs and restoration metadata remain in the workspace.

### Cloudflare Pages publication

`publish_report.py` verifies the package manifest and the existing Cloudflare
login. Its default operation is read-only. It reuses the `cf` OAuth login for
Wrangler's Pages uploader without saving a token in the project.

```sh
.venv/bin/python publish_report.py
.venv/bin/python publish_report.py --deploy
```

The deployment command publishes the complete package, including the unknown
structures, to the `food-molecules-prediction` Pages project. It updates that
project on subsequent runs and leaves other projects alone. The deployment
receipt is `results/report_package/cloudflare_deployment.json`; it records the
public URL, deployment ID and package checksum. Rebuild and verify the package
before publishing subsequent report edits.

`method_control_comparisons.py` assembles each method's separate andrographolide
and berberine comparisons from saved outputs and curated experiments. It keeps
the prediction, reported result, interpretation and sources together, including
measurement incompatibilities and cases without a matching experiment.
