# Research Desk — distributable local workflow (experimental / runnable kit)

> **Answer quality: 0 / 9 accepted** (2026-09-11 battery). Workflow can run; answers are not reliable. Earlier PASS = limited workflow checks only.

> **Status:** Experimental local workflow — answer-quality evaluation failed.  
> The September 11 review found unsupported publication approvals, false claim attributions and missed numerical inconsistencies. The earlier PASS refers to limited workflow checks, not answer reliability. This kit is for controlled experiments, not unattended research, publication approval or trading decisions. Outputs require checking against the original sources. No external action was executed in these tests. This is sequential local Ollama use, not simultaneous collaboration between cloud AI services.

**Not a SaaS product. Not a registration/payment flow.**  
This folder alone is the kit. Everything you need is relative to this directory.

**Honesty line:** sequential local Ollama-tag workflow ≠ automatic simultaneous multi-cloud.

**Synthetic disclaimer:** Verified=0 / Untested in the fixture sources ≠ a live Library registry claim.

## Prerequisites

- Linux/macOS shell
- Python 3.10+
- [Ollama](https://ollama.com) listening on `http://127.0.0.1:11434` **OR** willingness to use the manual handoff path below
- Model tags (create from included Modelfiles if missing):
  - `botshelf-deep-research-3b` — see `models/Modelfile.botshelf-deep-research-3b`
  - `botshelf-doc-review` — see `models/Modelfile.botshelf-doc-review`
- Disk write access for an output directory under this kit (e.g. `./out/`)
- **No** BotShelf account, payment, or cloud AI seat required for this local path

### Create models (if tags missing)

```bash
ollama pull llama3.2
ollama pull llama3.2:1b
ollama create botshelf-deep-research-3b -f models/Modelfile.botshelf-deep-research-3b
ollama create botshelf-doc-review -f models/Modelfile.botshelf-doc-review
```

Record digests on **your** host (do not reuse another host’s digests as proof):

```bash
ollama show botshelf-deep-research-3b
ollama show botshelf-doc-review
# or: curl -s http://127.0.0.1:11434/api/tags
```

Proof-host digests (reference only): `models/MODEL-TAGS.json`.  
If Ollama is unavailable → mark tags **UNAVAILABLE** and use **Manual handoff** below. Do not invent ChatGPT/Claude output.

## Synthetic input

- `input/synthetic-sources.json` — small synthetic/permissioned sources (safe to share)
- `input/ai1-prompt.txt` — AI-1 structure prompt bound to those sources
- `input/restrictions.txt` — carried-forward fixture restrictions (Monitoring Untested; data-analysis NOT cleared; WAITING_ON_HUMAN)
- Optional: replace with your own public paste; rebuild **both** agents’ inputs from the new source set. Do not pair a new source JSON with the old embedded synthetic AI-1 prompt. Do not include secrets.

## Run commands (automated sequential Ollama path)

```bash
set -eu

# 0) Confirm Ollama (skip if UNAVAILABLE → Manual handoff)
curl -s http://127.0.0.1:11434/api/tags | head

# 1) Output dirs
OUT=./out/research-desk-$(date -u +%Y%m%d)
mkdir -p "$OUT/handoffs" "$OUT/input" "$OUT/raw" "$OUT/corrected"
cp input/synthetic-sources.json input/ai1-prompt.txt input/restrictions.txt "$OUT/input/"

# 2) Init envelope
python3 bin/runner.py init \
  --run-dir "$OUT/handoffs" \
  --task-id RD-LOCAL-001 \
  --objective "Local Research Desk sequential two-model run" \
  --first-agent ollama:botshelf-deep-research-3b

# 3) AI-1 (Ollama) — RAW output (exits nonzero on transport/incomplete/empty failure)
python3 bin/runner.py ollama \
  --model botshelf-deep-research-3b \
  --prompt-file "$OUT/input/ai1-prompt.txt" \
  --out "$OUT/raw/02-ai1-structure.md"

# 4) Build AI-2 prompt from ORIGINAL sources + AI-1 draft + restrictions
python3 bin/runner.py build-review \
  --sources-file "$OUT/input/synthetic-sources.json" \
  --draft-file "$OUT/raw/02-ai1-structure.md" \
  --restrictions-file input/restrictions.txt \
  --out "$OUT/input/ai2-prompt.txt"

# 5) AI-2 (distinct Ollama model tag) — RAW critique (do not run if step 3 failed)
python3 bin/runner.py ollama \
  --model botshelf-doc-review \
  --prompt-file "$OUT/input/ai2-prompt.txt" \
  --out "$OUT/raw/04-ai2-critique.md"

# 6) Handoff envelopes (example)
python3 bin/runner.py handoff \
  --run-dir "$OUT/handoffs" --task-id RD-LOCAL-001 --step 2 \
  --from-agent ollama:botshelf-deep-research-3b \
  --to-agent ollama:botshelf-doc-review \
  --objective "AI-2 critique" --status WAITING_ON_AGENT \
  --next-step "run doc-review" --model-tag botshelf-deep-research-3b

# 7) Human writes CORRECTED final (required — model output stays UNREVIEWED)
#    Copy/edit into: $OUT/corrected/FINAL-research-brief.md
#    Review state remains WAITING_ON_HUMAN until an actual human edit/approval.
```

`set -eu` stops the pipeline after a failed `ollama` / `build-review` step. Successful model text is labeled `MODEL_OUTPUT_RECEIVED_UNREVIEWED` — never auto COMPLETE / factual PASS.

## Manual handoff steps (when Ollama tags UNAVAILABLE)

Label this path clearly in your RESULT: `MANUAL_HANDOFF` / `OLLAMA_UNAVAILABLE`.

1. Open `input/ai1-prompt.txt` + `input/synthetic-sources.json` in any chat UI (seat A).
2. Save seat A reply to `$OUT/raw/02-ai1-structure.md`. Record provider/model string honestly (or `unknown`).
3. **Seat B must receive all three:** `input/synthetic-sources.json` (original labeled sources) + `$OUT/raw/02-ai1-structure.md` (draft) + `input/restrictions.txt` (carried-forward restrictions). Do **not** paste the draft alone. Prefer generating the combined prompt with:
   ```bash
   python3 bin/runner.py build-review \
     --sources-file input/synthetic-sources.json \
     --draft-file "$OUT/raw/02-ai1-structure.md" \
     --restrictions-file input/restrictions.txt \
     --out "$OUT/input/ai2-prompt.txt"
   ```
   Then paste `$OUT/input/ai2-prompt.txt` into seat B.
4. Save seat B reply to `$OUT/raw/04-ai2-critique.md`.
5. Human-edit → `$OUT/corrected/FINAL-research-brief.md` (state `AWAITING_HUMAN_REVIEW` until actually human-edited).
6. Do **not** claim simultaneous multi-cloud automation.

## Output paths

| Kind | Path |
|---|---|
| Raw AI-1 | `$OUT/raw/02-ai1-structure.md` (+ `.meta.json` if runner used) |
| Built AI-2 prompt | `$OUT/input/ai2-prompt.txt` |
| Raw AI-2 | `$OUT/raw/04-ai2-critique.md` (+ `.meta.json`) |
| Envelopes | `$OUT/handoffs/*.json` |
| **Corrected final** | `$OUT/corrected/FINAL-research-brief.md` |

Sample illustrations (redacted) — see `sample-output/EVIDENCE-INDEX.md`:

- `sample-output/ai1-structure.REDACTED.md`
- `sample-output/ai2-critique.FAILED.REDACTED.md` — historical FAILED (Publish); **not accepted**
- `sample-output/ai2-critique.RERUN.REDACTED.md` — re-run; still not authority
- `sample-output/ai2-prompt.BUILT.txt` — example `build-review` output
- `sample-output/CONFLICTS.json` / `sample-output/EVIDENCE-INDEX.md`
- `sample-output/corrected/FINAL-research-brief.COORDINATOR-CORRECTED.AWAITING-HUMAN.md`
- `sample-output/handoffs/*.json` (host paths redacted)

## Handoffs

Sequential only (never claim simultaneous). Envelope fields are documented in `bin/runner.py` (`write_envelope`).

## Stop condition

Stop when:

- both raw model outputs + metas exist **and** corrected final is written with human review, OR
- any model missing / transport failure / incomplete generation → label `BLOCKED` / `FAILED` / `UNAVAILABLE` honestly (runner exits nonzero; do not invent cloud-provider output; do not run AI-2 after AI-1 failure)

## Known limits

- Site report pages are **proof/case-study pages**, not this runnable kit.
- One Ollama host with two model tags ≠ two cloud providers.
- Digests in `models/MODEL-TAGS.json` are from a proof host; re-record on your machine before citing them as your evidence.
- Small models (esp. 1b-class doc-review) can soft-invent “safer” claims — prefer source-grounded bullets + human edit.
- Sample outputs under `sample-output/` are synthetic/redacted illustrations.
- Original Phase-1 Research Desk used one Ollama model + process labels (`web-fetch-curl`, `local-file-audit`, `executor-synthesis`).
- Runner transport tests (`bin/test_runner_transport.py`) are **mocked-transport exit-status** checks — not live Ollama model-quality proof.

## Layout

```
README.md
index.html
bin/runner.py
bin/test_runner_transport.py
input/           synthetic-sources.json, ai1-prompt.txt, restrictions.txt
models/          Modelfiles + MODEL-TAGS.json
sample-output/   failed / rerun / corrected + CONFLICTS + evidence index
out/             (created by you; gitignored if you add one)
```

## Download

A downloadable ZIP of this kit (allowlisted public contents only) is linked from `index.html` as `/cross-ai/kits/research-desk-local.zip`.
