# Harbury Energy Q&A — App Specification

## What to build

A local web chatbot called the **Harbury Energy Q&A**. It answers questions about home and
community energy — but only from the fact sheets in `facts/`, never from what the model already
"knows." If a question isn't covered by the fact sheets, it must say so clearly instead of
guessing. Every answer that *is* given must show which fact sheet(s) it came from.

This is built in two phases: **index once, then serve.**

## Dependencies and constraints

- Python packages: `flask`, `voyageai`, `anthropic`, `python-dotenv`. Install these first —
  `pip install flask voyageai anthropic python-dotenv` — before writing any code. Beyond those,
  stick to the standard library (`json`, `os`, `math`, `pathlib`).
- **No vector database, and no `numpy`/`scipy`.** The index is small (a few dozen chunks) —
  cosine similarity over a list of floats is a handful of lines of plain Python. Don't reach for
  a library to do the arithmetic.
- API keys (`VOYAGE_API_KEY`, `ANTHROPIC_API_KEY`) are read from environment variables, loaded
  from a git-ignored `.env` via `python-dotenv`. See `.env.example`.
- The Claude model is **not** hardcoded — read it from the `CHAT_MODEL` environment variable
  (also in `.env`), which is set to `claude-sonnet-5`. Use that exact value; don't substitute
  a different model.

## Files to create

| File | Purpose |
|------|---------|
| `build_index.py` | One-time script — reads `facts/*.md`, splits each into chunks, embeds each chunk with Voyage AI, writes `index.json` |
| `app.py` | Flask server — loads `index.json`, serves the chat page, handles questions |
| `templates/chat.html` | The chat UI |
| `index.json` | Generated by `build_index.py` — do not hand-write this |

Do not modify anything in `facts/`.

## Phase 1 — `build_index.py`

1. Read every `.md` file in `facts/`.
2. Split each file into chunks by its `##` section headings (skip the `## Sources` section —
   its URLs become metadata instead, not an answerable chunk). Prefix each chunk's text with
   the file's `#` title so the chunk makes sense out of context.
3. For each chunk, also collect the source URLs listed under that file's `## Sources` section.
4. Call the Voyage AI embeddings API once per chunk to get its vector (batch the calls if the
   library supports it — one file's chunks in one request is fine).
5. Write `index.json`: a list of objects, each `{ "file": ..., "heading": ..., "text": ...,
   "source_urls": [...], "embedding": [...] }`.
6. Print how many chunks were indexed from how many files when done.

Run this once by hand before starting the server. It doesn't need to re-run automatically —
the fact sheets don't change during the session.

## Phase 2 — `app.py`

**`GET /`** — serves `templates/chat.html`: a simple chat interface (message list, a text input,
a send button). No page reload needed — use `fetch()` to call the API and append messages to
the page.

**`POST /api/ask`** — body `{"question": "..."}`. On each call:

1. Embed the question with the Voyage AI API (a query embedding, same model as indexing).
2. Compare it against every chunk's embedding in `index.json` using cosine similarity; take the
   top 3 chunks.
3. Build a context block from those chunks, each labelled with its source file and heading.
4. Call the Claude API (model from `CHAT_MODEL`) with a system prompt that:
   - Instructs it to answer **only** using the supplied context.
   - Instructs it to say plainly that it doesn't have grounded information on that topic if the
     context doesn't cover the question — never fill the gap from general knowledge. This
     refusal is the whole point of the app, not optional styling; without it you have a chatbot
     with extra steps, not a RAG app.
   - Instructs it to name which fact sheet(s) it drew on. Every answer must cite its source(s),
     even a short one.
5. Return JSON: `{"answer": "...", "used": [{"file": ..., "heading": ...}], "sources": [urls
   from the chunks actually used]}`.

`app.py` only **reads** `index.json` at startup — it must never re-embed the fact sheets on each
request. Indexing (Phase 1) happens once, by hand, before the server starts.

If a Voyage AI or Claude API call fails (bad key, network, rate limit), catch it and return a
clear JSON error the chat UI can display — e.g.
`{"error": "Couldn't reach the embeddings API — check VOYAGE_API_KEY"}` — not a raw 500 stack
trace in the browser.

The chat UI shows the answer, then a small "Sources" line listing the fact-sheet file(s) and
their URLs, styled distinctly from the answer text (e.g. smaller, muted).

## What "grounded" looks like in practice

Two questions to try once it's running, to see the whole point of the exercise:

- **One the library covers** — e.g. *"What does the Boiler Upgrade Scheme pay for?"* — should
  get a specific answer with a source citation.
- **One it doesn't** — e.g. *"Is a battery subscription better value than buying my own solar
  panels outright?"* (a judgement call the fact sheets don't make) — should get a clear "I don't
  have grounded information to answer that" instead of an invented comparison.

If the second one gets a confident-sounding made-up answer, the grounding isn't working yet —
that's the bug to chase, not a nice-to-have.
