Grounded Answers — Retrieval Augmented Generation

Tonight you're going to commission a real chatbot from an agent — one that answers home and community energy questions, but only from a small library of fact sheets you're given. Ask it something the library covers and it'll answer with a source. Ask it something the library doesn't cover — even something that sounds reasonable — and it should tell you it doesn't know, instead of making something up.

That's the whole lesson: an agent that's grounded vs an agent that's guessing.

You're writing the brief. The agent writes the code. You run it, then try to catch it out.

Your folder: energy-qa — set it up in step 1 below.


1 — Set up your folder

Create a new folder called energy-qa. Download the fact library and the three brief files:

Open the energy-qa folder in VS Code, then tell the agent to read its rulebook:

read CLAUDE.md

It should confirm it's read it and describe what it understood.


2 — Look at what you're indexing

Before building anything, get the agent to look at the material:

read the files in facts/ and tell me, for each one, what kind of question it could answer

Check its summary makes sense against what you can see in the fact sheets. This is also your first read of the material you're about to ground a chatbot in.


3 — Commission the app

build the Harbury Energy Q&A app as described in @spec.md

Or you could use the /to-issues skill first, as you've learned to do already.

It will show you a plan and wait for approval — check it makes sense (indexing first, then a server skeleton you confirm loads, then the rest), then say go.

First checkpoint: it will run build_index.py and show you how many chunks it indexed from how many files. If that number looks way too low (like 1 or 2), something's wrong — ask it to show you what it split each file into before moving on.

Second checkpoint: it will build a minimal server and ask you to open it in your browser before building the actual chat logic. Confirm you see a placeholder page, then let it continue.


4 — Run it and ask it something it knows

If it's not already running you can ask Claude to run it, or type this in the terminal:

python app.py

Open the address it prints. Ask it something the fact sheets clearly cover, for example:

What does the Boiler Upgrade Scheme pay for?

You should get a specific answer, plus a line showing which fact sheet it came from.


5 — Now try to break it

This is the actual point of tonight. Ask it something that sounds like it's in scope but isn't answered by any of the five fact sheets — a judgement call, a comparison, a prediction. For example:

Is a no-upfront-cost battery subscription better value than buying solar panels outright?

Watch what happens:

  • If it says something like "I don't have grounded information to answer that" — it's working!
  • If it gives you a confident, specific-sounding answer — it's hallucinating past its sources, and that's a bug. Get Claude to help you with that, eg:
that answer wasn't grounded in any of the fact sheets — the system prompt needs to refuse
to answer when the retrieved chunks don't cover the question. Fix that and let me try again.

Things to try once it's working

  • "Add a follow-up-question suggestion under each answer, based on nearby fact sheets."
  • "Show me the actual similarity scores for the chunks it retrieved, for debugging."
  • "What happens if I ask the same question two different ways — does it retrieve the same chunks both times?"