What we mean by deliberation

Models answer in one pass. We're interested in systems that can hold a problem, revisit it and notice when they're wrong.

Ask a model a hard question and it answers immediately, with the same confidence whether it's right or not. Reasoning models have improved this — they think longer before answering — but the thinking still happens in one sitting, and it's gone when the session ends.

People don't solve hard problems that way. We draft, leave it, come back, notice the flaw, and try again with a memory of what didn't work. That loop — hold, revisit, revise — is what we mean by deliberation.

Three questions we're exploring

Self-critique with memory. Can a model pause, critique its own draft and continue, keeping a record of what failed so it doesn't repeat it?

A scratchpad that persists. A working space that survives across sessions — hypotheses, dead ends, open threads. Closer to a lab notebook than a chat transcript.

Where it should live. Is deliberation something people should see and steer, or infrastructure that quietly makes other tools more careful? We don't know yet.

Why memory comes first

Deliberation over time needs memory. A model can't revisit a problem it can't recall, or learn from a mistake it has no record of. That's why our first applied work is on memory, and why Murmur exists. Deliberation is the longer question it leads to.

This is early, exploratory work, and we'd rather say that plainly than dress it up. When we have something worth showing, it will be here.