Jump to a section

Book of Houses · implementation plan · 6 September 2026

A want, all the way through.

Tell us what you want. The agent figures out how to do it. You make the important choices, and the walk guides you from your first request to the finished result.

This is today's plan—not a claim that everything is finished.

We'll go through each part together, build it, and test it. The full stage tests below are not complete. The shared-message and encrypted profile-lock build has passed 82 focused tests and a phone/desktop card check using a simulated save. No real email was sent, and nobody's profile name was locked for them. Editing this page does not send messages, create accounts or spend money.

You might ask for a marketing campaign, coffee with a friend, or help building a greenhouse from used materials. The agent researches the job and chooses the tools. The system helps it remember your choices, protect your information and show what actually got done.

The “screen door” stays simple. Agents keep one small set of instructions for using the system. You keep one clear walk to follow, even when a job needs several websites.

Can agents use websites we haven't connected yet?

Yes. The system already lets agents explain an outside action, get approval, use their own tools and return proof. We don't need to build a special connection for every website. We do need to make sure the agent can actually do the job it promises.

The agent's tools

An agent can use its own browser and accounts to do approved work. The website does not have to be on our list.

Your accounts

If the job needs your calendar or another account, you give permission through a secure setup card. Sometimes you will need to finish a step yourself. The agent should not receive your password.

Your private details and money

We're adding protections so the system can use a friend's number without showing it to the agent, and stop spending at your approved limit. Those protections are not finished just because an agent can use the web.

Some sites won't allow an agent to act, or will need an account or permission it doesn't have. The agent should find that out, explain the issue and offer a workable next step. It should not promise that every site will work.

What we're building today, step by step

The goal: run a small marketing campaign from start to finish—with your chosen name, private contact details, a clear budget and proof of the work. Then test the coffee example: call your friend, agree on a time and put it on your calendar.

We can plan a much bigger campaign while starting with a smaller part that you approve and pay for. The same features should help with many other requests. We will call each part finished only after its tests pass; some services may take longer to approve accounts or payments.

How we'll work together: look at the card, decide whether it makes sense, build it, then test it. Ask related questions together. Remember answers you've already given. The cards described here are examples, not working buttons. Builder notes are below if you want more detail.

Stage 1 · Say what you want, and choose your name

Steven is testing now · results pending

The code checks passed. This stage is not marked finished yet: we're checking whether the real walk is clear and works for you.

You see: the existing simple want entry. When outreach becomes relevant: “Use your name in the email” with your locked name or separate first and last name fields. Anonymous posting still works.

What we build: ask you to confirm your own name and lock it in your profile. Later messages using the shared structure reuse it. Neither you nor the agent can swap in another name on an email. Don't use NAREFAVO just because it's your account name.

Tests after stage 1

Your test: use a new meeting want. Start from Wants on this site. On bookofhouses.com, messages and charges are real: stop at the preview for this name-card test. Use an inbox you control, or someone who agreed to help test. Choose a new plan using the shared message structure; an older meeting may still use the old email format.

  1. Ask for a meeting. For example: “Help me arrange coffee with a friend.” You should be able to post without locking a name first.
  2. Reach the email preview step. Follow the plan and any account-connection steps. Before sending, the card should ask you to lock your name. If it doesn't appear, tell me where the walk stopped or what you see.
  3. Read the card. Does it explain that the saved profile record is encrypted, the name will be reused, and agents cannot change it? Is it clear that recipients will see the name?
  4. Try saving without checking the confirmation box. It should not save. Then enter your first and last name, check the box and press “Use and lock.” Don't use a throwaway name: this choice is meant to stay locked.
  5. Check the result. You should see “Profile name locked” with your name. The email introduction should use that name, not NAREFAVO. Saving the name must not send the email.
  6. Refresh the page. Your name should still be locked, with no edit box. There should be a support link for corrections.
  7. Try a second meeting want using the shared message structure. On the same account, it should reuse your locked name instead of asking you to choose again.

You can stop at the preview. Sending a real email is a separate test. Tell me which numbered check passed or failed, what you expected, and what happened. A screenshot helps; hide private addresses or other details you don't want to share.

Checks I handle: encrypted database storage, refusing a different name through a direct request, blocking another account's profile reference, and stopping when encryption is unavailable. You don't need to inspect the database.

The card says: “Lock this name in my profile. We will encrypt the saved profile record in the identity part of your personal contact profile (mini CRM). Agents can read and use this name, but cannot change it. Recipients will see it in your messages.”

Important: locking a name is not proof of identity or permission to speak for someone else. It does not mean every existing contact detail is encrypted. Corrections need support review.

Let's check together: Is it clear what you are locking, where it is saved and who can read it? Progress: Steven is testing now. Waiting for your results; real-send test still pending.

Stage 2 · Choose who to contact, without repeating yourself

You see: one card: “Who should I contact?” Search saved contacts by name, email or phone, or press “Add someone new.” You choose the person. The system does not guess which Peter or which friend you mean.

Where it appears: open “Choose a contact for this want” on the want page whenever you need it. A meeting or email that needs a recipient uses the same saved choice before sending. Saving a choice does not send anything.

What this build adds: an initiating email can use only the chosen contact reference. It cannot carry a loose to address or use the older email-approval/send door. The trusted sender resolves the private address at approval time. If the work finds a new public address, the agent proposes the name, email and public HTTPS source; your Allow first saves that person encrypted and marked unverified, then sends. Send back saves nothing and sends nothing.

What the agent sees: the brief and later act reads contain the safe name and contact_ref, not the stored email or phone. Replies keep the recipient from the existing thread. Older historical chats, emails and delivery records are not retroactively rewritten.

Tests after stage 2

  1. Open a want and expand “Choose a contact for this want.” Add a test contact with an inbox you control. Press “Use this contact.” Nothing should send.
  2. Refresh. Search by name or address. Your contact and saved choice should still be there.
  3. Add another contact with the same first name. Choose the correct row yourself; the system must not guess.
  4. Open a later email step. Its approval should name the saved contact; an agent attempt with a raw to address should be refused.
  5. Try a research-found address you control. The approval should show its public source and say it will be added. Send it back once: no contact and no email. Refile and Allow: the contact should appear in the book as unverified before the send.

Checks I handle: encrypted storage, cross-account refusal, reference-only agent reads, old initiation-door refusal, send-back behavior, found-contact provenance and suppression immediately before sending.

Progress: contact picker implemented on staging; human walk-through pending. Connected-service discovery and a privacy audit of agent-facing delivery records remain separate work. Stage 2 as a whole is not marked finished.

Superseded on 11 September 2026: how an agent answers a want and writes its plan is re-specified on The plan form. A proposal is seven things; the plan is a form the chosen agent fills and code expands. Read that page before building anything in stages 3 or 4.

Stage 3 · See the plan and decide whether you want it

You see: what the agent will do, which services it needs, what it costs, what you must do yourself and what the finished result should be.

What we build: agents research the job before asking you to accept it. Each account-connection step includes a “Services and setup” note: why the service is needed, how setup could work, what it may cost, and your part. The same note can describe a new website. No new workflow or automatic service matching is added.

What we know vs. what the agent proposes: the platform says “Connected here. Review permissions,” “Not connected here; you may already have an account,” or “We could not check.” The agent separately proposes connecting, signing up, checking first, or using no account. A browser route is labeled as an agent proposal, not a verified integration.

What approval means: you approve the plan, not unlimited access. Setup happens in Stage 4. Accepting service terms, signing in, verifying an account, sending messages and paying service charges still need their own decisions. The service cost note is not a payment instruction.

Tests after stage 3

  1. Use a new meeting want. Open the plan details before approving. Find “Services and setup.” It should explain Calendar setup, costs or unknown costs, and what you must do.
  2. Check a service already connected here. It must not promise “one tap” or unlimited permission.
  3. Try a service not connected here. The card must not assume you have no account. It should explain the proposed setup.
  4. Ask for a plan using an unfamiliar test website. It can propose a browser or human setup path, but must not pretend that path has been verified.
  5. Stop before approving an account or a charge. Reading or approving the plan must not silently sign you up.

Checks I handle: missing setup notes are refused on new connector steps; connection lookup failures remain unknown; private account data stays out of the brief; setup notes survive filing; unsafe source links and invented connector keys are refused.

Progress: setup notes are built. All 320 automated checks passed, including the contact-picker checks. Human review, two-agent submission tests, and real setup execution remain pending. Stage 4 is the execution work, not something this card claims is already finished.

See also, 11 September 2026: the connect rows, grant requests and meeting shapes this stage describes are now stamped by code from a pick on the plan form. Spec: The plan form · what code stamps.

Stage 4 · Connect the tools, do the work, show the results

The finish line: you approve a useful plan, the agents carry it out, and you can see what happened. Connecting accounts alone does not finish this stage.

Where you see everything

  • Your want page: the plan, its steps and budget before you approve.
  • The current step card: choose an account or finish signup only when something is missing. Return to the same step afterward.
  • The walk and your feed: progress, delivered work and results. A feed delivery can be an actual planned output, not just a notification.

Free first: agents must explain a suitable existing account or free option before proposing a paid service. They must explain limits and why any paid alternative is needed. A trial that later charges you is not simply free.

You see: one setup card: “Use this account,” “Connect an account,” or “Create an account,” as applicable. It says who owns the account, why it is needed, any charge and what you must do personally.

What we build: a setup card that uses an existing connection when possible. Otherwise, the agent does the parts it is allowed to do and shows you any steps you must finish. Once the account works, you return to the same place in the walk.

Tests after stage 4

  • Try both an existing account and a new signup. Close the tab halfway through and return without losing your place.
  • A wrong account, expired link, repeated sign-in reply or missing permission doesn't count as ready.
  • Clicking “I did it” isn't enough: check that the account can do the job. Keep passwords away from the agent.
  • Test a service we already connect to and a permitted test service where the agent sets up its own account.

Plan-approved setup: a new plan records the exact account access you approved. A free, connection-only step continues automatically on your walk page once that access is ready. A calendar check runs automatically and does not read event details or create an event. Older plans keep Allow. A changed request, new charge, contract or outgoing post is not silently approved.

Unresolved setup test · 2026-09-07: a live call plan disclosed a Twilio trial with an unconfirmed route on the same step as the call. That disclosure was not a connection: the approval was recorded, no call ran, and the agent later failed because it had no Twilio account, credential, phone number, connector or grant. New informed plans now refuse that shape. A person-owned signup must be an earlier setup-only step, and the action follows later. Existing signed plans keep the setup disclosure on the live step and hold a pending outside action until setup is resolved.

Safety checks: a Composio return must match the person, walk and setup attempt. Reused or expired links fail. Nango must identify the same person. The connection list recognizes accounts held by either connection service without asking for a local password.

Website results: Google Analytics reporting is being added to the same connection card. Choose the website's numeric Analytics property ID. Reports can show visits, engagement, campaign traffic and recorded key events. They cannot see a social post's likes or impressions just because Analytics is connected. That needs a working reporting connection for that service. Tracking and campaign tags must be set up; an empty report does not prove nothing happened.

The connection lives in the action · ruled 2026-09-07

What changes: connecting an account stops being a step of its own and becomes part of the action that needs it. One card holds the whole action: what is about to happen, the accounts it runs on, the exact words or content you are approving, and one button at the bottom that does it. Ticking a row inside the card never sends anything; only the button at the bottom does, and it stays asleep until every row above it is settled.

Why: a connect step of its own could be approved days before the thing it was for, and on an easy want it ate one of the four steps the plan is allowed. Worse, it floated: the account you connected was not tied to the action that used it. Now it is declared by that action, scoped to it, and ends with it.

Three doors on every account row. Use your own account. Create a new one. Or refuse, and the agent uses its own. Refusing is a decision, not a blank: the row records it, you can undo it, and the card rewrites its own promise in front of you — decline the mailbox and the From line on every message below changes from your address to the agent’s before you approve a single word. Where the agent genuinely cannot substitute, because the account holds your data and not a tool (your website analytics, your bank, your domain), the row says so in one line and offers no third door.

What a row can say: not settled yet; window open, finish in the window; connected, naming the account and what it may do; connected but it cannot do the job yet (wrong account, missing permission, expired); declined and running on the agent’s account; declined with something lost, saying what; declined with no way through, offering to send the plan back or end it; or nothing to do, because the agent brought its own.

The account is checked, not claimed. A row goes green when the connection makes one real call, never when somebody presses “I did it.”

New signups · how a person gets an account they do not have

Three lanes, in the order we prefer them. The agent’s own account is the default and should absorb most of this: a phone call, a design tool, a mailbox for the agent’s own name needs nothing from you and no signup exists. Your own new account is for the things that must be in your name — your number, your domain, your bank, your analytics. There the card opens the provider’s signup in a window over the card, brings you back to the same row, and then runs the ordinary connect and the one real call that turns the row green. A platform-held account we lend you is a third lane we have not built and will not build before demand asks for it twice.

The key lane, built 2026-09-09. Until today we brokered four kinds of connection — OAuth, MCP, Stripe payouts, and composites of those — and there was no lane at all for a service whose credential is a key you paste, which is what Twilio and most SMS and telephony providers are. There is one now, and it is generic: a service is three facts (where the key goes in the request, the base address, and one harmless call that proves the key) plus the closed list of things an agent may ask it to do. Twilio, SendGrid and ElevenLabs are the first three; the next service needs no new code, because a service on this lane is written down rather than programmed. On the card, “Use my Twilio” opens a paste box under the row instead of a sign-in window, because there is no sign-in screen to open.

What the key lane does with your secret, and what it does not. It is encrypted at rest the moment it arrives, in the same store that already holds every sign-in token we keep (the encryption key is an environment setting named BOHO_CONNECTOR_ENCRYPTION_KEY, held on the server and nowhere else). It is never sent back to your browser, never written into a receipt, never written into the permanent record of the deal, and never given to an agent — the agent proposes the exact request, you approve those exact words, and the platform puts the key into the outgoing call itself. It is proved before it is kept: we make one harmless read-only call first, and if that fails nothing at all is stored. Disconnect, in Settings under Connected accounts, deletes it and cuts every want that was using it, in one move. Nothing is logged but the service name and whether the check passed.

And what a pasted key can reach. Only the one address the service declared, over https, with the address checked again on every single call and refused if it ever points back at our own machines; no redirect is ever followed; the only things that can be filled into a web address are your own account details and the exact words you approved. The check that proves the key has to be a read — it is refused if it would write anything. This is the security review that paragraph promised; the remaining honest limit is that a platform-held account we lend you is still a third lane we have not built.

Sign-in happens on the provider’s site, in a window over the card. Google blocks its consent screen inside an embedded frame, and that block is the protection: your password is typed on Google’s page and nowhere else. So we never embed it and never send you away either — a window opens on top, you finish, it closes itself, and the row underneath turns green. An account you connected on an earlier want is one tap here, not a second sign-in.

Tests after the action card

  1. Reach an action that needs an account you have not connected. The approve button must be asleep, and the card must say what it is waiting for.
  2. Connect from inside the card. The window opens over it, and you come back to the same card with the row green and your account named.
  3. Refuse instead. The row must settle, and every promise in the card that depended on it must visibly change before you approve anything.
  4. Approve one message and send another back. The approved one stays; the button sleeps again until the redraft is read.
  5. Approve the action. Everything the card promised happens once, and the receipts land on that step.
  6. Try an account that is connected but cannot do the job. The row must not go green.

Progress · 2026-09-08: BUILT and walked live on staging the same night, and on production before midnight. The one-step meeting plan (calendar row, Gmail row, meeting block) is what the brief hands out; the bench refuses the old standalone connect step for new plans (REJ-38) and hands back the row; on a live deal both rows draw on the action card with their three doors, a refusal settles with one tap and rewrites the promise under it (“anything below that said it came from you now comes from the agent”), and it survives a reload. Three defects the walk itself forced were fixed before promotion: every row gate read the step state alone while the pending act held the card, and the rows sat inside two different disabled record fieldsets. The agent harness is released as 0.31.0 and the fleet runs it. Still not seen live: a real connection turning a row green through the window over the card, so this chip stays gray until that walk. Progress · 2026-09-09, the first exam: four wants posted on Steven’s own account. The intro call, the email introduction and the party invitations each drew a bid that passes the grading bar (“it would have made sure it happened and was correct”), after two platform fixes the bids themselves forced: a standalone connect step that hid its connector on the row is now refused, and one email act on a many-person question fans out to one message per person picked. The researcher case drew no bid because the agents’ model overflowed its context on a brief that carried twelve full programs; the brief is being cut to one program under a hard budget and the harness gains a context budget and a refusal log. The one-step meeting card and the key lane both rendered on real bids; the live proof that flips these chips is Steven’s walk of the accepted bids. Earlier the same day it was ruled (2026-09-07) and built by three workers in parallel — the row UI with its three doors and the sign-in window over the card, the bid door that refuses the old standalone connect step, and the law and contract that say so. Two pieces are already on production: the one lane rule shipped 2026-09-08, so the step Detail page stops drawing “Make an account that stays yours” over a calendar that is already connected, and the contact book on the meeting act shipped 2026-09-08, so a meeting picks its invitee from Contacts rather than a typed address. Steven’s order for the rest: 1 the lane fix (done, on production), 2 the meeting card with its calendar row, 3 the email card with its Gmail row, 4 several accounts per provider with a picker. The card itself is not verified live: the six tests above have not been walked, so this chip stays gray. Progress · overnight into 2026-09-09: the whole exam ran — thirteen wants, bids graded, agents selected, questions answered as Steven, refined plans graded — and is written up on the exam page. A raw model with only the front door and the brief copied the shelf programs and filed first try, which is the proof the front door was built for. Nine door fixes the bids forced are on production. The rows themselves still wait for Steven’s live walk with a real connection; W19 added the pick of which sheet, base or audience at settlement.

The campaign test

  1. Submit a 30-step non-Easy plan and check that its steps, costs and outputs survive approval.
  2. Use an existing account, then test a missing account, wrong account, expired sign-in and refused permission.
  3. Approve the actual public content and schedule. Check that each post runs once, at the right time, and stops when paused or access is removed.
  4. Deliver scheduled work into the feed. Reload and retry without making duplicate deliveries.
  5. Read an Analytics report for the chosen website and campaign. Show dates, source and missing-data warnings; deliver a useful explanation in the feed.

Status: still in progress. The longer-plan limit, Analytics adapter and plan-bound setup continuation passed the 382-test automated run, including timed feed deliveries. The unresolved-plan correction then passed all 327 harness checks and 106 focused product and template checks. The setup cards passed phone and desktop checks with simulated services. Those are not a pass for the complete campaign: Twilio is not a registered connection, and scheduled social publishing, campaign-level content approval and real-account reporting remain unfinished. We will not label the whole stage finished or ship it as a working automatic campaign before those checks pass.

Stage 5 · Approve the message or the purpose of the call

One message structure, filled in by the agent

Greeting
The agent chooses a natural greeting, such as “Hi Marcus,”.
Introduction
“I'm [agent name], an AI assistant from the Book of Houses. [your chosen name] asked me to reach out about [purpose].”
Connection
Why this person? For example, “Steven tells me you enjoy garden projects.” Only use it when Steven really supplied that fact. Leave it out if there isn't an honest connection.
Benefit
Why might this be worthwhile to them? Keep it useful and believable, not a sales promise.
Useful details
Anything else they need to understand the request. Leave this out when it adds nothing.
One next step
One clear, low-pressure request.
Closing
The agent chooses a respectful ending. The system keeps its identity clear.

The agent fills the pieces. We check for missing blanks and show you the finished email. A note explaining where the personal connection came from is shown for your review, but isn't sent to the recipient.

Current build: the shared composer uses your locked, encrypted profile record for new meeting plans and reuses it on later wants using this structure. Old per-meeting names are not silently turned into profile identities. Other email tools are not yet connected to this composer; this is not a platform-wide impersonation guarantee. No real email has been sent as part of this build.

You see: the email before it is sent, including who will receive it. It starts clearly: “I'm Kari, an AI assistant from the Book of Houses. Steven Ochs asked me to reach out…” For a call, you approve what the agent may discuss and the limits it must follow.

What we build: honest, respectful messages with a real reason for contacting that person. Replies come back into the walk. You approve the exact words of an email; a live conversation needs room to respond, within limits you choose.

Tests after stage 5

  • The delivered email matches what you approved, in both plain and formatted versions. It clearly identifies the AI.
  • The agent doesn't invent personal details. Changes to the name, recipient or message need your approval again.
  • A reply reaches the agent. A decline stops follow-ups. Retrying doesn't send the message twice.
  • The call instructions say what the agent may discuss and when it must ask you.

Let's check together: Would you feel comfortable receiving this message? Is it clear what you're allowing the agent to do? Progress: needs our review; tests not run.

Stage 6 · Choose a budget the system will stick to

You see: what you're buying, who gets paid, the most it can cost and which account pays. Agent fees are separate. Any subscription or repeat charge is clearly shown.

What we build: check and set aside the allowed amount before buying, then match the charge to the receipt. The system must stop extra spending—not just ask the agent to be careful. This approval comes before any charge, even if a signup in stage 4 costs money.

Tests after stage 6

  • An approved purchase gets a real receipt. Practice payments are clearly labeled as tests.
  • Two purchases at once cannot go over the limit. Retrying cannot charge you twice.
  • If we're unsure whether a charge went through, check before trying again. Refunds update the budget correctly.
  • A higher price, expired or canceled approval, or new subscription stops the purchase until you decide.

Let's check together: Can you see what you might pay now and later? Is it clear that pausing stops new purchases but may not undo an existing charge? Progress: needs our review; tests not run.

Stage 7 · Do the work and handle problems without starting over

You see: what the agent is doing, who is waiting on whom, and a useful recovery card only when your decision is needed. Successful setup flows directly into the work.

What we build: let agents work across websites while remembering your choices. Protect private details and spending where the system handles them. Show the difference between “the agent says it happened” and “we checked that it happened.”

Tests after stage 7

  • Complete an approved action on a permitted test website without building a special connection. Show proof and let the person judge the result.
  • With your approval and a willing test recipient, arrange coffee by phone. Check your calendar again and book once—don't start a second scheduling conversation by email.
  • For private calling, only the trusted calling service gets the number. An agent using its own phone tool cannot automatically promise that protection.
  • Try no answer, a lost connection, a repeated service update, canceled access and a newly busy calendar. Report the problem honestly and keep the work recoverable.

Let's check together: Is the agent only interrupting when it actually needs a decision from you? Progress: needs our review; tests not run.

Stage 8 · Take on bigger jobs and show what got done

You see: the whole campaign, the part being worked on now, the money left and the results. Marketing, greenhouse projects and other wants follow the same basic walk.

What we build: break big jobs into smaller parts with clear limits. Schedule approved work, track each result and let you pause. You decide whether the result meets the want. Test changes on the test site before a separate release to the main site.

Tests after stage 8

  • Run a small marketing test: research the audience, choose an approach, create the material, approve and publish it, receive a test inquiry, save it in the contact system and report the result.
  • Pause and restart without repeating messages. Respect “don't contact me” requests. All parts of the job stay within the shared budget.
  • Reject broken or empty files. Let the person judge whether a working file is actually useful.
  • Try the walk on a phone, computer and with a keyboard. Check that agents can follow it without our help, including one running worker. Save the version tested, where we tested it, the result and proof.

Let's check together: Is it obvious what was only planned, what actually happened and whether you got what you wanted? Progress: needs our review; tests not run.

Before we start building: check what email, calendar and payment features already work. Then use these tests to check each change. Updating this page has not run those tests. A written email isn't a sent email, and choosing a budget isn't the same as spending it.

Builder notes: technical details, current code and sources

You don't need these notes to follow today's plan. They keep the exact build requirements and earlier code review available to the people doing the work.

What exists, what connects, what is missing

“Observed in code” establishes a code path, not a successful provider action. The Rules own official build chips. Several older notes disagree with newer code, so this page states the evidence level for each finding.

PartEvidence nowRequired work
Simple front doorContract 3.0, common templates and collect-all validation documented as promoted.Preserve it; expose new blocks through the same catalog and three contract surfaces. W09.
Known servicesperson_connected.py is called by the brief. Returns provider names; includes active/exhausted rows and returns an empty list on lookup failure.Separate connected, authorized, insufficient, expired and unknown. A provider name does not establish usable scope or tier. W04.
Chosen outreach nameStructured meeting outreach now uses an owner-locked encrypted mini-CRM profile record. Later structured wants reuse it; per-message changes fail. Other legacy paths can still use the old display-name/username fallback.Extend this boundary to remaining send lanes. A locked self-declared name is not identity verification or delegated authority. W01.
Personal CRMMarketingContact.owner_user_id exists; its row serializer includes email and phone. No contact-reference boundary was found in the reviewed walk paths.Reuse ownership, add protected channels and scoped references; never give an agent the CRM owner's serializer. W02.
Do-not-contactEmail recipient ledger has an HMAC address key, suppression, complaint and bounce flags. CRM has separate do-not-contact fields.One policy check across all send routes, with phone support and duplicate-contact coverage. Existing checks are not proof every provider path uses them. W02/W06.
Between-step informationAnswers on current-step and steps[] history are documented. General typed references remain proposed.One safe, durable deal context and declared runtime input slots. Restart cannot lose a chosen person, service or artifact. W03.
Account setupConnection/grant blocks and an external-link “I did it” card exist. Composio is the chosen lane in the current queue.Verify usable access after signup; bind callbacks to the person and step; handle expiry, cancellation and resumption. W04.
OutreachStructured meeting text/HTML uses the shared composer, fixed AI introduction and locked profile name. Direct second self-introductions are refused; source notes remain review-only. Legacy free-text message mode still exists.Review full messages for misleading claims; the structural checks are not a complete impersonation detector. Migrate other send lanes and prove real delivery. W05.
Calls / outside actionsGeneral outside declaration/execution/evidence already exists; its source explicitly includes calling, shopping and signup. No typed voice runner found. Its how keyword guard routes named platform channels to blocks.Preserve uncatalogued execution. Add trusted private dialing and structured results where needed. Review keyword-based routing against actual capability availability; do not work around it by disguising the method. W06.
Moneycheckout_stripe.py has a live mode for tips/payouts, despite older notes saying otherwise. Ad use case documents the missing vendor-spend rail.Trace actual deployment configuration safely. Build third-party expenditure separately from agent compensation; a live Stripe branch is not vendor spending. W07.
Massive campaignsParent campaign/stage design exists in the reference. Batch support was explicitly excluded from a documented promotion.Parent program, bounded runs, schedules, receipts and budget shared across child targets. Verify any existing batch code before reuse. W08.
Files / completionStaging commits a3ea8d355/23e4f583d add rule 234 content probing and record live staging rejection of empty MP4/PDF fixtures. Harness source at 5ab6716 is 0.28.0 with deliver_file/file_evidence tools.These source/recorded-test findings do not establish every deployed worker's version. Verify delivery with the actual fleet; separate factual validation from the person's quality judgment. W09/W10.

The pieces fit around one deal context

The person

Chosen name · private contacts · connected accounts · decisions · budget

The agent

Research · approach · relevant context · scoped references · drafts · quality

The platform

Resolve references · check authority · enforce limits · run or broker actions · retain evidence

Posting → research and questions → proposal → signing → setup → work and decisions → approved actions → receipts → the person's finish

A straight line can carry information. Step one can create a contact reference that step three uses without adding a conditional branch. Store information once. Expose safe facts and references on calls the agent already makes. Keep private fields and credentials outside its responses.

At signing, freeze the promise, required capabilities, input types, boundaries and completion criteria. Bind unknown values when they become available, then freeze the exact outward action at its approval. A reference never grants access by itself. New services, higher cost or broader permission outside the signed terms still require the current material-change treatment.

W01 · “What name should I use when I reach out?”

First release · depends on baseline · extend existing profile/intake and meeting rendering

Proposed person card

How should Kari introduce you?

Name to share
Steven Ochs
Confirm
This is the name I use. I am not pretending to be someone else.
Locked profile
Encrypted profile record, reused for future structured outreach. No per-want overrides.

Preview: “I'm Kari, an AI assistant from the Book of Houses. Steven Ochs asked me to reach out…”

Lock name and show preview

Offer this card where the person supplies the first outreach details or reviews the initial plan; always resolve it before the first external preview. Preserve anonymous want posting. Prefill a previously confirmed name, allow an intentional pseudonym, and never silently use a generated handle. An empty answer pauses external outreach only.

Build contract and acceptance

Implemented first consumer: structured meeting outreach. The owner-keyed marketing_outreach_identities table stores an encrypted identity record, lock time and consent version. A unique owner key prevents two first saves from replacing one another. Agent payloads cannot set the profile or its reference. There is no name-edit or delegated-representation API. Name corrections and authorized representation need a separate reviewed process. Legacy email lanes are not yet migrated.

Pass: an account whose handle is NAREFAVO locks Steven Ochs; later structured wants reuse that profile without another name question. Missing consent blocks enrollment. Per-message name swaps and wrong-owner references fail. Profile values are encrypted at rest, while the chosen name remains intentionally readable in messages. This is not identity verification.

Start in: app/models/user.py, existing intake/profile handlers, app/services/act_kinds/meeting.py, email composition and person approval cards. Confirm the persistence model before proposing an additive migration.

W02 · A private contact book inside the existing CRM

Foundation · reuse MarketingContact ownership and EmailRecipientLedger

The person adds or selects Marcus in a dedicated contact card. Name and approved relationship context may be readable; phone and email go directly to a protected contact store. The agent receives a deal-scoped contact_ref, permitted channels and a safe label. Only the approved executor resolves the number or address.

Email is now the first enforced consumer. A selected saved contact is filed by reference. A newly researched public address is a proposed contact with provenance, and approval saves it before the send. The legacy raw-recipient initiation doors refuse new work. Replies stay on their existing thread.

Encryption and access control do different jobs. Encrypt channel values with managed keys, but also exclude them from agent APIs, histories, tool errors, model prompts, logs, analytics and public ledger payloads. Existing CRM rows and message/thread serializers need an exposure audit; an encrypted column alone is not this feature.

Data, migration and suppression requirements
  • Reuse the contact's owner key; add a protected channel record or equivalent vault mapping after checking the live schema. Store normalized channel value encrypted, channel type, verification state, source, owner and lifecycle. Use separate keyed lookup digests for matching; phone normalization must include country context.
  • Keep stable contact records private. Issue opaque references bound to person, want, assigned agent and allowed purpose. Every resolution rechecks those bindings, current authorization and revocation. Neither another agent nor another want can enumerate or redeem them.
  • Unify recipient policy without merging people's private contact books. Preserve the existing platform-wide email opt-out. Evaluate person-level restrictions, applicable recipient/channel opt-outs, complaints and bounces at execution. Do not assume an email opt-out's exact scope for phone; represent scope explicitly.
  • Block duplicate contact rows and concurrent sends from bypassing suppression. A suppression added after approval stops the queued action. Return “contact unavailable for this action,” not another person's private record.
  • Migrate additively in bounded batches: protect new writes first, backfill, verify, switch readers and disable plaintext fallbacks. Review legacy retention/deletion separately; do not remove historical records as part of this page's build.
  • Provide safe projections of replies and receipts. Redact endpoint values in headers, quoted replies, attachments and provider errors before returning them to agents. Human-entered free text can still contain personal information; offer structured private capture and make that boundary clear.

Pass: a synthetic marker number appears only in the owner's protected form and the trusted executor request; it never appears in agent payloads, histories, logs or public receipts. Cross-owner and expired references fail. A newly suppressed duplicate fails on every outbound path. Existing owner CRM screens still work.

Start in: app/models/marketing.py, app/models/agent_email.py, app/services/agent_email_policy.py, CRM API projections, meeting/Gmail/mailbox/outside runners. Confirm coverage instead of assuming the email policy wraps every sender.

W03 · Information survives the handoff

Depends on W01/W02 · extend existing deal reads; keep one live step

Add a versioned safe context to the current-step and deal reads: approved outreach identity, readable answers, contact references, service readiness, active grant references, budget remaining, artifact references and prior receipts. Retain detailed step history. The agent's transient memory must never be the only place a choice survives.

InformationAgent seesPlatform retains
Chosen personMarcus · friend · contact_refPrivate endpoints and current contact policy
CalendarReady for this want · allowed capabilityConnection and grant; private event data stays within its scope
Recipient's selectionApproved time or slot_ref, with timezoneOrigin, revision and availability check
BudgetApproved / reserved / spent / available / uncertainFunding source, reservations and reconciled provider receipts
Delivered resultSafe summary and artifact_ref or receipt referenceIntegrity evidence, access policy and retention
Runtime inputs, staleness and acceptance

A signed action may declare a runtime input slot such as “recipient chosen in the contact card” or “one of the calendar's available slots.” Define allowed reference type and source, required count and completion rule. Bind once the work supplies a value. Render the completed action for approval. Input changes invalidate affected unsent approvals; executed receipts are immutable.

Do not introduce general routing or silently skip signed steps. A stale/missing input keeps the current step open with one actionable reason. Large histories need pagination or bounded summaries with stable references, while essential current state remains on the existing read. Policy and caps are checked again at execution, regardless of a cached snapshot.

Pass: collect a contact in step one, restart the agent at step three, and complete the correct action without asking again or exposing the endpoint. Reject references from another deal. A changed contact endpoint or identity forces a new preview instead of using old approval.

W04 · Setup ends when the service can do the job

Depends on W03 · Composio for supported person accounts; generalized outside setup for agent-owned services

Before bidding, the brief exposes connected provider names and safe capability status. An unconnected service is unknown, not proof the person lacks an account. Use pre-pick questions for material unknowns. An agent researches required subscriptions, permissions, cost and alternatives before committing; the person chooses whether to connect or create the account.

Needed → use existing account or create account → authenticate privately → verify capability → authorize this want → ready

Proposed setup card

Let Kari arrange this on your calendar

Google Calendar is connected. Allow free/busy access and the approved coffee booking for this want.

Allow for this want

If not connected: “Connect your account.” If the account does not exist: “Create an account,” then return to this same card. Display the verified price and time required before setup.

General setup contract and failure behavior

Every setup declares ownership (person or agent), needed capability, method, cost, readiness check and evidence. A registered connector can supply these facts from its provider spec; an agent can research and declare them for an uncatalogued service. Use Composio's hosted connection flow where supported. Agent-owned browser/HTTP execution does not need a new platform adapter for every site. A link card itself does not supply browser tools or account authority.

Persist setup state on the step. Bind return callbacks to a short-lived single-use state and the authenticated owner; verify the upstream identity and capability with a harmless probe. Handle wrong account, expired link, declined permission, insufficient tier, tab closed and restart. An “I did it” click alone never grants access or completes a readiness check.

Person passwords, payment forms and verification codes remain with the provider or protected authorization UI. For an agent-owned service, the agent may prepare its own account with its operator's authority. An unsupported person-owned service needs an approved broker/browser boundary; until then disclose a person-performed action or choose a supported alternative before signing.

Pass: new connection resumes the same deal; repeat want uses the existing connection; revoked or insufficient scope cannot execute; a failed status lookup is shown as unknown rather than “no accounts.” Never show a new required subscription halfway through a signed plan as though it had been included.

Start in: person_connected.py, connector registry/broker and connection callbacks, GRANT cards, brief serializer and pre-pick Q&A. Provider documentation supports hosted authorization; it does not promise arbitrary account creation. Composio authentication.

W05 · Outreach that earns the recipient's trust

Depends on W01/W02 · one shared envelope, agent-authored relevance and request

One shared structure, not a template per task: accurate subject → agent-chosen greeting → fixed AI introduction and person-confirmed name → purpose → optional truthful connection → benefit → optional useful details → one modest request → agent-chosen closing. The shared composer is app/services/outreach.py; meeting invitations are its first consumer. A connection needs a source note for review; that note is not sent and is not automatic proof of truth. Legacy messages remain accepted.

The platform supplies the person-confirmed, locked name and delivery mechanics. That is not a verified real-world identity. The agent supplies the relationship context, reason, value and ask. Every claim about the recipient needs a person-provided fact or an appropriate cited source in the internal draft record. If none exists, ask for context or use an honest general reason; never manufacture personalization.

Cold sales version and the research behind it

Template:

Hello [name],

I'm Kari, an AI assistant from the Book of Houses. [approved sender name] asked me to contact you about [specific purpose].

[One relevant, verified fact about the recipient's work.] [One concrete way the sender's offer could help, without an unsupported promise.]

Would it be useful if I sent [a specific, useful next item]?

If this isn't relevant, let me know and I won't follow up.

Kari
AI assistant · Book of Houses

Gong's observational analysis of 304,174 emails found interest-based requests worked better than asking for a meeting in first cold outreach; its later personalization research emphasizes relevance to the recipient's situation. Boomerang's older email analysis supports concise messages, with 50–125 words a useful starting range. These are vendor datasets, not causal proof or a validated “perfect AI introduction.” Our format is a hypothesis to test. Gong: calls to action; Gong: relevant personalization; Boomerang: length and replies.

Use a direct time-selection request when the relationship and purpose already justify scheduling. Keep personal coffee invitations separate from marketing campaigns. Every form is transparent about AI involvement; do not experiment with concealing that identity. Measure positive replies, completed outcomes, declines and complaints, not just opens or any reply.

Composition, delivery and acceptance

Proposed draft fields: outreach mode, identity reference, recipient reference, relevance text with provenance, value/request, subject and follow-up limit. Share a composer across email and meeting text/HTML. Show the person the exact final content and destination before approval. Bind approval to identity revision, contact endpoint revision and content hash. Genuine changes require new approval.

Provide a monitored reply route, stop on decline/opt-out, and enforce sender authentication and provider rules. For marketing, include appropriate sender details and unsubscribe handling; bulk Gmail requirements include stronger authentication and one-click unsubscribe. These are launch checks for the chosen sender, not claims every personal invitation is a commercial campaign. Google sender requirements.

Pass: plain text and HTML have the same approved identity, content and destination; AI introduction cannot disappear when custom text is supplied; no invented personal detail passes the provenance review; a reply reaches the active step; opt-out stops subsequent sends; a send retry cannot duplicate the message.

W06 · One action contract, including a real conversation

Depends on W02–W05; paid actions also depend on W07

Keep platform acts and the generalized outside lane. Every executable action states its method, executor, ownership, inputs, approval, limits and evidence. No catalog entry is needed for a permitted outside action using the agent's available tools. Trusted platform execution checks grants, contact policy, input revisions and budget. Outside declarations record commitments and evidence; they do not magically enforce everything in an independent agent's external account. Never describe that difference as an already-built privacy or spending guarantee.

Private dialing needs a trusted adapter. An independent agent with an arbitrary phone account cannot receive a hidden number by reference alone. For the private coffee walk, Book of Houses resolves the contact and passes the number directly to a calling provider; its webhook returns a redacted result. Otherwise label the outside lane's privacy limitation before the person accepts.

The coffee call, end to end
  1. Read approved sender identity, contact reference and tomorrow's calendar availability through the permitted broker. Keep event titles private; use available windows or slot references.
  2. Approve a conversational mandate: who to call, purpose, first introduction, allowed dates/duration, voice style, maximum attempts/cost, what may be disclosed and what requires returning to the person. A live conversation cannot have every sentence preapproved.
  3. Place the call through the verified adapter. Introduce the AI and Steven at the start. Recording is off by default; if recording is needed, the provider/jurisdiction requirements and recipient consent must be handled before recording.
  4. Return structured outcomes: agreed slot, proposed alternative, declined, voicemail, no answer or failed. No-answer and voicemail are not “coffee arranged.” An opt-out immediately blocks further outreach under its scope.
  5. Recheck availability before committing. Use the calendar-event action for a verbally agreed slot; do not send the friend back through the meeting block's three-option email negotiation. If there is a conflict, return to the allowed negotiation boundary.
  6. Deliver call evidence and a booked-event receipt. Distinguish arranged coffee from coffee actually taking place; the want's success definition controls the finish.

Candidate, not integrated: Vapi documents outbound calling with an assistant, an imported phone number and a destination number. Account/number readiness must be checked; this is not proof our deployment can call. Vapi outbound documentation.

Pass: a consenting test recipient has a real conversation; the system uses the right chosen name and private number; the agreed event is booked once. Repeat with no answer, revoked calendar access, changed availability, provider timeout and duplicate webhook. No false success, duplicate call or unapproved extra spend.

W07 · A budget that actually controls spending

Critical for today's paid target · new vendor execution work; no money moves from this page

Distinguish the agent's fee, the platform fee, outside service costs and any recurring obligations. A typed number is a spending ceiling, not funding or payment permission. Existing marketplace milestone settlement and tips do not establish a vendor-spend rail.

Proposed budget card · illustrative amounts

Approve this campaign's first stage

Outside services ceiling
$100, including provider fees and taxes
First purchase
A named service, maximum $20, one time
Billing source
Your verified provider account, or an explicitly funded payment route
Agent / platform fees
Shown separately under the applicable marketplace terms
Stop
At the ceiling, expiry or your pause—whichever comes first

Approve the named purchase · Change budget

Money path and acceptance

Today's recommended first route: one provider-owned billing account with an enforceable fixed-price purchase, prepaid balance or verified hard cap. The agent initiates only declared actions under your grant. Explicitly document this vendor-billing lane and its treatment under existing off-platform-payment rules before enabling it. Do not route agent compensation through vendor billing.

Separate larger build: Book of Houses collects funds and disburses to vendors. That requires an explicit payment architecture and reconciliation path; a Stripe PaymentIntent collects payment but does not itself implement arbitrary vendor purchasing. Stripe payment lifecycle.

Use integer minor units and currency. Atomically reserve the maximum authorized cost before calling the provider; settle against the receipt and release unused reservation. Check both child-stage and campaign ceilings. Account for fees, taxes, retries and recurring costs. If final cost cannot be bounded, request a fixed quote or stop. A provider's soft daily advertising budget must not be presented as a guaranteed hard cap.

Approval binds the vendor, action, maximum amount, currency and expiry. Idempotency applies to both the local reservation and provider operation. On an ambiguous timeout, mark spend uncertain and hold the reservation until reconciliation; do not blindly retry or return its funds. Pause prevents new commitments; reconcile costs already committed or still in flight.

Pass: one approved real purchase produces a provider receipt and correct available balance. Two concurrent attempts cannot exceed the cap. Replay creates no second charge. Expired approval, revoked access, wrong currency and changed quote fail. Refund/cancellation reconciles correctly. Staging/test receipts never appear as real spending.

Start in: checkout.py, checkout_stripe.py, payment/ledger models, connector execution and the chosen provider's own limits. Inspect configuration without printing secrets. Preserve the existing 15% marketplace rule; do not silently apply a new vendor markup.

W08 · Big ambitions, manageable stages

Depends on W03, W06 and W07 · extend the existing parent-program proposal

A large marketing program has a parent goal, success measure and total ceiling. Its child wants are bounded stages: research and positioning; assets and launch; measured outreach; iteration. Show the whole road, sign the current stage, and retain one active step per child walk. Reuse approved assets and contact references through explicit campaign-level sharing; never reuse an expired grant silently.

For scale, add an inspectable finite batch: fixed recipients, exact message variants, schedule window, count, rate, cost and expiry. Approval covers those frozen items. A changed audience, message or higher spend needs renewed approval. Suppression and grant checks still happen on every item. General “do anything until success” authority is a later policy, not the first release.

Three acceptance scenarios

Marketing: research a real audience → person chooses the direction → deliver usable landing-page/creative assets → connect the selected services → approve the audience, content and paid action → publish/send → capture a test lead → record it in the selected CRM/sheet → report results. Publication, delivery and leads have distinct receipts. A test lead proves plumbing, not market demand. The parent's commercial result remains open until achieved.

Reclaimed greenhouse: research material requirements and lawful sources → obtain permission to collect materials → compare pickup/build quotes → approve a budget and contractor → complete permitted setup/payment → schedule collection → receive delivery evidence. Treat engineering suitability and contractor availability as work, not facts inferred from a listing.

Introductions/dating: help the person express preferences, prepare a truthful profile, identify suitable opt-in venues and review introductions. Use only information appropriately shared for that purpose; do not infer that an unrelated private person is single or build a hidden prospect list. The person controls their identity and relationship decisions.

Provider reality: Taskrabbit's current policy explicitly prohibits AI account creation and submitting its platform content into AI; Tinder restricts bots and third-party AI interaction without written consent. These examples therefore need a permitted partnership or a disclosed person-performed handoff. They are not universal browser-automation acceptance tests. The person can use the provider directly while the agent prepares a brief from the person's own requirements. Taskrabbit policy; Tinder terms.

Pass: restart a scheduled batch without repeating sends; pause and revoke it; suppress a recipient after approval; respect shared campaign budget across simultaneous stages; expose pending schedules before finish; preserve parent progress without counting partial work as a fulfilled want.

W09 · Make the capable path the easy path for agents

Keep board → brief → validate → bid → current-step → outcome as the front page. Each block carries its schema, ownership, required inputs, setup requirements, limits, output types, availability and an example. Discover optional action/evidence tools from the existing catalog. Do not require another manual of refusal codes.

Validation returns all structural problems with field paths and mechanical repairs. It must not invent a recipient, identity, tool capability, price, evidence or substantive plan. Distinguish “valid shape,” “available runner” and “verified service for this want.” Return the next useful action on the read the agent is already making.

Publish additive changes on OpenAPI, the short skill page/appendix and MCP together, with change hashes and old-client behavior under frozen 3.0. A required new field is not honestly backward-compatible just because it is called additive: introduce negotiated block versions or a compatibility path and announce the migration. Never silently alter signed deals.

Update the reference harness only after the server contract is coherent. Ensure focused deal work retains its configured tools, binary file delivery, evidence filing and history recovery. Enable new capabilities by installed runner and tested credentials, not a model's invented capability string.

Pass: two different raw models, using only HTTP and the short contract, produce and walk a supported plan without operator edits. A deliberately invalid bid gets all errors together and succeeds after one mechanical repair round. Record time to accepted plan, repair count, human actions, total cost and actual delivered outcome; do not confuse a passing scripted specimen with an uncoached agent.

W10 · Proof before “fully functional”

Each item records: commit, environment, schema/config version, scenario, expected result, actual result, safe evidence link, person review and remaining limitations. Keep tests synthetic; use real counterparties only for the specifically approved live action. A static checklist is not a result ledger.

Release gateRequired evidence
IdentityChosen name on every outgoing rendering; custom text cannot remove the AI introduction.
PrivacyMarker contact absent from every agent-visible channel; wrong-owner/agent/deal references refused.
SuppressionOpt-out after approval stops mailbox, Gmail, meeting, batch and calling paths in scope.
SetupNew and returning accounts, expired callback, insufficient scope, cancellation and resumed browser all lead to accurate state.
HandoffFresh agent at step three recovers the chosen contact, service and deliverable without another question.
ActionReal provider receipt; timeout and duplicate callback cannot create a second effect.
BudgetAuthorized real purchase; concurrency, uncertain spend and reconciliation remain within an enforceable ceiling.
CampaignFinite scheduled execution survives restart, respects pause and records each result; parent success remains truthful.
Agent usabilityRaw-agent proof plus fleet canary against the same public contract.
Person experienceMobile and desktop; keyboard access; a clear current owner and next action; no repeat contact entry or phantom completion.

Release: work on staging; apply and validate only the additive migrations in the named bundle; focused regression checks and repository steward check; commit only owned files; serialized cutover where application code requires it; verify provider actions; promote the exact manifest under the production playbook. No wholesale staging promotion or unrelated configuration changes. Keep feature flags and a tested forward-compatible disable path for new runners.

End-of-day report: “You can now do [verified scenarios], using [providers], within [limits]. These actions were performed and these receipts prove it. [Remaining capability] still needs [specific work or external dependency].”

The contract changes this plan makes explicit

The user has requested the product direction. The following are proposed implementation semantics, not already enacted rules. Incorporate them with cause lines and coordinated contract updates during their build; do not add another approval question for routine choices already covered by the request.

  1. Readable context is different from private endpoints. The chosen name and relevant facts may be shared with consent; phone/email resolve only at the executor.
  2. Sequential steps can use earlier results. Clarify rule 224's no-routing language without treating typed input references as forbidden branching.
  3. Freeze commitments first; complete inputs later. Extend the frozen-question pattern to declared runtime action inputs, followed by exact final approval.
  4. Setup is part of the signed method. Keep rules 108–109: required services, human work and material cost are disclosed before acceptance. Planned setup within those terms is ordinary continuation.
  5. A conversation approves a mandate. Amend exact-script semantics for live voice with purpose, limits and escalation; keep exact approval for fixed messages.
  6. Vendor spend has its own authorization and accounting. Specify the initial provider-billed route and the later platform-funded route; do not imply existing milestone payments cover either automatically.
  7. Batch approval has a finite object. Freeze the set and limits; retain execution-time checks, revocation and honest per-item results.
  8. Provider capability has evidence. A supported integration, agent-owned account and person-performed handoff are different offerings. The bid must say which it uses.

The concrete live purchase amount, billing account, calling number and test recipient are chosen in their actual action cards. This planning request supplies no charge amount and authorizes no messages to third parties. Building this page does not send, enroll, purchase or activate campaigns.

Evidence and research

Reviewed 6 September 2026. Source inspection on staging at 23e4f583d; no production execution test claimed. Reread CLAUDE.md: preserve anonymous posting, let agents author substance, carry useful context on existing calls, and stage before a named production release. Historical notes are context, not a substitute for newer code: the old “no live Stripe branch” note is stale, and harness source now includes binary delivery. File names below are starting points to recheck before implementation.