AI Agents
Ralph
The signal that a relationship is going cold shows up as silence, and silence means something different for every contact. The data to see it coming is in the CRM and nobody has time to cross-reference it.
- My role
- Owned the agent graph, the approval boundary, both surfaces, and the data model beneath them
- Maturity
- Deployed and live with a public demo. Test suite verified 2026-09-20: 192 unit and contract tests passing. Re-checked on schema or prompt changes.
- Decision
- The agent pauses at the tool boundary and offers three responses, not two. Enforcement lives in the graph router, not in the prompt.
- Boundary
- No email, campaign, or Slack message leaves the system without a person approving it. Three of four tools are gated.
3 of 4
Tools, Gated
192
Tests Passing
3
Review Responses
4 views
Agent SQL Surface
Built for Riverton Capital Advisory, for Sidera. I designed and built the agent graph, the API, both front-end surfaces, the database schema, and the brand context system, as a vertical AI product reference implementation.
The operating problem
A PE advisory firm's pipeline is relationships, not tickets. The signal that one is going cold shows up as silence.
But silence does not mean the same thing twice. A Managing Director who has not responded in ten days is a signal. A Portfolio CFO who has not responded in ten days is Tuesday. Every contact has a different cadence, and the threshold for "going dark" is different for each of them.
The data to see it coming is sitting in the CRM. Engagement dates, deal stages, persona, last touch. Reading it means someone opening reports and cross-referencing them on a morning when there are eleven other things to do.
Handing that to an agent is easy. Handing it to an agent that can also send email on your behalf is where it gets uncomfortable, and that discomfort is the correct instinct.
The system decisions
Five choices shaped this, and each one closed off something simpler. A regression suite built later tested every one of them, which is the only reason I can say which held.
A state machine, not a prompt chain or an n8n workflow. Both were available and both are less work. Neither fits, because Ralph's execution path depends on what the data says. Ask about cold contacts and it runs a query. Ask it to draft outreach and it creates a campaign, generates personalised emails, and queues them for review. A prompt chain cannot branch on its own findings. n8n can orchestrate a fixed flow but cannot reason mid-flow. LangGraph gives a graph where the agent decides which tools to call and the graph decides what requires permission.
Three responses at the gate, not two. Most human-in-the-loop implementations are binary: approve or cancel. That is a bad fit for drafted outreach, because the useful answer is usually neither. The subject line is wrong but the timing is right. The recipient is correct but the tone is off for that persona.
| Response | What happens |
|---|---|
| Continue | Approve as-is. The tool call executes |
| Update | Edit the arguments, for example the email subject, then execute |
| Feedback | Reject with a note. The agent receives the note and reasons again |
The third option is what makes it usable, and it is also the expensive one. Declining a draft means returning a ToolMessage for every pending tool call, not only the rejected one, because the API requires it. The cheap version of this feature is two buttons.
The gate in the router, not in the prompt. "The model usually asks first" is what a prompt instruction buys, and it is not a guarantee. assistant_router checks the pending call against state.protected_tools and sends it to human_tool_review_node before anything executes. A prompt can be argued with. A routing condition cannot. This is the decision that most needed testing and the one that held: the audit could not construct a path around it.
Four denormalized views, not the seven real tables. Giving an agent the actual schema is the obvious move and it fails. Pointed at normalized tables, Ralph invents column names. So it never sees them. It gets views that expose business questions instead of storage, including v_at_risk_with_open_deals, which filters to Cold, Dormant and At Risk, joins active deals and sorts by engagement score ascending. The payoff is visible as an absence: no hallucinated-column defect appears anywhere in the findings, because the decision removed the category.
Two surfaces on one data model, not one chat window. The person checking pipeline health at eight in the morning and the person drafting re-engagement at two in the afternoon are doing different jobs on the same data. So the Morning Briefing is a read-only dashboard with no write path at all. That also narrowed the security work later: when the time came to authenticate write endpoints, only one surface had any.
How the system works
Ralph is a LangGraph graph. Conversation runs through it, tools query Supabase for pipeline and engagement data, and the agent reasons about which relationships are at risk.
When the graph reaches a protected tool, human_tool_review_node calls LangGraph's interrupt() primitive and execution suspends. FastAPI streams the interrupt to the front end over SSE as a typed event, and the Next.js interface renders it as an approval card. The person answers, the resume endpoint takes the response, and the graph continues from where it stopped.
The Morning Briefing is the second surface: pipeline health score, relationship velocity with seven-day sparklines, a persona and engagement heat map, deal-risk flags, and recent campaign activity. Read-only, no conversation required, same database.
Brand voice is not a runtime configuration layer, and an earlier version of this page said it was. The voice rules, messaging pillars, banned phrases and CTA library are roughly 47 lines of literal text inside the system prompt at src/ralph/prompts.py, lines 127 to 174 of 260. The seventeen documents in brand/riverton/ are the source material that section was written from, by hand.
Adapting Ralph to a second client means authoring a new prompt section from that client's brand documents. A day of work, not a directory swap.
Operating boundary
Ralph has four tools. Three of them are protected.
| Tool | Gated |
|---|---|
query | No. Reads run immediately |
create_campaign | Yes |
send_campaign_email | Yes |
slack_post_message | Yes |
The system prompt reinforces the gate from the other side: Ralph is explicitly forbidden from writing its own approval language, no "Shall I proceed?", because the approval dialog fires from the graph and the agent does not manage that flow.
One deliberate bypass exists. yolo_mode is a state field, defaulting to false, that routes protected calls straight to execution. It is documented in state.py as development and testing only. I am naming it because it is in the code and because a system that claims no bypass exists is making a claim the reader can check. The honest version: the bypass is off by default, it is a per-run state field rather than an environment variable, and nothing in the deployed path sets it.
The briefing dashboard is read-only end to end and has no write path.
Evidence and verification
| Check | State |
|---|---|
| Public demo | marketing-intelligence-eight.vercel.app |
| Production deployment | Next.js on Vercel, FastAPI on Render, PostgreSQL on Supabase |
| Enforcement point | assistant_router checks state.protected_tools before dispatch, not the prompt |
| Tool surface | 4 tools, 3 protected |
| Agent SQL surface | 4 denormalized views, not the 7 raw tables |
| System prompt | 260 lines, assembled by hand from 17 brand documents |
| Streaming | SSE with typed events: chunk, interrupt, done, error |
| Test suite | 192 unit and contract tests passing, 2026-09-20. 55 integration tests require a local Postgres and were not run in that pass |
| Regression coverage | The interrupt payload, the SSE event contract, protected-tool routing, read-only enforcement, and the health-score formula each have dedicated test modules |
| Adversarial audit | docs/claims-verification.md cross-checks every claim on this page against the code at a named commit |
Designed behavior, and what the numbers are
Two numbers get called scoring here and they are different things.
canon_engagement_score is hand-authored per contact in the demo seed: thirty-eight contacts with a chosen score, status and day count, shaped to match the output of a nightly scorer that runs elsewhere in a separate Sidera project. In the demo it is a fixture, not a calculation, and it is labelled that way here because a reader who opens db/seed_b2b_data.py will find a literal dict.
The pipeline health score is computed:
base = sum(engagement_score x enterprise_value) / sum(enterprise_value)
penalty = mean((days_silent - 30) / 7) over contacts silent more than 30 days
score = clamp(base - penalty, 0, 100)
Enterprise-value weighting so a quiet ten-million-dollar relationship outranks a quiet small one. A thirty-day threshold before silence counts against anyone. One point of penalty per seven days past it.
Those three constants were chosen from how the firm actually works, not fitted to an outcome set. That is a deliberate position rather than an omission: with thirty to forty active relationships there is no sample to fit against, and a number tuned on a handful of deals would be a worse guide than a stated rule anyone can argue with.
Cache windows run fifteen minutes to four hours by panel. Three of the five are busted immediately when a protected tool is approved, so an approved send appears at once rather than waiting out its TTL. Health score and heat map read canon_fields, which no protected tool writes, so those wait the full two hours. The dashboard is a recent picture and should not be called real-time.
Limitations and what remains unverified
Scoring has not been tested against outcomes, and it structurally cannot be yet: there is no scores_history table, so the system cannot compare today's health score to yesterday's. I can say it surfaces quiet relationships. I cannot say those relationships closed at a different rate.
Portability to a second client is a design property, not a demonstrated one. The prompt rewrite path is clear and has not been walked.
The frontend has a test runner and two test files, and the approval modal's three resume actions were verified through a real browser against a stub backend rather than by tests. The in-memory cache is per-process and resets on restart, which is correct on one Render instance and wrong the moment it scales past one.
What the system demonstrates
The gate is architectural. It lives in the graph, which means no interface, integration, or future caller can route around it. Approval implemented in a front end is a suggestion.
The three-way response matches how people actually review drafted work, which is mostly by changing it.
The decisions held and the wiring did not, and that split is the useful result. Seven defects came out of the regression build. Not one of them was a design choice that turned out wrong. The router gate held under every path the audit could construct. The views meant no hallucinated-column bug existed to find. The read-only briefing meant the authentication work had one surface to cover instead of two. What failed was the implementation underneath good decisions, which is a cheaper thing to fix than a wrong architecture and a harder thing to find without tests.
The most serious of the seven was in the thing the whole system is built around. The approval node took tool_calls[-1] as the call to show, so when the agent proposed a batch, the approval card could headline a different tool than the one the user was about to authorize, and only the last call was editable. A human gate that shows you the wrong action is worse than no gate, because it manufactures consent. It is fixed, the primary is now the first call whose name is in state.protected_tools, and three tests hold it there.
Then I audited this page against the code and found my own README wrong. It claimed brand context was loaded from a directory at runtime. No file I/O exists anywhere in the project. Following that instruction, a second client would have deployed Riverton's voice to their own contacts. The audit is in docs/claims-verification.md and the correction made the limitation worse, not better, which is how it should read.
An unknown review action raises rather than falling through to a null that would surface later as a confusing LangGraph error. Approval flows are easy to demo and easy to get wrong in the branches nobody clicks.
And the same data model serves a conversational agent and a read-only analyst dashboard, because not every question deserves a chat interface. Some of them deserve a number on a screen before the first call of the day.
Repository and technical artifacts
LangGraph 0.3, Claude Sonnet via langchain-anthropic, FastAPI with SSE streaming, Next.js 15 with TypeScript and Tailwind, Supabase PostgreSQL via SQLAlchemy, LangSmith tracing, Slack SDK.
Deployed on Vercel and Render, with a public demo.
The repository is private. The claims audit, the findings from the regression build and the test suite are publishable: unlike Canon's, Ralph's tests assert approval routing and streaming contracts rather than proprietary data, so publishing them costs nothing.
Visible limitation
Ralph reasons well about the pipeline it can see. It has no view of anything outside the CRM, so a relationship that went quiet because the contact changed jobs looks identical to one that went quiet because the last email landed badly. The person approving the outreach is usually the one who knows which it was.