Agentic AI / Trust / Delegation

Aurora AI: Delegation You Can Inspect

Delegating a call was easy. Trusting what happened was not.

Product
AI agent for real-world calls
Role
Founding product designer candidate
Scope
Product thesis, interaction model, live status, ranked results, edge cases
Aurora AI: Delegation You Can Inspect case study cover

The tension

The moment a user leaves the app, the agent becomes invisible. That is where trust starts to break.

What was happening

Aurora could take a request and place calls in the background, but users still needed to know what the system understood, what was happening, and why a result deserved action.

What changed the brief

The waiting state was not dead time. It was the product surface where Aurora could expose progress, learn one useful preference, and earn permission to finish offscreen.

Success looked like

Users can confirm intent once, leave safely, return to ranked choices, and open the exact call evidence behind each recommendation.

The constraint

Calls have uncertain pickup rates, estimates can move, facts may be incomplete, and mid-run changes can invalidate work.

6-10min. expected range

Prototype benchmark, recalibrated as businesses answer.

5stages ask to action

Designed flow: ask, confirm, call, rank, and act.

100%ranked results

Validation target, not a shipped outcome.

The product story

The interface, one decision at a time.

01 / Intent

Teach specificity while confirming the delegation contract.

Aurora turns a loose request into a visible brief, asking only for the constraint that can change the result.

Aurora AI request clarification and editable call brief
The user can inspect what Aurora will ask before any call starts.Budget, service type, coverage, and location become one editable contract instead of hidden prompt interpretation.

02 / Live run

Make the wait useful, honest, and safe to leave.

A live run shows completed calls, the remaining range, and what changed. One optional preference question improves ranking without blocking the task.

Aurora AI live calling progress and preference tuning screen
Progress becomes a receipt, not a decorative loader.The estimate updates as businesses answer, notification is enabled, and preference play teaches the product what matters today.

03 / Ranked result

Compress calls into a decision without hiding uncertainty.

Verified options are ranked against the brief, while no-answer, over-budget, and missing-price outcomes remain visible.

Aurora AI ranked service-centre results inside chat
The best result says why it is first.Price, availability, distance, and requirement fit stay scannable, with incomplete evidence labelled instead of guessed.

04 / Proof

Put the evidence beside the recommendation.

Structured facts, a recording, and a transcript let the user audit the call before booking or getting directions.

Aurora AI business detail with verified call recording and transcript
Every ranked fact has a route back to the conversation.The detail view separates what was confirmed from what was inferred and keeps the next action close.

05 / Edge case

A changed constraint should not silently waste completed work.

When the budget changes mid-run, Aurora explains which answers still qualify, which calls adapt, and whether a restart is needed.

Aurora AI mid-run budget constraint update confirmation
The consequence appears before the new constraint is applied.Valid answers are preserved, active calls use the new limit, and undo remains available.

The trust test

The happy path is only half the product.

Businesses do not answer

The run distinguishes no answer from failure, recalibrates the estimate, and keeps verified progress intact.

Evidence is incomplete

Missing price or availability is labelled as not confirmed instead of being filled by model confidence.

The brief changes mid-run

Aurora explains the impact, reuses valid work, and asks before any change that would require a restart.

The decisions

Three moves carried the story.

01

Teach the brief progressively

Aurora asks for one decision-changing constraint at a time, then shows the complete call brief so better instructions feel useful, not like prompt homework.

02

Make waiting useful and honest

A time range, live receipts, and recalibration replace fake precision. Optional preference tuning learns what matters without holding completion hostage.

03

Rank with proof, not confidence theatre

Results explain why they rank, preserve incomplete outcomes, and place recordings and transcripts beside the facts they support.

What I would test next

I would test whether progress receipts reduce check-ins, whether one preference question improves ranking confidence, and whether visible proof increases booking action.