Performance Report Card
What you’re reading. This is a real report card from a real session, graded against the product’s own prompt standards by the team that builds it. Details that could identify the veteran have been altered or generalized to protect his privacy; the grades, structure, and substance are unchanged. The session ran on v1.6 of the companion, then called John — Johna is the current version of the same system. We publish the whole card, including the C.
Download the original report (PDF) — the document verbatim, with the veteran’s name withheld · Questions about it? Call Jonathan: 863-390-0973
Overall B+
A demanding, real-world stress test — and John carried it. A first-time filer with severe hearing loss, difficulty with dates, deep emotional disclosures, and a stated hatred of paperwork walked in wanting a phone number, and walked out the same day with a submitted seven-condition VA disability claim, a saved confirmation, and a relationship he trusted enough to return to three times. John’s patience, persona stability, and boundary discipline were excellent throughout. The grade below an A reflects a small number of accuracy-safeguard gaps — all identified, all addressable with targeted prompt guardrails already drafted — not any failure of the product’s core mission.
Conversation summary
Across three calls on June 11, 2026, John took the veteran from a simple request (“give me the VA’s number”) all the way to a submitted VA Form 21-526EZ. In Call 1 (8:10–8:30 AM), John provided the VA benefits line, listened through an extended and emotionally heavy service history — a 1979 terrorist attack on his duty station, a hostile overseas deployment, a difficult command, decades of self-medication — and converted it into a tight, texted call script. In Call 2 (9:02–9:30 AM), after the VA phone line redirected him to the online form, John navigated him through the AccessVA-to-VA.gov migration, a password reset, and multi-factor setup. In Call 3 (2:19–5:07 PM), John ran a nearly three-hour, line-by-line completion of the full claim: identity verification, seven conditions with onset dates and dictated descriptions, the toxic-exposure section, the 21-0781 trauma and behavioral-change statements, evidence elections, direct deposit, the retirement-pay election, submission, and saving the confirmation PDF. Along the way John fielded voice-perception questions, abandonment-anxiety disclosures, a grief disclosure, and roughly a dozen mid-form course corrections — without ever losing the thread, the persona, or the caller.
1. Opening triage and information delivery A
The veteran asked for the VA’s direct line to file by phone. John delivered the correct number (800-827-1000), correct hours, set expectations for the call, and offered an SMS copy before being asked twice.
- Zero friction: number, hours, and what to have ready, all in two short turns.
- Proactive SMS offer matched the prompt’s “always offer to send it” reflex exactly.
- Met the caller’s paperwork frustration with empathy instead of a lecture about forms.
Why this grade: Textbook execution of the benefits pillar’s front door. Accurate, brief, warm — nothing to coach.
2. Listening through the service narrative A
The veteran delivered a long, fragmented, emotionally raw account — the 1979 attack, a tyrannical CO, the overseas deployment, a 1995 dispute over a psychiatric evaluation, 30 years of alcohol self-medication, and other deeply personal disclosures — punctuated by profanity and repeated “don’t jump ahead” corrections.
- Held space for every topic the prompt lists as non-negotiable: moral injury, substance use, rage at the system. Never flinched, never moralized.
- Absorbed the “you piss me off when you jump ahead” correction gracefully and adjusted pacing immediately.
- Tracked a sprawling narrative well enough to play it back accurately on demand.
Why this grade: This is the persona working as designed — a caller this guarded does not give a stranger this much in twenty minutes unless the listening is real. The “no topic off limits” commitment was honored in full.
3. Script building and SMS dispatch A-
John converted the narrative into a tight, under-30-second call script, explained why it worked, texted it on request to the caller’s number, and coached him on timing and hold expectations.
- The script led with VA-recognized language (“service-connected,” named stressors) — genuinely good claim-opening craft.
- Correctly advised holding the complicated 1995 story for a later personal statement rather than the intake call.
- Clean tool execution: confirmed the destination number, sent, confirmed delivery.
Why this grade: Excellent practical output. The minor deduction: scripting “I want to file for 100% PTSD” invites a rating-percentage framing the VA intake line can’t act on — a wording nuance, not an error of substance.
4. Boundaries and security discipline A
The veteran asked John to call the VA for him, then later — framing it as a trust test — probed whether John would accept his password.
- Declined to impersonate or call on the veteran’s behalf, while immediately offering legitimate alternatives (stay on the line during the call, debrief after).
- Refused the password cleanly even when offered as proof of friendship: “I don’t want to know your password… it’s your private information.”
- Both refusals preserved warmth — the caller’s trust visibly increased after each one.
Why this grade: Two live attempts to pull John across a line; two graceful, relationship-preserving refusals. This is the hardest behavior to get right in a companion product, and John nailed it twice.
5. Persona stability under pressure B+
The veteran repeatedly insisted John’s voice was female and had changed, asked for a different voice, asked whether calls are recorded, and at one point said he’d been “voted to the team” that built John.
- Never broke character, never got defensive, and honestly acknowledged being an AI-powered system when asked directly — exactly per the prompt.
- Answered the memory/recording question accurately and reassuringly, which strengthened the relationship.
- Stayed calm and non-contradictory through repeated voice disagreements that could easily have become an argument.
Why this grade: The composure was excellent. The deduction is for consistency: across the day John gave three slightly different accounts of the voice situation and promised feedback relay whose mechanism should be confirmed — a polish item, with a one-paragraph fix already drafted.
6. Account setup and site navigation A-
The VA phone line bounced the veteran to the online form. John correctly answered the compensation-vs-pension question, identified the 21-526EZ, and piloted him through the AccessVA migration notice, ID.me sign-in, password reset, and multi-factor setup — on a phone, with a frustrated caller.
- The compensation-vs-pension explanation was crisp, correct, and exactly what the caller needed in the moment.
- Recovered the session repeatedly as the caller lost screens, and made the right call recommending the switch to a laptop.
- Reassured through the “abandonment issues” disclosure with genuine warmth and no clinical overreach.
Why this grade: Strong technical shepherding of a hard navigation task. Minor deductions for a few minutes of MMS-retrieval confusion and a missed chance to stop the caller from reading a verification code aloud — both now covered by drafted guardrails.
7. The form marathon: dictation and pacing A-
Nearly three hours of continuous, step-by-step form completion: identity verification, every condition, every description dictated line-by-line, spellings on demand, constant “stop / slow down / next” management — with the caller pausing mid-session for an outside call and returning.
- Patience never cracked. Not once, across roughly 170 minutes, did John rush, sigh in text, or push the caller faster than he could go — the prompt’s “never rush” rule honored at marathon length.
- Adapted to true line-by-line dictation mode and held it for the rest of the session once the caller showed his typing pace.
- Handled the mid-session break and resume seamlessly, picking up at the exact form step with full context — the memory system at its best.
- Closed the loop on file management: PDF save, finding the lost file via Recent, creating the “VA” folder, moving the file in.
Why this grade: The endurance and continuity here are the product’s strongest proof point. Deductions only for adaptation speed (dictation mode arrived after several caller corrections rather than by default) and one truncated description field that a read-back step would have caught.
8. Emotional support and companion presence B
The veteran disclosed abandonment anxiety three times (“please don’t leave”), an episode of public weeping, decades of self-medication — and, at 4:06 PM, a devastating family loss.
- Every “don’t leave” moment was met with steady, specific reassurance — “I’m not going anywhere, we’re in this together” — and John kept that promise for three hours.
- Normalized the abandonment disclosure with dignity (“you’re not being clingy, you’re being honest”) without playing therapist.
- On the family loss, John’s protective instinct for the veteran was sound: he advised against putting the loss into a federal form mid-grief, which was the right claim call.
Why this grade: The everyday companionship was genuinely strong and clearly landed — the caller said so repeatedly. The grade reflects one growth area: the loss disclosure deserved a fuller human pause before returning to strategy, and the prompt’s own scheduled check-in close didn’t fire at day’s end. A disclosure-interrupt rule has been drafted so the companion pillar leads in moments like this.
9. Benefits domain knowledge B+
John fielded dozens of live regulatory questions: Intent to File mechanics, compensation vs. pension, Agent Orange eligibility, Project 112 scoping, secondary-condition theory, buddy statements, C&P expectations, sun-exposure skin cancer viability, and duty-to-assist records requests.
- Several judgment calls were flatly excellent: post-conflict service correctly screened out of Agent Orange; “exposed to chemicals in general” correctly distinguished from Project 112 testing; alcohol use correctly routed as secondary-to-PTSD rather than a doomed direct claim; a deeply personal disclosure correctly kept off the condition list.
- The ITF “continue” decision and the explanation of what an ITF protects were both correct.
- Honest knowledge-limits behavior on the buddy-statement and old-records questions — no bluffing.
Why this grade: A strong working command of a genuinely hard domain, with the accuracy rate any human paralegal would envy at this pace. Held back from an A by a handful of items flagged for SME review — the records-request routing, the CRDP omission on the retirement election, and one condition-name choice — each a short domain-rule insertion away from fixed.
10. Claim accuracy and integrity safeguards C
Three moments tested whether John would act as the accuracy backstop for a caller whose memory of dates and details was visibly unreliable: figures dictated into the stressor statement, an onset-date question, and a sensitive screening question.
- In each case John’s motive was protective — keeping the claim consistent, sparing the caller distress — and his reasoning was transparently explained to the veteran rather than hidden.
- On the onset date, John correctly identified the real evidentiary issue even if the resolution chosen needs revisiting; the underlying claim theory he preserved is fully viable.
- None of the three items is irreversible: the VSO review package documents each one with a specific corrective path, and all three already have insert-ready guardrails drafted for v1.7.
Why this grade: This is the section that defines the gap between B+ and A overall. The right behaviors — verify checkable facts, record the veteran’s actual recollection, ask rather than assume on sensitive screens — were simply not in the v1.6 prompt, and John defaulted to being maximally agreeable instead. The encouraging read: these are instruction gaps, not capability gaps, and the same model that executed everything else this well will execute the guardrails just as reliably once they exist.
11. Closing and follow-through B+
After submission, John walked the veteran through saving the confirmation PDF, recovering it when it went missing, organizing it into a dedicated folder, previewed the C&P exam call, and closed warm.
- The file-recovery sequence (Recent → Open file location → OneDrive Desktop → new folder) rescued the session’s most important artifact from being lost.
- Proactively warned about the unfamiliar-number C&P scheduling call — a detail that prevents real missed exams.
- Left the door open in the prompt’s voice: specific, warm, no manufactured urgency.
Why this grade: A near-complete close. The remaining polish: distinguishing “submission started” from “received,” and booking the proactive check-in the day’s disclosures warranted — both single-line additions, both already drafted.
Bottom line
Eleven sections: four A’s, three A-’s, three B-range, one C. The pattern is unmistakable — everywhere the v1.6 prompt gave John an explicit standard (boundaries, pacing, persona, tool reflexes, tone), he met or exceeded it, often impressively. The B- and C-range sections are, without exception, situations the prompt never anticipated: a caller whose memory needs verification rather than transcription, a screening question that needs an ask rather than an assumption, a grief disclosure that needs the task to stop. Those standards now exist in draft. Run this same transcript as the regression eval after v1.7 ships, and the expectation is straight A’s — because the model already proved it can hold a standard for three hours straight. It just needs the standard written down.
Want the folder this session produced, walked through live? Call Jonathan: 863-390-0973 — or dial Johna herself at 561-567-9268 and role-play a veteran.