Scoreboard
Evaluated calls, verdict mix, verified accuracy, decision speed.
| Brand | Verdicts | Coverage | Human | Machine | Accuracy | p50 |
|---|
Call volume
Raw provider volume per bucket.
Verdict mix
Human, machine subtypes, Undecided, and unanswered share per engine.
Verified accuracy
Against operator-verified truth labels.
Detection latency
All brands · median audio consumed to decision (not wall clock). p95 and n in each row.
Active Call
Live state of the run currently in flight.
No active run.
Recent calls
Newest first, across every ingestion source.
- Loading recent calls…
Run History
Filtered by the controls above; search narrows the loaded page.
| Started | Status | Agreement | Verified | Confidence | Latency | Source | Direction | Language |
|---|
Call detail
Selected run: none
Engine comparison
Recording Playback
Replay feeds the saved recording through the live engines as PCM. It writes a new run and does not call Bland listen on a finished call.
Transcript
Provider transcript
No provider transcript.
AMD transcript
No AMD transcript.
Corrections
Every correction is written to an immutable audit row.
Verified truth
AMD verdict
Audit history
- No corrections recorded.
Delete Data
Irreversible. Type the selected run id to confirm.
Verified accuracy
Streamed calls someone listened to and labeled.
No calls verified yet. Judge a few below and this fills in.
Accuracy by call type
Verified calls
Every call someone listened to and labeled. Click a row to open it in Calls.
| When | Bland call ID | Truth |
|---|
Reverify call
Listen again and update the label if needed.
Bland has no playable recording for this call.
Judge a call
The queue is recording + at least 5 engines tagged human or machine. Empty is expected when a call is missing a tag or a playable recording. Listen, then answer. What the engines said stays covered until you have — seeing their verdict first turns this into agreement, which is the number we already have. Calls where JulsAMD's internal evidence sources conflict come first, least confident at the top. Space play · H person · M machine · N nobody answered · S skip
Bland could not stream that recording — skipping.
What it scores
Offline figures come from a frozen split the model never trained on. Verified figures come from calls a person judged — the only ones here that measure correctness rather than agreement.
What the field claims
Kept in its own column: a marketing adjective is not a measured rate and should never be read beside one.
Inside the model
How a call becomes a verdict
Each stage names what it removes or adds. "The model decided" is not an explanation anyone can act on.
Why this call
Paste a transcript, or load one from the queue. For a linear model the contributions are not an approximation of the reasoning — they add up to it exactly.
Eight axes
Loading…
Where the evidence comes from
Three separate piles. A call a person judged and a clip an upstream detector tagged are never counted in the same denominator.
What the field claims
Kept in its own column: a marketing adjective is not a measured rate and should never be read beside one.
Source documents
Every figure above is computed from these on each request. Nothing on this tab is typed in by hand.
Manual Evaluation
Sync the S3 call-audio corpus, set the human or machine gold label, then run every recording through the acoustic engines. Audio streams from S3 and is never stored on this host.
S3 metadata is ready to sync.
Sync S3 to inventory call recordings. This page does not scan S3 automatically.
Corpus recordings
Select a recording to open its transcript and label review.
| Recording | Transcript | Judgment | Latest evaluation |
|---|
Upload a corpus recording
A WAV upload streams directly to helloalexamd/calls/upload_<epoch>.wav. Uploading adds it to the corpus; it does not start evaluation.
JulsAMD, StephenAMD, GalaxyAMD, DomAMD, and LeeAMD consume the same PCM stream. BlandAMD is provider-only and therefore N/A for an audio file.
Current batch score
No corpus batch has been started.
Live updates: waiting for sign-in.
Launch Test Call
Dials a configured destination and records both verdicts.
Consent Notice
Shown to the operator before every launch.
Provider Reconciliation
Backfills provider calls that never reached a webhook.
Raw sync state
Loading sync status…
Bland ingestion diagnostics
Live poll state and redacted provider failures. This is the first place to check when a call is not heard.
Recent Bland errors
Loading Bland diagnostics…
Readiness
Config-derived; never green on unproven credentials.
Loading readiness…
Retention
Drops runs and recordings past their retention window.
No cleanup run yet.
Lifetime summary
Raw summary
Loading summary…