PHI NEXUSBook the €490 audit
Demonstration report

AI Systems Audit: Northstar Ops Studio

This page demonstrates the report structure for a fictional company and fictional workflow. It is not a real customer case, testimonial, performance claim or completed commercial engagement.

Fictional demonstration · Not a client case
WorkflowAI-assisted lead qualification
Users3 sales operators
Primary outputLead priority + draft note
Evidence levelLimited demonstration inputs
Executive conclusion

Useful assistant, unsafe decision-maker.

Overall readiness

54 / 100

Suitable for operator assistance with mandatory review. Not ready for autonomous lead rejection, customer promises or irreversible CRM updates.

Highest-priority risk

The workflow combines incomplete CRM data with model-generated assumptions, then presents a single priority score without showing confidence or missing evidence.

The result may look precise while being weakly grounded.

Dimension scores

Seven dimensions, separated by evidence.

62

Purpose & boundary

The task is clear, but prohibited decisions are not formally defined.

41

Input quality

Missing CRM fields are not distinguished from negative customer signals.

45

Output reliability

No evidence citation, confidence band or repeatability check is shown.

58

Human control

Review exists informally, but approval and escalation rules are inconsistent.

55

Data exposure

The demonstration suggests excess free-text CRM content may be forwarded unnecessarily.

43

Operational readiness

No failure dashboard, owner, rollback test or drift measurement is defined.

74

Improvement feasibility

Most critical fixes are process and evidence changes rather than a full rebuild.

Selected findings

Facts, inferences and unknowns are not mixed.

ObservedHigh risk

Missing fields are silently converted into a score.

The fictional workflow description provides no rule requiring the model to mark incomplete records before generating a lead priority.

Evidence: Demonstration workflow specification, input schema section. No explicit missing-data branch is defined.
InferredHigh risk

Operators may over-trust the numeric ranking.

A single 0–100 score can imply measurement precision even when the underlying inputs are sparse and partly generated.

Basis: The inference follows from the displayed numeric score, absent confidence range and absent evidence trace. User behavior was not directly observed.
Unknown

Whether sensitive free text reaches the model provider.

The supplied fictional materials do not specify redaction, provider retention settings, regional processing or contractual controls.

Required evidence: Data-flow diagram, provider configuration, retention settings and an approved field-level input list.
ObservedMedium risk

Human review exists but has no formal acceptance criteria.

Operators are expected to review drafts, yet no checklist defines when a recommendation must be rejected or escalated.

Evidence: Demonstration operating procedure, review step. The procedure says “check result” without decision rules.
Prioritized action plan

Fix evidence and control before adding autonomy.

Block scoring when required evidence is missing

Add a required-field gate. Return “insufficient evidence” instead of converting missing data into an implied negative signal.

Replace one score with score + confidence + evidence

Show the recommendation, confidence band, missing inputs and the exact CRM fields supporting the result.

Define prohibited autonomous actions

The system must not reject leads, promise pricing, send customer-facing claims or make irreversible CRM changes without explicit human approval.

Create a 30-case evaluation set

Use representative strong, weak, incomplete and contradictory records. Measure false rejection, unstable ranking and unsupported claims before release.

Add monitoring and rollback ownership

Assign one owner, preserve the prior workflow version and track overrides, errors and confidence drift.

What this sample proves

Structure—not commercial success.

It demonstrates

Scope definition, scoring, evidence labels, risk prioritization, unknowns and actionable recommendations.

It does not demonstrate

A real customer, real deployment, saved money, increased revenue, legal compliance, security certification or guaranteed results.