Case Study
AI Product Practice — Discovery, Evaluation & Responsible Deployment
Most AI product failure is not model failure. It is deploying into a workflow nobody mapped, against a definition of “good enough” nobody agreed, for users who were never asked whether they trusted it. My AI practice focuses on reducing those risks through disciplined discovery, feasibility assessment, evaluation design, and responsible deployment.
Feasibility, ethics & responsible deployment
I evaluate AI capability through the lens of workflow impact, academic integrity, significance thresholds, and customer trust. My work spans higher education, digital platforms, and public‑sector environments where the cost of getting it wrong lands on someone who did not choose the technology.
Model evaluation & significance
Evaluation is not about finding a better answer — it is about understanding how often the same answer comes back. Whether the input is a librarian or a language model, the discipline is identical: define “good enough,” know your significance threshold, and stop sampling when further data adds no value.
Case Studies
These examples illustrate how AI feasibility, ethics, evaluation, and workflow impact shaped product decisions and protected users.
Risk & Ethics — The Decision Not to Ship AI
Reading List & Discovery Platform — UK Higher Education
Context: Shortly after the public launch of GPT, my employer ran company‑wide AI training and a hack day to identify viable use cases. Two colleagues with doctorates in machine learning warned that hallucination risk was too high for academic users to trust the output.
What we found: Our senior AI engineer produced worked examples. Bibliographies and citation links were entirely fabricated — a critical failure mode for a product whose core function depends on citation accuracy.
Decision: I halted all AI capability until a tagging project was complete and a verified dataset existed. A brief I could not restate clearly became the first thing to fix, not a detail to sort out later.
Outcome: No AI shipped on unverified or fabricated data. A failure mode identified in a workshop rather than in front of a student.
Product Judgement — AI‑Assisted Market Evaluation
Acquisition Workflow — Academic Content Distributor Integrations
Context: A major academic content aggregator announced a shift to a subscription model part‑way through our planned integration. I used AI to model downstream effects at speed — workflow disruption, cost exposure per institution, and cost/benefit under three pricing outcomes.
What I did: Stopped the integration. The modelling showed librarian acquisition workflows would be destabilised for a benefit that could evaporate at the supplier’s next pricing decision. AI compressed a multi‑week evaluation into days; the decision itself rested on product judgement and customer absorption capacity.
Outcome: A reversible decision made early, protecting customer workflows during supplier‑driven market uncertainty.
Evaluation & Significance — When Enough Is Enough
Discovery Programme — University Library Sector
Context: I ran 17 discovery calls on a single problem. Pure discovery needed fewer; I ran the full set because each call was also a chance to meet customers and thank people who volunteered their time.
Finding: No single call proved the need was widespread. The recurrence across calls did. That is prevalence, not novelty — the difference between an anecdote and evidence you can prioritise a roadmap against.
Why this is an AI skill: Model evaluation follows the same logic. A prompt run ten times is not a hunt for a better answer — it is a measurement of how often the same answer comes back.
Outcome: A deliberate, defended trade‑off — spending past the significance threshold for relationship value, with reasoning made explicit and agreed in advance.
How I Sequence AI Work
I work in this order: understand the existing workflow before proposing to change it; define what “good enough” means and how it will be measured; establish where the model must not be trusted; and only then design for adoption. The uncertainty that would kill the product gets reduced first — not the one that is easiest to reduce.
Live Agentic Workflow — Shipped Example (Alpha)
LyttonOS Input: A shipped agentic workflow tool supporting structured analysis, sequencing, and document generation. Designed to demonstrate responsible agentic patterns in practice.
Alpha release — responsible deployment disclaimer applies.