Case Study

AI Product Practice — Discovery, Evaluation & Responsible Deployment

Most AI product failure is not model failure. It is deploying into a workflow nobody mapped, against a definition of “good enough” nobody agreed, for users who were never asked whether they trusted it. The work below shows how I reduce those three risks — in EdTech, higher education, heritage and public sector environments where the cost of getting it wrong lands on someone who did not choose the technology.

Client names have been withheld in line with standard consultancy confidentiality practice. References and further detail are available on request.

Legacy Workflow Modernisation — Adoption Proven

Reading list platform — UK higher education

Workflow Redesign Zero‑to‑One Feature Alpha/Beta Experimentation Adoption Tracking GDS Alpha/Beta/GA

Result: a bookmarking workflow that had been essentially unchanged for 15–20 years, rebuilt and taken to 94% adoption within six months across a platform serving 100+ universities.

Context: librarians lost hours to manual rework wherever the platform failed to talk to their discovery and acquisition systems. Structured discovery interviews across 17 universities set the priorities — not internal opinion about what needed fixing first.

Adoption had to be earned: the merged organisation needed evidence, not assumption, that shipped features landed. I ran alpha and beta experimentation frameworks with adoption tracking, so a feature demonstrated real uptake before general release rather than being declared a success at launch.

Alongside it: API integrations across library discovery and purchase acquisition systems, and accessibility governance progressed from WCAG 2.1 to 2.2 AA across product and interface text.

Risk & Ethics — The Decision Not to Ship AI

Reading list and discovery platform — UK higher education

Hallucination Risk Academic Integrity Feasibility Assessment Data Readiness

Context: June 2023, shortly after the public launch of GPT. My employer put every member of staff through AI training and then ran a company‑wide hack day to find real applications. I arrived without a clear brief.

What the room found: two colleagues with doctorates in machine learning told us it was not safe to deploy AI in our context — hallucination risk was too high for academic users to trust the output. Our senior AI engineer built worked examples to test that claim. The formal bibliographies and citation links in those examples were entirely fabricated.

Why that was disqualifying: on a reading list and discovery platform, citation accuracy is close to the core function of the product. Fabricated citations are not a quality defect to be tuned out later. They are an academic integrity risk, carried to students and academics who had no practical way to check.

The decision: I made the call that day. No AI capability would move forward until a tagging project was complete and a real dataset existed.

What it changed: colleagues had proposed “list health” as a use case and I went into the day without understanding what I was being asked to assess. I now treat a brief I cannot restate in my own words as the first thing to fix, not a detail to sort out on the day.

Product Judgement — AI‑Assisted Market Evaluation

Acquisition workflow — academic content distributor integrations

Market Modelling Risk Scenarios Workflow Impact Customer Protection

Context: a major academic content aggregator announced that a core product was moving to a subscription model, part‑way through our planned integration with it. I used AI to model the downstream effects at speed — workflow disruption, cost exposure per institution, and the cost/benefit under three different pricing outcomes.

What I did: I stopped the integration. The modelling showed librarian acquisition workflows would be destabilised for a benefit that could evaporate at the supplier’s next pricing decision. AI compressed a multi‑week evaluation into days; the decision itself rested on product judgement and on what our customers could absorb.

What it cost and saved: a planned roadmap item, withdrawn. Against that: engineering effort not spent, and customers not asked to re‑learn a workflow twice inside a year.

Evaluation & Significance — When Enough Is Enough

Discovery programme — university library sector

Evaluation Design Significance Thresholds Discovery at Scale Trade‑off Reasoning

Context: I ran 17 discovery calls with universities on a single problem. Pure discovery needed fewer and I knew that going in. I ran the full set anyway, because each call was also a chance to meet a customer and to thank people who had volunteered their time. I put that reasoning to my director in advance and he approved all 17.

The finding was the repetition. No single call told me the need was widespread. The recurrence across calls did. That is prevalence, not novelty — and it is the difference between an interesting anecdote and evidence you can prioritise a roadmap against.

Why this is an AI evaluation skill: the same logic governs model evaluation. A prompt run ten times is not a hunt for a better answer — it is a measurement of how often the same answer comes back. Whether the input is a librarian or a language model, the discipline is identical: know your significance threshold, know what you are measuring, and know the point at which further sampling is waste.

Zero to One — A Greenfield Assessment Platform

Digital textbook and assessment provider — UK higher education

Zero‑to‑One MVP Boundary Scope Focus‑Group Validation Embedded Analytics Go‑to‑Market

Context: a greenfield instructor assessment platform — no product, no defined scope, nothing internal to copy. I arrived as a product business analyst on the company’s eBook reader and was promoted to product manager partway through, on the strength of how I worked with the technical team and moved a complex backlog on my own.

What I did: I took it from concept to focus‑group‑ready. I drafted the roadmap, set the MVP boundary scope, prioritised the backlog and led delivery across functions. I demonstrated work in progress to librarian focus groups and to leadership and venture investors, and iterated on what came back.

Measurement built in, not bolted on: I wrote the analytics requirements on behalf of my team and handed them to the company’s Power BI developers, so engagement reporting was embedded inside the assessment tool rather than added afterwards. To know whether a product works, you specify the measurement at the same time as the feature — not once the feature is already live.

To market: with design and marketing I storyboarded the launch video and prepared the pre‑launch sales collateral. Ownership of a product from problem definition all the way through to how it would be sold is the part that settled it for me. This was the job I wanted.

How it ended: the release was paused when the company entered financial restructuring. The product itself never shipped. What happened next matters more: the same company later acquired a mature product in exactly that space. The problem was real and the opportunity was real — the timing and the balance sheet were not.

Capability Building — Practice That Outlasts Me

Mentoring Communities of Practice Retrospectives Inclusive Facilitation

Product practice does not scale by hiring alone. I have mentored 16 product managers, business analysts and interns, and built Centres of Excellence and Communities of Practice so that good technique becomes repeatable rather than personal to whoever happens to know it.

How I do it: be transparent about my own mistakes first. Run retrospectives, and run them on myself in front of the team. Support people publicly, in the room; take problems to a one‑to‑one. People only share what they do not know once they have watched someone senior do it and survive.

I also chair a Patient Participation Group — the same work in a civic setting. That group arrives with very different starting points and no obligation to agree. The task does not change: settle on what matters most.

How I Sequence AI Work

I work in this order: understand the existing workflow before proposing to change it; define what “good enough” means and how it will be measured, in writing, before anything is built; establish where the model must not be trusted; and only then design for adoption.

Roadmaps follow the same sequence. The uncertainty that would kill the product gets reduced first — not the one that is easiest to reduce. That ordering is uncomfortable, because the hardest uncertainty is usually the one nobody wants to look at yet. It is also the only ordering that stops a team spending six months on something that was never going to work.

Want to discuss a specific challenge?

Every engagement is different. If you'd like to talk through something that doesn't fit neatly into the above, I'm happy to have that conversation.

Get in touch