Case study — AI Tagging

◇ In progress — framed transparently

Building an AI-assisted pipeline to scale qualitative coding

Emerging Travel Group

Where the pipeline stands today
Codebooks unified
done
~2,700 coded by hand
done
AI comparison run
done
Production launch
blocked — next milestone
Problem
Recurring qualitative feedback was coded manually against five inconsistent, fragmented codebooks — a process that didn't scale and delayed insight delivery.
My role
Leading the initiative end-to-end, with minimal dependency on a dedicated AI/engineering team.
Methods
Codebook unification, historical-data benchmarking (AI-generated tags vs. ~2,700 comments coded by hand), confidence-scoring design for human-in-the-loop review.
Status
Codebook unified and validated against real historical data; tagging-rule governance in place; production rollout not yet launched.
01
Why this project, and why now

Feedback volume had outgrown what manual coding could sustain

Five separate, inconsistent codebooks had accumulated across different feedback sources. Rather than requesting engineering resources to build a bespoke tagging tool, I set out to build and validate an AI-assisted process myself — using this as an opportunity to develop hands-on fluency with applied AI tooling in a research context, which is increasingly expected at senior level.

02
What I did

A human baseline first, then a direct comparison

  • Consolidated five fragmented codebooks into one unified taxonomy of topics and subtopics.
  • Working with my research operations colleague, manually coded roughly 2,700 historical comments by hand first, to establish a clean human baseline — then ran the same comments through an AI-assisted tagging pass to directly compare the AI's categorization logic against our own.
  • Used the disagreements between AI and human coding to generate around 230 concrete tagging rules, then set up a structured review process to clean up duplicates and standardize naming before treating the ruleset as production-ready.
  • Designed the target workflow: an automated process that pulls in feedback from surveys and relevant internal channels, assigns a confidence score to each tag, and routes anything below a defined threshold to a human reviewer — rather than fully automating without a quality check.
  • Defined the rollout principles up front: start scoped to a single product area, keep a human in the loop via confidence thresholding, and only expand to other products once the approach is validated with domain experts in each area.
03
Where it stands

Validated against real data — production is the next step

The codebook and rule-governance layer are built and validated against real historical data. The production version — the automated system that tags new, incoming feedback without manual intervention — has not yet launched; that's the next milestone.

◇ Status: validation complete — production rollout is the next milestone.
The agreement-rate comparison between AI and human coding will be added here once available — likely the strongest single visual in this case study.
04
Why I'm including an unfinished project

Being transparent about what's shipped versus what's next is itself the point here — the market increasingly expects researchers to be building with AI, not just prompting it, and a credible in-progress project with a clear validation methodology, done together with the researcher I've been mentoring, says more about how I work than a polished but shallow one would.