↓ Skip to main content

How a Research Bottleneck Became an API

·6 mins·
Phil Clarke
Author
Journalist, Hacker, Musician

While I was working at Motorcycle News as a product writer, I encountered a number of workflow bottlenecks that were losing time. After my first workflow automation tool made it into the full workflow for the team, that became part of my regular duties. This is the largest and most recent tool I made for the editorial team — MCN Bike Reviews Search.

The Workflow Problem

MCN has the largest single archive of motorcycle reviews on the internet (if you know, you know). It's one of MCN's greatest digital assets, with long-tail organic traffic, and it's a gold mine of potential content ideas. That is, if you can search the bike reviews effectively.

A rudimentary archive search existed, but there was no way to filter bikes by their specs, like:

  • seat height
  • engine power
  • used price
  • fuel economy
  • running costs
  • etc

Research meant opening individual review pages and manually checking for bike specs, and that was taking hours per week of editorial time.

My pitch was simple. With one well-structured query, you can immediately find all the "Best Bargain Bikes for Speed Demons", or "Most Fuel-efficient but Motorway Capable Bikes" for whatever newsletter, buying guide, or social post you're currently working on. Or, if you're tasked with generating new content ideas, your source of inspiration should almost never run dry.

I prototyped a working demo within 24 hours, pitched it to editorial and IT teams, and gained their backing to develop it into a full version.

Explain the practical bottleneck:

  • MCN had lots of bike reviews.
  • Writers needed to compare bikes by specs.
  • Site search was not enough.
  • Research meant manually opening pages and checking specs.

Purpose: show this came from a real user/workflow need, not portfolio cosplay.

Why I Built an API, Not Just a Spreadsheet

Before you ask, yes, an API really was necessary for this tool. MCN were sitting on a potential gold mine, and it was my job to extract the most value out of it.

The main reason to make this an API is so it can become reusable infrastructure, not just a one-and-done tool. A well designed API gives a predictable way to query for other interfaces, scripts, AI agents and more. That makes future features or integrations much quicker to develop, and the value proposition will compound over time. In a large publishing environment, where systems often grow in isolation without talking to each other, that matters — it was the single biggest source of user complaints while I was there, hindering workflows for many moons before my arrival. My single API would act as that one small step for man, but a giant leap for publishing-kind.

The data we needed wasn't useful as plain text, only when queryable by attributes. Making associations between data points was a big part of the value the tool promised, so it needed to be stored as structured data.

As well as this, motorcycle specs needed to be filterable, sortable, and comparable. That's more than just a spreadsheet is capable of, and warranted its own query engine.

Data needed to remain structured so it could be truly portable — readable by any tool for any purposes. The option to add features later for different kinds of search was also crucial, as there are always new use cases to be discovered.

Explain the product/technical decision:

  • Specs needed filtering, sorting, comparison.
  • Data needed to stay structured.
  • Other tools/AI agents might need to query it later.
  • API makes the data reusable.

Purpose: show API judgement.

Turning Messy Pages into Structured Data

The first major problem I encountered was actually getting to the data. All the entries already existed in a MySQL database, but as an editorial colleague I wasn't given access to it directly. So, I thought if I can't get in through the back end, head to the front.

Inspecting HTML from the live web pages, the bike specs were rendered as HTML table elements. It looked quite straightforward at first — fetch the pages, parse the tables with CSS selectors, and normalise the results.

But the real work was making that process reliable. I had to handle sitemap diffing, async page fetching and parsing, rate limiting, retries, error handling, and cached records.

The main hiccup with parsing came down to the age of the archive. Motorcycle review pages have been through several different formatting changes over the years, so one parser had to handle multiple generations of page structure.

Each failure had to be kept inspectable, so I paid particular attention to errors around the system boundary and wrapped them with relevant metadata. That meant HTML fetch errors could retry automatically, and HTML parse errors wouldn't block the rest of the pipeline.

Once the pipeline could generate cached records reliably, the first hurdle was complete. It was now ready to support the query engine.

Explain the data layer:

  • fetched/parsing pages,
  • inconsistent fields/missing values,
  • how there's multiple types of pages and they all require different parsing strategies,
  • cached records,
  • normalising enough to query,
  • keeping failures inspectable.

Purpose: show real-world messiness.

TODO Designing the Query Shape

Show example filters:

  • seat height under X,
  • price under X,
  • sort by power,
  • combine filters with and / or / not.

Include one or two curl examples. Purpose: developer-facing tutorial proof.

Making It Safe Enough for Other Systems to Call

This is the most Tyk-shaped section. Cover:

  • validation,
  • malformed JSON,
  • unknown fields,
  • clause limits,
  • nesting-depth limits,
  • oversized bodies,
  • structured errors,
  • health checks.

Purpose: connect to API reliability and trust.

From API Endpoint to Editorial Workflow

Cover:

  • the AI agent front-end and what you built into it

    • tutorial
    • idea generation engine
    • how it interfaces with the API
    • the importance of docs

Documentation as Part of the Tool

Cover:

  • README/docs,
  • OpenAPI 3.1,
  • Postman collection,
  • examples,
  • matching docs to actual behaviour.

Purpose: DevRel/docs proof.

Why This Matters for AI Agents

Connect to your LinkedIn post:

  • agents need reliable actions against real systems,
  • prompts are not enough,
  • APIs need clear behaviour,
  • errors need to be explainable,
  • humans need inspectable source data.

Purpose: thought leadership + Tyk relevance.

What I'd Improve Next

Show maturity:

  • database instead of EDN cache,
  • auth/rate limits,
  • frontend/browser UI,
  • better observability/logging,
  • CI/CD,
  • broader test coverage.

Purpose: honest technical judgement.

Closing Lesson

End with the general principle: The hard part was not making bike reviews searchable. The hard part was turning messy editorial knowledge into an interface that people — and other systems — could trust.

Recommended Final Outline

Turning an Editorial Workflow Bottleneck into a Search API

The workflow problem hiding in plain sight Why site search wasn’t enough Why an API was the right shape Turning messy review pages into structured data Designing useful filters instead of keyword search Making the API safe for other systems to call Documentation is part of the product What this taught me about AI agents What I’d improve next The lesson: useful automation needs reliable interfaces