Skip to content

Education technology

Reference build

Fifty thousand courses, two hundred providers, one schema

A course discovery engine that normalises 200+ provider APIs into one canonical schema and ranks them with a scoring model tuned against real click behaviour.

50K+courses indexed

At a glance

Duration
52 weeks
Team size
5 people
Engagement
New build
Project type
Data platform
Industry
Education

The situation

The challenge

A Coursera specialization and a Udemy course are structurally different products, and 200+ providers meant 200+ inconsistent schemas updating on their own cadences. On top of normalising all of that, search had to feel like Google.

What we did

A plugin ingestion pipeline so each provider's weirdness stays contained, then a canonical schema and a scoring model that learns from behaviour.

The calls that mattered

  • One adapter per provider, one schema behind them

    Provider-specific adapters own auth, rate limiting and normalisation; everything downstream sees a canonical course with provider quirks in JSONB. Deduplication happens by fingerprint before indexing.

  • BM25 as a starting point, not an answer

    Custom Elasticsearch scoring layering rating, freshness, provider reputation and enrolment velocity on top of BM25 — with the weights set by A/B test rather than intuition.

What changed

courses indexed
50K+courses indexed
monthly active learners
500K+monthly active learners
search p95 latency
94mssearch p95 latency
recommendation click-through
3.2xrecommendation click-through
  • 50K+ — with under 0.1% normalisation errors
  • 3.2x — after A/B tested scoring weights

50,000+ courses from 200+ providers at under 0.1% normalisation error, 94ms p95 search, and a 3.2x lift in recommendation click-through.

Services used

  • Plugin ingestion pipeline with 200+ provider adapters
  • Canonical course schema with fingerprint deduplication
  • Custom Elasticsearch scoring model with A/B tested weights

What we would do differently

Every project has one of these. Publishing it is the point — a case study with no regrets in it is marketing, not evidence.

Relevance is a product problem before it is an infrastructure one. Instrumenting what people actually clicked turned out to be worth more than any further tuning of the retrieval layer.