
Education technology
Reference buildFifty thousand courses, two hundred providers, one schema
A course discovery engine that normalises 200+ provider APIs into one canonical schema and ranks them with a scoring model tuned against real click behaviour.
At a glance
- Duration
- 52 weeks
- Team size
- 5 people
- Engagement
- New build
- Project type
- Data platform
- Industry
- Education
The situation
The challenge
A Coursera specialization and a Udemy course are structurally different products, and 200+ providers meant 200+ inconsistent schemas updating on their own cadences. On top of normalising all of that, search had to feel like Google.
What we did
A plugin ingestion pipeline so each provider's weirdness stays contained, then a canonical schema and a scoring model that learns from behaviour.
The calls that mattered
One adapter per provider, one schema behind them
Provider-specific adapters own auth, rate limiting and normalisation; everything downstream sees a canonical course with provider quirks in JSONB. Deduplication happens by fingerprint before indexing.
BM25 as a starting point, not an answer
Custom Elasticsearch scoring layering rating, freshness, provider reputation and enrolment velocity on top of BM25 — with the weights set by A/B test rather than intuition.
What changed
- courses indexed
- 50K+courses indexed
- monthly active learners
- 500K+monthly active learners
- search p95 latency
- 94mssearch p95 latency
- recommendation click-through
- 3.2xrecommendation click-through
- 50K+ — with under 0.1% normalisation errors
- 3.2x — after A/B tested scoring weights
50,000+ courses from 200+ providers at under 0.1% normalisation error, 94ms p95 search, and a 3.2x lift in recommendation click-through.
Services used
- Plugin ingestion pipeline with 200+ provider adapters
- Canonical course schema with fingerprint deduplication
- Custom Elasticsearch scoring model with A/B tested weights
What we would do differently
Every project has one of these. Publishing it is the point — a case study with no regrets in it is marketing, not evidence.
Relevance is a product problem before it is an infrastructure one. Instrumenting what people actually clicked turned out to be worth more than any further tuning of the retrieval layer.
More work
- Internet Money · Crypto financial services$10M+monthly transaction volume
One wallet interface across five blockchains
A multi-chain crypto platform where every chain hides behind one interface, with a tamper-evident audit trail underneath it. $10M+ moves through it monthly.
Read case study - MagicTask · Project management SaaS100K+concurrent users
From latency spikes at 5,000 users to 100,000 concurrent
A gamified project management platform re-architected from a single Node server into an event-driven system that holds 100,000+ concurrent connections without the reward engine touching the critical path.
Read case study