مواد پر جائیں

MagicTask · Project management SaaS

From latency spikes at 5,000 users to 100,000 concurrent

A gamified project management platform re-architected from a single Node server into an event-driven system that holds 100,000+ concurrent connections without the reward engine touching the critical path.

100K+concurrent users

ایک نظر میں

مدت
78 ہفتے
ٹیم کا حجم
8 افراد
معاہدے کی نوعیت
پلیٹ فارم کی تبدیلی
پروجیکٹ کی قسم
ویب ایپلیکیشن
شعبہ
B2B SaaS

صورتحال

مسئلہ

One Node server handled everything — WebSocket connections, task mutations, reward calculations and database writes. It was already spiking at 5,000 concurrent users, and the target was 100,000 within twelve months. The reward algorithm was the worst of it: every completion, comment, mention, reaction and login triggered a multi-step scoring pass across a user's entire history, synchronously, at peak.

ہم نے کیا کیا

We took the reward engine off the critical path first, because it was the thing making every other interaction slow, then scaled the connection layer behind it.

وہ فیصلے جو اہم تھے

  • The reward engine became a queue, not a function call

    BullMQ over Redis, with a dedicated worker consuming events and writing XP to PostgreSQL in batches. The UI updates instantly and the arithmetic settles behind it within 200ms p99 — users were never waiting on the scoring, they were only ever waiting on the architecture.

  • WebSockets across instances instead of inside one

    In-memory Socket.IO rooms replaced with a Redis Pub/Sub adapter, sticky sessions at the ALB and heartbeat pruning. That is what turned one box into twelve EC2 instances sharing 100,000+ connections.

  • Leaderboards stopped being computed live

    Materialized views refreshed every five minutes, served from read replicas. A leaderboard that is five minutes stale is indistinguishable from a live one to a user, and enormously cheaper.

کیا بدلا

concurrent users
100K+concurrent users
API p95 latency
85msAPI p95 latency
user retention
+340%user retention
uptime
99.99%uptime
  • 100K+ — from 5,000 before the rebuild
  • 85ms — down from 1.2 seconds

The platform carries 100,000+ concurrent users and 2M+ events a day at 99.99% uptime, with p95 API latency down from 1.2 seconds to 85ms and retention up 340%.

استعمال شدہ خدمات

  • Event-driven reward engine on BullMQ with batched writes
  • Horizontally scalable WebSocket layer across twelve instances
  • Read-replica and materialized-view strategy for leaderboards
  • Load-test harness and capacity plan to 100K concurrent

ہم کیا مختلف کرتے

ہر منصوبے میں ایسی ایک بات ہوتی ہے۔ اسے شائع کرنا ہی اصل نکتہ ہے — جس کیس اسٹڈی میں کوئی پچھتاوا نہ ہو وہ ثبوت نہیں، تشہیر ہے۔

Knowing when to evolve an architecture without stopping product velocity is the actual skill. The original monolith was not a mistake — it was correct for its stage. The judgement is in avoiding premature optimisation and premature architectural pessimism at the same time, while continuing to ship features throughout.

شروع کرنے کے لیے تیار ہیں؟