Application Technical Spec

Reliable Notification System for 1M+ Users

A technical specification for a multi-channel notification system supporting push notifications, SMS, and email delivery with durable processing, idempotency, provider failover, retries, and observability.

Prepared by

Daniel Enwefah

Founder & Engineer, Axiom Labs Limited

Download Full Technical Spec PDF

Scale

1M+ users

Channels

Push, SMS, email

Reliability focus

No missed sends, no duplicate user-facing notifications

01

Problem Statement

The system needs to reliably send notifications to users across push, SMS, and email. These notifications may include transaction alerts, security updates, portfolio activity, payment confirmations, product updates, and operational messages.

At this scale, direct provider calls and a simple send-and-forget model are not sufficient. Providers fail, APIs timeout, queues redeliver jobs, and services can crash during processing. The platform must prevent missed notifications, avoid duplicate user-facing sends, and expose clear delivery state.

02

Goals

  • Reliable delivery using durable storage, queues, retries, and recovery workflows.
  • Duplicate prevention through idempotency keys, database constraints, and idempotent workers.
  • Multi-channel support for push, SMS, and email from one notification platform.
  • Provider failover for critical messages when a primary provider is unavailable.
  • Graceful degradation by prioritizing critical notifications over lower-priority messages.
  • Delivery tracking from creation to sent, delivered, failed, skipped, or manual review.
  • Operational visibility through logs, metrics, dashboards, and alerts.

03

Non-Goals

The first version is not intended to be a full customer engagement or marketing automation platform.

  • No drag-and-drop campaign builder.
  • No advanced customer segmentation or AI personalization.
  • No real-time chat or support messaging.
  • No complex A/B testing framework.
  • No full customer engagement suite.

04

Architecture Overview

Product services create notification requests through a Notification API or publish notification events. The notification platform validates the request, persists durable records, publishes queue jobs through an outbox, and processes delivery asynchronously.

Notification API
Notification database
Outbox publisher
Message queue
Channel workers
Provider adapters
Webhook processor
Reconciliation jobs
Dead-letter queue
Observability dashboards

05

Data Model

Notifications are stored before delivery is attempted. The database is the source of truth for notification state, recipient-channel state, provider references, retries, webhook events, and audit fields.

  • `notifications` stores the main notification request and idempotency key.
  • `notification_recipients` stores each user-channel delivery attempt.
  • `notification_outbox` stores durable events waiting to be published.
  • `notification_provider_events` deduplicates provider webhooks.
  • `user_notification_preferences` controls eligible user opt-ins and opt-outs.

06

Idempotency Strategy

The system assumes at-least-once delivery from queues and external providers. It does not depend on infrastructure-level exactly-once delivery.

Instead, it creates an effectively-once user experience through request-level idempotency keys, unique recipient-channel records, worker status checks before sending, provider webhook deduplication, and reconciliation for uncertain provider responses.

payment_success:user_123:transaction_789

07

Queue & Worker Design

Queues decouple notification creation from delivery and allow the system to control throughput, isolate failures, prioritize critical messages, and scale workers horizontally.

Workers are idempotent. Before sending, a worker loads the recipient record and exits if the status is already sent, delivered, skipped, or permanently failed.

Typical statuses include queued, processing, sent, delivered, pending_confirmation, retry_scheduled, failed_retryable, failed_permanent, skipped, and manual_review.

08

Provider Failover

Each channel should support primary and backup providers where practical. Critical notifications can fail over aggressively when a provider is unhealthy, while lower-priority messages can be delayed to control cost and reduce provider pressure.

Timeouts require special handling. A timeout may mean the provider accepted the request but failed to respond. Those sends should move to pending_confirmation and be reconciled before retrying.

09

Retry Strategy

Retries should be controlled, bounded, observable, and safe. Retryable failures include provider timeouts, network failures, 5xx responses, rate limits, temporary queue failures, and unknown provider responses.

Non-retryable failures include invalid destinations, unsubscribed users, blocked destinations, invalid push tokens, and template validation errors. Repeated failures move to manual review, failed_permanent, or the dead-letter queue depending on the error category.

10

Observability

Notification failures directly affect user trust, so the system needs dashboards and alerts that make delivery health visible without manual investigation.

  • Metrics: sends by channel, delivery rate, queue depth, latency, retry count, DLQ count, provider errors, and duplicate prevention count.
  • Alerts: critical queue delay, provider failure spikes, authentication failures, webhook failures, outbox lag, and DLQ growth.
  • Logs: notification ID, recipient ID, channel, provider, retry count, status, error category, and trace ID without raw sensitive user data.

11

Failure Scenarios

  • If the primary email provider is down, critical emails fail over while marketing emails are delayed.
  • If an SMS provider times out, the message moves to pending_confirmation before retry or failover.
  • If a queue delivers the same job twice, the worker checks state and skips the duplicate.
  • If the app crashes after creating a notification, the outbox event is still published after recovery.
  • If a provider sends a duplicate webhook, the provider event ID prevents reprocessing.
  • If the system is under heavy load, critical queues are prioritized and lower-priority sends are paused or delayed.

12

Tradeoffs

The design adds complexity through queues, outbox tables, state machines, provider adapters, retries, reconciliation jobs, and dead-letter queues. For financial products, that complexity is justified because missed or duplicate notifications can harm user trust.

The system does not promise infrastructure-level exactly-once delivery. It uses at-least-once infrastructure with application-level idempotency to provide an effectively-once user experience.

Conclusion

For over 1 million users, notification delivery should be treated as a reliability system. Durable storage, queues, idempotent workers, provider failover, retries, reconciliation, dead-letter queues, and observability reduce the risk of missed notifications and duplicate user-facing sends.

Download Full Technical Spec PDF