All writing

Writing

Idempotent WhatsApp Webhooks in NestJS: Receiving Messages You Can't Process Twice

May 20, 2026·8 min read

nestjswebhooksintegrationbackend

Meta's WhatsApp Cloud API has a rule that quietly shapes everything downstream of it: if your webhook endpoint doesn't answer with 200 fast enough, it assumes the delivery failed and sends the same message again. So the very first thing I had to accept building Matcha — the multi-tenant multichannel booking assistant I described in the API overview — is that the inbound side is not a stream of unique events. It's an at-least-once firehose, and a customer's "yes, book it" might arrive twice. The outbound provider port was the easy half; the hard half is receiving something an attacker might forge and the network might duplicate, and turning it into exactly one booking.

The two-endpoint contract Meta forces on you

A Cloud API webhook is really two HTTP routes living under one path, and Matcha's WhatsAppWebhookController implements both:

  1. A GET that Meta hits once, at subscription time, with hub.mode, hub.verify_token, and hub.challenge query params. You echo the challenge back as plain text only if the token matches.
  2. A POST that Meta hits forever after, carrying message and status payloads, signed with an X-Hub-Signature-256 header.

The verification handler is deliberately boring — match the mode, match the token, return the challenge:

@Get()
handleVerification(
  @Query('hub.mode') mode: string,
  @Query('hub.challenge') challenge: string,
  @Query('hub.verify_token') verifyToken: string,
): string {
  if (mode !== 'subscribe') throw new BadRequestException('Invalid mode');
  if (!this.webhookService.verifyToken(verifyToken)) {
    throw new ForbiddenException('Invalid verify token');
  }
  return challenge;
}

The controller carries a @SkipTenant() decorator, because there is no X-Business-Id header on a webhook — Meta doesn't know our tenant model. The business is resolved later, from the phone number the message arrived on. That's the inbound mirror of the header-only tenant rule: the one route where the tenant id can't come from a header is the one route that has to look it up instead.

Verifying the payload: HMAC, and why timing matters

The GET token proves we registered the webhook. It does nothing for the POST body — anyone who learns the URL could forge a booking. Meta signs every payload with an HMAC-SHA256 of the raw request bytes using the app secret, and the only way to verify it is to recompute the same HMAC and compare. Two details are load-bearing here:

validateSignature(rawBody: string, signatureHeader?: string): boolean {
  const { signingSecret } = this.securityConfig;
  if (!signingSecret) return true; // dev mode only — logs a loud warning
  if (!signatureHeader) return false;
 
  const expected = this.computeSignature(rawBody, signingSecret); // HMAC-SHA256 hex
  const candidate = signatureHeader.startsWith('sha256=')
    ? signatureHeader.slice(7) // Meta prefixes the hex with "sha256="
    : signatureHeader;
 
  const expectedBuffer = Buffer.from(expected, 'hex');
  const candidateBuffer = Buffer.from(candidate, 'hex');
  if (expectedBuffer.length !== candidateBuffer.length) return false;
  return crypto.timingSafeEqual(expectedBuffer, candidateBuffer);
}

A mismatch throws ForbiddenException before a single byte is parsed. The secrets themselves (WA_VERIFY_TOKEN, WA_APP_SECRET) live only in env vars and are read once into a securityConfig at construction — they never touch a log line or a response body.

A parser port, so the payload shape stays at the edge

Once the body is trusted, it still has to be understood — and Meta's envelope is a deeply nested entry[].changes[].value.messages[] structure that I did not want leaking into the conversation pipeline. So parsing sits behind a port, exactly like the outbound provider does in ports and adapters. The port's whole job is to flatten any provider's payload into two normalized arrays:

export interface WhatsAppWebhookParserPort {
  parseMessages(raw: unknown): WhatsAppInboundMessage[];
  parseStatuses(raw: unknown): WhatsAppInboundStatus[];
}
 
export const WHATSAPP_WEBHOOK_PARSER = Symbol('WHATSAPP_WEBHOOK_PARSER');

The application service depends only on that contract. Whether the bytes came from Meta or from MockWaParser (a flat { messages: [...], statuses: [...] } shape used in dev and tests, with no network and no app secret), the webhook service sees the same WhatsAppInboundMessage — and crucially, the same waMessageId. That field is the linchpin of everything that follows.

Idempotency: the unique index is the real guarantee

Here is the part I got wrong the first time. My instinct was to dedupe in code: look up the message, and if it exists, skip. Matcha does keep that read-check, because it's cheap and handles the common case:

async createIncomingFromWebhook(params: { /* ... */ waMessageId: string }) {
  const existing = await this.repo.findByWhatsappId(params.waMessageId);
  if (existing) {
    this.logger.debug(`Duplicate message ${params.waMessageId}, returning existing`);
    return { id: existing._id.toString(), waMessageId };
  }
  // ...create the message
}

But a read-then-write check has a window: two retried deliveries of the same message can both pass the findByWhatsappId check before either writes, and you get two rows. So the actual guarantee lives in the database, as a unique partial index on the WhatsApp message id:

MessageSchema.index(
  { whatsappMsgId: 1 },
  {
    unique: true,
    partialFilterExpression: { whatsappMsgId: { $type: 'string' } },
    background: true,
  },
);

The partialFilterExpression matters: outgoing messages and internal records have no whatsappMsgId, and a plain unique index would treat all those nulls as colliding. The partial filter scopes uniqueness to documents that have a string id, so only real WhatsApp messages compete for it.

On the webhook path, that read-check plus the index is the whole story. If two retries both slip past the read-check, the second insert collides with the unique index and throws a duplicate-key error — which the controller's always-200 handler quietly swallows. No second row, no second booking. Elsewhere in the codebase — the inbound message-processor pipeline — the same idea wears a tidier form: an atomic upsert that turns the losing racer into a clean no-op instead of a thrown error, and reports which call actually did the insert:

async claimIncomingOnce(dto: ClaimIncomingDto) {
  const res = await this.model.findOneAndUpdate(
    { whatsappMsgId: dto.whatsappMsgId },
    { $setOnInsert: { /* user, chatId, content, direction: 'incoming', ... */ } },
    { upsert: true, new: true, includeResultMetadata: true },
  );
  const wasUpserted = Boolean(res.lastErrorObject?.upserted);
  return { doc: res.value, wasUpserted };
}

Two concurrent retries both run this; the index lets exactly one win the insert. The winner gets wasUpserted: true and is allowed to trigger the AI pipeline. The loser gets the existing document and wasUpserted: false, and does nothing further — no second booking, no second credit charged.

Idempotency isn't "remembering not to repeat yourself." It's making the database refuse to let you, so a retry is a no-op even when two of them arrive at the same millisecond.

ConcernMechanismWhere it lives
Webhook is genuinely ourshub.verify_token match on GETWhatsAppWebhookService.verifyToken
Payload is genuinely Meta'sHMAC-SHA256 over raw body, timing-safevalidateSignature + rawBody: true
Payload shape stays at the edgeWhatsAppWebhookParserPortparser port + MockWaParser
A duplicate never double-booksread-check, then a unique partial indexfindByWhatsappId + the whatsappMsgId index

Dispatching into the conversation pipeline

A verified, deduplicated message still has work to do: handleWebhook resolves the businessId from the receiving phone via WhatsAppSettingsRepositoryPort.findByPhone, drops the message if no tenant owns that number, finds-or-creates the user, upserts the open chat, writes the incoming message, and bumps the unread count. Statuses (sent, delivered, read, failed) take a parallel path that updates an existing outgoing message and is a no-op if the id is unknown. And one more deliberate choice: when application processing throws, the controller still returns 200. An exception that returned 500 would just earn another retry of a message we already stored — so we swallow it, log it, and acknowledge.

What this buys, and what it costs

The payoff is that the scariest inbound failures become impossible rather than merely unlikely: a forged payload dies at the signature check, and a retried delivery dies at the unique index. The booking pipeline downstream can be written as if every message is unique and trusted, because the edge already guaranteed it. The cost is real, though — rawBody: true and a separate raw-body path are easy to forget and silently break verification; the partial-index filter is a subtlety you only discover when a second null collides in staging; and the read-check plus the unique index is, honestly, two mechanisms for one guarantee on the webhook path. I keep both on purpose: the read handles the 99% cheaply, and the index handles the 1% that would otherwise quietly double-book a customer. For an integration where the other side will retry you whether you're ready or not — and Meta's webhook docs are clear that it will — that redundancy is the difference between a calm conversation and two appointments for the same slot.