Skip to content

Monitoring ​

Metrics ​

With the Metrics plugin registered, elections emit automatically:

MetricLabelsMeaning
redisx_leader_elected_totalgroupTimes this instance became leader
redisx_leader_lost_totalgroup, reasonTimes it lost leadership (expired / stepdown / shutdown)

A custom MetricsPlugin({ prefix }) renames the series accordingly (e.g. myapp_leader_elected_total) — adjust the queries below to your prefix.

Useful PromQL:

promql
# Leadership churn per group (should be ~0 outside deployments)
sum by (group) (rate(redisx_leader_lost_total{reason="expired"}[15m]))

# Recent elections by instance (deployment/failover forensics)
sum by (instance, group) (increase(redisx_leader_elected_total[15m]))

Counters cannot answer "who leads right now" — a stable leader elected an hour ago shows zero increase in any window. Read the current leader from the status endpoint below (getLeaderId()), or export a gauge from your own onElected/onLost callbacks.

Frequent reason="expired" losses outside deployments mean renewals are failing — look at Redis latency and event-loop blocking, or widen the ttlMs/renewIntervalMs gap.

Lifecycle Callbacks ​

typescript
import { Module } from '@nestjs/common';
import { RedisModule } from '@nestjs-redisx/core';
import { LeaderPlugin } from '@nestjs-redisx/leader';
import { alertOps } from './types';

@Module({
  imports: [
    RedisModule.forRoot({
      clients: {
        host: 'localhost',
        port: 6379,
      },
      plugins: [
        new LeaderPlugin({
          // Extra elections started at bootstrap (besides 'default' and
          // the groups referenced by @LeaderOnly decorators)
          groups: ['reports', 'cleanup'],

          // Fire-and-forget lifecycle callbacks: errors are logged and
          // never break the election loop.
          onElected: (group) => {
            console.log(`This instance now leads "${group}"`);
          },
          onLost: (group, reason) => {
            // reason: 'expired' | 'stepdown' | 'shutdown'
            if (reason === 'expired') {
              alertOps(`Unexpectedly lost leadership of "${group}"`);
            }
          },
        }),
      ],
    }),
  ],
})
export class AppModule {}

Callbacks are fire-and-forget: a throwing listener is logged and never breaks the election loop.

onElected and onLost strictly alternate per group — every gained leadership is closed by exactly one loss event before the next election, so pairing them to start/stop a singleton worker is safe. A lease that expires without a tick observing it (event-loop stall, clock jump) reports onLost(group, 'expired') before any re-election, and a lapse watchdog fires that event on time even while the heartbeat's store call is still dangling on a hung connection — the stop signal does not wait for the socket to time out.

Log Lines ​

The service logs every transition with the instance identity:

[LeaderService] Instance "api-7f9c-1234" became leader of "default"
[LeaderService] Instance "api-7f9c-1234" lost leadership of "default" (stepdown)

Set a stable instanceId (pod name) to make these greppable across restarts.

Status Endpoint ​

Expose leadership state for dashboards and debugging:

typescript
import { Injectable, Inject } from '@nestjs/common';
import { LEADER_SERVICE, ILeaderService } from '@nestjs-redisx/leader';

@Injectable()
export class LeadershipStatusService {
  constructor(
    @Inject(LEADER_SERVICE)
    private readonly leaderService: ILeaderService,
  ) {}

  // Synchronous local view — safe to call on every request.
  isThisInstanceTheLeader(): boolean {
    return this.leaderService.isLeader();
  }

  // Works on ANY instance: reads the election key from Redis.
  async whoLeads(): Promise<string | null> {
    return this.leaderService.getLeaderId();
  }

  // Execute singleton work programmatically (without a decorator).
  async maybeCompact(): Promise<number | undefined> {
    return this.leaderService.runIfLeader(async () => {
      // ...heavy singleton work...
      return 42;
    });
  }

  status() {
    return {
      instanceId: this.leaderService.instanceId,
      isLeader: this.leaderService.isLeader(),
      groups: this.leaderService.getGroups(),
    };
  }
}

The example app ships a ready-made variant at GET /demo/leader/status.

What to Alert On ​

  • redisx_leader_lost_total{reason="expired"} increasing outside deploy windows — unstable leadership.
  • No instance reporting isLeader: true for longer than ttlMs + retryIntervalMs — elections stuck (Redis down: by design nobody leads).

Next Steps ​

Released under the MIT License.