Monitoring
Metrics
With the Metrics plugin registered, elections emit automatically:
| Metric | Labels | Meaning |
|---|---|---|
redisx_leader_elected_total | group | Times this instance became leader |
redisx_leader_lost_total | group, reason | Times it lost leadership (expired / stepdown / shutdown) |
A custom MetricsPlugin({ prefix }) renames the series accordingly (e.g. myapp_leader_elected_total) — adjust the queries below to your prefix.
Useful PromQL:
# Leadership churn per group (should be ~0 outside deployments)
sum by (group) (rate(redisx_leader_lost_total{reason="expired"}[15m]))
# Recent elections by instance (deployment/failover forensics)
sum by (instance, group) (increase(redisx_leader_elected_total[15m]))Counters cannot answer "who leads right now" — a stable leader elected an hour ago shows zero increase in any window. Read the current leader from the status endpoint below (getLeaderId()), or export a gauge from your own onElected/onLost callbacks.
Frequent reason="expired" losses outside deployments mean renewals are failing — look at Redis latency and event-loop blocking, or widen the ttlMs/renewIntervalMs gap.
Lifecycle Callbacks
import { Module } from '@nestjs/common';
import { RedisModule } from '@nestjs-redisx/core';
import { LeaderPlugin } from '@nestjs-redisx/leader';
import { alertOps } from './types';
@Module({
imports: [
RedisModule.forRoot({
clients: {
host: 'localhost',
port: 6379,
},
plugins: [
new LeaderPlugin({
// Extra elections started at bootstrap (besides 'default' and
// the groups referenced by @LeaderOnly decorators)
groups: ['reports', 'cleanup'],
// Fire-and-forget lifecycle callbacks: errors are logged and
// never break the election loop.
onElected: (group) => {
console.log(`This instance now leads "${group}"`);
},
onLost: (group, reason) => {
// reason: 'expired' | 'stepdown' | 'shutdown'
if (reason === 'expired') {
alertOps(`Unexpectedly lost leadership of "${group}"`);
}
},
}),
],
}),
],
})
export class AppModule {}Callbacks are fire-and-forget: a throwing listener is logged and never breaks the election loop.
onElected and onLost strictly alternate per group — every gained leadership is closed by exactly one loss event before the next election, so pairing them to start/stop a singleton worker is safe. A lease that expires without a tick observing it (event-loop stall, clock jump) reports onLost(group, 'expired') before any re-election, and a lapse watchdog fires that event on time even while the heartbeat's store call is still dangling on a hung connection — the stop signal does not wait for the socket to time out.
Log Lines
The service logs every transition with the instance identity:
[LeaderService] Instance "api-7f9c-1234" became leader of "default"
[LeaderService] Instance "api-7f9c-1234" lost leadership of "default" (stepdown)Set a stable instanceId (pod name) to make these greppable across restarts.
Status Endpoint
Expose leadership state for dashboards and debugging:
import { Injectable, Inject } from '@nestjs/common';
import { LEADER_SERVICE, ILeaderService } from '@nestjs-redisx/leader';
@Injectable()
export class LeadershipStatusService {
constructor(
@Inject(LEADER_SERVICE)
private readonly leaderService: ILeaderService,
) {}
// Synchronous local view — safe to call on every request.
isThisInstanceTheLeader(): boolean {
return this.leaderService.isLeader();
}
// Works on ANY instance: reads the election key from Redis.
async whoLeads(): Promise<string | null> {
return this.leaderService.getLeaderId();
}
// Execute singleton work programmatically (without a decorator).
async maybeCompact(): Promise<number | undefined> {
return this.leaderService.runIfLeader(async () => {
// ...heavy singleton work...
return 42;
});
}
status() {
return {
instanceId: this.leaderService.instanceId,
isLeader: this.leaderService.isLeader(),
groups: this.leaderService.getGroups(),
};
}
}The example app ships a ready-made variant at GET /demo/leader/status.
What to Alert On
redisx_leader_lost_total{reason="expired"}increasing outside deploy windows — unstable leadership.- No instance reporting
isLeader: truefor longer thanttlMs + retryIntervalMs— elections stuck (Redis down: by design nobody leads).
Next Steps
- Failover — expected timings during transitions
- Troubleshooting — diagnosing stuck or flapping elections