Nightly
On-call for AI agents

The on-call engineer for your AI agents

Nightly watches every run of your agents. When one breaks at 3 AM, it rolls back the bad release, proves the fix, and asks you in Slack only for what can’t be undone.

Every incident cited
NEATLOGS TRACEENTIRE CHECKPOINTCODING PROMPTDEPLOY LOGRELEASE DIFFCFO.AI MODELPOLICY v3REPLAY 3/3POSTMORTEMREGRESSION EVALFIX PRNEATLOGS TRACEENTIRE CHECKPOINTCODING PROMPTDEPLOY LOGRELEASE DIFFCFO.AI MODELPOLICY v3REPLAY 3/3POSTMORTEMREGRESSION EVALFIX PR

The same bad release, with and without Nightly

Without Nightly
03:07Bad release goes live. Every request still returns 200.
03:07–08:12No alert fires. $22,846 is lost before anyone wakes up.
08:12Someone notices a spike on a dashboard.
11:30Postmortem by hand. Nobody finds the prompt that caused it.
With Nightly● recorded · benchmark

One SDK in your agent. Nightly does the rest.

Your agents
HS+◐◇⌘
Nightly
detect · trace · price · fix
Does on its own
↩ Roll back⇄ Switch model⏻ Kill switch⧗ Cap steps
Asks you in Slack
✎ Fix PR$ Refunds
Evidence
neatlogs
what happened
Entire
why it happened
cfo.ai
what it costs

Works with the tools you already use

Slack

One message per incident, updated live. Approve or undo from the channel.

GitHub

Sign in with GitHub. Approved fixes open as revert PRs on your repo.

neatlogs

Your agent's traces. Nightly reads them back while it investigates.

Entire

Checkpoints link a bad deploy to the coding-agent prompt behind it.

cfo.ai

Your unit economics put a price on every minute of an incident.

pyPython SDK

pip install nightly-sdk. Zero dependencies, fails open.

◔Phone push

ntfy notifications for the incidents that really need you.

{}Webhooks

Send pages to PagerDuty, Opsgenie or your own tools.

—
incidents handled correctly
—
from detection to fix
—
losses prevented
—
per investigation

loading…

01neatlogs

It investigates like an engineer.

Reads the failing traces, the deploy log and the diff, and cites every trace it used.

loading…
02Entire

It finds the prompt behind the bug.

Follows the bad deploy to its commit, its Entire checkpoint, and the coding-agent prompt that wrote it.

commit → checkpoint → prompt

“Make refunds one step: payments already validates eligibility.”

release r46
⚠ The coding agent warned about this. Nobody read it.
03cfo.ai + Slack

It only wakes you when it’s worth it.

Prices every minute of the incident. Fixes what’s reversible, then sends one message.

NightlyAPP3:08 AM
✅ Handled. No action needed tonight

Cause: release r46, written by a Claude Code session. Its Entire checkpoint shows the coding agent warned about it.

Contained: rolled back 18s after detection. Verified: 5/5 failed tickets now pass.

Saved: ~$22,846 that would have been lost before anyone woke up.

Harbor · INC-1001 · status MONITORING
Open in Nightly

Fixes what’s reversible. Asks about what isn’t.

On its own
  • ↩︎ Roll back a bad release
  • ⇄ Switch to a backup model
  • ⏻ Turn off a risky tool
  • ⧗ Stop a runaway loop
Asks you in Slack
  • ✎ Ship a code fix (opens a PR)
  • $ Refund or email customers
  • ✕ Delete data
  • … anything your policy doesn’t allow

Monitoring tells you. Nightly handles it.

Agent observabilityAI SRE toolsNightly
Watches what your AI agent does✓ traces, evals, alerts— watches servers✓ traces and outcomes
Finds the prompt behind a bad change——✓ via the Entire checkpoint
Puts a price on the incidentLLM spend only—✓ $/min from your unit economics
Fixes it while you sleep—infrastructure only✓ rollback, model failover, kill switch
Proves the fix worked—service health✓ replays the failed runs
Asks before anything irreversiblen/avaries✓ always, in Slack

Go to sleep. It’s handled.

Connect an agent in five minutes and add Nightly to Slack.

Slack GitHubneatlogsEntirecfo.aipy Python SDK◔ Phone push{} Webhooks Slack GitHubneatlogsEntirecfo.aipy Python SDK◔ Phone push{} Webhooks