Put your agent on call
One SDK, your neatlogs key, five minutes.
How it works
Nightly watches your agent through two channels: its neatlogs traces and the outcome events your agent reports. When a window of runs goes bad (failures, tool errors, escalations, runaway loops, cost spikes), it investigates. It reads the traces, the deploy log, the diff, and the Entire checkpoint with the coding-agent prompt behind the change. Then it prices the damage with your unit economics and decides under your policy.
Containment is reversible: roll back the release, fail over the model, flip a kill switch, cap steps. It happens through the SDK's control plane, which your agent reads on every run. Every action is verified by replaying the failing inputs through your agent. Anything irreversible (a code fix PR, refunds, contacting customers) waits for your approval in Slack.
1. Install
pip install nightly-sdk neatlogs export NIGHTLY_API_URL=https://api.your-nightshift-host export NIGHTLY_AGENT_TOKEN=nsa_... # from "Connect an agent" export NEATLOGS_API_KEY=... # your neatlogs project key
The SDK has zero dependencies and fails open. If Nightly is unreachable, your agent keeps its last known config or its defaults.
2. Trace with neatlogs
Wrap one run of your agent in a root span. Nightly polls your project for traces with that root span name.
import neatlogs
from openai import OpenAI
neatlogs.init(api_key=NEATLOGS_API_KEY, workflow_name="support-agent")
client = neatlogs.wrap(OpenAI()) # every LLM call becomes a span
with neatlogs.trace("support_conversation", kind="WORKFLOW"):
... # your agent loop; raise inside a tool span to mark it errored3. Read the control plane
These are the levers Nightly pulls during an incident. Read them at the start of each run (they're cached for 5 s).
from nightly_sdk import Nightly
ns = Nightly()
release = ns.release("v1") # rollback target
model = ns.model("claude-haiku") # model failover
max_steps = ns.max_steps(8) # runaway-loop cap
refunds = ns.flag("refunds_enabled", True) # kill switch for a tool4. Report outcomes and deploys
Grade each run however you already do (a rubric, a checker, a customer thumbs-down). Include the run details: the run then counts even if its trace is late, and the investigator can read the input.
ns.record("failed", value_usd=12.0, # money lost on this run, if any
input=message, output=reply, steps=n_steps, tool_errors=n_tool_errors,
tokens=usage.total_tokens, latency_ms=elapsed_ms, note="refund above policy")
# in CI, after a deploy (the commit lets Nightly find the Entire checkpoint)
ns.report_deploy("v2", commit=os.environ["GIT_SHA"])5. Let Nightly verify fixes
After it acts, Nightly sends the failing inputs back to your agent to re-run on the new config. Run the handler without side effects (no refunds, no emails) and say whether it passed.
def replay(message: str) -> dict:
result = run_agent(message, dry_run=True)
return {"ok": result.ok and result.tool_errors == 0, "output": result.text}
ns.serve_replays(replay) # background thread6. Set the policy
Each agent has a versioned YAML policy. Nightly acts alone only on actions listed under auto, and every one of them must be reversible.
gates:
min_confidence: 0.6 # diagnosis confidence needed to act alone
min_projected_loss_usd: 25 # below this, it waits for morning
verify_replay_min_pass_rate: 0.8
auto:
rollback_release: { reversible: true }
failover_model: { reversible: true, allowed_models: [nova-lite, claude-sonnet-4.6] }
disable_flag: { reversible: true, allowed_flags: [refunds_enabled] }
cap_steps: { reversible: true, min_steps: 4 }
needs_approval: [deploy_code_fix, refund_customers, contact_customers, delete_data]7. Paging
Add Nightly to Slack and it posts one message per incident with Approve fix and Undo buttons, updated as it works. Approving a code fix opens a revert PR on your repo with the postmortem and a regression eval. ntfy (phone push) and a generic webhook are there too.
Full example
Scout, a docs assistant in the repo (agents/scout/scout.py), is wired exactly this way. Nightly caught its bad release, rolled it back and verified the fix in under three minutes. See evidence/scout/.