Watchdog Mode // Alerts and one-tap fixes

Your servers. Checks keep running.

mttrly Watchdog runs configured checks, alerts when it detects a supported problem, and offers scoped button-tap actions. See the evidence before you choose the next move.

Free tier available now under current limits • Upgrade to Bro ($39/mo) when ready
ssh session terminated
> Connection lost. Reconnecting... Failed.
⚠ CRITICAL ALERT
Service 'nginx' is down on prod-server-01
// Switched to mttrly approval flow
/restart nginx
[mttrly] ✅ Command executed successfully.
Service 'nginx' is running.
Result verified and added to the audit trail.
$

ERROR: CONNECTION_TIMEOUT

Sound familiar? Infrastructure fails at the worst moment, and you're tied to your laptop.

01_PANIC

Failure on the go

Server crashed while you're in the subway or stuck in traffic. SSH from phone is torture, and clients are already sending angry emails.

02_ROUTINE

Wasted time

Opening laptop, connecting VPN, entering passwords just to restart a process or clean /tmp.

03_ANXIETY

Anxiety

Fear of stepping away from computer for too long because "something might break." You become hostage to your own code.

Watchdog has your back

  • >Crash alerts + one-tap fixes for common services
  • >Alerts when configured checks detect a supported problem
  • >Approve fixes with button taps, no typing
  • >Prepared actions — quick, visible, audited
SYSTEM_STATUS
Server AlphaONLINE
DatabaseONLINE
Redis QueueONLINE
Last check: Just now via mttrly

Watchdog features

> /status

All servers — one tap

CPU, RAM, disk, and services in one view. Check the current state without opening a terminal and reconstructing context from memory.

> /logs

Logs in your pocket

View logs right in your messenger. Filter errors, search by keywords. SSH client no longer needed.

> One-tap fix

Approved recovery

App crashed? Watchdog alerts you and offers a restart or repair action you approve from chat. OOM killed your process? Use the suggested fix without SSH.

> Triggers

Rules: if X — do Y

CPU > 90%? Clear cache from chat. Disk full? Run a cleanup script after approval. You stay in the loop for operational changes.

Initialization SETUP_PROTOCOL

STEP 01: Account

Sign up with email, then choose where approvals should reach you.

STEP 02: Agent

Install the Node.js agent on your server with a one-line installer.

curl -sL https://mttrly.com/install.sh | bash -s -- -t YOUR_TOKEN

STEP 03: Control

Server is online. Manage it from the dashboard, Telegram, or your MCP-enabled IDE.

SECURITY_AUDIT

Giving a bot access to your server is a big step. That's why security is built into every layer.

Approval gates and bounded sessions

Runtime policy separates read-only and approval-required operations. A separately user-authorized Investigation may bypass per-command confirmation only for mttrly_execute_command, within its server, inactivity, hard-cap, and action limits.

No incoming ports

The agent initiates an outbound TLS/WebSocket connection. No new inbound mttrly control port is required; outbound connectivity must be allowed.

Audit logs

Approval decisions, supported actions, and execution results are recorded. Coverage and look-back depend on the current runtime and plan.

Watchdog vs alternatives

ParameterWatchdogSSH Mobile ClientsGrafana / Prometheus
Mobile convenienceHigh (Chat)Low (Console)Medium (Dashboards)
Approved recoveryYes (one-tap fixes)NoNot the primary job
Active actionsYes (button taps)Yes (typing)No
Setup complexityEmail + outbound agentKeys, VPNExisting monitoring stack
Standalone release in preparation • Not public yet

The standalone agent is being prepared.

The hosted agent code is private today. We plan to publish a separate MIT-licensed standalone mode for one server and your own Telegram bot, but it is not ready for public installation yet.

  • Planned MIT license for the standalone agent release
  • Planned standalone mode without the Central service
  • Planned support for your own Telegram bot
  • Current private registry: 80+ built-in playbooks
See the current status
release-status.txt
# Available now
Hosted email signup
Required outbound agent
Public MCP client (MIT)
# In preparation
Standalone agent is not downloadable yet
Hosted agent code remains private

Watchdog and Deployment Bro FAQ

Further Reading

Site Reliability Engineering: How Google Runs Production Systems

Betsy Beyer, Chris Jones, Jennifer Petoff, Niall Richard Murphy · Book (Free Online)

The foundational SRE text defining MTTR, monitoring, alerting, and incident response practices used by Google and adopted industry-wide.

The Site Reliability Workbook: Practical Ways to Implement SRE

Betsy Beyer, Niall Richard Murphy, David K. Rensin, et al. · Book (Free Online)

Hands-on companion to the SRE Book with concrete implementation patterns for alerting, on-call, and incident management.

Accelerate: The Science of Lean Software and DevOps

Nicole Forsgren, Jez Humble, Gene Kim · Book

Research-backed evidence that MTTR is one of the four key metrics predicting software delivery performance and organizational outcomes.

The Phoenix Project: A Novel About IT, DevOps, and Helping Your Business Win

Gene Kim, Kevin Behr, George Spafford · Book

Narrative introduction to DevOps principles — why reducing Mean Time to Recovery matters more than preventing all failures.

Let Watchdog handle the routine.

Start free. Add AI features when you need them.

Free tier available now under current limits • Deployment Bro $39/mo • Beta: 75% off the first 3 months with code BETA75