Your servers. Checks keep running.
mttrly Watchdog runs configured checks, alerts when it detects a supported problem, and offers scoped button-tap actions. See the evidence before you choose the next move.
Service 'nginx' is down on prod-server-01
Service 'nginx' is running.
Result verified and added to the audit trail.
ERROR: CONNECTION_TIMEOUT
Sound familiar? Infrastructure fails at the worst moment, and you're tied to your laptop.
Failure on the go
Server crashed while you're in the subway or stuck in traffic. SSH from phone is torture, and clients are already sending angry emails.
Wasted time
Opening laptop, connecting VPN, entering passwords just to restart a process or clean /tmp.
Anxiety
Fear of stepping away from computer for too long because "something might break." You become hostage to your own code.
Watchdog has your back
- >Crash alerts + one-tap fixes for common services
- >Alerts when configured checks detect a supported problem
- >Approve fixes with button taps, no typing
- >Prepared actions — quick, visible, audited
Watchdog features
> /status
All servers — one tap
CPU, RAM, disk, and services in one view. Check the current state without opening a terminal and reconstructing context from memory.
> /logs
Logs in your pocket
View logs right in your messenger. Filter errors, search by keywords. SSH client no longer needed.
> One-tap fix
Approved recovery
App crashed? Watchdog alerts you and offers a restart or repair action you approve from chat. OOM killed your process? Use the suggested fix without SSH.
> Triggers
Rules: if X — do Y
CPU > 90%? Clear cache from chat. Disk full? Run a cleanup script after approval. You stay in the loop for operational changes.
Initialization SETUP_PROTOCOL
STEP 01: Account
Sign up with email, then choose where approvals should reach you.
STEP 02: Agent
Install the Node.js agent on your server with a one-line installer.
STEP 03: Control
Server is online. Manage it from the dashboard, Telegram, or your MCP-enabled IDE.
SECURITY_AUDIT
Giving a bot access to your server is a big step. That's why security is built into every layer.
✓ Approval gates and bounded sessions
Runtime policy separates read-only and approval-required operations. A separately user-authorized Investigation may bypass per-command confirmation only for mttrly_execute_command, within its server, inactivity, hard-cap, and action limits.
✓ No incoming ports
The agent initiates an outbound TLS/WebSocket connection. No new inbound mttrly control port is required; outbound connectivity must be allowed.
✓ Audit logs
Approval decisions, supported actions, and execution results are recorded. Coverage and look-back depend on the current runtime and plan.
Watchdog vs alternatives
| Parameter | Watchdog | SSH Mobile Clients | Grafana / Prometheus |
|---|---|---|---|
| Mobile convenience | High (Chat) | Low (Console) | Medium (Dashboards) |
| Approved recovery | Yes (one-tap fixes) | No | Not the primary job |
| Active actions | Yes (button taps) | Yes (typing) | No |
| Setup complexity | Email + outbound agent | Keys, VPN | Existing monitoring stack |
The standalone agent is being prepared.
The hosted agent code is private today. We plan to publish a separate MIT-licensed standalone mode for one server and your own Telegram bot, but it is not ready for public installation yet.
- ✓Planned MIT license for the standalone agent release
- ✓Planned standalone mode without the Central service
- ✓Planned support for your own Telegram bot
- ✓Current private registry: 80+ built-in playbooks
Watchdog and Deployment Bro FAQ
Further Reading
Site Reliability Engineering: How Google Runs Production Systems ↗
Betsy Beyer, Chris Jones, Jennifer Petoff, Niall Richard Murphy · Book (Free Online)
The foundational SRE text defining MTTR, monitoring, alerting, and incident response practices used by Google and adopted industry-wide.
The Site Reliability Workbook: Practical Ways to Implement SRE ↗
Betsy Beyer, Niall Richard Murphy, David K. Rensin, et al. · Book (Free Online)
Hands-on companion to the SRE Book with concrete implementation patterns for alerting, on-call, and incident management.
Accelerate: The Science of Lean Software and DevOps
Nicole Forsgren, Jez Humble, Gene Kim · Book
Research-backed evidence that MTTR is one of the four key metrics predicting software delivery performance and organizational outcomes.
The Phoenix Project: A Novel About IT, DevOps, and Helping Your Business Win
Gene Kim, Kevin Behr, George Spafford · Book
Narrative introduction to DevOps principles — why reducing Mean Time to Recovery matters more than preventing all failures.
Let Watchdog handle the routine.
Start free. Add AI features when you need them.
Free tier available now under current limits • Deployment Bro $39/mo • Beta: 75% off the first 3 months with code BETA75