Production-style Linux operations lab focused on configuration management, application delivery, infrastructure monitoring, incident response, access administration, log management, rollback, and observability.
All failures and incidents in this repository were deliberately reproduced in an isolated lab. They are not production incidents or customer data.
- 3 Ubuntu VMs on KVM/libvirt: application, monitoring, and isolated SNMP target nodes
- Ansible roles for Linux baseline configuration, application deployment, monitoring, developer access, log rotation, and SNMP
- Docker + Nginx + FastAPI application path
- Zabbix + Grafana + PostgreSQL monitoring stack
- Custom Zabbix low-level discovery (LLD) with filesystem item and trigger prototypes
- SNMPv2 monitoring with a dedicated Linux target, availability alerting, and network discovery
- Grafana operations dashboard + Grafana-managed filesystem alerting backed by the Zabbix API
- 5 documented application/operations incidents plus observability failure-and-recovery drills
- Deployment health validation + rollback using immutable image tags
- GitHub Actions validation for Ansible and exported Zabbix templates
WSL2 Ubuntu
Ansible Control Node
|
| SSH
+---------------+----------------+
| |
v v
devops-app-01 devops-monitor-01
Ubuntu 22.04 Ubuntu 22.04
| |
| +-- PostgreSQL
| +-- Zabbix Server
| +-- Zabbix Web
| +-- Grafana
|
+-- Nginx :80
+-- Docker
+-- FastAPI
+-- Zabbix Agent
+-- logrotate
|
| SNMPv2 / UDP 161
v
devops-snmp-01
Ubuntu 22.04
snmpd target
Virtual machines are hosted with KVM/libvirt. Ansible manages the desired state over SSH. Zabbix monitors the application node with an agent and the dedicated SNMP node over UDP/161.
| Area | Technologies |
|---|---|
| Linux / virtualization | Ubuntu, KVM, libvirt, cloud-init |
| Configuration management | Ansible roles, variables, handlers |
| Application | FastAPI, Python |
| Runtime / proxy | Docker, Nginx |
| Monitoring | Zabbix Server, Zabbix Agent, SNMPv2, Grafana |
| Observability | Zabbix LLD, item prototypes, trigger prototypes, discovery actions, Grafana dashboards, Grafana-managed alerting |
| Data | PostgreSQL |
| Operations | SSH, systemd, logrotate, journalctl, SNMP CLI |
| Delivery validation | Ansible post-deployment health checks, GitHub Actions |
The monitoring environment was extended beyond basic host availability to demonstrate reusable monitoring design and alert-quality work.
The repository includes an exportable Zabbix 7.0 template:
monitoring/zabbix/templates/template-linux-filesystem-capacity.yaml
It contains:
- filesystem low-level discovery using
vfs.fs.discovery - a
vfs.fs.size[{#FSNAME},pused]item prototype - a reusable
{$FS.PUSED.WARN}threshold macro - a sustained 10-minute trigger prototype to reduce short-spike noise
- component, mount, and capacity tags
A third VM, devops-snmp-01, is configured through the snmp_target Ansible role.
The SNMP template collects:
- system name
- system uptime
- interface count
- no-data detection for lost SNMP telemetry
The community string is supplied through the SNMP_COMMUNITY environment variable and is not stored in the repository.
Zabbix network discovery was configured for the lab subnet to locate the SNMP target. A discovery action then associates the discovered device with the Linux servers host group and the custom SNMP template.
Grafana reads operational data through the Zabbix API using a dedicated API token rather than a stored administrator password.
The dashboard combines:
- root filesystem utilization
- SNMP system uptime
- SNMP interface count
- recent infrastructure problems
A Grafana-managed rule named High Root Filesystem Utilization evaluates the Zabbix-backed root filesystem metric for devops-app-01.
Normal operating configuration:
- threshold: above 85%
- evaluation interval: 1 minute
- pending period: 2 minutes
- keep-firing period: 0 seconds
For controlled validation, the threshold was temporarily lowered to 20% and the pending period to 0 seconds while the filesystem was at approximately 28.7%. This safely forced the rule into Firing without filling the disk. The rule was then restored to 85% / 2m, and Grafana returned it to Normal.
Notification transport was intentionally left out of scope; this drill validates Grafana rule evaluation, state transition, and recovery.
Full implementation notes and evidence: Observability expansion.
| Incident | Scenario | Detection / evidence | Recovery |
|---|---|---|---|
| INC001 | Application container outage | Zabbix HIGH alert and failed health check | Restarted container; HTTP 200 and Zabbix recovery |
| INC002 | Invalid Nginx deployment | Ansible handler failed on nginx -t |
Restored valid configuration without reloading the bad one |
| INC003 | Developer SSH access failure | Permission denied (publickey) and missing authorized_keys |
Reapplied Ansible access role and validated SSH login |
| INC004 | Excessive log growth / disk pressure | df, du, and file-size investigation |
Added Ansible-managed logrotate policy and validated rotation |
| INC005 | Broken application release | Failed deployment health check, HTTP 502, Zabbix HIGH alert | Rolled back to known-good v1; HTTP 200 and automatic alert recovery |
INC005 models a production-style release regression.
A known-good image was preserved as:
devops-demo-api:v1
A deliberately broken release was then deployed:
devops-demo-api:v2-broken
The container started, but the application listened on port 9000 while the Docker/Nginx path expected port 8000.
container running
|
+--> Uvicorn listening on :9000
|
Nginx / deployment path expects :8000
|
+--> HTTP 502
|
+--> Zabbix HIGH alert
Investigation used docker ps, docker inspect, docker logs, and curl.
The release was rolled back with the Ansible deployment playbook to devops-demo-api:v1. The health endpoint returned HTTP 200 and Zabbix recorded recovery.
Provides the Linux baseline across managed hosts: administration packages, operations group membership, lab directories, and environment identification.
Configures Docker, Nginx, FastAPI deployment, Nginx validation, container lifecycle, and application health checks.
Automates account provisioning, SSH authorized keys, access revocation, process termination, and home-directory cleanup.
Deploys PostgreSQL, Zabbix Server, Zabbix Web, and Grafana.
The database password is supplied with the ZABBIX_DB_PASSWORD environment variable and is not committed.
Configures agent-based monitoring on the application host.
Configures the isolated SNMP target with snmpd, restricted read-only access, service management, and UDP/161 verification.
Controls application log growth with size-based rotation, retention, compression, and copytruncate.
Application path
Client -> Nginx :80 -> Docker :8000 -> FastAPI /health
|
+-> Zabbix web scenario
+-> problem / recovery events
Host metrics
devops-app-01 -> Zabbix Agent -> Zabbix Server
SNMP telemetry
devops-snmp-01 :161/udp -> Zabbix Server -> Grafana
The observability extension includes recruiter-facing evidence for:
- filesystem discovery and prototypes
- filesystem alert and recovery
- live SNMP data
- SNMP outage detection
- SNMP recovery
- network discovery
- final Grafana infrastructure dashboard
- Grafana-native alert firing and recovery history
See docs/evidence/observability/.
linux-devops-operations-lab/
├── .github/workflows/
│ └── ansible-validation.yml
├── ansible/
│ ├── roles/
│ │ ├── common/
│ │ ├── app_server/
│ │ ├── developer_access/
│ │ ├── monitoring_stack/
│ │ ├── zabbix_agent/
│ │ ├── snmp_target/
│ │ └── log_management/
│ ├── inventory.example.ini
│ ├── site.yml
│ ├── developer-access.yml
│ └── deploy-release.yml
├── app/
├── cloud-init/
├── docs/
│ ├── observability-expansion.md
│ └── evidence/observability/
├── incidents/
│ ├── INC001-application-outage/
│ ├── INC002-ansible-nginx-deployment-failure/
│ ├── INC003-developer-ssh-access-failure/
│ ├── INC004-log-growth-disk-pressure/
│ └── INC005-production-deployment-regression/
├── monitoring/zabbix/templates/
│ ├── template-linux-filesystem-capacity.yaml
│ └── template-snmp-linux-lab.yaml
├── .env.example
├── .gitignore
└── README.md
- Linux/WSL2 control node
- Ansible
- KVM/libvirt
- three reachable Ubuntu VMs
- SSH key-based access
- Docker-capable monitoring/application guests
Copy the example inventory:
cp ansible/inventory.example.ini ansible/inventory.iniUpdate the VM IP addresses and SSH key path for your environment.
Set secrets locally:
export ZABBIX_DB_PASSWORD='replace-with-your-local-password'
export SNMP_COMMUNITY='replace-with-a-non-default-lab-value'Do not commit the real values.
cd ansible
ansible-playbook -i inventory.ini site.yml --syntax-check
ansible-playbook -i inventory.ini site.ymlVerify connectivity:
ansible all -i inventory.ini -m pingValidate the application:
curl http://<APP_VM_IP>/healthExpected:
{"status":"ok"}Deploy a known-good release:
cd ansible
ansible-playbook -i inventory.ini deploy-release.yml \
-e release_image=devops-demo-api:v1The playbook performs a post-deployment health check and fails the deployment workflow if the application does not return HTTP 200.
Rollback uses the same playbook with the previous known-good image tag.
GitHub Actions validates:
- the main Ansible playbook
- the release deployment playbook
- the exported Zabbix 7.0 template YAML files
The CI workflow does not attempt to reproduce KVM/libvirt or destructive incident scenarios.
The public repository intentionally excludes private SSH keys, local inventory, .env files, API tokens, SNMP community secrets, and database passwords. Example files contain placeholders only.
- Linux administration and troubleshooting
- KVM/libvirt virtualization
- Ansible configuration management
- Docker and Nginx operations
- Zabbix agent monitoring
- custom Zabbix templates
- low-level discovery and prototypes
- alert-threshold tuning
- SNMP monitoring and service-loss detection
- Zabbix network discovery and discovery actions
- Grafana dashboarding through the Zabbix API
- Grafana-managed alert rules and firing/recovery validation
- incident investigation and recovery validation
- Linux user and SSH access management
- log and disk troubleshooting
- deployment health checks and rollback
This repository is a controlled technical lab built to demonstrate hands-on Linux, monitoring, and DevOps operations. It does not claim production ownership, customer incidents, or high-availability production experience.


