The lab / Engineering project

A deliberate homelab architecture

Designing a resilient self-hosted environment that separates durable data from compute, reduces operational complexity, and supports experimentation without compromising production services.

This overview details the design and scaling of a self-hosted infrastructure lab covering virtualization, networking, storage, observability, and automated resilience.
Configuration details like credentials and specific network addresses are excluded for security

Overview

What began as a local storage and media repository has evolved into a deliberately structured infrastructure platform that supports mission-critical services, sandbox experimentation, automated workflows, and disaster recovery.

The architecture is anchored by a core operational principle: Monolith owns the data. Nexus consumes the data. By establishing Monolith as the authoritative storage layer while dedicating Nexus to primary compute and application workloads, the design enforces clear system boundaries. This separation ensures cleaner ownership domains, predictable recovery paths, and maximum flexibility when individual services require upgrades, rebuilds, or migration.

Goal / Problem

The primary challenge went far beyond hosting more applications; it meant engineering an infrastructure platform that could scale without introducing structural fragility.

The environment needed to centralize durable data assets, cleanly decouple storage from compute, host independently managed services, ensure robust monitoring and backup pipelines, minimize external attack surfaces, and streamline root-cause analysis and recovery.

Critically, the design had to support continuous, hands-on experimentation without degrading production workload stability.

Architecture

The environment is structured around two core systems with clearly segregated responsibilities:

Monolith/Unraid (Authoritative Storage Platform): Manages persistent media, backups, application states, and durable data, ensuring information survives the loss or replacement of any compute node.

Nexus/Proxmox (Production Compute Layer): Hosts virtual machines, containerized workloads, application services, monitoring pipelines, and automation engines that consume data from Monolith.

The supporting infrastructure integrates segmented networking, controlled public ingress, zero-trust remote access, automated backup orchestration, and power resilience. Public-facing services are exposed selectively via encrypted reverse proxies, while administrative control planes remain restricted to trusted private networks.

SYSTEM ARCHITECTURESANITIZED DIAGRAM

System architecture/dependency model

Technology Stack

The platform integrates storage, virtualization, enterprise networking, observability, security, automation, and recovery technologies into a cohesive, managed environment.

  • Core Infrastructure: Unraid (authoritative storage), Proxmox VE (virtualization and compute), and UniFi (network management, routing, and micro-segmentation), supported by isolated container and VM workloads alongside secure reverse-proxy and private-access tunnels.
  • Supporting Services: Media delivery, personal cloud applications, credential management, infrastructure metrics, uptime tracking, backup orchestration, and UPS-integrated power shutdown coordination.

Design Principle: Tool selection is strictly intentional; every component fulfills a defined architectural requirement within the broader system rather than serving as a disparate utility.

Implementation

I deployed the environment incrementally rather than through a high-risk, big-bang rollout.

The initial phase focused entirely on establishing authoritative storage and a disciplined data structure. Once secured, I decoupled compute workloads from the data plane, ensuring applications could be migrated, modified, or rebuilt without impacting underlying storage states. I layered in network micro-segmentation, secure ingress paths, telemetry, automated backups, and power resilience as platform maturity dictated. I migrated services individually, validated them in place, and stabilized them before introducing subsequent changes.

This staged execution adheres to a strict operating rule: verify the context, execute the smallest viable change, validate the outcome, and only then proceed.

Challenges

The most complex engineering obstacles rarely lived in a single application.

Stateful services introduced intricate dependencies linking compute nodes, storage arrays, network routing, storage mounts, access permissions, and boot sequencing. Enforcing strict filesystem path consistency became critical for reliable automation. Migrating active services required meticulous handling of persistent state, while balancing high availability against system overhead demanded careful triage. Furthermore, providing seamless remote access required constant vigilance to minimize external surfaces.

These hurdles reinforced a fundamental systems principle: true troubleshooting requires mapping the entire dependency chain, not inspecting components in isolation.

Decisions

As the environment matured, several core decisions became foundational to its stability:

  • Separation of Concerns: Persistent data is anchored in the Monolith, while Nexus handles execution. Disposable staging data is intentionally excluded from backup pipelines.
  • Environments: Production workloads are strictly segregated from experimental sandboxes, and administrative interfaces remain private by default.
  • Data Protection & Hygiene: Backups are isolated from the primary workloads they protect, credentials are systematically excluded from documentation, and modifications are applied via targeted, deliberate configuration changes rather than sweeping updates.

These guiding principles eliminate ambiguity during failures, ensuring absolute clarity across ownership domains, dependency chains, and recovery vectors.

Lessons Learned

The most resilient architecture is ultimately the one that makes operational responsibilities entirely self-evident.

Separating durable data from compute drastically reduced both technical complexity and recovery risk. Standardizing filesystem paths eliminated avoidable automation bottlenecks, while monitoring proved exponentially more valuable when it mapped system dependencies rather than merely reporting binary uptime states. Furthermore, the work reaffirmed a critical truth about data protection: a backup holds little value until the restore pathway is fully understood and empirically tested.

Above all, the project reinforced a core systems principle: true reliability stems from deliberate boundaries, comprehensive observability, and repeatable operating disciplines—never simply from adding more technology.

Current Status

The platform is fully operational and continuously evolving.

Monolith anchors the authoritative storage layer, while Nexus drives the primary production compute environment. Core capabilities—spanning media management, telemetry, automation, personal cloud apps, and infrastructure services—are isolated into manageable workloads, reinforced by robust backup and power-protection workflows.

Current R&D is intentionally restrained: focus has shifted away from feature bloat and additive services toward hardening resilience, optimizing off-site data protection, streamlining automation, tightening security postures, and simplifying day-to-day operations.

Planned visual documentation will include original rack photography, a sanitized architecture diagram, network and service relationships, backup and power-flow diagrams, and selected infrastructure-monitoring views.

Keep exploring.