Flagship project / AI + infrastructure
AI with access.
Humans with judgment.
I use AI assistance to administer Gaia, my multi-node Proxmox homelab. The interesting engineering problem isn't simply generating commands. It's defining permission boundaries, recognizing failures across system layers, and retaining a reliable way back.
01 / Context
The environment
Gaia combines compute nodes Sierra, K2, Fuji and Olympus. The lab spans containerized services, storage, network access, authentication, monitoring and local AI. Olympus provides the NVIDIA GPU capacity used by Ollama in a Linux container.
Conceptual architecture. This does not represent live infrastructure telemetry or publish the underlying access configuration.
02 / Migration journey
Change is a sequence of decisions.
- 01
Define the boundary
Set human-directed objectives and scoped permissions across the Proxmox servers. An agent recommendation is not the same as authorization to make unrestricted changes.
- 02
Plan for recovery
Build the backup and recovery approach before the major upgrade, including Proxmox Backup Server and a plan for portable NanoKVM console access.
- 03
Upgrade the cluster
Use AI assistance to navigate the Proxmox VE 8 to 9 migration, iterating through compatibility concerns and verifying services throughout.
- 04
Restore the workload
Preserve NVIDIA GPU access through host, device and LXC layers so that Ollama can continue using accelerated local inference.
- 05
Validate the result
Check the actual post-change behavior rather than treating successful commands or device visibility alone as proof that dependent services work.
03 / The hard parts
GPU access is a chain, not a checkbox.
The Ollama workload depended on compatible behavior across the Proxmox host, NVIDIA components, container access and the application itself. Troubleshooting needed to distinguish a detected device from actual usable model acceleration.
Observed engineering challenge
Maintaining local GPU inference through a hypervisor upgrade and related container/driver changes required iterative investigation across multiple layers.
Operating principle
Validate each dependency at its own boundary, then verify the end-to-end workload. A command that exits successfully is evidence of one step, not proof of the overall outcome.
04 / Recovery discipline
Build a way back.
Proxmox Backup Server was established as part of the recovery approach. A NanoKVM USB was planned as a portable crash cart for direct console access when the network administration path is unavailable; this is not a claim that a full disaster-recovery drill has been completed.
AI can accelerate investigation and execution. Ownership of scope, risk, verification and recovery remains with the operator.
05 / What I learned
Orchestration is engineering work.
Useful AI-assisted administration requires careful system context, limited permissions, explicit acceptance criteria and the willingness to interrogate an apparent success. Those habits transfer directly to technical product work: define intent, understand dependencies and make outcomes observable.
← All technical projects