← Back to Projects

Private AI Infrastructure

Building a Private AI Environment

A hands-on infrastructure project focused on local language models, GPU acceleration, Linux administration, containers, remote access, and security-minded AI workflows.

Ubuntu Server LTSNVIDIA RTX 5070 TiDockerOllamaOpen WebUILocal LLMsSSH10 GbE

Overview

Why I Built It

I wanted a private environment where I could run local language models, experiment with GPU-accelerated workloads, and learn how AI infrastructure behaves outside of hosted cloud services.

The project also gives me a practical environment for Linux, Docker, networking, remote administration, security testing, troubleshooting, and automation.

My goal is not simply to run models locally. I want to understand the full system around them — hardware, operating system, networking, security, deployment, monitoring, and recovery.

Architecture

System Design

01

Compute

Intel Xeon E5-2699 v4 server platform with NVIDIA RTX 5070 Ti acceleration for local AI workloads.

02

Operating System

Headless Ubuntu Server environment administered remotely through SSH.

03

AI Runtime

Ollama and containerized services for running, testing, and managing local language models.

04

Network

10 GbE networking via the TP-Link TX401, with remote administration and future segmentation in mind.

Hardware

Current Server Configuration

Dedicated AI Server

CPU

Intel Xeon E5-2699 v4

Motherboard

ASUS X99-DELUXE II

Memory

128 GB DDR4

GPU

PNY GeForce RTX 5070 Ti

Networking

TP-Link TX401 10 GbE

Operating System

Ubuntu Server LTS

Storage

Multiple NVMe SSDs

AI Stack

Ollama · Docker · Open WebUI

Implementation

Building the Stack

The environment is built in layers so each component can be tested independently and maintained without treating the system as one large application.

Current State

The dedicated AI server is built around the ASUS X99-DELUXE II, Intel Xeon E5-2699 v4, 128 GB of DDR4 memory, and an NVIDIA RTX 5070 Ti. Core Linux, GPU acceleration, model runtime, container, and remote administration workflows are established, with additional monitoring and security automation continuing to evolve.

Implemented

Linux Foundation

Configured a headless Ubuntu Server environment with remote administration and a stable network baseline.

Implemented

GPU Acceleration

Configured NVIDIA drivers and verified RTX 5070 Ti GPU availability for local AI workloads.

Implemented

Model Runtime

Installed Ollama and tested local language models directly from the server.

Implemented

Container Services

Using Docker for supporting services and Open WebUI for browser-based model access and administration.

Implemented

Remote Administration

Built the system to operate headlessly through SSH and remote management tools.

Implemented

Dedicated AI Hardware

Deployed the ASUS X99-DELUXE II and Intel Xeon E5-2699 v4 platform with 128 GB DDR4 and an RTX 5070 Ti.

Planned

Security Automation

Add security-focused workflows for monitoring, event analysis, response assistance, and repeatable administration.

Planned

Observability

Expand logging, performance monitoring, GPU utilization tracking, and service health visibility.

Troubleshooting

Problems Became Part of the Project

The build was not just about getting services online. Each failure became an opportunity to isolate layers, validate assumptions, and improve the environment.

Troubleshooting Approach

I separate hardware, networking, operating system, drivers, containers, and model runtime issues instead of treating every failure as an application problem.

01

Docker Permissions

Problem

Open WebUI deployment failed because the account could not access the Docker socket.

Response

Isolated the problem as a permissions issue rather than an application failure and worked through Docker group and service access.

02

DNS Resolution

Problem

The server experienced name-resolution problems even while basic network connectivity remained available.

Response

Separated DNS from general connectivity and validated network, resolver, and operating-system configuration independently.

03

GPU Validation

Problem

AI services depend on more than the GPU simply appearing in the system.

Response

Verified NVIDIA driver availability and confirmed that the RTX 5070 Ti could actually be used by local AI workloads.

04

Model Validation

Problem

A successfully loaded model does not automatically mean the environment is useful.

Response

Tested models with logic questions, technical prompts, and infrastructure scenarios to evaluate behavior and practical usefulness.

05

Headless Administration

Problem

The server needs to remain manageable without relying on a local desktop environment.

Response

Built the system around SSH and remote administration so services can be maintained, restarted, and diagnosed remotely.

06

Layered Diagnosis

Problem

AI infrastructure failures can originate from hardware, drivers, networking, containers, or the model runtime.

Response

Used a layer-by-layer troubleshooting process to reduce guesswork and identify the actual point of failure.

“The goal is not just to make the system work. It is to understand why it failed, how it recovered, and how to make the next failure easier to diagnose.”

Security

Security Built Into
the Environment

I designed the environment with the expectation that remote administration, local AI services, containers, and future automation all create security considerations that need to be addressed from the beginning.

Design Principle

Security is treated as part of the infrastructure design, not as a separate layer added after the system is already running.

Security Priorities

Limit privileged access

Keep management paths controlled

Segment services where practical

Maintain visibility through logging

Document recovery and rebuild procedures

01

Remote Administration

Designed the server for headless administration through controlled remote access instead of depending on a local desktop session.

02

Access Control

Separate administrative access from application access and limit privileged operations to the accounts and services that require them.

03

Service Isolation

Use containers and separated services to reduce unnecessary dependencies and keep individual components easier to manage and secure.

04

Network Segmentation

Plan the AI environment around dedicated network boundaries so management interfaces and AI services can be separated from general client traffic.

05

System Hardening

Reduce unnecessary services, maintain the operating system and drivers, and keep the server focused on its intended role.

06

Logging & Monitoring

Build toward centralized visibility for system health, authentication activity, container events, GPU utilization, and service failures.

07

Local Data Control

Keep selected AI workloads and data processing inside infrastructure I control instead of automatically sending them to a hosted model provider.

08

Recovery & Documentation

Document configuration changes and deployment steps so services can be rebuilt, troubleshot, and recovered without relying on memory.

Protect

Reduce unnecessary exposure through controlled access, hardened services, and network boundaries.

Detect

Improve visibility into authentication, service health, failures, and unusual system behavior.

Recover

Keep the environment documented and repeatable so failed components can be rebuilt with less downtime and guesswork.

“The objective is not to make a lab complicated. It is to make every layer understandable, controllable, and recoverable.”

Lessons Learned

What the Build Taught Me

The most useful lessons came from treating the server as a complete infrastructure system rather than just a machine for running models.

01

Start with a stable hardware and network baseline before adding AI services.

02

Treat GPU drivers, containers, model runtimes, and remote access as separate layers when troubleshooting.

03

Document configuration changes so failures are easier to reverse and systems are easier to rebuild.

04

Build secure remote administration into the environment from the beginning.

Key Takeaway

The project reinforced that successful AI infrastructure depends just as much on systems administration, networking, security, documentation, and recovery as it does on model performance.

Next Steps

Where the Lab Goes Next

The platform is intentionally expandable. Future work will focus less on adding hardware and more on improving visibility, automation, repeatability, and practical security use cases.

01

Expand model testing and benchmarking

02

Improve monitoring and logging

03

Add more security-focused automation

04

Document repeatable deployment workflows

05

Publish a video walkthrough

Project Media

Video walkthrough coming later

A future walkthrough will document the hardware, deployment process, model environment, remote administration, and security architecture.

More Projects

Explore the rest of the lab.

Back to Projects →