Product engineering · Kubernetes platform

Production ecosystem

RedoxForge Ecosystem

A public information platform backed by self-managed Kubernetes, GitOps delivery, persistent storage, and isolated game services.

Interactive project overview

Explore the system in operation.

These demonstrations present the project's state, timing, and key tradeoffs. The full case study and implementation details follow below.

01

Argo CD namespace explorer

Inspect four resource trees and watch Next.js pods spawn, pass readiness, receive traffic, and replace the previous revision without interrupting availability.

Preparing cluster topology…

Problem and objectives.

Players relied on Discord for announcements, game mechanics, and staff support. Durable information was difficult to navigate, and people without an account could not access it.

The first dedicated-server deployment solved the content problem but introduced manual container updates, brief outages, and growing operational overhead as more applications and stateful services were added.

System architecture.

Explore the public product, web request path, GitOps delivery, game workloads, persistent services, and network-policy boundaries.

Interactive architecture

RedoxForge product and platform

The public website and game services share a delivery platform, but use separate request paths, workloads, and state boundaries.

Network Data Approval
PRODUCTPublic information with privileged editing and managed application data.
Public information with privileged editing and managed application data.01Players & visitorsPublic access02Next.js content UIPublic product03Draft.js editorPrivileged product04SupabasePostgreSQL + auth

Selected component

Players & visitors

Announcements, wiki content, and support guidance are readable without requiring a Discord account.

Accessible component list

My responsibilities.

The components I designed, implemented, tested, or integrated.

  1. 01

    Designed the RedoxForge website around public announcements, a wiki, support guidance, user profiles, and authenticated staff editing backed by Supabase.

  2. 02

    Containerized the application and built GitHub Actions workflows that publish multi-architecture images after changes reach the main branch.

  3. 03

    Architected and self-managed a four-node Kubernetes environment after managed cloud offerings proved outside the available budget.

  4. 04

    Used Argo CD and digest-based image reconciliation to replace manual Docker Compose refreshes and reduce deployment interruption.

  5. 05

    Provisioned stateful game, PostgreSQL, and SFTP workloads with Longhorn-backed claims and redundant object-storage backups.

  6. 06

    Defined ingress, service-account, RBAC, and Calico policy boundaries while excluding private addresses and connection details from this case study.

Key engineering decisions.

The constraints and tradeoffs that shaped the implementation.

01

Make information public by default

Constraint. The existing information source required an account and mixed durable guidance with real-time conversation.

Decision. Publish announcements, mechanics, and support guidance on the web while keeping staff editing authenticated.

02

Move from host deployment to reconciliation

Constraint. A pushed image still required a manual Docker Compose update and produced a short outage.

Decision. Use Kubernetes for replica management and Argo CD reconciliation so declared state, rather than an operator command, controls the running version.

03

Self-manage the cluster

Constraint. Managed Kubernetes services did not fit the operating budget, but multiple applications still needed scale, isolation, and shared delivery practices.

Decision. Build the platform on owned Ubuntu servers and accept the added networking, storage, policy, and maintenance work as a deliberate tradeoff.

04

Treat state as a first-class workload concern

Constraint. Game worlds, account data, and PostgreSQL cannot disappear when a pod is rescheduled.

Decision. Bind stateful services to Longhorn persistent volumes, expose only the required volumes through SFTP, and back claims up to object storage.

Development timeline.

Important revisions, technical pivots, and lessons from each stage.

01 / Product

Replaced a chat-only information model

The first milestone established a central Next.js site with announcements, wiki material, support guidance, rich-text editing, and role-aware content management.

02 / Container

Isolated the application from the host

A Docker image and GitHub Actions build removed the need to run the web application directly on the Ubuntu host, but updates still required a manual Compose refresh.

03 / Cluster

Expanded a short estimate into a platform build

A two-to-three-week estimate expanded to roughly two and a half months as the work grew to include cluster architecture, ingress, load balancing, storage, access, and policy.

04 / GitOps

Made image state declarative

Argo CD and image digest checks connected registry changes to controlled reconciliation, removing the manual update step from the normal release path.

05 / Services

Added durable game and operator services

Velocity, game workloads, PostgreSQL, SFTP, tunnel clients, persistent claims, backups, and least-privilege access became one managed ecosystem.

Results and system metrics.

Measurements, configuration boundaries, and outcomes that show the scope of the work.

Cluster shape

3 + 1 nodes

The generalized Calico configuration represents three control-plane nodes and one worker node.

Web workload

2 replicas

The current Helm deployment configures two Next.js replicas and an HPA range of one to four.

Edge layers

2–8 replicas

The NGINX/ModSecurity and Cloudflare tunnel layers each define a two-to-eight autoscaling range.

Platform build

≈2.5 months

Recorded duration for the Kubernetes platform and supporting services.

Image targets

amd64 + arm64

The production image workflow publishes both architectures from one Buildx job.

Implementation highlights.

Focused excerpts paired with the engineering behavior each one implements.

01Persist structured rich textsrc/components/Editor.js

The shared Draft.js editor converts content between the stored raw representation and an editable state, supporting both announcements and wiki content.

Open source file
typescript
01if (rawContent === null) {02  this.state = {03    editorState: EditorState.createEmpty(),04  };05}06else {07  const contentState = convertFromRaw(rawContent);08 09  this.state = {10    editorState: EditorState.createWithContent(11      contentState12    ),13  };14}15 16const serializedContent = JSON.stringify(17  convertToRaw(editorState.getCurrentContent())18);
02Publish one image for two CPU architectures.github/workflows/docker-deploy.yml

The production workflow uses Buildx to push both amd64 and arm64 variants, matching the mixed server architecture managed by the platform.

Open source file
yaml
01- name: Set up Docker Buildx02  uses: docker/[email protected]03 04- name: Push Next Docker Image05  uses: docker/[email protected]06  with:07    context: ./08    file: ./prod.Dockerfile09    push: true10    platforms: linux/amd64,linux/arm6411    tags: |12      registry/redoxforge-web:commit13      registry/redoxforge-web:latest
03Scale the Next.js workload inside boundshelm/templates/redoxforge-nextjs.yaml

The chart runs two initial web replicas and defines a CPU-driven horizontal autoscaler with explicit minimum and maximum capacity.

Open source file
yaml
01apiVersion: autoscaling/v202kind: HorizontalPodAutoscaler03metadata:04  name: nextjs-hpa05spec:06  scaleTargetRef:07    apiVersion: apps/v108    kind: Deployment09    name: nextjs10  minReplicas: 111  maxReplicas: 412  metrics:13    - type: Resource14      resource:15        name: cpu16        target:17          type: Utilization18          averageUtilization: 50
04Remove unnecessary container privilegeshelm/templates/redoxforge-nextjs.yaml

The deployment disables service-account token mounting, runs as a non-root user, prevents privilege escalation, and drops every Linux capability.

Open source file
yaml
01spec:02  automountServiceAccountToken: false03  containers:04    - name: nextjs05      securityContext:06        runAsUser: 100007        runAsGroup: 100008        runAsNonRoot: true09        allowPrivilegeEscalation: false10        privileged: false11        capabilities:12          drop:13            - ALL
05Start namespace traffic from default denyredoxforge-mc/templates/default/default-network-policy.yaml

The game workload namespace begins with a deny-all policy. Separate policies then grant only the required Velocity, game-server, database, SFTP, and tunnel paths.

yaml
01apiVersion: networking.k8s.io/v102kind: NetworkPolicy03metadata:04  name: default-deny-all05spec:06  podSelector: {}07  policyTypes:08    - Ingress09    - Egress

Sources and boundaries.

Public links open in a new tab. Private code and project artifacts are summarized without exposing infrastructure details or credentials.

repository

RedoxForge website, workflow, and Helm repositoryrevision 631eb2e

artifact

Game and critical-services Helm charts

artifact

Generalized Calico policy and host-role configuration

owner statement

Platform evolution and project timeline