Open to AI infrastructure, LLMOps & backend roles
KHAS-ERDENE TSOGTSAIKHAN

AI infrastructure.Built. Scaled.

I build production AI infrastructure and products — from multi-cloud GPU orchestration and open-source backend systems to evaluation gates, retrieval agents, and multimodal apps.

1 SkyPilot PR merged · 5 activeBerkeley RISELab · NutrioMN#2 App Store Mongolia
Signal
RoleAI infrastructure engineer · founder
BasedBerkeley, California
StudyingUC Berkeley EECS · Expected May 2028
NowNutrioMN · Berkeley RISELab
Distributed AI infrastructureMulti-cloud orchestrationAgent evaluationLLMOpsKubernetesBackend systemsMultimodal AIDistributed AI infrastructureMulti-cloud orchestrationAgent evaluationLLMOpsKubernetesBackend systemsMultimodal AI
00 / Proof

Outcomes, not adjectives.

SkyPilot PRMerged upstream · 5 active
NutrioMN usersProduction adoption
App Store rankingMongolia App Store
PrimitiveBench starsOpen-source pull
01 / Selected work

AI infrastructure and products.

Production systems, accepted open-source work, measurable products, and public proof where available.

02 / Philosophy

I care about the version of AI that survives outside the demo — where real users, real money, and real failure modes are on the line.

I've merged a production infrastructure change into SkyPilot through Berkeley RISELab, with five more contributions active, and built NutrioMN, an angel-funded multimodal nutrition app with 50K+ users, $20K+ MRR, and a #2 App Store ranking in Mongolia.

The throughline is reliable infrastructure around probabilistic systems: scheduling, failure handling, structured outputs, golden datasets, traces, deterministic fallbacks, and CI gates. AI becomes a product only when the system around it is measurable and debuggable.

I move at founder speed but build with engineering restraint — because taste, distribution, and reliability are all part of systems design.

03 / Capabilities

Technical depth, made legible.

Core patterns behind measurable, debuggable agentic systems.

01

Distributed AI infrastructure

Multi-cloud GPU scheduling, backend services, fault tolerance, and performance.

02

Evaluation infrastructure

Golden datasets, agent traces, structured validators, and deployment gates.

03

Production AI products

From multimodal inference to subscriptions, observability, releases, and growth.

04

Agentic systems

Tool calling, retrieval, ranking, constrained outputs, and external integrations.

04 / Experience

Founder speed, engineering discipline.

Production ownership inside small teams and real constraints.

01 DEC 2025—NOW

Founder & CTO

NutrioMN · Multimodal AI Nutrition
  • Built Dockerized FastAPI services on AWS with PostgreSQL/Supabase authentication, subscriptions, and scan history for an angel-funded nutrition platform, scaling to 50K+ users and $20K+ MRR with a #2 App Store ranking in Mongolia.

  • Optimized AI/ML inference through prompt compression, caching, and constrained outputs; led 6 interns through Agile sprints, code reviews, monitoring, and production releases while driving 3M+ views and securing 3 gym sponsorships.

FastAPIAWSPostgreSQLSupabaseDocker
02 FEB 2026—NOW

Software Engineer

Berkeley RISELab
Berkeley RISELab · SkyPilot

Merged one production PR into SkyPilot main, with five additional contributions in development or review, while developing systems research on cross-cloud inference-state recovery.

PythonLinux CLIDistributed systems
03 APR 2026—NOW

Co-founder & Lead Engineer

PrimitiveBench · AI Evaluation Infrastructure

Built an open-source, vendor-neutral evaluation platform with async execution, multi-tenant job APIs, tracing, and CI/CD gates for production regressions.

PythonFastAPIPostgreSQL
04 MAY—JUL 2026

Software Engineering Intern

CourseLynx

Owned reliable course ingestion for 40,000+ users across 16+ university catalogs, reaching 99.4% reliability and reducing p95 latency from 820 ms to 510 ms.

PythonPostgreSQLRedis
05 DEC 2025—MAR 2026

Software Engineering Intern

Curio AI · Learning Platform

Developed APIs, structured recommendation logic, and a LangGraph agent for dynamic content retrieval.

LangGraphREST APIsJSON Schema
05 / Stack

A practical technical range.

Tools for connecting model quality, system reliability, and product delivery.

AI Infrastructure

  • Distributed orchestration
  • Multi-cloud systems
  • GPU workloads
  • Kubernetes
  • Fault tolerance
  • Observability

AI / LLMOps

  • Agent orchestration
  • RAG evaluation
  • Structured outputs
  • VLMs
  • Vector search
  • Evaluation gates

Languages

  • Python
  • Go
  • Java
  • C++
  • TypeScript
  • SQL

Backend / Data

  • FastAPI
  • Node.js
  • PostgreSQL
  • Redis
  • REST APIs
  • ETL pipelines

Product / Infra

  • React / Next.js
  • Docker
  • AWS
  • Linux
  • GitHub Actions
  • CI/CD
Open to ambitious AI work

Building something that needs to work outside the demo?

Open to AI infrastructure, distributed systems, LLMOps, backend, and high-ownership startup engineering opportunities. Email me directly, or verify the work through LinkedIn, GitHub, and PrimitiveBench.

KHAS-ERDENE TSOGTSAIKHANBerkeley · California© 2026