Kuberns logo

Aiko Tanaka

About

Aiko Tanaka is an AI/ML Infrastructure Engineer based in Tokyo with 6 years of experience building the systems that take machine learning models from research notebooks into production. She has worked on LLM serving infrastructure for AI-native product teams in Japan and the US, with hands-on experience deploying and optimising models on GPU clusters using A100 and H100 instances. Aiko's work sits at the intersection that most tutorials ignore: not how to train a model, but how to serve it reliably, cheaply, and at scale. She has built inference APIs with FastAPI, implemented model quantisation to reduce GPU memory requirements by up to 60%, integrated vector databases for RAG architectures, and debugged cold start latency issues on serverless AI deployments that standard cloud platforms were not built to handle. At Kuberns, Aiko writes about AI application infrastructure, the gap between a working model and a deployed product, what most cloud platforms get wrong about ML workloads, and how agentic deployment changes the calculus for teams building AI-native applications.

Articles