Mochammad AlamsyahAvailable for new engagements · ~20 hrs/week
I fix LLM infrastructure costs and productionize agent systems —fractionally.
Software engineer in Tokyo, nine years building consumer-scale platforms: a super-app at scale, a data-warehouse bill cut 90%, production AI on the Claude Agent SDK and AWS. Platform judgment for LLM-heavy teams, without the headcount.
Blueprint of an LLM request path: a client hits a router, which either returns from cache or sends the miss to a model, then through an eval harness, before a response. Cost is measured on the model path.
[ 01 — What I take off your plate ]
- 01
LLM spend, measured down
Caching, routing, model right-sizing, batch vs realtime. Your bill should track usage, not outgrow it — the same discipline that cut a company-wide data-warehouse spend by 90%.
- 02
From demo to boring reliability
Retries, guardrails, state, observability — agent pipelines and retrieval built with the discipline of any other distributed system. The gap between a great demo and a system nobody worries about is exactly the work I do.
- 03
An eval harness, not a gamble
Every prompt or model change is a coin flip until there's a harness. I build the measurement that turns changes into measured rollouts.
- 04
Distributed systems & SRE
Go, Kubernetes, AWS and GCP. Monolith decomposition, incident response, alert hygiene — the unglamorous plumbing that keeps products alive.
[ 02 — The work ]
Nine-plus years across fintech, ride-hailing, healthcare, and AI — leading distributed teams across Japan, India, Indonesia, and US time zones. Names withheld; specifics available in conversation.
- 012017 — 2018
withheldcompany name withheldE-commerce startup — Jakarta
Co-founder & CTO
- Built the MVP of a bidding-based e-commerce platform across web and iOS
- WebSocket bidding engine and multi-gateway payment integrations
- Go, React, Node.js, Swift, Redis, Amazon SQS
FounderGoWebSocket
- 022018 — 2023
withheldcompany name withheldRide-hailing super-app — Bangalore & Jakarta
Lead Software Engineer
- Decomposed a core monolith into microservices; helped migrate primary services from VMs to Kubernetes
- Built a Go distributed-locking library on Redis Cluster; cut alert volume from 30+ a week to under 5
- Designed the ride-hailing homepage on a BFF architecture; shipped a phone-number masking feature later adopted org-wide
GoKubernetesDistributed systems
- 032023 — 2025
withheldcompany name withheldFinancial-services conglomerate — Jakarta
Lead → Principal, AI & Data Platform
- Built an AI assistant for C-level executives: LLM function-calling over the company data warehouse for natural-language business metrics
- Migrated the company-wide data warehouse to cloud-native serverless analytics — 90% cost reduction
- Architected and led a loan-application system for the used-vehicle market on GCP, development through production
AI platformBigQueryGCP
- 042023 — present · concurrent
withheldcompany name withheldGlobal healthcare platform — remote
Senior Software Engineer (Consultant)
- Production AI tooling on the Claude Agent SDK, AWS Lambda, and Go
- Key contributor scaling the platform from a single market to dozens of countries — multi-currency, multi-language, multi-payment-gateway
- Prototyped a bulk-upload validation engine that turned a ticket-analysis task into a working MVP
GoClaude Agent SDKAWS
[ 03 — Proof ]
Group-wide data warehouse, re-architected
Apache Hive on ~100 VMs to an S3 data lake + Redshift Serverless — ≈$20–30k/month down to ≈$2k, half the scheduled reports retired first. Three months, zero cutover incidents.
- Monthly bill
- $20–30k → $2k
- Timeline
- 3 months
- Scheduled reports
- 30 → 15
- VMs retired
- ~100
Super-app home screen on one BFF call
Six-plus API calls collapsed into one Go service — 100M+ calls/day at <50ms p90, with parallel fan-out, per-service timeouts, and a zero-incident migration.
- Calls per screen open
- 6 → 1
- Daily volume
- 100M+
- Latency
- <50ms p90
- Migration
- ~1 mo to 90%
Built in the open
Open source and certifications
- mindgraph-mcp
Graph-backed memory MCP server for Claude — cross-session memory with semantic search, full-text search, and graph traversal. Go.
- leakfix
Remediation agent for secret-leak findings — per-provider revocation runbooks and review-ready PRs, with side effects spelled out, not hidden. Go.
- nihongo.malamsyah.com
Japanese learning platform with spaced repetition and AI conversation practice. Next.js.
Google Cloud Professional ML EngineerCertified Kubernetes AdministratorAWS Solutions ArchitectGoogle Cloud Associate Engineer
[ 04 — How I work ]
Ways of working
Advisory
A standing line to a senior platform engineer. Architecture reviews, cost audits, hiring help, written decisions you can forward to the team.
Embedded fractional
Hands in the codebase a few days a week. I ship the platform work alongside your team and leave the runbooks behind.
Fixed scope
A defined deliverable — an eval harness, a cost-reduction pass, a productionized agent pipeline — with a written scope before any contract.
Paid pilot · Every engagement starts with a short, paid pilot: tight scope, written findings, a decision at the end. No pressure to continue — the pilot has to earn it.
- 01
Async-first, from Tokyo
A track record of leading distributed teams across Japan, India, Indonesia, and US time zones. Async hours plus one weekly call — my mornings overlap US afternoons, so you get written decisions, not meetings.
- 02
Evals before scale
Nothing ships to more traffic until there's a measurement that says it should.
- 03
Boring is the goal
Production AI systems should be the least exciting part of your stack.
- 04
Leave the runbooks
Every engagement ends with your team able to run what I built without me.
Currently taking conversations
Tell me what's breaking — or what's about to. I reply within a day with three questions and an honest read on fit; if it makes sense, we scope a short paid pilot on a call.
hi@malamsyah.comReferences available on request · github.com/malamsyah