Levon Barseghyan

Machine Learning Engineer

Levon Barseghyan

I'm a machine learning engineer who likes hard design problems: LLM agents, reinforcement learning, and the training and infrastructure behind them. Day to day, I work on a multi-agent AI platform at Amazon.

About

I like big, difficult design problems: reinforcement learning, training, and the infrastructure behind it.

I'm a machine learning engineer at Amazon in Los Angeles. I got here through mobile. For almost four years I built iOS and Android mapping features for delivery drivers, and in January 2026 I moved to the ML side, where I work on a platform that runs groups of LLM agents together. I enjoy all of machine learning engineering, but what I like most is reinforcement learning, training, infrastructure, and anything that comes down to a big design problem with no obvious answer.

I started programming in high school robotics, writing the autonomous code for our team, and later mentored a middle-school team of my own. At Berkeley I was in the Machine Learning club. Since then I've worked on a government procurement platform, been CTO of a startup, co-founded another, and led a team on a freight-tracking app.

I'm finishing my M.S. in Computer Science at Georgia Tech in December 2026, mostly in machine learning, deep learning, reinforcement learning and agents, and I co-authored a paper there on multi-agent text-to-SQL.

Outside work I play golf and pickup sports, travel, and try to eat at the best restaurant in whatever city I'm in. I speak Armenian, Russian, Ukrainian and English.

Portrait of Levon Barseghyan
LEVON BARSEGHYAN · LOS ANGELES · LB

Experience

Full resume
  1. Mar 2022 – Present

    Amazon

    Software Development Engineer, Machine LearningJan 2026 – Present

    Built and lead the evaluation and observability setup for Amazon's internal platform where groups of LLM agents work together on geospatial inference. It traces every LLM call and agent hand-off and scores agent setups against ground truth, including with LLM-as-judge. It's the org's first agent-evaluation tool, used by four teams across five projects. I also moved the trace store to DynamoDB and S3 (about 60% cheaper), built the dataset behind the platform's distillation pipeline, and automated on-call work with agents that resolve tickets, lifting SLA attainment by 30%.

    LLM agents · Evaluation · OpenTelemetry · Bedrock · Python

    Software Development EngineerMar 2022 – Jan 2026

    Built offline map caching for iOS and Android (40% less on-device storage) and a real-time location-validation algorithm in Kotlin and Swift. My mapping work was part of a team initiative worth about $93MM in annual savings, and I worked across orgs on parking guidance that contributes $100MM+. I also fixed 95% of map-rendering issues in three months and cut incident response time by 25%.

    Kotlin · Swift · iOS · Android · Geospatial

  2. Mar 2024 – Aug 2024 · nights and weekends

    Co-founder & CTO · LeagueAtlas

    Co-founded a startup with two friends to cut no-shows on sports facility rentals, and was the main engineer. I built two web apps, one for players and one for facilities, on the T3 stack with Postgres, going from nothing to a working MVP in three months. No-shows dropped 25% among our early beta testers. I left after five months to start my M.S.

    Next.js · tRPC · Prisma · Postgres · Vercel

  3. Oct 2022 – Jan 2024 · nights and weekends

    Senior Software Engineer (Contract) · FreightRight

    Led four engineers building a web app for tracking international freight shipments, designed for 5,000+ concurrent users. I built the user management on Spring Boot, the React frontend, and a fully automated CI/CD pipeline on AWS.

    Spring Boot · React · AWS · CI/CD

  4. Jun 2021 – Jun 2022 · nights and weekends

    CTO · Pogbet

    Pogbet was a startup building a prediction platform for streamers on Twitch and YouTube. Before a game, a streamer opens a live prediction ("will I get more than 10 kills this round?") and their audience bets on it. I hired and managed five engineers across the US, Ukraine and Armenia, wrote a lot of the code (React, Spring Boot, and a Python server running our ML model), and built a model that generated League of Legends predictions automatically. We unfortunately weren't able to raise our seed round.

    React · Spring Boot · Python · ML inference

  5. Apr 2020 – Mar 2022

    Software Developer · PlanetBids

    PlanetBids runs bidding and vendor registration for public agencies. When Flash was retired, we rebuilt the application on Ember and Django in under ten months. The old app had years of one-off agency requests built in, so much of the work was finding those edge cases and keeping behavior identical. I worked on both the vendor and agency sides, built an internal admin app and a user-guide website, and set up QA that reached 100% of the client's acceptance criteria.

    Ember · Django · SQL · Accessibility

  6. Sep 2018 – 2020

    Software Mentor · Piedmont Pioneers Robotics

    Mentored ten middle-school students building a robot for the FIRST Tech Challenge, teaching them Java and precise autonomous programming. Their robot beat teams with more experience.

    Java · FIRST Tech Challenge

Projects & Research

projectIn progress · Solo

Deep RL for Bazar Blot

Bazar Blot is the Armenian version of Belote, and the card game I grew up playing with my family. I'm training agents to play it with self-play reinforcement learning.

It's a hard problem to learn. Four players in two teams, 24 of the 32 cards hidden, a bidding round that sets the contract, and a partner you can only talk to through your bids and the cards you play. How much you should bid depends on how well you'll play, and how you should play depends on what was bid. The two have to be learned together.

What's built

  • A rules engine for the Blot Star ruleset, with 330 tests, checked against a million random deals.
  • A solver that finds the best possible play when all four hands are visible. I use it as ground truth.
  • A training environment, plus baseline bots: random, a rule-based player, and one that samples possible hidden hands and solves each.
  • An evaluation setup that replays the same deals with the seats swapped, so luck cancels out.
  • A PPO self-play agent, and a web app where you can watch the bots or play against them.
+67
pts/deal vs heuristic
+206
pts/deal vs random
330
tests, 1M-deal check

After about 28 minutes of training on one CPU core, the agent beats the rule-based bot by about 67 points a deal and random play by about 206. The confidence intervals stayed above zero at every check during training.

My first self-play run broke: all four seats learned to always pass, so no hand was ever played. I fixed it with two changes. A round where everyone passes now counts as a real zero-point result instead of being thrown away, and the opponent pool includes fixed baseline bots. I haven't yet checked whether its bidding is actually good, or only good enough to beat weak opponents.

Next: try different ways of training the bidder and the player together, then add search on top of the learned policy.

PyTorch · PPO · Self-play · PettingZoo · Python

researchGeorgia Tech · 2025

MaGS: Multi-Agent Geospatial SQL

Language models are bad at spatial SQL. They don't know the PostGIS functions or how geometry types behave, so queries come back wrong or fail to run, especially on joins across several tables.

We built MaGS, a team of LLM agents run by a supervisor. One agent picks the relevant tables and columns, one breaks the question into smaller steps and writes PostGIS SQL, one checks and fixes the query and runs it, and one draws the result on a map. The agents look up PostGIS documentation and example queries before writing, instead of guessing function names. There's no fine-tuning: it's all prompts and retrieval.

I co-authored the paper, designed the supervisor architecture, and built the map-making agent.

83%
accuracy, vs 67% baseline
60%
complex joins, vs 40%

We tested against a LangChain SQL agent on Austin Airbnb and census data, using both GPT and DeepSeek models. The biggest gap was on complex multi-table spatial joins.

Stacy Liu, Levon Barseghyan, Vijay Madisetti · Manuscript, 2025

Multi-agent LLMs · RAG · PostGIS · LangChain · Text-to-SQL

Skills

LLMs & agents
LLM agents · Multi-agent systems · RAG · Prompt engineering · Agent evaluation · Knowledge distillation · LangChain · LangGraph · AWS Bedrock
ML & training
PyTorch · TensorFlow · Reinforcement learning (PPO, self-play) · Deep learning · LLM-as-judge · A/B testing & experimentation · NumPy · Pandas · Matplotlib
Data & infrastructure
Postgres / PostGIS · DynamoDB · S3 · Vector stores · AWS (EKS, Elastic Beanstalk, CloudFront) · Kubernetes · Docker · Jenkins · CI/CD · OpenTelemetry
Mobile & web
Kotlin · Swift · iOS · Android · React · Next.js · Spring Boot · Django · Ember
Languages
Python · Java · JavaScript · TypeScript · C · Kotlin · Swift · SQL · HTML/CSS
Spoken
Armenian (fluent) · Russian (fluent) · Ukrainian (fluent) · English (fluent) · Spanish (basic) · German (basic)

Education