Autonomous Driving · Applied AI · Cloud · Distributed Systems
I build the large-scale systems that validate the software running in production vehicles, and the AI agents — models, tools, and plugins — that operate them. Over 10+ years, from securing the internet at scale to self-driving, the throughline is the same: turning hard data-and-compute problems into reliable production infrastructure.
I'm a Staff Engineer at Qualcomm, on the critical path of the company's autonomous-driving program — the large-scale infrastructure that replays recorded real-world drives to validate the software running in production vehicles. I designed and operate that platform, and I was a key collaborator on the reprocessing pipeline behind the jointly developed Qualcomm–BMW Snapdragon Ride Pilot automated-driving system, launched in 2025.
Increasingly my work is applied AI: building and running production agents, and the models, tools, and plugins around them, that operate mission-critical infrastructure — from an agent that autonomously triages tens of thousands of pipeline failures a day to custom LLM tooling that helps engineering teams move faster.
Before Qualcomm, I spent ~6 years at Webroot / OpenText building internet-scale cybersecurity systems — distributed crawlers that safely processed 17M+ URLs a day, real-time threat classification, IP threat intelligence, and multi-cloud honeypots — work that produced two granted U.S. patents. Whether it's defending the internet or validating self-driving AI, the throughline holds: secure, reliable, massively parallel data pipelines that turn raw data into trustworthy results.
I care about systems that are correct under pressure, observable, and increasingly self-operating — where AI does the repetitive diagnosis so engineers can focus on the hard problems.
The systems I've designed and operate at Qualcomm, plus the patented work that preceded it.
Designed, built, and operate two large on-premises Kubernetes clusters of Qualcomm's QCR-100 accelerators — the same Snapdragon Ride automotive silicon that ships in vehicles — for bit-accurate ADAS/AD reprocessing. Purpose-built silicon runs this ~100–300× faster than general-purpose compute, fast enough to replay recorded drives at close to real time. Built as Reprocessing-as-a-Service: one interface running identically in cloud or on-prem, with event-driven Argo Workflows, DAG dependency resolution, and parallel petabyte-scale data transfer. This is the data that both validates today's software and trains the next generation of driving models.
As project lead, built and run an AI system that autonomously triages tens of thousands of pipeline failures/day. Self-hosted embeddings + pgvector cosine-similarity suppress duplicates, so only novel failures reach a Claude Agent SDK + AWS Bedrock agent that diagnoses root cause and opens Jira tickets or GitHub fix-PRs — shifting engineers from investigating every failure to reviewing a diagnosed problem with a fix already proposed.
Beyond triage, I build the broader applied-AI layer: custom Claude skills & plugins and MCP tool integrations that wire LLMs into real engineering workflows, agentic-coding tooling (Claude Code, Codex), and self-hosted embedding models (Qwen3, ONNX int8) — choosing the right model and the right amount of autonomy for each task so teams ship faster with humans in the loop.
Re-architected the cloud reprocessing pipeline onto GCP Dataflow to run at massive parallelism — redesigning how work is partitioned and scheduled to raise concurrency and cut end-to-end turnaround while improving output accuracy. Cross-account, event-driven AWS underneath: ECS Fargate, SQS FIFO, EventBridge Pipes, RDS/PostgreSQL, provisioned as code with AWS CDK.
Prometheus-based fleet monitoring tracking each accelerator's health, firmware, and Real-Time Factor (how closely reprocessing matches ground truth), plus low-level IPMI/BMC power control and safe Kubernetes node lifecycle — servicing hardware without disrupting running jobs.
Led distributed clusters crawling 17M+ URLs/day and co-invented "DeepCrawl" — pivoting from a known malicious link to uncover related malicious infrastructure and grade its threat with an ML model. The basis of two granted U.S. patents.
Built an IP threat-classification system categorizing malicious hosts by activity (phishing, botnet, malware, proxy, spam), a cloud-reputation scoring system, and multi-cloud honeypots (AWS, GCP, Azure) that captured live attacks to continuously feed the threat-analysis pipelines.
Kubernetes-native, event-driven architectures that stay reliable from one node to petabyte-scale fleets.
Agents, models, tools, and plugins that diagnose, remediate, and accelerate real engineering work — the right model and the right autonomy for each task, humans in the loop.
Bit-accuracy, observability, and safe operations for systems where the output has to be trusted.
Named co-inventor on two granted U.S. patents, plus first-author peer-reviewed research. All links open the primary source.
I spend my days validating lane-keeping software. Here's the keeps-you-honest version — steer to stay on the road, dodge the markers, and make it to the space station at the end. A rocket's waiting. No cloud required.
Runs entirely in your browser — no data leaves the page. Best score is remembered on this device only.
Open to conversations on autonomous-driving infrastructure, large-scale AI systems, and hard engineering problems. Reach out directly.