About
My work lies at the intersection of large language models, intelligent agents, and production operations. I am interested in building reliable AI systems that combine retrieval, memory, reasoning, and domain knowledge to support complex analytical tasks.
Building on practical experience in reliability engineering for large-scale online financial systems and my current work and life in the San Francisco Bay Area, I am especially interested in the application of LLMs to operational intelligence, incident investigation, troubleshooting workflows, and evidence-based root cause analysis for distributed systems.
Research Interests
- Large Language Models
- Retrieval-Augmented Generation
- Agent Systems
- Long-Term Memory Retrieval
- Conversational Question Answering
- AIOps and LLM4Ops
- Site Reliability Engineering
- Root Cause Analysis
- Machine Learning
Experience
I have 5 years of industry experience in Site Reliability Engineering, with a primary focus on reliability engineering, high availability, and infrastructure resilience for large-scale online financial systems. My professional practice has involved production reliability, distributed systems, operational troubleshooting, and service support in complex, high-demand environments.
My current research explores LLM-centered systems for memory, retrieval, reasoning, and operations, especially in AIOps, LLM4Ops, troubleshooting, and diagnostic workflows for complex environments.
Selected Publications
A complete list of publications can also be found on my Google Scholar profile.