WeiRan Yan

WeiRan Yan

Independent Researcher / Senior Engineer

I am a senior site reliability engineer based in the San Francisco Bay Area, with five years of professional experience in Site Reliability Engineering (SRE), with a primary focus on reliability engineering, high availability, and infrastructure resilience in large-scale online financial systems. I have led the design and operation of a multi-data-center disaster recovery architecture for online payment systems serving hundreds of millions of users, covering both intra-city and cross-region resilience. This architecture proved effective during real-world failures and helped prevent business interruption. My recent research focuses on intelligent systems built with large language models, including long-term conversational question answering, tool-augmented memory retrieval, agentic fallback, parameter-efficient small models, AIOps, and LLM4Ops. I am particularly interested in applying LLMs and agent-based methods to reliability engineering, operational troubleshooting, and root cause analysis in complex production systems. My earlier work includes audio-based signal processing and machine learning.

About

My work lies at the intersection of large language models, intelligent agents, and production operations. I am interested in building reliable AI systems that combine retrieval, memory, reasoning, and domain knowledge to support complex analytical tasks.

Building on practical experience in reliability engineering for large-scale online financial systems and my current work and life in the San Francisco Bay Area, I am especially interested in the application of LLMs to operational intelligence, incident investigation, troubleshooting workflows, and evidence-based root cause analysis for distributed systems.

Research Interests

Experience

Professional Background

I have 5 years of industry experience in Site Reliability Engineering, with a primary focus on reliability engineering, high availability, and infrastructure resilience for large-scale online financial systems. My professional practice has involved production reliability, distributed systems, operational troubleshooting, and service support in complex, high-demand environments.

Current Directions

My current research explores LLM-centered systems for memory, retrieval, reasoning, and operations, especially in AIOps, LLM4Ops, troubleshooting, and diagnostic workflows for complex environments.

Selected Publications

TA-Mem: Tool-Augmented Autonomous Memory Retrieval for LLM in Long-Term Conversational QA
M Yuan, J Liu, J Yang, X Li, W Yan, Y Wu, P Liang
arXiv preprint arXiv:2603.09297, 2026
Tiny-Critic RAG: Empowering Agentic Fallback with Parameter-Efficient Small Language Models
Y Wu, P Liang, Y Xiang, M Yuan, J Liu, J Yang, X Li, W Yan
arXiv preprint arXiv:2603.00846, 2026
Automatically Predicting Giant Panda Mating Success Based on Acoustic Features
W Yan, M Tang, Z Chen, P Chen, Q Zhao, P Que, K Wu, R Hou, Z Zhang
Global Ecology and Conservation 24, e01301, 2020
Audio-based Automatic Mating Success Prediction of Giant Pandas
W R Yan, M L Tang, Q Zhao, P Chen, D Qi, R Hou, Z Zhang
arXiv preprint arXiv:1912.11333, 2019
Multi-Scale Residual CRNN with Data Augmentation for DCASE 2020 Task 4
M Tang, L Guo, Y Zhang, W Yan, Q Zhao
Conference paper, 2020

A complete list of publications can also be found on my Google Scholar profile.

Contact

yanwr2016@gmail.com

Google Scholar