Publications
2026
-
- AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility2026Preprint
-
- FaultLoc: Evaluating Coding Agents for Fault LocalizationIn Second Workshop on Agents in the Wild: Safety, Security, and Beyond (AIWILD), ICML, 2026
-
A Comprehensive Survey of Evaluating Multimodal Foundation Models: Hierarchical Perspective and Extensive Applications2026Under review at ARR 2026 - CyberCycle: A Scalable Real-World Benchmark for AI Agents’ End-to-End Cybersecurity CapabilitiesIn Proceedings of the International Conference on Machine Learning, 2026
- FICO: Evaluating Vision-Language Models under Visual Fidelity and Compression at ScaleIn Proceedings of the Annual Meeting of the Association for Computational Linguistics, 2026
2025
-
MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language ModelsIn Proceedings of the 3rd Workshop on Towards Knowledgeable Foundation Models (KnowFM), Aug 2025 -
Predicting Task Performance with Context-aware Scaling LawsIn Proceedings of the 3rd Workshop on Towards Knowledgeable Foundation Models (KnowFM), Aug 2025
2024
-
⭐ Failure in a population: Tauopathy disrupts homeostatic set-points in emergent dynamics despite stability in the constituent neuronsNeuron, Aug 2024Cover Paper