The IAP Cornell Workshop on the Future of AI in the Cloud was Conducted on Friday, May 15, 2026 on the Cornell Campus.
Time: 8:30am - 4:00pm EDT
Venue: CIS 142, Computing and Information Science Building, 127 Hoy Road, Cornell University, Ithaca, NY
This workshop was hosted and co-organized by Prof. José Martínez (ECE), Prof. Robbert van Renesse (CS), Prof. Hakim Weatherspoon (CS) and the IAP.
Time: 8:30am - 4:00pm EDT
Venue: CIS 142, Computing and Information Science Building, 127 Hoy Road, Cornell University, Ithaca, NY
This workshop was hosted and co-organized by Prof. José Martínez (ECE), Prof. Robbert van Renesse (CS), Prof. Hakim Weatherspoon (CS) and the IAP.
Participants included faculty, postdocs, students, industry scientists and engineers. Participating companies include AMD, ByteDance, Futurewei, IBM, Marvell, and Micron. Cornell Faculty presenting their research included Prof. Robbert van Renesse, Prof. Rachee Singh, Prof. Jiaxin Lin, Prof. Giulia Guidi, Prof. Hakim Weatherspoon and Prof. Kilian Weinberger.
The student poster session was conducted during the break at lunch. The Best Poster Award winner was PhD student Salman Abid.
Agenda - Videos of the Presentations
8:30-8:55 – Badge Pick-up – Coffee/Tea and Breakfast Food/Snacks
8:55-9:00 – Welcome – Prof. Hakim Weatherspoon
9:00-9:30 – Prof. Robbert van Renesse, Cornell, “Latency, Consistency, and the Myth of the Linearizable Log”
9:30-10:00 – Dr. Gloire Rubambiza, IBM, “Advancing AI Platforms for Agents (kagenti) and Distributed LLM Inferencing (llm-d)”
10:00-10:30 – Prof. Rachee Singh, Cornell, “When Bandwidth Isn't the Bottleneck: Breaking Synchronization Barriers in Distributed ML”
10:30-11:00 – Prof. Hakim Weatherspoon, Cornell, "AI and Optimal Oblivious Reconfigurable Networks"
11-11:30 – Lightning Round for Student Posters
11:30-12:30 – Lunch and Poster Viewing
12:30-1:00 – Prof. Kilian Weinberger, Cornell, “Think Before you Speak: Next Gen LLMs with Global Reasoning and External Memory”
1:00-1:30 – Prof. Jiaxin Lin, Cornell, “Intelligent Data Movement for Specialized Scale-up and Scale-out Networks”
1:30-2:00 – Craig Barner, Marvell, “Securing the AI Cloud: Silicon-Rooted Trust in the Age of Intelligent Infrastructure”
2:00-2:30 – Prof. Giulia Guidi, Cornell, “Overcoming Parallelism Challenges in Irregular Data Analysis using Sparse Linear Algebra”
2:30-3:00 – Best Poster Award & Reception
Please see Abstracts and Bios below the Testimonials from Previous Workshops.
Testimonials from Previous Workshops
Professor David Patterson, the Pardee Professor of Computer Science, UC Berkeley, “I saw strong participation at the Cloud Workshop, with some high energy and enthusiasm; and I was delighted to see industry engineers bring and describe actual hardware, representing some of the newest innovations in the data center.”
Professor Christos Kozyrakis, Professor of Electrical Engineering & Computer Science, Stanford University, “As a starting point, I think of these IAP workshops an intersection of industry’s newest solutions in hardware with academic research in computer architecture; but more so, these workshops additionally cover new subsystems and applications, and in a smaller venue where it is easy to discuss ideas and cross-cutting approaches with colleagues.”
Professor Hakim Weatherspoon, Professor of Computer Science, Cornell University, “I have participated in three IAP Workshops since the first one at Cornell in 2013 and it is great to see that the IAP premise is a success now as it was then, bringing together industry and academia in a focused all-day exchange of ideas. It was a fantastic experience and I look forward to the next one!”
Dr. Carole-Jean Wu, Research Scientist, AI Infrastructure, Facebook Research, and Professor of CSE, Arizona State University, “IAP Workshops provide valuable interactions among faculty, students and industry. The smaller venue and the poster session foster an interactive environment for in-depth discussions and spark new collaborative opportunities. Thank you for organizing this wonderful event! It was very well run.”
Dr. Pankaj Mehra, VP Product Planning, Samsung (currently CEO Elephance Memory), "Terrifically organized Workshops that give all parties -- students, faculty, industry -- valuable insights to take back."
Abstracts and Bios (alphabetically by last name). Please Check Back for Updates.
Craig Barner, Marvell, “Securing the AI Cloud: Silicon-Rooted Trust in the Age of Intelligent Infrastructure”
Abstract: Artificial intelligence is reshaping cloud infrastructure faster than the underlying security foundations can keep pace. The traditional model of trusting the software and defending the perimeter is breaking down. In its place, a more fundamental shift is emerging, one that redefines where trust begins.
This talk explores how the industry is rethinking trust from the ground up: why the chip itself has become the new security boundary, how that changes everything from data centers to AI agents, and what it means for the future of sovereign cloud and data privacy. We’ll also look at how companies like Marvell are turning these ideas into silicon that runs at the heart of the world’s largest clouds, and what it takes to build security infrastructure at that scale.
If you’ve ever wondered what actually keeps AI systems secure, or whether they are, this talk is for you.
Bio: Craig Barner is a Senior Distinguished Engineer at Marvell Technology, where he leads the architecture and development of security solutions for custom silicon platforms. With over 25 years of experience in hardware security, he focuses on building cryptographic technologies that safeguard data across cloud platforms, AI-driven data centers, and the distributed infrastructure enabling security at scale. His work centers on some of the most important shifts in cloud security today, including the transition to post-quantum cryptography, the move toward hardware-rooted Zero Trust, and the growing need for confidential computing at scale. As new standards such as ML-KEM and ML-DSA take shape, Craig is helping drive the development of crypto-agile silicon designed to stay ahead of evolving threats. He holds over a dozen patents spanning hardware security and a wide spectrum of hardware innovations, and has authored widely adopted technical specifications in areas such as cryptographic acceleration, secure key management, firmware integrity, and system-level security. More broadly, his work focuses on embedding strong security directly into the silicon layer to help enable trusted, large-scale computing.
Prof. Giulia Guidi, Cornell, Overcoming Parallelism Challenges in Irregular Data Analysis using Sparse Linear Algebra
Abstract: Today, scaling AI beyond single-node limits has shifted workloads to deeply hierarchical multi-GPU systems, where communication becomes the bottleneck. Central operations in modern large-scale AI including GNN training, embedding clustering, and sparse attention, reduce to the same core primitives: sparse and dense matrix multiplications. Yet, scaling these primitives in isolation is not enough; the real challenge is efficiently composing them, as communication costs and data layouts rarely align across operations.
In this talk, we show how to (a) map clustering to sparse and dense linear algebra for efficient single-GPU computation and (b) scale these primitives efficiently on large-scale systems. In particular, we present communication-efficient distributed-memory algorithms for Kernel K-Means that combine GEMM and SpMM with partitioning schemes designed to minimize communication. Our 1.5D algorithm enables exact Kernel K-Means on million-scale datasets [PPoPP '25, IPDPS '26]. This is especially important for AI for science, where data irregularity (graphs, sequences, sparse observations) is intrinsic, and where co-designed sparse and dense linear algebra with the architecture provides a natural path forward.
Bio: Giulia Guidi is an Assistant Professor in the Department of Computer Science and a Graduate Field Faculty in the Department of Computational Biology, the Center for Applied Math, and the Department of Electrical and Computer Engineering at Cornell. She received her PhD in Computer Science from UC Berkeley. She works in the field of high-performance computing for large-scale computational sciences. She is interested in developing algorithms and software infrastructures on parallel machines to speed up data processing without sacrificing programming productivity and to make high-performance computing more accessible. Dr. Guidi received the 2024 SIAM Activity Group on Supercomputing Early Career Prize.
Prof. Jiaxin Lin, Cornell, “Intelligent Data Movement for Specialized Scale-up and Scale-out Networks”
Abstract: Data movement has become a critical bottleneck for next-generation data center workloads, driven largely by the massive communication demands of modern AI applications. To overcome these limitations, distributed systems are increasingly relying on specialized network devices and advanced interconnects to efficiently move data between distributed GPUs and CPUs. These domain-specialized devices and interconnects have unique characteristics and create complex system-level problems. How can we harness the massive performance benefits of specialization without being overwhelmed by its complexity?
In this talk, I will demonstrate how we can solve this by re-architecting the data movement stack to fully harness the benefits of specialized networks while making them accessible. I will explore this through two distinct examples in the scale-out and scale-up domains. First, I will introduce Alkali, a compiler framework that tackles the complexities of programming and utilizing domain-specialized smart network interface cards (SmartNICs). Next, I will present DirectKernel, an end-to-end framework designed for scale-up, GPU-to-remote memory data movement that tackles the LLM memory wall.
Bio: Jiaxin Lin is an assistant professor at Cornell ECE. She received her Ph.D. from UT Austin. Her research aims to design innovative software and hardware for data center networks. She is widely recognized for her influential work on programmable networks. She has received a Google Junior Faculty Award in 2025, as well as Google and Meta PhD Fellowships in Computer Networking, and was selected as an MIT EECS Rising Star in 2024.
Prof. Robbert van Renesse, Cornell, “Latency, Consistency, and the Myth of the Linearizable Log”
Abstract: We often treat linearizability as the gold standard for distributed systems. Yet in shared logs, the workhorses behind replication, coordination, and transactions, this ideal becomes a liability. Enforcing linearizable appends forces global synchronization into the critical path, bloating tail latencies into the millisecond range. This talk explores a different design point: logs that guarantee total order but not real-time precedence. I’ll introduce Ziplog, a coordination-free ordered log that achieves 16 µs end-to-end latency, and Zipkvs, a transactional key-value store built atop it that delivers stronger semantics than today’s RDMA systems at higher performance.
Bio: Robbert van Renesse is a Professor of Computer Science at Cornell University, where his research focuses on distributed systems, fault tolerance, and operating systems. His work has shaped modern thinking about scalable and reliable distributed computing, from early group-communication systems to recent advances in low-latency coordination and consistency. He is co-Editor-in-Chief of ACM Transactions on Computer Systems and a former Chair of ACM SIGOPS. An ACM Fellow, Robbert has served as Program Chair for SOSP, OSDI, and EuroSys.
Gloire Rubambiza, IBM, “Advancing AI Platforms for Agents (kagenti) and Distributed LLM Inferencing (llm-d)”
Abstract: Agentic AI, where AI agents leverage large language models (LLMs) and tools for coding and other personal tasks, is emerging as an exciting new category of cloud computing workload. The effectiveness of generative AI agents critically depends on the diversity and reasoning capabilities of the LLMs used by the agents. However, serving these dynamic AI workloads on enterprise clouds faces numerous challenges such limited GPU resources, delegation and security of agents/tools, cold start for model startup and inferencing, etc. In this talk, I will present two threads of work on building cloud-native platforms for the agentic era. First, I will briefly describe our work on agent workload lifecycle management in the kagenti project. Second, the exemplary workloads in kagenti will serve as motivation for a deeper dive into our work on fast model actuation (FMA) for distributed LLM inferencing on Kubernetes in the llm-d project. Specifically, I will describe different FMA techniques that have gradually enabled us to host multiple models on a single GPU and reduce startup time of a Llama-3.1-8B-Instruct model by an order of magnitude from 1 minute 32 seconds to 3 seconds.
Bio: Dr. Gloire Rubambiza is a Research Scientist at IBM, where he advances the state of the art in hybrid cloud infrastructure for AI agents and LLM inferencing. Before rejoining the "big blue" family, he led research and engineering efforts to build distributed cloud infrastructure for AI in space at Satlyt, a startup in Silicon Valley. Prior to his startup journey, he was a PhD student and postdoc in Computer Science at Cornell University, where he conducted research in hybrid cloud computing for digital agriculture with an emphasis on societal impact. His research approach strategically combines techniques from computer networking and critical human-computer interaction (HCI) into a single approach called trilingualism, which continuously integrates (1) networked systems building, (2) critical domain applications, and (3) critical reflections on the technical decisions.
Prof. Rachee Singh, Cornell, “When Bandwidth Isn't the Bottleneck: Breaking Synchronization Barriers in Distributed ML”
Abstract: The future of efficient distributed ML lies not only in increasing bandwidth, but in addressing delays introduced by bulk synchronous communication. We present two key advances in this direction. The first is a redesign of collective communication algorithms that eliminates global synchronization barriers while maintaining bandwidth optimality. The second replaces traditional collectives with asynchronous device-initiated communication in megakernels for communication-intensive MoE models. Together, these advances make the case that without addressing the synchronization tax, scaling network bandwidth alone leaves performance on the table.
Bio: Rachee Singh is an assistant professor of Computer Science at Cornell University. Her research improves the efficiency of communication over photonic interconnects. Singh's research group develops systems and algorithms that use these gains to improve the performance of distributed machine learning and large-scale cloud applications. Her work has been recognized with awards from Amazon, Cisco and Google. Her PhD dissertation won the SIGCOMM doctoral dissertation award.
Prof. Hakim Weatherspoon, Cornell, "AI and Optimal Oblivious Reconfigurable Networks"
Abstract: As machine learning models scale to hundreds of billions of parameters and beyond, distributed training has become increasingly communication-bound, making cost-efficient network design critical. However, as Moore's Law slows down, packet switch capabilities are falling behind datacenter demands. Recent hardware advances have enabled the new switching technology of nanosecond-scale rapid circuit switches. Combined with novel network designs, these have the potential to fully replace packet switches. This talk presents the Oblivious Reconfigurable Network (ORN) design paradigm which is ideally suited to this new switching technology. I describe how to design ORNs that work for AI training and at datacenter scale, supporting tens of thousands of network nodes. A paper of this works appears in SIGCOMM 2024: "Shale: A Practical, Scalable Oblivious Reconfigurable Networks"
Bio: Hakim Weatherspoon is a Professor in the Department of Computer Science at Cornell University, Co-Director of the Cornell Institute for Digital Agriculture (CIDA), Cornell Bowers College of Computer and Information Science Associate Dean for Diversity, Equity, Inclusion, and Belonging, and the Chief Scientist of Exostellar, Inc (http://exostellar.io). His research interests cover various aspects of fault-tolerance, reliability, security, and performance of internet-scale data systems such as cloud and distributed systems. Weatherspoon received his PhD from University of California, Berkeley. Weatherspoon has received awards for his many contributions, including the University of Washington, Allen School of Computer Science and Engineering, Alumni Achievement Award; Alfred P. Sloan Research Fellowship; and a Kavli Fellowship from the National Academy of Sciences. Weatherspoon has also been recognized for his work to promote diversity, earning the University of Washington, College of Engineering, Diamond award and Cornell's Zellman Warhaft Commitment to Diversity Award.
Prof. Kilian Weinberger, Cornell, “Think Before you Speak: Next Gen LLMs with Global Reasoning and External Memory”
Abstract: The dominant paradigm in language modeling—scaling next-token prediction with parametric knowledge storage—delivers impressive capabilities but also fundamental limitations: brittle factual memory, inefficient parameters, and myopic reasoning. Progress requires a shift toward external memory and architectures that reason globally before committing to tokens.
In this talk I present two recent directions that support this claim. First, Limited-Memory Language Models externalize factual knowledge during pre-training, yielding models that are more controllable, verifiable, and parameter-efficient. Second, latent diffusion–augmented language models demonstrate how planning in continuous latent space overcomes the foresight limitations of next-token prediction, improving reasoning and coherence.
Bio: Kilian Weinberger is a Professor in the Department of Computer Science at Cornell University. He received his Ph.D. from the University of Pennsylvania in Machine Learning under the supervision of Lawrence Saul and his undergraduate degree in Mathematics and Computing from the University of Oxford. During his career he has won several best paper awards at ICML (2004), CVPR (2004, 2017), AISTATS (2005) and KDD (2014, runner-up award). In 2011 he was awarded the Outstanding AAAI Senior Program Chair Award and in 2012 he received an NSF CAREER award. He is the recipient of the Daniel M Lazar ‘29 Excellence in Teaching Award (2016) and the Ann S. Bowers Teaching and Advising Excellence Award (2024). As of 2024 he is an ACM and AAAI fellow and in 2021 became a Blavatnik National Awards Finalists. He was elected co-Program Chair for ICML 2016 and for AAAI 2018 and has been on the ICML board since 2016. He became president of the ICML society in 2023. Since 2024 he has been a member of the Sloan Research Fellowships Selection Committee. Kilian Weinberger’s research focuses on Machine Learning and its applications. In particular, he has worked on learning under resource constraints, metric learning, AI in Science, computer vision, autonomous vehicles, Gaussian Processes, and deep learning. Before joining Cornell University, he was an Associate Professor at Washington University in St. Louis and before that he worked as a research scientist at Yahoo! Research in Santa Clara.
The student poster session was conducted during the break at lunch. The Best Poster Award winner was PhD student Salman Abid.
Agenda - Videos of the Presentations
8:30-8:55 – Badge Pick-up – Coffee/Tea and Breakfast Food/Snacks
8:55-9:00 – Welcome – Prof. Hakim Weatherspoon
9:00-9:30 – Prof. Robbert van Renesse, Cornell, “Latency, Consistency, and the Myth of the Linearizable Log”
9:30-10:00 – Dr. Gloire Rubambiza, IBM, “Advancing AI Platforms for Agents (kagenti) and Distributed LLM Inferencing (llm-d)”
10:00-10:30 – Prof. Rachee Singh, Cornell, “When Bandwidth Isn't the Bottleneck: Breaking Synchronization Barriers in Distributed ML”
10:30-11:00 – Prof. Hakim Weatherspoon, Cornell, "AI and Optimal Oblivious Reconfigurable Networks"
11-11:30 – Lightning Round for Student Posters
11:30-12:30 – Lunch and Poster Viewing
12:30-1:00 – Prof. Kilian Weinberger, Cornell, “Think Before you Speak: Next Gen LLMs with Global Reasoning and External Memory”
1:00-1:30 – Prof. Jiaxin Lin, Cornell, “Intelligent Data Movement for Specialized Scale-up and Scale-out Networks”
1:30-2:00 – Craig Barner, Marvell, “Securing the AI Cloud: Silicon-Rooted Trust in the Age of Intelligent Infrastructure”
2:00-2:30 – Prof. Giulia Guidi, Cornell, “Overcoming Parallelism Challenges in Irregular Data Analysis using Sparse Linear Algebra”
2:30-3:00 – Best Poster Award & Reception
Please see Abstracts and Bios below the Testimonials from Previous Workshops.
Testimonials from Previous Workshops
Professor David Patterson, the Pardee Professor of Computer Science, UC Berkeley, “I saw strong participation at the Cloud Workshop, with some high energy and enthusiasm; and I was delighted to see industry engineers bring and describe actual hardware, representing some of the newest innovations in the data center.”
Professor Christos Kozyrakis, Professor of Electrical Engineering & Computer Science, Stanford University, “As a starting point, I think of these IAP workshops an intersection of industry’s newest solutions in hardware with academic research in computer architecture; but more so, these workshops additionally cover new subsystems and applications, and in a smaller venue where it is easy to discuss ideas and cross-cutting approaches with colleagues.”
Professor Hakim Weatherspoon, Professor of Computer Science, Cornell University, “I have participated in three IAP Workshops since the first one at Cornell in 2013 and it is great to see that the IAP premise is a success now as it was then, bringing together industry and academia in a focused all-day exchange of ideas. It was a fantastic experience and I look forward to the next one!”
Dr. Carole-Jean Wu, Research Scientist, AI Infrastructure, Facebook Research, and Professor of CSE, Arizona State University, “IAP Workshops provide valuable interactions among faculty, students and industry. The smaller venue and the poster session foster an interactive environment for in-depth discussions and spark new collaborative opportunities. Thank you for organizing this wonderful event! It was very well run.”
Dr. Pankaj Mehra, VP Product Planning, Samsung (currently CEO Elephance Memory), "Terrifically organized Workshops that give all parties -- students, faculty, industry -- valuable insights to take back."
Abstracts and Bios (alphabetically by last name). Please Check Back for Updates.
Craig Barner, Marvell, “Securing the AI Cloud: Silicon-Rooted Trust in the Age of Intelligent Infrastructure”
Abstract: Artificial intelligence is reshaping cloud infrastructure faster than the underlying security foundations can keep pace. The traditional model of trusting the software and defending the perimeter is breaking down. In its place, a more fundamental shift is emerging, one that redefines where trust begins.
This talk explores how the industry is rethinking trust from the ground up: why the chip itself has become the new security boundary, how that changes everything from data centers to AI agents, and what it means for the future of sovereign cloud and data privacy. We’ll also look at how companies like Marvell are turning these ideas into silicon that runs at the heart of the world’s largest clouds, and what it takes to build security infrastructure at that scale.
If you’ve ever wondered what actually keeps AI systems secure, or whether they are, this talk is for you.
Bio: Craig Barner is a Senior Distinguished Engineer at Marvell Technology, where he leads the architecture and development of security solutions for custom silicon platforms. With over 25 years of experience in hardware security, he focuses on building cryptographic technologies that safeguard data across cloud platforms, AI-driven data centers, and the distributed infrastructure enabling security at scale. His work centers on some of the most important shifts in cloud security today, including the transition to post-quantum cryptography, the move toward hardware-rooted Zero Trust, and the growing need for confidential computing at scale. As new standards such as ML-KEM and ML-DSA take shape, Craig is helping drive the development of crypto-agile silicon designed to stay ahead of evolving threats. He holds over a dozen patents spanning hardware security and a wide spectrum of hardware innovations, and has authored widely adopted technical specifications in areas such as cryptographic acceleration, secure key management, firmware integrity, and system-level security. More broadly, his work focuses on embedding strong security directly into the silicon layer to help enable trusted, large-scale computing.
Prof. Giulia Guidi, Cornell, Overcoming Parallelism Challenges in Irregular Data Analysis using Sparse Linear Algebra
Abstract: Today, scaling AI beyond single-node limits has shifted workloads to deeply hierarchical multi-GPU systems, where communication becomes the bottleneck. Central operations in modern large-scale AI including GNN training, embedding clustering, and sparse attention, reduce to the same core primitives: sparse and dense matrix multiplications. Yet, scaling these primitives in isolation is not enough; the real challenge is efficiently composing them, as communication costs and data layouts rarely align across operations.
In this talk, we show how to (a) map clustering to sparse and dense linear algebra for efficient single-GPU computation and (b) scale these primitives efficiently on large-scale systems. In particular, we present communication-efficient distributed-memory algorithms for Kernel K-Means that combine GEMM and SpMM with partitioning schemes designed to minimize communication. Our 1.5D algorithm enables exact Kernel K-Means on million-scale datasets [PPoPP '25, IPDPS '26]. This is especially important for AI for science, where data irregularity (graphs, sequences, sparse observations) is intrinsic, and where co-designed sparse and dense linear algebra with the architecture provides a natural path forward.
Bio: Giulia Guidi is an Assistant Professor in the Department of Computer Science and a Graduate Field Faculty in the Department of Computational Biology, the Center for Applied Math, and the Department of Electrical and Computer Engineering at Cornell. She received her PhD in Computer Science from UC Berkeley. She works in the field of high-performance computing for large-scale computational sciences. She is interested in developing algorithms and software infrastructures on parallel machines to speed up data processing without sacrificing programming productivity and to make high-performance computing more accessible. Dr. Guidi received the 2024 SIAM Activity Group on Supercomputing Early Career Prize.
Prof. Jiaxin Lin, Cornell, “Intelligent Data Movement for Specialized Scale-up and Scale-out Networks”
Abstract: Data movement has become a critical bottleneck for next-generation data center workloads, driven largely by the massive communication demands of modern AI applications. To overcome these limitations, distributed systems are increasingly relying on specialized network devices and advanced interconnects to efficiently move data between distributed GPUs and CPUs. These domain-specialized devices and interconnects have unique characteristics and create complex system-level problems. How can we harness the massive performance benefits of specialization without being overwhelmed by its complexity?
In this talk, I will demonstrate how we can solve this by re-architecting the data movement stack to fully harness the benefits of specialized networks while making them accessible. I will explore this through two distinct examples in the scale-out and scale-up domains. First, I will introduce Alkali, a compiler framework that tackles the complexities of programming and utilizing domain-specialized smart network interface cards (SmartNICs). Next, I will present DirectKernel, an end-to-end framework designed for scale-up, GPU-to-remote memory data movement that tackles the LLM memory wall.
Bio: Jiaxin Lin is an assistant professor at Cornell ECE. She received her Ph.D. from UT Austin. Her research aims to design innovative software and hardware for data center networks. She is widely recognized for her influential work on programmable networks. She has received a Google Junior Faculty Award in 2025, as well as Google and Meta PhD Fellowships in Computer Networking, and was selected as an MIT EECS Rising Star in 2024.
Prof. Robbert van Renesse, Cornell, “Latency, Consistency, and the Myth of the Linearizable Log”
Abstract: We often treat linearizability as the gold standard for distributed systems. Yet in shared logs, the workhorses behind replication, coordination, and transactions, this ideal becomes a liability. Enforcing linearizable appends forces global synchronization into the critical path, bloating tail latencies into the millisecond range. This talk explores a different design point: logs that guarantee total order but not real-time precedence. I’ll introduce Ziplog, a coordination-free ordered log that achieves 16 µs end-to-end latency, and Zipkvs, a transactional key-value store built atop it that delivers stronger semantics than today’s RDMA systems at higher performance.
Bio: Robbert van Renesse is a Professor of Computer Science at Cornell University, where his research focuses on distributed systems, fault tolerance, and operating systems. His work has shaped modern thinking about scalable and reliable distributed computing, from early group-communication systems to recent advances in low-latency coordination and consistency. He is co-Editor-in-Chief of ACM Transactions on Computer Systems and a former Chair of ACM SIGOPS. An ACM Fellow, Robbert has served as Program Chair for SOSP, OSDI, and EuroSys.
Gloire Rubambiza, IBM, “Advancing AI Platforms for Agents (kagenti) and Distributed LLM Inferencing (llm-d)”
Abstract: Agentic AI, where AI agents leverage large language models (LLMs) and tools for coding and other personal tasks, is emerging as an exciting new category of cloud computing workload. The effectiveness of generative AI agents critically depends on the diversity and reasoning capabilities of the LLMs used by the agents. However, serving these dynamic AI workloads on enterprise clouds faces numerous challenges such limited GPU resources, delegation and security of agents/tools, cold start for model startup and inferencing, etc. In this talk, I will present two threads of work on building cloud-native platforms for the agentic era. First, I will briefly describe our work on agent workload lifecycle management in the kagenti project. Second, the exemplary workloads in kagenti will serve as motivation for a deeper dive into our work on fast model actuation (FMA) for distributed LLM inferencing on Kubernetes in the llm-d project. Specifically, I will describe different FMA techniques that have gradually enabled us to host multiple models on a single GPU and reduce startup time of a Llama-3.1-8B-Instruct model by an order of magnitude from 1 minute 32 seconds to 3 seconds.
Bio: Dr. Gloire Rubambiza is a Research Scientist at IBM, where he advances the state of the art in hybrid cloud infrastructure for AI agents and LLM inferencing. Before rejoining the "big blue" family, he led research and engineering efforts to build distributed cloud infrastructure for AI in space at Satlyt, a startup in Silicon Valley. Prior to his startup journey, he was a PhD student and postdoc in Computer Science at Cornell University, where he conducted research in hybrid cloud computing for digital agriculture with an emphasis on societal impact. His research approach strategically combines techniques from computer networking and critical human-computer interaction (HCI) into a single approach called trilingualism, which continuously integrates (1) networked systems building, (2) critical domain applications, and (3) critical reflections on the technical decisions.
Prof. Rachee Singh, Cornell, “When Bandwidth Isn't the Bottleneck: Breaking Synchronization Barriers in Distributed ML”
Abstract: The future of efficient distributed ML lies not only in increasing bandwidth, but in addressing delays introduced by bulk synchronous communication. We present two key advances in this direction. The first is a redesign of collective communication algorithms that eliminates global synchronization barriers while maintaining bandwidth optimality. The second replaces traditional collectives with asynchronous device-initiated communication in megakernels for communication-intensive MoE models. Together, these advances make the case that without addressing the synchronization tax, scaling network bandwidth alone leaves performance on the table.
Bio: Rachee Singh is an assistant professor of Computer Science at Cornell University. Her research improves the efficiency of communication over photonic interconnects. Singh's research group develops systems and algorithms that use these gains to improve the performance of distributed machine learning and large-scale cloud applications. Her work has been recognized with awards from Amazon, Cisco and Google. Her PhD dissertation won the SIGCOMM doctoral dissertation award.
Prof. Hakim Weatherspoon, Cornell, "AI and Optimal Oblivious Reconfigurable Networks"
Abstract: As machine learning models scale to hundreds of billions of parameters and beyond, distributed training has become increasingly communication-bound, making cost-efficient network design critical. However, as Moore's Law slows down, packet switch capabilities are falling behind datacenter demands. Recent hardware advances have enabled the new switching technology of nanosecond-scale rapid circuit switches. Combined with novel network designs, these have the potential to fully replace packet switches. This talk presents the Oblivious Reconfigurable Network (ORN) design paradigm which is ideally suited to this new switching technology. I describe how to design ORNs that work for AI training and at datacenter scale, supporting tens of thousands of network nodes. A paper of this works appears in SIGCOMM 2024: "Shale: A Practical, Scalable Oblivious Reconfigurable Networks"
Bio: Hakim Weatherspoon is a Professor in the Department of Computer Science at Cornell University, Co-Director of the Cornell Institute for Digital Agriculture (CIDA), Cornell Bowers College of Computer and Information Science Associate Dean for Diversity, Equity, Inclusion, and Belonging, and the Chief Scientist of Exostellar, Inc (http://exostellar.io). His research interests cover various aspects of fault-tolerance, reliability, security, and performance of internet-scale data systems such as cloud and distributed systems. Weatherspoon received his PhD from University of California, Berkeley. Weatherspoon has received awards for his many contributions, including the University of Washington, Allen School of Computer Science and Engineering, Alumni Achievement Award; Alfred P. Sloan Research Fellowship; and a Kavli Fellowship from the National Academy of Sciences. Weatherspoon has also been recognized for his work to promote diversity, earning the University of Washington, College of Engineering, Diamond award and Cornell's Zellman Warhaft Commitment to Diversity Award.
Prof. Kilian Weinberger, Cornell, “Think Before you Speak: Next Gen LLMs with Global Reasoning and External Memory”
Abstract: The dominant paradigm in language modeling—scaling next-token prediction with parametric knowledge storage—delivers impressive capabilities but also fundamental limitations: brittle factual memory, inefficient parameters, and myopic reasoning. Progress requires a shift toward external memory and architectures that reason globally before committing to tokens.
In this talk I present two recent directions that support this claim. First, Limited-Memory Language Models externalize factual knowledge during pre-training, yielding models that are more controllable, verifiable, and parameter-efficient. Second, latent diffusion–augmented language models demonstrate how planning in continuous latent space overcomes the foresight limitations of next-token prediction, improving reasoning and coherence.
Bio: Kilian Weinberger is a Professor in the Department of Computer Science at Cornell University. He received his Ph.D. from the University of Pennsylvania in Machine Learning under the supervision of Lawrence Saul and his undergraduate degree in Mathematics and Computing from the University of Oxford. During his career he has won several best paper awards at ICML (2004), CVPR (2004, 2017), AISTATS (2005) and KDD (2014, runner-up award). In 2011 he was awarded the Outstanding AAAI Senior Program Chair Award and in 2012 he received an NSF CAREER award. He is the recipient of the Daniel M Lazar ‘29 Excellence in Teaching Award (2016) and the Ann S. Bowers Teaching and Advising Excellence Award (2024). As of 2024 he is an ACM and AAAI fellow and in 2021 became a Blavatnik National Awards Finalists. He was elected co-Program Chair for ICML 2016 and for AAAI 2018 and has been on the ICML board since 2016. He became president of the ICML society in 2023. Since 2024 he has been a member of the Sloan Research Fellowships Selection Committee. Kilian Weinberger’s research focuses on Machine Learning and its applications. In particular, he has worked on learning under resource constraints, metric learning, AI in Science, computer vision, autonomous vehicles, Gaussian Processes, and deep learning. Before joining Cornell University, he was an Associate Professor at Washington University in St. Louis and before that he worked as a research scientist at Yahoo! Research in Santa Clara.