跳到主要导航 跳到搜索 跳到主要内容

Rocket: Warming Serverless Inference via Hierarchical ML Artifact Pre-loading and Sharing

  • Beijing Institute of Technology
  • Temple University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Serverless computing is a promising method to serve Machine Learning (ML) inference via on-demand functions. Due to the time- and memory-consuming ML library and model (i.e., ML artifact) loading, serverless inference endures notable startup overhead and memory waste issues. In this paper, we advocate for hierarchical ML artifact pre-loading and sharing to balance loading and memory efficiency. Building on this, we propose Rocket, a serverless ML inference system that accelerates function startup while reducing memory waste. Rocket dynamically pre-loads partial, shared ML artifacts, each implying a hierarchy of trade-offs between the loading latency and memory usage. Specifically, with a dual-timescale invocation prediction, Rocket first estimates the pre-loading timing for each function, and then schedules them via a sharing-aware agglomerative clustering to improve ML artifact sharing efficiency. In particular, Rocket learns to make the online hierarchical pre-loading decision for function containers based on a lightweight contextual bandit algorithm. Finally, we implement Rocket and evaluate it with realistic workloads. Experimental results display that Rocket outperforms existing solutions by up to 38.7% on startup latency and up to 43.8% on memory saving.

源语言英语
主期刊名INFOCOM 2026 - IEEE Conference on Computer Communications
出版商Institute of Electrical and Electronics Engineers Inc.
ISBN(电子版)9798331549619
DOI
出版状态已出版 - 2026
已对外发布
活动2026 IEEE Conference on Computer Communications, INFOCOM 2026 - Tokyo, 日本
期限: 18 5月 202621 5月 2026

丛书

姓名Proceedings - IEEE INFOCOM
ISSN(印刷版)0743-166X

会议

会议2026 IEEE Conference on Computer Communications, INFOCOM 2026
国家/地区日本
Tokyo
时期18/05/2621/05/26

学术指纹

探究 'Rocket: Warming Serverless Inference via Hierarchical ML Artifact Pre-loading and Sharing' 的科研主题。它们共同构成独一无二的学术指纹。

引用此