DataPerf: Benchmarks for Data-Centric AI Development

Mazumder, Mark; Banbury, Colby; Yao, Xiaozhe; Karlaš, Bojan; Gaviria Rojas, William; Diamos, Sudnya; Diamos, Greg; He, Lynn; Parrish, Alicia; Kirk, Hannah Rose; Quaye, Jessica; Rastogi, Charvi; Kiela, Douwe; Jurado, David; Kanter, David; Mosquera, Rafael; Cukierski, Will; Ciro, Juan; Aroyo, Lora; Acun, Bilge; Chen, Lingjiao; Raje, Mehul; Bartolo, Max; Eyuboglu, Evan Sabri; Ghorbani, Amirata; Goodman, Emmett; Howard, Addison; Inel, Oana; Kane, Tariq; Kirkpatrick, Christine R.; Sculley, D.; Kuo, Tzu-Sheng; Mueller, Jonas W.; Thrush, Tristan; Vanschoren, Joaquin; Warren, Margaret; Williams, Adina; Yeung, Serena; Ardalani, Newsha; Paritosh, Praveen; Zhang, Ce; Zou, James Y.; Wu, Carole-Jean; Coleman, Cody; Ng, Andrew; Mattson, Peter; Janapa Reddi, Vijay

DataPerf: Benchmarks for Data-Centric AI Development

Part of Advances in Neural Information Processing Systems 36 (NeurIPS 2023) Datasets and Benchmarks Track

Authors

Mark Mazumder, Colby Banbury, Xiaozhe Yao, Bojan Karlaš, William Gaviria Rojas, Sudnya Diamos, Greg Diamos, Lynn He, Alicia Parrish, Hannah Rose Kirk, Jessica Quaye, Charvi Rastogi, Douwe Kiela, David Jurado, David Kanter, Rafael Mosquera, Will Cukierski, Juan Ciro, Lora Aroyo, Bilge Acun, Lingjiao Chen, Mehul Raje, Max Bartolo, Evan Sabri Eyuboglu, Amirata Ghorbani, Emmett Goodman, Addison Howard, Oana Inel, Tariq Kane, Christine R. Kirkpatrick, D. Sculley, Tzu-Sheng Kuo, Jonas W Mueller, Tristan Thrush, Joaquin Vanschoren, Margaret Warren, Adina Williams, Serena Yeung, Newsha Ardalani, Praveen Paritosh, Ce Zhang, James Y Zou, Carole-Jean Wu, Cody Coleman, Andrew Y. Ng, Peter Mattson, Vijay Janapa Reddi

Abstract

Machine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithfulness of the underlying problems. Neglecting the fundamental importance of data has given rise to inaccuracy, bias, and fragility in real-world applications, and research is hindered by saturation across existing dataset benchmarks. In response, we present DataPerf, a community-led benchmark suite for evaluating ML datasets and data-centric algorithms. We aim to foster innovation in data-centric AI through competition, comparability, and reproducibility. We enable the ML community to iterate on datasets, instead of just architectures, and we provide an open, online platform with multiple rounds of challenges to support this iterative development. The first iteration of DataPerf contains five benchmarks covering a wide spectrum of data-centric techniques, tasks, and modalities in vision, speech, acquisition, debugging, and diffusion prompting, and we support hosting new contributed benchmarks from the community. The benchmarks, online evaluation platform, and baseline implementations are open source, and the MLCommons Association will maintain DataPerf to ensure long-term benefits to academia and industry.

DataPerf: Benchmarks for Data-Centric AI Development

Authors

Abstract

Name Change Policy