By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
AIModelKitAIModelKitAIModelKit
  • Home
  • News
    NewsShow More
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    SpaceXAI’s Grok Tool Uploading Users’ Entire Codebase to Cloud Storage: What You Need to Know
    4 Min Read
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    New York Leads the Way: First State to Enforce One-Year Moratorium on New AI Data Centers
    4 Min Read
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    AI Replacing New York Nurses: Why Patients Should be Concerned About Quality of Care
    5 Min Read
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    Navigating AI Agent Crawlers and Cloudflare’s New Rules: A Comprehensive Guide
    5 Min Read
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    How Apple’s Self-Driving Car Program Paved the Way for Advanced AI Chip Technology
    4 Min Read
  • Open-Source Models
    Open-Source ModelsShow More
    Exploring How Mobility Enhances Language Models’ Understanding of Location
    Exploring How Mobility Enhances Language Models’ Understanding of Location
    5 Min Read
    Optimize Candidate Biomarkers with Our AI Tool for Wearable Sensor Data Analysis
    Optimize Candidate Biomarkers with Our AI Tool for Wearable Sensor Data Analysis
    4 Min Read
    Beyond BMI: Assessing Cardiometabolic Risk Using Smartphone Images
    Beyond BMI: Assessing Cardiometabolic Risk Using Smartphone Images
    5 Min Read
    Overcoming Recall Challenges: The Impact of Empty Shelves and Lost Keys on Parametric Factuality
    Overcoming Recall Challenges: The Impact of Empty Shelves and Lost Keys on Parametric Factuality
    6 Min Read
    Enhancing AMIE for Expert-Level Audio-Visual Clinical Consultations
    Enhancing AMIE for Expert-Level Audio-Visual Clinical Consultations
    5 Min Read
  • Guides
    GuidesShow More
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    Your Comprehensive Guide to Practical Constraint Decoding: Basics and Applications
    6 Min Read
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    KDnuggets Weekly Data Science News Roundup: Highlights from July 20, 2026
    4 Min Read
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    Unlock Your AI Potential with Kaggle and Google’s Free 5-Day Agentic AI Course
    6 Min Read
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    Top 5 High-Performance MCP Servers for Optimal Agentic Development
    6 Min Read
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    Top 5 Free Resources for Understanding Agentic AI: Unlock Your Knowledge
    6 Min Read
  • Tools
    ToolsShow More
    Optimizing LFM2.5 Q4_0 Checkpoints through Quantization-Aware Distillation Techniques
    Optimizing LFM2.5 Q4_0 Checkpoints through Quantization-Aware Distillation Techniques
    4 Min Read
    Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
    Deploy Qwen 3.8-2.4T-A95B: A Configurable 2.4T Parameter Model on NVIDIA GB300 NVL72 for Enhanced Reasoning
    6 Min Read
    Optimize Your AI Models with Baseten on Hugging Face Inference Providers 🔥
    Optimize Your AI Models with Baseten on Hugging Face Inference Providers 🔥
    5 Min Read
    July 2026 Security Incident Disclosure: Key Insights and Updates
    July 2026 Security Incident Disclosure: Key Insights and Updates
    6 Min Read
    Boosting Performance with Native-Speed vLLM Transformers for Enhanced Modeling Backend
    Boosting Performance with Native-Speed vLLM Transformers for Enhanced Modeling Backend
    5 Min Read
  • Events
    EventsShow More
    Empowering Veteran Students: Effective Teaching Strategies in Technology and Learning
    Empowering Veteran Students: Effective Teaching Strategies in Technology and Learning
    4 Min Read
    NVIDIA Partners with NSF to Enhance AI Research and Education Through State and Regional AI Hubs Across the US
    NVIDIA Partners with NSF to Enhance AI Research and Education Through State and Regional AI Hubs Across the US
    5 Min Read
    South Korea Unveils AI Future at AI Summit with NVIDIA and Strategic Partners
    South Korea Unveils AI Future at AI Summit with NVIDIA and Strategic Partners
    5 Min Read
    NVIDIA Launches First Open-Source GPU-Accelerated Framework for Medical Physics Simulations
    NVIDIA Launches First Open-Source GPU-Accelerated Framework for Medical Physics Simulations
    5 Min Read
    Unlocking the Power of Open Models at Nemotron Labs: Discover the Advantage
    Unlocking the Power of Open Models at Nemotron Labs: Discover the Advantage
    7 Min Read
  • Ethics
    EthicsShow More
    Why Law Enforcement Has Been Advised to Suspend AI Use in Court Cases
    Why Law Enforcement Has Been Advised to Suspend AI Use in Court Cases
    6 Min Read
    Exploring Space Threats from Mirrors and Recognizing AI Drug Innovations: The Download
    Exploring Space Threats from Mirrors and Recognizing AI Drug Innovations: The Download
    5 Min Read
    Understanding AI Bias: How Human Decisions Shape Algorithmic Errors
    Understanding AI Bias: How Human Decisions Shape Algorithmic Errors
    5 Min Read
    How This Company’s Space Mirror Plans Could Threaten the Night Sky for Everyone
    How This Company’s Space Mirror Plans Could Threaten the Night Sky for Everyone
    5 Min Read
    Understanding Orphan Risks in Artificial Intelligence: Insights from Diverging Safety and Compliance Frameworks on AI Companies’ Risk Prioritization
    Understanding Orphan Risks in Artificial Intelligence: Insights from Diverging Safety and Compliance Frameworks on AI Companies’ Risk Prioritization
    5 Min Read
  • Comparisons
    ComparisonsShow More
    Unlocking Self-Knowledge: SKILL-RAG for Enhanced Learning and Filtering in Retrieval-Augmented Generation
    Unlocking Self-Knowledge: SKILL-RAG for Enhanced Learning and Filtering in Retrieval-Augmented Generation
    4 Min Read
    Understanding Decentralization: An Ontological Exploration and Definition
    Understanding Decentralization: An Ontological Exploration and Definition
    5 Min Read
    Microsoft Transitions AI Governance from Policy Frameworks to Real-time Enforcement
    Microsoft Transitions AI Governance from Policy Frameworks to Real-time Enforcement
    6 Min Read
    Optimizing Multi-Turn Reasoning in LLM Agents with Fine-Grained Reward Structures and Effective Credit Assignment Strategies
    Optimizing Multi-Turn Reasoning in LLM Agents with Fine-Grained Reward Structures and Effective Credit Assignment Strategies
    6 Min Read
    Analyzing Prompt-Induced Waste in Coding Agents: Optimizing Reasoning, Effort, Design, and End-to-End Costs
    Analyzing Prompt-Induced Waste in Coding Agents: Optimizing Reasoning, Effort, Design, and End-to-End Costs
    6 Min Read
Search
  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
Reading: Reflecting on the Past and Anticipating the Future
Share
Notification Show More
Font ResizerAa
AIModelKitAIModelKit
Font ResizerAa
  • 🏠
  • 🚀
  • 📰
  • 💡
  • 📚
  • ⭐
Search
  • Home
  • News
  • Models
  • Guides
  • Tools
  • Ethics
  • Events
  • Comparisons
Follow US
  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events
© 2025 AI Model Kit. All Rights Reserved.
AIModelKit > Tools > Reflecting on the Past and Anticipating the Future
Tools

Reflecting on the Past and Anticipating the Future

aimodelkit
Last updated: April 15, 2025 11:16 pm
aimodelkit
Share
Reflecting on the Past and Anticipating the Future
SHARE

Data Is Better Together: Empowering Open-Source Dataset Creation

In the rapidly evolving landscape of machine learning, the collaboration between Hugging Face and Argilla has birthed an innovative initiative known as Data Is Better Together (DIBT). This initiative harnesses the collective power of the open-source community to create impactful datasets that can drive advancements in machine learning models. This article delves into the achievements, community involvement, and tools designed to facilitate collaborative dataset creation.

Contents
  • Community Efforts
  • Cookbook Efforts
  • What Have We Learned?
  • How Can You Get Involved?

Community Efforts

At the heart of the DIBT initiative lies a commitment to fostering community engagement. Our initial focus was on the Prompt Ranking Project, which aimed to compile a dataset of 10,000 prompts—both synthetic and human-generated—ranked by quality. The response from the community was overwhelming:

  • Within days, over 385 individuals joined the initiative.
  • We successfully launched the DIBT/10k_prompts_ranked dataset, which is tailored for prompt ranking tasks and synthetic data generation.
  • This dataset has already been instrumental in developing new models, such as SPIN.

Recognizing the need for inclusivity, we acknowledged that English-centric data was not enough. To address the lack of language-specific benchmarks for open Large Language Models (LLMs), we initiated the Multilingual Prompt Evaluation Project (MPEP). The goal of MPEP is to create a leaderboard that evaluates prompts across multiple languages.

From this project, we achieved several milestones:

  • A curated selection of 500 high-quality prompts from the DIBT/10k_prompts_ranked dataset was translated into various languages.
  • More than 18 language leaders took the initiative to create spaces for these translations.
  • Completed translations have been achieved in Dutch, Russian, and Spanish, with ongoing efforts to expand these translations.

The establishment of a community of dataset builders on Discord has also been a significant achievement, providing a platform for collaboration and knowledge sharing.

More Read

Elevate Your Projects in Spaces with Gradio: A Complete Guide
Explore the New Open Source Qwen3-Next Models: Hybrid MoE Architecture for Enhanced Accuracy and Faster Parallel Processing on NVIDIA Platforms
Microsoft and Hugging Face Strengthen Partnership to Advance AI Collaboration
Master Long Document Processing with Mistral Medium 3 and NVIDIA NIM: A Guide to Building Effective Agents
How to Stream AR Experiences to Your Apple iPad Using NVIDIA Omniverse

Cookbook Efforts

Beyond community involvement, the DIBT initiative is dedicated to equipping individuals with the resources needed to create high-quality datasets independently. This is encapsulated in our Cookbook Efforts, which provide guides and tools that empower users to build valuable datasets tailored to their unique needs.

Some key projects within the cookbook efforts include:

  • Domain Specific Dataset: Designed to jumpstart the creation of domain-specific datasets, this project connects engineers with domain experts to enhance the relevance of the data produced.
  • DPO/ORPO Dataset: Aimed at encouraging the community to produce more DPO-style datasets across various languages and domains, fostering diversity in dataset creation.
  • KTO Dataset: A resource to assist the community in developing their own KTO datasets, enabling a broader range of datasets for different tasks.

What Have We Learned?

Throughout the development of these initiatives, several key insights have emerged:

  • Eagerness to Participate: The community’s response has demonstrated a strong desire to engage in collaborative efforts focused on dataset creation.
  • Addressing Inequalities: Our work has highlighted existing disparities in the availability of comprehensive benchmarks. Certain languages, domains, and tasks remain underrepresented in the open-source community, necessitating targeted efforts to rectify these gaps.
  • Tools for Collaboration: We have identified that many of the necessary tools for effective collaboration already exist. The challenge now lies in harnessing these tools to build valuable datasets collectively.

How Can You Get Involved?

The DIBT initiative is open for continued participation and collaboration. If you’re interested in contributing to the cookbook efforts, here are several ways to get involved:

  • Follow the Project Instructions: Each project has a README file with guidelines on how to contribute. This is your starting point for getting involved.
  • Share Your Datasets: If you have created datasets or have results to share, please contribute them to the community.
  • Provide New Guides and Tools: Your insights and expertise can help others in the community. Offering new guides or tools can significantly enhance the dataset-building process.

For those eager to join this collaborative effort, we invite you to participate in the #data-is-better-together channel on the Hugging Face Discord. This is a space where you can connect with like-minded individuals and share your ideas on what can be developed together.

The strength of the open-source community lies in its ability to collaborate and innovate. With your contributions, we can continue to build better datasets and drive the future of machine learning forward. Join us in this exciting journey of collective dataset creation!

Inspired by: Source

PyTorch Foundation Introduces vLLM as a New Hosted Project
Maximizing Test-Time Compute Performance: How to Secure a Gold Medal at IOI 2025 Using Open-Weight Models
How AI Technology Safeguards Marine Life by Locating Abandoned Fishing Nets in Oceans
Discover the Latest Features in TensorFlow 2.15: Updates from the TensorFlow Blog
Unlocking the Power of Pull Requests and Discussions: A Comprehensive Guide 🥳

Sign Up For Daily Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

By signing up, you agree to our Terms of Use and acknowledge the data practices in our Privacy Policy. You may unsubscribe at any time.
Share This Article
Facebook Copy Link Print
Previous Article How a Small US City is Using AI to Discover Residents’ Needs and Preferences How a Small US City is Using AI to Discover Residents’ Needs and Preferences
Next Article Enhanced Retrieval-Based Explainable Multimodal Modeling for Brain Evaluation and Neurodegenerative Diagnosis in Zero- and Few-Shot Scenarios Enhanced Retrieval-Based Explainable Multimodal Modeling for Brain Evaluation and Neurodegenerative Diagnosis in Zero- and Few-Shot Scenarios

Stay Connected

XFollow
PinterestPin
TelegramFollow
LinkedInFollow

							banner							
							banner
Explore Top AI Tools Instantly
Discover, compare, and choose the best AI tools in one place. Easy search, real-time updates, and expert-picked solutions.
Browse AI Tools

Latest News

Unlocking Self-Knowledge: SKILL-RAG for Enhanced Learning and Filtering in Retrieval-Augmented Generation
Unlocking Self-Knowledge: SKILL-RAG for Enhanced Learning and Filtering in Retrieval-Augmented Generation
Comparisons
Understanding Decentralization: An Ontological Exploration and Definition
Understanding Decentralization: An Ontological Exploration and Definition
Comparisons
Why Law Enforcement Has Been Advised to Suspend AI Use in Court Cases
Why Law Enforcement Has Been Advised to Suspend AI Use in Court Cases
Ethics
Microsoft Transitions AI Governance from Policy Frameworks to Real-time Enforcement
Microsoft Transitions AI Governance from Policy Frameworks to Real-time Enforcement
Comparisons
//

Leading global tech insights for 20M+ innovators

Quick Link

  • Latest News
  • Model Comparisons
  • Tutorials & Guides
  • Open-Source Tools
  • Community Events

Support

  • Privacy Policy
  • Terms of Service
  • Contact Us
  • FAQ / Help Center
  • Advertise With Us

Sign Up for Our Newsletter

Get AI news first! Join our newsletter for fresh updates on open-source models.

AIModelKitAIModelKit
Follow US
© 2025 AI Model Kit. All Rights Reserved.
Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?