Databricks presents Dolly 2.0 — an open source language model for enterprises and startups
Data protection and compliance are critical topics for businesses today, especially for startups and small to medium-sized companies. As a data-privacy-as-a-service startup, heyData offers an all-in-one platform solution that helps companies manage their data privacy and compliance requirements efficiently. In this context, Dolly 2.0, Databricks' latest language model, is highly relevant.
As a data-privacy-as-a-service startup, we're constantly on the lookout for innovative solutions to help startups, companies, and founders meet their data privacy and compliance requirements. Today, we're excited to introduce a groundbreaking innovation: Dolly 2.0, the world's first open and instruction-tuned language model (LLM), developed by Databricks.
Background
In the world of AI models, Databricks has made a remarkable breakthrough with Dolly 2.0. Dolly 2.0 is a ChatGPT-like language model trained for less than $30. It's built on EleutherAI's Pythia model family and was fine-tuned by Databricks employees using a crowdsourced, human-generated instruction dataset licensed for research and commercial use. This unique model is now available as open source, giving companies and startups a cost-effective way to build and customize powerful language models for their conversational applications.
Open Source: Free Use for Enterprises and Startups
Dolly 2.0 is released as open source, meaning organizations can build, customize, and own their own language models without paying for costly API services or sharing data with third parties. This is a groundbreaking development for companies and startups looking for cost-effective solutions for their conversational applications. With Dolly 2.0, they have the freedom to adapt and extend the model to meet their specific needs.
A Unique Instruction Dataset for Fine-Tuning
A standout feature of Dolly 2.0 is its instruction dataset, crowdsourced from Databricks employees. The Databricks-Dolly-15k dataset, containing 15,000 prompt/response pairs, was specifically developed to fine-tune large language models on instructions and is available under a Creative Commons Attribution-ShareAlike 3.0 Unported license. This means it can be used, modified, or extended by anyone, including for commercial applications. This dataset is the first open source, human-generated instruction dataset that enables large language models to demonstrate ChatGPT-like interactivity. It contains natural, expressive training records representing a range of behaviors such as brainstorming, content creation, information extraction, and summarization.
Motivation
The creation of the new Dolly 2.0 dataset was motivated by user requests asking whether they could use Dolly commercially, in order to get around the commercial-use restrictions of the original Dolly 1.0 model. Databricks responded to these requests by creating a new dataset specifically suited for commercial applications. This underscores Databricks' commitment to its users' needs and gives companies access to powerful language models for their commercial applications. The motivation for creating this new dataset stemmed from user requests asking whether they could use Dolly commercially, in order to get around the commercial-use restrictions of the original Dolly 1.0 model. Dolly 1.0 was trained on a dataset from the Stanford Alpaca team using the OpenAI API, which brought commercial-use restrictions due to the terms of service. Databricks therefore decided to create a new dataset that was not "contaminated" and could be used for commercial purposes. To do this, the team drew inspiration from OpenAI's InstructGPT research paper and had Databricks employees compete to generate an original, high-quality dataset covering tasks such as open- and closed-ended Q&A, information extraction and summarization from Wikipedia, classification, and creative writing.
Inspired by InstructGPT
The development of Dolly 2.0 was inspired by OpenAI's groundbreaking research paper on InstructGPT. InstructGPT is a language model specifically trained to follow instructions and complete complex tasks. Dolly 2.0 is built on EleutherAI's Pythia model family and was trained on the Databricks-Dolly-15k dataset to develop similar instruction-following capabilities. This enables Dolly 2.0 to handle a wide range of tasks such as brainstorming, content creation, information extraction, and summarization, providing real support for users across various application areas.
Benefits for Startups and Enterprises
Releasing Dolly 2.0 as open source offers numerous benefits for startups and enterprises. Here are some of the most important:
- Cost efficiency: Since Dolly 2.0 is available as open source, startups and companies can use the software for free without paying expensive licensing fees or subscriptions. This lets them allocate resources to other important aspects of their business.
- Flexibility: As open source software, Dolly 2.0 gives users the freedom to adapt and customize it to their own needs. Startups and companies can tailor Dolly 2.0's functions and features to their specific requirements to build customized solutions.
- Community engagement: The open source community is known for its collaboration and sharing of knowledge and resources. By releasing Dolly 2.0 as open source, startups and companies can benefit from working with the developer community to fix bugs, implement new features, and continue improving the software.
- Faster innovation: Open source lets startups and companies build on an existing codebase, enabling them to develop innovative solutions faster. By using Dolly 2.0 as open source, they can benefit from the work of other developers and build their own innovations on a proven platform.
- Interoperability: As open source software, Dolly 2.0 can be integrated with a variety of technologies and systems, giving startups and companies the ability to interact with other products and services and extend their functionality.
- Transparency and trust: Because Dolly 2.0's source code is available as open source, startups and companies can review the code and verify that it's secure and trustworthy. This can help strengthen customer and user trust in the software.
- Shared resources: Using open source lets startups and companies share resources and connect with other developers and organizations. This can lead to more efficient use of resources and create synergies for working together on new solutions.
- Adaptability: Dolly 2.0's open source nature allows startups and companies to adapt the software to new technologies, market demands, or business models. This lets them respond with agility and continuously improve their solutions to stay competitive.
In summary, Databricks' Dolly 2.0 gives companies a powerful solution for building language models with advanced capabilities such as transfer learning, cultural adaptation, and monitoring features. It enables companies to build high-quality, adaptable language models and integrate them into their existing workflows and data processing pipelines. With Dolly 2.0, companies can leverage the benefits of language AI technology to enhance their use cases, optimize communication with their target audience, and streamline their business processes.
Data Protection and Compliance With Databricks and Dolly 2.0: Ensuring Conformity
When it comes to data protection and compliance, Databricks ensures that the use of Dolly 2.0 aligns with applicable data protection policies and regulations. Companies can implement their own data protection policies and ensure that data processing with Dolly 2.0 complies with relevant regulations. This is critically important, as data protection and compliance are becoming increasingly significant, and companies are required to adequately protect their customers' and users' data.
Overall, Databricks' Dolly 2.0 gives companies a powerful and flexible solution for building language models tailored to their individual needs. By combining open source capabilities, a comprehensive dataset, and a scalable cloud platform, Dolly 2.0 enables companies to use advanced language models to streamline their business processes, develop customer-focused solutions, and meet data protection and compliance requirements.
Conclusion
Overall, Dolly 2.0 is a groundbreaking development in the world of language models and AI technology. As the first open source language model trained on human-generated instruction datasets and suited for commercial use, Dolly 2.0 gives startups, companies, and founders a unique opportunity to build and customize powerful language models for conversational applications, without relying on paid API access or having to share sensitive data with third parties.
With the Databricks-Dolly-15k dataset, also available as open source and containing over 15,000 high-quality prompt/response pairs created by Databricks employees, developers can access and use a wide range of behaviors to further improve their models and tailor them to their specific needs. The release of this dataset was a response to user concerns and requests regarding the commercial use of Dolly, and it demonstrates Databricks' commitment to advancing open source and supporting the developer community.
Dolly 2.0 also exemplifies the advances being made in AI technology and the possibilities that emerge from combining human creativity with machine learning. With a high-quality instruction dataset created by Databricks employees, Dolly 2.0 offers an impressive level of interactivity and versatility, enabling companies to develop innovative applications for content creation, information extraction, summarization, and more.
Overall, Dolly 2.0 is a milestone in the development of language models, giving companies the ability to build and customize their own models to interact naturally with users and achieve their business goals. With Dolly 2.0 now available as open source, developers have new opportunities and resources to build innovative solutions and continue pushing the boundaries of AI technology.







