Databricks Releases Commercially Viable Open Source ChatGPT Alternative

Databricks launches Dolly 2.0, an open source text-generating AI model that avoids restrictive licensing and proprietary data constraints. The model relies entirely on employee-generated training data to power basic chatbot and search applications.

Databricks releases Dolly 2.0, a new open source text-generating AI model that allows commercial use for both independent developers and companies. Unlike many competing open source models, Dolly 2.0 avoids using output data from OpenAI, ensuring that users do not violate any terms of service. CEO Ali Ghodsi states that this release promotes transparency and enables organizations to build AI-powered applications using their own proprietary data.

To train this new model, Databricks relies on a unique dataset of 15,000 records generated entirely by its own employees. This volunteer-sourced data guides an existing open source model called GPT-J-6B to follow instructions in a chatbot-like manner. While the company frames this move as a push for industry openness, the strategy also encourages developers to build and host their AI applications directly on the Databricks platform.

Despite its innovative training approach, Dolly 2.0 inherits several limitations from its base model. It currently only generates text in English and carries the risk of producing toxic or offensive responses due to its training on internet-scraped data. Additionally, early testing reveals that the model struggles with factual consistency, sometimes providing stereotyped or inaccurate answers to user prompts.

Read More at the original source →