Chinese language enterprise capitalist Kai-Fu Lee’s synthetic intelligence startup 01.AI has been embroiled in an issue over usages of open-source giant language mannequin structure. The startup, already price over US$1 billion since being based in July, issued assertion in the present day to make clear its usages of present mannequin structure.
01.AI, based by chairman and CEO of Chinese language enterprise agency Innovation Works, Kai-Fu Lee, full a spherical of financing by Alibaba Cloud with valuation reportedly exceeding US$1 billion. It launched two open-sourced pre-trained giant fashions, Yi-34B and Yi-6B, final month on the open-source neighborhood Hugging Face.
After the disclosing of the 2 fashions, Jia Yangqing, former Vice President of Expertise at Alibaba and the inventor of the deep studying framework Caffe, hinted that the Yi fashions only a shell of the LLaMA structure.
LLaMA is a big language mannequin created by Meta, launched in July of this yr and absolutely open-sourced. Some builders have said that aside from two tensors being renamed, Yi fully used the LLaMA structure.
In response to those doubts, 01.AI launched an announcement concerning the coaching technique of Yi-34B, mentioning that “the core of the continual growth and breakthrough of enormous fashions lies not solely within the structure but in addition within the parameters obtained by coaching.”
It clarified that within the course of of coaching the mannequin, 01.AI used the essential structure of GPT/LLaMA, which allowed for a fast begin and was extra developer-friendly. However that Yi-34B and Yi-6B fashions have been educated from scratch by 01.AI and underwent loads of unique optimization work.
Concerning the oversight of renaming some inference code borrowed from LLaMA after experimentation, the startup stated the unique intention was to completely take a look at the mannequin and conduct comparative experiments. The renaming of some inference parameters will not be supposed to intentionally conceal something.
To additional dispel the controversy, 01.AI stated in the present day that after a number of weeks of inner authorized evaluation, the corporate has confirmed that it’s not concerned in any shell or plagiarism points.
After the open-source launch of Yi-34B, a developer named Eric Hartford urged that the Yi mannequin ought to reverse the renaming of the 2 tensors with a purpose to keep consistency in tensor names throughout all fashions primarily based on the open-sourced LLaMA structure.
Following this, 01.AI resubmitted the mannequin and code to numerous open-source platforms, reversing the renaming of the 2 tensor names.
Nevertheless, critics nonetheless level out that the controversy was about how 01.AI offered its Yi fashions. The corporate got here out stating that the Yi fashions are “the primary Chinese language-made mannequin to high the worldwide open-source giant mannequin rankings.” The corporate emphasised how its fashions are “domestically-made”, whereas mentioning nothing about using present open-sourced fashions or gave any credit score to the LLaMA structure.
At the moment, the Yi mannequin has been downloaded 168,000 occasions within the Hugging Face neighborhood. It has gained over 4,900 Stars on GitHub. A number of well-known corporations and establishments have additionally launched fine-tuned fashions primarily based on the Yi mannequin platform.
For instance, OrionStar, a subsidiary of Cheetah Cell, launched the OrionStar-Yi-34B-Chat mannequin, and the Cognitive Computing and Pure Language Analysis Heart of Southern College of Science and Expertise and the Guangdong-Hong Kong-Macao Better Bay Space Digital Financial system Analysis Institute collectively launched the SUS-Chat-34B.
**FAQ**
**1. What’s the controversy surrounding 01.AI’s open-source fashions, Yi-34B and Yi-6B?**
The controversy stems from allegations that 01.AI’s Yi fashions are primarily based on the open-source LLaMA structure created by Meta, with out correct attribution or acknowledgment. Critics have additionally raised considerations concerning the renaming of some inference parameters.
**2. How did 01.AI reply to the doubts and allegations?
01.AI launched an announcement addressing the coaching technique of Yi-34B, emphasizing that whereas the essential structure of GPT/LLaMA was used for a fast begin, the Yi fashions have been educated from scratch and underwent unique optimization work. The corporate additionally said that the renaming of some inference parameters was not supposed to intentionally conceal something.
**3. What steps did 01.AI take to handle the controversy?
After inner authorized evaluation, 01.AI confirmed that it was not concerned in any shell or plagiarism points. Moreover, the corporate reversed the renaming of two tensor names in response to developer suggestions and resubmitted the mannequin and code to open-source platforms.
