Los Angeles-based startup Moonvalley has launched Marey, a video-generating AI model that its creators say was trained exclusively on public domain films, a departure from the industry's reliance on copyrighted material. The tool, which became publicly available after a limited debut in March, uses a credit-based payment system similar to other AI video services.
The announcement comes amid widespread debate over the legality and ethics of training generative AI on proprietary works. While some tech executives argue that access to copyrighted books, music, and video is necessary for progress, Moonvalley's approach offers an alternative that avoids the legal and ethical pitfalls of unlicensed data scraping.
Marey is described as a "3D-aware" video synthesis model, a technical feature that sets it apart from many text-to-video tools. The company has not disclosed the full scope of its training dataset, but the claim of 100% public domain sourcing is central to its pitch to filmmakers and studios.
Hollywood Veteran Joins the Team
Moonvalley's ethical stance has attracted notable industry figures. Ed Ulbrich, a VFX artist and producer known for his work on "Titanic," "The Curious Case of Benjamin Button," and "Top Gun: Maverick," joined the company in June as a liaison to film studios. Ulbrich, who had previously been skeptical of generative AI, said Moonvalley's "clean model" changed his mind.
"No stolen pixels, no scraping of the internet," Ulbrich told Deadline, emphasizing the importance of an ethically sourced and trained AI system. His endorsement lends credibility to Moonvalley's claims, though independent verification of the training data has not yet been conducted.
Public Domain AI: A Growing Trend
Moonvalley is not alone in exploring public domain and openly licensed data for AI training. In June, a team of over two dozen researchers built a large language model (LLM) using only openly licensed or public domain material. The project required sifting through over eight terabytes of data—roughly equivalent to 1,685,461 Bibles—to ensure copyright compliance. The resulting model performed comparably to Meta's Llama 1 and 2 7B, which are several years old.
These efforts challenge the narrative that AI development must rely on mass data collection from copyrighted sources. While the legal landscape remains unsettled, the public domain approach offers a path that respects intellectual property rights.
Moonvalley's success could influence how other AI companies approach training data, particularly in the creative industries where copyright infringement concerns are acute. As the technology matures, the demand for ethically sourced AI models may grow, potentially reshaping industry standards.
For now, Marey is available to the public, and its reception will be closely watched by both technologists and filmmakers. Whether it can compete with models trained on vast proprietary datasets remains an open question, but its existence proves that an alternative is feasible.