Yan Dai, Maryam Farboodi, Negin Golrezaei, Sepehr Shahshahani
A theoretical study analyzing how both free-for-all and strong copyright protection in the AI training data market hinder creative incentives and AI progress, proposing an alternative via a data intermediary.
How to design a market for human-generated content used in training AI models that simultaneously enables technological progress and preserves individual incentives for high-quality content creation? What problems arise from the existing 'fair use-based free model' and the 'strong intellectual property rights model'?
We analyze the two extreme copyright regimes using a static Stackelberg game model and a dynamic model. In the static model, we discover the 'originality penalty' where strong copyright disadvantages innovative creators. In the dynamic model, we find the 'curse of precision': AI-assisted creation homogenizes content, which when fed back into training degrades AI performance. To solve this, we propose a market design with a data intermediary that internalizes cross-creator externalities and subsidizes innovative contributions.
We prove that both extreme copyright regimes lead to market failure, and identify novel phenomena: the 'originality penalty' against innovative creators and the 'curse of precision' causing AI performance degradation. The proposed data intermediary model is shown to restore efficiency. This research provides important theoretical contributions to the regulation and policy design of AI training data markets.