Researchers at MIT CSAIL have found that as generative image models are trained on larger datasets, it becomes increasingly difficult to trace their outputs back to any single source image. The team describes this effect as attribution decay.
Published in Nature Communications, the study tested 24 model ensembles built from datasets ranging from 256 images to more than 160,000. To measure influence with greater precision, the researchers retrained models from scratch after removing individual images, rather than relying on estimates.
According to the team, the results suggest that removing one image at a time often does not change the final output, even at scale. MIT professor David Gifford said the method offers an exact way to test whether specific training examples truly shape a model's result.
The findings also point to a broader conversation about how AI-generated visuals are understood in relation to originality, fair use, and authorship. The study showed that the new approach performed comparably to conventional diffusion models by standard quality measures.
As generative systems continue to evolve, research like this may help define a clearer framework for how creative AI is measured, credited, and integrated into the future of digital culture.