Write a description of a photo with varying amounts of words.
Like, 10 words, 50 words, 100 words, 500 hundred words, a thousand words, and 5000 words. They all describe the same picture.
You have a group of people read the description and try to recreate the picture (which they have not seen, just read the description)
With more words, we should see the recreations become more consistent with the original and each other. Once everyone can recreate the image accurately, then we know how many words it’s worth. Beyond that point, I imagine extra words would produce diminishing returns.
Alternatively, make an annotated dataset with 1000 words per picture. Then, use that dataset to train an image generating AI. Next, delete 50 words from each picture, and train another AI with that new dataset. Keep on going like that until you have a bunch of image AIs. Test all of them to see how many words do you really need to make a working model. Will there also be a point of diminishing returns, or was the 1000 word model clearly the best one?
Here’s my experiment.
Write a description of a photo with varying amounts of words.
Like, 10 words, 50 words, 100 words, 500 hundred words, a thousand words, and 5000 words. They all describe the same picture.
You have a group of people read the description and try to recreate the picture (which they have not seen, just read the description)
With more words, we should see the recreations become more consistent with the original and each other. Once everyone can recreate the image accurately, then we know how many words it’s worth. Beyond that point, I imagine extra words would produce diminishing returns.
Alternatively, make an annotated dataset with 1000 words per picture. Then, use that dataset to train an image generating AI. Next, delete 50 words from each picture, and train another AI with that new dataset. Keep on going like that until you have a bunch of image AIs. Test all of them to see how many words do you really need to make a working model. Will there also be a point of diminishing returns, or was the 1000 word model clearly the best one?