Write a description of a photo with varying amounts of words.
Like, 10 words, 50 words, 100 words, 500 hundred words, a thousand words, and 5000 words. They all describe the same picture.
You have a group of people read the description and try to recreate the picture (which they have not seen, just read the description)
With more words, we should see the recreations become more consistent with the original and each other. Once everyone can recreate the image accurately, then we know how many words it's worth. Beyond that point, I imagine extra words would produce diminishing returns.
Nice experiment if all stories were created equal. Some people will need a book to describe something that someone else can describe in two lines. Also, some people never need a dictionary for anything where as someone else might only understand a few thousand words. I'd be interested in the data from the experiment, but it will be really hard to come up with something verifiable.
Artistic skill definitely should NOT be a factor. If the photo is a rural landscape, they could put a blob of red in the corner and label it a barn. As long as they have it in the right spot, and it's the right color, I say it counts.
We're not testing their ability to create output, we're testing the description's ability to relay a very specific and consistent message.
Alternatively, make an annotated dataset with 1000 words per picture. Then, use that dataset to train an image generating AI. Next, delete 50 words from each picture, and train another AI with that new dataset. Keep on going like that until you have a bunch of image AIs. Test all of them to see how many words do you really need to make a working model. Will there also be a point of diminishing returns, or was the 1000 word model clearly the best one?
10 replies
Or if we define a picture as 1000 words, we can calculate exactly by how much blind people get scammed
Here's my experiment.
Write a description of a photo with varying amounts of words.
Like, 10 words, 50 words, 100 words, 500 hundred words, a thousand words, and 5000 words. They all describe the same picture.
You have a group of people read the description and try to recreate the picture (which they have not seen, just read the description)
With more words, we should see the recreations become more consistent with the original and each other. Once everyone can recreate the image accurately, then we know how many words it's worth. Beyond that point, I imagine extra words would produce diminishing returns.
Nice experiment if all stories were created equal. Some people will need a book to describe something that someone else can describe in two lines. Also, some people never need a dictionary for anything where as someone else might only understand a few thousand words. I'd be interested in the data from the experiment, but it will be really hard to come up with something verifiable.
Man, I could be staring at the picture and I still couldn't re-create it.
Artistic skill definitely should NOT be a factor. If the photo is a rural landscape, they could put a blob of red in the corner and label it a barn. As long as they have it in the right spot, and it's the right color, I say it counts.
We're not testing their ability to create output, we're testing the description's ability to relay a very specific and consistent message.
Alternatively, make the picture more and more blurry and have people try to recreate the text.
Alternatively, make an annotated dataset with 1000 words per picture. Then, use that dataset to train an image generating AI. Next, delete 50 words from each picture, and train another AI with that new dataset. Keep on going like that until you have a bunch of image AIs. Test all of them to see how many words do you really need to make a working model. Will there also be a point of diminishing returns, or was the 1000 word model clearly the best one?
I'm (not very actively, but better than nothing) working on a new version of Lemvotes, maybe I could add this as one of the sitewide stats
CatsStandingUp Georg who only puts "cat" on every image containing a cat is skewing the average.
He's doing the Lord's work.