849

Please generate an image with NO dogs (lemmy.world)

submitted 7 months ago by isyasad@lemmy.world to c/onehundredninetysix

155 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[-] lvxferre@mander.xyz 90 points 7 months ago

[-] user224@lemmy.sdf.org 66 points 7 months ago

As full as it gets:

Prompts (2):

1. Overflowing wine glass of arch linux femboy essence
2. Make it more furry (as in furry fandom)

I am gonna have fun with this.

[-] uuldika@lemmy.ml 21 points 7 months ago

why do all the femboys run Arch? I'm a NixOS girl and I refuse to convert for any boy no matter how cute he is.

[-] AuroraB 4 points 7 months ago

I use Debian btw. Sometimes even ubuntu, but the snap thing is annoying, so I may switch to another distro at some point.

load more comments (1 replies)

[-] lvxferre@mander.xyz 15 points 7 months ago

It's actually really good, considering the odd request!

[-] QuantumSparkles@sh.itjust.works 10 points 7 months ago

Fiberglass🤤

[-] Rai@lemmy.dbzer0.com 2 points 7 months ago

That’s really good! Could I ask what type of AI this is generated with?

[-] user224@lemmy.sdf.org 4 points 7 months ago

Also Gemini.

[-] Rai@lemmy.dbzer0.com 1 points 7 months ago

Thank you!

[-] lvxferre@mander.xyz 40 points 7 months ago

It gets even worse, but I'll need to translate this one.

[Input 1] Generate a picture containing a copo completely full of wine. The copo must be completely full, with no space to add more wine.
[Output 1] Sure! (Gemini provides a picture containing a taça [stemmed glass] only partially full of wine.)
[Input 2] The picture provided does not fulfill the request. Generate a picture of a copo (not a taça) completely full of wine, with no available space for more wine.
[Output 2] Sure! (Gemini provides yet another half-full taça)

For context, Portuguese uses different words for what English calls a drinking glass:

copo ['kɔ.po]~['kɔ.pu] - non-stemmed drinking glass. The one you likely use everyday.
taça ['tä.sɐ] - stemmed drinking glass, like the ones you'd use with wine.

Both requests demand a full copo but Gemini is rather insistent on outputting half-full taças.

The reason for that is as @will_steal_your_username@lemmy.blahaj.zone pointed out: just like there's practically no training data containing full glasses, there's none for non-stemmed glasses with wine.

[-] Arkhive 5 points 7 months ago

I wonder is something like “a mason jar full to the brim with wine” would do anything interesting. As someone else pointed out the training data for containers of wine is probably disproportionately biased toward stemmed wine glasses that are filled to about the standard restaurant pour.

[-] lvxferre@mander.xyz 3 points 7 months ago

It refuses to generate it!

[Input] Generate a picture containing a mason jar full to the brim with wine.
[Output] I'm still learning how to generate certain kinds of images, so I might not be able to create exactly what you're looking for yet or it may go against my guidelines. If you'd like to ask for something else, just let me know!

[-] brucethemoose@lemmy.world 4 points 7 months ago* (last edited 7 months ago)

This is a misconception. Sort of.

I think the problem is misguided attention. The word "glass of wine" and all the previous context is so strong that it "blows out" the "full glass of wine" as the actual intent. Also, LLMs are still pretty crap at multi turn multimedia understanding. They work are especially prone to repeating previous conversation.

It should be better if you word it like "an overflowing glass with wine splashing out." And clear the history.

I hate to ramble, but this is what I hate most about the way big corpos present "AI." They are narrow tools the user needs to learn how to operate, like photoshop or something, not magic genie lamps like they are trying to sell.

[-] lvxferre@mander.xyz 4 points 7 months ago

There's no previous context to speak of; each screenshot shows a self-contained "conversation", with no earlier input or output. And there's no history to clear, since Gemini app activity is not even turned on.

And even with your suggested prompt, one of the issues is still there:

The other issue is not being tested in this shot as it's language-specific, but it is relevant here because it reinforces that the issue is in the training, not in the context window.

[-] brucethemoose@lemmy.world 4 points 7 months ago

Was just a guess. The AI is still shitty, lol.

What I am trying to get at is the misconception: AI can generate novel content not in its training dataset. An astronaut riding a horse is the classic test case, which did not exist anywhere before diffusion models, and it should be able to extrapolate a fuller wine glass. It’s just too dumb to do it, lol.

[-] Spider2013@lemmy.dbzer0.com 2 points 7 months ago* (last edited 7 months ago)

What if you prompt glass with water , then you paint/tint the water with red

[-] HelterSkeletor@lemmy.world 13 points 7 months ago* (last edited 7 months ago)

Alex O'Connor did an interesting video on this, he's got other videos exploring the shortcomings of LLM 's.

https://youtu.be/160F8F8mXlo

[-] Cassa 10 points 7 months ago

Tbh that is a full glass of wine... it's not supposed to be filled all the way

[-] lvxferre@mander.xyz 14 points 7 months ago

It is not a completely full glass.

it’s not supposed to be filled all the way

What I requested is not what you're "supposed" to do, indeed. You aren't supposed to drink wine from glasses that are completely full. Except when really drunk. But then might as well drink straight from the bottle.

...fuck, I played myself now. I really want some booze.

load more comments (1 replies)

[-] NOT_RICK@lemmy.world 11 points 7 months ago

Probably why it won’t put more in it. How much training data of wine in a glass will have it filled to the brim? Probably next to none.

[-] will_steal_your_username 9 points 7 months ago

You can't tell it to fill it to the brim or be a quarter full either, though. It doesn't have the training data for it

[-] user224@lemmy.sdf.org 6 points 7 months ago

Hmm, I didn't know Gemini could generate images already. My bad, I trusted it to know whether it can do that (it still says it can't when asked).

[-] lvxferre@mander.xyz 3 points 7 months ago

It does for a while already. Frankly, it's the only reason why I'd use Gemini on first place (DDG version of GPT 4-o mini doesn't have a built-in image generator).

[-] Draconic_NEO@lemmy.dbzer0.com 5 points 7 months ago

I wonder, does AI horde also have this problem too?

@aihorde@lemmy.dbzer0.com draw for me a wine glass completely filled to the top style:flux

[-] aihorde@lemmy.dbzer0.com 7 points 7 months ago

Here are some images matching your request

Prompt: a wine glass completely filled to the top

Style: flux

Image with seed 2155117656 generated via AI Horde through @aihorde@lemmy.dbzer0.com. Prompt: a wine glass completely filled to the top

[-] Draconic_NEO@lemmy.dbzer0.com 14 points 7 months ago

Yup Horde still suffers from this issue, though it seems to have more promise than the others considering the second glass is way closer to being full than anything I've sen from openAI or Gemini demonstrations. Maybe there's hope to fix this issue here.

I only tried one model so if you know of a different horde model which works better for this and actually gives a full glass please reply below letting me know, maybe even ask the horde bot to generate it right here.

[-] lvxferre@mander.xyz 5 points 7 months ago

I have considerably less experience with image generation than text generators, but I kind of expect the issue to be only truly fixed if people train the model with a bunch of pictures of glasses full of wine.

I'll run a test using a local tree, that is supposed to look like this:

@aihorde@lemmy.dbzer0.com draw for me a picture of three Araucaria angustifolia trees style:flux

[-] aihorde@lemmy.dbzer0.com 4 points 7 months ago

Here are some images matching your request

Prompt: a picture of three Araucaria angustifolia trees

Style: flux

Image with seed 2535437189 generated via AI Horde through @aihorde@lemmy.dbzer0.com. Prompt: a picture of three Araucaria angustifolia trees

[-] lvxferre@mander.xyz 9 points 7 months ago* (last edited 7 months ago)

Bingo - this tree is non-existent outside my homeland, so people barely speak about it in English - and odds are that the model was trained with almost no pictures of it. However one of the names you see for it in English is Paraná pine, so it's modelling it after images of European pines - because odds are those are plenty in its training set.

[-] joshchandra@midwest.social 2 points 7 months ago

So we could keep having it generate these and poison its own training data!

[-] KeenFlame@feddit.nu 1 points 7 months ago

No. What are you on about?

[-] joshchandra@midwest.social 1 points 7 months ago* (last edited 7 months ago)

What I mean is that if people keep making it produce garbage tied to some keyword or phrase and people publish said garbage, that'll only strengthen AIs' neural network between the bad data and that keyword, so AI results for such trees will drift even further away from the truth.

[-] KeenFlame@feddit.nu 1 points 7 months ago

Publishing fake data that outweighs the data on the real plant is a way, but that doesn't require a plant, you can publish bad images today on any subject

[-] joshchandra@midwest.social 1 points 7 months ago* (last edited 7 months ago)

Right, but I think it'd be harder to get it to unlearn the wrong data if the topic itself is obscure.

[-] KeenFlame@feddit.nu 2 points 7 months ago

Ah. I think I get you, but unfortunately it would probably be a lot easier to unlearn an obscure topic, not the other way around. Poisoning is done at a pixel level if that makes sense

load more comments (1 replies)

[-] Indivisability9559@lemm.ee 3 points 7 months ago

That fourth picture is just four penguins in a trenchcoat

[-] Focal@pawb.social 2 points 7 months ago

Wait, this seems incredible. Do you have to be in the same instance or does it work anywhere? @aihorde@lemmy.dbzer0.com Can you draw a smart phone without a rotary phone dial?

[-] Draconic_NEO@lemmy.dbzer0.com 3 points 7 months ago

It works on any instance that is federated to dbzer0. You have to use annotated mentions though since that's what the bot uses. Like this:
@aihorde@lemmy.dbzer0.com draw for me a smart phone without a rotary phone dial

[-] aihorde@lemmy.dbzer0.com 6 points 7 months ago

Here are some images matching your request

Prompt: a smart phone without a rotary phone dial

Style: flux

Image with seed 2926357957 generated via AI Horde through @aihorde@lemmy.dbzer0.com. Prompt: a smart phone without a rotary phone dial

[-] Draconic_NEO@lemmy.dbzer0.com 3 points 7 months ago

Guess AIhorde had some trouble understanding the prompt too...

[-] Focal@pawb.social 2 points 7 months ago* (last edited 7 months ago)

Thank you very much. I'll give it another shot with the annotation.

@aihorde@lemmy.dbzer0.com

Draw a picture of a poker table without any poker chips what so ever

I think I messed up the annotation

[-] Draconic_NEO@lemmy.dbzer0.com 3 points 7 months ago

Yeah, you also have to say draw for me. I don't think the bot recognizes queries otherwise. Also editing mentions doesn't work, they have to be new, fresh posts with the mention. Just a quirk with Lemmy and how mentions work here.

[-] Focal@pawb.social 3 points 7 months ago

I appreciate your patience with me here :P

@aihorde@lemmy.dbzer0.com draw for me a picture taken at night with a trail camera with absolutely no washing machines roaming free

[-] aihorde@lemmy.dbzer0.com 3 points 7 months ago

Here are some images matching your request

Prompt: a picture taken at night with a trail camera with absolutely no washing machines roaming free

Style: flux

Image with seed 949982511 generated via AI Horde through @aihorde@lemmy.dbzer0.com. Prompt: a picture taken at night with a trail camera with absolutely no washing machines roaming free Image with seed 1973116260 generated via AI Horde through @aihorde@lemmy.dbzer0.com. Prompt: a picture taken at night with a trail camera with absolutely no washing machines roaming free

[-] Draconic_NEO@lemmy.dbzer0.com 3 points 7 months ago

Seems to have understood the negative prompt very well this time. At least AI horde did on the flux model.

CC: @Focal@pawb.social

[-] Focal@pawb.social 3 points 7 months ago

You're right! It did! Thanks a lot, pal! You taught me something super neat and I appreciate it!

[-] Draconic_NEO@lemmy.dbzer0.com 2 points 7 months ago

Glad I could help.

[-] AlienContact2049@lemmy.ca 5 points 7 months ago

I think the AI is just trying to promote healthy drinking habits. /S

[-] Harvey656@lemmy.world 3 points 7 months ago

Full is relatively apparently.

[-] Pofski@lemmy.world 2 points 7 months ago

Ask it to generate a room full of clocks with all of them having the hands at different times. You'll see that all (or almost) all the clocks will say it is 10:10.

this post was submitted on 20 Mar 2025

849 points (100.0% liked)

196

4661 readers

1603 users here now

Community Rules

You must post before you leave

Be nice. Assume others have good intent (within reason).

Block or ignore posts, comments, and users that irritate you in some way rather than engaging. Report if they are actually breaking community rules.

Use content warnings and/or mark as NSFW when appropriate. Most posts with content warnings likely need to be marked NSFW.

Most 196 posts are memes, shitposts, cute images, or even just recent things that happened, etc. There is no real theme, but try to avoid posts that are very inflammatory, offensive, very low quality, or very "off topic".

Bigotry is not allowed, this includes (but is not limited to): Homophobia, Transphobia, Racism, Sexism, Abelism, Classism, or discrimination based on things like Ethnicity, Nationality, Language, or Religion.

Avoid shilling for corporations, posting advertisements, or promoting exploitation of workers.

Proselytization, support, or defense of authoritarianism is not welcome. This includes but is not limited to: imperialism, nationalism, genocide denial, ethnic or racial supremacy, fascism, Nazism, Marxism-Leninism, Maoism, etc.

Avoid AI generated content.

Avoid misinformation.

Avoid incomprehensible posts.

No threats or personal attacks.

No spam.

Moderator Guidelines

Moderator Guidelines

Don’t be mean to users. Be gentle or neutral.
Most moderator actions which have a modlog message should include your username.
When in doubt about whether or not a user is problematic, send them a DM.
Don’t waste time debating/arguing with problematic users.
Assume the best, but don’t tolerate sealioning/just asking questions/concern trolling.
Ask another mod to take over cases you struggle with, if you get tired, or when things get personal.
Ask the other mods for advice when things get complicated.
Share everything you do in the mod matrix, both so several mods aren't unknowingly handling the same issues, but also so you can receive feedback on what you intend to do.
Don't rush mod actions. If a case doesn't need to be handled right away, consider taking a short break before getting to it. This is to say, cool down and make room for feedback.
Don’t perform too much moderation in the comments, except if you want a verdict to be public or to ask people to dial a convo down/stop. Single comment warnings are okay.
Send users concise DMs about verdicts about them, such as bans etc, except in cases where it is clear we don’t want them at all, such as obvious transphobes. No need to notify someone they haven’t been banned of course.
Explain to a user why their behavior is problematic and how it is distressing others rather than engage with whatever they are saying. Ask them to avoid this in the future and send them packing if they do not comply.
First warn users, then temp ban them, then finally perma ban them when they break the rules or act inappropriately. Skip steps if necessary.
Use neutral statements like “this statement can be considered transphobic” rather than “you are being transphobic”.
No large decisions or actions without community input (polls or meta posts f.ex.).
Large internal decisions (such as ousting a mod) might require a vote, needing more than 50% of the votes to pass. Also consider asking the community for feedback.
Remember you are a voluntary moderator. You don’t get paid. Take a break when you need one. Perhaps ask another moderator to step in if necessary.

founded 9 months ago

MODERATORS

SoleInvictus

will_steal_your_username

TheCoolerMia

kittenzrulz123

rockSlayer@lemmy.world

WillStealYourUsername@piefed.blahaj.zone

kittenzrulz123@piefed.blahaj.zone