Calling open-weights as open-source in marketing materials is the usual misrepresentation. But now with the restriction on commercial use (which is against opensource definition) it is not even open-weights, technically it would be more accurate to call it weights-available.
unrented7977 43 minutes ago [-]
I'm willing to bet a nonzero amount of its training material is GPL, so I'll treat it as GPL licensed instead and use it however the fuck I want.
If AI labs get to ignore licenses, so do we.
user43928 12 minutes ago [-]
It's not going to matter unless you plan to commercially deploy the model, as far as I see.
If you were to generate outputs for commercial use, I think it would still violate this research license, but it's not like they are going to know, are they?
That said, I am disappointed that the model is not actually open-weights as I expected based on the headline.
JaggerJo 17 minutes ago [-]
Agreed.
gregoriol 1 hours ago [-]
Was going to post about this: the last image models with Apache 2.0 license seem to be from 2025, recent Qwen models are "non-commercial use".
bloaf 53 minutes ago [-]
I love the non-commercial clauses because of how many people are using these for deceptive ads and “virtual staging” and fake social media accounts. Anything that makes those guys lives harder while still letting me make silly pictures for my kids and tapestries for my D&D campaign feel fine by me.
user43928 9 minutes ago [-]
This achieves absolutely nothing to that end.
People can continue to use closed SOTA models to generate outputs for commercial or malicious purposes.
What this research license achieves is that we cannot use this model in applications we publish.
tenuousemphasis 38 minutes ago [-]
You think they care about the probably unenforceable license terms?
Luker88 50 minutes ago [-]
Companies can use llm to license-wash open source code regardless of license.
How difficult would it be to use this model to create a second model without licensing issues?
Zambyte 45 minutes ago [-]
Why would you even do that? Just... use it? There hasn't been any legal precedent on if models can even be copyright restricted. Labs just keep publishing license documents as if they matter.
zdragnar 35 minutes ago [-]
Well, it is an indication that it matters to the lab, so if you don't want legal fees to be the first one to set precedent, then it does matter a great deal.
plufz 29 minutes ago [-]
What are the top image models that still use a less restrictive license today?
vunderba 11 minutes ago [-]
Boogu-Image has the Apache 2.0 License [1] (good coherence, but outputs can look synthetic).
And Krea 2 has a community license [2] that is fairly permissive - I think commercial usage is allowed under $1 million.
Boogu-Image scored 6/15 and Krea 2 scored 7/15 on my GenAI Showdown benchmark [3] - only Ideogram4 eclipses them in terms of local models, but its got a far more restrictive license and the JSON structured inputs can be a pain to work with.
My first impression is that it's not so good at following prompt directions. I asked it to place a 3D text made of glass in a particular city. It instead gave me a broken 3D text on a white background. Maybe with different seeds it gets better, but it's more of a trial and error process than reliable results.
fishfasell 2 hours ago [-]
The capabilities of local LLM text-to-image is honestly pretty damn impressive. IMO, I think local image generation is currently ahead of local code generation. I can get an image in seconds locally with the quality being way higher than what I'd expect from a local model. However with coding it's much slower and much less impressive. I'm sure there's a reason for this and I'm not an AI expert so I'll let the smarter folks tell me why, but that's just been my observation thus far.
mft_ 2 hours ago [-]
I've played with diffusion models on and off since the first release of Stable Diffusion - just for amusement, without a particular goal.
Recently, I've been helping a friend's wife with some basic vector images for her sewing hobby (she has what is essentially a CNC sewing machine) and have been super-impressed with FLUX.1-Kontext, which I've been running on my Macbook Pro with mflux. Its ability to (for example) take a photo of a human or an animal and return a line drawing which is recognisably them (rather than just a generic similarish image as I've experienced with other models) is excellent.
It's an older model now, but (AIUI) has the text-handling features baked in, and in my various testing is very reliable at giving me the outputs that I want, without the randomness I've experienced previously. It's big and relatively slow (~3 mins per 512x512 image edit on my M1 Max Mac) but excellent to work with. It's also very straightforward to set up, without the harness complexity of e.g. comfyui.
jLaForest 54 minutes ago [-]
is the cnc sewing machine an off the shelf model or something DIY? I'd love to hear more
mft_ 50 minutes ago [-]
Off the shelf - it’s a Brother. It prints via a proprietary file format (.PES) but there’s an extension for Inkscape that supports creation and export.
gavmor 11 minutes ago [-]
Remember that quality output is a necessary but insufficient property of a generative model.
Prompt-adherence is really hit-or-miss—especially if one lacks the visual vocabulary. Likewise with coding, I find junior devs don't think to prompt re: respecting this-or-that interface, or refactoring to point-free style, etc.
So, as others have said, the artist knows better.
victorbjorklund 2 hours ago [-]
I mean I’m sure it’s the reverse for an artist. They would be less impressed with the image and more impressed with the code quality
26d0 32 minutes ago [-]
The point I think is interesting is that this is just 7B. The current SOTA 7B LLMs are barely usable for quite simple coding.
gedy 1 hours ago [-]
To generalize, LLMs are great at what you are not skilled at.
k__ 26 minutes ago [-]
That's how they're sold, isn't it?
fishfasell 1 hours ago [-]
That's a fair statement, I agree. I'm quite an abysmal artist so I could be a victim of my own bias here
mdp2021 2 hours ago [-]
How do you use this model locally, similarly to using `llama-server -m <model>`?
(I mean: outside direct or substantial use of Python, and running the Neural Network in the most efficient way.)
rwmj 7 minutes ago [-]
Additional question is what kind of local hardware would be required for this? 7B parameters sounds very light weight, but I'm not sure.
Iolaum 1 hours ago [-]
There is difussion.cpp which is intended for those types of models. I set up krea-2-turbo with the help of ChatGPT 2 months ago, if you have a capable computer that's what I would suggest once it becomes supported.
fp64 1 hours ago [-]
on the linked GitHub page they list support Diffusers, ComfyUI, vLLM-Omni, SGLang, and LightX2V with links to each
mdp2021 30 minutes ago [-]
> Diffusers, ComfyUI, vLLM-Omni, SGLang, and LightX2V
I think that's all Python (not a direct executable).
You could just do (see the "Quick Start") four `pip install` and have a dozen lines script to generate the image. But `llama.cpp` and similar do not require e.g. installing Torch (or PyTorch) - you can use `llama.cpp` on a non-specialized machine.
embedding-shape 2 hours ago [-]
Probably ComfyUI is one of the easiest way to get started with local image/video models. Or perhaps vLLM, if they have support for it already, would be something like `vllm serve <model> --omni --port 9080`
> Currently, we support image, audio and video input.
utopiah 45 minutes ago [-]
Seems I'm missing something. Does this model support other inputs?
Image outputs are supported, videos I'm not sure but I don't think that's an output, just a preview of the equirectangular example, so, same question here, what does this model outputs that isn't supported?
mdp2021 14 minutes ago [-]
Not all architectures are supported by llama.cpp . The GGUF format encodes the NN in a standardized way, but then you need code that can use that NN structure.
I understand that llama.cpp could only output text, last time I checked (I do not know how to find a good source for that though).
It's a diffusion model, completely different from autoregressive attention models.
gunalx 1 hours ago [-]
Its happy to see a new open image model from qwen. But the license is a let down. And it dosent even beat their closed qwen3 image wich is already a bit old.
A 7B diffusion model can now render CJK text better than Microsoft Windows.
doctorpangloss 31 minutes ago [-]
Ideogram 4 has been around for a while haha
tomjen3 1 hours ago [-]
Just think about how recently we got that feature in the official ChatGPT image gen. And now we have that running locally — assuming that is, I can figure out how to get this running on my Mac — blows my mind.
jimmydoe 45 minutes ago [-]
Is ChatGPT really that good?
Back in Apr, ChatGPT Images 2.0 has some broken Chinese texts in its featured examples, and they later removed that from blog post. Is 2.5 better now?
Havoc 58 minutes ago [-]
Pretty sure comfyui has a mac executable
samayashar 17 minutes ago [-]
Qwen and Alibaba are the biggest competitor for basically every model out there. They're beating the benchmarks like top-frontier models, focused on open-source and much cheaper than the competitors.
Excited to see what the future holds for them!
d2kx 2 hours ago [-]
God I love the Qwen team. Easily the most diverse set of models from all the Chinese labs. Only Gemini/DeepMind comes close.
hgufj 1 hours ago [-]
I am really grateful to the Chinese Labs for open sourcing their best models. If it was left to the Americans, we would be forced to pay obscene API fees to use them.
jfoster 1 hours ago [-]
Note that the license on this has this in it:
> You shall not use the Materials for any commercial purpose without obtaining a separate commercial license from us.
It probably will be much cheaper to use than other image models, but it seems that will be up to the whims of Qwen/Alibaba rather than just being the cost of putting it in a cloud provider.
They finally fixed their VAE. It really held back their models over the last 2 years.
mdp2021 38 minutes ago [-]
> finally fixed their VAE
Can you share the sources?
trentor 15 minutes ago [-]
It's right there in the hugging face link?
latents go from 16ch @ 8x compression to 64ch @ 16x, so roughly the same total latent budget but much more channel heavy. It’s also deeper/wider, and the old 2x2 transformer patching is gone.
On some images it still produces artifacts but can't say if it's the transformer or the VAE yet.
Hard_Space 2 hours ago [-]
Interesting in the example of assembling the Cheers team how the otherwise great result genericizes Shelley Long.
hughc 1 hours ago [-]
The result seems a pretty good representation given the source image wasn't that great. I think that Woody Harrelson comes across much worse.
TomGarden 2 hours ago [-]
Very impressive, and kind of worrying a 7B model can have such capabilities. The implications are huge. And Qwen does no watermarking (yet) yeah?
trentor 1 hours ago [-]
They always had a fourier space mark in their models even without the VAEs are usually pretty easy to detect.
TomGarden 1 hours ago [-]
Ah I wasn't aware, thank you
spottedmarley 2 hours ago [-]
Boy do I love waking up to find a new awesome toy from the Qwen team waiting for me to play with! Pulling it now
1 hours ago [-]
bknight1983 55 minutes ago [-]
While I'm impressed with the Bluey example, the lack of Muffin disappoints me.
weee322 2 hours ago [-]
[flagged]
BlackGlory 1 hours ago [-]
[flagged]
reedf1 1 hours ago [-]
Context?
JimDabell 1 hours ago [-]
People say ChatGPT generates images with a yellow tint. The person you are replying to is suggesting that these images have a yellow tint and therefore this model is distilled from ChatGPT.
hn45e7pbij 2 hours ago [-]
Image gen you eyeball one frame and stop, code needs hundreds of tokens all correct in sequence, one bad line and the whole thing fails.
https://en.wikipedia.org/wiki/Qwen#List_of_models
Unfortunately, it looks like this model is using a much more restrictive license:
https://github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE
If AI labs get to ignore licenses, so do we.
If you were to generate outputs for commercial use, I think it would still violate this research license, but it's not like they are going to know, are they?
That said, I am disappointed that the model is not actually open-weights as I expected based on the headline.
People can continue to use closed SOTA models to generate outputs for commercial or malicious purposes.
What this research license achieves is that we cannot use this model in applications we publish.
How difficult would it be to use this model to create a second model without licensing issues?
And Krea 2 has a community license [2] that is fairly permissive - I think commercial usage is allowed under $1 million.
Boogu-Image scored 6/15 and Krea 2 scored 7/15 on my GenAI Showdown benchmark [3] - only Ideogram4 eclipses them in terms of local models, but its got a far more restrictive license and the JSON structured inputs can be a pain to work with.
[1] - https://github.com/Boogu-Project/Boogu-Image
[2] - https://www.krea.ai/krea-2-licensing
[3] - https://genai-showdown.specr.net/?models=fd,hd,kd,qi,f2d,zt,...
Recently, I've been helping a friend's wife with some basic vector images for her sewing hobby (she has what is essentially a CNC sewing machine) and have been super-impressed with FLUX.1-Kontext, which I've been running on my Macbook Pro with mflux. Its ability to (for example) take a photo of a human or an animal and return a line drawing which is recognisably them (rather than just a generic similarish image as I've experienced with other models) is excellent.
It's an older model now, but (AIUI) has the text-handling features baked in, and in my various testing is very reliable at giving me the outputs that I want, without the randomness I've experienced previously. It's big and relatively slow (~3 mins per 512x512 image edit on my M1 Max Mac) but excellent to work with. It's also very straightforward to set up, without the harness complexity of e.g. comfyui.
Prompt-adherence is really hit-or-miss—especially if one lacks the visual vocabulary. Likewise with coding, I find junior devs don't think to prompt re: respecting this-or-that interface, or refactoring to point-free style, etc.
So, as others have said, the artist knows better.
(I mean: outside direct or substantial use of Python, and running the Neural Network in the most efficient way.)
I think that's all Python (not a direct executable).
You could just do (see the "Quick Start") four `pip install` and have a dozen lines script to generate the image. But `llama.cpp` and similar do not require e.g. installing Torch (or PyTorch) - you can use `llama.cpp` on a non-specialized machine.
> Currently, we support image, audio and video input.
Image outputs are supported, videos I'm not sure but I don't think that's an output, just a preview of the equirectangular example, so, same question here, what does this model outputs that isn't supported?
I understand that llama.cpp could only output text, last time I checked (I do not know how to find a good source for that though).
See https://github.com/ggml-org/llama.cpp/blob/master/src/llama-... , the
...Back in Apr, ChatGPT Images 2.0 has some broken Chinese texts in its featured examples, and they later removed that from blog post. Is 2.5 better now?
Excited to see what the future holds for them!
> You shall not use the Materials for any commercial purpose without obtaining a separate commercial license from us.
It probably will be much cheaper to use than other image models, but it seems that will be up to the whims of Qwen/Alibaba rather than just being the cost of putting it in a cloud provider.
https://github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE
Can you share the sources?
latents go from 16ch @ 8x compression to 64ch @ 16x, so roughly the same total latent budget but much more channel heavy. It’s also deeper/wider, and the old 2x2 transformer patching is gone.
On some images it still produces artifacts but can't say if it's the transformer or the VAE yet.