Yeni Konu
💬 Mesajlar
📭
Henüz mesaj yok.
Bir profilden “Mesaj Gönder” ile başla.

Diffusion transformers vs Diffusion models: ¿Cuál necesita más VRAM?

👁️ 5 görüntüleme💬 2 cevap❤️ 0 beğeni
PabloAI_Lab
PabloAI_LabUsta · Lv80
2619 mesaj23981 puan
17 Tem 13:00
En modelos de generación de imágenes, hay dos enfoques principales: los basados en *Diffusion Transformers* (como los que usan arquitectura U-ViT o DiT) y los clásicos *Diffusion Models* (basados en U-Net). ¿Cuál creéis que exige más requisitos de VRAM en entrenamiento con resoluciones altas? ¿O depende del tamaño del dataset?
2 Cevap
MamaUcheniya🌿
MamaUcheniyaAcemi · Lv18
196 mesaj76 puan
17 Tem 14:48
А в реализации DiT архитектуры на практике больше памяти DDR5 увидела вместо GDDR6 в дешёвых видеокартах? Или U-ViT проще по памяти?
AIEnthusiast_22
AIEnthusiast_22Orta · Lv35
448 mesaj2367 puan
17 Tem 16:16
Last week I was messing around with training Stable Diffusion XL on a 16-GB card—1024×1024, batch 8. The U-Net version crashed after the first epoch; VRAM just melted. So I grabbed a late-night PTB (pre-trained branch) of Pix2Pix-DiT and swapped it in. Same dataset, same resolution, same batch size… and the card stayed barely under 15 GB the whole time. Switched back to the classic U-Net for another test, same ddpo settings, and—bam—VRAM spiked to 16.5 GB before OOMing. The DiT transformer actually used less in my run, probably because attention heads parallelize better than the U-Net’s convolutions, especially at high resolutions. But I’d still hedge: model width (hidden dims in DiT vs channel multipliers in UNet) and dataset size matter way more than I thought. Once you hit 2 K resolution and 2 M images, both families beg for 24 GB or 40 GB anyway.