Welcome to Kerala Project Center.

UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation

UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation

UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation project

Text-to-image generation has exploded in popularity, but masked generative transformers (MGTs) — the architecture behind models like Muse — struggle with compositional accuracy: they frequently mis-assign attributes (color, count, position) to the wrong objects in complex prompts. This project, UNCAGE, implements a breakthrough training-free technique that fixes compositional failures without any model retraining or fine-tuning. The method works by analyzing the transformer's internal cross-attention maps during the unmasking process, applying contrastive attention guidance to sharpen the distinction between tokens for different objects and their corresponding image regions. By optimizing which tokens get unmasked first and enforcing accurate text-image token binding, UNCAGE dramatically improves attribute binding, object placement, and multi-object scene coherence. Because it is training-free, the technique plugs directly into pre-trained MGT models with zero additional compute cost for training — making it immediately practical. Evaluation on standard compositional benchmarks (T2I-CompBench) shows measurable gains in attribute binding and text-image alignment. Implemented with Python and PyTorch using pre-trained open MGT checkpoints, this is a frontier generative AI project for researchers in text-to-image synthesis and attention mechanism research.

Components



Python 3.8+
PyTorch
Pre-trained masked generative transformer checkpoints
Cross-attention analysis modules
T2I-CompBench evaluation suite

Key Features


  • Training-free — plugs directly into pre-trained MGT models
  • Contrastive cross-attention guidance during the unmasking process
  • Accurate attribute binding between text tokens and image regions
  • Improved multi-object scene coherence and object placement
  • Measured gains on the T2I-CompBench compositional benchmark

Applications


Generative AI researchers, text-to-image developers and advanced deep learning students — a frontier research topic.


UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation – hexcodeplus ads

Hours

Monday - Saturday: 9:00 AM - 5:00 PM
Sunday: Not Working

Location

2nd Floor, Comptron Arcade, Kallattumukku,
Thiruvananthapuram, Kerala 695012

Book Now

+91 9633118080