A Survey on Quality Metrics for Text-to-Image Generation

Taxonomy of quality criteria and metrics for text-to-image generation Figure from the author-uploaded arXiv version.

At a glance

  • Surveys metrics for evaluating both visual quality and text–image alignment.
  • Proposes a taxonomy for organizing metric families, datasets, and their intended uses.
  • Identifies open challenges and offers guidance for evaluating text-to-image systems.
Publication
IEEE Transactions on Visualization and Computer Graphics (TVCG)

Motivation

Text-to-image systems must be judged on more than photorealism: an image can look plausible while failing to depict the requested objects, relationships, or style. Classical image-quality measures alone do not capture this joint visual-and-linguistic requirement.

Method

This survey organizes text-to-image evaluation metrics around two main criteria: compositional image quality and semantic consistency with the prompt. It reviews metric families, the datasets used to validate them, and how evaluation choices interact with the intended generation task.

Evaluation

Beyond cataloguing metrics, the paper discusses the evidence behind them, surveys available evaluation datasets, and examines a selection of human-preference metrics in supplementary experiments. The result is a practical map of what a reported score can—and cannot—support.

Limitations

No single automatic metric captures all of human judgment. Metric selection remains task-dependent, and the survey emphasizes the need to combine appropriate automatic measures with carefully designed human evaluation when stakes are high.

Dominik Engel
Dominik Engel
Deep Learning Researcher