Figure from the author-uploaded arXiv version.Text-to-image systems must be judged on more than photorealism: an image can look plausible while failing to depict the requested objects, relationships, or style. Classical image-quality measures alone do not capture this joint visual-and-linguistic requirement.
This survey organizes text-to-image evaluation metrics around two main criteria: compositional image quality and semantic consistency with the prompt. It reviews metric families, the datasets used to validate them, and how evaluation choices interact with the intended generation task.
Beyond cataloguing metrics, the paper discusses the evidence behind them, surveys available evaluation datasets, and examines a selection of human-preference metrics in supplementary experiments. The result is a practical map of what a reported score can—and cannot—support.
No single automatic metric captures all of human judgment. Metric selection remains task-dependent, and the survey emphasizes the need to combine appropriate automatic measures with carefully designed human evaluation when stakes are high.