Description
Parti, the Pathways Autoregressive Text-to-Image model, is a novel approach to generating photorealistic images from textual descriptions. It frames text-to-image generation as a sequence-to-sequence modeling problem, drawing parallels to machine translation and benefiting from advancements in large language models, particularly those unlocked by scaling data and model sizes. Parti utilizes the ViT-VQGAN image tokenizer to convert images into sequences of discrete tokens, enabling it to reconstruct these sequences into high-quality, visually diverse images.
This model demonstrates consistent quality improvements as its encoder-decoder scales up to 20 billion parameters. Parti has achieved state-of-the-art zero-shot FID scores of 7.23 and finetuned FID scores of 3.22 on the MS-COCO benchmark. Its effectiveness spans a wide variety of categories and challenges, as analyzed on the newly released PartiPrompts benchmark, a holistic collection of over 1600 English prompts designed to measure model capabilities across diverse aspects.
Parti excels at handling long, complex prompts that require accurate reflection of world knowledge, composition of numerous participants and objects with fine-grained details and interactions, and adherence to specific image formats and styles. The model's capabilities are evident in its ability to generate images based on abstract concepts, incorporate world knowledge, render specific perspectives, and even produce text and symbols within images. The research also highlights limitations and areas for future improvement, such as the handling of negation and indications of absence.
Developed using Lingvo and scaled with GSPMD on TPU v4 hardware, Parti's architecture allows for significant scaling. Comparisons between models ranging from 350 million to 20 billion parameters show substantial improvements in model capabilities and output image quality. Human evaluators consistently preferred larger models, particularly for image realism, quality, and text-image alignment. The 20 billion parameter model shows particular strength in prompts requiring abstract reasoning, world knowledge, specific viewpoints, and the rendering of text and symbols.
While Parti offers exciting possibilities for creativity and visual communication, the research acknowledges significant risks, including the potential for encoding harmful stereotypes and representations due to biases in training data. Similar to other text-to-image models, Parti may reflect Western biases and underrepresent certain backgrounds. The potential for creating deepfakes and propagating misinformation is also a concern. Due to these risks, Google Research has decided not to release the Parti models, code, or data publicly without further safeguards. Instead, they provide a Parti watermark on released images and plan to focus on bias measurement, mitigation strategies, and coordinating with artists to explore responsible applications.
Parti: Pathways Autoregressive Text-to-Image Model Highlights
Autoregressive text-to-image generation
High-fidelity photorealistic image generation
Supports content-rich synthesis with complex compositions
Leverages large language model advancements
Sequence-to-sequence modeling approach
Utilizes ViT-VQGAN for image tokenization
Scalable architecture up to 20 billion parameters
Achieves state-of-the-art FID scores on MS-COCO
Effective across diverse prompt categories and difficulty aspects
Handles abstract prompts and world knowledge
Renders specific perspectives and text/symbols
Implemented in Lingvo and scaled with GSPMD on TPU v4
Getting Started with Parti: Pathways Autoregressive Text-to-Image Model
Access model: Understand the research paper and its findings on Parti.
Explore capabilities: Review example prompts and generated images to gauge model performance.
Analyze benchmark results: Examine FID scores and performance on the PartiPrompts benchmark.
Understand limitations: Familiarize yourself with identified failure modes and areas for improvement.
Consider responsible use: Be aware of the ethical considerations and risks associated with text-to-image models.
Review data card: Access information on data collection and model development practices.
Parti: Pathways Autoregressive Text-to-Image Model's Use Cases
- Photorealistic Image Generation
- Complex Scene Synthesis
- World Knowledge Integration
- Abstract Concept Visualization
- Artistic Style Emulation
- Text and Symbol Rendering
- Creative Content Creation






