Description
Welcome to the official documentation for Intel® Gaudi® AI accelerators, covering versions up to v1.24. This resource is designed to empower developers and researchers in leveraging the full potential of Intel® Gaudi® 2 and Intel® Gaudi® 3 AI accelerators.
The documentation provides in-depth information to facilitate your journey with Gaudi hardware. You will find detailed guides on migrating your existing AI models to the Gaudi platform, ensuring a smooth transition. Practical code samples are included to illustrate various functionalities and accelerate your development process.
Furthermore, the documentation delves into best practices for debugging and optimizing your AI workloads on Gaudi. This includes performance tuning tips, architectural insights, and strategies for achieving maximum efficiency. Comprehensive API references are available to help you understand and utilize the full range of Gaudi's capabilities.
Key sections cover software installation and platform access, enabling you to set up your environment correctly. Training and scaling functionalities are explained, with specific guidance on training PyTorch models, running large models using DeepSpeed, and scaling multi-node training with distributed training techniques. For inference, the documentation details how to run PyTorch inference, utilize vLLM and SGLang for efficient model serving, and work with quantized models using FP8 or INT4 datatypes.
This documentation serves as a central hub for all information related to Intel Gaudi, ensuring users have the necessary tools and knowledge to build and deploy cutting-edge AI applications. Access to older documentation versions is available via the 'Versions' button. Feedback on your experience is encouraged to help improve future iterations.
Intel Gaudi Documentation's Core Features
Documentation for Intel Gaudi 2 and Gaudi 3 AI accelerators
Guidance on migrating AI models to Gaudi
Code samples for various functionalities
Best practices for debugging and optimization
API references for Gaudi hardware and software
Software installation and platform access guides
Instructions for training PyTorch models
Support for large model training with DeepSpeed
Multi-node distributed training capabilities
Guidance on PyTorch inference
Inference using vLLM and SGLang
Support for FP8 and INT4 quantized models
Model serving capabilities
Getting Started with Intel Gaudi Documentation
Install Software: Follow the Installation Guide to set up the necessary software and system configurations.
Access Platform: Learn how to run the Intel Gaudi Docker Image on your chosen platform.
Start Training: Begin training PyTorch models and explore options for running large models with DeepSpeed.
Scale Training: Implement multi-node training using distributed training techniques.
Run Inference: Get started with running PyTorch inference and explore advanced inference engines like vLLM and SGLang.
Optimize Models: Utilize FP8 or INT4 datatypes for running quantized models.
Deploy Models: Explore options for model serving.
Intel Gaudi Documentation's Use Cases
- Model Migration
- AI Model Training
- Large Model Deployment
- Optimized Inference
- Performance Tuning
- Platform Setup




