Description
Scribd, Inc. has partnered with Databricks to harness the power of generative AI, significantly enhancing content discovery and operational efficiency. With a library exceeding 250 million documents, Scribd faced challenges in managing fragmented data workflows and limited content insights. By integrating Databricks, Scribd unified its data and AI operations, enabling rapid innovation and smarter search capabilities.
The collaboration has led to a 7% increase in user sign-ups and a 90% reduction in generative AI costs. Databricks' platform supports Scribd's entire data and AI lifecycle, from data ingestion to real-time inference, allowing for seamless experimentation and deployment. This integration has streamlined operations, reduced context switching, and facilitated collaboration across teams.
Scribd utilizes Databricks for batch ETL, real-time data processing, and model development, which accelerates iteration cycles. The platform's flexibility allows Scribd to choose the appropriate models for various use cases, such as auto-generating metadata and powering semantic search.
The partnership extends beyond technology, with Databricks providing support during critical moments, such as tuning model performance and managing GPU resources. This support has been crucial in accelerating delivery and maintaining focus on building AI-driven solutions.
Looking ahead, Scribd plans to develop more AI-native features, such as intelligent topic extraction and slide-level search, leveraging the scalable infrastructure provided by Databricks. This collaboration represents a significant step towards embedding AI into the core of Scribd's product experience, enhancing user engagement and operational efficiency.
Key Takeaways
Unified data and AI platform
Batch ETL and real-time data processing
Model development and deployment
Generative AI cost reduction
Semantic search capabilities
Auto-generation of content metadata
Collaboration across teams
Support for real-time and batch inference
What This Case Study Demonstrates
- Content Discovery
- Metadata Generation
- Semantic Search
- AI Model Deployment
- Operational Efficiency







