ytsaurus/ytsaurus
YTsaurus is a scalable and fault-tolerant open-source big data platform.
About ytsaurus/ytsaurus
ytsaurus/ytsaurus is an open-source project on GitHub, mainly written in C++. YTsaurus is a scalable and fault-tolerant open-source big data platform. It currently holds 2,283 stars and 233 forks with 511 open issues, and was last pushed on 2026-10-08 (repository created 2022-12-05).
Project Overview
Git Homed tracks it on the Today's Trending board, currently at rank #88 with 31 new stars today.
GitHub Repository Details
README

YTsaurus
Website | Documentation | YouTube
YTsaurus is a distributed storage and processing platform for big data with support for MapReduce model, a distributed file system and a NoSQL key-value database.
You can read post about YTsaurus or check video:
Advantages of the platform
Multitenant ecosystem
- A set of interrelated subsystems: MapReduce, an SQL query engine, a job schedule, and a key-value store for OLTP workloads.
- Support for large numbers of users that eliminates multiple installations and streamlines hardware usage
Reliability and stability
- No single point of failure
- Automated replication between servers
- Updates with no loss of computing progress
Scalability
- Up to 1 million CPU cores and thousands of GPUs
- Exabytes of data on different media: HDD, SSD, NVME, RAM
- Tens of thousands of nodes
- Automated server up and down-scaling
Rich functionality
- Expansive MapReduce module
- Distributed ACID transactions
- A variety of SDKs and APIs
- Secure isolation for compute resources and storage
- User-friendly and easy-to-use UI
YTsaurus Flow
- Native framework for real-time data stream processing
- Efficient stateful computations for recommendation and content systems
- Complex pipeline topologies with automatic partitioning and balancing
- Exactly-once processing by default
CHYT powered by ClickHouse®
- A well-known SQL dialect and familiar functionality
- Fast analytic queries
- Integration with popular BI solutions via JDBC and ODBC
SPYT powered by Apache Spark
- A set of popular tools for writing ETL processes
- Launch and support for multiple mini SPYT clusters
- Easy migration for ready-made solutions
Getting Started
Try YTsaurus cluster using Kubernetes or try our online demo.
How to Build from Source Code
- Build from source code.
How to Contribute
We are glad to welcome new contributors!
Please read the contributor's guide and the styleguide for more details.
