Granica Revolutionizes AI Data Management with Cloud-Native Platform
Granica Company has developed a revolutionary AI data management platform that processes data directly in cloud data lakehouses while maintaining full security and control within customer environments. The company's products, Crunch, Screen, and Signal, enable safe and efficient AI development at scale through advanced compression techniques, sophisticated privacy detection algorithms, and model-aware data selection. Granica's architecture achieves higher than 99.99% availability across all instances, including spot instances, through a sophisticated design that combines multiple availability zones with Kubernetes and VPC peering. The platform processes data directly in the cloud data lakehouses and their underlying object stores without requiring any application integration, achieving 10:1 reduction in S3 API costs while maintaining performance.
Granica Company is an information and computer science firm founded in 2019 with a mission to solve complex data challenges through advanced AI technologies. The company's products include Crunch for cost optimization, Screen for data privacy, and Signal for data selection, designed to enable safe and efficient AI development at scale.
At the core of Granica's technology is their AI data platform, which processes data in cloud data lakehouses while maintaining full security and control within customer environments. The platform achieves higher than 99.99% availability across all instances, including spot instances, through a sophisticated architecture that combines multiple availability zones with Kubernetes and VPC peering.
The company's culture emphasizes joy, asynchronous communication, outcome-driven focus, kindness, and high velocity. This unique approach has helped the company maintain a passionate and productive workforce as they tackle some of the most challenging problems in data management today. Their success has attracted significant investment, including $45 million in funding from leading AI, data, and cloud investors.
The platform's architecture runs entirely within the customer's cloud environment, meaning that the company's data never leaves the customer's secure environment. Granica's technology processes data directly in the cloud data lakehouses and their underlying object stores, making the data AI-ready without requiring any application integration. This allows for continuous background processing of data, reading from and writing to cloud storage seamlessly.
A key component of Granica's architecture is its ability to maintain data security through built-in encryption for both data at rest and in transit. The company uses native cloud provider services for encryption, ensuring that all data remains within the Virtual Private Cloud (VPC) and complies with customer security policies. Granica's approach to security also includes fine-grained access controls that respect all security policies, running as a single tenant in a dedicated account/project and VPC.
The platform's availability stands at more than 99.99% across all instances, including spot instances. This reliability is achieved through a sophisticated architecture that combines multiple availability zones with Kubernetes and VPC peering. Granica's architecture uses Kubernetes pods running across a cluster of compute instances, with a minimum configuration of a 2-node cluster consisting of 24/7 on-demand instances. To handle increased load, the platform elastically spins up additional spot instances and service pods, managed by a Broker pod that distributes requests to distributed service pods across the cluster instances.
Data processing occurs in multiple availability zones, automatically leveraging spot instances where possible to maintain cost-effectiveness. The platform uses Amazon Web Services (AWS) EKS Availability Groups to build its infrastructure primitives, ensuring high availability for their products. When managing spot instances, Granica employs a rolling upgrade and rollback approach across all service pods and Kubernetes cluster infrastructure, maintaining full performance and availability even during upgrades. This rolling deployment strategy also includes non-disruptive upgrade capabilities, allowing automatic version updates while providing manual upgrade options via the granica update command.
Granica's three core products—Crunch, Screen, and Signal—represent significant advancements in AI data management across cost optimization, privacy, and selection. These tools enable AI teams to work effectively with large, complex datasets while maintaining the highest standards of security and efficiency.
Crunch represents the industry's first comprehensive solution for optimizing large-scale tabular and columnar datasets, particularly optimized for Apache Parquet files. By introducing revolutionary compression techniques, Crunch can reduce the physical size of Parquet files by up to 60%, significantly lowering storage and transfer costs while boosting query performance by up to 56%. This transformative approach maintains full functionality without requiring any changes to existing applications.
Screen represents the most advanced data privacy platform available today, featuring unparalleled accuracy in detecting sensitive information and harmful content across multiple terabytes of data weekly. Utilizing sophisticated machine learning algorithms, Screen achieves state-of-the-art detection accuracy while processing data 5-10 times faster than traditional methods. The platform supports over 100 languages and 50 named entities, including advanced classifiers for phone numbers, SSNs, VINs, and custom data types.
Signal introduces the first model-aware data selection service, directly addressing the challenge of "signal in the noise" in large training datasets. By analyzing the entire dataset and existing models, Signal helps teams focus on the most impactful samples, delivering up to 30% improvements in model performance while reducing training cycles by 20-30%. The platform continuously updates training data by automatically adding relevant new samples and removing outdated information, ensuring models remain effective over time.
Granica's architecture runs entirely within customer environments, ensuring that all data remains secure and compliant with existing policies. The platform processes data in cloud data lakehouses and their underlying object stores through background operations, requiring no changes to existing applications. For real-time data safety, Screen leverages APIs and SDKs to integrate seamlessly with LLM-enabled applications.
The architecture maintains >99.99% availability across all instances, including spot instances, through a sophisticated multi-AZ design backed by Kubernetes and VPC peering. This robust foundation allows the platform to scale seamlessly while maintaining optimal performance. Granica's cloud-prem approach combines traditional VPC security with SaaS benefits, offering the best of both worlds while achieving cost savings of 10:1 on S3 API usage without compromising performance.
The platform utilizes a sophisticated cloud-prem architecture that self-deploys as a single tenant within the customer's dedicated account/project and VPC. Its design enables multiple availability zones (AZs) while minimizing cross-AZ charges through VPC peering connections between API-based services and applications.
The data processing flow begins with background operations reading from and writing to cloud storage without requiring application integration. For real-time data safety, Screen leverages APIs and SDKs to integrate seamlessly with LLM-enabled applications. Granica employs a multi-level validation system including pre-Crunch file validation and post-Crunch object integrity checks using checksum comparisons to maintain data consistency with native formats.
The company's architecture demonstrates significant cost efficiency, achieving 10:1 reduction in S3 API costs while maintaining performance. This optimization enables the platform to process 150MBps per node with fully elastic scaling capabilities that dynamically adjust from zero to multiple nodes as needed. Spot instances play a crucial role in maintaining performance while minimizing costs, allowing the platform to achieve >99.99% availability through AWS EKS Availability Groups.
In terms of security, Granica implements encryption for both data at rest and in transit using native cloud provider services. The platform's internal security mechanisms include Open Policy Agent and Gatekeeper services that enforce global policies while allowing application-defined access controls. This approach ensures bucket security for objects with public access more effectively than AWS's traditional ACL method.
The rolling upgrade mechanism maintains high availability by automatically managing service pod and Kubernetes infrastructure updates. This capability ensures consistent performance with no impact on availability, as upgrades can be configured to require manual initiation through the granica update command. The platform's shared infrastructure provides multiple layers of validation and security, stopping processing and alerting teams in case of integrity failures.
Data safety remains a cornerstone of Granica's platform, with the company's AI technology adept at identifying Personally Identifiable Information (PII), biased content, and toxic elements within data lakes and LLM prompts. The platform's accuracy stands at state-of-the-art levels across multiple terabytes of data weekly, while processing information 5-10 times faster than traditional methods thanks to its sophisticated machine learning algorithms. The system supports over 100 languages and 50 named entities, including advanced classifiers for phone numbers, Social Security Numbers (SSN), Vehicle Identification Numbers (VIN), and custom data types.
Storage efficiency represents another significant breakthrough, with Granica's lakehouse-native compression technology reducing the physical size of Parquet files by up to 60%. This compression achieves this substantial reduction while simultaneously speeding up query access and data loading times by 56%, based on TPC-DS benchmarks. The platform's cost optimization extends beyond storage, as it enables customers to achieve a remarkable 10:1 reduction in S3 API costs without any performance impact. These savings are realized through a combination of improved data management practices and the platform's efficient processing capabilities.
Model performance improvements demonstrate the platform's value for AI teams working with large datasets. Granica Signal introduces groundbreaking capabilities in model-aware data refinement, helping teams focus on the most impactful samples while automatically managing the selection process. This focused approach delivers up to 30% improvements in model performance while reducing training cycles by 20-30%. The platform's internal data management processes further enhance these benefits through continuous analysis of training datasets and existing models, automatically prioritizing the most representative samples for optimal training outcomes.